Skip to content

perf: accelerate local validation with bounded concurrency - #3644

Merged
kunchenguid merged 4 commits into
mainfrom
fm/fm-validation-suite-speed-r1
Sep 3, 2026
Merged

kunchenguid merged 4 commits into
mainfrom
fm/fm-validation-suite-speed-r1

Conversation

@kunchenguid

@kunchenguid kunchenguid commented Sep 3, 2026 •

Copy link
Copy Markdown
Owner

Intent

Cut Firstmate LOCAL validation wall-clock by >=50% via CONCURRENCY plus POLL-SLEEP REDUCTION. Captain-authorized 2026-09-03 after the timeout scout. SCOPED to these levers only - do NOT broaden.

What Changed

  • Route the no-mistakes changed-file test baseline through fm-test-run.sh, excluding live Herdr tests, and enable bounded automatic concurrency for explicit script lists.
  • Admit the proven pr-forge family at up to four workers while separating family-scoped concurrency into isolated phases and keeping unproven tests serial.
  • Add scheduler regression coverage and document concurrency evidence, refused family proofs, and why coarser polling intervals were not adopted.

Risk Assessment

✅ Low: The concurrency changes are bounded by recorded isolation families, cross-family work is separated into phases, and the timeout fix restores unbounded plain script-list behavior as requested.

Testing

Inspected the base-to-target change, ran the focused runner suite, and exercised the end-user named-script CLI flow; two one-second scripts began concurrently, completed in about 1.3 seconds with four resolved workers, and required no timeout helper, while focused coverage also confirmed serial override, mixed serial tails, explicit timeout behavior, and isolation-family phase separation. The worktree was clean after removing the temporary fixture.

Evidence: Named-script automatic concurrency CLI transcript

Source: Named-script automatic concurrency CLI transcript

END-USER COMMAND: bin/fm-test-run.sh tests/fm-cd-pretool-check.test.sh tests/fm-pr-merge.test.sh
FIXTURE CONTRACT: no bin/fm-timeout-lib.sh is present; both named scripts sleep for one second.
FM_TEST_BEGIN 2026-09-03T16:35:57Z tests/fm-cd-pretool-check.test.sh family=pure-contract-unit expected_gate_skip=none
FM_TEST_BEGIN 2026-09-03T16:35:57Z tests/fm-pr-merge.test.sh family=pr-forge expected_gate_skip=none
ok - named-script fixture completed
FM_TEST_END 2026-09-03T16:35:58Z tests/fm-cd-pretool-check.test.sh exit=0 duration_ms=1170 gate_skip=false
ok - named-script fixture completed
FM_TEST_END 2026-09-03T16:35:58Z tests/fm-pr-merge.test.sh exit=0 duration_ms=1173 gate_skip=false
FM_TEST_SUMMARY total=2 failed=0 skipped_gate=0 duration_ms=1306
FM_TEST_SUMMARY_FAMILY family=pr-forge count=1 duration_ms=1173 failed=0
FM_TEST_SUMMARY_FAMILY family=pure-contract-unit count=1 duration_ms=1170 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-pr-merge.test.sh duration_ms=1173
FM_TEST_SLOWEST rank=2 script=tests/fm-cd-pretool-check.test.sh duration_ms=1170

TIMING ARTIFACT:
{
    "families": [
        {
            "count": 1,
            "duration_ms": 1173,
            "failed": 0,
            "name": "pr-forge"
        },
        {
            "count": 1,
            "duration_ms": 1170,
            "failed": 0,
            "name": "pure-contract-unit"
        }
    ],
    "finished_at": "2026-09-03T16:35:58Z",
    "run_id": "fm-test-run-1788453357237-58892",
    "scripts": [
        {
            "duration_ms": 1170,
            "exit": 0,
            "expected_gate_skip": "none",
            "family": "pure-contract-unit",
            "gate_skip": false,
            "path": "tests/fm-cd-pretool-check.test.sh"
        },
        {
            "duration_ms": 1173,
            "exit": 0,
            "expected_gate_skip": "none",
            "family": "pr-forge",
            "gate_skip": false,
            "path": "tests/fm-pr-merge.test.sh"
        }
    ],
    "selection": "scripts;jobs=4",
    "started_at": "2026-09-03T16:35:57Z",
    "summary": {
        "duration_ms": 1306,
        "failed": 0,
        "skipped_gate": 0,
        "total": 2
    }
}
Evidence: Named-script timing artifact

Source: Named-script timing artifact

{
  "families": [
    {
      "count": 1,
      "duration_ms": 1173,
      "failed": 0,
      "name": "pr-forge"
    },
    {
      "count": 1,
      "duration_ms": 1170,
      "failed": 0,
      "name": "pure-contract-unit"
    }
  ],
  "finished_at": "2026-09-03T16:35:58Z",
  "run_id": "fm-test-run-1788453357237-58892",
  "scripts": [
    {
      "duration_ms": 1170,
      "exit": 0,
      "expected_gate_skip": "none",
      "family": "pure-contract-unit",
      "gate_skip": false,
      "path": "tests/fm-cd-pretool-check.test.sh"
    },
    {
      "duration_ms": 1173,
      "exit": 0,
      "expected_gate_skip": "none",
      "family": "pr-forge",
      "gate_skip": false,
      "path": "tests/fm-pr-merge.test.sh"
    }
  ],
  "selection": "scripts;jobs=4",
  "started_at": "2026-09-03T16:35:57Z",
  "summary": {
    "duration_ms": 1306,
    "failed": 0,
    "skipped_gate": 0,
    "total": 2
  }
}

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 2 issues found → auto-fixed (2) ✅
  • 🚨 CONTRIBUTING.md:112 - The authoritative intent requires the >=50% reduction via “CONCURRENCY plus POLL-SLEEP REDUCTION,” but this change explicitly directs maintainers to leave the 0.1s poll sleeps unchanged. The authorized poll-sleep lever is therefore absent and needs user approval before replacing it with concurrency alone.
  • 🚨 bin/fm-test-run.sh:446 - Adding pr-forge to the family registry makes it globally compatible with every other admitted family, although its proof only covers pr-forge members together. For example, a plain list containing fm-pr-check-security.test.sh and fm-calm-pi-extension.test.sh now runs concurrently despite belonging to separately proven families; the isolation record explicitly excludes the former from the mixed pool. Enforce compatibility at the batch boundary by running each family-proof group in a separate concurrent phase, while only globally proven scripts may mix across families.

🔧 Fix: Separate concurrent runs by isolation proof family
2 errors still open:

  • 🚨 CONTRIBUTING.md:112 - Intent requires “>=50% via CONCURRENCY plus POLL-SLEEP REDUCTION,” but the change explicitly preserves every sleep 0.1; its directly reported verification benchmark is only a 48% reduction. Replacing the required poll-sleep lever with concurrency alone needs user approval.
  • 🚨 bin/fm-test-run.sh:1861 - The intent says “SCOPED to these levers only - do NOT broaden,” but plain script-path runs now also acquire a new automatic 900-second timeout. For example, fm-test-run.sh tests/slow.test.sh changes from unbounded execution to termination with exit 124 despite no timeout being requested. This non-concurrency behavior expansion requires user approval.

🔧 Fix: Limit automatic timeouts to changed-file validation
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • git diff 3d2a08b2097dd24f9ce03fdafe6501e39b79dbd0..58a5db6d9b71a3116dab61537be5dd89bb450880
  • bash tests/fm-test-run.test.sh
  • bin/fm-test-isolation-proof.sh --list
  • bin/fm-test-run.sh --help
  • Ran a fixture through bin/fm-test-run.sh tests/fm-cd-pretool-check.test.sh tests/fm-pr-merge.test.sh --json <evidence-path> with no timeout helper installed
  • Parsed the generated timing JSON and verified selection=scripts;jobs=4, successful exits, and concurrent wall-clock behavior
  • git status --short and transient-fixture cleanup check
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.


Measurement evidence

All figures measured on the captain's Mac, 2026-09-03, 0 failures on both sides of every comparison unless noted. The machine carried unrelated fleet load throughout, so paired comparisons were run back to back in the same window.

The number that changes a verification round

A round of four scripts, run the way an agent runs one today versus named to the runner:

wall failures
chained bash tests/X.test.sh, one at a time 448s 0
the same four named to bin/fm-test-run.sh 231s 0
cut 48.4%

Scripts: fm-captain-hold-lifecycle, fm-pr-check-security, fm-teardown, fm-x-mode.

Per-family, one worker versus four

family j1 j4 speedup cut
watcher-wake-lock (already admitted) 1311.1s 539.3s 2.43x 58.9%
pr-forge (admitted here) 409.2s 237.9s 1.72x 41.9%
pure-contract-unit (already admitted, prior measurement) 510.7s 252.4s 2.02x 50.6%
secondmate (refused) 1432.1s ~478s at proof ~3.0x not admitted

pr-forge cannot beat 2.1x: its longest script alone runs 198.5s of the 409.2s.

The gate's configured baseline, run exactly as configured: bin/fm-test-run.sh --changed --exclude-family real-herdr-gated selected 33 scripts and finished in 199.8s with 0 failures.

Families admitted and refused

pr-forge is admitted on two consecutive clean proofs at four workers: total=6 failed=0 duration_ms=198594 and total=6 failed=0 duration_ms=186796.

Two families were proven and refused, and docs/fm-test-isolation-proof.md records the exact script and reason for each so the refusals are actionable:

  • secondmate — 14.3 minutes of the suite and it scales about 3x, so it is the largest remaining prize. One script blocks it: tests/fm-backlog-handoff.test.sh failed in two runs of three, both times on its crash-recovery case. That case SIGKILLs a real handoff mid-operation; under four workers the kill lands after the move instead of before it, so recovery reports Task "pre-move-crash" not found in this backlog. The other 19 scripts passed in every run, so the blocker is one crash-injection race, not shared secondmate state.
  • session-bootstrap — one failure, tests/fm-session-start.test.sh reporting the digest waited 9s for inactive reconciliation's 8s state read. That assertion measures elapsed time rather than shared state, and the host's five-minute load average was above 12 while the proof ran. Recorded as a refusal because admission requires a passing proof and the harness never retries a failure into green; re-run it on an idle host before deciding.

unclassified (12.0 minutes) was targeted and not attempted: it currently holds tests/fm-backend-herdr-focus-flash-e2e.test.sh, a real-Herdr lab regression whose family should be real-herdr-gated, and an opt-in live-harness script that gate-skips on its first line. Worth noting separately: because that Herdr regression sits in the portable serial lane today, Linux CI gate-skips it and it never actually runs.

Whole-suite result, and why it is short of 50%

Against the scout's 77.4-minute serial baseline:

family serial after basis
watcher-wake-lock 15.1 6.2 measured 2.43x
pure-contract-unit 6.2 3.1 measured 2.02x
pr-forge 5.8 3.4 measured 1.72x
everything else 50.3 50.3 unproven, stays serial
total 77.4 63.0 18.6% cut

With the three refused families admitted the same arithmetic gives about 41 minutes, a 47% cut — which is exactly what the scout's revised section 11 predicted for concurrency alone at the four-worker ceiling. Concurrency was never going to reach 50% on the whole suite by itself.

The second lever was meant to close that gap, and the measurement refused it. Raising the bounded-wait sample interval from 0.1s to 0.5s, charging each sample proportionally so no budget shrank:

fm-watch-triage.test.sh secs
unchanged 390, 393
coarser sample interval 440, 435

Back to back, same window: 11.8% slower. The scout named this exact falsifier — if 394s does not fall to roughly 300s, the per-sleep penalty is not recoverable overhead. Those sleeps are not overhead added to the clock; they are how a test waits for a subject that only moves on fm-watch.sh's own one-second FM_POLL cadence, so sampling less often only delays detection. The same change also broke tests/fm-watcher-lock.test.sh, which catches a transient (a watcher alive and holding the lock it just reclaimed) rather than waiting for a settled condition — it failed in 16s where the unchanged script passed in 116s.

The lever is therefore not in this change, and CONTRIBUTING.md records the result so the experiment is not repeated.

…unner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.
@greptile-apps

greptile-apps Bot commented Sep 3, 2026

Copy link
Copy Markdown

Confidence Score: 5/5

The PR appears safe to merge, with no concrete changed-code defect remaining after review.

The scheduler preserves family proof boundaries through separate phases, leaves unproven scripts serial, and retains dedicated CI coverage for the locally excluded Herdr family.

Reviews (1): Last reviewed commit: "no-mistakes(document): Clarify validatio..." | Re-trigger Greptile

@kunchenguid
kunchenguid merged commit 7dcf072 into main Sep 3, 2026
16 checks passed
@kunchenguid
kunchenguid deleted the fm/fm-validation-suite-speed-r1 branch September 3, 2026 17:02
AgardnerAU added a commit to AgardnerAU/firstmate that referenced this pull request Sep 4, 2026
* refactor(quota): extract mid-task polling and candidate selection into dedicated scripts (#3221)

* Add quota exhaustion detection and safe fallback helpers

- bin/fm-procevent-quota.sh: generic procevent adapter that arms a
  recurring quota-axi --json poll and wakes firstmate when a tracked
  provider's effectivePercentRemaining drops below a threshold or its
  runway.status becomes exhausted_now.
- bin/fm-quota-choose.sh: worker-side helper that picks the first ranked
  harness:model candidate with positive effectivePercentRemaining.
- AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document
  the new helpers and the mid-task quota-exhaustion wake path.
- tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON
  source.

* no-mistakes(review): Fix quota polling and scope bounds

* no-mistakes(review): Enforce safe default quota selection

* no-mistakes(review): Handle decimal quota values safely

* no-mistakes(review): Fail closed on invalid quota inputs

* no-mistakes(review): Reject empty quota candidate segments

* no-mistakes(review): Harden quota parsing and timeout ownership

* no-mistakes(review): Reuse captured quota snapshots consistently

* no-mistakes(review): Match quota using explicit candidate providers

* no-mistakes(review): Centralize fail-closed quota schema validation

* no-mistakes(review): Reject out-of-range quota percentages

* no-mistakes(review): Validate quota runway status enum

* no-mistakes(review): Tighten quota scope and status contracts

* no-mistakes(review): Preserve unknown quota and exact product bounds

* no-mistakes(review): Preserve provider-level unknown quota

* no-mistakes(review): Reuse canonical verified harness validation

* no-mistakes(document): Document mid-task quota handling

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(docs): restore default routing contract, keep quota helper optional

Restore the AGENTS.md section 4 always-loaded routing paragraph the PR
had deleted, so the standing TOON-first intake, spendPriority ranker,
every-candidate accounting, and load-trigger contract stay exactly as
before this PR. The mid-task quota wake is optional and must not alter
default routing.

Restore the quota-array-dispatch skill ownership line to section 4 as
the always-loaded intake boundary owner; keep the worker-side helper
section as an addition only, without rewiring ownership or load
triggers to section 13.

* fix(bin): use harness-keyed quota matching in optional helper

Revert fm-quota-choose.sh from harness:provider:model tuples back to
harness:model candidates with harness-keyed provider matching, per the
resolved ask-user finding. The helper is optional; authoritative
multi-provider routing (provider discovery from the harness catalog and
quota matching by that explicit provider) stays owned by AGENTS.md
section 4 and the quota-array-dispatch skill intake procedure, not the
helper.

Document the multi-provider limitation in the helper header and the
quota-array-dispatch skill: the helper maps each harness to one primary
provider family only, so a candidate whose established provider differs
from that primary family is checked against the wrong quota row. Use it
only when the brief fixed the candidate order and every candidate's
provider is the harness's primary family.

The helper still consumes one already-captured default-TOON or JSON
snapshot via stdin or --snapshot and never calls quota-axi itself, so
it selects from the same quota state as the intake.

* no-mistakes(review): Fix Muse quota mapping and helper contract docs

* no-mistakes(review): Reject known-empty quotas and map quota tests explicitly

* no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse

* no-mistakes(review): Fix quota retirement and dependent regression coverage

* no-mistakes(review): Accept zero-row quota TOON snapshots

* no-mistakes(review): Enforce quota semantics status consistency

* no-mistakes(review): Veto dispatch on any exhausted applicable scope

* no-mistakes(review): Record exhausted quota scope in wake details

* no-mistakes(review): Fix quota help and control dependency coverage

* no-mistakes(review): Decode quoted TOON fields and document quota wakes

* no-mistakes(review): Validate zero-row TOON and map timeout coverage

* no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes

* no-mistakes(review): Validate complete nonzero TOON envelopes

* no-mistakes(review): Accept producer-shaped quota TOON envelopes

* no-mistakes(review): Support empty quota arrays and validate counted rows

* no-mistakes(review): Harden TOON completion, scopes, and quoted fields

* no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields

* no-mistakes(review): Allow unknown headroom under known semantics

* no-mistakes(review): Reject noncanonical quota identities

* no-mistakes(review): Preserve empty quota polling and validate attention identities

* no-mistakes(review): Reject noncanonical provider watches

* no-mistakes(review): Validate all candidates before quota selection

* no-mistakes(document): Correct quota helper safety documentation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix: surface comments on Lavish annotations (#3371)

* fix(bin): keep typed Lavish comments when an element is also annotated

read preferred element text over prompt, so an annotate-and-comment
item dropped the captain's words. Surface prompt as its own field.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Filter non-comment prompts from Lavish reader output

* no-mistakes(document): Clarify Lavish comment presentation contract

* no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure

* no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure

* fix(bin): always emit Lavish comments and use real annotation fixtures

Stop inferring comment provenance from prompt==text. Real pure
annotations have no prompt, so always-emit does not duplicate.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: support first public-followup registration on Bash 3.2 (#3420)

* Fix public-followup register crashing on empty lock arrays under bash 3.2.

bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock.
The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there.

* no-mistakes(document): Document stock Bash registration coverage

* no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5

* no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression

* fix(bin): isolate new Herdr server environments (#2792)

* fix(herdr): isolate server launch environment

* no-mistakes(review): Clear inherited supervision model from Herdr launches

* no-mistakes(document): Document Herdr server launch environment isolation

* fix: surface inbound Relay media to responding agents (#3442)

* fix: surface inbound Relay attachments to the responding agent

A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.

Fix it where the gap is, in prose:

- Read the complete payload object rather than a fixed field list, so
  media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
  mention and on every chain entry, and call out the common shape where
  only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
  (Discord: cdn.discordapp.com, media.discordapp.net,
  images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
  pbs.twimg.com, video.twimg.com), report a blocked host instead of
  working around it, and treat everything fetched as untrusted public
  input on the same terms as the surrounding thread text.

The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.

The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.

* no-mistakes(review): Preserve media authority and enforce poll-only fetching

* no-mistakes(document): Clarify Relay attachment safety prose

* fix(bin): defer inactive reconciliation during startup (#3480)

* Defer inactive startup reconciliation

* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably

* no-mistakes(review): Require worker phases to cover startup requests

* no-mistakes(review): Make diagnostic wakes safely acknowledgeable

* no-mistakes(document): Document deferred startup phase coverage

* fix(bin): bound wake drain presentation lock waits (#3475)

* fix: bound status presentation lock waits

* no-mistakes(review): Distinguish malformed presentation locks from live contention

* no-mistakes(review): Bound no-ack drain queue lock acquisition

* no-mistakes(document): Document bounded presentation-lock drain behavior

* no-mistakes(lint): Annotate bounded lock output global

* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite

* fix(bin): retire public follow-ups in remote homes (#3479)

* fix(relay): close a public loop whose work lives in a remote secondmate home

A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.

The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.

Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.

Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.

* no-mistakes(review): Guard remote link clears by request identity

* no-mistakes(review): Fail guarded clears on unreadable remote state

* no-mistakes(review): Reject guarded clears on non-writable remote state

* no-mistakes(review): Allow no-link retirement in non-writable remote state

* no-mistakes(document): Correct public-followup verification guarantee count

* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh

* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint

* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks

* fix(relay): bound the guarded remote link clear so it refuses instead of hanging

The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.

The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.

The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.

The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.

* no-mistakes(review): Harden lock-timeout regression with independent deadline

* no-mistakes(review): Restore no-op guarded clears on read-only state

* no-mistakes(document): Clarify remote public-followup cleanup contract

* fix(bin): support process events under symlinked homes (#3484)

* fix(bin): resolve process-event state roots before validating them

The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.

Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.

This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.

* fix(bin): pin the external capture staging boundary to its physical path

The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.

The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.

* no-mistakes(review): Propagate canonical process-event state roots

* no-mistakes(review): Propagate canonical state to process-event adapters

* no-mistakes(document): Document physical process-event state roots

* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)

* fix(pi): persist captain outcomes visibly

* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition

* no-mistakes(document): Document cold-start captain-outcome recovery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(pi): process captain outcomes through a sequence-keyed turn

PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.

The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.

Add the processing half on top of the persistence half:

- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
  cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
  only advances through an explicit sequence-bound acknowledgement, never
  past the read cursor and never backwards; an absent marker reads as zero
  and `processed-init` migrates delivered history once so an upgraded home
  is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
  every still-unprocessed captain row to main as one hidden, typed
  `fm-branch-process` request listing each `[seq N] task: summary`, opening
  exactly one main turn. Main closes it only by calling the new
  `fm_branch_processed` tool with the highest sequence listed. An unrelated,
  empty, or paraphrased answer leaves the sequence open, and the same request
  is presented again at the end of the next main run and at session start.
  The first two presentations of a sequence set open a turn of their own;
  after that the request rides the captain's next prompt so an ignored
  request cannot loop, and a session replacement resets that budget.
  Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
  scripts: an empty answer and an unrelated prior answer neither advance the
  marker nor stop re-presentation, the acknowledgement is refused beyond the
  read cursor and outside lock ownership, a partial acknowledgement keeps the
  newer sequence open, and #3312's own assertions now forbid an unkeyed turn
  rather than any turn. The store suite pins the marker's bounds and the
  migration; the real-SDK guard for appendEntry persistence and model
  exclusion is unchanged.

Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.

* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements

* no-mistakes(review): Harden outcome state validation and request pacing

* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores

* no-mistakes(review): Validate canonical mark-read cursor state

* no-mistakes(review): Guard cursor advancement against corrupt processed state

* no-mistakes(review): Bind acknowledgements to active processing requests

* no-mistakes(review): Reset pacing when processing sequence membership changes

* no-mistakes(review): Enforce silent outcome invariants at storage boundary

* no-mistakes(document): Document hardened captain outcome processing contracts

---------

Co-authored-by: kunchenguid <kun@kunchenguid.com>

* feat: add bounded concurrent Bearings ledger collection (#3481)

* feat: bound Bearings remote ledger collection

* no-mistakes(review): Clarify default remote-ledger collection behavior

* no-mistakes(review): Detach reconcile delivery from watcher loop

* no-mistakes(review): Enforce bounded snapshot and request captures

* no-mistakes(review): Bound legacy summary capture before parsing

* no-mistakes(review): Bound primary remote ledger captures

* no-mistakes(document): Correct snapshot and reconcile documentation

* no-mistakes(lint): Fix ShellCheck quoting in bounded collector

* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks

* test: await reconcile request retirement

* no-mistakes(review): Avoid empty reconcile queue process churn

* no-mistakes(review): Read ledger summaries from immutable snapshots

* no-mistakes(review): Reject multi-document home ledger streams

* no-mistakes(review): Coalesce durable reconcile requests per target

* no-mistakes(review): Unify reconcile keys and reject snapshot streams

* no-mistakes(review): Key reconcile requests by stable target ID

* no-mistakes(document): Document per-target reconcile request coalescing

* no-mistakes(lint): Remove unused snapshot summary file variable

* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass

* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks

* ci: rebalance portable serial test shards (#3489)

* fix(ci): rebalance the portable serial shards on measured durations

The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.

Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.

Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.

Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.

No test changes what it asserts and no test stops running; only the
partition across shards changes.

* no-mistakes(document): Clarify conservative shard timing aggregate

* fix(pi): fall back on incomplete supervision branch prompts (#3491)

* fix(pi): fall back after settled branch errors

* no-mistakes(review): Detect provider errors across prompt compaction

* no-mistakes(review): Preserve in-flight branch state across selection changes

* fix(pi): re-probe supervision branch after cooldown (#3497)

* fix(pi): recover supervision branch after cooldown

* no-mistakes(review): Defer branch recovery until prompt settlement

* no-mistakes(document): Clarify supervision cooldown recovery contract

* fix(bin): remove legacy remote snapshot reads (#3501)

* refactor: remove legacy remote summary reads

* no-mistakes(document): Document ledger-only snapshot reads

* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass

* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean

* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux

* fix(pi): preserve watcher continuity across session replacement (#3498)

* fix(pi): rearm watcher after session replacement

* no-mistakes(review): Queue actionable closes across Pi session replacement

* no-mistakes(review): Stop replacement arm when handoff persistence fails

* no-mistakes(review): Preserve actionable wakes through branch and late child races

* no-mistakes(review): Surface late handoff failures without crashing Pi

* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens

* no-mistakes(review): Retry stale deliveries and release settled claims

* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup

* no-mistakes(review): Deduplicate persistent handoff cleanup alerts

* no-mistakes(review): Acknowledge watcher follow-ups only when consumed

* no-mistakes(review): Persist idle follow-ups until agent consumption

* no-mistakes(review): Preserve pending outcomes when handoff persistence fails

* no-mistakes(review): Arm replacement before awaiting prior delivery settlement

* no-mistakes(review): Adopt pending handoffs after lock reclamation

* no-mistakes(review): Prevent stale generations from adopting replacement handoffs

* no-mistakes(review): Scope replacement handoffs by watcher state

* no-mistakes(document): Clarify replacement handoff documentation

* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks

* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes

* no-mistakes(document): Document watcher-owned replacement handoffs

* no-mistakes(document): Verify replacement handoff documentation

* test(pi): cover watcher-owned branch fallback

* no-mistakes(document): Refresh watcher-owned fallback documentation

* fix(bin): resurface task statuses missed by wake handling (#3495)

* fix(bin): resurface terminal statuses lost after branch handling

* test(watch): canonicalize process-event fixture homes

* no-mistakes(review): Index branch outcomes by causal status position

* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses

* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics

* no-mistakes(review): Keep unclassifiable oversized statuses silent

* no-mistakes(document): Document lost-wake outcome backstop

* no-mistakes(document): Update outcome backstop documentation

* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally

* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes

* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift

* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state

* fix(bin): collect follow-up results from remote work homes (#3503)

* fix(bin): deliver typed terminal results from remote work homes

A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.

The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.

A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.

This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.

* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes

* no-mistakes(review): Fail collection when remote outbox is unreadable

* no-mistakes(review): Surface reassigned remote routes during empty collection

* no-mistakes(review): Fail remote collection on invalid registrations

* no-mistakes(review): Reject unsafe registration entries during remote collection

* no-mistakes(review): Restore healthy empty remote collection behavior

* no-mistakes(review): Skip remote collection for delivered registrations

* no-mistakes(review): Skip delivered registrations before route validation

* no-mistakes(document): Document remote follow-up collection semantics

* fix(bin): exclude secondmates from home-summary validity (#3504)

* fix(bin): exclude secondmates from home-summary child inventory

kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.

* no-mistakes(review): Cover terminal secondmate in-flight exclusion

* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds

* fix(bin): self-heal outcome indexes on first drain (#3509)

* fix(bin): self-heal status-outcome indexes on every drain

Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.

* no-mistakes(review): Guard held-lock initialization and fail marker writes

* no-mistakes(document): Document cross-harness outcome-index self-healing

* fix(bearings): keep active children underway during captain holds (#3505)

* fix(bearings): keep active children underway beside a captain hold

Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.

* no-mistakes(review): Preserve Underway repos and disclose child truncation

* no-mistakes(review): Fall back to task project for Underway repos

* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean

* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)

* fix(pi): settle watcher delivery on Pi accepting the follow-up

A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.

The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.

Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(pi): retry a verified successor that fails during wake delivery

A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.

The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.

The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(bin): bound repeat stale wakes for parked workers (#3532)

* fix(bin): bound repeat stale wakes for a parked but live worker

A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.

pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.

Two places let that churn re-alarm:

- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
  declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
  have suppressed it. The throttle was never read on this path and was advanced
  by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
  whenever the classification came back `none`, so each tick also bought the same
  declared wait a fresh window. Fixing only the first site changes nothing.

Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.

First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.

Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.

* fix(document): Clarify declared-wait wake cadence documentation

* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor

* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed

* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)

* fix(turnend): accept the away-mode daemon as the supervision owner

While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.

Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.

The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.

The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.

* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage

* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md

* fix(backlog): omit --file from row probes for non-markdown backends (#3582)

* fix(backlog): omit markdown file for beads probes

* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes

* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)

* fix(bin): classify progress updates on requested work as routine (#3589)

The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.

The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.

* fix(bin): preserve captain calls during teardown (#3595)

* fix(bin): never close a captain call during cleanup

A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.

bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.

Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.

The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.

Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.

Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np

* no-mistakes(review): Serialize captain holds and fix backend-aware listing

* no-mistakes(document): Update captain-call retention documentation

* no-mistakes(document): Fix relocated captain-hold backlog diagnostics

* fix(bin): deliver secondmate outcomes to the parent channel (#3592)

* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts

A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:

- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
  exact-line append-once; the merge outcome path and the inactive-outcome
  scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
  watcher poll in a secondmate home: a direct child's whole terminal done or
  failed line is delivered at once with its note, recorded PR, mode, merge
  posture, and scout report pointer, keyed and receipted so it is delivered
  once, and the inactive path yields to it. `report <task-id>` runs the same
  delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
  registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
  and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
  record and refuses, retaining every record, while the channel cannot be
  written.
- The charter opens with the parent-channel rule and confines the mate's own
  appends to judgement; AGENTS.md carries the carve-out at the persona
  address rule and the escalation list.

docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.

* no-mistakes(review): Fix parent outcome retries and reconciliation locking

* no-mistakes(review): Prevent busy children from starving ledger delivery

* no-mistakes(review): Correct ledger metadata and hold occurrence handling

* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons

* no-mistakes(review): Close ledger races and preserve teardown records

* no-mistakes(document): Correct parent-channel receipt and scanner documentation

* no-mistakes(lint): Quote done arguments for ShellCheck compliance

* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks

* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure

* fix(bin): sync remote second mates to primary commit (#3599)

* fix(bin): sync remote second-mate homes to the parent primary commit

Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.

The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.

The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.

/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.

* no-mistakes(document): Document primary-targeted remote secondmate synchronization

* fix(bin): separate captain intent from firstmate specs (#3597)

* fix(bin): split brief task into captain intent and firstmate spec

Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.

* fix(bin): stop task-subsection copies at the next heading

Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.

* no-mistakes(review): Validate brief content and preserve nested specifications

* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies

* no-mistakes(review): Ignore fenced subsection headings during brief validation

* no-mistakes(review): Preserve captain intent across scout promotion

* no-mistakes(review): Enforce safe intent boundaries for legacy promotions

* no-mistakes(review): Allow marked legacy intent and reject empty promotions

* no-mistakes(review): Scope task parsing and overlay legacy intent contracts

* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns

* no-mistakes(review): Preserve later captain clarifications in intent overlays

* no-mistakes(document): Document brief intent enforcement and ownership

* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed

* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint

* fix: start a fresh supervision branch for every main session (#3600)

* fix(pi): start a new supervision branch conversation per main session

The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.

The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.

The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.

* no-mistakes(document): Document fresh Pi supervision conversations

* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks

* feat: restart second mates after instruction updates (#3614)

* feat(update): restart second mates whose instructions changed

/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.

An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.

Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.

fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.

Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.

* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting

* no-mistakes(review): Parallelize relaunches and classify replacement incarnations

* no-mistakes(review): Gate restart actions on live agent state

* no-mistakes(review): Handle failed restart workers without hanging

* no-mistakes(review): Nudge legacy remotes and preserve persist recovery

* no-mistakes(review): Document one-time secondmate restart rollout

* no-mistakes(review): Honor arrived replies and refresh remote profiles

* no-mistakes(review): Revert remote parent profile reconciliation

* no-mistakes(review): Reset remote profile defaults and honor published results

* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates

* no-mistakes(document): Document second-mate restart update flow

* no-mistakes(lint): Fix ShellCheck warnings in restart scripts

* perf: accelerate local validation with bounded concurrency (#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation

* fix: copy PR URLs from durable records (#3648)

* fix: copy PR URLs from records or abstain, never assemble them

Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.

Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:

- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
  or abstain" section requires a URL to be copied verbatim from a durable
  record (the done: PR <url> status line, pr= metadata, or the backlog note),
  forbids assembling owner, repository, host, or number from memory, and has
  the branch report only the identifier it actually holds when no record names
  the URL yet, leaving the PR check unarmed until the worker's ready line
  arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for
  main in place of the bare full-URL mandate.

- Worker briefs (bin/fm-brief.sh, ship and scout rules) require …
Valentino-Sole added a commit to Valentino-Sole/firstmate that referenced this pull request Sep 4, 2026
* fix: start a fresh supervision branch for every main session (kunchenguid#3600)

* fix(pi): start a new supervision branch conversation per main session

The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.

The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.

The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.

* no-mistakes(document): Document fresh Pi supervision conversations

* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks

* feat: restart second mates after instruction updates (kunchenguid#3614)

* feat(update): restart second mates whose instructions changed

/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.

An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.

Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.

fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.

Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.

* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting

* no-mistakes(review): Parallelize relaunches and classify replacement incarnations

* no-mistakes(review): Gate restart actions on live agent state

* no-mistakes(review): Handle failed restart workers without hanging

* no-mistakes(review): Nudge legacy remotes and preserve persist recovery

* no-mistakes(review): Document one-time secondmate restart rollout

* no-mistakes(review): Honor arrived replies and refresh remote profiles

* no-mistakes(review): Revert remote parent profile reconciliation

* no-mistakes(review): Reset remote profile defaults and honor published results

* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates

* no-mistakes(document): Document second-mate restart update flow

* no-mistakes(lint): Fix ShellCheck warnings in restart scripts

* perf: accelerate local validation with bounded concurrency (kunchenguid#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation

* fix: copy PR URLs from durable records (kunchenguid#3648)

* fix: copy PR URLs from records or abstain, never assemble them

Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.

Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:

- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
  or abstain" section requires a URL to be copied verbatim from a durable
  record (the done: PR <url> status line, pr= metadata, or the backlog note),
  forbids assembling owner, repository, host, or number from memory, and has
  the branch report only the identifier it actually holds when no record names
  the URL yet, leaving the PR check unarmed until the worker's ready line
  arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for
  main in place of the bare full-URL mandate.

- Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full
  https:// URL wherever a PR is mentioned - status line, terminal, or summary -
  never a bare "PR 108", so the link is in view as early as the number is.

- bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that
  the task's own done lines contradict, printing both spellings; a log naming
  no URL still records the argument as before. fm_pr_status_ready_urls in
  bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches
  bin/fm-pr-merge.sh, so nothing merges under a contradicted URL.

Tests cover the offline refusal with zero side effects, the recorded spelling
being accepted, markdown-wrapped and punctuated URLs, working lines not
counting, the merge wrapper propagation, a self-hosted merge request with no
forge call, the prompt carrying the rule, and the brief carrying the worker
rule.

* no-mistakes(review): Remove stale PR URL enforcement

* no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`

* fix(bin): disable Claude feedback drafts for fleet launches (kunchenguid#3661)

* fix(bin): disable Claude's feedback-draft flow for fleet-launched agents

Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched
Claude crewmate and secondmate, so /bug and /feedback never queue or
submit a bug report on the captain's behalf. feedbackDrafts is the
documented settings key (Claude Code changelog 2.1.247); the
per-launch CLI flag never touches the captain's global settings.json.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts

* no-mistakes(document): Fix Claude feedback documentation formatting

* fix(bin): layer both feedback-draft controls for defense in depth

The prior --settings-only fix can be overridden by a managed Claude
settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0
alongside --settings '{"feedbackDrafts":"off"}': either control alone
disables the SendFeedback tool, so a managed override of one still
leaves the other in force.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(document): Document Claude feedback-draft suppression ownership

* feat(tests): run three more validation families concurrently (kunchenguid#3662)

* perf(tests): admit three more families to concurrent validation

The three families that `docs/fm-test-isolation-proof.md` recorded as refused
were not refused for concurrency. Each blocker was a test that decided a
property by wall clock, or a script filed where it cannot run. Fixing those
three things admits all three families and recovers 28.6 minutes of local
validation with no assertion removed or weakened.

- `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the
  handoff, sleeping a fixed second, then delegating the move to the real
  binary. Nothing ever killed the fake, so on a host slow enough for the case's
  next assertions to take longer than a second, the orphan woke and completed
  the very move the case requires left undone, and recovery then failed with
  `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs
  during the injected crash showed exactly that, the item moving one second
  after the crash. All four crash injections in the file now go through a new
  `fm_fake_crash_injector` shim that signals the target and returns only once
  it is observably gone, and the pre-move fake never delegates the move at all.

- `tests/fm-session-start.test.sh` proved the startup digest does not block on
  a slow current-state read by timing the whole digest against a fixed
  eight-second sleep, which a loaded host exceeds without the property being
  violated. It now holds that read open until the case releases it and asserts,
  the moment the digest returns, that the read has not finished. A digest that
  waited would wait indefinitely rather than for an interval a slow host can
  out-run, so the assertion is stronger than the bound it replaces. Its scan
  budget moves to the maximum, because the old value left two seconds of margin
  over the fixed sleep and measured the host rather than the deadline that
  `tests/fm-inactive-reconcile.test.sh` owns.

- `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all,
  which put it in the portable serial lane, where Linux CI gate-skips it: that
  real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and
  the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its
  opt-in variable and moves to `live-harness-optin`.

The 28 remaining ungrouped scripts become an enumerated `standalone` family
instead of admitting `unclassified` itself. `unclassified` is the family map's
`*)` arm, so admitting it would silently grant concurrency to every test added
afterwards, which is exactly the population with no proof. A new test still
lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh`
covers that split behaviorally.

Each family passes two consecutive four-worker proofs with zero failures. On
the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap`
756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s
recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock
against 121 minutes of summed script time.

* no-mistakes(document): Refresh concurrent validation and shard documentation

* no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2

* feat: structure no-mistakes ask-user escalations (kunchenguid#3670)

* feat(brief): structure no-mistakes ask-user escalation as event + snapshot file

Crewmates escalating a no-mistakes ask-user gate now report one status
event naming every finding id plus a snapshot file holding the gate's
axi finding records verbatim (id, severity, file, line, description,
authority), using the same shape even for a single finding. The status
line never paraphrases. The format is defined once in fm-dod-lib.sh and
rendered into both the scout and ship rule 6 in fm-brief.sh, so a
promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets
the identical contract as a freshly-spawned no-mistakes ship worker.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei

* no-mistakes(review): Preserve ask-user escalation output contract

* no-mistakes(review): Align escalation format test expectation

* no-mistakes(review): Scope ask-user escalation instructions correctly

* no-mistakes(review): Remove ask-user from generic decision rules

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix(bin): require self-sufficient no-mistakes intent (kunchenguid#3671)

* fix(bin): require a self-sufficient no-mistakes intent

A no-mistakes worker's --intent is only as useful as the string it
passes. PR kunchenguid#3604 shipped with an intent that was only "do 1, 2, 3, 7
from the report": the real contract lived in a private scout report and
never reached --intent, so nobody holding that string plus the codebase
could have derived the specification.

This is pure instruction at the contract's one owner; no spawn-side or
promotion-side check is added.

- bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now
  states that the --intent string must be self-sufficient (the string
  plus the codebase reconstructs roughly the same specification) and
  tells the worker to write the substance of any report, decision, or
  PR the captain's intent refers to into --intent rather than the
  pointer, while Firstmate build instructions and the worker's own
  decisions still stay out. The spawn-time overlay points back at that
  rule so its "supersedes" wording cannot cancel it, and the header's
  owner statement carries the rule.
- AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to
  include the substance of referenced material when filling
  ## Captain's intent, and section 11 points at the owner of the rule.
- tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the
  rendered brief and launch contract carry the rule.

Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim

* no-mistakes(document): Replace incident-specific intent test commentary

* fix: accelerate local Bearings snapshot composition (kunchenguid#3499)

* Speed local fleet snapshot composition

* no-mistakes(review): Stabilize task inventory during concurrent snapshot composition

* no-mistakes(document): Document local snapshot observation concurrency

* no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks

* no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks

* no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks

* no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky

* no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks

* fix(snapshot): keep live observations generation-coherent

* no-mistakes(review): Keep secondmate observations generation-bound without copying reports

* no-mistakes(document): Document generation-coherent snapshot observations

* test(bearings): measure local read overlap instead of wall-clock budget

The large-local-snapshot regression asserted that a whole snapshot
composed in under five seconds. That bound measures how loaded the host
is, not whether the per-task reads actually overlap, so it failed
intermittently on a contended machine: one run in six on a box at load
16-20, landing exactly on the five second boundary.

Time a serialized run and a concurrent run of the same workload instead
and require the concurrent one to save at least two seconds. Both runs
pay the same composition overhead, so the difference isolates the
overlap this change delivers. Five one-second reads serialize into five
seconds and overlap into about one, and re-serializing the reads
collapses the saving to roughly zero, so the assertion still fails
loudly if the concurrency regresses.

Also bump the pinned Bearings test count to 48, since rebasing onto the
current default branch picked up its captain-hold test.

* no-mistakes(review): Restore JSON-derived decision flags

* no-mistakes(review): Unify status-derived snapshot observations

* no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass

* fix: prevent stale supervision wake loops (kunchenguid#3672)

* fix(bin): stop the supervision branch's stale-ack and ghost-report loops

Clean-slate implementation of the four authorized recommendations from the
supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal
form, superseding PR kunchenguid#3604:

- fm_branch_report refuses a task the wake being handled never named. The
  extension fixes the reportable task set from the eligible rows before each
  prompt (signal and stale rows resolve to their tasks, a heartbeat allows any
  task with a live record, fleet is always allowed), so a report typed from
  memory about a task whose records teardown already removed is never stored
  or delivered.
- An acknowledgement that consumes nothing says "nothing was acknowledged
  through N" and prints the exact --ack-through / --recovery-generation
  command for the current presented wake, instead of "re-run the drain",
  which re-fed the same stale acknowledgement in a loop.
- bin/fm-guard.sh no longer tells the branch actor to drain queued wakes
  while it is handling them; it names the granted rows instead.
- Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and
  descendants; the index rebuild and the append-side index write both skip a
  task with neither a live record nor a status log, so the branch's report of
  a teardown it just performed is stored without recreating the index.

No new locking, no spawn-generation binding, and no retired-task refusal: the
branch can still report the outcome of a task it just tore down, and the
teardown test now proves that path end to end.

* fix(bin): narrow the branch report scope and guard silence to the minimal form

Apply the four review decisions on the clean-slate branch:

- A signal or stale prompt may report only the tasks its own rows resolve
  to; fleet is refused there too. A heartbeat review is not scoped by task
  at all, so the extension no longer tracks live task records and refuses
  nothing by task id during a fleet review.
- The outcome-index rebuild no longer skips retired tasks; the append-side
  skip alone keeps a torn-down task's index from being recreated.
- bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor
  instead of printing a replacement note.

* no-mistakes(document): Align supervision docs with scoped wake handling

* fix(bin): avoid fleet snapshot argument limits (kunchenguid#3677)

* Fix fleet snapshot large JSON transport

* no-mistakes(review): Captain: file-back fleet snapshot transport safely

* no-mistakes(review): Captain: file-back parent summary aggregation

* no-mistakes(ci): Rebased the PR's three commits onto f4d7875 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor

* fix(bin): attribute active runs with unfetched pipeline heads (kunchenguid#3681)

* fix(bin): recognize active pipeline fix rounds with unfetched run heads

A no-mistakes fix round advances the run head beyond the submitted head,
and the pipeline commits in its own checkout, so the task copy never
receives the new commit object. fm-crew-state's strict head rule rejected
the active row, the coarse runs-list scan skipped it and matched the
older failed row at the submitted head, and an active validation read as
failed (observed on model-routing-benchmark-hardening: active head
ac61c64 vs task copy at fb47636d).

fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns
runs-ledger attribution: the branch's newest row alone decides, and a
newest row whose head cannot resolve locally is recognized only as a
provable pipeline-owned continuation - active (running) and anchored by
the immediately older row for the same branch having ended at exactly
this worktree's HEAD. The reader keeps the axi TOON as full detail for
that proven same-branch run. Unanchored, ancestor-anchored, and terminal
unresolvable rows stay unattributed, so branch-name coincidence and other
tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact
prior semantics for teardown (verified by the full teardown suite).

Tests: reproduction regression for the unfetched active fix head (reads
working via full run-step detail), coarse-path continuation when axi
answers another branch, and negative controls for the unanchored active
row and the unresolvable terminal row with the historical fallback
preserved.

Ported onto upstream/main f4d7875, where kunchenguid#3194 independently added the
branch_sync custody exemption on the full axi-status path: both mechanisms
now coexist, each owning one surface (TOON custody on the full path, the
runs ledger on the coarse path). The port deletes the superseded coarse
scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers
(fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the
exemption comment's "the one exemption" phrasing now that a second
complementary exemption exists, and points the stale
FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree
(judge follow-up #1). The parent coarse-guard test's fixture is the
ledger-anchored continuation shape, so its expectation flips to the fixed
behavior (working via run-step, never the older failed row); a new
mismatched-anchor coarse negative control preserves that guard's original
no-anchor protection (pane answers, never the older row).

* no-mistakes(document): Clarify pipeline attribution documentation

* fix(bin): pre-register claude workspace trust at spawn time (kunchenguid#3663)

* fix(bin): pre-register claude workspace trust for task worktrees

A claude crewmate launched into a fresh task worktree met Claude Code's
interactive workspace-trust dialog before it ever read its brief, and firstmate
could not answer it: the key plane carries only Enter, Escape, and C-c with no
arrow navigation, and the dialog's selection starts on "No, exit", so the
documented Enter recipe ended the session instead of accepting it. Two workers
wedged this way and were unblocked only by hand-seeding the trust store per
path.

--dangerously-skip-permissions does not cover that gate. `claude --help`
records the dialog as skipped only in non-interactive mode, through -p or a
non-TTY stdout, and a crewmate pane is interactive, so there is no launch flag
to reach for.

fm-spawn now pre-registers the worktree through bin/fm-claude-trust.sh in the
existing claude branch, before the project settings that the same gate would
otherwise block, and refuses the spawn when that write fails rather than
launching a worker that would wedge.

The scope test is the safety property and is structural rather than a path
policy: the path must be a linked git worktree, sharing the spawning project's
common dir, whose top level is exactly the resolved argument. Git is the ground
truth, so the argument is never trusted on its own word, and a primary
checkout, an unrelated repo, a worktree subdirectory, a plain directory, and a
home directory are each refused rather than warned about or skipped. A
treehouse or orca path prefix was deliberately avoided because treehouse's root
is configurable, which would make a prefix both wrong and a new policy surface.
One structural test covers both worktree providers.

tests/fm-claude-trust.test.sh pins both halves, including a case where HOME is
itself a valid linked worktree so the home guard is proven load-bearing rather
than passing vacuously, plus the spawn-level proof that a claude spawn trusts
its worktree and launches with the brief pointed at the same store.

The adapter reference no longer tells a firstmate to press Enter on that
dialog, and the shared trust reference now names every harness surface: which
harnesses gate, which suppress at launch, which dodge the gate, which now
pre-registers, and that a claude secondmate is excluded by design.

The spawn fixture runs each spawn against a throwaway HOME so the suite cannot
write the developer's real store, isolating through HOME rather than
CLAUDE_CONFIG_DIR because the spawn forwards a set CLAUDE_CONFIG_DIR onto the
launch command that launch-shape assertions read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v

* fix(bin): create the staged trust store exclusively

The staged store was written to a predictable pid-based path with a plain
write, which follows a symlink. Where the Claude config directory is writable
by another local account, that account could pre-create the path as a symlink
and redirect the write into another file the launching user owns.

The staged name now carries random bytes and is created with an exclusive
"wx" open, so an existing path is refused outright instead of followed. The
happy-path test also asserts no staged store survives the rename.

The durability comment now states the residual window plainly: the readback
proves the entry landed, not that it survives, because a vendor session that
rewrites the whole store afterwards can still drop it and no lock closes that
window when the writer is Claude itself. The worker then meets the dialog and
stalls, which reaches firstmate as the ordinary stale wake rather than as
silent success.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v

* no-mistakes(review): neutralise CDPATH in claude trust scope guard

* no-mistakes(review): sandbox HOME in spawn tests, drop out-of-scope artifacts

* no-mistakes(review): refuse unresolvable git dir, compact store, fix secondmate doc

* no-mistakes(review): clear git env overrides, resolve symlinked store target

* no-mistakes(review): degrade without node, fix Pi gate claim, record trust proof

* no-mistakes(review): refuse without node, pin CLAUDE_CONFIG_DIR in spawn tests

* no-mistakes(review): refuse relative config dir and concurrent store modification

* no-mistakes(review): correct orca worktree claim, clean staged store on failure

* no-mistakes(review): restore pretty-printed store, correct trust dialog docs

* no-mistakes(review): arm trust gate before busy state to avoid orphans

* no-mistakes(document): record claude trust pre-registration in its owner docs

* no-mistakes(document): note orca limit for claude trust pre-registration

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-spawn.sh by moving the Claude trust gate earlier rather than adding cleanup machinery. Diagnosis: Greptile reported that when Claude trust registration fails on tmux/Zellij/cmux/non-projected Herdr, the exit runs after the backend endpoint and /tmp/fm-<id> were created, and the abort trap cleans neither. The endpoint half is pre-existing, deliberate architecture — the two refusals immediately above the gate (the 60s `treehouse get` timeout at fm-spawn.sh:2550 and `validate_spawn_worktree` at :2487) also exit with the endpoint live and direct the operator with "inspect window $T"; spawn_abort_cleanup only reclaims orca endpoints (already covered via ORCA_ABORT_CLEANUP) and herdr projections. The temp-root half was genuinely introduced by this PR: the gate was placed beside the busy-state arm, ~30 lines after `mkdir -p "$TASK_TMP/gotmp"`, and fm-teardown can only find that root through `tasktmp=` in a meta record a refused spawn never publishes. Root-cause fix (smallest correct change, no new subsystem): - bin/fm-spawn.sh — moved the `claude*` trust gate from inside the busy-arm block up to the first point $WT is known, immediately after the `freshen_spawn_worktree_base` block and before TASK_TMP creation, the STATE setup, and the relaunch `clear_relaunch_harness_wiring` retirement. A refusal now leaves no temp root, no retired relaunch wiring, and no busy record; only the endpoint remains, in the same class as the two refusals just above it. - bin/fm-spawn.sh — the refusal message now ends with "inspect window $T", matching the existing convention so control/teardown can identify the endpoint. $T is set for every backend on the non-secondmate path. - bin/fm-spawn.sh:196 — header note corrected from "before any state is armed" to "before any per-task state exists". - tests/fm-claude-trust.test.sh — the existing refused-spawn test's own comment claimed "before any task state exists" but only asserted busy state. Renamed to test_refused_spawn_leaves_no_task_state and added an assertion that /tmp/fm-<id> is absent, with the task id suffixed by the test process pid so the assertion reads only this run's path (a stale /tmp/fm-refusedspawn from the fixed-id version was in fact present on this box). No assertions on implementation source bytes. Verification run locally: - The new assertion fails against the pre-fix bin/fm-spawn.sh ("not ok - a refused spawn stranded a temp root no teardown can find") and passes after — a real before/after regression proof. - tests/fm-claude-trust.test.sh: 20/20 ok. - tests/fm-backend.test.sh, fm-backend-orca, fm-control-relaunch, fm-spawn-dispatch-profile, fm-trace-context-spawn, fm-gotmp: all pass. - tests/fm-backlog-atomicity.test.sh: rc=0, 79 assertions ok. - bin/fm-lint.sh (repo's single lint owner, pinned ShellCheck 0.11.0 + actionlint 1.7.12): clean. - No /tmp/fm-refusedspawn* leftovers after the runs. Scope respected: no trust subsystem, no policy layer, no config surface, no endpoint-cleanup mechanism added; the change is an ordering move plus one error-message clause and the test that pins it. Adapter references and docs made no ordering claim, so none needed updating. Changes are left uncommitted in the worktree for the outer executor

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix: restart every live second mate after updates (kunchenguid#3690)

* feat(update): restart every live second mate after a successful update

/updatefirstmate only restarted a second mate when that pass advanced its
AGENTS.md or .agents/skills. An already-current home was skipped entirely, a
bin/-only advance was steered instead, and a remote host that could not report
its instruction diff was downgraded to a re-read. A running agent also freezes
its launch-time wiring - turn-end hooks, harness flags, per-harness feature
switches - and none of that is derivable from a file diff, so an unchanged
tracked surface is not evidence the agent is already on the current behavior.

Restart is now unconditional on a successful update of that home. Every live
second mate the pass leaves on the target commit is restarted, whether it
advanced or was already there.

The safety contract is unchanged: open records are persisted before the agent is
replaced, nothing is forced, stashed, or discarded, a home the pass had to skip
is not restarted at all, and a mate whose runtime cannot prove a restart keeps
the honest re-read path and is never reported as reloaded.

bin/fm-ff-lib.sh gains a settled-state hook that fires for a home left at the
base whether it advanced or was already there, and never for a skipped one; the
instruction-gated hook the session-start convergence sweep uses is untouched.

Regressions: fm-update pins the already-current mate into the restart set and
the unprovable one into the nudge set, and fm-secondmate-restart drives both
real commands end to end - an already-current home is named, persisted, and
genuinely replaced with its checkout untouched, while the unprovable one keeps
its running agent.

* no-mistakes(document): Document unconditional secondmate restarts

* fix(bin): close pending-reply decisions via resolve-key (kunchenguid#3696)

* fix(bin): close reserved pending-reply keys via fm-send --resolve-key

fm-send wrote answered: notes that the reserved-key fold ignores, so
operator closes exited 0 while OPEN DECISIONS kept the decision open.
Speak the owning library's close vocabulary on that path, and refuse
when a reserved close cannot take effect.

* no-mistakes(review): Safely quote manual decision-close recovery commands

* no-mistakes(review): Reject unclosable overlong decision keys before sending

* no-mistakes(review): Remove contract suffix from open decisions hint

* no-mistakes(document): Document resolve-key line-cap refusal

* fix(bin): prevent false missed-reply escalations (kunchenguid#3697)

* fix(bin): stop false missed-reply escalations for same-basename self-home answers

A healthy secondmate that wrote corr= to its own state/<id>.status never matched the parent channel, so recovery confirmed and the record escalated as pending-reply-missed. Make the report helper resolve the parent channel itself, skip parent-replies.status as wrong-home, put a readable sighting path on the missed line, and restatement-copy only that same-basename self-home file onto the parent channel.

* no-mistakes(review): Resolve late replies before recovery escalation

* no-mistakes(review): Tighten reply routing and regression coverage

* no-mistakes(review): Preserve reply paths and require explicit home

* no-mistakes(review): Encode wrong-home paths before persistence

* no-mistakes(document): Document corrected secondmate reply routing

* no-mistakes(lint): Fix pending-reply ShellCheck warnings

* feat: add verified Gemini crewmate runtime (kunchenguid#3695)

* feat(harness): verify gemini as a crewmate runtime adapter

Adds Gemini CLI as a fourth dispatch target alongside claude, codex, and
grok, scoped to crewmate and scout work only. Every axis was proven against
gemini-cli 0.58.0 rather than inferred; docs/verification/runtime-backends.md
carries the dated evidence and names what stayed unverified.

Busy state is semantic, not rendered: BeforeAgent opens a turn and AfterAgent
and SessionEnd close it. AfterAgent also fires on a manual interrupt, so a
cancelled turn closes its own record.

Three findings shaped the wiring rather than a config line:

- --skip-trust and GEMINI_CLI_TRUST_WORKSPACE=true are presented by the CLI
  as equivalents and are not. A controlled A/B showed --skip-trust leaves
  project configuration unloaded, so workspace skills never load.
- The worktree's .gemini/settings.json is the PROJECT's committed settings
  file, unlike claude's settings.local.json. Firstmate's hooks therefore go
  to a firstmate-owned state/<id>.gemini-settings.json reached through
  GEMINI_CLI_SYSTEM_SETTINGS_PATH, which also works untrusted and merges
  with a project's own hooks instead of replacing them.
- The shipped CLI is a node bundle whose live process reports comm as
  MainThread, so ancestry cannot see it. GEMINI_CLI=1 is load-bearing and is
  tested before an inherited CLAUDECODE, and pane liveness identifies gemini
  from the script argument through the new bin/fm-gemini-lib.sh.

Gemini is refused for secondmates: it has no primary supervision protocol.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L

* test: clear gemini's marker in launch and detection expectations

Every non-gemini launch now clears GEMINI_CLI the way it already clears
cursor's markers, so the two tests that pin the exact launch prefix are
updated to match. The harness-detection tests that scrub foreign markers
before probing ancestry scrub GEMINI_CLI too, so running the suite from
inside a gemini session cannot produce a false verdict.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L

* docs: classify the gemini harness reference

The documentation inventory is the single classification owner for maintained
prose surfaces, and every surface must appear in it exactly once. The new
harness reference is agent-runtime, matching its siblings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L

* no-mistakes(review): Narrow Gemini ancestry detection

* no-mistakes(review): Restrict Gemini hooks to canonical launches

* no-mistakes(document): Document Gemini adapter support boundaries

* no-mistakes(ci): Fixed Gemini process identity when interpreter or script paths contain whitespace. Tmux liveness now uses NUL-delimited /proc argv on Linux, with the existing flattened ps fallback elsewhere. Added a real-process regression test. Verified with the Gemini harness test suite, full fm-lint, ShellCheck, and git diff --check. The CI and Require no-mistakes runs were action_required/attestation outcomes rather than code failures

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(teardown): conclude parked runs advanced past task copy (kunchenguid#3704)

* conclude parked runs the pipeline advanced past the task copy

A no-mistakes fix round commits in the daemon's own gate-repo clone, so a
run parked at a gate can carry a head whose object the task copy never
received. Teardown's strict object-local identity rule then declined to
conclude the run, and cleanup left it parked forever holding a fleet slot
(observed 2026-09-03; the same masking condition PR 3681 fixed on the
read path, now closing the teardown half its scope boundary deferred).

task_status_is_own_parked_run now falls back - only when the reported
head resolves to no local object - to the one shared runs-ledger
attribution rule fm_nm_runs_status_for_worktree (bin/fm-nm-run-lib.sh),
whose anchored continuation proof binds the branch's newest active row
to this worktree's exact submitted head. Foreign branches, stale
history, terminal rows, ancestor-only anchors, diverged newer rows, and
ambiguous multi-row shapes all still refuse, and runs that are actively
running, fixing, or in CI remain untouched: only the parked-at-a-gate
determination ever reaches the abort. No sqlite access, no fetches into
another task copy, no custody changes, no duplicated matching logic.

* tighten the parked-run ledger fallback and pin both judge corrections

The teardown ledger fallback now authorizes concluding this task's parked
run only when the shared runs-ledger rule's proved answer is the explicitly
active word (running): a terminal newest row - even anchored at exactly the
worktree's head - is finished history and never an abort authorization.
The read path may classify the same owner's answer; teardown's abort must
never fire for a run that already ended.

Two bounded pre-validation corrections from the implementation review:
- a fetched-object counterfactual pins the strict-rule path: a pipeline fix
  head fetched into the task copy aborts through object-local identity
  alone, with an empty ledger and a proof the runs query never fired;
- a negative fixture pins the tightened boundary: an unresolvable reported
  head with a terminal newest same-branch row anchored at the worktree head
  engages the ledger fallback and still refuses, so the refusal is the
  terminal-word boundary and not an earlier guard.

* no-mistakes(review): Bind teardown ledger fallback to validated run heads

* no-mistakes(review): Restore validated advanced-head ledger continuation

* no-mistakes(review): Reject invalid ledger dates and terminal statuses

* no-mistakes(document): Document teardown ledger scan limit

* feat(bin): show requested vs effective model in Herdr agent view

Track spawn-config requested_model separately from runtime-verified
effective_model, probe Claude/Pi transcripts for exact API ids, push
compact display metadata to Herdr, and preserve verified models across
relaunch/compaction hooks without inferring aliases as truth.

* fix(bin): keep re-probing effective model after first exact reading

fm-model-sync.sh only probed for the runtime-verified effective model
while it was still pending/UNKNOWN, so a session that later switched
models (manual switch, provider fallback) kept displaying the first
verified model forever and never appended a fallback-history entry.
Probe unconditionally instead; fm_model_record_effective already
no-ops when the probed value is unchanged, so this stays cheap.

Addresses the Greptile P1 finding on PR kunchenguid#3705's fm-model-sync.sh.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm

* fix(bin): distinguish Cursor Grok, direct xAI Grok, and Anthropic Claude in the display

Kapitänskorrektur: harness alone conflated Cursor-hosted Grok models
(cursor-grok-4.6-*) and direct xAI Grok models (xai/grok-4.6) under
one generic label, and displayed Anthropic Claude without naming the
provider. Add fm_model_source_label, pattern-matched on the verified
exact model id, so the compact display always reads Cursor · Grok,
xAI · Grok, or Anthropic · Claude with the exact model id appended.
Falls back to the existing harness label for every other model. No
routing change: this only affects display strings.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm

* fix(bin): wire model-sync into the Pi extension's turn lifecycle

fm-model-sync.sh was only invoked from Claude's SessionStart/
UserPromptSubmit/Stop hooks; the Pi harness's own extension
(state/<id>.pi-ext.ts) never called it, so a Pi-hosted session (e.g.
a pi/xai-grok crewmate) never refreshed its effective model after the
first probe and Herdr kept showing the stale value with no
fallback-history entry. Call fm-model-sync.sh from the same
agent_start/turn_end boundaries Pi already uses for busy-state and the
turn-end notification touch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm

* fix(bin): serialize fm-model-sync.sh's meta read-probe-write

Overlapping lifecycle events (Pi's agent_start/turn_end, Claude's
SessionStart/UserPromptSubmit/Stop) can invoke fm-model-sync.sh
concurrently for the same task. The unlocked read-probe-write let
interleaved runs revert a newer effective model, mismatch its
source, or duplicate a model-history entry. Serialize the critical
section through the same per-task meta lock fm-spawn.sh already uses
(fm_meta_lock_path + fm_lock_acquire_wait/fm_lock_release), released
before every exit path.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Arthur Haro <38157909+haroarthur@users.noreply.github.com>
Co-authored-by: Nicolas Payette <nicolas.payette@specira.ai>
Co-authored-by: Jon Roosevelt <rooseveltadvisors@gmail.com>
Co-authored-by: att430 <41454889+att430@users.noreply.github.com>
Co-authored-by: Valentino-Sole <171032438+Valentino-Sole@users.noreply.github.com>
kunchenguid added a commit that referenced this pull request Sep 6, 2026
* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR #823 added and PR #1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
eduardstan added a commit to eduardstan/firstmate that referenced this pull request Sep 7, 2026
* fix(bin): wait out transient primary-checkout reads in the spawn worktree poll (#3834)

* fix(spawn): keep the worktree poll from adopting the repository primary

After `treehouse get` is sent, the worktree-discovery poll reads the pane's
foreground-process cwd. While treehouse is still fetching and checking a slot
out, the foreground process is treehouse itself and it reports the repository's
PRIMARY checkout as its cwd for several seconds. The poll accepted any path
that merely differed from the spawning project, so from a linked spawning home
- whose project is itself a worktree of that repository - it adopted the
primary, and the isolation guard then refused a launch whose slot treehouse
went on to create normally.

Screen every candidate with the isolation guard's own conditions, extracted as
spawn_worktree_isolated, so a read the guard would reject stays a transient the
poll keeps waiting through. The two-consecutive-reads rule and the guard as
final backstop are unchanged; a pane that never reaches an isolated worktree
still fails at the existing 60s deadline, now naming the last path it reported.

The already-settled timing assertion counted whole-spawn wall time against a
5s budget and failed on unmodified HEAD on slower machines; it now counts pane
reads, which is what "one confirming read, not an extra cycle" actually means.

* fix(spawn): say which path the worktree wait rejected, and why

Screening every discovery-poll candidate means a host that never reaches an
isolated worktree spends the whole 60s window before refusing. That wait is
deliberate - separating a transient from a terminal misconfiguration needs
machinery this path does not want - so the refusal explains itself instead:
the isolation check records why a candidate failed, and the deadline names the
last path seen together with that reason. Message and diagnostics only; the
poll's control flow is unchanged.

Two suites asserted the guard's wording on paths the poll now rejects rather
than adopts, so their refusal arrives from the deadline instead: realign
fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also
asserting the stated reason, and the second the metadata absence it was
missing) and the herdr projection e2e's forced non-worktree cwd.

* no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test

* no-mistakes(document): document spawn poll isolation screen in fm-spawn header

* test(spawn): make the non-git isolation case non-git anywhere

The refusal-reason assertion for a path outside any repository assumed TMPDIR
is not inside a git repository. Where it is, git walks up from the temporary
directory, finds that repository, and the spawn reports the subdirectory cause
instead - so the case passed or failed on a property of the host rather than on
the behaviour under test.

Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES,
which git documents as not chdir-ing up into a listed directory while looking
for a repository. Git never excludes the directory being searched, so the
ceiling is the parent of the path handed to the spawn.

The assertions pin which cause fired rather than the sentence that explains it,
leaving the operator wording free to improve.

* no-mistakes(document): point spawn poll comment at the isolation screen's comparison

* no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure

* fix: support stock macOS Bash 3.2 paths (#3732)

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): read orphaned green ci monitor and daemon-down failed record as not failed (#3846)

* fix(bin): read an orphaned green ci monitor as held-for-merge, not failed

A no-mistakes run held for a captain merge decision keeps its ci step
polling until merged or closed; when the shared daemon restarts under
that poll, the run is recorded failed although every substantive step
completed and GitHub reports the PR green. A monitor whose only
remaining job is to observe a human decision must not convert the
absence of that decision into a failure verdict.

fm-crew-state.sh now reclassifies a terminal failed run as done
(held-for-merge), surfacing the run's PR URL, when the steps table
shows every step completed except exactly ci failed and the ci log's
last recognized marker reads checks green. A genuinely red check, an
unreadable ci log, or a second failed step keeps the failure.

* no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed

* no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification

* fix(spawn): verify preserved backlog state after interrupted spawn delivery (#3852)

* fix(bin): read preserved spawn state back before the interrupted exit claims it

The deferred-signal exit path asserted the paired task record and
In-flight backlog state were preserved without reading either back,
exactly when a reader is least able to check (fm-yi4j evidence,
2026-09-05). The commit's exit status alone has been observed to agree
with a row that did not actually move.

The exit path now re-reads the record and the row under the same
per-task lock as the commit, repairs a row the commit believed it
moved, and phrases the error as exactly what was verified or attempted
- verified preserved, repaired and verified, or an explicit
preservation-could-not-be-verified with the reason and hand-closeout
instruction. Two behavior tests drive a lying tasks-axi start through a
real interrupted spawn and assert the printed claim and the real
backlog state agree.

* no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask

* no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean

* docs: correct runtime-backend maturity labels for Herdr (#3821)

* docs: correct stale tmux/herdr backend maturity claims

Herdr now has 21 test files, its own required CI job (tests-herdr)
that installs a pinned build and hard-fails on "skip: herdr not
found", while tmux has 3 test files and is only required as a
dependency of the portable-serial e2e lane. zellij, orca, and cmux
still have no CI lane at all. AGENTS.md and docs/herdr-backend.md
still called Herdr merely "experimental" alongside those three,
misleading every session and reader about actual coverage.

Update AGENTS.md's config/backend entry, the opening lines of
docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend
section of docs/configuration.md, and the matching claims in
docs/architecture.md, CONTRIBUTING.md, and README.md so they agree
and distinguish tmux (default), herdr (own required CI lane, largest
suite, Windows still spike-only), and zellij/orca/cmux (still
experimental, no CI lane). No behavior, selection order, or
dispatch logic changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj

* no-mistakes(review): docs: fix stale herdr label and CI-lane wording

* no-mistakes(document): docs: align tmux adapter label in scripts.md

* no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page

* no-mistakes(review): docs: drop windows claim, align contributing backend wording

* no-mistakes(review): docs: trim duplicated CI claim from herdr opening line

* no-mistakes(review): docs: drop unguarded largest-test-suite superlative

* no-mistakes(review): docs: restore tmux verified label and README experimental scope

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: restore intent-targeted no-mistakes validation (#3865)

* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR #823 added and PR #1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: verify Treehouse slot ownership before teardown (#3837)

* fix(bin): verify pool-slot ownership before returning a worktree slot

Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.

Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.

* no-mistakes(review): Protect slots across cloned Firstmate homes

* no-mistakes(test): Gate teardown locking on genuine Treehouse slots

* no-mistakes(test): Clarify pooled descendant slot gating

* no-mistakes(test): Synchronize watcher re-arm test on process exit

* no-mistakes(test): Wait for watcher cleanup before timeout escalation

* no-mistakes(document): Document pool-slot ownership safeguards

* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion

* fix(bin): resolve relative origins from repository root

* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout

* feat(bin): add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary (#3867)

* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary

Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.

Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test: prove the omp guard continuation through a guard spy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* fix(spawn): clear the gemini marker at the omp launch boundary

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): force the guard stage by freezing the watcher and clear lint findings

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): reap the live lab by path and record omp's rpc shutdown as a note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces

Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin

* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs

* no-mistakes(review): omp: pin config-model validation with a test, trim overlay

* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family

* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes

* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux

* no-mistakes(document): docs: add omp subagent-guard row, fix live test header

* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π|󰵗|pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)

* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|󰵗)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix(pi): invoke Bash helpers correctly on native Windows (#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>

* test(bin): pin teardown outcomes for squash-merged rebased branches (#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>

* fix(bin): keep supervision armed for registered custom checks (#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes #3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner

* fix(bin): resolve Treehouse locks for remote secondmate homes (#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership

* fix(bearings): repair board listening and decision reconciliation (#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable state - refuses
rather than acting.

Regression coverage fails without each fix, and pins every leak path: a bare
reconcile, a standalone close or note with no pending request, an any-channel
reconcile, an annotated selection from a freeform card, and a
generation-skewed authorization. An opt-in guard re-proves the lavish-axi
shapes and the reopen against the installed tool.

* fix(bin): quote the done comparison in the reconcile intake

shellcheck SC1010 reads the bare word as the loop keyword. The failed run
never reached its lint step, so this shipped in the recovered content.

* no-mistakes(review): Publish reconciled parent resolution before request retirement

* no-mistakes(review): Clarify committed cleanup and reconcile reservation scope

* no-mistakes(review): Preserve remote cards and legacy answer compatibility

* no-mistakes(test): Separate live claim release from stale reclamation

* no-mistakes(test): Allow terminal self-retirement during active capture

* no-mistakes(document): Document Bearings repair contracts

* no-mistakes(ci): Stabilized the failing Herdr presentation E2E by serializing test-harness Treehouse allocator calls, preventing concurrent recovery spawns from claiming the same pool slot while preserving Herdr concurrency coverage. Verified with the full E2E suite on Herdr 0.8.2, bash syntax checks, ShellCheck, and git diff checks

* fix(bin): allow pooled spawns without a git origin (#3885)

* fix(bin): skip pooled-worktree freshness fetch when no origin is configured

An origin-less local-only project has nothing remote to be stale
against, so fm-spawn's freshen_spawn_worktree_base refused to launch
crews for it. Detect a missing origin remote and skip the fetch
freshness gate entirely; an existing-but-unreachable origin keeps
refusing as before.

* no-mistakes(review): Preserve pool safety for absent and unusable origins

* no-mistakes(review): Refuse empty origin configurations during pooled spawn

* no-mistakes(review): Detect empty origin sections across config includes

* no-mistakes(review): Honor globbed includes when detecting origin configuration

* no-mistakes(review): Document conservative conditional include handling

* no-mistakes(review): Use Git-resolved config files for origin detection

* no-mistakes(review): Document included empty-origin detection boundary

* no-mistakes(document): Document originless pooled spawn behavior

* feat(tests): run live harness guards by default when available (#3889)

* feat(tests): run live harness guards by default where the harness is installed

The 24 live-harness guards each opened with their own env check, so on the
machine that has every harness - the one the product and its validation
actually run on - all of them skipped and passed. Fourteen had never been run
by the pipeline at all.

tests/lib.sh gains fm_live_gate as the single owner of that decision: a guard
that spends no model tokens runs wherever its tools are installed, a guard that
submits prompts stays opt-in, an absent tool is a named capability skip, and a
guard's own variable or FM_LIVE forces it on (turning an absent tool into a
failure) or off. Every live guard now opens with it, which also carries the
test-suite gate-refusal bypass into the guards that never sourced the shared
helpers and were therefore refused whenever a gate agent ran them.

bin/fm-test-run.sh records what a skip means: the family's expected class is
live-capability rather than a bare env opt-in, and each gate skip's reason is
logged and written to the timing artifact, so a lane can say which tool this
host could not exercise.

Only the token-free guards flip to default-on: composer-matrix, the harness
liveness drift guard, and the Herdr version floor. cursor-primary submits three
prompts, so it stays opt-in.

Running the drift guard unasked immediately found a real defect it existed to
catch: it resolved the harness through a generic `command -v cursor`, which on
a machine that also has the Cursor editor finds the editor launcher rather than
cursor-agent. That binary exits at once, leaving a bare shell in the pane and a
liveness-drift failure no classifier change could fix. It now asks
fm_cursor_resolve_binary first, the same verified owner fm-spawn uses.

CI installs the public Pi package in the portable serial lane and fails on its
skip token, so the Pi extension tests stop passing silently against a package
that is not there. No secret is added.

Verified on macOS 26.5.2 arm64: the drift guard runs with no variable set and
classifies 8 installed harnesses alive; the Herdr version-floor guard runs by
default and checks 4 real releases; every live guard refuses together under
FM_LIVE=0.

* fix(tests): keep the composer-matrix guard opt-in

Running it unasked is red on a healthy machine for reasons no code change here
removes: a harness that has not trusted this checkout sits on its own trust
dialog, which the guard treats as an unreadable composer and correctly fails.
The opencode 1.18.29 and grok 1.0.13 composer drift it also surfaced reproduces
identically on main and is filed as separate work.

So this token-free guard stays opt-in with the reason stated in its header, and
the coding guidelines record the narrow exception: a guard whose verdict
depends on host state that installing its tools does not establish may stay
opt-in, because one that is permanently red is one the fleet learns to ignore.
The other two token-free guards keep running by default.

* no-mistakes(review): Wire bearings guard and remove composer exception policy

* no-mistakes(review): Run Pi responsiveness guard by default

* no-mistakes(review): Gate AFK Pi Herdr through authoritative family sweep

* no-mistakes(review): Sanitize live gate test environments

* no-mistakes(document): Document default-on live guard behavior

* no-mistakes(ci): Fixed CI by installing the Pi package in portable-parallel-1, where fm-pi-primary-types.test.sh runs, and enforcing its package-missing gate skip there. Verified with fm-lint.sh, workflow actionlint, coverage partition checks, lane membership, and git diff checks

* no-mistakes(ci): Fixed CI’s Pi typecheck skip enforcement by giving npm, tsc, and Pi-package capability skips a shared prefix and configuring both relevant CI lanes to fail on that prefix. Verified missing tsc emits the expected skip, missing Pi package becomes a runner failure, and actionlint, ShellCheck, and git diff checks pass

* fix(bin): refuse test runs in the primary checkout when a task marker is set (#3891)

* fix(bin): refuse the behavior suite in the repository primary checkout

A task worker's isolated worktree placement is verified exactly once, when
its task starts, and nothing re-checks it afterwards. A worker that later
changes directory into the repository's primary checkout runs its Git
commands, and this branch-switching suite, against the one checkout every
linked worktree resolves against and every landing merges into. A run that
dies mid-suite can leave that checkout on a stray branch.

bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task
worker and the runner resolves to the primary checkout, every executing mode
exits non-zero before selecting a suite, with one line naming the primary
path and pointing at the assigned task worktree. The predicate is the one
bin/fm-spawn.sh already uses for launch placement: the working tree's own
git dir is the repository's common git dir, which separates the primary from
every linked worktree even when their top levels differ. A run with no
FM_TASK_ID set is unchanged, and so are the inspection modes, which execute
nothing. When git resolves neither directory - a non-repository fixture, a
detached copy - nothing proves this is the primary, so the run proceeds.

bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID
into the pane shell on the same pre-launch channel as GOTMPDIR, and the name
joins the sanitized launch environment allowlist so an isolated launch keeps
it.

* no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT

* no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal

---------

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): bound stale alarms for backlog captain holds (#3842)

* fix(bin): bound a stale alarm with the backlog hold, not only the status line

A legitimate wait has two records and the stale alarm reads only one.
`status_is_paused_or_captain_held` takes a status line, so it sees a wait the
worker declared. It cannot see the wait firstmate records when it hands work to
the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and
leaves the status log alone, so a delivered task keeps `done: PR ...` as its
last line for the whole time the captain is deciding.

Both stale branches were blind to it, and each churned a new pane hash back into
its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale
branch, while a held task whose last line is `working:` reaches
`surface_nonterminal_stale` and fails its declared-wait test.

Consult that second record where the watcher is about to alarm, through
`bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and
bound the alarm on the shared `.paused-resurfaced-<key>` marker and
`PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first
sight still alarms, the window's end alarms once more, and a held crew that goes
genuinely silent still escalates through the wedge timer.

Only an established open captain call bounds anything: an unreadable backlog, an
absent or incompatible tasks-axi, a row this home does not carry, and every task
with no hold keep alarming exactly as before. The backlog hold is deliberately
not recorded as a declared pause, because the loop-top reconciliation and
`pause_state_class` both read the status line and would clear a flag that line
does not support.

Extends the fix in #3443, which closed the forms of this loop that the status
line itself can express.

* fix(bin): identify the captain call a stale alarm is bounded by

Three gaps in the bound added by the previous commit, all in how the throttle is
scoped and where the backlog is consulted.

The scope carried only the status-log signature. A task can be held, answered
with `--release`, and re-held as a genuinely different captain call without any
status append, so the second call inherited the first one's marker and its first
sight was absorbed - the one thing this bound must never do. The task id is not
the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the
call's own lifecycle - its hold-set stamp and the number of recorded answers -
on an exit 0 and only then, leaving the silent predicate every existing caller
reads unchanged. The throttle scope now carries that identity.

The terminal path recorded the throttle before publishing the durable wake. A
failed append exits the watcher with nothing queued, and the next sighting then
read that fresh marker and absorbed the retry, turning a delayed alarm into a
lost one. Recording moves behind the append, as the non-terminal path already
had it, and the comment claiming the marker could not outlive its wake is gone
because it was false.

The backlog was consulted only on a new terminal pane hash. A captain call can
open after a hash was absorbed as provably working, changing neither the pane nor
the status log, so nothing re-read the backlog and the wedge timer kept firing
possible-wedge alarms through a legitimate wait. That timer now consults the call
at its own alarm boundary and takes the same bounded cadence - and only at that
boundary, so an ordinary repeat poll under the bound stays the local-only read it
was.

Regression coverage for each, all driving churn through one watcher process
rather than relaunching per pane change: relaunch cost dominated the earlier
shape, and an absorbing watcher stays in its poll loop across churn in production
anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer
with no open captain call still escalates as a possible wedge.

* fix(review): Compose stale throttles with captain-call lifecycle identity

* fix(review): Preserve bounded same-hash captain-call resurfacing

* revert(bin): narrow the captain-hold stale bound to its observed defect

Lifts the lifecycle-identity and cadence-ownership work back out, leaving the
change at the shape that matches the defect actually observed: the stale alarm
did not consult the backlog captain hold, on either stale branch.

Reviewing the wider version surfaced a series of adjacent gaps in the watcher's
alarm state machine - a call opening after the first alarm, marker invalidation
at the hold lifecycle boundary, and which deadline a terminal timer represents.
They are real, but fixing them turns a small extension into a state-machine
change to the alarm path, which is a different review on a subsystem that is
being actively reworked. They are named as known limitations rather than carried
here, and none of them is load-bearing for what remains: the bound does strictly
less than the reverted version, leaves the wedge path escalating on
STALE_ESCALATE_SECS exactly as before, and introduces no silence that the
existing terminal-alarm path did not already have.

Kept from the reverted work is the record-after-append ordering, because that is
a defect in the code being shipped rather than an adjacent one: recording the
cadence marker before publishing the durable wake let a failed append lose an
alarm outright instead of delaying it.

History is preserved: the earlier commits stay on the branch and this removal
sits on top of them.

* fix(review): Document secondmate captain-hold scope boundary

* fix(document): Document captain-hold stale alarm scope

* fix(bin): bind the stale throttle to the captain call, not the status log

The throttle this change introduces was scoped to the task's status-log
signature. Answering a call with `--release` and holding the task again creates a
genuinely different captain call without necessarily appending to that log, so
the second call inherited the first one's marker and its first sight was
absorbed.

That is the one alarm this bound must never swallow. A delivery announced twice
is noise; a decision waiting on the captain that is never surfaced is invisible,
because nobody asks for what they do not know to ask for.

Measured rather than assumed, on the same fixture - a delivered task held for the
captain, released, and re-held with no status append, driven through bin/fm-watch.sh:

  base c499f84    call-1 first=ALARM  call-1 churn=ALARM     new call first sight=ALARM
  before this fix call-1 first=ALARM  call-1 churn=absorbed  new call first sight=absorbed
  after           call-1 first=ALARM  call-1 churn=absorbed  new call first sight=ALARM

Base never suppresses the new call, so the suppression came from this change and
closing it completes the fix rather than widening it.

`bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle -
its hold-set stamp and count of recorded answers - on an exit 0 and only then, so
the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope
carries that identity beside the status signature.

The sibling case was measured too and is NOT included: on the status-declared
path, where the last line is `captain-held:`, base already absorbs a re-held
call's first sight. That behaviour predates this change and stays documented as a
known limitation rather than repaired here.

* fix(document): Document captain-call throttle lifecycle scope

* fix(ci): isolate the Herdr restart fixtures from a claimed worktree

The Herdr behaviour test intermittently reused a local worktree still claimed
by an earlier fixture after a restart.
The restart scenarios now use an isolated Treehouse project.

The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as
well.

* feat(pi): resolve extension-registered providers in the supervision branch (#3871)

* Let the supervision branch resolve extension-registered providers

The isolated branch ModelRuntime cannot see providers an extension
registered into main's runtime at run time, so a pin on pi-devin-auth's
devin/swe-1-7 (or an unpinned branch following a main session on devin)
failed with "unavailable to the isolated branch runtime".

Capture main's ModelRegistry alongside mainModel and copy each
extension-registered provider config into the branch runtime at
model-resolution time. The config carries the provider's own streamSimple
and oauth wiring by reference, so the custom gRPC transport reaches the
branch unchanged instead of being reimplemented. The /supervision-model
picker uses the same copy so those models are offered.

Update configuration.md and pi-supervision-branch.md, which previously
stated extension-registered providers were not offered.

* no-mistakes(document): docs: own devin provider carve-out in branch architecture doc

* no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered

* ci: retrigger flaky Herdr/serial-1 lanes

* no-mistakes(document): docs already cover branch extension-provider copy

---------

Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: 3264studios <3264studios@gmail.com>
Co-authored-by: Talon Stark <talonstark@gmail.com>
Co-authored-by: Jon Roosevelt <rooseveltadvisors@gmail.com>
Co-authored-by: kerry morrison <141675980+kmorebetter@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Daniel Kuykendall IV <danielkuykendall23@gmail.com>
Co-authored-by: Cristian Rosescu <crosescu@gmail.com>
Co-authored-by: test <test@example.invalid>
Co-authored-by: MortenGad <55603022+MortenGad@users.noreply.github.com>
Co-authored-by: Morten Gad <mogad@itm8.com>
Co-authored-by: Gyute <9tempo@gmail.com>
Co-authored-by: Mickaël Rémond <mremond@process-one.net>
Co-authored-by: Pedro Guimarães <pedroguim@pm.me>
Valentino-Sole added a commit to Valentino-Sole/firstmate that referenced this pull request Sep 8, 2026
* fix: start a fresh supervision branch for every main session (kunchenguid#3600)

* fix(pi): start a new supervision branch conversation per main session

The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.

The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.

The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.

* no-mistakes(document): Document fresh Pi supervision conversations

* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks

* feat: restart second mates after instruction updates (kunchenguid#3614)

* feat(update): restart second mates whose instructions changed

/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.

An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.

Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.

fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.

Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.

* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting

* no-mistakes(review): Parallelize relaunches and classify replacement incarnations

* no-mistakes(review): Gate restart actions on live agent state

* no-mistakes(review): Handle failed restart workers without hanging

* no-mistakes(review): Nudge legacy remotes and preserve persist recovery

* no-mistakes(review): Document one-time secondmate restart rollout

* no-mistakes(review): Honor arrived replies and refresh remote profiles

* no-mistakes(review): Revert remote parent profile reconciliation

* no-mistakes(review): Reset remote profile defaults and honor published results

* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates

* no-mistakes(document): Document second-mate restart update flow

* no-mistakes(lint): Fix ShellCheck warnings in restart scripts

* perf: accelerate local validation with bounded concurrency (kunchenguid#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation

* fix: copy PR URLs from durable records (kunchenguid#3648)

* fix: copy PR URLs from records or abstain, never assemble them

Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.

Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:

- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
  or abstain" section requires a URL to be copied verbatim from a durable
  record (the done: PR <url> status line, pr= metadata, or the backlog note),
  forbids assembling owner, repository, host, or number from memory, and has
  the branch report only the identifier it actually holds when no record names
  the URL yet, leaving the PR check unarmed until the worker's ready line
  arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for
  main in place of the bare full-URL mandate.

- Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full
  https:// URL wherever a PR is mentioned - status line, terminal, or summary -
  never a bare "PR 108", so the link is in view as early as the number is.

- bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that
  the task's own done lines contradict, printing both spellings; a log naming
  no URL still records the argument as before. fm_pr_status_ready_urls in
  bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches
  bin/fm-pr-merge.sh, so nothing merges under a contradicted URL.

Tests cover the offline refusal with zero side effects, the recorded spelling
being accepted, markdown-wrapped and punctuated URLs, working lines not
counting, the merge wrapper propagation, a self-hosted merge request with no
forge call, the prompt carrying the rule, and the brief carrying the worker
rule.

* no-mistakes(review): Remove stale PR URL enforcement

* no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`

* fix(bin): disable Claude feedback drafts for fleet launches (kunchenguid#3661)

* fix(bin): disable Claude's feedback-draft flow for fleet-launched agents

Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched
Claude crewmate and secondmate, so /bug and /feedback never queue or
submit a bug report on the captain's behalf. feedbackDrafts is the
documented settings key (Claude Code changelog 2.1.247); the
per-launch CLI flag never touches the captain's global settings.json.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts

* no-mistakes(document): Fix Claude feedback documentation formatting

* fix(bin): layer both feedback-draft controls for defense in depth

The prior --settings-only fix can be overridden by a managed Claude
settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0
alongside --settings '{"feedbackDrafts":"off"}': either control alone
disables the SendFeedback tool, so a managed override of one still
leaves the other in force.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(document): Document Claude feedback-draft suppression ownership

* feat(tests): run three more validation families concurrently (kunchenguid#3662)

* perf(tests): admit three more families to concurrent validation

The three families that `docs/fm-test-isolation-proof.md` recorded as refused
were not refused for concurrency. Each blocker was a test that decided a
property by wall clock, or a script filed where it cannot run. Fixing those
three things admits all three families and recovers 28.6 minutes of local
validation with no assertion removed or weakened.

- `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the
  handoff, sleeping a fixed second, then delegating the move to the real
  binary. Nothing ever killed the fake, so on a host slow enough for the case's
  next assertions to take longer than a second, the orphan woke and completed
  the very move the case requires left undone, and recovery then failed with
  `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs
  during the injected crash showed exactly that, the item moving one second
  after the crash. All four crash injections in the file now go through a new
  `fm_fake_crash_injector` shim that signals the target and returns only once
  it is observably gone, and the pre-move fake never delegates the move at all.

- `tests/fm-session-start.test.sh` proved the startup digest does not block on
  a slow current-state read by timing the whole digest against a fixed
  eight-second sleep, which a loaded host exceeds without the property being
  violated. It now holds that read open until the case releases it and asserts,
  the moment the digest returns, that the read has not finished. A digest that
  waited would wait indefinitely rather than for an interval a slow host can
  out-run, so the assertion is stronger than the bound it replaces. Its scan
  budget moves to the maximum, because the old value left two seconds of margin
  over the fixed sleep and measured the host rather than the deadline that
  `tests/fm-inactive-reconcile.test.sh` owns.

- `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all,
  which put it in the portable serial lane, where Linux CI gate-skips it: that
  real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and
  the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its
  opt-in variable and moves to `live-harness-optin`.

The 28 remaining ungrouped scripts become an enumerated `standalone` family
instead of admitting `unclassified` itself. `unclassified` is the family map's
`*)` arm, so admitting it would silently grant concurrency to every test added
afterwards, which is exactly the population with no proof. A new test still
lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh`
covers that split behaviorally.

Each family passes two consecutive four-worker proofs with zero failures. On
the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap`
756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s
recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock
against 121 minutes of summed script time.

* no-mistakes(document): Refresh concurrent validation and shard documentation

* no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2

* feat: structure no-mistakes ask-user escalations (kunchenguid#3670)

* feat(brief): structure no-mistakes ask-user escalation as event + snapshot file

Crewmates escalating a no-mistakes ask-user gate now report one status
event naming every finding id plus a snapshot file holding the gate's
axi finding records verbatim (id, severity, file, line, description,
authority), using the same shape even for a single finding. The status
line never paraphrases. The format is defined once in fm-dod-lib.sh and
rendered into both the scout and ship rule 6 in fm-brief.sh, so a
promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets
the identical contract as a freshly-spawned no-mistakes ship worker.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei

* no-mistakes(review): Preserve ask-user escalation output contract

* no-mistakes(review): Align escalation format test expectation

* no-mistakes(review): Scope ask-user escalation instructions correctly

* no-mistakes(review): Remove ask-user from generic decision rules

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix(bin): require self-sufficient no-mistakes intent (kunchenguid#3671)

* fix(bin): require a self-sufficient no-mistakes intent

A no-mistakes worker's --intent is only as useful as the string it
passes. PR kunchenguid#3604 shipped with an intent that was only "do 1, 2, 3, 7
from the report": the real contract lived in a private scout report and
never reached --intent, so nobody holding that string plus the codebase
could have derived the specification.

This is pure instruction at the contract's one owner; no spawn-side or
promotion-side check is added.

- bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now
  states that the --intent string must be self-sufficient (the string
  plus the codebase reconstructs roughly the same specification) and
  tells the worker to write the substance of any report, decision, or
  PR the captain's intent refers to into --intent rather than the
  pointer, while Firstmate build instructions and the worker's own
  decisions still stay out. The spawn-time overlay points back at that
  rule so its "supersedes" wording cannot cancel it, and the header's
  owner statement carries the rule.
- AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to
  include the substance of referenced material when filling
  ## Captain's intent, and section 11 points at the owner of the rule.
- tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the
  rendered brief and launch contract carry the rule.

Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim

* no-mistakes(document): Replace incident-specific intent test commentary

* fix: accelerate local Bearings snapshot composition (kunchenguid#3499)

* Speed local fleet snapshot composition

* no-mistakes(review): Stabilize task inventory during concurrent snapshot composition

* no-mistakes(document): Document local snapshot observation concurrency

* no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks

* no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks

* no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks

* no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky

* no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks

* fix(snapshot): keep live observations generation-coherent

* no-mistakes(review): Keep secondmate observations generation-bound without copying reports

* no-mistakes(document): Document generation-coherent snapshot observations

* test(bearings): measure local read overlap instead of wall-clock budget

The large-local-snapshot regression asserted that a whole snapshot
composed in under five seconds. That bound measures how loaded the host
is, not whether the per-task reads actually overlap, so it failed
intermittently on a contended machine: one run in six on a box at load
16-20, landing exactly on the five second boundary.

Time a serialized run and a concurrent run of the same workload instead
and require the concurrent one to save at least two seconds. Both runs
pay the same composition overhead, so the difference isolates the
overlap this change delivers. Five one-second reads serialize into five
seconds and overlap into about one, and re-serializing the reads
collapses the saving to roughly zero, so the assertion still fails
loudly if the concurrency regresses.

Also bump the pinned Bearings test count to 48, since rebasing onto the
current default branch picked up its captain-hold test.

* no-mistakes(review): Restore JSON-derived decision flags

* no-mistakes(review): Unify status-derived snapshot observations

* no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass

* fix: prevent stale supervision wake loops (kunchenguid#3672)

* fix(bin): stop the supervision branch's stale-ack and ghost-report loops

Clean-slate implementation of the four authorized recommendations from the
supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal
form, superseding PR kunchenguid#3604:

- fm_branch_report refuses a task the wake being handled never named. The
  extension fixes the reportable task set from the eligible rows before each
  prompt (signal and stale rows resolve to their tasks, a heartbeat allows any
  task with a live record, fleet is always allowed), so a report typed from
  memory about a task whose records teardown already removed is never stored
  or delivered.
- An acknowledgement that consumes nothing says "nothing was acknowledged
  through N" and prints the exact --ack-through / --recovery-generation
  command for the current presented wake, instead of "re-run the drain",
  which re-fed the same stale acknowledgement in a loop.
- bin/fm-guard.sh no longer tells the branch actor to drain queued wakes
  while it is handling them; it names the granted rows instead.
- Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and
  descendants; the index rebuild and the append-side index write both skip a
  task with neither a live record nor a status log, so the branch's report of
  a teardown it just performed is stored without recreating the index.

No new locking, no spawn-generation binding, and no retired-task refusal: the
branch can still report the outcome of a task it just tore down, and the
teardown test now proves that path end to end.

* fix(bin): narrow the branch report scope and guard silence to the minimal form

Apply the four review decisions on the clean-slate branch:

- A signal or stale prompt may report only the tasks its own rows resolve
  to; fleet is refused there too. A heartbeat review is not scoped by task
  at all, so the extension no longer tracks live task records and refuses
  nothing by task id during a fleet review.
- The outcome-index rebuild no longer skips retired tasks; the append-side
  skip alone keeps a torn-down task's index from being recreated.
- bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor
  instead of printing a replacement note.

* no-mistakes(document): Align supervision docs with scoped wake handling

* fix(bin): avoid fleet snapshot argument limits (kunchenguid#3677)

* Fix fleet snapshot large JSON transport

* no-mistakes(review): Captain: file-back fleet snapshot transport safely

* no-mistakes(review): Captain: file-back parent summary aggregation

* no-mistakes(ci): Rebased the PR's three commits onto f4d7875 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor

* fix(bin): attribute active runs with unfetched pipeline heads (kunchenguid#3681)

* fix(bin): recognize active pipeline fix rounds with unfetched run heads

A no-mistakes fix round advances the run head beyond the submitted head,
and the pipeline commits in its own checkout, so the task copy never
receives the new commit object. fm-crew-state's strict head rule rejected
the active row, the coarse runs-list scan skipped it and matched the
older failed row at the submitted head, and an active validation read as
failed (observed on model-routing-benchmark-hardening: active head
ac61c64 vs task copy at fb47636d).

fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns
runs-ledger attribution: the branch's newest row alone decides, and a
newest row whose head cannot resolve locally is recognized only as a
provable pipeline-owned continuation - active (running) and anchored by
the immediately older row for the same branch having ended at exactly
this worktree's HEAD. The reader keeps the axi TOON as full detail for
that proven same-branch run. Unanchored, ancestor-anchored, and terminal
unresolvable rows stay unattributed, so branch-name coincidence and other
tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact
prior semantics for teardown (verified by the full teardown suite).

Tests: reproduction regression for the unfetched active fix head (reads
working via full run-step detail), coarse-path continuation when axi
answers another branch, and negative controls for the unanchored active
row and the unresolvable terminal row with the historical fallback
preserved.

Ported onto upstream/main f4d7875, where kunchenguid#3194 independently added the
branch_sync custody exemption on the full axi-status path: both mechanisms
now coexist, each owning one surface (TOON custody on the full path, the
runs ledger on the coarse path). The port deletes the superseded coarse
scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers
(fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the
exemption comment's "the one exemption" phrasing now that a second
complementary exemption exists, and points the stale
FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree
(judge follow-up #1). The parent coarse-guard test's fixture is the
ledger-anchored continuation shape, so its expectation flips to the fixed
behavior (working via run-step, never the older failed row); a new
mismatched-anchor coarse negative control preserves that guard's original
no-anchor protection (pane answers, never the older row).

* no-mistakes(document): Clarify pipeline attribution documentation

* fix(bin): pre-register claude workspace trust at spawn time (kunchenguid#3663)

* fix(bin): pre-register claude workspace trust for task worktrees

A claude crewmate launched into a fresh task worktree met Claude Code's
interactive workspace-trust dialog before it ever read its brief, and firstmate
could not answer it: the key plane carries only Enter, Escape, and C-c with no
arrow navigation, and the dialog's selection starts on "No, exit", so the
documented Enter recipe ended the session instead of accepting it. Two workers
wedged this way and were unblocked only by hand-seeding the trust store per
path.

--dangerously-skip-permissions does not cover that gate. `claude --help`
records the dialog as skipped only in non-interactive mode, through -p or a
non-TTY stdout, and a crewmate pane is interactive, so there is no launch flag
to reach for.

fm-spawn now pre-registers the worktree through bin/fm-claude-trust.sh in the
existing claude branch, before the project settings that the same gate would
otherwise block, and refuses the spawn when that write fails rather than
launching a worker that would wedge.

The scope test is the safety property and is structural rather than a path
policy: the path must be a linked git worktree, sharing the spawning project's
common dir, whose top level is exactly the resolved argument. Git is the ground
truth, so the argument is never trusted on its own word, and a primary
checkout, an unrelated repo, a worktree subdirectory, a plain directory, and a
home directory are each refused rather than warned about or skipped. A
treehouse or orca path prefix was deliberately avoided because treehouse's root
is configurable, which would make a prefix both wrong and a new policy surface.
One structural test covers both worktree providers.

tests/fm-claude-trust.test.sh pins both halves, including a case where HOME is
itself a valid linked worktree so the home guard is proven load-bearing rather
than passing vacuously, plus the spawn-level proof that a claude spawn trusts
its worktree and launches with the brief pointed at the same store.

The adapter reference no longer tells a firstmate to press Enter on that
dialog, and the shared trust reference now names every harness surface: which
harnesses gate, which suppress at launch, which dodge the gate, which now
pre-registers, and that a claude secondmate is excluded by design.

The spawn fixture runs each spawn against a throwaway HOME so the suite cannot
write the developer's real store, isolating through HOME rather than
CLAUDE_CONFIG_DIR because the spawn forwards a set CLAUDE_CONFIG_DIR onto the
launch command that launch-shape assertions read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v

* fix(bin): create the staged trust store exclusively

The staged store was written to a predictable pid-based path with a plain
write, which follows a symlink. Where the Claude config directory is writable
by another local account, that account could pre-create the path as a symlink
and redirect the write into another file the launching user owns.

The staged name now carries random bytes and is created with an exclusive
"wx" open, so an existing path is refused outright instead of followed. The
happy-path test also asserts no staged store survives the rename.

The durability comment now states the residual window plainly: the readback
proves the entry landed, not that it survives, because a vendor session that
rewrites the whole store afterwards can still drop it and no lock closes that
window when the writer is Claude itself. The worker then meets the dialog and
stalls, which reaches firstmate as the ordinary stale wake rather than as
silent success.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v

* no-mistakes(review): neutralise CDPATH in claude trust scope guard

* no-mistakes(review): sandbox HOME in spawn tests, drop out-of-scope artifacts

* no-mistakes(review): refuse unresolvable git dir, compact store, fix secondmate doc

* no-mistakes(review): clear git env overrides, resolve symlinked store target

* no-mistakes(review): degrade without node, fix Pi gate claim, record trust proof

* no-mistakes(review): refuse without node, pin CLAUDE_CONFIG_DIR in spawn tests

* no-mistakes(review): refuse relative config dir and concurrent store modification

* no-mistakes(review): correct orca worktree claim, clean staged store on failure

* no-mistakes(review): restore pretty-printed store, correct trust dialog docs

* no-mistakes(review): arm trust gate before busy state to avoid orphans

* no-mistakes(document): record claude trust pre-registration in its owner docs

* no-mistakes(document): note orca limit for claude trust pre-registration

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-spawn.sh by moving the Claude trust gate earlier rather than adding cleanup machinery. Diagnosis: Greptile reported that when Claude trust registration fails on tmux/Zellij/cmux/non-projected Herdr, the exit runs after the backend endpoint and /tmp/fm-<id> were created, and the abort trap cleans neither. The endpoint half is pre-existing, deliberate architecture — the two refusals immediately above the gate (the 60s `treehouse get` timeout at fm-spawn.sh:2550 and `validate_spawn_worktree` at :2487) also exit with the endpoint live and direct the operator with "inspect window $T"; spawn_abort_cleanup only reclaims orca endpoints (already covered via ORCA_ABORT_CLEANUP) and herdr projections. The temp-root half was genuinely introduced by this PR: the gate was placed beside the busy-state arm, ~30 lines after `mkdir -p "$TASK_TMP/gotmp"`, and fm-teardown can only find that root through `tasktmp=` in a meta record a refused spawn never publishes. Root-cause fix (smallest correct change, no new subsystem): - bin/fm-spawn.sh — moved the `claude*` trust gate from inside the busy-arm block up to the first point $WT is known, immediately after the `freshen_spawn_worktree_base` block and before TASK_TMP creation, the STATE setup, and the relaunch `clear_relaunch_harness_wiring` retirement. A refusal now leaves no temp root, no retired relaunch wiring, and no busy record; only the endpoint remains, in the same class as the two refusals just above it. - bin/fm-spawn.sh — the refusal message now ends with "inspect window $T", matching the existing convention so control/teardown can identify the endpoint. $T is set for every backend on the non-secondmate path. - bin/fm-spawn.sh:196 — header note corrected from "before any state is armed" to "before any per-task state exists". - tests/fm-claude-trust.test.sh — the existing refused-spawn test's own comment claimed "before any task state exists" but only asserted busy state. Renamed to test_refused_spawn_leaves_no_task_state and added an assertion that /tmp/fm-<id> is absent, with the task id suffixed by the test process pid so the assertion reads only this run's path (a stale /tmp/fm-refusedspawn from the fixed-id version was in fact present on this box). No assertions on implementation source bytes. Verification run locally: - The new assertion fails against the pre-fix bin/fm-spawn.sh ("not ok - a refused spawn stranded a temp root no teardown can find") and passes after — a real before/after regression proof. - tests/fm-claude-trust.test.sh: 20/20 ok. - tests/fm-backend.test.sh, fm-backend-orca, fm-control-relaunch, fm-spawn-dispatch-profile, fm-trace-context-spawn, fm-gotmp: all pass. - tests/fm-backlog-atomicity.test.sh: rc=0, 79 assertions ok. - bin/fm-lint.sh (repo's single lint owner, pinned ShellCheck 0.11.0 + actionlint 1.7.12): clean. - No /tmp/fm-refusedspawn* leftovers after the runs. Scope respected: no trust subsystem, no policy layer, no config surface, no endpoint-cleanup mechanism added; the change is an ordering move plus one error-message clause and the test that pins it. Adapter references and docs made no ordering claim, so none needed updating. Changes are left uncommitted in the worktree for the outer executor

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix: restart every live second mate after updates (kunchenguid#3690)

* feat(update): restart every live second mate after a successful update

/updatefirstmate only restarted a second mate when that pass advanced its
AGENTS.md or .agents/skills. An already-current home was skipped entirely, a
bin/-only advance was steered instead, and a remote host that could not report
its instruction diff was downgraded to a re-read. A running agent also freezes
its launch-time wiring - turn-end hooks, harness flags, per-harness feature
switches - and none of that is derivable from a file diff, so an unchanged
tracked surface is not evidence the agent is already on the current behavior.

Restart is now unconditional on a successful update of that home. Every live
second mate the pass leaves on the target commit is restarted, whether it
advanced or was already there.

The safety contract is unchanged: open records are persisted before the agent is
replaced, nothing is forced, stashed, or discarded, a home the pass had to skip
is not restarted at all, and a mate whose runtime cannot prove a restart keeps
the honest re-read path and is never reported as reloaded.

bin/fm-ff-lib.sh gains a settled-state hook that fires for a home left at the
base whether it advanced or was already there, and never for a skipped one; the
instruction-gated hook the session-start convergence sweep uses is untouched.

Regressions: fm-update pins the already-current mate into the restart set and
the unprovable one into the nudge set, and fm-secondmate-restart drives both
real commands end to end - an already-current home is named, persisted, and
genuinely replaced with its checkout untouched, while the unprovable one keeps
its running agent.

* no-mistakes(document): Document unconditional secondmate restarts

* fix(bin): close pending-reply decisions via resolve-key (kunchenguid#3696)

* fix(bin): close reserved pending-reply keys via fm-send --resolve-key

fm-send wrote answered: notes that the reserved-key fold ignores, so
operator closes exited 0 while OPEN DECISIONS kept the decision open.
Speak the owning library's close vocabulary on that path, and refuse
when a reserved close cannot take effect.

* no-mistakes(review): Safely quote manual decision-close recovery commands

* no-mistakes(review): Reject unclosable overlong decision keys before sending

* no-mistakes(review): Remove contract suffix from open decisions hint

* no-mistakes(document): Document resolve-key line-cap refusal

* fix(bin): prevent false missed-reply escalations (kunchenguid#3697)

* fix(bin): stop false missed-reply escalations for same-basename self-home answers

A healthy secondmate that wrote corr= to its own state/<id>.status never matched the parent channel, so recovery confirmed and the record escalated as pending-reply-missed. Make the report helper resolve the parent channel itself, skip parent-replies.status as wrong-home, put a readable sighting path on the missed line, and restatement-copy only that same-basename self-home file onto the parent channel.

* no-mistakes(review): Resolve late replies before recovery escalation

* no-mistakes(review): Tighten reply routing and regression coverage

* no-mistakes(review): Preserve reply paths and require explicit home

* no-mistakes(review): Encode wrong-home paths before persistence

* no-mistakes(document): Document corrected secondmate reply routing

* no-mistakes(lint): Fix pending-reply ShellCheck warnings

* feat: add verified Gemini crewmate runtime (kunchenguid#3695)

* feat(harness): verify gemini as a crewmate runtime adapter

Adds Gemini CLI as a fourth dispatch target alongside claude, codex, and
grok, scoped to crewmate and scout work only. Every axis was proven against
gemini-cli 0.58.0 rather than inferred; docs/verification/runtime-backends.md
carries the dated evidence and names what stayed unverified.

Busy state is semantic, not rendered: BeforeAgent opens a turn and AfterAgent
and SessionEnd close it. AfterAgent also fires on a manual interrupt, so a
cancelled turn closes its own record.

Three findings shaped the wiring rather than a config line:

- --skip-trust and GEMINI_CLI_TRUST_WORKSPACE=true are presented by the CLI
  as equivalents and are not. A controlled A/B showed --skip-trust leaves
  project configuration unloaded, so workspace skills never load.
- The worktree's .gemini/settings.json is the PROJECT's committed settings
  file, unlike claude's settings.local.json. Firstmate's hooks therefore go
  to a firstmate-owned state/<id>.gemini-settings.json reached through
  GEMINI_CLI_SYSTEM_SETTINGS_PATH, which also works untrusted and merges
  with a project's own hooks instead of replacing them.
- The shipped CLI is a node bundle whose live process reports comm as
  MainThread, so ancestry cannot see it. GEMINI_CLI=1 is load-bearing and is
  tested before an inherited CLAUDECODE, and pane liveness identifies gemini
  from the script argument through the new bin/fm-gemini-lib.sh.

Gemini is refused for secondmates: it has no primary supervision protocol.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L

* test: clear gemini's marker in launch and detection expectations

Every non-gemini launch now clears GEMINI_CLI the way it already clears
cursor's markers, so the two tests that pin the exact launch prefix are
updated to match. The harness-detection tests that scrub foreign markers
before probing ancestry scrub GEMINI_CLI too, so running the suite from
inside a gemini session cannot produce a false verdict.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L

* docs: classify the gemini harness reference

The documentation inventory is the single classification owner for maintained
prose surfaces, and every surface must appear in it exactly once. The new
harness reference is agent-runtime, matching its siblings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L

* no-mistakes(review): Narrow Gemini ancestry detection

* no-mistakes(review): Restrict Gemini hooks to canonical launches

* no-mistakes(document): Document Gemini adapter support boundaries

* no-mistakes(ci): Fixed Gemini process identity when interpreter or script paths contain whitespace. Tmux liveness now uses NUL-delimited /proc argv on Linux, with the existing flattened ps fallback elsewhere. Added a real-process regression test. Verified with the Gemini harness test suite, full fm-lint, ShellCheck, and git diff --check. The CI and Require no-mistakes runs were action_required/attestation outcomes rather than code failures

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(teardown): conclude parked runs advanced past task copy (kunchenguid#3704)

* conclude parked runs the pipeline advanced past the task copy

A no-mistakes fix round commits in the daemon's own gate-repo clone, so a
run parked at a gate can carry a head whose object the task copy never
received. Teardown's strict object-local identity rule then declined to
conclude the run, and cleanup left it parked forever holding a fleet slot
(observed 2026-09-03; the same masking condition PR 3681 fixed on the
read path, now closing the teardown half its scope boundary deferred).

task_status_is_own_parked_run now falls back - only when the reported
head resolves to no local object - to the one shared runs-ledger
attribution rule fm_nm_runs_status_for_worktree (bin/fm-nm-run-lib.sh),
whose anchored continuation proof binds the branch's newest active row
to this worktree's exact submitted head. Foreign branches, stale
history, terminal rows, ancestor-only anchors, diverged newer rows, and
ambiguous multi-row shapes all still refuse, and runs that are actively
running, fixing, or in CI remain untouched: only the parked-at-a-gate
determination ever reaches the abort. No sqlite access, no fetches into
another task copy, no custody changes, no duplicated matching logic.

* tighten the parked-run ledger fallback and pin both judge corrections

The teardown ledger fallback now authorizes concluding this task's parked
run only when the shared runs-ledger rule's proved answer is the explicitly
active word (running): a terminal newest row - even anchored at exactly the
worktree's head - is finished history and never an abort authorization.
The read path may classify the same owner's answer; teardown's abort must
never fire for a run that already ended.

Two bounded pre-validation corrections from the implementation review:
- a fetched-object counterfactual pins the strict-rule path: a pipeline fix
  head fetched into the task copy aborts through object-local identity
  alone, with an empty ledger and a proof the runs query never fired;
- a negative fixture pins the tightened boundary: an unresolvable reported
  head with a terminal newest same-branch row anchored at the worktree head
  engages the ledger fallback and still refuses, so the refusal is the
  terminal-word boundary and not an earlier guard.

* no-mistakes(review): Bind teardown ledger fallback to validated run heads

* no-mistakes(review): Restore validated advanced-head ledger continuation

* no-mistakes(review): Reject invalid ledger dates and terminal statuses

* no-mistakes(document): Document teardown ledger scan limit

* feat(bin): show requested vs effective model in Herdr agent view

Track spawn-config requested_model separately from runtime-verified
effective_model, probe Claude/Pi transcripts for exact API ids, push
compact display metadata to Herdr, and preserve verified models across
relaunch/compaction hooks without inferring aliases as truth.

* fix(bin): keep re-probing effective model after first exact reading

fm-model-sync.sh only probed for the runtime-verified effective model
while it was still pending/UNKNOWN, so a session that later switched
models (manual switch, provider fallback) kept displaying the first
verified model forever and never appended a fallback-history entry.
Probe unconditionally instead; fm_model_record_effective already
no-ops when the probed value is unchanged, so this stays cheap.

Addresses the Greptile P1 finding on PR kunchenguid#3705's fm-model-sync.sh.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm

* fix(bin): distinguish Cursor Grok, direct xAI Grok, and Anthropic Claude in the display

Kapitänskorrektur: harness alone conflated Cursor-hosted Grok models
(cursor-grok-4.6-*) and direct xAI Grok models (xai/grok-4.6) under
one generic label, and displayed Anthropic Claude without naming the
provider. Add fm_model_source_label, pattern-matched on the verified
exact model id, so the compact display always reads Cursor · Grok,
xAI · Grok, or Anthropic · Claude with the exact model id appended.
Falls back to the existing harness label for every other model. No
routing change: this only affects display strings.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm

* fix(bin): wire model-sync into the Pi extension's turn lifecycle

fm-model-sync.sh was only invoked from Claude's SessionStart/
UserPromptSubmit/Stop hooks; the Pi harness's own extension
(state/<id>.pi-ext.ts) never called it, so a Pi-hosted session (e.g.
a pi/xai-grok crewmate) never refreshed its effective model after the
first probe and Herdr kept showing the stale value with no
fallback-history entry. Call fm-model-sync.sh from the same
agent_start/turn_end boundaries Pi already uses for busy-state and the
turn-end notification touch.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm

* fix(bin): serialize fm-model-sync.sh's meta read-probe-write

Overlapping lifecycle events (Pi's agent_start/turn_end, Claude's
SessionStart/UserPromptSubmit/Stop) can invoke fm-model-sync.sh
concurrently for the same task. The unlocked read-probe-write let
interleaved runs revert a newer effective model, mismatch its
source, or duplicate a model-history entry. Serialize the critical
section through the same per-task meta lock fm-spawn.sh already uses
(fm_meta_lock_path + fm_lock_acquire_wait/fm_lock_release), released
before every exit path.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Co-authored-by: Arthur Haro <38157909+haroarthur@users.noreply.github.com>
Co-authored-by: Nicolas Payette <nicolas.payette@specira.ai>
Co-authored-by: Jon Roosevelt <rooseveltadvisors@gmail.com>
Co-authored-by: att430 <41454889+att430@users.noreply.github.com>
Co-authored-by: Valentino-Sole <171032438+Valentino-Sole@users.noreply.github.com>
lytv pushed a commit to lytv/mymate that referenced this pull request Sep 8, 2026
…id#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation
lytv pushed a commit to lytv/mymate that referenced this pull request Sep 8, 2026
* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR kunchenguid#3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR kunchenguid#823 added and PR kunchenguid#1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
Mauryanx added a commit to Mauryanx/firstmate that referenced this pull request Sep 9, 2026
* fix(bin): scope the worker role contract for ship and scout launches (#3797)

* fix(brief): scope Firstmate workers to their launch contract

* no-mistakes(document): Clarify supervisor scope and worker contract ownership

* no-mistakes(ci): Removed the heading-based bypass so every ship/scout launch receives the current worker-role contract. Added a regression that failed before the fix and passes afterward. Dispatch, brief, and delivery suites, focused ShellCheck, and git diff --check all passed

* no-mistakes(review): make launch overlay sole owner of worker role contract

* no-mistakes(review): narrow heading test dimension and fix publish error wording

* no-mistakes(review): gate role supersession, fix render guard, drop AGENTS twin

* no-mistakes(document): align architecture AGENTS.md scope and spawn launch-brief header

* fix(bin): stop ringing steering doorbells into dead panes (#3823)

* fix(bin): stop ringing steering doorbells into dead panes

The steering-inbox doorbell was a plain sentence plus Enter typed into a
worker's pane, and the watcher re-rang it on the assumption that a ring is
free. In a pane whose agent has exited that line is a shell command, and the
re-ring ladder kept typing it into a shell that can never acknowledge it.

- Prefix the doorbell with the shell no-op `: ` so a bare shell executes
  nothing while a live worker still reads the same self-describing line.
  `#` is not used because interactive zsh does not treat it as a comment by
  default and the claude harness binds it to memory mode.
- fm_task_inbox_ring skips the pane (return 3) when the backend positively
  classifies the agent as dead; missing, ambiguous, unreadable, and unverified
  endpoints still ring so a blind classifier never starves a live worker.
- The watcher caps the ladder for a dead pane: one stale wake for recovery,
  no ring, no ladder walk, and the durable record stays for
  stuck-crewmate-recovery. fm-send and the remote steer leg report the skip.

Tests cover the no-op in real shells, the dead/live/unclassifiable ring
verdicts, and the single-surfacing watcher path.

* no-mistakes(review): Quote doorbell paths against shell injection

* no-mistakes(review): Reject terminal-control paths before ringing

* no-mistakes(review): Document accepted partial doorbell delivery race

* no-mistakes(review): Skip unavailable endpoints before busy-state handling

* no-mistakes(test): Respect shell startup PATH in environment allowlist test

* no-mistakes(test): Fix doorbell test fixtures for endpoint liveness

* no-mistakes(test): Prioritize confirmed restarts and clean shell test syntax

* no-mistakes(document): Document dead and missing doorbell recovery

* no-mistakes(ci): Fixed the persistence-reply timeout race by rechecking for a correlated reply immediately before falling back to a nudge. Added a deterministic regression covering replies arriving between the preliminary resolution pass and timeout handling. Verified with the targeted restart suite, project lint, coverage guard, bash syntax checks, and diff checks

* fix(bin): wait out transient primary-checkout reads in the spawn worktree poll (#3834)

* fix(spawn): keep the worktree poll from adopting the repository primary

After `treehouse get` is sent, the worktree-discovery poll reads the pane's
foreground-process cwd. While treehouse is still fetching and checking a slot
out, the foreground process is treehouse itself and it reports the repository's
PRIMARY checkout as its cwd for several seconds. The poll accepted any path
that merely differed from the spawning project, so from a linked spawning home
- whose project is itself a worktree of that repository - it adopted the
primary, and the isolation guard then refused a launch whose slot treehouse
went on to create normally.

Screen every candidate with the isolation guard's own conditions, extracted as
spawn_worktree_isolated, so a read the guard would reject stays a transient the
poll keeps waiting through. The two-consecutive-reads rule and the guard as
final backstop are unchanged; a pane that never reaches an isolated worktree
still fails at the existing 60s deadline, now naming the last path it reported.

The already-settled timing assertion counted whole-spawn wall time against a
5s budget and failed on unmodified HEAD on slower machines; it now counts pane
reads, which is what "one confirming read, not an extra cycle" actually means.

* fix(spawn): say which path the worktree wait rejected, and why

Screening every discovery-poll candidate means a host that never reaches an
isolated worktree spends the whole 60s window before refusing. That wait is
deliberate - separating a transient from a terminal misconfiguration needs
machinery this path does not want - so the refusal explains itself instead:
the isolation check records why a candidate failed, and the deadline names the
last path seen together with that reason. Message and diagnostics only; the
poll's control flow is unchanged.

Two suites asserted the guard's wording on paths the poll now rejects rather
than adopts, so their refusal arrives from the deadline instead: realign
fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also
asserting the stated reason, and the second the metadata absence it was
missing) and the herdr projection e2e's forced non-worktree cwd.

* no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test

* no-mistakes(document): document spawn poll isolation screen in fm-spawn header

* test(spawn): make the non-git isolation case non-git anywhere

The refusal-reason assertion for a path outside any repository assumed TMPDIR
is not inside a git repository. Where it is, git walks up from the temporary
directory, finds that repository, and the spawn reports the subdirectory cause
instead - so the case passed or failed on a property of the host rather than on
the behaviour under test.

Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES,
which git documents as not chdir-ing up into a listed directory while looking
for a repository. Git never excludes the directory being searched, so the
ceiling is the parent of the path handed to the spawn.

The assertions pin which cause fired rather than the sentence that explains it,
leaving the operator wording free to improve.

* no-mistakes(document): point spawn poll comment at the isolation screen's comparison

* no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure

* fix: support stock macOS Bash 3.2 paths (#3732)

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): read orphaned green ci monitor and daemon-down failed record as not failed (#3846)

* fix(bin): read an orphaned green ci monitor as held-for-merge, not failed

A no-mistakes run held for a captain merge decision keeps its ci step
polling until merged or closed; when the shared daemon restarts under
that poll, the run is recorded failed although every substantive step
completed and GitHub reports the PR green. A monitor whose only
remaining job is to observe a human decision must not convert the
absence of that decision into a failure verdict.

fm-crew-state.sh now reclassifies a terminal failed run as done
(held-for-merge), surfacing the run's PR URL, when the steps table
shows every step completed except exactly ci failed and the ci log's
last recognized marker reads checks green. A genuinely red check, an
unreadable ci log, or a second failed step keeps the failure.

* no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed

* no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification

* fix(spawn): verify preserved backlog state after interrupted spawn delivery (#3852)

* fix(bin): read preserved spawn state back before the interrupted exit claims it

The deferred-signal exit path asserted the paired task record and
In-flight backlog state were preserved without reading either back,
exactly when a reader is least able to check (fm-yi4j evidence,
2026-09-05). The commit's exit status alone has been observed to agree
with a row that did not actually move.

The exit path now re-reads the record and the row under the same
per-task lock as the commit, repairs a row the commit believed it
moved, and phrases the error as exactly what was verified or attempted
- verified preserved, repaired and verified, or an explicit
preservation-could-not-be-verified with the reason and hand-closeout
instruction. Two behavior tests drive a lying tasks-axi start through a
real interrupted spawn and assert the printed claim and the real
backlog state agree.

* no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask

* no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean

* docs: correct runtime-backend maturity labels for Herdr (#3821)

* docs: correct stale tmux/herdr backend maturity claims

Herdr now has 21 test files, its own required CI job (tests-herdr)
that installs a pinned build and hard-fails on "skip: herdr not
found", while tmux has 3 test files and is only required as a
dependency of the portable-serial e2e lane. zellij, orca, and cmux
still have no CI lane at all. AGENTS.md and docs/herdr-backend.md
still called Herdr merely "experimental" alongside those three,
misleading every session and reader about actual coverage.

Update AGENTS.md's config/backend entry, the opening lines of
docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend
section of docs/configuration.md, and the matching claims in
docs/architecture.md, CONTRIBUTING.md, and README.md so they agree
and distinguish tmux (default), herdr (own required CI lane, largest
suite, Windows still spike-only), and zellij/orca/cmux (still
experimental, no CI lane). No behavior, selection order, or
dispatch logic changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj

* no-mistakes(review): docs: fix stale herdr label and CI-lane wording

* no-mistakes(document): docs: align tmux adapter label in scripts.md

* no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page

* no-mistakes(review): docs: drop windows claim, align contributing backend wording

* no-mistakes(review): docs: trim duplicated CI claim from herdr opening line

* no-mistakes(review): docs: drop unguarded largest-test-suite superlative

* no-mistakes(review): docs: restore tmux verified label and README experimental scope

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: restore intent-targeted no-mistakes validation (#3865)

* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR #823 added and PR #1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: verify Treehouse slot ownership before teardown (#3837)

* fix(bin): verify pool-slot ownership before returning a worktree slot

Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.

Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.

* no-mistakes(review): Protect slots across cloned Firstmate homes

* no-mistakes(test): Gate teardown locking on genuine Treehouse slots

* no-mistakes(test): Clarify pooled descendant slot gating

* no-mistakes(test): Synchronize watcher re-arm test on process exit

* no-mistakes(test): Wait for watcher cleanup before timeout escalation

* no-mistakes(document): Document pool-slot ownership safeguards

* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion

* fix(bin): resolve relative origins from repository root

* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout

* feat(bin): add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary (#3867)

* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary

Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.

Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test: prove the omp guard continuation through a guard spy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* fix(spawn): clear the gemini marker at the omp launch boundary

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): force the guard stage by freezing the watcher and clear lint findings

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): reap the live lab by path and record omp's rpc shutdown as a note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces

Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin

* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs

* no-mistakes(review): omp: pin config-model validation with a test, trim overlay

* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family

* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes

* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux

* no-mistakes(document): docs: add omp subagent-guard row, fix live test header

* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π|󰵗|pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)

* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|󰵗)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix(pi): invoke Bash helpers correctly on native Windows (#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>

* test(bin): pin teardown outcomes for squash-merged rebased branches (#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>

* fix(bin): keep supervision armed for registered custom checks (#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes #3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner

* fix(bin): resolve Treehouse locks for remote secondmate homes (#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership

* fix(bearings): repair board listening and decision reconciliation (#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable state - refuses
rather than acting.

Regression coverage fails without each fix, and pins every leak path: a bare
reconcile, a standalone close or note with no pending request, an any-channel
reconcile, an annotated selection from a freeform card, and a
generation-skewed authorization. An opt-in guard re-proves the lavish-axi
shapes and the reopen against the installed tool.

* fix(bin): quote the done comparison in the reconcile intake

shellcheck SC1010 reads the bare word as the loop keyword. The failed run
never reached its lint step, so this shipped in the recovered content.

* no-mistakes(review): Publish reconciled parent resolution before request retirement

* no-mistakes(review): Clarify committed cleanup and reconcile reservation scope

* no-mistakes(review): Preserve remote cards and legacy answer compatibility

* no-mistakes(test): Separate live claim release from stale reclamation

* no-mistakes(test): Allow terminal self-retirement during active capture

* no-mistakes(document): Document Bearings repair contracts

* no-mistakes(ci): Stabilized the failing Herdr presentation E2E by serializing test-harness Treehouse allocator calls, preventing concurrent recovery spawns from claiming the same pool slot while preserving Herdr concurrency coverage. Verified with the full E2E suite on Herdr 0.8.2, bash syntax checks, ShellCheck, and git diff checks

* fix(bin): allow pooled spawns without a git origin (#3885)

* fix(bin): skip pooled-worktree freshness fetch when no origin is configured

An origin-less local-only project has nothing remote to be stale
against, so fm-spawn's freshen_spawn_worktree_base refused to launch
crews for it. Detect a missing origin remote and skip the fetch
freshness gate entirely; an existing-but-unreachable origin keeps
refusing as before.

* no-mistakes(review): Preserve pool safety for absent and unusable origins

* no-mistakes(review): Refuse empty origin configurations during pooled spawn

* no-mistakes(review): Detect empty origin sections across config includes

* no-mistakes(review): Honor globbed includes when detecting origin configuration

* no-mistakes(review): Document conservative conditional include handling

* no-mistakes(review): Use Git-resolved config files for origin detection

* no-mistakes(review): Document included empty-origin detection boundary

* no-mistakes(document): Document originless pooled spawn behavior

* feat(tests): run live harness guards by default when available (#3889)

* feat(tests): run live harness guards by default where the harness is installed

The 24 live-harness guards each opened with their own env check, so on the
machine that has every harness - the one the product and its validation
actually run on - all of them skipped and passed. Fourteen had never been run
by the pipeline at all.

tests/lib.sh gains fm_live_gate as the single owner of that decision: a guard
that spends no model tokens runs wherever its tools are installed, a guard that
submits prompts stays opt-in, an absent tool is a named capability skip, and a
guard's own variable or FM_LIVE forces it on (turning an absent tool into a
failure) or off. Every live guard now opens with it, which also carries the
test-suite gate-refusal bypass into the guards that never sourced the shared
helpers and were therefore refused whenever a gate agent ran them.

bin/fm-test-run.sh records what a skip means: the family's expected class is
live-capability rather than a bare env opt-in, and each gate skip's reason is
logged and written to the timing artifact, so a lane can say which tool this
host could not exercise.

Only the token-free guards flip to default-on: composer-matrix, the harness
liveness drift guard, and the Herdr version floor. cursor-primary submits three
prompts, so it stays opt-in.

Running the drift guard unasked immediately found a real defect it existed to
catch: it resolved the harness through a generic `command -v cursor`, which on
a machine that also has the Cursor editor finds the editor launcher rather than
cursor-agent. That binary exits at once, leaving a bare shell in the pane and a
liveness-drift failure no classifier change could fix. It now asks
fm_cursor_resolve_binary first, the same verified owner fm-spawn uses.

CI installs the public Pi package in the portable serial lane and fails on its
skip token, so the Pi extension tests stop passing silently against a package
that is not there. No secret is added.

Verified on macOS 26.5.2 arm64: the drift guard runs with no variable set and
classifies 8 installed harnesses alive; the Herdr version-floor guard runs by
default and checks 4 real releases; every live guard refuses together under
FM_LIVE=0.

* fix(tests): keep the composer-matrix guard opt-in

Running it unasked is red on a healthy machine for reasons no code change here
removes: a harness that has not trusted this checkout sits on its own trust
dialog, which the guard treats as an unreadable composer and correctly fails.
The opencode 1.18.29 and grok 1.0.13 composer drift it also surfaced reproduces
identically on main and is filed as separate work.

So this token-free guard stays opt-in with the reason stated in its header, and
the coding guidelines record the narrow exception: a guard whose verdict
depends on host state that installing its tools does not establish may stay
opt-in, because one that is permanently red is one the fleet learns to ignore.
The other two token-free guards keep running by default.

* no-mistakes(review): Wire bearings guard and remove composer exception policy

* no-mistakes(review): Run Pi responsiveness guard by default

* no-mistakes(review): Gate AFK Pi Herdr through authoritative family sweep

* no-mistakes(review): Sanitize live gate test environments

* no-mistakes(document): Document default-on live guard behavior

* no-mistakes(ci): Fixed CI by installing the Pi package in portable-parallel-1, where fm-pi-primary-types.test.sh runs, and enforcing its package-missing gate skip there. Verified with fm-lint.sh, workflow actionlint, coverage partition checks, lane membership, and git diff checks

* no-mistakes(ci): Fixed CI’s Pi typecheck skip enforcement by giving npm, tsc, and Pi-package capability skips a shared prefix and configuring both relevant CI lanes to fail on that prefix. Verified missing tsc emits the expected skip, missing Pi package becomes a runner failure, and actionlint, ShellCheck, and git diff checks pass

* fix(bin): refuse test runs in the primary checkout when a task marker is set (#3891)

* fix(bin): refuse the behavior suite in the repository primary checkout

A task worker's isolated worktree placement is verified exactly once, when
its task starts, and nothing re-checks it afterwards. A worker that later
changes directory into the repository's primary checkout runs its Git
commands, and this branch-switching suite, against the one checkout every
linked worktree resolves against and every landing merges into. A run that
dies mid-suite can leave that checkout on a stray branch.

bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task
worker and the runner resolves to the primary checkout, every executing mode
exits non-zero before selecting a suite, with one line naming the primary
path and pointing at the assigned task worktree. The predicate is the one
bin/fm-spawn.sh already uses for launch placement: the working tree's own
git dir is the repository's common git dir, which separates the primary from
every linked worktree even when their top levels differ. A run with no
FM_TASK_ID set is unchanged, and so are the inspection modes, which execute
nothing. When git resolves neither directory - a non-repository fixture, a
detached copy - nothing proves this is the primary, so the run proceeds.

bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID
into the pane shell on the same pre-launch channel as GOTMPDIR, and the name
joins the sanitized launch environment allowlist so an isolated launch keeps
it.

* no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT

* no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal

---------

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): bound stale alarms for backlog captain holds (#3842)

* fix(bin): bound a stale alarm with the backlog hold, not only the status line

A legitimate wait has two records and the stale alarm reads only one.
`status_is_paused_or_captain_held` takes a status line, so it sees a wait the
worker declared. It cannot see the wait firstmate records when it hands work to
the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and
leaves the status log alone, so a delivered task keeps `done: PR ...` as its
last line for the whole time the captain is deciding.

Both stale branches were blind to it, and each churned a new pane hash back into
its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale
branch, while a held task whose last line is `working:` reaches
`surface_nonterminal_stale` and fails its declared-wait test.

Consult that second record where the watcher is about to alarm, through
`bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and
bound the alarm on the shared `.paused-resurfaced-<key>` marker and
`PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first
sight still alarms, the window's end alarms once more, and a held crew that goes
genuinely silent still escalates through the wedge timer.

Only an established open captain call bounds anything: an unreadable backlog, an
absent or incompatible tasks-axi, a row this home does not carry, and every task
with no hold keep alarming exactly as before. The backlog hold is deliberately
not recorded as a declared pause, because the loop-top reconciliation and
`pause_state_class` both read the status line and would clear a flag that line
does not support.

Extends the fix in #3443, which closed the forms of this loop that the status
line itself can express.

* fix(bin): identify the captain call a stale alarm is bounded by

Three gaps in the bound added by the previous commit, all in how the throttle is
scoped and where the backlog is consulted.

The scope carried only the status-log signature. A task can be held, answered
with `--release`, and re-held as a genuinely different captain call without any
status append, so the second call inherited the first one's marker and its first
sight was absorbed - the one thing this bound must never do. The task id is not
the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the
call's own lifecycle - its hold-set stamp and the number of recorded answers -
on an exit 0 and only then, leaving the silent predicate every existing caller
reads unchanged. The throttle scope now carries that identity.

The terminal path recorded the throttle before publishing the durable wake. A
failed append exits the watcher with nothing queued, and the next sighting then
read that fresh marker and absorbed the retry, turning a delayed alarm into a
lost one. Recording moves behind the append, as the non-terminal path already
had it, and the comment claiming the marker could not outlive its wake is gone
because it was false.

The backlog was consulted only on a new terminal pane hash. A captain call can
open after a hash was absorbed as provably working, changing neither the pane nor
the status log, so nothing re-read the backlog and the wedge timer kept firing
possible-wedge alarms through a legitimate wait. That timer now consults the call
at its own alarm boundary and takes the same bounded cadence - and only at that
boundary, so an ordinary repeat poll under the bound stays the local-only read it
was.

Regression coverage for each, all driving churn through one watcher process
rather than relaunching per pane change: relaunch cost dominated the earlier
shape, and an absorbing watcher stays in its poll loop across churn in production
anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer
with no open captain call still escalates as a possible wedge.

* fix(review): Compose stale throttles with captain-call lifecycle identity

* fix(review): Preserve bounded same-hash captain-call resurfacing

* revert(bin): narrow the captain-hold stale bound to its observed defect

Lifts the lifecycle-identity and cadence-ownership work back out, leaving the
change at the shape that matches the defect actually observed: the stale alarm
did not consult the backlog captain hold, on either stale branch.

Reviewing the wider version surfaced a series of adjacent gaps in the watcher's
alarm state machine - a call opening after the first alarm, marker invalidation
at the hold lifecycle boundary, and which deadline a terminal timer represents.
They are real, but fixing them turns a small extension into a state-machine
change to the alarm path, which is a different review on a subsystem that is
being actively reworked. They are named as known limitations rather than carried
here, and none of them is load-bearing for what remains: the bound does strictly
less than the reverted version, leaves the wedge path escalating on
STALE_ESCALATE_SECS exactly as before, and introduces no silence that the
existing terminal-alarm path did not already have.

Kept from the reverted work is the record-after-append ordering, because that is
a defect in the code being shipped rather than an adjacent one: recording the
cadence marker before publishing the durable wake let a failed append lose an
alarm outright instead of delaying it.

History is preserved: the earlier commits stay on the branch and this removal
sits on top of them.

* fix(review): Document secondmate captain-hold scope boundary

* fix(document): Document captain-hold stale alarm scope

* fix(bin): bind the stale throttle to the captain call, not the status log

The throttle this change introduces was scoped to the task's status-log
signature. Answering a call with `--release` and holding the task again creates a
genuinely different captain call without necessarily appending to that log, so
the second call inherited the first one's marker and its first sight was
absorbed.

That is the one alarm this bound must never swallow. A delivery announced twice
is noise; a decision waiting on the captain that is never surfaced is invisible,
because nobody asks for what they do not know to ask for.

Measured rather than assumed, on the same fixture - a delivered task held for the
captain, released, and re-held with no status append, driven through bin/fm-watch.sh:

  base c499f84    call-1 first=ALARM  call-1 churn=ALARM     new call first sight=ALARM
  before this fix call-1 first=ALARM  call-1 churn=absorbed  new call first sight=absorbed
  after           call-1 first=ALARM  call-1 churn=absorbed  new call first sight=ALARM

Base never suppresses the new call, so the suppression came from this change and
closing it completes the fix rather than widening it.

`bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle -
its hold-set stamp and count of recorded answers - on an exit 0 and only then, so
the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope
carries that identity beside the status signature.

The sibling case was measured too and is NOT included: on the status-declared
path, where the last line is `captain-held:`, base already absorbs a re-held
call's first sight. That behaviour predates this change and stays documented as a
known limitation rather than repaired here.

* fix(document): Document captain-call throttle lifecycle scope

* fix(ci): isolate the Herdr restart fixtures from a claimed worktree

The Herdr behaviour test intermittently reused a local worktree still claimed
by an earlier fixture after a restart.
The restart scenarios now use an isolated Treehouse project.

The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as
well.

* feat(pi): resolve extension-registered providers in the supervision branch (#3871)

* Let the supervision branch resolve extension-registered providers

The isolated branch ModelRuntime cannot see providers an extension
registered into main's runtime at run time, so a pin on pi-devin-auth's
devin/swe-1-7 (or an unpinned branch following a main session on devin)
failed with "unavailable to the isolated branch runtime".

Capture main's ModelRegistry alongside mainModel and copy each
extension-registered provider config into the branch runtime at
model-resolution time. The config carries the provider's own streamSimple
and oauth wiring by reference, so the custom gRPC transport reaches the
branch unchanged instead of being reimplemented. The /supervision-model
picker uses the same copy so those models are offered.

Update configuration.md and pi-supervision-branch.md, which previously
stated extension-registered providers were not offered.

* no-mistakes(document): docs: own devin provider carve-out in branch architecture doc

* no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered

* ci: retrigger flaky Herdr/serial-1 lanes

* no-mistakes(document): docs already cover branch extension-provider copy

* fix(bin): gate secondmate wake-loop stall alerts on real queue no-progress (#3943)

* fix(watch): detect stalled secondmate queue progress

* no-mistakes(review): gate secondmate stall on active turns and progress episodes

* no-mistakes(document): align secondmate wake-stall docs with progress-episode detector

* no-mistakes(document): clarify active-turn gate in wake-stall config docs

* no-mistakes(ci): Fixed the one real defect behind the failing checks. ROOT CAUSE (Greptile P1, real code defect in this PR): `secondmate_wake_stall_tick` in bin/fm-watch.sh reset the no-progress timer only when the oldest actionable queue sequence INCREASED (`[ "$seq" -gt "$observed_seq" ]`). When a secondmate is retired and reprovisioned under the same task ID, its fresh home's queue sequence restarts BELOW the recorded position, so the comparison is false, no reset happens, and the new queue inherits the retired generation's already-expired idle interval — emitting a false `secondmate wake-loop stalled` on its very first observation. That is precisely the false-alarm class the user intent requires this PR to remove. FIX (smallest, removal-first): bin/fm-watch.sh:749 now resets when the drain position MOVES AT ALL (`-ne` instead of `-gt`). Draining moves it up, reprovisioning moves it down; neither is a continued no-progress episode. The asymmetric `-gt` branch is removed rather than special-cased or hardened. Updated the function header comment plus the two doc sentences in docs/architecture.md and docs/configuration.md that stated the old advance-only semantics. REGRESSION TEST: added `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` to tests/fm-wake-queue.test.sh (registered in the invocation list). It drives the real watcher through retired generation (seq 9) -> reprovision (seq 3, later clock) -> freeze, asserting observable wake-queue output, no source-text inspection. VERIFICATION: - Fails before / passes after: with the fix reverted the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted`; with the fix it passes, and its third leg confirms the restarted generation still escalates on a genuine freeze (row=3 idle=2s), so the fix does not merely mute the alarm. - `bash tests/fm-wake-queue.test.sh`: exit 0, 38/38 pass, all five secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs (34152740610, 34152740604) both ended with conclusion `action_required` — workflow approval pending, not a test/build failure. Separately, `bin/fm-test-run.sh --check-coverage` exits 1 in this environment, but I confirmed by stashing my changes that it fails identically on the unmodified base tree (locale-related `comm: input is not in sorted order`); it is pre-existing and this change adds no new test file for the partition to account for. Changes are left uncommitted in the worktree

* no-mistakes(ci): Fixed the one real code defect behind the failing checks. ROOT CAUSE (Greptile P1, second round, on the head commit e1304e6): `secondmate_wake_stall_tick` in bin/fm-watch.sh identified the queue's drain position by the sequence number ALONE. The previous round changed the comparison from `-gt` to `-ne`, which handles a reprovisioned queue that restarts BELOW the recorded position, but not one that restarts ON it. A mate retired and reprovisioned under the same task id gets a fresh home whose wake-queue sequence counter restarts at 1 — and the retained parent progress marker very plausibly holds a low sequence too (a queue frozen on its first row records seq 1). Equal sequence ⇒ no reset ⇒ the brand-new queue inherits the retired generation's long-expired idle interval and emits a false `secondmate wake-loop stalled` on its very first observation. That is exactly the false-alarm class this PR exists to remove. FIX (smallest, removal-first): the file already defines the identity of a queue row once, as `row_key="$epoch-$seq"` (used for stall receipts, the stall marker, and the notify key). The progress marker's separate, weaker seq-only identity is removed: `row_key` is now computed once right after the row is parsed, stored in the progress marker, and compared with `!=`. Across generations the epoch differs (the new generation's rows are appended later), so no sequence collision can carry a stale interval; within a generation the key is stable exactly while the position does not move. bin/fm-wake-lib.sh's `fm_wake_secondmate_progress_marker_write` now takes `<oldest-row-key>` and validates it the same way the two neighbouring row-key writers do. Updated the function header comment and the two doc sentences (docs/architecture.md, docs/configuration.md) that described the old sequence-only semantics. REGRESSION TEST: `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` in tests/fm-wake-queue.test.sh now drives the reported case — the reprovisioned generation restarts on the SAME sequence 9 (epoch 200) that the retired generation recorded (epoch 100), at a later clock — and asserts observable watcher output only. Its third leg still confirms the restarted generation escalates on a genuine freeze (row=9 idle=2s), so the fix does not merely mute the alarm. Three seeded progress markers in the symlink, crash-window and prefix-receipt tests were updated to the epoch-sequence form. VERIFICATION: - Fails before / passes after: with bin/fm-watch.sh and bin/fm-wake-lib.sh reverted to HEAD and the new test in place, the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted` (exit 1); with the fix, `bash tests/fm-wake-queue.test.sh` exits 0 with 38/38 pass, all six secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). - I also started `bin/fm-test-run.sh tests/fm-watch-checkpoint.test.sh tests/fm-watch-triage.test.sh tests/fm-watch-recovery-loop.test.sh` as a blast-radius check; it was still running when this phase had to return, so its result is not included. No other suite references the stall detector or the progress marker (grep over tests/ for `wake-loop stall|SECONDMATE_WAKE_STALL|secondmate-wake-progress` matches only fm-wake-queue.test.sh), and the changed lib function has exactly one caller. CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs on the head commit (34154091191, 34154091236, 34154091945) all ended with conclusion `action_required` — pending workflow approval, not a test/…
tiago-peixoto added a commit to tiago-peixoto/firstmate that referenced this pull request Sep 10, 2026
* fix(bin): scope the worker role contract for ship and scout launches (#3797)

* fix(brief): scope Firstmate workers to their launch contract

* no-mistakes(document): Clarify supervisor scope and worker contract ownership

* no-mistakes(ci): Removed the heading-based bypass so every ship/scout launch receives the current worker-role contract. Added a regression that failed before the fix and passes afterward. Dispatch, brief, and delivery suites, focused ShellCheck, and git diff --check all passed

* no-mistakes(review): make launch overlay sole owner of worker role contract

* no-mistakes(review): narrow heading test dimension and fix publish error wording

* no-mistakes(review): gate role supersession, fix render guard, drop AGENTS twin

* no-mistakes(document): align architecture AGENTS.md scope and spawn launch-brief header

* fix(bin): stop ringing steering doorbells into dead panes (#3823)

* fix(bin): stop ringing steering doorbells into dead panes

The steering-inbox doorbell was a plain sentence plus Enter typed into a
worker's pane, and the watcher re-rang it on the assumption that a ring is
free. In a pane whose agent has exited that line is a shell command, and the
re-ring ladder kept typing it into a shell that can never acknowledge it.

- Prefix the doorbell with the shell no-op `: ` so a bare shell executes
  nothing while a live worker still reads the same self-describing line.
  `#` is not used because interactive zsh does not treat it as a comment by
  default and the claude harness binds it to memory mode.
- fm_task_inbox_ring skips the pane (return 3) when the backend positively
  classifies the agent as dead; missing, ambiguous, unreadable, and unverified
  endpoints still ring so a blind classifier never starves a live worker.
- The watcher caps the ladder for a dead pane: one stale wake for recovery,
  no ring, no ladder walk, and the durable record stays for
  stuck-crewmate-recovery. fm-send and the remote steer leg report the skip.

Tests cover the no-op in real shells, the dead/live/unclassifiable ring
verdicts, and the single-surfacing watcher path.

* no-mistakes(review): Quote doorbell paths against shell injection

* no-mistakes(review): Reject terminal-control paths before ringing

* no-mistakes(review): Document accepted partial doorbell delivery race

* no-mistakes(review): Skip unavailable endpoints before busy-state handling

* no-mistakes(test): Respect shell startup PATH in environment allowlist test

* no-mistakes(test): Fix doorbell test fixtures for endpoint liveness

* no-mistakes(test): Prioritize confirmed restarts and clean shell test syntax

* no-mistakes(document): Document dead and missing doorbell recovery

* no-mistakes(ci): Fixed the persistence-reply timeout race by rechecking for a correlated reply immediately before falling back to a nudge. Added a deterministic regression covering replies arriving between the preliminary resolution pass and timeout handling. Verified with the targeted restart suite, project lint, coverage guard, bash syntax checks, and diff checks

* fix(bin): wait out transient primary-checkout reads in the spawn worktree poll (#3834)

* fix(spawn): keep the worktree poll from adopting the repository primary

After `treehouse get` is sent, the worktree-discovery poll reads the pane's
foreground-process cwd. While treehouse is still fetching and checking a slot
out, the foreground process is treehouse itself and it reports the repository's
PRIMARY checkout as its cwd for several seconds. The poll accepted any path
that merely differed from the spawning project, so from a linked spawning home
- whose project is itself a worktree of that repository - it adopted the
primary, and the isolation guard then refused a launch whose slot treehouse
went on to create normally.

Screen every candidate with the isolation guard's own conditions, extracted as
spawn_worktree_isolated, so a read the guard would reject stays a transient the
poll keeps waiting through. The two-consecutive-reads rule and the guard as
final backstop are unchanged; a pane that never reaches an isolated worktree
still fails at the existing 60s deadline, now naming the last path it reported.

The already-settled timing assertion counted whole-spawn wall time against a
5s budget and failed on unmodified HEAD on slower machines; it now counts pane
reads, which is what "one confirming read, not an extra cycle" actually means.

* fix(spawn): say which path the worktree wait rejected, and why

Screening every discovery-poll candidate means a host that never reaches an
isolated worktree spends the whole 60s window before refusing. That wait is
deliberate - separating a transient from a terminal misconfiguration needs
machinery this path does not want - so the refusal explains itself instead:
the isolation check records why a candidate failed, and the deadline names the
last path seen together with that reason. Message and diagnostics only; the
poll's control flow is unchanged.

Two suites asserted the guard's wording on paths the poll now rejects rather
than adopts, so their refusal arrives from the deadline instead: realign
fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also
asserting the stated reason, and the second the metadata absence it was
missing) and the herdr projection e2e's forced non-worktree cwd.

* no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test

* no-mistakes(document): document spawn poll isolation screen in fm-spawn header

* test(spawn): make the non-git isolation case non-git anywhere

The refusal-reason assertion for a path outside any repository assumed TMPDIR
is not inside a git repository. Where it is, git walks up from the temporary
directory, finds that repository, and the spawn reports the subdirectory cause
instead - so the case passed or failed on a property of the host rather than on
the behaviour under test.

Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES,
which git documents as not chdir-ing up into a listed directory while looking
for a repository. Git never excludes the directory being searched, so the
ceiling is the parent of the path handed to the spawn.

The assertions pin which cause fired rather than the sentence that explains it,
leaving the operator wording free to improve.

* no-mistakes(document): point spawn poll comment at the isolation screen's comparison

* no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure

* fix: support stock macOS Bash 3.2 paths (#3732)

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): read orphaned green ci monitor and daemon-down failed record as not failed (#3846)

* fix(bin): read an orphaned green ci monitor as held-for-merge, not failed

A no-mistakes run held for a captain merge decision keeps its ci step
polling until merged or closed; when the shared daemon restarts under
that poll, the run is recorded failed although every substantive step
completed and GitHub reports the PR green. A monitor whose only
remaining job is to observe a human decision must not convert the
absence of that decision into a failure verdict.

fm-crew-state.sh now reclassifies a terminal failed run as done
(held-for-merge), surfacing the run's PR URL, when the steps table
shows every step completed except exactly ci failed and the ci log's
last recognized marker reads checks green. A genuinely red check, an
unreadable ci log, or a second failed step keeps the failure.

* no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed

* no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification

* fix(spawn): verify preserved backlog state after interrupted spawn delivery (#3852)

* fix(bin): read preserved spawn state back before the interrupted exit claims it

The deferred-signal exit path asserted the paired task record and
In-flight backlog state were preserved without reading either back,
exactly when a reader is least able to check (fm-yi4j evidence,
2026-09-05). The commit's exit status alone has been observed to agree
with a row that did not actually move.

The exit path now re-reads the record and the row under the same
per-task lock as the commit, repairs a row the commit believed it
moved, and phrases the error as exactly what was verified or attempted
- verified preserved, repaired and verified, or an explicit
preservation-could-not-be-verified with the reason and hand-closeout
instruction. Two behavior tests drive a lying tasks-axi start through a
real interrupted spawn and assert the printed claim and the real
backlog state agree.

* no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask

* no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean

* docs: correct runtime-backend maturity labels for Herdr (#3821)

* docs: correct stale tmux/herdr backend maturity claims

Herdr now has 21 test files, its own required CI job (tests-herdr)
that installs a pinned build and hard-fails on "skip: herdr not
found", while tmux has 3 test files and is only required as a
dependency of the portable-serial e2e lane. zellij, orca, and cmux
still have no CI lane at all. AGENTS.md and docs/herdr-backend.md
still called Herdr merely "experimental" alongside those three,
misleading every session and reader about actual coverage.

Update AGENTS.md's config/backend entry, the opening lines of
docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend
section of docs/configuration.md, and the matching claims in
docs/architecture.md, CONTRIBUTING.md, and README.md so they agree
and distinguish tmux (default), herdr (own required CI lane, largest
suite, Windows still spike-only), and zellij/orca/cmux (still
experimental, no CI lane). No behavior, selection order, or
dispatch logic changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj

* no-mistakes(review): docs: fix stale herdr label and CI-lane wording

* no-mistakes(document): docs: align tmux adapter label in scripts.md

* no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page

* no-mistakes(review): docs: drop windows claim, align contributing backend wording

* no-mistakes(review): docs: trim duplicated CI claim from herdr opening line

* no-mistakes(review): docs: drop unguarded largest-test-suite superlative

* no-mistakes(review): docs: restore tmux verified label and README experimental scope

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: restore intent-targeted no-mistakes validation (#3865)

* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR #823 added and PR #1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: verify Treehouse slot ownership before teardown (#3837)

* fix(bin): verify pool-slot ownership before returning a worktree slot

Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.

Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.

* no-mistakes(review): Protect slots across cloned Firstmate homes

* no-mistakes(test): Gate teardown locking on genuine Treehouse slots

* no-mistakes(test): Clarify pooled descendant slot gating

* no-mistakes(test): Synchronize watcher re-arm test on process exit

* no-mistakes(test): Wait for watcher cleanup before timeout escalation

* no-mistakes(document): Document pool-slot ownership safeguards

* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion

* fix(bin): resolve relative origins from repository root

* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout

* feat(bin): add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary (#3867)

* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary

Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.

Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test: prove the omp guard continuation through a guard spy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* fix(spawn): clear the gemini marker at the omp launch boundary

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): force the guard stage by freezing the watcher and clear lint findings

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): reap the live lab by path and record omp's rpc shutdown as a note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces

Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin

* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs

* no-mistakes(review): omp: pin config-model validation with a test, trim overlay

* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family

* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes

* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux

* no-mistakes(document): docs: add omp subagent-guard row, fix live test header

* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π|󰵗|pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)

* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|󰵗)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix(pi): invoke Bash helpers correctly on native Windows (#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>

* test(bin): pin teardown outcomes for squash-merged rebased branches (#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>

* fix(bin): keep supervision armed for registered custom checks (#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes #3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner

* fix(bin): resolve Treehouse locks for remote secondmate homes (#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership

* fix(bearings): repair board listening and decision reconciliation (#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable state - refuses
rather than acting.

Regression coverage fails without each fix, and pins every leak path: a bare
reconcile, a standalone close or note with no pending request, an any-channel
reconcile, an annotated selection from a freeform card, and a
generation-skewed authorization. An opt-in guard re-proves the lavish-axi
shapes and the reopen against the installed tool.

* fix(bin): quote the done comparison in the reconcile intake

shellcheck SC1010 reads the bare word as the loop keyword. The failed run
never reached its lint step, so this shipped in the recovered content.

* no-mistakes(review): Publish reconciled parent resolution before request retirement

* no-mistakes(review): Clarify committed cleanup and reconcile reservation scope

* no-mistakes(review): Preserve remote cards and legacy answer compatibility

* no-mistakes(test): Separate live claim release from stale reclamation

* no-mistakes(test): Allow terminal self-retirement during active capture

* no-mistakes(document): Document Bearings repair contracts

* no-mistakes(ci): Stabilized the failing Herdr presentation E2E by serializing test-harness Treehouse allocator calls, preventing concurrent recovery spawns from claiming the same pool slot while preserving Herdr concurrency coverage. Verified with the full E2E suite on Herdr 0.8.2, bash syntax checks, ShellCheck, and git diff checks

* fix(bin): allow pooled spawns without a git origin (#3885)

* fix(bin): skip pooled-worktree freshness fetch when no origin is configured

An origin-less local-only project has nothing remote to be stale
against, so fm-spawn's freshen_spawn_worktree_base refused to launch
crews for it. Detect a missing origin remote and skip the fetch
freshness gate entirely; an existing-but-unreachable origin keeps
refusing as before.

* no-mistakes(review): Preserve pool safety for absent and unusable origins

* no-mistakes(review): Refuse empty origin configurations during pooled spawn

* no-mistakes(review): Detect empty origin sections across config includes

* no-mistakes(review): Honor globbed includes when detecting origin configuration

* no-mistakes(review): Document conservative conditional include handling

* no-mistakes(review): Use Git-resolved config files for origin detection

* no-mistakes(review): Document included empty-origin detection boundary

* no-mistakes(document): Document originless pooled spawn behavior

* feat(tests): run live harness guards by default when available (#3889)

* feat(tests): run live harness guards by default where the harness is installed

The 24 live-harness guards each opened with their own env check, so on the
machine that has every harness - the one the product and its validation
actually run on - all of them skipped and passed. Fourteen had never been run
by the pipeline at all.

tests/lib.sh gains fm_live_gate as the single owner of that decision: a guard
that spends no model tokens runs wherever its tools are installed, a guard that
submits prompts stays opt-in, an absent tool is a named capability skip, and a
guard's own variable or FM_LIVE forces it on (turning an absent tool into a
failure) or off. Every live guard now opens with it, which also carries the
test-suite gate-refusal bypass into the guards that never sourced the shared
helpers and were therefore refused whenever a gate agent ran them.

bin/fm-test-run.sh records what a skip means: the family's expected class is
live-capability rather than a bare env opt-in, and each gate skip's reason is
logged and written to the timing artifact, so a lane can say which tool this
host could not exercise.

Only the token-free guards flip to default-on: composer-matrix, the harness
liveness drift guard, and the Herdr version floor. cursor-primary submits three
prompts, so it stays opt-in.

Running the drift guard unasked immediately found a real defect it existed to
catch: it resolved the harness through a generic `command -v cursor`, which on
a machine that also has the Cursor editor finds the editor launcher rather than
cursor-agent. That binary exits at once, leaving a bare shell in the pane and a
liveness-drift failure no classifier change could fix. It now asks
fm_cursor_resolve_binary first, the same verified owner fm-spawn uses.

CI installs the public Pi package in the portable serial lane and fails on its
skip token, so the Pi extension tests stop passing silently against a package
that is not there. No secret is added.

Verified on macOS 26.5.2 arm64: the drift guard runs with no variable set and
classifies 8 installed harnesses alive; the Herdr version-floor guard runs by
default and checks 4 real releases; every live guard refuses together under
FM_LIVE=0.

* fix(tests): keep the composer-matrix guard opt-in

Running it unasked is red on a healthy machine for reasons no code change here
removes: a harness that has not trusted this checkout sits on its own trust
dialog, which the guard treats as an unreadable composer and correctly fails.
The opencode 1.18.29 and grok 1.0.13 composer drift it also surfaced reproduces
identically on main and is filed as separate work.

So this token-free guard stays opt-in with the reason stated in its header, and
the coding guidelines record the narrow exception: a guard whose verdict
depends on host state that installing its tools does not establish may stay
opt-in, because one that is permanently red is one the fleet learns to ignore.
The other two token-free guards keep running by default.

* no-mistakes(review): Wire bearings guard and remove composer exception policy

* no-mistakes(review): Run Pi responsiveness guard by default

* no-mistakes(review): Gate AFK Pi Herdr through authoritative family sweep

* no-mistakes(review): Sanitize live gate test environments

* no-mistakes(document): Document default-on live guard behavior

* no-mistakes(ci): Fixed CI by installing the Pi package in portable-parallel-1, where fm-pi-primary-types.test.sh runs, and enforcing its package-missing gate skip there. Verified with fm-lint.sh, workflow actionlint, coverage partition checks, lane membership, and git diff checks

* no-mistakes(ci): Fixed CI’s Pi typecheck skip enforcement by giving npm, tsc, and Pi-package capability skips a shared prefix and configuring both relevant CI lanes to fail on that prefix. Verified missing tsc emits the expected skip, missing Pi package becomes a runner failure, and actionlint, ShellCheck, and git diff checks pass

* fix(bin): refuse test runs in the primary checkout when a task marker is set (#3891)

* fix(bin): refuse the behavior suite in the repository primary checkout

A task worker's isolated worktree placement is verified exactly once, when
its task starts, and nothing re-checks it afterwards. A worker that later
changes directory into the repository's primary checkout runs its Git
commands, and this branch-switching suite, against the one checkout every
linked worktree resolves against and every landing merges into. A run that
dies mid-suite can leave that checkout on a stray branch.

bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task
worker and the runner resolves to the primary checkout, every executing mode
exits non-zero before selecting a suite, with one line naming the primary
path and pointing at the assigned task worktree. The predicate is the one
bin/fm-spawn.sh already uses for launch placement: the working tree's own
git dir is the repository's common git dir, which separates the primary from
every linked worktree even when their top levels differ. A run with no
FM_TASK_ID set is unchanged, and so are the inspection modes, which execute
nothing. When git resolves neither directory - a non-repository fixture, a
detached copy - nothing proves this is the primary, so the run proceeds.

bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID
into the pane shell on the same pre-launch channel as GOTMPDIR, and the name
joins the sanitized launch environment allowlist so an isolated launch keeps
it.

* no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT

* no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal

---------

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): bound stale alarms for backlog captain holds (#3842)

* fix(bin): bound a stale alarm with the backlog hold, not only the status line

A legitimate wait has two records and the stale alarm reads only one.
`status_is_paused_or_captain_held` takes a status line, so it sees a wait the
worker declared. It cannot see the wait firstmate records when it hands work to
the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and
leaves the status log alone, so a delivered task keeps `done: PR ...` as its
last line for the whole time the captain is deciding.

Both stale branches were blind to it, and each churned a new pane hash back into
its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale
branch, while a held task whose last line is `working:` reaches
`surface_nonterminal_stale` and fails its declared-wait test.

Consult that second record where the watcher is about to alarm, through
`bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and
bound the alarm on the shared `.paused-resurfaced-<key>` marker and
`PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first
sight still alarms, the window's end alarms once more, and a held crew that goes
genuinely silent still escalates through the wedge timer.

Only an established open captain call bounds anything: an unreadable backlog, an
absent or incompatible tasks-axi, a row this home does not carry, and every task
with no hold keep alarming exactly as before. The backlog hold is deliberately
not recorded as a declared pause, because the loop-top reconciliation and
`pause_state_class` both read the status line and would clear a flag that line
does not support.

Extends the fix in #3443, which closed the forms of this loop that the status
line itself can express.

* fix(bin): identify the captain call a stale alarm is bounded by

Three gaps in the bound added by the previous commit, all in how the throttle is
scoped and where the backlog is consulted.

The scope carried only the status-log signature. A task can be held, answered
with `--release`, and re-held as a genuinely different captain call without any
status append, so the second call inherited the first one's marker and its first
sight was absorbed - the one thing this bound must never do. The task id is not
the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the
call's own lifecycle - its hold-set stamp and the number of recorded answers -
on an exit 0 and only then, leaving the silent predicate every existing caller
reads unchanged. The throttle scope now carries that identity.

The terminal path recorded the throttle before publishing the durable wake. A
failed append exits the watcher with nothing queued, and the next sighting then
read that fresh marker and absorbed the retry, turning a delayed alarm into a
lost one. Recording moves behind the append, as the non-terminal path already
had it, and the comment claiming the marker could not outlive its wake is gone
because it was false.

The backlog was consulted only on a new terminal pane hash. A captain call can
open after a hash was absorbed as provably working, changing neither the pane nor
the status log, so nothing re-read the backlog and the wedge timer kept firing
possible-wedge alarms through a legitimate wait. That timer now consults the call
at its own alarm boundary and takes the same bounded cadence - and only at that
boundary, so an ordinary repeat poll under the bound stays the local-only read it
was.

Regression coverage for each, all driving churn through one watcher process
rather than relaunching per pane change: relaunch cost dominated the earlier
shape, and an absorbing watcher stays in its poll loop across churn in production
anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer
with no open captain call still escalates as a possible wedge.

* fix(review): Compose stale throttles with captain-call lifecycle identity

* fix(review): Preserve bounded same-hash captain-call resurfacing

* revert(bin): narrow the captain-hold stale bound to its observed defect

Lifts the lifecycle-identity and cadence-ownership work back out, leaving the
change at the shape that matches the defect actually observed: the stale alarm
did not consult the backlog captain hold, on either stale branch.

Reviewing the wider version surfaced a series of adjacent gaps in the watcher's
alarm state machine - a call opening after the first alarm, marker invalidation
at the hold lifecycle boundary, and which deadline a terminal timer represents.
They are real, but fixing them turns a small extension into a state-machine
change to the alarm path, which is a different review on a subsystem that is
being actively reworked. They are named as known limitations rather than carried
here, and none of them is load-bearing for what remains: the bound does strictly
less than the reverted version, leaves the wedge path escalating on
STALE_ESCALATE_SECS exactly as before, and introduces no silence that the
existing terminal-alarm path did not already have.

Kept from the reverted work is the record-after-append ordering, because that is
a defect in the code being shipped rather than an adjacent one: recording the
cadence marker before publishing the durable wake let a failed append lose an
alarm outright instead of delaying it.

History is preserved: the earlier commits stay on the branch and this removal
sits on top of them.

* fix(review): Document secondmate captain-hold scope boundary

* fix(document): Document captain-hold stale alarm scope

* fix(bin): bind the stale throttle to the captain call, not the status log

The throttle this change introduces was scoped to the task's status-log
signature. Answering a call with `--release` and holding the task again creates a
genuinely different captain call without necessarily appending to that log, so
the second call inherited the first one's marker and its first sight was
absorbed.

That is the one alarm this bound must never swallow. A delivery announced twice
is noise; a decision waiting on the captain that is never surfaced is invisible,
because nobody asks for what they do not know to ask for.

Measured rather than assumed, on the same fixture - a delivered task held for the
captain, released, and re-held with no status append, driven through bin/fm-watch.sh:

  base c499f84    call-1 first=ALARM  call-1 churn=ALARM     new call first sight=ALARM
  before this fix call-1 first=ALARM  call-1 churn=absorbed  new call first sight=absorbed
  after           call-1 first=ALARM  call-1 churn=absorbed  new call first sight=ALARM

Base never suppresses the new call, so the suppression came from this change and
closing it completes the fix rather than widening it.

`bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle -
its hold-set stamp and count of recorded answers - on an exit 0 and only then, so
the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope
carries that identity beside the status signature.

The sibling case was measured too and is NOT included: on the status-declared
path, where the last line is `captain-held:`, base already absorbs a re-held
call's first sight. That behaviour predates this change and stays documented as a
known limitation rather than repaired here.

* fix(document): Document captain-call throttle lifecycle scope

* fix(ci): isolate the Herdr restart fixtures from a claimed worktree

The Herdr behaviour test intermittently reused a local worktree still claimed
by an earlier fixture after a restart.
The restart scenarios now use an isolated Treehouse project.

The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as
well.

* feat(pi): resolve extension-registered providers in the supervision branch (#3871)

* Let the supervision branch resolve extension-registered providers

The isolated branch ModelRuntime cannot see providers an extension
registered into main's runtime at run time, so a pin on pi-devin-auth's
devin/swe-1-7 (or an unpinned branch following a main session on devin)
failed with "unavailable to the isolated branch runtime".

Capture main's ModelRegistry alongside mainModel and copy each
extension-registered provider config into the branch runtime at
model-resolution time. The config carries the provider's own streamSimple
and oauth wiring by reference, so the custom gRPC transport reaches the
branch unchanged instead of being reimplemented. The /supervision-model
picker uses the same copy so those models are offered.

Update configuration.md and pi-supervision-branch.md, which previously
stated extension-registered providers were not offered.

* no-mistakes(document): docs: own devin provider carve-out in branch architecture doc

* no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered

* ci: retrigger flaky Herdr/serial-1 lanes

* no-mistakes(document): docs already cover branch extension-provider copy

* fix(bin): gate secondmate wake-loop stall alerts on real queue no-progress (#3943)

* fix(watch): detect stalled secondmate queue progress

* no-mistakes(review): gate secondmate stall on active turns and progress episodes

* no-mistakes(document): align secondmate wake-stall docs with progress-episode detector

* no-mistakes(document): clarify active-turn gate in wake-stall config docs

* no-mistakes(ci): Fixed the one real defect behind the failing checks. ROOT CAUSE (Greptile P1, real code defect in this PR): `secondmate_wake_stall_tick` in bin/fm-watch.sh reset the no-progress timer only when the oldest actionable queue sequence INCREASED (`[ "$seq" -gt "$observed_seq" ]`). When a secondmate is retired and reprovisioned under the same task ID, its fresh home's queue sequence restarts BELOW the recorded position, so the comparison is false, no reset happens, and the new queue inherits the retired generation's already-expired idle interval — emitting a false `secondmate wake-loop stalled` on its very first observation. That is precisely the false-alarm class the user intent requires this PR to remove. FIX (smallest, removal-first): bin/fm-watch.sh:749 now resets when the drain position MOVES AT ALL (`-ne` instead of `-gt`). Draining moves it up, reprovisioning moves it down; neither is a continued no-progress episode. The asymmetric `-gt` branch is removed rather than special-cased or hardened. Updated the function header comment plus the two doc sentences in docs/architecture.md and docs/configuration.md that stated the old advance-only semantics. REGRESSION TEST: added `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` to tests/fm-wake-queue.test.sh (registered in the invocation list). It drives the real watcher through retired generation (seq 9) -> reprovision (seq 3, later clock) -> freeze, asserting observable wake-queue output, no source-text inspection. VERIFICATION: - Fails before / passes after: with the fix reverted the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted`; with the fix it passes, and its third leg confirms the restarted generation still escalates on a genuine freeze (row=3 idle=2s), so the fix does not merely mute the alarm. - `bash tests/fm-wake-queue.test.sh`: exit 0, 38/38 pass, all five secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs (34152740610, 34152740604) both ended with conclusion `action_required` — workflow approval pending, not a test/build failure. Separately, `bin/fm-test-run.sh --check-coverage` exits 1 in this environment, but I confirmed by stashing my changes that it fails identically on the unmodified base tree (locale-related `comm: input is not in sorted order`); it is pre-existing and this change adds no new test file for the partition to account for. Changes are left uncommitted in the worktree

* no-mistakes(ci): Fixed the one real code defect behind the failing checks. ROOT CAUSE (Greptile P1, second round, on the head commit e1304e6): `secondmate_wake_stall_tick` in bin/fm-watch.sh identified the queue's drain position by the sequence number ALONE. The previous round changed the comparison from `-gt` to `-ne`, which handles a reprovisioned queue that restarts BELOW the recorded position, but not one that restarts ON it. A mate retired and reprovisioned under the same task id gets a fresh home whose wake-queue sequence counter restarts at 1 — and the retained parent progress marker very plausibly holds a low sequence too (a queue frozen on its first row records seq 1). Equal sequence ⇒ no reset ⇒ the brand-new queue inherits the retired generation's long-expired idle interval and emits a false `secondmate wake-loop stalled` on its very first observation. That is exactly the false-alarm class this PR exists to remove. FIX (smallest, removal-first): the file already defines the identity of a queue row once, as `row_key="$epoch-$seq"` (used for stall receipts, the stall marker, and the notify key). The progress marker's separate, weaker seq-only identity is removed: `row_key` is now computed once right after the row is parsed, stored in the progress marker, and compared with `!=`. Across generations the epoch differs (the new generation's rows are appended later), so no sequence collision can carry a stale interval; within a generation the key is stable exactly while the position does not move. bin/fm-wake-lib.sh's `fm_wake_secondmate_progress_marker_write` now takes `<oldest-row-key>` and validates it the same way the two neighbouring row-key writers do. Updated the function header comment and the two doc sentences (docs/architecture.md, docs/configuration.md) that described the old sequence-only semantics. REGRESSION TEST: `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` in tests/fm-wake-queue.test.sh now drives the reported case — the reprovisioned generation restarts on the SAME sequence 9 (epoch 200) that the retired generation recorded (epoch 100), at a later clock — and asserts observable watcher output only. Its third leg still confirms the restarted generation escalates on a genuine freeze (row=9 idle=2s), so the fix does not merely mute the alarm. Three seeded progress markers in the symlink, crash-window and prefix-receipt tests were updated to the epoch-sequence form. VERIFICATION: - Fails before / passes after: with bin/fm-watch.sh and bin/fm-wake-lib.sh reverted to HEAD and the new test in place, the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted` (exit 1); with the fix, `bash tests/fm-wake-queue.test.sh` exits 0 with 38/38 pass, all six secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). - I also started `bin/fm-test-run.sh tests/fm-watch-checkpoint.test.sh tests/fm-watch-triage.test.sh tests/fm-watch-recovery-loop.test.sh` as a blast-radius check; it was still running when this phase had to return, so its result is not included. No other suite references the stall detector or the progress marker (grep over tests/ for `wake-loop stall|SECONDMATE_WAKE_STALL|secondmate-wake-progress` matches only fm-wake-queue.test.sh), and the changed lib function has exactly one caller. CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs on the head commit (34154091191, 34154091236, 34154091945) all ended with conclusion `action_required` — pending workflow approval, not a test/bui…
tiago-peixoto added a commit to tiago-peixoto/firstmate that referenced this pull request Sep 12, 2026
* fix(bin): scope the worker role contract for ship and scout launches (#3797)

* fix(brief): scope Firstmate workers to their launch contract

* no-mistakes(document): Clarify supervisor scope and worker contract ownership

* no-mistakes(ci): Removed the heading-based bypass so every ship/scout launch receives the current worker-role contract. Added a regression that failed before the fix and passes afterward. Dispatch, brief, and delivery suites, focused ShellCheck, and git diff --check all passed

* no-mistakes(review): make launch overlay sole owner of worker role contract

* no-mistakes(review): narrow heading test dimension and fix publish error wording

* no-mistakes(review): gate role supersession, fix render guard, drop AGENTS twin

* no-mistakes(document): align architecture AGENTS.md scope and spawn launch-brief header

* fix(bin): stop ringing steering doorbells into dead panes (#3823)

* fix(bin): stop ringing steering doorbells into dead panes

The steering-inbox doorbell was a plain sentence plus Enter typed into a
worker's pane, and the watcher re-rang it on the assumption that a ring is
free. In a pane whose agent has exited that line is a shell command, and the
re-ring ladder kept typing it into a shell that can never acknowledge it.

- Prefix the doorbell with the shell no-op `: ` so a bare shell executes
  nothing while a live worker still reads the same self-describing line.
  `#` is not used because interactive zsh does not treat it as a comment by
  default and the claude harness binds it to memory mode.
- fm_task_inbox_ring skips the pane (return 3) when the backend positively
  classifies the agent as dead; missing, ambiguous, unreadable, and unverified
  endpoints still ring so a blind classifier never starves a live worker.
- The watcher caps the ladder for a dead pane: one stale wake for recovery,
  no ring, no ladder walk, and the durable record stays for
  stuck-crewmate-recovery. fm-send and the remote steer leg report the skip.

Tests cover the no-op in real shells, the dead/live/unclassifiable ring
verdicts, and the single-surfacing watcher path.

* no-mistakes(review): Quote doorbell paths against shell injection

* no-mistakes(review): Reject terminal-control paths before ringing

* no-mistakes(review): Document accepted partial doorbell delivery race

* no-mistakes(review): Skip unavailable endpoints before busy-state handling

* no-mistakes(test): Respect shell startup PATH in environment allowlist test

* no-mistakes(test): Fix doorbell test fixtures for endpoint liveness

* no-mistakes(test): Prioritize confirmed restarts and clean shell test syntax

* no-mistakes(document): Document dead and missing doorbell recovery

* no-mistakes(ci): Fixed the persistence-reply timeout race by rechecking for a correlated reply immediately before falling back to a nudge. Added a deterministic regression covering replies arriving between the preliminary resolution pass and timeout handling. Verified with the targeted restart suite, project lint, coverage guard, bash syntax checks, and diff checks

* fix(bin): wait out transient primary-checkout reads in the spawn worktree poll (#3834)

* fix(spawn): keep the worktree poll from adopting the repository primary

After `treehouse get` is sent, the worktree-discovery poll reads the pane's
foreground-process cwd. While treehouse is still fetching and checking a slot
out, the foreground process is treehouse itself and it reports the repository's
PRIMARY checkout as its cwd for several seconds. The poll accepted any path
that merely differed from the spawning project, so from a linked spawning home
- whose project is itself a worktree of that repository - it adopted the
primary, and the isolation guard then refused a launch whose slot treehouse
went on to create normally.

Screen every candidate with the isolation guard's own conditions, extracted as
spawn_worktree_isolated, so a read the guard would reject stays a transient the
poll keeps waiting through. The two-consecutive-reads rule and the guard as
final backstop are unchanged; a pane that never reaches an isolated worktree
still fails at the existing 60s deadline, now naming the last path it reported.

The already-settled timing assertion counted whole-spawn wall time against a
5s budget and failed on unmodified HEAD on slower machines; it now counts pane
reads, which is what "one confirming read, not an extra cycle" actually means.

* fix(spawn): say which path the worktree wait rejected, and why

Screening every discovery-poll candidate means a host that never reaches an
isolated worktree spends the whole 60s window before refusing. That wait is
deliberate - separating a transient from a terminal misconfiguration needs
machinery this path does not want - so the refusal explains itself instead:
the isolation check records why a candidate failed, and the deadline names the
last path seen together with that reason. Message and diagnostics only; the
poll's control flow is unchanged.

Two suites asserted the guard's wording on paths the poll now rejects rather
than adopts, so their refusal arrives from the deadline instead: realign
fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also
asserting the stated reason, and the second the metadata absence it was
missing) and the herdr projection e2e's forced non-worktree cwd.

* no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test

* no-mistakes(document): document spawn poll isolation screen in fm-spawn header

* test(spawn): make the non-git isolation case non-git anywhere

The refusal-reason assertion for a path outside any repository assumed TMPDIR
is not inside a git repository. Where it is, git walks up from the temporary
directory, finds that repository, and the spawn reports the subdirectory cause
instead - so the case passed or failed on a property of the host rather than on
the behaviour under test.

Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES,
which git documents as not chdir-ing up into a listed directory while looking
for a repository. Git never excludes the directory being searched, so the
ceiling is the parent of the path handed to the spawn.

The assertions pin which cause fired rather than the sentence that explains it,
leaving the operator wording free to improve.

* no-mistakes(document): point spawn poll comment at the isolation screen's comparison

* no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure

* fix: support stock macOS Bash 3.2 paths (#3732)

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): read orphaned green ci monitor and daemon-down failed record as not failed (#3846)

* fix(bin): read an orphaned green ci monitor as held-for-merge, not failed

A no-mistakes run held for a captain merge decision keeps its ci step
polling until merged or closed; when the shared daemon restarts under
that poll, the run is recorded failed although every substantive step
completed and GitHub reports the PR green. A monitor whose only
remaining job is to observe a human decision must not convert the
absence of that decision into a failure verdict.

fm-crew-state.sh now reclassifies a terminal failed run as done
(held-for-merge), surfacing the run's PR URL, when the steps table
shows every step completed except exactly ci failed and the ci log's
last recognized marker reads checks green. A genuinely red check, an
unreadable ci log, or a second failed step keeps the failure.

* no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed

* no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification

* fix(spawn): verify preserved backlog state after interrupted spawn delivery (#3852)

* fix(bin): read preserved spawn state back before the interrupted exit claims it

The deferred-signal exit path asserted the paired task record and
In-flight backlog state were preserved without reading either back,
exactly when a reader is least able to check (fm-yi4j evidence,
2026-09-05). The commit's exit status alone has been observed to agree
with a row that did not actually move.

The exit path now re-reads the record and the row under the same
per-task lock as the commit, repairs a row the commit believed it
moved, and phrases the error as exactly what was verified or attempted
- verified preserved, repaired and verified, or an explicit
preservation-could-not-be-verified with the reason and hand-closeout
instruction. Two behavior tests drive a lying tasks-axi start through a
real interrupted spawn and assert the printed claim and the real
backlog state agree.

* no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask

* no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean

* docs: correct runtime-backend maturity labels for Herdr (#3821)

* docs: correct stale tmux/herdr backend maturity claims

Herdr now has 21 test files, its own required CI job (tests-herdr)
that installs a pinned build and hard-fails on "skip: herdr not
found", while tmux has 3 test files and is only required as a
dependency of the portable-serial e2e lane. zellij, orca, and cmux
still have no CI lane at all. AGENTS.md and docs/herdr-backend.md
still called Herdr merely "experimental" alongside those three,
misleading every session and reader about actual coverage.

Update AGENTS.md's config/backend entry, the opening lines of
docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend
section of docs/configuration.md, and the matching claims in
docs/architecture.md, CONTRIBUTING.md, and README.md so they agree
and distinguish tmux (default), herdr (own required CI lane, largest
suite, Windows still spike-only), and zellij/orca/cmux (still
experimental, no CI lane). No behavior, selection order, or
dispatch logic changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj

* no-mistakes(review): docs: fix stale herdr label and CI-lane wording

* no-mistakes(document): docs: align tmux adapter label in scripts.md

* no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page

* no-mistakes(review): docs: drop windows claim, align contributing backend wording

* no-mistakes(review): docs: trim duplicated CI claim from herdr opening line

* no-mistakes(review): docs: drop unguarded largest-test-suite superlative

* no-mistakes(review): docs: restore tmux verified label and README experimental scope

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: restore intent-targeted no-mistakes validation (#3865)

* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR #823 added and PR #1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: verify Treehouse slot ownership before teardown (#3837)

* fix(bin): verify pool-slot ownership before returning a worktree slot

Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.

Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.

* no-mistakes(review): Protect slots across cloned Firstmate homes

* no-mistakes(test): Gate teardown locking on genuine Treehouse slots

* no-mistakes(test): Clarify pooled descendant slot gating

* no-mistakes(test): Synchronize watcher re-arm test on process exit

* no-mistakes(test): Wait for watcher cleanup before timeout escalation

* no-mistakes(document): Document pool-slot ownership safeguards

* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion

* fix(bin): resolve relative origins from repository root

* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout

* feat(bin): add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary (#3867)

* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary

Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.

Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test: prove the omp guard continuation through a guard spy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* fix(spawn): clear the gemini marker at the omp launch boundary

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): force the guard stage by freezing the watcher and clear lint findings

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): reap the live lab by path and record omp's rpc shutdown as a note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces

Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin

* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs

* no-mistakes(review): omp: pin config-model validation with a test, trim overlay

* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family

* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes

* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux

* no-mistakes(document): docs: add omp subagent-guard row, fix live test header

* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π|󰵗|pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)

* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|󰵗)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix(pi): invoke Bash helpers correctly on native Windows (#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>

* test(bin): pin teardown outcomes for squash-merged rebased branches (#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>

* fix(bin): keep supervision armed for registered custom checks (#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes #3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner

* fix(bin): resolve Treehouse locks for remote secondmate homes (#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership

* fix(bearings): repair board listening and decision reconciliation (#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable state - refuses
rather than acting.

Regression coverage fails without each fix, and pins every leak path: a bare
reconcile, a standalone close or note with no pending request, an any-channel
reconcile, an annotated selection from a freeform card, and a
generation-skewed authorization. An opt-in guard re-proves the lavish-axi
shapes and the reopen against the installed tool.

* fix(bin): quote the done comparison in the reconcile intake

shellcheck SC1010 reads the bare word as the loop keyword. The failed run
never reached its lint step, so this shipped in the recovered content.

* no-mistakes(review): Publish reconciled parent resolution before request retirement

* no-mistakes(review): Clarify committed cleanup and reconcile reservation scope

* no-mistakes(review): Preserve remote cards and legacy answer compatibility

* no-mistakes(test): Separate live claim release from stale reclamation

* no-mistakes(test): Allow terminal self-retirement during active capture

* no-mistakes(document): Document Bearings repair contracts

* no-mistakes(ci): Stabilized the failing Herdr presentation E2E by serializing test-harness Treehouse allocator calls, preventing concurrent recovery spawns from claiming the same pool slot while preserving Herdr concurrency coverage. Verified with the full E2E suite on Herdr 0.8.2, bash syntax checks, ShellCheck, and git diff checks

* fix(bin): allow pooled spawns without a git origin (#3885)

* fix(bin): skip pooled-worktree freshness fetch when no origin is configured

An origin-less local-only project has nothing remote to be stale
against, so fm-spawn's freshen_spawn_worktree_base refused to launch
crews for it. Detect a missing origin remote and skip the fetch
freshness gate entirely; an existing-but-unreachable origin keeps
refusing as before.

* no-mistakes(review): Preserve pool safety for absent and unusable origins

* no-mistakes(review): Refuse empty origin configurations during pooled spawn

* no-mistakes(review): Detect empty origin sections across config includes

* no-mistakes(review): Honor globbed includes when detecting origin configuration

* no-mistakes(review): Document conservative conditional include handling

* no-mistakes(review): Use Git-resolved config files for origin detection

* no-mistakes(review): Document included empty-origin detection boundary

* no-mistakes(document): Document originless pooled spawn behavior

* feat(tests): run live harness guards by default when available (#3889)

* feat(tests): run live harness guards by default where the harness is installed

The 24 live-harness guards each opened with their own env check, so on the
machine that has every harness - the one the product and its validation
actually run on - all of them skipped and passed. Fourteen had never been run
by the pipeline at all.

tests/lib.sh gains fm_live_gate as the single owner of that decision: a guard
that spends no model tokens runs wherever its tools are installed, a guard that
submits prompts stays opt-in, an absent tool is a named capability skip, and a
guard's own variable or FM_LIVE forces it on (turning an absent tool into a
failure) or off. Every live guard now opens with it, which also carries the
test-suite gate-refusal bypass into the guards that never sourced the shared
helpers and were therefore refused whenever a gate agent ran them.

bin/fm-test-run.sh records what a skip means: the family's expected class is
live-capability rather than a bare env opt-in, and each gate skip's reason is
logged and written to the timing artifact, so a lane can say which tool this
host could not exercise.

Only the token-free guards flip to default-on: composer-matrix, the harness
liveness drift guard, and the Herdr version floor. cursor-primary submits three
prompts, so it stays opt-in.

Running the drift guard unasked immediately found a real defect it existed to
catch: it resolved the harness through a generic `command -v cursor`, which on
a machine that also has the Cursor editor finds the editor launcher rather than
cursor-agent. That binary exits at once, leaving a bare shell in the pane and a
liveness-drift failure no classifier change could fix. It now asks
fm_cursor_resolve_binary first, the same verified owner fm-spawn uses.

CI installs the public Pi package in the portable serial lane and fails on its
skip token, so the Pi extension tests stop passing silently against a package
that is not there. No secret is added.

Verified on macOS 26.5.2 arm64: the drift guard runs with no variable set and
classifies 8 installed harnesses alive; the Herdr version-floor guard runs by
default and checks 4 real releases; every live guard refuses together under
FM_LIVE=0.

* fix(tests): keep the composer-matrix guard opt-in

Running it unasked is red on a healthy machine for reasons no code change here
removes: a harness that has not trusted this checkout sits on its own trust
dialog, which the guard treats as an unreadable composer and correctly fails.
The opencode 1.18.29 and grok 1.0.13 composer drift it also surfaced reproduces
identically on main and is filed as separate work.

So this token-free guard stays opt-in with the reason stated in its header, and
the coding guidelines record the narrow exception: a guard whose verdict
depends on host state that installing its tools does not establish may stay
opt-in, because one that is permanently red is one the fleet learns to ignore.
The other two token-free guards keep running by default.

* no-mistakes(review): Wire bearings guard and remove composer exception policy

* no-mistakes(review): Run Pi responsiveness guard by default

* no-mistakes(review): Gate AFK Pi Herdr through authoritative family sweep

* no-mistakes(review): Sanitize live gate test environments

* no-mistakes(document): Document default-on live guard behavior

* no-mistakes(ci): Fixed CI by installing the Pi package in portable-parallel-1, where fm-pi-primary-types.test.sh runs, and enforcing its package-missing gate skip there. Verified with fm-lint.sh, workflow actionlint, coverage partition checks, lane membership, and git diff checks

* no-mistakes(ci): Fixed CI’s Pi typecheck skip enforcement by giving npm, tsc, and Pi-package capability skips a shared prefix and configuring both relevant CI lanes to fail on that prefix. Verified missing tsc emits the expected skip, missing Pi package becomes a runner failure, and actionlint, ShellCheck, and git diff checks pass

* fix(bin): refuse test runs in the primary checkout when a task marker is set (#3891)

* fix(bin): refuse the behavior suite in the repository primary checkout

A task worker's isolated worktree placement is verified exactly once, when
its task starts, and nothing re-checks it afterwards. A worker that later
changes directory into the repository's primary checkout runs its Git
commands, and this branch-switching suite, against the one checkout every
linked worktree resolves against and every landing merges into. A run that
dies mid-suite can leave that checkout on a stray branch.

bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task
worker and the runner resolves to the primary checkout, every executing mode
exits non-zero before selecting a suite, with one line naming the primary
path and pointing at the assigned task worktree. The predicate is the one
bin/fm-spawn.sh already uses for launch placement: the working tree's own
git dir is the repository's common git dir, which separates the primary from
every linked worktree even when their top levels differ. A run with no
FM_TASK_ID set is unchanged, and so are the inspection modes, which execute
nothing. When git resolves neither directory - a non-repository fixture, a
detached copy - nothing proves this is the primary, so the run proceeds.

bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID
into the pane shell on the same pre-launch channel as GOTMPDIR, and the name
joins the sanitized launch environment allowlist so an isolated launch keeps
it.

* no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT

* no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal

---------

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): bound stale alarms for backlog captain holds (#3842)

* fix(bin): bound a stale alarm with the backlog hold, not only the status line

A legitimate wait has two records and the stale alarm reads only one.
`status_is_paused_or_captain_held` takes a status line, so it sees a wait the
worker declared. It cannot see the wait firstmate records when it hands work to
the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and
leaves the status log alone, so a delivered task keeps `done: PR ...` as its
last line for the whole time the captain is deciding.

Both stale branches were blind to it, and each churned a new pane hash back into
its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale
branch, while a held task whose last line is `working:` reaches
`surface_nonterminal_stale` and fails its declared-wait test.

Consult that second record where the watcher is about to alarm, through
`bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and
bound the alarm on the shared `.paused-resurfaced-<key>` marker and
`PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first
sight still alarms, the window's end alarms once more, and a held crew that goes
genuinely silent still escalates through the wedge timer.

Only an established open captain call bounds anything: an unreadable backlog, an
absent or incompatible tasks-axi, a row this home does not carry, and every task
with no hold keep alarming exactly as before. The backlog hold is deliberately
not recorded as a declared pause, because the loop-top reconciliation and
`pause_state_class` both read the status line and would clear a flag that line
does not support.

Extends the fix in #3443, which closed the forms of this loop that the status
line itself can express.

* fix(bin): identify the captain call a stale alarm is bounded by

Three gaps in the bound added by the previous commit, all in how the throttle is
scoped and where the backlog is consulted.

The scope carried only the status-log signature. A task can be held, answered
with `--release`, and re-held as a genuinely different captain call without any
status append, so the second call inherited the first one's marker and its first
sight was absorbed - the one thing this bound must never do. The task id is not
the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the
call's own lifecycle - its hold-set stamp and the number of recorded answers -
on an exit 0 and only then, leaving the silent predicate every existing caller
reads unchanged. The throttle scope now carries that identity.

The terminal path recorded the throttle before publishing the durable wake. A
failed append exits the watcher with nothing queued, and the next sighting then
read that fresh marker and absorbed the retry, turning a delayed alarm into a
lost one. Recording moves behind the append, as the non-terminal path already
had it, and the comment claiming the marker could not outlive its wake is gone
because it was false.

The backlog was consulted only on a new terminal pane hash. A captain call can
open after a hash was absorbed as provably working, changing neither the pane nor
the status log, so nothing re-read the backlog and the wedge timer kept firing
possible-wedge alarms through a legitimate wait. That timer now consults the call
at its own alarm boundary and takes the same bounded cadence - and only at that
boundary, so an ordinary repeat poll under the bound stays the local-only read it
was.

Regression coverage for each, all driving churn through one watcher process
rather than relaunching per pane change: relaunch cost dominated the earlier
shape, and an absorbing watcher stays in its poll loop across churn in production
anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer
with no open captain call still escalates as a possible wedge.

* fix(review): Compose stale throttles with captain-call lifecycle identity

* fix(review): Preserve bounded same-hash captain-call resurfacing

* revert(bin): narrow the captain-hold stale bound to its observed defect

Lifts the lifecycle-identity and cadence-ownership work back out, leaving the
change at the shape that matches the defect actually observed: the stale alarm
did not consult the backlog captain hold, on either stale branch.

Reviewing the wider version surfaced a series of adjacent gaps in the watcher's
alarm state machine - a call opening after the first alarm, marker invalidation
at the hold lifecycle boundary, and which deadline a terminal timer represents.
They are real, but fixing them turns a small extension into a state-machine
change to the alarm path, which is a different review on a subsystem that is
being actively reworked. They are named as known limitations rather than carried
here, and none of them is load-bearing for what remains: the bound does strictly
less than the reverted version, leaves the wedge path escalating on
STALE_ESCALATE_SECS exactly as before, and introduces no silence that the
existing terminal-alarm path did not already have.

Kept from the reverted work is the record-after-append ordering, because that is
a defect in the code being shipped rather than an adjacent one: recording the
cadence marker before publishing the durable wake let a failed append lose an
alarm outright instead of delaying it.

History is preserved: the earlier commits stay on the branch and this removal
sits on top of them.

* fix(review): Document secondmate captain-hold scope boundary

* fix(document): Document captain-hold stale alarm scope

* fix(bin): bind the stale throttle to the captain call, not the status log

The throttle this change introduces was scoped to the task's status-log
signature. Answering a call with `--release` and holding the task again creates a
genuinely different captain call without necessarily appending to that log, so
the second call inherited the first one's marker and its first sight was
absorbed.

That is the one alarm this bound must never swallow. A delivery announced twice
is noise; a decision waiting on the captain that is never surfaced is invisible,
because nobody asks for what they do not know to ask for.

Measured rather than assumed, on the same fixture - a delivered task held for the
captain, released, and re-held with no status append, driven through bin/fm-watch.sh:

  base c499f84    call-1 first=ALARM  call-1 churn=ALARM     new call first sight=ALARM
  before this fix call-1 first=ALARM  call-1 churn=absorbed  new call first sight=absorbed
  after           call-1 first=ALARM  call-1 churn=absorbed  new call first sight=ALARM

Base never suppresses the new call, so the suppression came from this change and
closing it completes the fix rather than widening it.

`bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle -
its hold-set stamp and count of recorded answers - on an exit 0 and only then, so
the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope
carries that identity beside the status signature.

The sibling case was measured too and is NOT included: on the status-declared
path, where the last line is `captain-held:`, base already absorbs a re-held
call's first sight. That behaviour predates this change and stays documented as a
known limitation rather than repaired here.

* fix(document): Document captain-call throttle lifecycle scope

* fix(ci): isolate the Herdr restart fixtures from a claimed worktree

The Herdr behaviour test intermittently reused a local worktree still claimed
by an earlier fixture after a restart.
The restart scenarios now use an isolated Treehouse project.

The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as
well.

* feat(pi): resolve extension-registered providers in the supervision branch (#3871)

* Let the supervision branch resolve extension-registered providers

The isolated branch ModelRuntime cannot see providers an extension
registered into main's runtime at run time, so a pin on pi-devin-auth's
devin/swe-1-7 (or an unpinned branch following a main session on devin)
failed with "unavailable to the isolated branch runtime".

Capture main's ModelRegistry alongside mainModel and copy each
extension-registered provider config into the branch runtime at
model-resolution time. The config carries the provider's own streamSimple
and oauth wiring by reference, so the custom gRPC transport reaches the
branch unchanged instead of being reimplemented. The /supervision-model
picker uses the same copy so those models are offered.

Update configuration.md and pi-supervision-branch.md, which previously
stated extension-registered providers were not offered.

* no-mistakes(document): docs: own devin provider carve-out in branch architecture doc

* no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered

* ci: retrigger flaky Herdr/serial-1 lanes

* no-mistakes(document): docs already cover branch extension-provider copy

* fix(bin): gate secondmate wake-loop stall alerts on real queue no-progress (#3943)

* fix(watch): detect stalled secondmate queue progress

* no-mistakes(review): gate secondmate stall on active turns and progress episodes

* no-mistakes(document): align secondmate wake-stall docs with progress-episode detector

* no-mistakes(document): clarify active-turn gate in wake-stall config docs

* no-mistakes(ci): Fixed the one real defect behind the failing checks. ROOT CAUSE (Greptile P1, real code defect in this PR): `secondmate_wake_stall_tick` in bin/fm-watch.sh reset the no-progress timer only when the oldest actionable queue sequence INCREASED (`[ "$seq" -gt "$observed_seq" ]`). When a secondmate is retired and reprovisioned under the same task ID, its fresh home's queue sequence restarts BELOW the recorded position, so the comparison is false, no reset happens, and the new queue inherits the retired generation's already-expired idle interval — emitting a false `secondmate wake-loop stalled` on its very first observation. That is precisely the false-alarm class the user intent requires this PR to remove. FIX (smallest, removal-first): bin/fm-watch.sh:749 now resets when the drain position MOVES AT ALL (`-ne` instead of `-gt`). Draining moves it up, reprovisioning moves it down; neither is a continued no-progress episode. The asymmetric `-gt` branch is removed rather than special-cased or hardened. Updated the function header comment plus the two doc sentences in docs/architecture.md and docs/configuration.md that stated the old advance-only semantics. REGRESSION TEST: added `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` to tests/fm-wake-queue.test.sh (registered in the invocation list). It drives the real watcher through retired generation (seq 9) -> reprovision (seq 3, later clock) -> freeze, asserting observable wake-queue output, no source-text inspection. VERIFICATION: - Fails before / passes after: with the fix reverted the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted`; with the fix it passes, and its third leg confirms the restarted generation still escalates on a genuine freeze (row=3 idle=2s), so the fix does not merely mute the alarm. - `bash tests/fm-wake-queue.test.sh`: exit 0, 38/38 pass, all five secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs (34152740610, 34152740604) both ended with conclusion `action_required` — workflow approval pending, not a test/build failure. Separately, `bin/fm-test-run.sh --check-coverage` exits 1 in this environment, but I confirmed by stashing my changes that it fails identically on the unmodified base tree (locale-related `comm: input is not in sorted order`); it is pre-existing and this change adds no new test file for the partition to account for. Changes are left uncommitted in the worktree

* no-mistakes(ci): Fixed the one real code defect behind the failing checks. ROOT CAUSE (Greptile P1, second round, on the head commit e1304e6): `secondmate_wake_stall_tick` in bin/fm-watch.sh identified the queue's drain position by the sequence number ALONE. The previous round changed the comparison from `-gt` to `-ne`, which handles a reprovisioned queue that restarts BELOW the recorded position, but not one that restarts ON it. A mate retired and reprovisioned under the same task id gets a fresh home whose wake-queue sequence counter restarts at 1 — and the retained parent progress marker very plausibly holds a low sequence too (a queue frozen on its first row records seq 1). Equal sequence ⇒ no reset ⇒ the brand-new queue inherits the retired generation's long-expired idle interval and emits a false `secondmate wake-loop stalled` on its very first observation. That is exactly the false-alarm class this PR exists to remove. FIX (smallest, removal-first): the file already defines the identity of a queue row once, as `row_key="$epoch-$seq"` (used for stall receipts, the stall marker, and the notify key). The progress marker's separate, weaker seq-only identity is removed: `row_key` is now computed once right after the row is parsed, stored in the progress marker, and compared with `!=`. Across generations the epoch differs (the new generation's rows are appended later), so no sequence collision can carry a stale interval; within a generation the key is stable exactly while the position does not move. bin/fm-wake-lib.sh's `fm_wake_secondmate_progress_marker_write` now takes `<oldest-row-key>` and validates it the same way the two neighbouring row-key writers do. Updated the function header comment and the two doc sentences (docs/architecture.md, docs/configuration.md) that described the old sequence-only semantics. REGRESSION TEST: `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` in tests/fm-wake-queue.test.sh now drives the reported case — the reprovisioned generation restarts on the SAME sequence 9 (epoch 200) that the retired generation recorded (epoch 100), at a later clock — and asserts observable watcher output only. Its third leg still confirms the restarted generation escalates on a genuine freeze (row=9 idle=2s), so the fix does not merely mute the alarm. Three seeded progress markers in the symlink, crash-window and prefix-receipt tests were updated to the epoch-sequence form. VERIFICATION: - Fails before / passes after: with bin/fm-watch.sh and bin/fm-wake-lib.sh reverted to HEAD and the new test in place, the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted` (exit 1); with the fix, `bash tests/fm-wake-queue.test.sh` exits 0 with 38/38 pass, all six secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). - I also started `bin/fm-test-run.sh tests/fm-watch-checkpoint.test.sh tests/fm-watch-triage.test.sh tests/fm-watch-recovery-loop.test.sh` as a blast-radius check; it was still running when this phase had to return, so its result is not included. No other suite references the stall detector or the progress marker (grep over tests/ for `wake-loop stall|SECONDMATE_WAKE_STALL|secondmate-wake-progress` matches only fm-wake-queue.test.sh), and the changed lib function has exactly one caller. CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs on the head commit (34154091191, 34154091236, 34154091945) all ended with conclusion `action_required` — pending workflow approval, not a test/bui…
BenWilcox8 pushed a commit to BenWilcox8/firstmate that referenced this pull request Sep 12, 2026
…id#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation
BenWilcox8 pushed a commit to BenWilcox8/firstmate that referenced this pull request Sep 12, 2026
* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR kunchenguid#3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR kunchenguid#823 added and PR kunchenguid#1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
eduardstan added a commit to eduardstan/firstmate that referenced this pull request Sep 13, 2026
* fix(bin): wait out transient primary-checkout reads in the spawn worktree poll (#3834)

* fix(spawn): keep the worktree poll from adopting the repository primary

After `treehouse get` is sent, the worktree-discovery poll reads the pane's
foreground-process cwd. While treehouse is still fetching and checking a slot
out, the foreground process is treehouse itself and it reports the repository's
PRIMARY checkout as its cwd for several seconds. The poll accepted any path
that merely differed from the spawning project, so from a linked spawning home
- whose project is itself a worktree of that repository - it adopted the
primary, and the isolation guard then refused a launch whose slot treehouse
went on to create normally.

Screen every candidate with the isolation guard's own conditions, extracted as
spawn_worktree_isolated, so a read the guard would reject stays a transient the
poll keeps waiting through. The two-consecutive-reads rule and the guard as
final backstop are unchanged; a pane that never reaches an isolated worktree
still fails at the existing 60s deadline, now naming the last path it reported.

The already-settled timing assertion counted whole-spawn wall time against a
5s budget and failed on unmodified HEAD on slower machines; it now counts pane
reads, which is what "one confirming read, not an extra cycle" actually means.

* fix(spawn): say which path the worktree wait rejected, and why

Screening every discovery-poll candidate means a host that never reaches an
isolated worktree spends the whole 60s window before refusing. That wait is
deliberate - separating a transient from a terminal misconfiguration needs
machinery this path does not want - so the refusal explains itself instead:
the isolation check records why a candidate failed, and the deadline names the
last path seen together with that reason. Message and diagnostics only; the
poll's control flow is unchanged.

Two suites asserted the guard's wording on paths the poll now rejects rather
than adopts, so their refusal arrives from the deadline instead: realign
fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also
asserting the stated reason, and the second the metadata absence it was
missing) and the herdr projection e2e's forced non-worktree cwd.

* no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test

* no-mistakes(document): document spawn poll isolation screen in fm-spawn header

* test(spawn): make the non-git isolation case non-git anywhere

The refusal-reason assertion for a path outside any repository assumed TMPDIR
is not inside a git repository. Where it is, git walks up from the temporary
directory, finds that repository, and the spawn reports the subdirectory cause
instead - so the case passed or failed on a property of the host rather than on
the behaviour under test.

Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES,
which git documents as not chdir-ing up into a listed directory while looking
for a repository. Git never excludes the directory being searched, so the
ceiling is the parent of the path handed to the spawn.

The assertions pin which cause fired rather than the sentence that explains it,
leaving the operator wording free to improve.

* no-mistakes(document): point spawn poll comment at the isolation screen's comparison

* no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure

* fix: support stock macOS Bash 3.2 paths (#3732)

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): read orphaned green ci monitor and daemon-down failed record as not failed (#3846)

* fix(bin): read an orphaned green ci monitor as held-for-merge, not failed

A no-mistakes run held for a captain merge decision keeps its ci step
polling until merged or closed; when the shared daemon restarts under
that poll, the run is recorded failed although every substantive step
completed and GitHub reports the PR green. A monitor whose only
remaining job is to observe a human decision must not convert the
absence of that decision into a failure verdict.

fm-crew-state.sh now reclassifies a terminal failed run as done
(held-for-merge), surfacing the run's PR URL, when the steps table
shows every step completed except exactly ci failed and the ci log's
last recognized marker reads checks green. A genuinely red check, an
unreadable ci log, or a second failed step keeps the failure.

* no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed

* no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification

* fix(spawn): verify preserved backlog state after interrupted spawn delivery (#3852)

* fix(bin): read preserved spawn state back before the interrupted exit claims it

The deferred-signal exit path asserted the paired task record and
In-flight backlog state were preserved without reading either back,
exactly when a reader is least able to check (fm-yi4j evidence,
2026-09-05). The commit's exit status alone has been observed to agree
with a row that did not actually move.

The exit path now re-reads the record and the row under the same
per-task lock as the commit, repairs a row the commit believed it
moved, and phrases the error as exactly what was verified or attempted
- verified preserved, repaired and verified, or an explicit
preservation-could-not-be-verified with the reason and hand-closeout
instruction. Two behavior tests drive a lying tasks-axi start through a
real interrupted spawn and assert the printed claim and the real
backlog state agree.

* no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask

* no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean

* docs: correct runtime-backend maturity labels for Herdr (#3821)

* docs: correct stale tmux/herdr backend maturity claims

Herdr now has 21 test files, its own required CI job (tests-herdr)
that installs a pinned build and hard-fails on "skip: herdr not
found", while tmux has 3 test files and is only required as a
dependency of the portable-serial e2e lane. zellij, orca, and cmux
still have no CI lane at all. AGENTS.md and docs/herdr-backend.md
still called Herdr merely "experimental" alongside those three,
misleading every session and reader about actual coverage.

Update AGENTS.md's config/backend entry, the opening lines of
docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend
section of docs/configuration.md, and the matching claims in
docs/architecture.md, CONTRIBUTING.md, and README.md so they agree
and distinguish tmux (default), herdr (own required CI lane, largest
suite, Windows still spike-only), and zellij/orca/cmux (still
experimental, no CI lane). No behavior, selection order, or
dispatch logic changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj

* no-mistakes(review): docs: fix stale herdr label and CI-lane wording

* no-mistakes(document): docs: align tmux adapter label in scripts.md

* no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page

* no-mistakes(review): docs: drop windows claim, align contributing backend wording

* no-mistakes(review): docs: trim duplicated CI claim from herdr opening line

* no-mistakes(review): docs: drop unguarded largest-test-suite superlative

* no-mistakes(review): docs: restore tmux verified label and README experimental scope

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: restore intent-targeted no-mistakes validation (#3865)

* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR #823 added and PR #1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: verify Treehouse slot ownership before teardown (#3837)

* fix(bin): verify pool-slot ownership before returning a worktree slot

Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.

Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.

* no-mistakes(review): Protect slots across cloned Firstmate homes

* no-mistakes(test): Gate teardown locking on genuine Treehouse slots

* no-mistakes(test): Clarify pooled descendant slot gating

* no-mistakes(test): Synchronize watcher re-arm test on process exit

* no-mistakes(test): Wait for watcher cleanup before timeout escalation

* no-mistakes(document): Document pool-slot ownership safeguards

* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion

* fix(bin): resolve relative origins from repository root

* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout

* feat(bin): add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary (#3867)

* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary

Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.

Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test: prove the omp guard continuation through a guard spy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* fix(spawn): clear the gemini marker at the omp launch boundary

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): force the guard stage by freezing the watcher and clear lint findings

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): reap the live lab by path and record omp's rpc shutdown as a note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces

Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin

* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs

* no-mistakes(review): omp: pin config-model validation with a test, trim overlay

* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family

* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes

* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux

* no-mistakes(document): docs: add omp subagent-guard row, fix live test header

* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π|󰵗|pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)

* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|󰵗)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix(pi): invoke Bash helpers correctly on native Windows (#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>

* test(bin): pin teardown outcomes for squash-merged rebased branches (#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>

* fix(bin): keep supervision armed for registered custom checks (#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes #3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner

* fix(bin): resolve Treehouse locks for remote secondmate homes (#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership

* fix(bearings): repair board listening and decision reconciliation (#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable state - refuses
rather than acting.

Regression coverage fails without each fix, and pins every leak path: a bare
reconcile, a standalone close or note with no pending request, an any-channel
reconcile, an annotated selection from a freeform card, and a
generation-skewed authorization. An opt-in guard re-proves the lavish-axi
shapes and the reopen against the installed tool.

* fix(bin): quote the done comparison in the reconcile intake

shellcheck SC1010 reads the bare word as the loop keyword. The failed run
never reached its lint step, so this shipped in the recovered content.

* no-mistakes(review): Publish reconciled parent resolution before request retirement

* no-mistakes(review): Clarify committed cleanup and reconcile reservation scope

* no-mistakes(review): Preserve remote cards and legacy answer compatibility

* no-mistakes(test): Separate live claim release from stale reclamation

* no-mistakes(test): Allow terminal self-retirement during active capture

* no-mistakes(document): Document Bearings repair contracts

* no-mistakes(ci): Stabilized the failing Herdr presentation E2E by serializing test-harness Treehouse allocator calls, preventing concurrent recovery spawns from claiming the same pool slot while preserving Herdr concurrency coverage. Verified with the full E2E suite on Herdr 0.8.2, bash syntax checks, ShellCheck, and git diff checks

* fix(bin): allow pooled spawns without a git origin (#3885)

* fix(bin): skip pooled-worktree freshness fetch when no origin is configured

An origin-less local-only project has nothing remote to be stale
against, so fm-spawn's freshen_spawn_worktree_base refused to launch
crews for it. Detect a missing origin remote and skip the fetch
freshness gate entirely; an existing-but-unreachable origin keeps
refusing as before.

* no-mistakes(review): Preserve pool safety for absent and unusable origins

* no-mistakes(review): Refuse empty origin configurations during pooled spawn

* no-mistakes(review): Detect empty origin sections across config includes

* no-mistakes(review): Honor globbed includes when detecting origin configuration

* no-mistakes(review): Document conservative conditional include handling

* no-mistakes(review): Use Git-resolved config files for origin detection

* no-mistakes(review): Document included empty-origin detection boundary

* no-mistakes(document): Document originless pooled spawn behavior

* feat(tests): run live harness guards by default when available (#3889)

* feat(tests): run live harness guards by default where the harness is installed

The 24 live-harness guards each opened with their own env check, so on the
machine that has every harness - the one the product and its validation
actually run on - all of them skipped and passed. Fourteen had never been run
by the pipeline at all.

tests/lib.sh gains fm_live_gate as the single owner of that decision: a guard
that spends no model tokens runs wherever its tools are installed, a guard that
submits prompts stays opt-in, an absent tool is a named capability skip, and a
guard's own variable or FM_LIVE forces it on (turning an absent tool into a
failure) or off. Every live guard now opens with it, which also carries the
test-suite gate-refusal bypass into the guards that never sourced the shared
helpers and were therefore refused whenever a gate agent ran them.

bin/fm-test-run.sh records what a skip means: the family's expected class is
live-capability rather than a bare env opt-in, and each gate skip's reason is
logged and written to the timing artifact, so a lane can say which tool this
host could not exercise.

Only the token-free guards flip to default-on: composer-matrix, the harness
liveness drift guard, and the Herdr version floor. cursor-primary submits three
prompts, so it stays opt-in.

Running the drift guard unasked immediately found a real defect it existed to
catch: it resolved the harness through a generic `command -v cursor`, which on
a machine that also has the Cursor editor finds the editor launcher rather than
cursor-agent. That binary exits at once, leaving a bare shell in the pane and a
liveness-drift failure no classifier change could fix. It now asks
fm_cursor_resolve_binary first, the same verified owner fm-spawn uses.

CI installs the public Pi package in the portable serial lane and fails on its
skip token, so the Pi extension tests stop passing silently against a package
that is not there. No secret is added.

Verified on macOS 26.5.2 arm64: the drift guard runs with no variable set and
classifies 8 installed harnesses alive; the Herdr version-floor guard runs by
default and checks 4 real releases; every live guard refuses together under
FM_LIVE=0.

* fix(tests): keep the composer-matrix guard opt-in

Running it unasked is red on a healthy machine for reasons no code change here
removes: a harness that has not trusted this checkout sits on its own trust
dialog, which the guard treats as an unreadable composer and correctly fails.
The opencode 1.18.29 and grok 1.0.13 composer drift it also surfaced reproduces
identically on main and is filed as separate work.

So this token-free guard stays opt-in with the reason stated in its header, and
the coding guidelines record the narrow exception: a guard whose verdict
depends on host state that installing its tools does not establish may stay
opt-in, because one that is permanently red is one the fleet learns to ignore.
The other two token-free guards keep running by default.

* no-mistakes(review): Wire bearings guard and remove composer exception policy

* no-mistakes(review): Run Pi responsiveness guard by default

* no-mistakes(review): Gate AFK Pi Herdr through authoritative family sweep

* no-mistakes(review): Sanitize live gate test environments

* no-mistakes(document): Document default-on live guard behavior

* no-mistakes(ci): Fixed CI by installing the Pi package in portable-parallel-1, where fm-pi-primary-types.test.sh runs, and enforcing its package-missing gate skip there. Verified with fm-lint.sh, workflow actionlint, coverage partition checks, lane membership, and git diff checks

* no-mistakes(ci): Fixed CI’s Pi typecheck skip enforcement by giving npm, tsc, and Pi-package capability skips a shared prefix and configuring both relevant CI lanes to fail on that prefix. Verified missing tsc emits the expected skip, missing Pi package becomes a runner failure, and actionlint, ShellCheck, and git diff checks pass

* fix(bin): refuse test runs in the primary checkout when a task marker is set (#3891)

* fix(bin): refuse the behavior suite in the repository primary checkout

A task worker's isolated worktree placement is verified exactly once, when
its task starts, and nothing re-checks it afterwards. A worker that later
changes directory into the repository's primary checkout runs its Git
commands, and this branch-switching suite, against the one checkout every
linked worktree resolves against and every landing merges into. A run that
dies mid-suite can leave that checkout on a stray branch.

bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task
worker and the runner resolves to the primary checkout, every executing mode
exits non-zero before selecting a suite, with one line naming the primary
path and pointing at the assigned task worktree. The predicate is the one
bin/fm-spawn.sh already uses for launch placement: the working tree's own
git dir is the repository's common git dir, which separates the primary from
every linked worktree even when their top levels differ. A run with no
FM_TASK_ID set is unchanged, and so are the inspection modes, which execute
nothing. When git resolves neither directory - a non-repository fixture, a
detached copy - nothing proves this is the primary, so the run proceeds.

bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID
into the pane shell on the same pre-launch channel as GOTMPDIR, and the name
joins the sanitized launch environment allowlist so an isolated launch keeps
it.

* no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT

* no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal

---------

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): bound stale alarms for backlog captain holds (#3842)

* fix(bin): bound a stale alarm with the backlog hold, not only the status line

A legitimate wait has two records and the stale alarm reads only one.
`status_is_paused_or_captain_held` takes a status line, so it sees a wait the
worker declared. It cannot see the wait firstmate records when it hands work to
the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and
leaves the status log alone, so a delivered task keeps `done: PR ...` as its
last line for the whole time the captain is deciding.

Both stale branches were blind to it, and each churned a new pane hash back into
its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale
branch, while a held task whose last line is `working:` reaches
`surface_nonterminal_stale` and fails its declared-wait test.

Consult that second record where the watcher is about to alarm, through
`bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and
bound the alarm on the shared `.paused-resurfaced-<key>` marker and
`PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first
sight still alarms, the window's end alarms once more, and a held crew that goes
genuinely silent still escalates through the wedge timer.

Only an established open captain call bounds anything: an unreadable backlog, an
absent or incompatible tasks-axi, a row this home does not carry, and every task
with no hold keep alarming exactly as before. The backlog hold is deliberately
not recorded as a declared pause, because the loop-top reconciliation and
`pause_state_class` both read the status line and would clear a flag that line
does not support.

Extends the fix in #3443, which closed the forms of this loop that the status
line itself can express.

* fix(bin): identify the captain call a stale alarm is bounded by

Three gaps in the bound added by the previous commit, all in how the throttle is
scoped and where the backlog is consulted.

The scope carried only the status-log signature. A task can be held, answered
with `--release`, and re-held as a genuinely different captain call without any
status append, so the second call inherited the first one's marker and its first
sight was absorbed - the one thing this bound must never do. The task id is not
the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the
call's own lifecycle - its hold-set stamp and the number of recorded answers -
on an exit 0 and only then, leaving the silent predicate every existing caller
reads unchanged. The throttle scope now carries that identity.

The terminal path recorded the throttle before publishing the durable wake. A
failed append exits the watcher with nothing queued, and the next sighting then
read that fresh marker and absorbed the retry, turning a delayed alarm into a
lost one. Recording moves behind the append, as the non-terminal path already
had it, and the comment claiming the marker could not outlive its wake is gone
because it was false.

The backlog was consulted only on a new terminal pane hash. A captain call can
open after a hash was absorbed as provably working, changing neither the pane nor
the status log, so nothing re-read the backlog and the wedge timer kept firing
possible-wedge alarms through a legitimate wait. That timer now consults the call
at its own alarm boundary and takes the same bounded cadence - and only at that
boundary, so an ordinary repeat poll under the bound stays the local-only read it
was.

Regression coverage for each, all driving churn through one watcher process
rather than relaunching per pane change: relaunch cost dominated the earlier
shape, and an absorbing watcher stays in its poll loop across churn in production
anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer
with no open captain call still escalates as a possible wedge.

* fix(review): Compose stale throttles with captain-call lifecycle identity

* fix(review): Preserve bounded same-hash captain-call resurfacing

* revert(bin): narrow the captain-hold stale bound to its observed defect

Lifts the lifecycle-identity and cadence-ownership work back out, leaving the
change at the shape that matches the defect actually observed: the stale alarm
did not consult the backlog captain hold, on either stale branch.

Reviewing the wider version surfaced a series of adjacent gaps in the watcher's
alarm state machine - a call opening after the first alarm, marker invalidation
at the hold lifecycle boundary, and which deadline a terminal timer represents.
They are real, but fixing them turns a small extension into a state-machine
change to the alarm path, which is a different review on a subsystem that is
being actively reworked. They are named as known limitations rather than carried
here, and none of them is load-bearing for what remains: the bound does strictly
less than the reverted version, leaves the wedge path escalating on
STALE_ESCALATE_SECS exactly as before, and introduces no silence that the
existing terminal-alarm path did not already have.

Kept from the reverted work is the record-after-append ordering, because that is
a defect in the code being shipped rather than an adjacent one: recording the
cadence marker before publishing the durable wake let a failed append lose an
alarm outright instead of delaying it.

History is preserved: the earlier commits stay on the branch and this removal
sits on top of them.

* fix(review): Document secondmate captain-hold scope boundary

* fix(document): Document captain-hold stale alarm scope

* fix(bin): bind the stale throttle to the captain call, not the status log

The throttle this change introduces was scoped to the task's status-log
signature. Answering a call with `--release` and holding the task again creates a
genuinely different captain call without necessarily appending to that log, so
the second call inherited the first one's marker and its first sight was
absorbed.

That is the one alarm this bound must never swallow. A delivery announced twice
is noise; a decision waiting on the captain that is never surfaced is invisible,
because nobody asks for what they do not know to ask for.

Measured rather than assumed, on the same fixture - a delivered task held for the
captain, released, and re-held with no status append, driven through bin/fm-watch.sh:

  base c499f84    call-1 first=ALARM  call-1 churn=ALARM     new call first sight=ALARM
  before this fix call-1 first=ALARM  call-1 churn=absorbed  new call first sight=absorbed
  after           call-1 first=ALARM  call-1 churn=absorbed  new call first sight=ALARM

Base never suppresses the new call, so the suppression came from this change and
closing it completes the fix rather than widening it.

`bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle -
its hold-set stamp and count of recorded answers - on an exit 0 and only then, so
the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope
carries that identity beside the status signature.

The sibling case was measured too and is NOT included: on the status-declared
path, where the last line is `captain-held:`, base already absorbs a re-held
call's first sight. That behaviour predates this change and stays documented as a
known limitation rather than repaired here.

* fix(document): Document captain-call throttle lifecycle scope

* fix(ci): isolate the Herdr restart fixtures from a claimed worktree

The Herdr behaviour test intermittently reused a local worktree still claimed
by an earlier fixture after a restart.
The restart scenarios now use an isolated Treehouse project.

The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as
well.

* feat(pi): resolve extension-registered providers in the supervision branch (#3871)

* Let the supervision branch resolve extension-registered providers

The isolated branch ModelRuntime cannot see providers an extension
registered into main's runtime at run time, so a pin on pi-devin-auth's
devin/swe-1-7 (or an unpinned branch following a main session on devin)
failed with "unavailable to the isolated branch runtime".

Capture main's ModelRegistry alongside mainModel and copy each
extension-registered provider config into the branch runtime at
model-resolution time. The config carries the provider's own streamSimple
and oauth wiring by reference, so the custom gRPC transport reaches the
branch unchanged instead of being reimplemented. The /supervision-model
picker uses the same copy so those models are offered.

Update configuration.md and pi-supervision-branch.md, which previously
stated extension-registered providers were not offered.

* no-mistakes(document): docs: own devin provider carve-out in branch architecture doc

* no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered

* ci: retrigger flaky Herdr/serial-1 lanes

* no-mistakes(document): docs already cover branch extension-provider copy

* fix(bin): gate secondmate wake-loop stall alerts on real queue no-progress (#3943)

* fix(watch): detect stalled secondmate queue progress

* no-mistakes(review): gate secondmate stall on active turns and progress episodes

* no-mistakes(document): align secondmate wake-stall docs with progress-episode detector

* no-mistakes(document): clarify active-turn gate in wake-stall config docs

* no-mistakes(ci): Fixed the one real defect behind the failing checks. ROOT CAUSE (Greptile P1, real code defect in this PR): `secondmate_wake_stall_tick` in bin/fm-watch.sh reset the no-progress timer only when the oldest actionable queue sequence INCREASED (`[ "$seq" -gt "$observed_seq" ]`). When a secondmate is retired and reprovisioned under the same task ID, its fresh home's queue sequence restarts BELOW the recorded position, so the comparison is false, no reset happens, and the new queue inherits the retired generation's already-expired idle interval — emitting a false `secondmate wake-loop stalled` on its very first observation. That is precisely the false-alarm class the user intent requires this PR to remove. FIX (smallest, removal-first): bin/fm-watch.sh:749 now resets when the drain position MOVES AT ALL (`-ne` instead of `-gt`). Draining moves it up, reprovisioning moves it down; neither is a continued no-progress episode. The asymmetric `-gt` branch is removed rather than special-cased or hardened. Updated the function header comment plus the two doc sentences in docs/architecture.md and docs/configuration.md that stated the old advance-only semantics. REGRESSION TEST: added `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` to tests/fm-wake-queue.test.sh (registered in the invocation list). It drives the real watcher through retired generation (seq 9) -> reprovision (seq 3, later clock) -> freeze, asserting observable wake-queue output, no source-text inspection. VERIFICATION: - Fails before / passes after: with the fix reverted the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted`; with the fix it passes, and its third leg confirms the restarted generation still escalates on a genuine freeze (row=3 idle=2s), so the fix does not merely mute the alarm. - `bash tests/fm-wake-queue.test.sh`: exit 0, 38/38 pass, all five secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs (34152740610, 34152740604) both ended with conclusion `action_required` — workflow approval pending, not a test/build failure. Separately, `bin/fm-test-run.sh --check-coverage` exits 1 in this environment, but I confirmed by stashing my changes that it fails identically on the unmodified base tree (locale-related `comm: input is not in sorted order`); it is pre-existing and this change adds no new test file for the partition to account for. Changes are left uncommitted in the worktree

* no-mistakes(ci): Fixed the one real code defect behind the failing checks. ROOT CAUSE (Greptile P1, second round, on the head commit e1304e6): `secondmate_wake_stall_tick` in bin/fm-watch.sh identified the queue's drain position by the sequence number ALONE. The previous round changed the comparison from `-gt` to `-ne`, which handles a reprovisioned queue that restarts BELOW the recorded position, but not one that restarts ON it. A mate retired and reprovisioned under the same task id gets a fresh home whose wake-queue sequence counter restarts at 1 — and the retained parent progress marker very plausibly holds a low sequence too (a queue frozen on its first row records seq 1). Equal sequence ⇒ no reset ⇒ the brand-new queue inherits the retired generation's long-expired idle interval and emits a false `secondmate wake-loop stalled` on its very first observation. That is exactly the false-alarm class this PR exists to remove. FIX (smallest, removal-first): the file already defines the identity of a queue row once, as `row_key="$epoch-$seq"` (used for stall receipts, the stall marker, and the notify key). The progress marker's separate, weaker seq-only identity is removed: `row_key` is now computed once right after the row is parsed, stored in the progress marker, and compared with `!=`. Across generations the epoch differs (the new generation's rows are appended later), so no sequence collision can carry a stale interval; within a generation the key is stable exactly while the position does not move. bin/fm-wake-lib.sh's `fm_wake_secondmate_progress_marker_write` now takes `<oldest-row-key>` and validates it the same way the two neighbouring row-key writers do. Updated the function header comment and the two doc sentences (docs/architecture.md, docs/configuration.md) that described the old sequence-only semantics. REGRESSION TEST: `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` in tests/fm-wake-queue.test.sh now drives the reported case — the reprovisioned generation restarts on the SAME sequence 9 (epoch 200) that the retired generation recorded (epoch 100), at a later clock — and asserts observable watcher output only. Its third leg still confirms the restarted generation escalates on a genuine freeze (row=9 idle=2s), so the fix does not merely mute the alarm. Three seeded progress markers in the symlink, crash-window and prefix-receipt tests were updated to the epoch-sequence form. VERIFICATION: - Fails before / passes after: with bin/fm-watch.sh and bin/fm-wake-lib.sh reverted to HEAD and the new test in place, the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted` (exit 1); with the fix, `bash tests/fm-wake-queue.test.sh` exits 0 with 38/38 pass, all six secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). - I also started `bin/fm-test-run.sh tests/fm-watch-checkpoint.test.sh tests/fm-watch-triage.test.sh tests/fm-watch-recovery-loop.test.sh` as a blast-radius check; it was still running when this phase had to return, so its result is not included. No other suite references the stall detector or the progress marker (grep over tests/ for `wake-loop stall|SECONDMATE_WAKE_STALL|secondmate-wake-progress` matches only fm-wake-queue.test.sh), and the changed lib function has exactly one caller. CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs on the head commit (34154091191, 34154091236, 34154091945) all ended with conclusion `action_required` — pending workflow approval, not a test/build failure. `bin/fm-test-run.sh --check-coverage` still exits 1 in this environment for the pre-existing locale reason recorded in the previous phase (`comm: input is not in sorted order` on the unmodified base tree); this change adds no new test file. Changes are left uncommitted in the worktree: bin/fm-watch.sh, bin/fm-wake-lib.sh, docs/architecture.md, docs/configuration.md, tests/fm-wake-queue.test.sh

---------

Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>

* test(calm): harden the export-DOM render step and record Pi 0.85.1 evidence (#3952)

* fix(tests): make the Calm export-DOM render step retry and report

The Calm suite's rendered-export-DOM assertion started breaking CI with a
bare "could not render calm-mode HTML export DOM", which read like a Pi
0.85 rendering change. It is not one. Calm's rendered rows are identical
across Pi 0.84.4, 0.85.0, and 0.85.1, and the CI break appeared in exactly
one of the thirteen most recent runs, all on the same Pi 0.85.1, with the
main runs immediately before and after it passing.

What actually failed is headless Chrome's start-up. The render step made a
single unattended attempt and discarded both Chrome's stderr and its exit
status, so the log held nothing to tell a Chrome crash apart from a real
change in Pi's export shape.

Rendering is a vendor-tool step; the DOM assertions that follow it are what
protect the Calm conversation boundary. So the step now retries a bounded
number of Chrome start-ups on a fresh profile, drops Chrome's background
network and /dev/shm dependencies without changing what a local file renders
to, and, when every attempt fails, reports the Chrome binary, its version,
the installed Pi version, each attempt's exit status, and Chrome's own
stderr. test_export_dom_render_guard pins that with real processes and no
browser: one clean render, one that only succeeds after a start-up failure,
and one that never renders and must report enough to diagnose itself.

The verification record adds the 0.85.1 evidence this contract is now
pinned to, the cross-version comparison run through isolated installs, and
the Pi 0.85.0 packaging gap - its dist/experimental/server.js statically
imports @earendil-works/pi-server, which 0.85.0 does not declare - that
made the contract look version-sensitive in the first place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y7rr2DHf51MjMvy7sMRauo

* no-mistakes(review): docs: attribute Pi 0.85 calm contract adaptation to renderer change

* no-mistakes(review): tests: drop inert chrome flags, report render timeouts

* no-mistakes(document): docs: fix stale Pi version facts and doc-lint link

---------

Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix(bin): make every counted wake queue row presentable or retired (#3950)

* fix(bin): make every counted wake queue row presentable or retired

A wake row could be counted as queued while no drain wo…
jorguez96 added a commit to jorguez96/firstmate that referenced this pull request Sep 17, 2026
…s resolved (#14)

* fix(bin): defer inactive reconciliation during startup (#3480)

* Defer inactive startup reconciliation

* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably

* no-mistakes(review): Require worker phases to cover startup requests

* no-mistakes(review): Make diagnostic wakes safely acknowledgeable

* no-mistakes(document): Document deferred startup phase coverage

* fix(bin): bound wake drain presentation lock waits (#3475)

* fix: bound status presentation lock waits

* no-mistakes(review): Distinguish malformed presentation locks from live contention

* no-mistakes(review): Bound no-ack drain queue lock acquisition

* no-mistakes(document): Document bounded presentation-lock drain behavior

* no-mistakes(lint): Annotate bounded lock output global

* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite

* fix(bin): retire public follow-ups in remote homes (#3479)

* fix(relay): close a public loop whose work lives in a remote secondmate home

A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.

The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.

Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.

Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.

* no-mistakes(review): Guard remote link clears by request identity

* no-mistakes(review): Fail guarded clears on unreadable remote state

* no-mistakes(review): Reject guarded clears on non-writable remote state

* no-mistakes(review): Allow no-link retirement in non-writable remote state

* no-mistakes(document): Correct public-followup verification guarantee count

* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh

* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint

* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks

* fix(relay): bound the guarded remote link clear so it refuses instead of hanging

The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.

The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.

The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.

The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.

* no-mistakes(review): Harden lock-timeout regression with independent deadline

* no-mistakes(review): Restore no-op guarded clears on read-only state

* no-mistakes(document): Clarify remote public-followup cleanup contract

* fix(bin): support process events under symlinked homes (#3484)

* fix(bin): resolve process-event state roots before validating them

The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.

Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.

This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.

* fix(bin): pin the external capture staging boundary to its physical path

The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.

The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.

* no-mistakes(review): Propagate canonical process-event state roots

* no-mistakes(review): Propagate canonical state to process-event adapters

* no-mistakes(document): Document physical process-event state roots

* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)

* fix(pi): persist captain outcomes visibly

* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition

* no-mistakes(document): Document cold-start captain-outcome recovery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(pi): process captain outcomes through a sequence-keyed turn

PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.

The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.

Add the processing half on top of the persistence half:

- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
  cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
  only advances through an explicit sequence-bound acknowledgement, never
  past the read cursor and never backwards; an absent marker reads as zero
  and `processed-init` migrates delivered history once so an upgraded home
  is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
  every still-unprocessed captain row to main as one hidden, typed
  `fm-branch-process` request listing each `[seq N] task: summary`, opening
  exactly one main turn. Main closes it only by calling the new
  `fm_branch_processed` tool with the highest sequence listed. An unrelated,
  empty, or paraphrased answer leaves the sequence open, and the same request
  is presented again at the end of the next main run and at session start.
  The first two presentations of a sequence set open a turn of their own;
  after that the request rides the captain's next prompt so an ignored
  request cannot loop, and a session replacement resets that budget.
  Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
  scripts: an empty answer and an unrelated prior answer neither advance the
  marker nor stop re-presentation, the acknowledgement is refused beyond the
  read cursor and outside lock ownership, a partial acknowledgement keeps the
  newer sequence open, and #3312's own assertions now forbid an unkeyed turn
  rather than any turn. The store suite pins the marker's bounds and the
  migration; the real-SDK guard for appendEntry persistence and model
  exclusion is unchanged.

Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.

* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements

* no-mistakes(review): Harden outcome state validation and request pacing

* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores

* no-mistakes(review): Validate canonical mark-read cursor state

* no-mistakes(review): Guard cursor advancement against corrupt processed state

* no-mistakes(review): Bind acknowledgements to active processing requests

* no-mistakes(review): Reset pacing when processing sequence membership changes

* no-mistakes(review): Enforce silent outcome invariants at storage boundary

* no-mistakes(document): Document hardened captain outcome processing contracts

---------

Co-authored-by: kunchenguid <kun@kunchenguid.com>

* feat: add bounded concurrent Bearings ledger collection (#3481)

* feat: bound Bearings remote ledger collection

* no-mistakes(review): Clarify default remote-ledger collection behavior

* no-mistakes(review): Detach reconcile delivery from watcher loop

* no-mistakes(review): Enforce bounded snapshot and request captures

* no-mistakes(review): Bound legacy summary capture before parsing

* no-mistakes(review): Bound primary remote ledger captures

* no-mistakes(document): Correct snapshot and reconcile documentation

* no-mistakes(lint): Fix ShellCheck quoting in bounded collector

* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks

* test: await reconcile request retirement

* no-mistakes(review): Avoid empty reconcile queue process churn

* no-mistakes(review): Read ledger summaries from immutable snapshots

* no-mistakes(review): Reject multi-document home ledger streams

* no-mistakes(review): Coalesce durable reconcile requests per target

* no-mistakes(review): Unify reconcile keys and reject snapshot streams

* no-mistakes(review): Key reconcile requests by stable target ID

* no-mistakes(document): Document per-target reconcile request coalescing

* no-mistakes(lint): Remove unused snapshot summary file variable

* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass

* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks

* ci: rebalance portable serial test shards (#3489)

* fix(ci): rebalance the portable serial shards on measured durations

The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.

Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.

Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.

Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.

No test changes what it asserts and no test stops running; only the
partition across shards changes.

* no-mistakes(document): Clarify conservative shard timing aggregate

* fix(pi): fall back on incomplete supervision branch prompts (#3491)

* fix(pi): fall back after settled branch errors

* no-mistakes(review): Detect provider errors across prompt compaction

* no-mistakes(review): Preserve in-flight branch state across selection changes

* fix(pi): re-probe supervision branch after cooldown (#3497)

* fix(pi): recover supervision branch after cooldown

* no-mistakes(review): Defer branch recovery until prompt settlement

* no-mistakes(document): Clarify supervision cooldown recovery contract

* fix(bin): remove legacy remote snapshot reads (#3501)

* refactor: remove legacy remote summary reads

* no-mistakes(document): Document ledger-only snapshot reads

* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass

* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean

* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux

* fix(pi): preserve watcher continuity across session replacement (#3498)

* fix(pi): rearm watcher after session replacement

* no-mistakes(review): Queue actionable closes across Pi session replacement

* no-mistakes(review): Stop replacement arm when handoff persistence fails

* no-mistakes(review): Preserve actionable wakes through branch and late child races

* no-mistakes(review): Surface late handoff failures without crashing Pi

* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens

* no-mistakes(review): Retry stale deliveries and release settled claims

* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup

* no-mistakes(review): Deduplicate persistent handoff cleanup alerts

* no-mistakes(review): Acknowledge watcher follow-ups only when consumed

* no-mistakes(review): Persist idle follow-ups until agent consumption

* no-mistakes(review): Preserve pending outcomes when handoff persistence fails

* no-mistakes(review): Arm replacement before awaiting prior delivery settlement

* no-mistakes(review): Adopt pending handoffs after lock reclamation

* no-mistakes(review): Prevent stale generations from adopting replacement handoffs

* no-mistakes(review): Scope replacement handoffs by watcher state

* no-mistakes(document): Clarify replacement handoff documentation

* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks

* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes

* no-mistakes(document): Document watcher-owned replacement handoffs

* no-mistakes(document): Verify replacement handoff documentation

* test(pi): cover watcher-owned branch fallback

* no-mistakes(document): Refresh watcher-owned fallback documentation

* fix(bin): resurface task statuses missed by wake handling (#3495)

* fix(bin): resurface terminal statuses lost after branch handling

* test(watch): canonicalize process-event fixture homes

* no-mistakes(review): Index branch outcomes by causal status position

* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses

* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics

* no-mistakes(review): Keep unclassifiable oversized statuses silent

* no-mistakes(document): Document lost-wake outcome backstop

* no-mistakes(document): Update outcome backstop documentation

* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally

* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes

* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift

* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state

* fix(bin): collect follow-up results from remote work homes (#3503)

* fix(bin): deliver typed terminal results from remote work homes

A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.

The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.

A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.

This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.

* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes

* no-mistakes(review): Fail collection when remote outbox is unreadable

* no-mistakes(review): Surface reassigned remote routes during empty collection

* no-mistakes(review): Fail remote collection on invalid registrations

* no-mistakes(review): Reject unsafe registration entries during remote collection

* no-mistakes(review): Restore healthy empty remote collection behavior

* no-mistakes(review): Skip remote collection for delivered registrations

* no-mistakes(review): Skip delivered registrations before route validation

* no-mistakes(document): Document remote follow-up collection semantics

* fix(bin): exclude secondmates from home-summary validity (#3504)

* fix(bin): exclude secondmates from home-summary child inventory

kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.

* no-mistakes(review): Cover terminal secondmate in-flight exclusion

* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds

* fix(bin): self-heal outcome indexes on first drain (#3509)

* fix(bin): self-heal status-outcome indexes on every drain

Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.

* no-mistakes(review): Guard held-lock initialization and fail marker writes

* no-mistakes(document): Document cross-harness outcome-index self-healing

* fix(bearings): keep active children underway during captain holds (#3505)

* fix(bearings): keep active children underway beside a captain hold

Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.

* no-mistakes(review): Preserve Underway repos and disclose child truncation

* no-mistakes(review): Fall back to task project for Underway repos

* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean

* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)

* fix(pi): settle watcher delivery on Pi accepting the follow-up

A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.

The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.

Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(pi): retry a verified successor that fails during wake delivery

A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.

The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.

The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(bin): bound repeat stale wakes for parked workers (#3532)

* fix(bin): bound repeat stale wakes for a parked but live worker

A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.

pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.

Two places let that churn re-alarm:

- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
  declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
  have suppressed it. The throttle was never read on this path and was advanced
  by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
  whenever the classification came back `none`, so each tick also bought the same
  declared wait a fresh window. Fixing only the first site changes nothing.

Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.

First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.

Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.

* fix(document): Clarify declared-wait wake cadence documentation

* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor

* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed

* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)

* fix(turnend): accept the away-mode daemon as the supervision owner

While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.

Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.

The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.

The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.

* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage

* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md

* fix(backlog): omit --file from row probes for non-markdown backends (#3582)

* fix(backlog): omit markdown file for beads probes

* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes

* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)

* fix(bin): classify progress updates on requested work as routine (#3589)

The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.

The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.

* fix(bin): preserve captain calls during teardown (#3595)

* fix(bin): never close a captain call during cleanup

A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.

bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.

Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.

The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.

Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.

Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np

* no-mistakes(review): Serialize captain holds and fix backend-aware listing

* no-mistakes(document): Update captain-call retention documentation

* no-mistakes(document): Fix relocated captain-hold backlog diagnostics

* fix(bin): deliver secondmate outcomes to the parent channel (#3592)

* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts

A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:

- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
  exact-line append-once; the merge outcome path and the inactive-outcome
  scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
  watcher poll in a secondmate home: a direct child's whole terminal done or
  failed line is delivered at once with its note, recorded PR, mode, merge
  posture, and scout report pointer, keyed and receipted so it is delivered
  once, and the inactive path yields to it. `report <task-id>` runs the same
  delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
  registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
  and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
  record and refuses, retaining every record, while the channel cannot be
  written.
- The charter opens with the parent-channel rule and confines the mate's own
  appends to judgement; AGENTS.md carries the carve-out at the persona
  address rule and the escalation list.

docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.

* no-mistakes(review): Fix parent outcome retries and reconciliation locking

* no-mistakes(review): Prevent busy children from starving ledger delivery

* no-mistakes(review): Correct ledger metadata and hold occurrence handling

* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons

* no-mistakes(review): Close ledger races and preserve teardown records

* no-mistakes(document): Correct parent-channel receipt and scanner documentation

* no-mistakes(lint): Quote done arguments for ShellCheck compliance

* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks

* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure

* fix(bin): sync remote second mates to primary commit (#3599)

* fix(bin): sync remote second-mate homes to the parent primary commit

Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.

The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.

The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.

/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.

* no-mistakes(document): Document primary-targeted remote secondmate synchronization

* fix(bin): separate captain intent from firstmate specs (#3597)

* fix(bin): split brief task into captain intent and firstmate spec

Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.

* fix(bin): stop task-subsection copies at the next heading

Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.

* no-mistakes(review): Validate brief content and preserve nested specifications

* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies

* no-mistakes(review): Ignore fenced subsection headings during brief validation

* no-mistakes(review): Preserve captain intent across scout promotion

* no-mistakes(review): Enforce safe intent boundaries for legacy promotions

* no-mistakes(review): Allow marked legacy intent and reject empty promotions

* no-mistakes(review): Scope task parsing and overlay legacy intent contracts

* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns

* no-mistakes(review): Preserve later captain clarifications in intent overlays

* no-mistakes(document): Document brief intent enforcement and ownership

* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed

* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint

* fix: start a fresh supervision branch for every main session (#3600)

* fix(pi): start a new supervision branch conversation per main session

The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.

The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.

The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.

* no-mistakes(document): Document fresh Pi supervision conversations

* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks

* feat: restart second mates after instruction updates (#3614)

* feat(update): restart second mates whose instructions changed

/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.

An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.

Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.

fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.

Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.

* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting

* no-mistakes(review): Parallelize relaunches and classify replacement incarnations

* no-mistakes(review): Gate restart actions on live agent state

* no-mistakes(review): Handle failed restart workers without hanging

* no-mistakes(review): Nudge legacy remotes and preserve persist recovery

* no-mistakes(review): Document one-time secondmate restart rollout

* no-mistakes(review): Honor arrived replies and refresh remote profiles

* no-mistakes(review): Revert remote parent profile reconciliation

* no-mistakes(review): Reset remote profile defaults and honor published results

* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates

* no-mistakes(document): Document second-mate restart update flow

* no-mistakes(lint): Fix ShellCheck warnings in restart scripts

* perf: accelerate local validation with bounded concurrency (#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation

* fix: copy PR URLs from durable records (#3648)

* fix: copy PR URLs from records or abstain, never assemble them

Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.

Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:

- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
  or abstain" section requires a URL to be copied verbatim from a durable
  record (the done: PR <url> status line, pr= metadata, or the backlog note),
  forbids assembling owner, repository, host, or number from memory, and has
  the branch report only the identifier it actually holds when no record names
  the URL yet, leaving the PR check unarmed until the worker's ready line
  arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for
  main in place of the bare full-URL mandate.

- Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full
  https:// URL wherever a PR is mentioned - status line, terminal, or summary -
  never a bare "PR 108", so the link is in view as early as the number is.

- bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that
  the task's own done lines contradict, printing both spellings; a log naming
  no URL still records the argument as before. fm_pr_status_ready_urls in
  bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches
  bin/fm-pr-merge.sh, so nothing merges under a contradicted URL.

Tests cover the offline refusal with zero side effects, the recorded spelling
being accepted, markdown-wrapped and punctuated URLs, working lines not
counting, the merge wrapper propagation, a self-hosted merge request with no
forge call, the prompt carrying the rule, and the brief carrying the worker
rule.

* no-mistakes(review): Remove stale PR URL enforcement

* no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`

* fix(bin): disable Claude feedback drafts for fleet launches (#3661)

* fix(bin): disable Claude's feedback-draft flow for fleet-launched agents

Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched
Claude crewmate and secondmate, so /bug and /feedback never queue or
submit a bug report on the captain's behalf. feedbackDrafts is the
documented settings key (Claude Code changelog 2.1.247); the
per-launch CLI flag never touches the captain's global settings.json.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts

* no-mistakes(document): Fix Claude feedback documentation formatting

* fix(bin): layer both feedback-draft controls for defense in depth

The prior --settings-only fix can be overridden by a managed Claude
settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0
alongside --settings '{"feedbackDrafts":"off"}': either control alone
disables the SendFeedback tool, so a managed override of one still
leaves the other in force.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(document): Document Claude feedback-draft suppression ownership

* feat(tests): run three more validation families concurrently (#3662)

* perf(tests): admit three more families to concurrent validation

The three families that `docs/fm-test-isolation-proof.md` recorded as refused
were not refused for concurrency. Each blocker was a test that decided a
property by wall clock, or a script filed where it cannot run. Fixing those
three things admits all three families and recovers 28.6 minutes of local
validation with no assertion removed or weakened.

- `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the
  handoff, sleeping a fixed second, then delegating the move to the real
  binary. Nothing ever killed the fake, so on a host slow enough for the case's
  next assertions to take longer than a second, the orphan woke and completed
  the very move the case requires left undone, and recovery then failed with
  `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs
  during the injected crash showed exactly that, the item moving one second
  after the crash. All four crash injections in the file now go through a new
  `fm_fake_crash_injector` shim that signals the target and returns only once
  it is observably gone, and the pre-move fake never delegates the move at all.

- `tests/fm-session-start.test.sh` proved the startup digest does not block on
  a slow current-state read by timing the whole digest against a fixed
  eight-second sleep, which a loaded host exceeds without the property being
  violated. It now holds that read open until the case releases it and asserts,
  the moment the digest returns, that the read has not finished. A digest that
  waited would wait indefinitely rather than for an interval a slow host can
  out-run, so the assertion is stronger than the bound it replaces. Its scan
  budget moves to the maximum, because the old value left two seconds of margin
  over the fixed sleep and measured the host rather than the deadline that
  `tests/fm-inactive-reconcile.test.sh` owns.

- `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all,
  which put it in the portable serial lane, where Linux CI gate-skips it: that
  real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and
  the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its
  opt-in variable and moves to `live-harness-optin`.

The 28 remaining ungrouped scripts become an enumerated `standalone` family
instead of admitting `unclassified` itself. `unclassified` is the family map's
`*)` arm, so admitting it would silently grant concurrency to every test added
afterwards, which is exactly the population with no proof. A new test still
lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh`
covers that split behaviorally.

Each family passes two consecutive four-worker proofs with zero failures. On
the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap`
756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s
recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock
against 121 minutes of summed script time.

* no-mistakes(document): Refresh concurrent validation and shard documentation

* no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2

* feat: structure no-mistakes ask-user escalations (#3670)

* feat(brief): structure no-mistakes ask-user escalation as event + snapshot file

Crewmates escalating a no-mistakes ask-user gate now report one status
event naming every finding id plus a snapshot file holding the gate's
axi finding records verbatim (id, severity, file, line, description,
authority), using the same shape even for a single finding. The status
line never paraphrases. The format is defined once in fm-dod-lib.sh and
rendered into both the scout and ship rule 6 in fm-brief.sh, so a
promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets
the identical contract as a freshly-spawned no-mistakes ship worker.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei

* no-mistakes(review): Preserve ask-user escalation output contract

* no-mistakes(review): Align escalation format test expectation

* no-mistakes(review): Scope ask-user escalation instructions correctly

* no-mistakes(review): Remove ask-user from generic decision rules

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix(bin): require self-sufficient no-mistakes intent (#3671)

* fix(bin): require a self-sufficient no-mistakes intent

A no-mistakes worker's --intent is only as useful as the string it
passes. PR #3604 shipped with an intent that was only "do 1, 2, 3, 7
from the report": the real contract lived in a private scout report and
never reached --intent, so nobody holding that string plus the codebase
could have derived the specification.

This is pure instruction at the contract's one owner; no spawn-side or
promotion-side check is added.

- bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now
  states that the --intent string must be self-sufficient (the string
  plus the codebase reconstructs roughly the same specification) and
  tells the worker to write the substance of any report, decision, or
  PR the captain's intent refers to into --intent rather than the
  pointer, while Firstmate build instructions and the worker's own
  decisions still stay out. The spawn-time overlay points back at that
  rule so its "supersedes" wording cannot cancel it, and the header's
  owner statement carries the rule.
- AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to
  include the substance of referenced material when filling
  ## Captain's intent, and section 11 points at the owner of the rule.
- tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the
  rendered brief and launch contract carry the rule.

Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim

* no-mistakes(document): Replace incident-specific intent test commentary

* fix: accelerate local Bearings snapshot composition (#3499)

* Speed local fleet snapshot composition

* no-mistakes(review): Stabilize task inventory during concurrent snapshot composition

* no-mistakes(document): Document local snapshot observation concurrency

* no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks

* no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash s…
sctru added a commit to sctru/firstmate that referenced this pull request Sep 19, 2026
* fix(bin): safely unregister custom checks (#3369)

* fix(bin): add a safe owner for custom-check retirement

Agents were improvising rm of check files with unset STATE/ID, which wedges
headless panes. Unregister validates the id and state directory first.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Refuse explicitly empty custom-check state overrides

* no-mistakes(document): Document custom-check retirement safety contract

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(quota): extract mid-task polling and candidate selection into dedicated scripts (#3221)

* Add quota exhaustion detection and safe fallback helpers

- bin/fm-procevent-quota.sh: generic procevent adapter that arms a
  recurring quota-axi --json poll and wakes firstmate when a tracked
  provider's effectivePercentRemaining drops below a threshold or its
  runway.status becomes exhausted_now.
- bin/fm-quota-choose.sh: worker-side helper that picks the first ranked
  harness:model candidate with positive effectivePercentRemaining.
- AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document
  the new helpers and the mid-task quota-exhaustion wake path.
- tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON
  source.

* no-mistakes(review): Fix quota polling and scope bounds

* no-mistakes(review): Enforce safe default quota selection

* no-mistakes(review): Handle decimal quota values safely

* no-mistakes(review): Fail closed on invalid quota inputs

* no-mistakes(review): Reject empty quota candidate segments

* no-mistakes(review): Harden quota parsing and timeout ownership

* no-mistakes(review): Reuse captured quota snapshots consistently

* no-mistakes(review): Match quota using explicit candidate providers

* no-mistakes(review): Centralize fail-closed quota schema validation

* no-mistakes(review): Reject out-of-range quota percentages

* no-mistakes(review): Validate quota runway status enum

* no-mistakes(review): Tighten quota scope and status contracts

* no-mistakes(review): Preserve unknown quota and exact product bounds

* no-mistakes(review): Preserve provider-level unknown quota

* no-mistakes(review): Reuse canonical verified harness validation

* no-mistakes(document): Document mid-task quota handling

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(docs): restore default routing contract, keep quota helper optional

Restore the AGENTS.md section 4 always-loaded routing paragraph the PR
had deleted, so the standing TOON-first intake, spendPriority ranker,
every-candidate accounting, and load-trigger contract stay exactly as
before this PR. The mid-task quota wake is optional and must not alter
default routing.

Restore the quota-array-dispatch skill ownership line to section 4 as
the always-loaded intake boundary owner; keep the worker-side helper
section as an addition only, without rewiring ownership or load
triggers to section 13.

* fix(bin): use harness-keyed quota matching in optional helper

Revert fm-quota-choose.sh from harness:provider:model tuples back to
harness:model candidates with harness-keyed provider matching, per the
resolved ask-user finding. The helper is optional; authoritative
multi-provider routing (provider discovery from the harness catalog and
quota matching by that explicit provider) stays owned by AGENTS.md
section 4 and the quota-array-dispatch skill intake procedure, not the
helper.

Document the multi-provider limitation in the helper header and the
quota-array-dispatch skill: the helper maps each harness to one primary
provider family only, so a candidate whose established provider differs
from that primary family is checked against the wrong quota row. Use it
only when the brief fixed the candidate order and every candidate's
provider is the harness's primary family.

The helper still consumes one already-captured default-TOON or JSON
snapshot via stdin or --snapshot and never calls quota-axi itself, so
it selects from the same quota state as the intake.

* no-mistakes(review): Fix Muse quota mapping and helper contract docs

* no-mistakes(review): Reject known-empty quotas and map quota tests explicitly

* no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse

* no-mistakes(review): Fix quota retirement and dependent regression coverage

* no-mistakes(review): Accept zero-row quota TOON snapshots

* no-mistakes(review): Enforce quota semantics status consistency

* no-mistakes(review): Veto dispatch on any exhausted applicable scope

* no-mistakes(review): Record exhausted quota scope in wake details

* no-mistakes(review): Fix quota help and control dependency coverage

* no-mistakes(review): Decode quoted TOON fields and document quota wakes

* no-mistakes(review): Validate zero-row TOON and map timeout coverage

* no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes

* no-mistakes(review): Validate complete nonzero TOON envelopes

* no-mistakes(review): Accept producer-shaped quota TOON envelopes

* no-mistakes(review): Support empty quota arrays and validate counted rows

* no-mistakes(review): Harden TOON completion, scopes, and quoted fields

* no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields

* no-mistakes(review): Allow unknown headroom under known semantics

* no-mistakes(review): Reject noncanonical quota identities

* no-mistakes(review): Preserve empty quota polling and validate attention identities

* no-mistakes(review): Reject noncanonical provider watches

* no-mistakes(review): Validate all candidates before quota selection

* no-mistakes(document): Correct quota helper safety documentation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix: surface comments on Lavish annotations (#3371)

* fix(bin): keep typed Lavish comments when an element is also annotated

read preferred element text over prompt, so an annotate-and-comment
item dropped the captain's words. Surface prompt as its own field.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Filter non-comment prompts from Lavish reader output

* no-mistakes(document): Clarify Lavish comment presentation contract

* no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure

* no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure

* fix(bin): always emit Lavish comments and use real annotation fixtures

Stop inferring comment provenance from prompt==text. Real pure
annotations have no prompt, so always-emit does not duplicate.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: support first public-followup registration on Bash 3.2 (#3420)

* Fix public-followup register crashing on empty lock arrays under bash 3.2.

bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock.
The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there.

* no-mistakes(document): Document stock Bash registration coverage

* no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5

* no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression

* fix(bin): isolate new Herdr server environments (#2792)

* fix(herdr): isolate server launch environment

* no-mistakes(review): Clear inherited supervision model from Herdr launches

* no-mistakes(document): Document Herdr server launch environment isolation

* fix: surface inbound Relay media to responding agents (#3442)

* fix: surface inbound Relay attachments to the responding agent

A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.

Fix it where the gap is, in prose:

- Read the complete payload object rather than a fixed field list, so
  media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
  mention and on every chain entry, and call out the common shape where
  only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
  (Discord: cdn.discordapp.com, media.discordapp.net,
  images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
  pbs.twimg.com, video.twimg.com), report a blocked host instead of
  working around it, and treat everything fetched as untrusted public
  input on the same terms as the surrounding thread text.

The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.

The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.

* no-mistakes(review): Preserve media authority and enforce poll-only fetching

* no-mistakes(document): Clarify Relay attachment safety prose

* fix(bin): defer inactive reconciliation during startup (#3480)

* Defer inactive startup reconciliation

* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably

* no-mistakes(review): Require worker phases to cover startup requests

* no-mistakes(review): Make diagnostic wakes safely acknowledgeable

* no-mistakes(document): Document deferred startup phase coverage

* fix(bin): bound wake drain presentation lock waits (#3475)

* fix: bound status presentation lock waits

* no-mistakes(review): Distinguish malformed presentation locks from live contention

* no-mistakes(review): Bound no-ack drain queue lock acquisition

* no-mistakes(document): Document bounded presentation-lock drain behavior

* no-mistakes(lint): Annotate bounded lock output global

* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite

* fix(bin): retire public follow-ups in remote homes (#3479)

* fix(relay): close a public loop whose work lives in a remote secondmate home

A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.

The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.

Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.

Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.

* no-mistakes(review): Guard remote link clears by request identity

* no-mistakes(review): Fail guarded clears on unreadable remote state

* no-mistakes(review): Reject guarded clears on non-writable remote state

* no-mistakes(review): Allow no-link retirement in non-writable remote state

* no-mistakes(document): Correct public-followup verification guarantee count

* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh

* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint

* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks

* fix(relay): bound the guarded remote link clear so it refuses instead of hanging

The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.

The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.

The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.

The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.

* no-mistakes(review): Harden lock-timeout regression with independent deadline

* no-mistakes(review): Restore no-op guarded clears on read-only state

* no-mistakes(document): Clarify remote public-followup cleanup contract

* fix(bin): support process events under symlinked homes (#3484)

* fix(bin): resolve process-event state roots before validating them

The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.

Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.

This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.

* fix(bin): pin the external capture staging boundary to its physical path

The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.

The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.

* no-mistakes(review): Propagate canonical process-event state roots

* no-mistakes(review): Propagate canonical state to process-event adapters

* no-mistakes(document): Document physical process-event state roots

* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)

* fix(pi): persist captain outcomes visibly

* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition

* no-mistakes(document): Document cold-start captain-outcome recovery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(pi): process captain outcomes through a sequence-keyed turn

PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.

The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.

Add the processing half on top of the persistence half:

- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
  cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
  only advances through an explicit sequence-bound acknowledgement, never
  past the read cursor and never backwards; an absent marker reads as zero
  and `processed-init` migrates delivered history once so an upgraded home
  is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
  every still-unprocessed captain row to main as one hidden, typed
  `fm-branch-process` request listing each `[seq N] task: summary`, opening
  exactly one main turn. Main closes it only by calling the new
  `fm_branch_processed` tool with the highest sequence listed. An unrelated,
  empty, or paraphrased answer leaves the sequence open, and the same request
  is presented again at the end of the next main run and at session start.
  The first two presentations of a sequence set open a turn of their own;
  after that the request rides the captain's next prompt so an ignored
  request cannot loop, and a session replacement resets that budget.
  Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
  scripts: an empty answer and an unrelated prior answer neither advance the
  marker nor stop re-presentation, the acknowledgement is refused beyond the
  read cursor and outside lock ownership, a partial acknowledgement keeps the
  newer sequence open, and #3312's own assertions now forbid an unkeyed turn
  rather than any turn. The store suite pins the marker's bounds and the
  migration; the real-SDK guard for appendEntry persistence and model
  exclusion is unchanged.

Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.

* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements

* no-mistakes(review): Harden outcome state validation and request pacing

* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores

* no-mistakes(review): Validate canonical mark-read cursor state

* no-mistakes(review): Guard cursor advancement against corrupt processed state

* no-mistakes(review): Bind acknowledgements to active processing requests

* no-mistakes(review): Reset pacing when processing sequence membership changes

* no-mistakes(review): Enforce silent outcome invariants at storage boundary

* no-mistakes(document): Document hardened captain outcome processing contracts

---------

Co-authored-by: kunchenguid <kun@kunchenguid.com>

* feat: add bounded concurrent Bearings ledger collection (#3481)

* feat: bound Bearings remote ledger collection

* no-mistakes(review): Clarify default remote-ledger collection behavior

* no-mistakes(review): Detach reconcile delivery from watcher loop

* no-mistakes(review): Enforce bounded snapshot and request captures

* no-mistakes(review): Bound legacy summary capture before parsing

* no-mistakes(review): Bound primary remote ledger captures

* no-mistakes(document): Correct snapshot and reconcile documentation

* no-mistakes(lint): Fix ShellCheck quoting in bounded collector

* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks

* test: await reconcile request retirement

* no-mistakes(review): Avoid empty reconcile queue process churn

* no-mistakes(review): Read ledger summaries from immutable snapshots

* no-mistakes(review): Reject multi-document home ledger streams

* no-mistakes(review): Coalesce durable reconcile requests per target

* no-mistakes(review): Unify reconcile keys and reject snapshot streams

* no-mistakes(review): Key reconcile requests by stable target ID

* no-mistakes(document): Document per-target reconcile request coalescing

* no-mistakes(lint): Remove unused snapshot summary file variable

* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass

* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks

* ci: rebalance portable serial test shards (#3489)

* fix(ci): rebalance the portable serial shards on measured durations

The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.

Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.

Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.

Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.

No test changes what it asserts and no test stops running; only the
partition across shards changes.

* no-mistakes(document): Clarify conservative shard timing aggregate

* fix(pi): fall back on incomplete supervision branch prompts (#3491)

* fix(pi): fall back after settled branch errors

* no-mistakes(review): Detect provider errors across prompt compaction

* no-mistakes(review): Preserve in-flight branch state across selection changes

* fix(pi): re-probe supervision branch after cooldown (#3497)

* fix(pi): recover supervision branch after cooldown

* no-mistakes(review): Defer branch recovery until prompt settlement

* no-mistakes(document): Clarify supervision cooldown recovery contract

* fix(bin): remove legacy remote snapshot reads (#3501)

* refactor: remove legacy remote summary reads

* no-mistakes(document): Document ledger-only snapshot reads

* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass

* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean

* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux

* fix(pi): preserve watcher continuity across session replacement (#3498)

* fix(pi): rearm watcher after session replacement

* no-mistakes(review): Queue actionable closes across Pi session replacement

* no-mistakes(review): Stop replacement arm when handoff persistence fails

* no-mistakes(review): Preserve actionable wakes through branch and late child races

* no-mistakes(review): Surface late handoff failures without crashing Pi

* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens

* no-mistakes(review): Retry stale deliveries and release settled claims

* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup

* no-mistakes(review): Deduplicate persistent handoff cleanup alerts

* no-mistakes(review): Acknowledge watcher follow-ups only when consumed

* no-mistakes(review): Persist idle follow-ups until agent consumption

* no-mistakes(review): Preserve pending outcomes when handoff persistence fails

* no-mistakes(review): Arm replacement before awaiting prior delivery settlement

* no-mistakes(review): Adopt pending handoffs after lock reclamation

* no-mistakes(review): Prevent stale generations from adopting replacement handoffs

* no-mistakes(review): Scope replacement handoffs by watcher state

* no-mistakes(document): Clarify replacement handoff documentation

* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks

* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes

* no-mistakes(document): Document watcher-owned replacement handoffs

* no-mistakes(document): Verify replacement handoff documentation

* test(pi): cover watcher-owned branch fallback

* no-mistakes(document): Refresh watcher-owned fallback documentation

* fix(bin): resurface task statuses missed by wake handling (#3495)

* fix(bin): resurface terminal statuses lost after branch handling

* test(watch): canonicalize process-event fixture homes

* no-mistakes(review): Index branch outcomes by causal status position

* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses

* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics

* no-mistakes(review): Keep unclassifiable oversized statuses silent

* no-mistakes(document): Document lost-wake outcome backstop

* no-mistakes(document): Update outcome backstop documentation

* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally

* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes

* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift

* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state

* fix(bin): collect follow-up results from remote work homes (#3503)

* fix(bin): deliver typed terminal results from remote work homes

A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.

The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.

A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.

This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.

* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes

* no-mistakes(review): Fail collection when remote outbox is unreadable

* no-mistakes(review): Surface reassigned remote routes during empty collection

* no-mistakes(review): Fail remote collection on invalid registrations

* no-mistakes(review): Reject unsafe registration entries during remote collection

* no-mistakes(review): Restore healthy empty remote collection behavior

* no-mistakes(review): Skip remote collection for delivered registrations

* no-mistakes(review): Skip delivered registrations before route validation

* no-mistakes(document): Document remote follow-up collection semantics

* fix(bin): exclude secondmates from home-summary validity (#3504)

* fix(bin): exclude secondmates from home-summary child inventory

kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.

* no-mistakes(review): Cover terminal secondmate in-flight exclusion

* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds

* fix(bin): self-heal outcome indexes on first drain (#3509)

* fix(bin): self-heal status-outcome indexes on every drain

Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.

* no-mistakes(review): Guard held-lock initialization and fail marker writes

* no-mistakes(document): Document cross-harness outcome-index self-healing

* fix(bearings): keep active children underway during captain holds (#3505)

* fix(bearings): keep active children underway beside a captain hold

Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.

* no-mistakes(review): Preserve Underway repos and disclose child truncation

* no-mistakes(review): Fall back to task project for Underway repos

* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean

* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)

* fix(pi): settle watcher delivery on Pi accepting the follow-up

A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.

The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.

Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(pi): retry a verified successor that fails during wake delivery

A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.

The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.

The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(bin): bound repeat stale wakes for parked workers (#3532)

* fix(bin): bound repeat stale wakes for a parked but live worker

A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.

pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.

Two places let that churn re-alarm:

- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
  declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
  have suppressed it. The throttle was never read on this path and was advanced
  by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
  whenever the classification came back `none`, so each tick also bought the same
  declared wait a fresh window. Fixing only the first site changes nothing.

Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.

First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.

Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.

* fix(document): Clarify declared-wait wake cadence documentation

* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor

* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed

* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)

* fix(turnend): accept the away-mode daemon as the supervision owner

While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.

Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.

The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.

The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.

* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage

* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md

* fix(backlog): omit --file from row probes for non-markdown backends (#3582)

* fix(backlog): omit markdown file for beads probes

* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes

* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)

* fix(bin): classify progress updates on requested work as routine (#3589)

The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.

The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.

* fix(bin): preserve captain calls during teardown (#3595)

* fix(bin): never close a captain call during cleanup

A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.

bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.

Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.

The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.

Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.

Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np

* no-mistakes(review): Serialize captain holds and fix backend-aware listing

* no-mistakes(document): Update captain-call retention documentation

* no-mistakes(document): Fix relocated captain-hold backlog diagnostics

* fix(bin): deliver secondmate outcomes to the parent channel (#3592)

* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts

A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:

- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
  exact-line append-once; the merge outcome path and the inactive-outcome
  scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
  watcher poll in a secondmate home: a direct child's whole terminal done or
  failed line is delivered at once with its note, recorded PR, mode, merge
  posture, and scout report pointer, keyed and receipted so it is delivered
  once, and the inactive path yields to it. `report <task-id>` runs the same
  delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
  registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
  and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
  record and refuses, retaining every record, while the channel cannot be
  written.
- The charter opens with the parent-channel rule and confines the mate's own
  appends to judgement; AGENTS.md carries the carve-out at the persona
  address rule and the escalation list.

docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.

* no-mistakes(review): Fix parent outcome retries and reconciliation locking

* no-mistakes(review): Prevent busy children from starving ledger delivery

* no-mistakes(review): Correct ledger metadata and hold occurrence handling

* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons

* no-mistakes(review): Close ledger races and preserve teardown records

* no-mistakes(document): Correct parent-channel receipt and scanner documentation

* no-mistakes(lint): Quote done arguments for ShellCheck compliance

* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks

* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure

* fix(bin): sync remote second mates to primary commit (#3599)

* fix(bin): sync remote second-mate homes to the parent primary commit

Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.

The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.

The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.

/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.

* no-mistakes(document): Document primary-targeted remote secondmate synchronization

* fix(bin): separate captain intent from firstmate specs (#3597)

* fix(bin): split brief task into captain intent and firstmate spec

Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.

* fix(bin): stop task-subsection copies at the next heading

Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.

* no-mistakes(review): Validate brief content and preserve nested specifications

* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies

* no-mistakes(review): Ignore fenced subsection headings during brief validation

* no-mistakes(review): Preserve captain intent across scout promotion

* no-mistakes(review): Enforce safe intent boundaries for legacy promotions

* no-mistakes(review): Allow marked legacy intent and reject empty promotions

* no-mistakes(review): Scope task parsing and overlay legacy intent contracts

* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns

* no-mistakes(review): Preserve later captain clarifications in intent overlays

* no-mistakes(document): Document brief intent enforcement and ownership

* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed

* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint

* fix: start a fresh supervision branch for every main session (#3600)

* fix(pi): start a new supervision branch conversation per main session

The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.

The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.

The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.

* no-mistakes(document): Document fresh Pi supervision conversations

* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks

* feat: restart second mates after instruction updates (#3614)

* feat(update): restart second mates whose instructions changed

/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.

An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.

Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.

fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.

Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.

* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting

* no-mistakes(review): Parallelize relaunches and classify replacement incarnations

* no-mistakes(review): Gate restart actions on live agent state

* no-mistakes(review): Handle failed restart workers without hanging

* no-mistakes(review): Nudge legacy remotes and preserve persist recovery

* no-mistakes(review): Document one-time secondmate restart rollout

* no-mistakes(review): Honor arrived replies and refresh remote profiles

* no-mistakes(review): Revert remote parent profile reconciliation

* no-mistakes(review): Reset remote profile defaults and honor published results

* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates

* no-mistakes(document): Document second-mate restart update flow

* no-mistakes(lint): Fix ShellCheck warnings in restart scripts

* perf: accelerate local validation with bounded concurrency (#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation

* fix: copy PR URLs from durable records (#3648)

* fix: copy PR URLs from records or abstain, never assemble them

Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.

Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:

- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
  or abstain" section requires a URL to be copied verbatim from a dura…
peterOC26 added a commit to peterOC26/firstmate that referenced this pull request Sep 20, 2026
* fix: surface inbound Relay media to responding agents (#3442)

* fix: surface inbound Relay attachments to the responding agent

A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.

Fix it where the gap is, in prose:

- Read the complete payload object rather than a fixed field list, so
  media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
  mention and on every chain entry, and call out the common shape where
  only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
  (Discord: cdn.discordapp.com, media.discordapp.net,
  images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
  pbs.twimg.com, video.twimg.com), report a blocked host instead of
  working around it, and treat everything fetched as untrusted public
  input on the same terms as the surrounding thread text.

The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.

The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.

* no-mistakes(review): Preserve media authority and enforce poll-only fetching

* no-mistakes(document): Clarify Relay attachment safety prose

* fix(bin): defer inactive reconciliation during startup (#3480)

* Defer inactive startup reconciliation

* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably

* no-mistakes(review): Require worker phases to cover startup requests

* no-mistakes(review): Make diagnostic wakes safely acknowledgeable

* no-mistakes(document): Document deferred startup phase coverage

* fix(bin): bound wake drain presentation lock waits (#3475)

* fix: bound status presentation lock waits

* no-mistakes(review): Distinguish malformed presentation locks from live contention

* no-mistakes(review): Bound no-ack drain queue lock acquisition

* no-mistakes(document): Document bounded presentation-lock drain behavior

* no-mistakes(lint): Annotate bounded lock output global

* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite

* fix(bin): retire public follow-ups in remote homes (#3479)

* fix(relay): close a public loop whose work lives in a remote secondmate home

A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.

The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.

Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.

Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.

* no-mistakes(review): Guard remote link clears by request identity

* no-mistakes(review): Fail guarded clears on unreadable remote state

* no-mistakes(review): Reject guarded clears on non-writable remote state

* no-mistakes(review): Allow no-link retirement in non-writable remote state

* no-mistakes(document): Correct public-followup verification guarantee count

* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh

* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint

* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks

* fix(relay): bound the guarded remote link clear so it refuses instead of hanging

The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.

The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.

The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.

The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.

* no-mistakes(review): Harden lock-timeout regression with independent deadline

* no-mistakes(review): Restore no-op guarded clears on read-only state

* no-mistakes(document): Clarify remote public-followup cleanup contract

* fix(bin): support process events under symlinked homes (#3484)

* fix(bin): resolve process-event state roots before validating them

The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.

Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.

This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.

* fix(bin): pin the external capture staging boundary to its physical path

The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.

The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.

* no-mistakes(review): Propagate canonical process-event state roots

* no-mistakes(review): Propagate canonical state to process-event adapters

* no-mistakes(document): Document physical process-event state roots

* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)

* fix(pi): persist captain outcomes visibly

* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition

* no-mistakes(document): Document cold-start captain-outcome recovery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(pi): process captain outcomes through a sequence-keyed turn

PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.

The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.

Add the processing half on top of the persistence half:

- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
  cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
  only advances through an explicit sequence-bound acknowledgement, never
  past the read cursor and never backwards; an absent marker reads as zero
  and `processed-init` migrates delivered history once so an upgraded home
  is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
  every still-unprocessed captain row to main as one hidden, typed
  `fm-branch-process` request listing each `[seq N] task: summary`, opening
  exactly one main turn. Main closes it only by calling the new
  `fm_branch_processed` tool with the highest sequence listed. An unrelated,
  empty, or paraphrased answer leaves the sequence open, and the same request
  is presented again at the end of the next main run and at session start.
  The first two presentations of a sequence set open a turn of their own;
  after that the request rides the captain's next prompt so an ignored
  request cannot loop, and a session replacement resets that budget.
  Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
  scripts: an empty answer and an unrelated prior answer neither advance the
  marker nor stop re-presentation, the acknowledgement is refused beyond the
  read cursor and outside lock ownership, a partial acknowledgement keeps the
  newer sequence open, and #3312's own assertions now forbid an unkeyed turn
  rather than any turn. The store suite pins the marker's bounds and the
  migration; the real-SDK guard for appendEntry persistence and model
  exclusion is unchanged.

Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.

* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements

* no-mistakes(review): Harden outcome state validation and request pacing

* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores

* no-mistakes(review): Validate canonical mark-read cursor state

* no-mistakes(review): Guard cursor advancement against corrupt processed state

* no-mistakes(review): Bind acknowledgements to active processing requests

* no-mistakes(review): Reset pacing when processing sequence membership changes

* no-mistakes(review): Enforce silent outcome invariants at storage boundary

* no-mistakes(document): Document hardened captain outcome processing contracts

---------

Co-authored-by: kunchenguid <kun@kunchenguid.com>

* feat: add bounded concurrent Bearings ledger collection (#3481)

* feat: bound Bearings remote ledger collection

* no-mistakes(review): Clarify default remote-ledger collection behavior

* no-mistakes(review): Detach reconcile delivery from watcher loop

* no-mistakes(review): Enforce bounded snapshot and request captures

* no-mistakes(review): Bound legacy summary capture before parsing

* no-mistakes(review): Bound primary remote ledger captures

* no-mistakes(document): Correct snapshot and reconcile documentation

* no-mistakes(lint): Fix ShellCheck quoting in bounded collector

* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks

* test: await reconcile request retirement

* no-mistakes(review): Avoid empty reconcile queue process churn

* no-mistakes(review): Read ledger summaries from immutable snapshots

* no-mistakes(review): Reject multi-document home ledger streams

* no-mistakes(review): Coalesce durable reconcile requests per target

* no-mistakes(review): Unify reconcile keys and reject snapshot streams

* no-mistakes(review): Key reconcile requests by stable target ID

* no-mistakes(document): Document per-target reconcile request coalescing

* no-mistakes(lint): Remove unused snapshot summary file variable

* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass

* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks

* ci: rebalance portable serial test shards (#3489)

* fix(ci): rebalance the portable serial shards on measured durations

The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.

Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.

Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.

Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.

No test changes what it asserts and no test stops running; only the
partition across shards changes.

* no-mistakes(document): Clarify conservative shard timing aggregate

* fix(pi): fall back on incomplete supervision branch prompts (#3491)

* fix(pi): fall back after settled branch errors

* no-mistakes(review): Detect provider errors across prompt compaction

* no-mistakes(review): Preserve in-flight branch state across selection changes

* fix(pi): re-probe supervision branch after cooldown (#3497)

* fix(pi): recover supervision branch after cooldown

* no-mistakes(review): Defer branch recovery until prompt settlement

* no-mistakes(document): Clarify supervision cooldown recovery contract

* fix(bin): remove legacy remote snapshot reads (#3501)

* refactor: remove legacy remote summary reads

* no-mistakes(document): Document ledger-only snapshot reads

* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass

* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean

* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux

* fix(pi): preserve watcher continuity across session replacement (#3498)

* fix(pi): rearm watcher after session replacement

* no-mistakes(review): Queue actionable closes across Pi session replacement

* no-mistakes(review): Stop replacement arm when handoff persistence fails

* no-mistakes(review): Preserve actionable wakes through branch and late child races

* no-mistakes(review): Surface late handoff failures without crashing Pi

* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens

* no-mistakes(review): Retry stale deliveries and release settled claims

* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup

* no-mistakes(review): Deduplicate persistent handoff cleanup alerts

* no-mistakes(review): Acknowledge watcher follow-ups only when consumed

* no-mistakes(review): Persist idle follow-ups until agent consumption

* no-mistakes(review): Preserve pending outcomes when handoff persistence fails

* no-mistakes(review): Arm replacement before awaiting prior delivery settlement

* no-mistakes(review): Adopt pending handoffs after lock reclamation

* no-mistakes(review): Prevent stale generations from adopting replacement handoffs

* no-mistakes(review): Scope replacement handoffs by watcher state

* no-mistakes(document): Clarify replacement handoff documentation

* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks

* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes

* no-mistakes(document): Document watcher-owned replacement handoffs

* no-mistakes(document): Verify replacement handoff documentation

* test(pi): cover watcher-owned branch fallback

* no-mistakes(document): Refresh watcher-owned fallback documentation

* fix(bin): resurface task statuses missed by wake handling (#3495)

* fix(bin): resurface terminal statuses lost after branch handling

* test(watch): canonicalize process-event fixture homes

* no-mistakes(review): Index branch outcomes by causal status position

* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses

* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics

* no-mistakes(review): Keep unclassifiable oversized statuses silent

* no-mistakes(document): Document lost-wake outcome backstop

* no-mistakes(document): Update outcome backstop documentation

* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally

* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes

* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift

* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state

* fix(bin): collect follow-up results from remote work homes (#3503)

* fix(bin): deliver typed terminal results from remote work homes

A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.

The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.

A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.

This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.

* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes

* no-mistakes(review): Fail collection when remote outbox is unreadable

* no-mistakes(review): Surface reassigned remote routes during empty collection

* no-mistakes(review): Fail remote collection on invalid registrations

* no-mistakes(review): Reject unsafe registration entries during remote collection

* no-mistakes(review): Restore healthy empty remote collection behavior

* no-mistakes(review): Skip remote collection for delivered registrations

* no-mistakes(review): Skip delivered registrations before route validation

* no-mistakes(document): Document remote follow-up collection semantics

* fix(bin): exclude secondmates from home-summary validity (#3504)

* fix(bin): exclude secondmates from home-summary child inventory

kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.

* no-mistakes(review): Cover terminal secondmate in-flight exclusion

* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds

* fix(bin): self-heal outcome indexes on first drain (#3509)

* fix(bin): self-heal status-outcome indexes on every drain

Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.

* no-mistakes(review): Guard held-lock initialization and fail marker writes

* no-mistakes(document): Document cross-harness outcome-index self-healing

* fix(bearings): keep active children underway during captain holds (#3505)

* fix(bearings): keep active children underway beside a captain hold

Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.

* no-mistakes(review): Preserve Underway repos and disclose child truncation

* no-mistakes(review): Fall back to task project for Underway repos

* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean

* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)

* fix(pi): settle watcher delivery on Pi accepting the follow-up

A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.

The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.

Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(pi): retry a verified successor that fails during wake delivery

A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.

The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.

The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(bin): bound repeat stale wakes for parked workers (#3532)

* fix(bin): bound repeat stale wakes for a parked but live worker

A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.

pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.

Two places let that churn re-alarm:

- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
  declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
  have suppressed it. The throttle was never read on this path and was advanced
  by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
  whenever the classification came back `none`, so each tick also bought the same
  declared wait a fresh window. Fixing only the first site changes nothing.

Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.

First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.

Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.

* fix(document): Clarify declared-wait wake cadence documentation

* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor

* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed

* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)

* fix(turnend): accept the away-mode daemon as the supervision owner

While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.

Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.

The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.

The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.

* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage

* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md

* fix(backlog): omit --file from row probes for non-markdown backends (#3582)

* fix(backlog): omit markdown file for beads probes

* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes

* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)

* fix(bin): classify progress updates on requested work as routine (#3589)

The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.

The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.

* fix(bin): preserve captain calls during teardown (#3595)

* fix(bin): never close a captain call during cleanup

A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.

bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.

Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.

The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.

Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.

Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np

* no-mistakes(review): Serialize captain holds and fix backend-aware listing

* no-mistakes(document): Update captain-call retention documentation

* no-mistakes(document): Fix relocated captain-hold backlog diagnostics

* fix(bin): deliver secondmate outcomes to the parent channel (#3592)

* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts

A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:

- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
  exact-line append-once; the merge outcome path and the inactive-outcome
  scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
  watcher poll in a secondmate home: a direct child's whole terminal done or
  failed line is delivered at once with its note, recorded PR, mode, merge
  posture, and scout report pointer, keyed and receipted so it is delivered
  once, and the inactive path yields to it. `report <task-id>` runs the same
  delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
  registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
  and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
  record and refuses, retaining every record, while the channel cannot be
  written.
- The charter opens with the parent-channel rule and confines the mate's own
  appends to judgement; AGENTS.md carries the carve-out at the persona
  address rule and the escalation list.

docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.

* no-mistakes(review): Fix parent outcome retries and reconciliation locking

* no-mistakes(review): Prevent busy children from starving ledger delivery

* no-mistakes(review): Correct ledger metadata and hold occurrence handling

* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons

* no-mistakes(review): Close ledger races and preserve teardown records

* no-mistakes(document): Correct parent-channel receipt and scanner documentation

* no-mistakes(lint): Quote done arguments for ShellCheck compliance

* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks

* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure

* fix(bin): sync remote second mates to primary commit (#3599)

* fix(bin): sync remote second-mate homes to the parent primary commit

Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.

The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.

The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.

/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.

* no-mistakes(document): Document primary-targeted remote secondmate synchronization

* fix(bin): separate captain intent from firstmate specs (#3597)

* fix(bin): split brief task into captain intent and firstmate spec

Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.

* fix(bin): stop task-subsection copies at the next heading

Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.

* no-mistakes(review): Validate brief content and preserve nested specifications

* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies

* no-mistakes(review): Ignore fenced subsection headings during brief validation

* no-mistakes(review): Preserve captain intent across scout promotion

* no-mistakes(review): Enforce safe intent boundaries for legacy promotions

* no-mistakes(review): Allow marked legacy intent and reject empty promotions

* no-mistakes(review): Scope task parsing and overlay legacy intent contracts

* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns

* no-mistakes(review): Preserve later captain clarifications in intent overlays

* no-mistakes(document): Document brief intent enforcement and ownership

* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed

* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint

* fix: start a fresh supervision branch for every main session (#3600)

* fix(pi): start a new supervision branch conversation per main session

The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.

The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.

The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.

* no-mistakes(document): Document fresh Pi supervision conversations

* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks

* feat: restart second mates after instruction updates (#3614)

* feat(update): restart second mates whose instructions changed

/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.

An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.

Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.

fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.

Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.

* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting

* no-mistakes(review): Parallelize relaunches and classify replacement incarnations

* no-mistakes(review): Gate restart actions on live agent state

* no-mistakes(review): Handle failed restart workers without hanging

* no-mistakes(review): Nudge legacy remotes and preserve persist recovery

* no-mistakes(review): Document one-time secondmate restart rollout

* no-mistakes(review): Honor arrived replies and refresh remote profiles

* no-mistakes(review): Revert remote parent profile reconciliation

* no-mistakes(review): Reset remote profile defaults and honor published results

* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates

* no-mistakes(document): Document second-mate restart update flow

* no-mistakes(lint): Fix ShellCheck warnings in restart scripts

* perf: accelerate local validation with bounded concurrency (#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation

* fix: copy PR URLs from durable records (#3648)

* fix: copy PR URLs from records or abstain, never assemble them

Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.

Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:

- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
  or abstain" section requires a URL to be copied verbatim from a durable
  record (the done: PR <url> status line, pr= metadata, or the backlog note),
  forbids assembling owner, repository, host, or number from memory, and has
  the branch report only the identifier it actually holds when no record names
  the URL yet, leaving the PR check unarmed until the worker's ready line
  arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for
  main in place of the bare full-URL mandate.

- Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full
  https:// URL wherever a PR is mentioned - status line, terminal, or summary -
  never a bare "PR 108", so the link is in view as early as the number is.

- bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that
  the task's own done lines contradict, printing both spellings; a log naming
  no URL still records the argument as before. fm_pr_status_ready_urls in
  bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches
  bin/fm-pr-merge.sh, so nothing merges under a contradicted URL.

Tests cover the offline refusal with zero side effects, the recorded spelling
being accepted, markdown-wrapped and punctuated URLs, working lines not
counting, the merge wrapper propagation, a self-hosted merge request with no
forge call, the prompt carrying the rule, and the brief carrying the worker
rule.

* no-mistakes(review): Remove stale PR URL enforcement

* no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`

* fix(bin): disable Claude feedback drafts for fleet launches (#3661)

* fix(bin): disable Claude's feedback-draft flow for fleet-launched agents

Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched
Claude crewmate and secondmate, so /bug and /feedback never queue or
submit a bug report on the captain's behalf. feedbackDrafts is the
documented settings key (Claude Code changelog 2.1.247); the
per-launch CLI flag never touches the captain's global settings.json.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts

* no-mistakes(document): Fix Claude feedback documentation formatting

* fix(bin): layer both feedback-draft controls for defense in depth

The prior --settings-only fix can be overridden by a managed Claude
settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0
alongside --settings '{"feedbackDrafts":"off"}': either control alone
disables the SendFeedback tool, so a managed override of one still
leaves the other in force.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(document): Document Claude feedback-draft suppression ownership

* feat(tests): run three more validation families concurrently (#3662)

* perf(tests): admit three more families to concurrent validation

The three families that `docs/fm-test-isolation-proof.md` recorded as refused
were not refused for concurrency. Each blocker was a test that decided a
property by wall clock, or a script filed where it cannot run. Fixing those
three things admits all three families and recovers 28.6 minutes of local
validation with no assertion removed or weakened.

- `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the
  handoff, sleeping a fixed second, then delegating the move to the real
  binary. Nothing ever killed the fake, so on a host slow enough for the case's
  next assertions to take longer than a second, the orphan woke and completed
  the very move the case requires left undone, and recovery then failed with
  `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs
  during the injected crash showed exactly that, the item moving one second
  after the crash. All four crash injections in the file now go through a new
  `fm_fake_crash_injector` shim that signals the target and returns only once
  it is observably gone, and the pre-move fake never delegates the move at all.

- `tests/fm-session-start.test.sh` proved the startup digest does not block on
  a slow current-state read by timing the whole digest against a fixed
  eight-second sleep, which a loaded host exceeds without the property being
  violated. It now holds that read open until the case releases it and asserts,
  the moment the digest returns, that the read has not finished. A digest that
  waited would wait indefinitely rather than for an interval a slow host can
  out-run, so the assertion is stronger than the bound it replaces. Its scan
  budget moves to the maximum, because the old value left two seconds of margin
  over the fixed sleep and measured the host rather than the deadline that
  `tests/fm-inactive-reconcile.test.sh` owns.

- `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all,
  which put it in the portable serial lane, where Linux CI gate-skips it: that
  real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and
  the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its
  opt-in variable and moves to `live-harness-optin`.

The 28 remaining ungrouped scripts become an enumerated `standalone` family
instead of admitting `unclassified` itself. `unclassified` is the family map's
`*)` arm, so admitting it would silently grant concurrency to every test added
afterwards, which is exactly the population with no proof. A new test still
lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh`
covers that split behaviorally.

Each family passes two consecutive four-worker proofs with zero failures. On
the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap`
756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s
recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock
against 121 minutes of summed script time.

* no-mistakes(document): Refresh concurrent validation and shard documentation

* no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2

* feat: structure no-mistakes ask-user escalations (#3670)

* feat(brief): structure no-mistakes ask-user escalation as event + snapshot file

Crewmates escalating a no-mistakes ask-user gate now report one status
event naming every finding id plus a snapshot file holding the gate's
axi finding records verbatim (id, severity, file, line, description,
authority), using the same shape even for a single finding. The status
line never paraphrases. The format is defined once in fm-dod-lib.sh and
rendered into both the scout and ship rule 6 in fm-brief.sh, so a
promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets
the identical contract as a freshly-spawned no-mistakes ship worker.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei

* no-mistakes(review): Preserve ask-user escalation output contract

* no-mistakes(review): Align escalation format test expectation

* no-mistakes(review): Scope ask-user escalation instructions correctly

* no-mistakes(review): Remove ask-user from generic decision rules

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix(bin): require self-sufficient no-mistakes intent (#3671)

* fix(bin): require a self-sufficient no-mistakes intent

A no-mistakes worker's --intent is only as useful as the string it
passes. PR #3604 shipped with an intent that was only "do 1, 2, 3, 7
from the report": the real contract lived in a private scout report and
never reached --intent, so nobody holding that string plus the codebase
could have derived the specification.

This is pure instruction at the contract's one owner; no spawn-side or
promotion-side check is added.

- bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now
  states that the --intent string must be self-sufficient (the string
  plus the codebase reconstructs roughly the same specification) and
  tells the worker to write the substance of any report, decision,…
sctru added a commit to sctru/firstmate that referenced this pull request Sep 21, 2026
* fix(bin): safely unregister custom checks (#3369)

* fix(bin): add a safe owner for custom-check retirement

Agents were improvising rm of check files with unset STATE/ID, which wedges
headless panes. Unregister validates the id and state directory first.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Refuse explicitly empty custom-check state overrides

* no-mistakes(document): Document custom-check retirement safety contract

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(quota): extract mid-task polling and candidate selection into dedicated scripts (#3221)

* Add quota exhaustion detection and safe fallback helpers

- bin/fm-procevent-quota.sh: generic procevent adapter that arms a
  recurring quota-axi --json poll and wakes firstmate when a tracked
  provider's effectivePercentRemaining drops below a threshold or its
  runway.status becomes exhausted_now.
- bin/fm-quota-choose.sh: worker-side helper that picks the first ranked
  harness:model candidate with positive effectivePercentRemaining.
- AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document
  the new helpers and the mid-task quota-exhaustion wake path.
- tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON
  source.

* no-mistakes(review): Fix quota polling and scope bounds

* no-mistakes(review): Enforce safe default quota selection

* no-mistakes(review): Handle decimal quota values safely

* no-mistakes(review): Fail closed on invalid quota inputs

* no-mistakes(review): Reject empty quota candidate segments

* no-mistakes(review): Harden quota parsing and timeout ownership

* no-mistakes(review): Reuse captured quota snapshots consistently

* no-mistakes(review): Match quota using explicit candidate providers

* no-mistakes(review): Centralize fail-closed quota schema validation

* no-mistakes(review): Reject out-of-range quota percentages

* no-mistakes(review): Validate quota runway status enum

* no-mistakes(review): Tighten quota scope and status contracts

* no-mistakes(review): Preserve unknown quota and exact product bounds

* no-mistakes(review): Preserve provider-level unknown quota

* no-mistakes(review): Reuse canonical verified harness validation

* no-mistakes(document): Document mid-task quota handling

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(docs): restore default routing contract, keep quota helper optional

Restore the AGENTS.md section 4 always-loaded routing paragraph the PR
had deleted, so the standing TOON-first intake, spendPriority ranker,
every-candidate accounting, and load-trigger contract stay exactly as
before this PR. The mid-task quota wake is optional and must not alter
default routing.

Restore the quota-array-dispatch skill ownership line to section 4 as
the always-loaded intake boundary owner; keep the worker-side helper
section as an addition only, without rewiring ownership or load
triggers to section 13.

* fix(bin): use harness-keyed quota matching in optional helper

Revert fm-quota-choose.sh from harness:provider:model tuples back to
harness:model candidates with harness-keyed provider matching, per the
resolved ask-user finding. The helper is optional; authoritative
multi-provider routing (provider discovery from the harness catalog and
quota matching by that explicit provider) stays owned by AGENTS.md
section 4 and the quota-array-dispatch skill intake procedure, not the
helper.

Document the multi-provider limitation in the helper header and the
quota-array-dispatch skill: the helper maps each harness to one primary
provider family only, so a candidate whose established provider differs
from that primary family is checked against the wrong quota row. Use it
only when the brief fixed the candidate order and every candidate's
provider is the harness's primary family.

The helper still consumes one already-captured default-TOON or JSON
snapshot via stdin or --snapshot and never calls quota-axi itself, so
it selects from the same quota state as the intake.

* no-mistakes(review): Fix Muse quota mapping and helper contract docs

* no-mistakes(review): Reject known-empty quotas and map quota tests explicitly

* no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse

* no-mistakes(review): Fix quota retirement and dependent regression coverage

* no-mistakes(review): Accept zero-row quota TOON snapshots

* no-mistakes(review): Enforce quota semantics status consistency

* no-mistakes(review): Veto dispatch on any exhausted applicable scope

* no-mistakes(review): Record exhausted quota scope in wake details

* no-mistakes(review): Fix quota help and control dependency coverage

* no-mistakes(review): Decode quoted TOON fields and document quota wakes

* no-mistakes(review): Validate zero-row TOON and map timeout coverage

* no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes

* no-mistakes(review): Validate complete nonzero TOON envelopes

* no-mistakes(review): Accept producer-shaped quota TOON envelopes

* no-mistakes(review): Support empty quota arrays and validate counted rows

* no-mistakes(review): Harden TOON completion, scopes, and quoted fields

* no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields

* no-mistakes(review): Allow unknown headroom under known semantics

* no-mistakes(review): Reject noncanonical quota identities

* no-mistakes(review): Preserve empty quota polling and validate attention identities

* no-mistakes(review): Reject noncanonical provider watches

* no-mistakes(review): Validate all candidates before quota selection

* no-mistakes(document): Correct quota helper safety documentation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix: surface comments on Lavish annotations (#3371)

* fix(bin): keep typed Lavish comments when an element is also annotated

read preferred element text over prompt, so an annotate-and-comment
item dropped the captain's words. Surface prompt as its own field.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Filter non-comment prompts from Lavish reader output

* no-mistakes(document): Clarify Lavish comment presentation contract

* no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure

* no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure

* fix(bin): always emit Lavish comments and use real annotation fixtures

Stop inferring comment provenance from prompt==text. Real pure
annotations have no prompt, so always-emit does not duplicate.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: support first public-followup registration on Bash 3.2 (#3420)

* Fix public-followup register crashing on empty lock arrays under bash 3.2.

bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock.
The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there.

* no-mistakes(document): Document stock Bash registration coverage

* no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5

* no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression

* fix(bin): isolate new Herdr server environments (#2792)

* fix(herdr): isolate server launch environment

* no-mistakes(review): Clear inherited supervision model from Herdr launches

* no-mistakes(document): Document Herdr server launch environment isolation

* fix: surface inbound Relay media to responding agents (#3442)

* fix: surface inbound Relay attachments to the responding agent

A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.

Fix it where the gap is, in prose:

- Read the complete payload object rather than a fixed field list, so
  media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
  mention and on every chain entry, and call out the common shape where
  only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
  (Discord: cdn.discordapp.com, media.discordapp.net,
  images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
  pbs.twimg.com, video.twimg.com), report a blocked host instead of
  working around it, and treat everything fetched as untrusted public
  input on the same terms as the surrounding thread text.

The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.

The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.

* no-mistakes(review): Preserve media authority and enforce poll-only fetching

* no-mistakes(document): Clarify Relay attachment safety prose

* fix(bin): defer inactive reconciliation during startup (#3480)

* Defer inactive startup reconciliation

* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably

* no-mistakes(review): Require worker phases to cover startup requests

* no-mistakes(review): Make diagnostic wakes safely acknowledgeable

* no-mistakes(document): Document deferred startup phase coverage

* fix(bin): bound wake drain presentation lock waits (#3475)

* fix: bound status presentation lock waits

* no-mistakes(review): Distinguish malformed presentation locks from live contention

* no-mistakes(review): Bound no-ack drain queue lock acquisition

* no-mistakes(document): Document bounded presentation-lock drain behavior

* no-mistakes(lint): Annotate bounded lock output global

* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite

* fix(bin): retire public follow-ups in remote homes (#3479)

* fix(relay): close a public loop whose work lives in a remote secondmate home

A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.

The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.

Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.

Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.

* no-mistakes(review): Guard remote link clears by request identity

* no-mistakes(review): Fail guarded clears on unreadable remote state

* no-mistakes(review): Reject guarded clears on non-writable remote state

* no-mistakes(review): Allow no-link retirement in non-writable remote state

* no-mistakes(document): Correct public-followup verification guarantee count

* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh

* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint

* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks

* fix(relay): bound the guarded remote link clear so it refuses instead of hanging

The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.

The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.

The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.

The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.

* no-mistakes(review): Harden lock-timeout regression with independent deadline

* no-mistakes(review): Restore no-op guarded clears on read-only state

* no-mistakes(document): Clarify remote public-followup cleanup contract

* fix(bin): support process events under symlinked homes (#3484)

* fix(bin): resolve process-event state roots before validating them

The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.

Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.

This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.

* fix(bin): pin the external capture staging boundary to its physical path

The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.

The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.

* no-mistakes(review): Propagate canonical process-event state roots

* no-mistakes(review): Propagate canonical state to process-event adapters

* no-mistakes(document): Document physical process-event state roots

* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)

* fix(pi): persist captain outcomes visibly

* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition

* no-mistakes(document): Document cold-start captain-outcome recovery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(pi): process captain outcomes through a sequence-keyed turn

PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.

The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.

Add the processing half on top of the persistence half:

- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
  cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
  only advances through an explicit sequence-bound acknowledgement, never
  past the read cursor and never backwards; an absent marker reads as zero
  and `processed-init` migrates delivered history once so an upgraded home
  is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
  every still-unprocessed captain row to main as one hidden, typed
  `fm-branch-process` request listing each `[seq N] task: summary`, opening
  exactly one main turn. Main closes it only by calling the new
  `fm_branch_processed` tool with the highest sequence listed. An unrelated,
  empty, or paraphrased answer leaves the sequence open, and the same request
  is presented again at the end of the next main run and at session start.
  The first two presentations of a sequence set open a turn of their own;
  after that the request rides the captain's next prompt so an ignored
  request cannot loop, and a session replacement resets that budget.
  Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
  scripts: an empty answer and an unrelated prior answer neither advance the
  marker nor stop re-presentation, the acknowledgement is refused beyond the
  read cursor and outside lock ownership, a partial acknowledgement keeps the
  newer sequence open, and #3312's own assertions now forbid an unkeyed turn
  rather than any turn. The store suite pins the marker's bounds and the
  migration; the real-SDK guard for appendEntry persistence and model
  exclusion is unchanged.

Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.

* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements

* no-mistakes(review): Harden outcome state validation and request pacing

* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores

* no-mistakes(review): Validate canonical mark-read cursor state

* no-mistakes(review): Guard cursor advancement against corrupt processed state

* no-mistakes(review): Bind acknowledgements to active processing requests

* no-mistakes(review): Reset pacing when processing sequence membership changes

* no-mistakes(review): Enforce silent outcome invariants at storage boundary

* no-mistakes(document): Document hardened captain outcome processing contracts

---------

Co-authored-by: kunchenguid <kun@kunchenguid.com>

* feat: add bounded concurrent Bearings ledger collection (#3481)

* feat: bound Bearings remote ledger collection

* no-mistakes(review): Clarify default remote-ledger collection behavior

* no-mistakes(review): Detach reconcile delivery from watcher loop

* no-mistakes(review): Enforce bounded snapshot and request captures

* no-mistakes(review): Bound legacy summary capture before parsing

* no-mistakes(review): Bound primary remote ledger captures

* no-mistakes(document): Correct snapshot and reconcile documentation

* no-mistakes(lint): Fix ShellCheck quoting in bounded collector

* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks

* test: await reconcile request retirement

* no-mistakes(review): Avoid empty reconcile queue process churn

* no-mistakes(review): Read ledger summaries from immutable snapshots

* no-mistakes(review): Reject multi-document home ledger streams

* no-mistakes(review): Coalesce durable reconcile requests per target

* no-mistakes(review): Unify reconcile keys and reject snapshot streams

* no-mistakes(review): Key reconcile requests by stable target ID

* no-mistakes(document): Document per-target reconcile request coalescing

* no-mistakes(lint): Remove unused snapshot summary file variable

* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass

* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks

* ci: rebalance portable serial test shards (#3489)

* fix(ci): rebalance the portable serial shards on measured durations

The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.

Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.

Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.

Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.

No test changes what it asserts and no test stops running; only the
partition across shards changes.

* no-mistakes(document): Clarify conservative shard timing aggregate

* fix(pi): fall back on incomplete supervision branch prompts (#3491)

* fix(pi): fall back after settled branch errors

* no-mistakes(review): Detect provider errors across prompt compaction

* no-mistakes(review): Preserve in-flight branch state across selection changes

* fix(pi): re-probe supervision branch after cooldown (#3497)

* fix(pi): recover supervision branch after cooldown

* no-mistakes(review): Defer branch recovery until prompt settlement

* no-mistakes(document): Clarify supervision cooldown recovery contract

* fix(bin): remove legacy remote snapshot reads (#3501)

* refactor: remove legacy remote summary reads

* no-mistakes(document): Document ledger-only snapshot reads

* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass

* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean

* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux

* fix(pi): preserve watcher continuity across session replacement (#3498)

* fix(pi): rearm watcher after session replacement

* no-mistakes(review): Queue actionable closes across Pi session replacement

* no-mistakes(review): Stop replacement arm when handoff persistence fails

* no-mistakes(review): Preserve actionable wakes through branch and late child races

* no-mistakes(review): Surface late handoff failures without crashing Pi

* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens

* no-mistakes(review): Retry stale deliveries and release settled claims

* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup

* no-mistakes(review): Deduplicate persistent handoff cleanup alerts

* no-mistakes(review): Acknowledge watcher follow-ups only when consumed

* no-mistakes(review): Persist idle follow-ups until agent consumption

* no-mistakes(review): Preserve pending outcomes when handoff persistence fails

* no-mistakes(review): Arm replacement before awaiting prior delivery settlement

* no-mistakes(review): Adopt pending handoffs after lock reclamation

* no-mistakes(review): Prevent stale generations from adopting replacement handoffs

* no-mistakes(review): Scope replacement handoffs by watcher state

* no-mistakes(document): Clarify replacement handoff documentation

* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks

* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes

* no-mistakes(document): Document watcher-owned replacement handoffs

* no-mistakes(document): Verify replacement handoff documentation

* test(pi): cover watcher-owned branch fallback

* no-mistakes(document): Refresh watcher-owned fallback documentation

* fix(bin): resurface task statuses missed by wake handling (#3495)

* fix(bin): resurface terminal statuses lost after branch handling

* test(watch): canonicalize process-event fixture homes

* no-mistakes(review): Index branch outcomes by causal status position

* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses

* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics

* no-mistakes(review): Keep unclassifiable oversized statuses silent

* no-mistakes(document): Document lost-wake outcome backstop

* no-mistakes(document): Update outcome backstop documentation

* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally

* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes

* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift

* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state

* fix(bin): collect follow-up results from remote work homes (#3503)

* fix(bin): deliver typed terminal results from remote work homes

A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.

The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.

A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.

This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.

* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes

* no-mistakes(review): Fail collection when remote outbox is unreadable

* no-mistakes(review): Surface reassigned remote routes during empty collection

* no-mistakes(review): Fail remote collection on invalid registrations

* no-mistakes(review): Reject unsafe registration entries during remote collection

* no-mistakes(review): Restore healthy empty remote collection behavior

* no-mistakes(review): Skip remote collection for delivered registrations

* no-mistakes(review): Skip delivered registrations before route validation

* no-mistakes(document): Document remote follow-up collection semantics

* fix(bin): exclude secondmates from home-summary validity (#3504)

* fix(bin): exclude secondmates from home-summary child inventory

kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.

* no-mistakes(review): Cover terminal secondmate in-flight exclusion

* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds

* fix(bin): self-heal outcome indexes on first drain (#3509)

* fix(bin): self-heal status-outcome indexes on every drain

Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.

* no-mistakes(review): Guard held-lock initialization and fail marker writes

* no-mistakes(document): Document cross-harness outcome-index self-healing

* fix(bearings): keep active children underway during captain holds (#3505)

* fix(bearings): keep active children underway beside a captain hold

Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.

* no-mistakes(review): Preserve Underway repos and disclose child truncation

* no-mistakes(review): Fall back to task project for Underway repos

* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean

* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)

* fix(pi): settle watcher delivery on Pi accepting the follow-up

A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.

The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.

Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(pi): retry a verified successor that fails during wake delivery

A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.

The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.

The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(bin): bound repeat stale wakes for parked workers (#3532)

* fix(bin): bound repeat stale wakes for a parked but live worker

A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.

pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.

Two places let that churn re-alarm:

- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
  declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
  have suppressed it. The throttle was never read on this path and was advanced
  by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
  whenever the classification came back `none`, so each tick also bought the same
  declared wait a fresh window. Fixing only the first site changes nothing.

Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.

First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.

Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.

* fix(document): Clarify declared-wait wake cadence documentation

* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor

* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed

* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)

* fix(turnend): accept the away-mode daemon as the supervision owner

While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.

Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.

The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.

The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.

* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage

* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md

* fix(backlog): omit --file from row probes for non-markdown backends (#3582)

* fix(backlog): omit markdown file for beads probes

* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes

* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)

* fix(bin): classify progress updates on requested work as routine (#3589)

The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.

The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.

* fix(bin): preserve captain calls during teardown (#3595)

* fix(bin): never close a captain call during cleanup

A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.

bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.

Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.

The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.

Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.

Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np

* no-mistakes(review): Serialize captain holds and fix backend-aware listing

* no-mistakes(document): Update captain-call retention documentation

* no-mistakes(document): Fix relocated captain-hold backlog diagnostics

* fix(bin): deliver secondmate outcomes to the parent channel (#3592)

* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts

A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:

- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
  exact-line append-once; the merge outcome path and the inactive-outcome
  scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
  watcher poll in a secondmate home: a direct child's whole terminal done or
  failed line is delivered at once with its note, recorded PR, mode, merge
  posture, and scout report pointer, keyed and receipted so it is delivered
  once, and the inactive path yields to it. `report <task-id>` runs the same
  delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
  registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
  and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
  record and refuses, retaining every record, while the channel cannot be
  written.
- The charter opens with the parent-channel rule and confines the mate's own
  appends to judgement; AGENTS.md carries the carve-out at the persona
  address rule and the escalation list.

docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.

* no-mistakes(review): Fix parent outcome retries and reconciliation locking

* no-mistakes(review): Prevent busy children from starving ledger delivery

* no-mistakes(review): Correct ledger metadata and hold occurrence handling

* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons

* no-mistakes(review): Close ledger races and preserve teardown records

* no-mistakes(document): Correct parent-channel receipt and scanner documentation

* no-mistakes(lint): Quote done arguments for ShellCheck compliance

* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks

* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure

* fix(bin): sync remote second mates to primary commit (#3599)

* fix(bin): sync remote second-mate homes to the parent primary commit

Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.

The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.

The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.

/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.

* no-mistakes(document): Document primary-targeted remote secondmate synchronization

* fix(bin): separate captain intent from firstmate specs (#3597)

* fix(bin): split brief task into captain intent and firstmate spec

Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.

* fix(bin): stop task-subsection copies at the next heading

Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.

* no-mistakes(review): Validate brief content and preserve nested specifications

* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies

* no-mistakes(review): Ignore fenced subsection headings during brief validation

* no-mistakes(review): Preserve captain intent across scout promotion

* no-mistakes(review): Enforce safe intent boundaries for legacy promotions

* no-mistakes(review): Allow marked legacy intent and reject empty promotions

* no-mistakes(review): Scope task parsing and overlay legacy intent contracts

* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns

* no-mistakes(review): Preserve later captain clarifications in intent overlays

* no-mistakes(document): Document brief intent enforcement and ownership

* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed

* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint

* fix: start a fresh supervision branch for every main session (#3600)

* fix(pi): start a new supervision branch conversation per main session

The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.

The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.

The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.

* no-mistakes(document): Document fresh Pi supervision conversations

* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks

* feat: restart second mates after instruction updates (#3614)

* feat(update): restart second mates whose instructions changed

/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.

An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.

Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.

fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.

Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.

* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting

* no-mistakes(review): Parallelize relaunches and classify replacement incarnations

* no-mistakes(review): Gate restart actions on live agent state

* no-mistakes(review): Handle failed restart workers without hanging

* no-mistakes(review): Nudge legacy remotes and preserve persist recovery

* no-mistakes(review): Document one-time secondmate restart rollout

* no-mistakes(review): Honor arrived replies and refresh remote profiles

* no-mistakes(review): Revert remote parent profile reconciliation

* no-mistakes(review): Reset remote profile defaults and honor published results

* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates

* no-mistakes(document): Document second-mate restart update flow

* no-mistakes(lint): Fix ShellCheck warnings in restart scripts

* perf: accelerate local validation with bounded concurrency (#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation

* fix: copy PR URLs from durable records (#3648)

* fix: copy PR URLs from records or abstain, never assemble them

Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.

Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:

- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
  or abstain" section requires a URL to be copied verbatim from a …
friesentius pushed a commit to friesentius/firstmate that referenced this pull request Sep 21, 2026
…id#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation
friesentius pushed a commit to friesentius/firstmate that referenced this pull request Sep 21, 2026
* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR kunchenguid#3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR kunchenguid#823 added and PR kunchenguid#1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------
friesentius pushed a commit to friesentius/firstmate that referenced this pull request Sep 21, 2026
…id#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation
friesentius pushed a commit to friesentius/firstmate that referenced this pull request Sep 21, 2026
* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR kunchenguid#3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR kunchenguid#823 added and PR kunchenguid#1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
friesentius added a commit to friesentius/firstmate that referenced this pull request Sep 21, 2026
…g handling (#6)

* feat: restart second mates after instruction updates (#3614)

* feat(update): restart second mates whose instructions changed

/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.

An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.

Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.

fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.

Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.

* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting

* no-mistakes(review): Parallelize relaunches and classify replacement incarnations

* no-mistakes(review): Gate restart actions on live agent state

* no-mistakes(review): Handle failed restart workers without hanging

* no-mistakes(review): Nudge legacy remotes and preserve persist recovery

* no-mistakes(review): Document one-time secondmate restart rollout

* no-mistakes(review): Honor arrived replies and refresh remote profiles

* no-mistakes(review): Revert remote parent profile reconciliation

* no-mistakes(review): Reset remote profile defaults and honor published results

* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates

* no-mistakes(document): Document second-mate restart update flow

* no-mistakes(lint): Fix ShellCheck warnings in restart scripts

* perf: accelerate local validation with bounded concurrency (#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation

* fix: copy PR URLs from durable records (#3648)

* fix: copy PR URLs from records or abstain, never assemble them

Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.

Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:

- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
  or abstain" section requires a URL to be copied verbatim from a durable
  record (the done: PR <url> status line, pr= metadata, or the backlog note),
  forbids assembling owner, repository, host, or number from memory, and has
  the branch report only the identifier it actually holds when no record names
  the URL yet, leaving the PR check unarmed until the worker's ready line
  arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for
  main in place of the bare full-URL mandate.

- Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full
  https:// URL wherever a PR is mentioned - status line, terminal, or summary -
  never a bare "PR 108", so the link is in view as early as the number is.

- bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that
  the task's own done lines contradict, printing both spellings; a log naming
  no URL still records the argument as before. fm_pr_status_ready_urls in
  bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches
  bin/fm-pr-merge.sh, so nothing merges under a contradicted URL.

Tests cover the offline refusal with zero side effects, the recorded spelling
being accepted, markdown-wrapped and punctuated URLs, working lines not
counting, the merge wrapper propagation, a self-hosted merge request with no
forge call, the prompt carrying the rule, and the brief carrying the worker
rule.

* no-mistakes(review): Remove stale PR URL enforcement

* no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`

* fix(bin): disable Claude feedback drafts for fleet launches (#3661)

* fix(bin): disable Claude's feedback-draft flow for fleet-launched agents

Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched
Claude crewmate and secondmate, so /bug and /feedback never queue or
submit a bug report on the captain's behalf. feedbackDrafts is the
documented settings key (Claude Code changelog 2.1.247); the
per-launch CLI flag never touches the captain's global settings.json.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts

* no-mistakes(document): Fix Claude feedback documentation formatting

* fix(bin): layer both feedback-draft controls for defense in depth

The prior --settings-only fix can be overridden by a managed Claude
settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0
alongside --settings '{"feedbackDrafts":"off"}': either control alone
disables the SendFeedback tool, so a managed override of one still
leaves the other in force.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(document): Document Claude feedback-draft suppression ownership

* feat(tests): run three more validation families concurrently (#3662)

* perf(tests): admit three more families to concurrent validation

The three families that `docs/fm-test-isolation-proof.md` recorded as refused
were not refused for concurrency. Each blocker was a test that decided a
property by wall clock, or a script filed where it cannot run. Fixing those
three things admits all three families and recovers 28.6 minutes of local
validation with no assertion removed or weakened.

- `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the
  handoff, sleeping a fixed second, then delegating the move to the real
  binary. Nothing ever killed the fake, so on a host slow enough for the case's
  next assertions to take longer than a second, the orphan woke and completed
  the very move the case requires left undone, and recovery then failed with
  `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs
  during the injected crash showed exactly that, the item moving one second
  after the crash. All four crash injections in the file now go through a new
  `fm_fake_crash_injector` shim that signals the target and returns only once
  it is observably gone, and the pre-move fake never delegates the move at all.

- `tests/fm-session-start.test.sh` proved the startup digest does not block on
  a slow current-state read by timing the whole digest against a fixed
  eight-second sleep, which a loaded host exceeds without the property being
  violated. It now holds that read open until the case releases it and asserts,
  the moment the digest returns, that the read has not finished. A digest that
  waited would wait indefinitely rather than for an interval a slow host can
  out-run, so the assertion is stronger than the bound it replaces. Its scan
  budget moves to the maximum, because the old value left two seconds of margin
  over the fixed sleep and measured the host rather than the deadline that
  `tests/fm-inactive-reconcile.test.sh` owns.

- `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all,
  which put it in the portable serial lane, where Linux CI gate-skips it: that
  real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and
  the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its
  opt-in variable and moves to `live-harness-optin`.

The 28 remaining ungrouped scripts become an enumerated `standalone` family
instead of admitting `unclassified` itself. `unclassified` is the family map's
`*)` arm, so admitting it would silently grant concurrency to every test added
afterwards, which is exactly the population with no proof. A new test still
lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh`
covers that split behaviorally.

Each family passes two consecutive four-worker proofs with zero failures. On
the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap`
756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s
recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock
against 121 minutes of summed script time.

* no-mistakes(document): Refresh concurrent validation and shard documentation

* no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2

* feat: structure no-mistakes ask-user escalations (#3670)

* feat(brief): structure no-mistakes ask-user escalation as event + snapshot file

Crewmates escalating a no-mistakes ask-user gate now report one status
event naming every finding id plus a snapshot file holding the gate's
axi finding records verbatim (id, severity, file, line, description,
authority), using the same shape even for a single finding. The status
line never paraphrases. The format is defined once in fm-dod-lib.sh and
rendered into both the scout and ship rule 6 in fm-brief.sh, so a
promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets
the identical contract as a freshly-spawned no-mistakes ship worker.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei

* no-mistakes(review): Preserve ask-user escalation output contract

* no-mistakes(review): Align escalation format test expectation

* no-mistakes(review): Scope ask-user escalation instructions correctly

* no-mistakes(review): Remove ask-user from generic decision rules

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix(bin): require self-sufficient no-mistakes intent (#3671)

* fix(bin): require a self-sufficient no-mistakes intent

A no-mistakes worker's --intent is only as useful as the string it
passes. PR #3604 shipped with an intent that was only "do 1, 2, 3, 7
from the report": the real contract lived in a private scout report and
never reached --intent, so nobody holding that string plus the codebase
could have derived the specification.

This is pure instruction at the contract's one owner; no spawn-side or
promotion-side check is added.

- bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now
  states that the --intent string must be self-sufficient (the string
  plus the codebase reconstructs roughly the same specification) and
  tells the worker to write the substance of any report, decision, or
  PR the captain's intent refers to into --intent rather than the
  pointer, while Firstmate build instructions and the worker's own
  decisions still stay out. The spawn-time overlay points back at that
  rule so its "supersedes" wording cannot cancel it, and the header's
  owner statement carries the rule.
- AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to
  include the substance of referenced material when filling
  ## Captain's intent, and section 11 points at the owner of the rule.
- tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the
  rendered brief and launch contract carry the rule.

Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim

* no-mistakes(document): Replace incident-specific intent test commentary

* fix: accelerate local Bearings snapshot composition (#3499)

* Speed local fleet snapshot composition

* no-mistakes(review): Stabilize task inventory during concurrent snapshot composition

* no-mistakes(document): Document local snapshot observation concurrency

* no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks

* no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks

* no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks

* no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky

* no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks

* fix(snapshot): keep live observations generation-coherent

* no-mistakes(review): Keep secondmate observations generation-bound without copying reports

* no-mistakes(document): Document generation-coherent snapshot observations

* test(bearings): measure local read overlap instead of wall-clock budget

The large-local-snapshot regression asserted that a whole snapshot
composed in under five seconds. That bound measures how loaded the host
is, not whether the per-task reads actually overlap, so it failed
intermittently on a contended machine: one run in six on a box at load
16-20, landing exactly on the five second boundary.

Time a serialized run and a concurrent run of the same workload instead
and require the concurrent one to save at least two seconds. Both runs
pay the same composition overhead, so the difference isolates the
overlap this change delivers. Five one-second reads serialize into five
seconds and overlap into about one, and re-serializing the reads
collapses the saving to roughly zero, so the assertion still fails
loudly if the concurrency regresses.

Also bump the pinned Bearings test count to 48, since rebasing onto the
current default branch picked up its captain-hold test.

* no-mistakes(review): Restore JSON-derived decision flags

* no-mistakes(review): Unify status-derived snapshot observations

* no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass

* fix: prevent stale supervision wake loops (#3672)

* fix(bin): stop the supervision branch's stale-ack and ghost-report loops

Clean-slate implementation of the four authorized recommendations from the
supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal
form, superseding PR #3604:

- fm_branch_report refuses a task the wake being handled never named. The
  extension fixes the reportable task set from the eligible rows before each
  prompt (signal and stale rows resolve to their tasks, a heartbeat allows any
  task with a live record, fleet is always allowed), so a report typed from
  memory about a task whose records teardown already removed is never stored
  or delivered.
- An acknowledgement that consumes nothing says "nothing was acknowledged
  through N" and prints the exact --ack-through / --recovery-generation
  command for the current presented wake, instead of "re-run the drain",
  which re-fed the same stale acknowledgement in a loop.
- bin/fm-guard.sh no longer tells the branch actor to drain queued wakes
  while it is handling them; it names the granted rows instead.
- Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and
  descendants; the index rebuild and the append-side index write both skip a
  task with neither a live record nor a status log, so the branch's report of
  a teardown it just performed is stored without recreating the index.

No new locking, no spawn-generation binding, and no retired-task refusal: the
branch can still report the outcome of a task it just tore down, and the
teardown test now proves that path end to end.

* fix(bin): narrow the branch report scope and guard silence to the minimal form

Apply the four review decisions on the clean-slate branch:

- A signal or stale prompt may report only the tasks its own rows resolve
  to; fleet is refused there too. A heartbeat review is not scoped by task
  at all, so the extension no longer tracks live task records and refuses
  nothing by task id during a fleet review.
- The outcome-index rebuild no longer skips retired tasks; the append-side
  skip alone keeps a torn-down task's index from being recreated.
- bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor
  instead of printing a replacement note.

* no-mistakes(document): Align supervision docs with scoped wake handling

* fix(bin): avoid fleet snapshot argument limits (#3677)

* Fix fleet snapshot large JSON transport

* no-mistakes(review): Captain: file-back fleet snapshot transport safely

* no-mistakes(review): Captain: file-back parent summary aggregation

* no-mistakes(ci): Rebased the PR's three commits onto f4d7875824ecc5e274b4bb896f10c1e1f207b7e4 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor

* fix(bin): attribute active runs with unfetched pipeline heads (#3681)

* fix(bin): recognize active pipeline fix rounds with unfetched run heads

A no-mistakes fix round advances the run head beyond the submitted head,
and the pipeline commits in its own checkout, so the task copy never
receives the new commit object. fm-crew-state's strict head rule rejected
the active row, the coarse runs-list scan skipped it and matched the
older failed row at the submitted head, and an active validation read as
failed (observed on model-routing-benchmark-hardening: active head
ac61c64b vs task copy at fb47636d).

fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns
runs-ledger attribution: the branch's newest row alone decides, and a
newest row whose head cannot resolve locally is recognized only as a
provable pipeline-owned continuation - active (running) and anchored by
the immediately older row for the same branch having ended at exactly
this worktree's HEAD. The reader keeps the axi TOON as full detail for
that proven same-branch run. Unanchored, ancestor-anchored, and terminal
unresolvable rows stay unattributed, so branch-name coincidence and other
tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact
prior semantics for teardown (verified by the full teardown suite).

Tests: reproduction regression for the unfetched active fix head (reads
working via full run-step detail), coarse-path continuation when axi
answers another branch, and negative controls for the unanchored active
row and the unresolvable terminal row with the historical fallback
preserved.

Ported onto upstream/main f4d78758, where #3194 independently added the
branch_sync custody exemption on the full axi-status path: both mechanisms
now coexist, each owning one surface (TOON custody on the full path, the
runs ledger on the coarse path). The port deletes the superseded coarse
scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers
(fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the
exemption comment's "the one exemption" phrasing now that a second
complementary exemption exists, and points the stale
FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree
(judge follow-up #1). The parent coarse-guard test's fixture is the
ledger-anchored continuation shape, so its expectation flips to the fixed
behavior (working via run-step, never the older failed row); a new
mismatched-anchor coarse negative control preserves that guard's original
no-anchor protection (pane answers, never the older row).

* no-mistakes(document): Clarify pipeline attribution documentation

* fix(bin): pre-register claude workspace trust at spawn time (#3663)

* fix(bin): pre-register claude workspace trust for task worktrees

A claude crewmate launched into a fresh task worktree met Claude Code's
interactive workspace-trust dialog before it ever read its brief, and firstmate
could not answer it: the key plane carries only Enter, Escape, and C-c with no
arrow navigation, and the dialog's selection starts on "No, exit", so the
documented Enter recipe ended the session instead of accepting it. Two workers
wedged this way and were unblocked only by hand-seeding the trust store per
path.

--dangerously-skip-permissions does not cover that gate. `claude --help`
records the dialog as skipped only in non-interactive mode, through -p or a
non-TTY stdout, and a crewmate pane is interactive, so there is no launch flag
to reach for.

fm-spawn now pre-registers the worktree through bin/fm-claude-trust.sh in the
existing claude branch, before the project settings that the same gate would
otherwise block, and refuses the spawn when that write fails rather than
launching a worker that would wedge.

The scope test is the safety property and is structural rather than a path
policy: the path must be a linked git worktree, sharing the spawning project's
common dir, whose top level is exactly the resolved argument. Git is the ground
truth, so the argument is never trusted on its own word, and a primary
checkout, an unrelated repo, a worktree subdirectory, a plain directory, and a
home directory are each refused rather than warned about or skipped. A
treehouse or orca path prefix was deliberately avoided because treehouse's root
is configurable, which would make a prefix both wrong and a new policy surface.
One structural test covers both worktree providers.

tests/fm-claude-trust.test.sh pins both halves, including a case where HOME is
itself a valid linked worktree so the home guard is proven load-bearing rather
than passing vacuously, plus the spawn-level proof that a claude spawn trusts
its worktree and launches with the brief pointed at the same store.

The adapter reference no longer tells a firstmate to press Enter on that
dialog, and the shared trust reference now names every harness surface: which
harnesses gate, which suppress at launch, which dodge the gate, which now
pre-registers, and that a claude secondmate is excluded by design.

The spawn fixture runs each spawn against a throwaway HOME so the suite cannot
write the developer's real store, isolating through HOME rather than
CLAUDE_CONFIG_DIR because the spawn forwards a set CLAUDE_CONFIG_DIR onto the
launch command that launch-shape assertions read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v

* fix(bin): create the staged trust store exclusively

The staged store was written to a predictable pid-based path with a plain
write, which follows a symlink. Where the Claude config directory is writable
by another local account, that account could pre-create the path as a symlink
and redirect the write into another file the launching user owns.

The staged name now carries random bytes and is created with an exclusive
"wx" open, so an existing path is refused outright instead of followed. The
happy-path test also asserts no staged store survives the rename.

The durability comment now states the residual window plainly: the readback
proves the entry landed, not that it survives, because a vendor session that
rewrites the whole store afterwards can still drop it and no lock closes that
window when the writer is Claude itself. The worker then meets the dialog and
stalls, which reaches firstmate as the ordinary stale wake rather than as
silent success.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v

* no-mistakes(review): neutralise CDPATH in claude trust scope guard

* no-mistakes(review): sandbox HOME in spawn tests, drop out-of-scope artifacts

* no-mistakes(review): refuse unresolvable git dir, compact store, fix secondmate doc

* no-mistakes(review): clear git env overrides, resolve symlinked store target

* no-mistakes(review): degrade without node, fix Pi gate claim, record trust proof

* no-mistakes(review): refuse without node, pin CLAUDE_CONFIG_DIR in spawn tests

* no-mistakes(review): refuse relative config dir and concurrent store modification

* no-mistakes(review): correct orca worktree claim, clean staged store on failure

* no-mistakes(review): restore pretty-printed store, correct trust dialog docs

* no-mistakes(review): arm trust gate before busy state to avoid orphans

* no-mistakes(document): record claude trust pre-registration in its owner docs

* no-mistakes(document): note orca limit for claude trust pre-registration

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-spawn.sh by moving the Claude trust gate earlier rather than adding cleanup machinery. Diagnosis: Greptile reported that when Claude trust registration fails on tmux/Zellij/cmux/non-projected Herdr, the exit runs after the backend endpoint and /tmp/fm-<id> were created, and the abort trap cleans neither. The endpoint half is pre-existing, deliberate architecture — the two refusals immediately above the gate (the 60s `treehouse get` timeout at fm-spawn.sh:2550 and `validate_spawn_worktree` at :2487) also exit with the endpoint live and direct the operator with "inspect window $T"; spawn_abort_cleanup only reclaims orca endpoints (already covered via ORCA_ABORT_CLEANUP) and herdr projections. The temp-root half was genuinely introduced by this PR: the gate was placed beside the busy-state arm, ~30 lines after `mkdir -p "$TASK_TMP/gotmp"`, and fm-teardown can only find that root through `tasktmp=` in a meta record a refused spawn never publishes. Root-cause fix (smallest correct change, no new subsystem): - bin/fm-spawn.sh — moved the `claude*` trust gate from inside the busy-arm block up to the first point $WT is known, immediately after the `freshen_spawn_worktree_base` block and before TASK_TMP creation, the STATE setup, and the relaunch `clear_relaunch_harness_wiring` retirement. A refusal now leaves no temp root, no retired relaunch wiring, and no busy record; only the endpoint remains, in the same class as the two refusals just above it. - bin/fm-spawn.sh — the refusal message now ends with "inspect window $T", matching the existing convention so control/teardown can identify the endpoint. $T is set for every backend on the non-secondmate path. - bin/fm-spawn.sh:196 — header note corrected from "before any state is armed" to "before any per-task state exists". - tests/fm-claude-trust.test.sh — the existing refused-spawn test's own comment claimed "before any task state exists" but only asserted busy state. Renamed to test_refused_spawn_leaves_no_task_state and added an assertion that /tmp/fm-<id> is absent, with the task id suffixed by the test process pid so the assertion reads only this run's path (a stale /tmp/fm-refusedspawn from the fixed-id version was in fact present on this box). No assertions on implementation source bytes. Verification run locally: - The new assertion fails against the pre-fix bin/fm-spawn.sh ("not ok - a refused spawn stranded a temp root no teardown can find") and passes after — a real before/after regression proof. - tests/fm-claude-trust.test.sh: 20/20 ok. - tests/fm-backend.test.sh, fm-backend-orca, fm-control-relaunch, fm-spawn-dispatch-profile, fm-trace-context-spawn, fm-gotmp: all pass. - tests/fm-backlog-atomicity.test.sh: rc=0, 79 assertions ok. - bin/fm-lint.sh (repo's single lint owner, pinned ShellCheck 0.11.0 + actionlint 1.7.12): clean. - No /tmp/fm-refusedspawn* leftovers after the runs. Scope respected: no trust subsystem, no policy layer, no config surface, no endpoint-cleanup mechanism added; the change is an ordering move plus one error-message clause and the test that pins it. Adapter references and docs made no ordering claim, so none needed updating. Changes are left uncommitted in the worktree for the outer executor

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix: restart every live second mate after updates (#3690)

* feat(update): restart every live second mate after a successful update

/updatefirstmate only restarted a second mate when that pass advanced its
AGENTS.md or .agents/skills. An already-current home was skipped entirely, a
bin/-only advance was steered instead, and a remote host that could not report
its instruction diff was downgraded to a re-read. A running agent also freezes
its launch-time wiring - turn-end hooks, harness flags, per-harness feature
switches - and none of that is derivable from a file diff, so an unchanged
tracked surface is not evidence the agent is already on the current behavior.

Restart is now unconditional on a successful update of that home. Every live
second mate the pass leaves on the target commit is restarted, whether it
advanced or was already there.

The safety contract is unchanged: open records are persisted before the agent is
replaced, nothing is forced, stashed, or discarded, a home the pass had to skip
is not restarted at all, and a mate whose runtime cannot prove a restart keeps
the honest re-read path and is never reported as reloaded.

bin/fm-ff-lib.sh gains a settled-state hook that fires for a home left at the
base whether it advanced or was already there, and never for a skipped one; the
instruction-gated hook the session-start convergence sweep uses is untouched.

Regressions: fm-update pins the already-current mate into the restart set and
the unprovable one into the nudge set, and fm-secondmate-restart drives both
real commands end to end - an already-current home is named, persisted, and
genuinely replaced with its checkout untouched, while the unprovable one keeps
its running agent.

* no-mistakes(document): Document unconditional secondmate restarts

* fix(bin): close pending-reply decisions via resolve-key (#3696)

* fix(bin): close reserved pending-reply keys via fm-send --resolve-key

fm-send wrote answered: notes that the reserved-key fold ignores, so
operator closes exited 0 while OPEN DECISIONS kept the decision open.
Speak the owning library's close vocabulary on that path, and refuse
when a reserved close cannot take effect.

* no-mistakes(review): Safely quote manual decision-close recovery commands

* no-mistakes(review): Reject unclosable overlong decision keys before sending

* no-mistakes(review): Remove contract suffix from open decisions hint

* no-mistakes(document): Document resolve-key line-cap refusal

* fix(bin): prevent false missed-reply escalations (#3697)

* fix(bin): stop false missed-reply escalations for same-basename self-home answers

A healthy secondmate that wrote corr= to its own state/<id>.status never matched the parent channel, so recovery confirmed and the record escalated as pending-reply-missed. Make the report helper resolve the parent channel itself, skip parent-replies.status as wrong-home, put a readable sighting path on the missed line, and restatement-copy only that same-basename self-home file onto the parent channel.

* no-mistakes(review): Resolve late replies before recovery escalation

* no-mistakes(review): Tighten reply routing and regression coverage

* no-mistakes(review): Preserve reply paths and require explicit home

* no-mistakes(review): Encode wrong-home paths before persistence

* no-mistakes(document): Document corrected secondmate reply routing

* no-mistakes(lint): Fix pending-reply ShellCheck warnings

* feat: add verified Gemini crewmate runtime (#3695)

* feat(harness): verify gemini as a crewmate runtime adapter

Adds Gemini CLI as a fourth dispatch target alongside claude, codex, and
grok, scoped to crewmate and scout work only. Every axis was proven against
gemini-cli 0.58.0 rather than inferred; docs/verification/runtime-backends.md
carries the dated evidence and names what stayed unverified.

Busy state is semantic, not rendered: BeforeAgent opens a turn and AfterAgent
and SessionEnd close it. AfterAgent also fires on a manual interrupt, so a
cancelled turn closes its own record.

Three findings shaped the wiring rather than a config line:

- --skip-trust and GEMINI_CLI_TRUST_WORKSPACE=true are presented by the CLI
  as equivalents and are not. A controlled A/B showed --skip-trust leaves
  project configuration unloaded, so workspace skills never load.
- The worktree's .gemini/settings.json is the PROJECT's committed settings
  file, unlike claude's settings.local.json. Firstmate's hooks therefore go
  to a firstmate-owned state/<id>.gemini-settings.json reached through
  GEMINI_CLI_SYSTEM_SETTINGS_PATH, which also works untrusted and merges
  with a project's own hooks instead of replacing them.
- The shipped CLI is a node bundle whose live process reports comm as
  MainThread, so ancestry cannot see it. GEMINI_CLI=1 is load-bearing and is
  tested before an inherited CLAUDECODE, and pane liveness identifies gemini
  from the script argument through the new bin/fm-gemini-lib.sh.

Gemini is refused for secondmates: it has no primary supervision protocol.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L

* test: clear gemini's marker in launch and detection expectations

Every non-gemini launch now clears GEMINI_CLI the way it already clears
cursor's markers, so the two tests that pin the exact launch prefix are
updated to match. The harness-detection tests that scrub foreign markers
before probing ancestry scrub GEMINI_CLI too, so running the suite from
inside a gemini session cannot produce a false verdict.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L

* docs: classify the gemini harness reference

The documentation inventory is the single classification owner for maintained
prose surfaces, and every surface must appear in it exactly once. The new
harness reference is agent-runtime, matching its siblings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L

* no-mistakes(review): Narrow Gemini ancestry detection

* no-mistakes(review): Restrict Gemini hooks to canonical launches

* no-mistakes(document): Document Gemini adapter support boundaries

* no-mistakes(ci): Fixed Gemini process identity when interpreter or script paths contain whitespace. Tmux liveness now uses NUL-delimited /proc argv on Linux, with the existing flattened ps fallback elsewhere. Added a real-process regression test. Verified with the Gemini harness test suite, full fm-lint, ShellCheck, and git diff --check. The CI and Require no-mistakes runs were action_required/attestation outcomes rather than code failures

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* fix(teardown): conclude parked runs advanced past task copy (#3704)

* conclude parked runs the pipeline advanced past the task copy

A no-mistakes fix round commits in the daemon's own gate-repo clone, so a
run parked at a gate can carry a head whose object the task copy never
received. Teardown's strict object-local identity rule then declined to
conclude the run, and cleanup left it parked forever holding a fleet slot
(observed 2026-09-03; the same masking condition PR 3681 fixed on the
read path, now closing the teardown half its scope boundary deferred).

task_status_is_own_parked_run now falls back - only when the reported
head resolves to no local object - to the one shared runs-ledger
attribution rule fm_nm_runs_status_for_worktree (bin/fm-nm-run-lib.sh),
whose anchored continuation proof binds the branch's newest active row
to this worktree's exact submitted head. Foreign branches, stale
history, terminal rows, ancestor-only anchors, diverged newer rows, and
ambiguous multi-row shapes all still refuse, and runs that are actively
running, fixing, or in CI remain untouched: only the parked-at-a-gate
determination ever reaches the abort. No sqlite access, no fetches into
another task copy, no custody changes, no duplicated matching logic.

* tighten the parked-run ledger fallback and pin both judge corrections

The teardown ledger fallback now authorizes concluding this task's parked
run only when the shared runs-ledger rule's proved answer is the explicitly
active word (running): a terminal newest row - even anchored at exactly the
worktree's head - is finished history and never an abort authorization.
The read path may classify the same owner's answer; teardown's abort must
never fire for a run that already ended.

Two bounded pre-validation corrections from the implementation review:
- a fetched-object counterfactual pins the strict-rule path: a pipeline fix
  head fetched into the task copy aborts through object-local identity
  alone, with an empty ledger and a proof the runs query never fired;
- a negative fixture pins the tightened boundary: an unresolvable reported
  head with a terminal newest same-branch row anchored at the worktree head
  engages the ledger fallback and still refuses, so the refusal is the
  terminal-word boundary and not an earlier guard.

* no-mistakes(review): Bind teardown ledger fallback to validated run heads

* no-mistakes(review): Restore validated advanced-head ledger continuation

* no-mistakes(review): Reject invalid ledger dates and terminal statuses

* no-mistakes(document): Document teardown ledger scan limit

* fix: classify captain holds from structured state (#3508)

* Keep parked and aged undated captain holds off live Captain's Call.

Bearings was treating undated parked-style holds as live calls; mark those phrasings deferred and project holds older than a configurable 14-day since date as Charted Next gates instead.

* no-mistakes(review): Bound parked marker matching to lexical tokens

* no-mistakes(review): Age undated holds from durable hold-set dates

* no-mistakes(review): Reset re-held timestamps and scan full bodies

* no-mistakes(review): Preserve timestamp precision and prioritize parked suppression

* no-mistakes(document): Document undated captain-hold aging

* no-mistakes(ci): Fixed stock Bash CI test-count expectations (16 snapshot, 45 Bearings). Prevented fresh holds on old tasks from aging via stale `since` dates by aging only stamped holds. Added behavioral regressions and verified both suites plus Bash 3.2 parsing

* no-mistakes(ci): account for rebased snapshot regression

* no-mistakes(review): Restore legacy hold aging and mandate wrapper

* no-mistakes(review): Restrict hold stamps to canonical leading lines

* no-mistakes(review): Exclude historical answers and deduplicate revealed holds

* no-mistakes(document): Correct captain-hold projection documentation

* no-mistakes(ci): Rebased onto 8988af2 and resolved Bearings conflicts. Fixed the hold timestamp race by persisting and verifying the timestamp before publishing the captain hold; failures now leave the task unheld. Added behavioral coverage for ordering and failure handling. Preserved the required parked-phrase projection behavior. Relevant snapshot, Bearings, lifecycle, syntax, and ShellCheck validations pass

* no-mistakes(review): Bound current prose before historical resolutions

* no-mistakes(review): Preserve hold age across interrupted answers

* no-mistakes(review): Preserve leading hold stamps until answer closure

* no-mistakes(review): Normalize answer bodies on matching retries

* no-mistakes(review): Document concurrent re-hold age-basis limitation

* no-mistakes(document): Refresh captain hold lifecycle documentation

* no-mistakes(ci): Fixed both CI failures. Updated the macOS Bash snapshot expectation from 45 to 46 Bearings tests. Narrowed parked-style deferral matching to explicit hold-reason prefixes while preserving legacy explicit markers and preventing contextual prose from hiding active decisions. Added behavioral regression coverage. Verified with stock Bash 3.2: 17 fleet snapshot tests and 46 Bearings tests pass; full lint and workflow validation also pass

* no-mistakes(ci): Fixed Greptile’s P1 finding by restricting parked-style deferral phrases to complete hold-reason markers. Contextual reasons beginning with “not urgent,” “queued opportunity,” or “captain-gated” now remain visible decisions. Added behavioral coverage through the real fleet and Bearings snapshot paths and updated documentation. Verified both snapshot suites under Bash 3.2 (17 fleet tests and 46 Bearings tests), syntax checks, and git diff checks. The no-mistakes attestation failure is external/stale and requires the outer pipeline to refresh it for the new head

* no-mistakes(ci): Fixed parked-style undated captain holds disappearing from the default Bearings board. They now project to Charted Next with omitted[] disclosure, while --all-decisions reveals them and removes the safety gate. Added behavioral coverage for the reported “not urgent” case and aligned documentation. Verified fm-bearings-snapshot, fleet snapshot view, and captain-hold lifecycle tests; shellcheck, bash syntax, and git diff checks pass

* no-mistakes(test): Stabilize concurrency budget and provision timeout tests

* no-mistakes(document): Correct captain hold documentation details

* no-mistakes(ci): Fixed hold-reason parsing so commas in contextual reasons are preserved and do not incorrectly defer live Captain's Call decisions. Added end-to-end fleet/Bearings regression coverage. Reworked the flaky Herdr timeout test to assert observable late-launch behavior rather than process-ID liveness. Verified both snapshot suites, Herdr test 5 consecutive times, shell syntax, shellcheck, and git diff checks

* Restore the Herdr lab timeout test to its main version.

The stabilization rounds reworked tests/fm-herdr-lab.test.sh while chasing a
load-induced flake, replacing the fake server's wall-clock delay with a
SIGSTOP'd process and asserting that the blocked process is gone after a
timed-out provision. A stopped process does not die from SIGTERM, so that
assertion fails on Linux and the portable parallel shard stayed red.

That test is unrelated to the undated captain-hold projection this branch
delivers and was identical to main before these rounds, so restore main's
version exactly. It still proves that a timed-out provision cancels its late
launch before teardown.

* no-mistakes(review): Preserve metadata-like prose in captain hold reasons

* no-mistakes(review): Resurface due dated captain holds

* no-mistakes(review): Distinguish parked holds from explicit deferrals

* no-mistakes(review): Invalidate legacy secondmate summary caches

* no-mistakes(review): Keep blocked deferred holds in Charted Next

* no-mistakes(review): Count blocked deferred holds in omission disclosure

* no-mistakes(document): Correct captain-hold projection documentation

* no-mistakes(ci): Fixed both CI failures. Updated the macOS Bearings test count to 51. Preserved the v1 summary schema for compatibility while rejecting hold-bearing summaries missing the new aging fields, preventing stale caches from restoring noisy calls. Verified fleet snapshot, Bearings snapshot (51 tests), home-summary refresh, secondmate reconciliation, Bash 3.2 parsing, and diff checks

* no-mistakes(ci): Fixed Greptile’s valid finding: `--all-decisions` now reveals deferred/aged captain holds even when blocked, for both main and secondmate homes, and removes their duplicate Charted Next gates. Added behavioral regression coverage and updated documentation. The prose-classifier finding was not applied because exact complete-phrase matching is explicitly required by the author intent; contextual wording remains live. Verified with Bearings and fleet snapshot tests, `bin/fm-lint.sh`, Bash syntax checking, and `git diff --check`

* no-mistakes(ci): Fixed the actionable-state bug in Bearings: an arrived parked-style hold is live only when it is not explicitly non-actionable, so blocked due holds remain gated by default and are revealed by --all-decisions. Added behavioral regression coverage for that case. Preserved complete-reason parked-style classification as required by the author intent. Verified with tests/fm-bearings-snapshot.test.sh, bin/fm-lint.sh, and git diff --check

* Show why a revealed captain hold is deferred.

Under --all-decisions a deferred hold is revealed and its Charted Next gate
is removed, but the revealed row carried only the bare hold reason. A
date-deferred or blocked hold therefore read exactly like a genuine live
decision, because the until date, the age, and the blocking work only ever
appeared on the gate row that the reveal replaces.

Annotate a row that is revealed because it is deferred with the same
vocabulary the gate uses - until <date>, held <n>d, and the blocking work -
so the expanded view reads as deferred-but-shown. A genuinely live call is
left unannotated, and the default board is unchanged.

* Classify captain holds from structured fields alone.

Bucket membership was decided by several independent expressions, and two of
them matched hold reason or body prose. That produced a recurring class of
defects: holds that fell through every bucket and vanished from the board, and
live decisions silently suppressed because their wording happened to contain a
marker word - a reason of "non-deferred release choice" matched DEFERRED and
disappeared.

Replace all of it with one total classifier over structured fields only:
hold_kind, state, hold_until, unresolved_blocker_ids, and the machine-written
hold-set timestamp. Every captain hold gets exactly one hold_bucket - blocked,
dated, aged, or live - so no hold can fall through and none can match two.
captain_actionable is exactly the live bucket, and the --all-decisions reveal
is a property of the bucket rather than a second filter.

No hold reason or body prose is matched anywhere in the projection, so wording
can no longer hide, reveal, or reclassify a decision. A hold that is superseded
or no longer required is closed through the hold lifecycle instead of lingering
as an open hold flagged by a keyword.

* no-mistakes(review): Preserve working captain holds across bucket surfaces

* no-mistakes(review): Reject pre-classifier secondmate summary caches

* no-mistakes(review): Preserve complete live hold summaries

* no-mistakes(review): Clarify working hold decision bucket semantics

* no-mistakes(review): Reveal bounded remote holds and preserve blocker notes

* no-mistakes(review): Make blocker overflow explicit in hold summaries

* no-mistakes(document): Correct captain-hold projection documentation

* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 17 to 18 tests. Verified the suite under Bash 3.2.57: all 18 tests pass. `git diff --check` also passes

* no-mistakes(ci): Updated the stock macOS Bash CI expectation from 51 to 53 Bearings tests. Verified all 53 pass under Bash 3.2.57; git diff --check passes

* fix(pi): keep supervision outcome delivery responsive (#3767)

* fix(pi): deliver supervision outcomes off Pi's render thread

The supervision branch runs inside the captain's own Pi process, and Pi
runs extensions, their tools, and their event handlers on the single
JavaScript thread that also draws the TUI and reads the keyboard. Every
delivered outcome ran roughly five bash script invocations plus several
`ps` calls through spawnSync on that thread, so the TUI could not repaint
or echo a keystroke for the whole chain - the subsecond freeze the
captain saw every time a routine or captain-facing outcome arrived.

Convert the delivery path's subprocess calls to an awaited spawn behind a
serializing queue. lib/fm-async-exec.ts is the single owner of the
awaited-spawn replacement and returns the same capture shape and failure
verdicts spawnSync returned. Awaiting yields the thread, so what the
single thread used to guarantee for free is now an explicit queue: every
delivery, acknowledgement, and turn-boundary reconciliation runs as one
unit of it, preserving the durable append before anything visible, one
delivery at a time in sequence order, the read cursor advanced before the
next reader sees a row, and one ownership activation per generation.
Cancellation is preserved by the generation and lock-ownership rechecks
the awaits are placed around.

Two reads stay synchronous because Pi's own API is synchronous there, not
as an optimization: its bash spawn hook is typed as a plain function, and
the watcher reads offer.accepted the moment its dispatch event returns, so
a session that does not own the fleet lock must still refuse a wake
without waiting. Both walk the lock's process ancestry in full every time,
never cached, because reparenting and pid reuse can invalidate a
remembered chain and that answer decides ownership rather than hinting at
it. The store scripts and their durability contracts are unchanged.

Measured through the real fm_branch_report tool and real bin/ scripts with
a 1 ms interval timer, the largest block of the JS thread falls from 273
to 2.0 ms for a routine outcome, 286 to 2.0 ms for a captain outcome, and
134 to 1.9 ms for main's acknowledgement, against a 1.3-2.2 ms idle floor.
In a real Pi 0.82.0 TUI the worst keystroke echo while two outcomes arrive
falls from 676.9 ms to 36.8 ms, against a 22.6 ms extension-free floor.

Regressions: a delivery must leave the event loop running (zero timer
ticks before this change, in 250 ms), interleaved reports stay ordered and
exactly once, a session replaced mid-delivery neither loses nor duplicates
an outcome, and a failing store script surfaces without losing or doubling
one. The real-TUI half is an opt-in live guard that types into an isolated
Pi pane while outcomes are delivered and fails if echo leaves the class of
the same machine's own floor.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013bzoWyr2EcJGBKuoUjVRSp

* no-mistakes(review): Revalidate ownership and bound asynchronous subprocess output

* no-mistakes(ci): Fixed CI defects: routine outcomes now persist a sequence-keyed delivery receipt before awaiting cursor advancement, preventing duplicate delivery after mark-read failure. Corrected the session-replacement test to exercise an actual asynchronous ps ancestry lookup. Targeted behavioral tests, strict Pi typecheck, ShellCheck, and diff checks pass. The full extension test remains locally blocked by an unrelated stock-render assertion under the installed Pi runtime. The no-mistakes attestation failure is external pipeline state (test was previously skipped), not a source defect

* fix(pi): keep the declined routine receipt out and skip the renderer case below its Pi floor

Four follow-ups on the same branch, plus one revert.

Revert the routine-delivery receipt a CI auto-fix round added. It introduced a
new persisted `fm-branch-routine-delivery` entry, written into the captain's
transcript for every routine note, to deduplicate a note whose cursor write
failed. That is a change to the delivery contract, which this task is not
authorized to make: the approved work is the asynchronous conversion with the
existing durability contract preserved. The ownership re-read and output
bounding from the review round are kept - both are genuine asynchronous
correctness, not contract changes - as is that round's use of a real parent pid
so the replacement regression traverses an actual ps subprocess.

Record the routine gap instead of closing it. A routine note is a plain
message with no sequence-keyed record, so a mark-read failure after delivery
makes the next reconciliation send it once more; a captain row cannot
duplicate that way because its visible entry is found by store sequence. That
asymmetry predates moving delivery off the render thread. It is now stated at
the call site and in the delivery-contract docs, tracked as
fm-pi-routine-delivery-idempotency-followup-r1, and pinned by a regression
that proves the routine note is re-delivered exactly once more and never
again, the captain entry stays single, and the store keeps both rows.

Give the stock-renderer case a Pi version floor. It compares the extension's
renderers against Pi's stock rendering, so its verdict only means anything
against the contract those renderers target: since 0.84.4 the stock renderer
no longer supplies an implicit reset at multiline boundaries and the extension
emits that reset itself, so an older installed Pi differs legitimately. It now
names the installed version and the floor and skips, while a package whose
version cannot be read at all still fails.

Make the responsiveness regression's second signal a fraction rather than a
millisecond budget. A loaded machine that deschedules the process inflates an
absolute stall budget into a false failure, but it inflates the delivery's own
wall time too, so requiring the worst stall to be a minority of that wall time
holds under load. Synchronous delivery sits near 1.0 there whatever the load,
and the tick-count signal still reads zero on it.

Replace the test-family mapping for the Pi extension libraries with per-script
targeting. Routing them to whole families - or leaving them unmapped, which
widens through the reference scan to each referencing suite's entire family -
selected dozens of suites with nothing to do with Pi and pulled an unrelated
flake into the run. The changed-file selection drops from 112 scripts to 61.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013bzoWyr2EcJGBKuoUjVRSp

* no-mistakes(document): Clarify asynchronous execution documentation

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix: avoid duplicate AGENTS.md governance for marked projects (#3763)

* fix(memory): honor explicit project maintenance guidance

* no-mistakes(test): Blocked by pre-existing Bash and Muse fixture failures

* no-mistakes(test): Remove accidentally tracked test attribution report

* no-mistakes(ci): Restricted the marker to the exact first line, preventing fenced examples from suppressing governance, and corrected the documentation. Regression failed before the fix; all 18 helper tests, focused ShellCheck, documentation validation, and diff checks pass. CI and Require no-mistakes report action_required with zero jobs executed; those external checks remain unresolved

* fix: protect primary checkout when spawning from linked homes (#3783)

* fix(spawn): refuse repository primary from linked spawning homes

Compare the resolved task git directory with the spawning repository's
common git directory before refreshing a fresh copy or relaunching a task.
This protects the primary even when the spawning project is a linked home.
Keep pooled copies accepted and preserve recorded work on relaunch.

Fixes #3741.

Verification for the pipeline PR body:
- Red on origin/main 1820316b66ac2c68e244dd04a02512859ee8c1f4 with the new
  regression and unchanged production code: bin/fm-test-run.sh
  tests/fm-spawn-pool-base-freshen.test.sh exited 1 with
  "linked spawning home accepted primary as a disposable copy".
- Green after the guard: the complete pool-base-freshen and control-relaunch
  suites passed through bin/fm-test-run.sh, covering primary and symlink
  refusal before fetch/reset, spawning-directory refusal, scout acceptance,
  and committed plus unfinished work preserved during linked-home relaunch.
- The worktree-settle suite passed on pristine main and the final branch.
  An earlier loaded-host run exceeded its five-second assertion (6s);
  the final retry passed without changing code or the assertion.
- Test fixture commits ran with GIT_CONFIG_COUNT=1,
  GIT_CONFIG_KEY_0=commit.gpgsign, GIT_CONFIG_VALUE_0=false.
- bin/fm-lint.sh and /bin/bash -n for all three changed scripts passed.

The upstream cwd-selection cause remains outside this change.

* no-mistakes(document): Clarify spawn isolation ownership and relaunch preservation

* fix(bin): stop reading an unanswered backend probe as a dead endpoint (#3785)

* fix: read a failed herdr CLI as unreachable, not a gone backend target

The no-run fallback in bin/fm-crew-state.sh collapsed every failed pane
capture into 'backend target gone', which downstream consumers treat as
positive death evidence - so a herdr CLI that errors or stalls under load
briefly scored dozens of live claims dead on a busy box. Only a successful
herdr answer proving the pane absent (fm_backend_agent_state's 'missing',
backed by pane get answering pane_not_found) may now read as gone; every
other verdict reports 'backend unreachable' with the endpoint state, which
is never positive death evidence. Adds a behavior test: an always-failing
fake herdr reads unknown/unreachable, never gone.

* test: pin the herdr suite's ambient home to a marker-free fixture

FM_HOME defaults to the suite's own root when unset, and any secondmate-
marked checkout (every treehouse crew home carries .fm-secondmate-home)
flips the default workspace label to 2ndmate-*, so the ambiguous-label
placement test found zero firstmate matches and fell into the create path
instead of refusing (expected exit 3, got 1) - deterministically green in
CI, deterministically red from a crew home. Export a marker-free ambient
FM_HOME fixture; per-test FM_HOME prefixes still override it.

* fix: classify herdr endpoint answers instead of every non-missing verdict

Review decision (firstmate, 2026-09-05): a failed pane capture is not
itself evidence of death, but neither is every non-missing classifier
verdict a failed answer. missing (pane get answered pane_not_found) and
dead (pane present, agent_not_found husk) keep gone-class text so a
stale-claim sweep may still reclaim them; an alive answer falls through
to the normal busy/state flow instead of being discarded when only the
heavy 200-line scrollback read failed; only when the cheap pane get /
agent get calls themselves fail to answer does the line read
'backend unreachable'. Adds the two missing cases: alive with a failed
scrollback read stays live, and a husk pane still reads gone.

* no-mistakes(review): route tmux through agent-state classifier; drop test stall

* no-mistakes(review): narrow inaccurate tmux socket and alive-arm fallback comments

* no-mistakes(document): document classifier-backed endpoint verdicts in crew-state contract

* fix(bin): preserve subshell lock ownership on Bash 3.2 (#3789)

* fix: distinguish subshell wake-lock owners on stock Bash

Restore distinct process ownership for issue #3743 using the existing PID helper, consistently across lock publication, reclaim, release, role checks, and bounded handoff.

The existing wake-queue regression fails on pristine upstream Bash 3.2 with rc=13. The complete suite now passes on Bash 3.2.57 and Bash 5.3.15, with added coverage for ownership when BASHPID is unset. Canonical lint and stock-Bash syntax checks pass.

* no-mistakes(document): Correct lock grace-period documentation

* no-mistakes(ci): Captain, fixed all 14 SC2031 false positives with nine ShellCheck source-boundary annotations across three tests. Full CI-mode lint and the complete wake-queue suite on stock Bash 3.2 passed. Runtime behavior is unchanged

* fix(bin): resolve captain holds and legacy teardowns on non-markdown backends (#3782)

* fix(bin): close legacy records on the Beads backend honestly

Two pre-Beads reads blocked honest closure of leftover records:

1. fm-captain-hold.sh complete/verify resolved attested legacy hold ids
   only against the live backend and the pre-collapse derived identity, so
   a home whose holds fm-hold-migration rehomed under fm- ids failed with
   an empty-name absence message (the resolve failure was swallowed by the
   command substitution feeding verify_hold_durable). Resolution now falls
  …
jorguez96 added a commit to jorguez96/firstmate that referenced this pull request Sep 25, 2026
…ts resolved (#19)

* fix(bin): support process events under symlinked homes (#3484)

* fix(bin): resolve process-event state roots before validating them

The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.

Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.

This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.

* fix(bin): pin the external capture staging boundary to its physical path

The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.

The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.

* no-mistakes(review): Propagate canonical process-event state roots

* no-mistakes(review): Propagate canonical state to process-event adapters

* no-mistakes(document): Document physical process-event state roots

* fix(pi): deliver captain outcomes as deterministic transcript entries (#3312)

* fix(pi): persist captain outcomes visibly

* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition

* no-mistakes(document): Document cold-start captain-outcome recovery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(pi): process captain outcomes through a sequence-keyed turn

PR #3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.

The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.

Add the processing half on top of the persistence half:

- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
  cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
  only advances through an explicit sequence-bound acknowledgement, never
  past the read cursor and never backwards; an absent marker reads as zero
  and `processed-init` migrates delivered history once so an upgraded home
  is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
  every still-unprocessed captain row to main as one hidden, typed
  `fm-branch-process` request listing each `[seq N] task: summary`, opening
  exactly one main turn. Main closes it only by calling the new
  `fm_branch_processed` tool with the highest sequence listed. An unrelated,
  empty, or paraphrased answer leaves the sequence open, and the same request
  is presented again at the end of the next main run and at session start.
  The first two presentations of a sequence set open a turn of their own;
  after that the request rides the captain's next prompt so an ignored
  request cannot loop, and a session replacement resets that budget.
  Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
  scripts: an empty answer and an unrelated prior answer neither advance the
  marker nor stop re-presentation, the acknowledgement is refused beyond the
  read cursor and outside lock ownership, a partial acknowledgement keeps the
  newer sequence open, and #3312's own assertions now forbid an unkeyed turn
  rather than any turn. The store suite pins the marker's bounds and the
  migration; the real-SDK guard for appendEntry persistence and model
  exclusion is unchanged.

Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.

* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements

* no-mistakes(review): Harden outcome state validation and request pacing

* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores

* no-mistakes(review): Validate canonical mark-read cursor state

* no-mistakes(review): Guard cursor advancement against corrupt processed state

* no-mistakes(review): Bind acknowledgements to active processing requests

* no-mistakes(review): Reset pacing when processing sequence membership changes

* no-mistakes(review): Enforce silent outcome invariants at storage boundary

* no-mistakes(document): Document hardened captain outcome processing contracts

---------

Co-authored-by: kunchenguid <kun@kunchenguid.com>

* feat: add bounded concurrent Bearings ledger collection (#3481)

* feat: bound Bearings remote ledger collection

* no-mistakes(review): Clarify default remote-ledger collection behavior

* no-mistakes(review): Detach reconcile delivery from watcher loop

* no-mistakes(review): Enforce bounded snapshot and request captures

* no-mistakes(review): Bound legacy summary capture before parsing

* no-mistakes(review): Bound primary remote ledger captures

* no-mistakes(document): Correct snapshot and reconcile documentation

* no-mistakes(lint): Fix ShellCheck quoting in bounded collector

* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks

* test: await reconcile request retirement

* no-mistakes(review): Avoid empty reconcile queue process churn

* no-mistakes(review): Read ledger summaries from immutable snapshots

* no-mistakes(review): Reject multi-document home ledger streams

* no-mistakes(review): Coalesce durable reconcile requests per target

* no-mistakes(review): Unify reconcile keys and reject snapshot streams

* no-mistakes(review): Key reconcile requests by stable target ID

* no-mistakes(document): Document per-target reconcile request coalescing

* no-mistakes(lint): Remove unused snapshot summary file variable

* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass

* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks

* ci: rebalance portable serial test shards (#3489)

* fix(ci): rebalance the portable serial shards on measured durations

The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.

Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.

Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.

Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.

No test changes what it asserts and no test stops running; only the
partition across shards changes.

* no-mistakes(document): Clarify conservative shard timing aggregate

* fix(pi): fall back on incomplete supervision branch prompts (#3491)

* fix(pi): fall back after settled branch errors

* no-mistakes(review): Detect provider errors across prompt compaction

* no-mistakes(review): Preserve in-flight branch state across selection changes

* fix(pi): re-probe supervision branch after cooldown (#3497)

* fix(pi): recover supervision branch after cooldown

* no-mistakes(review): Defer branch recovery until prompt settlement

* no-mistakes(document): Clarify supervision cooldown recovery contract

* fix(bin): remove legacy remote snapshot reads (#3501)

* refactor: remove legacy remote summary reads

* no-mistakes(document): Document ledger-only snapshot reads

* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass

* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean

* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux

* fix(pi): preserve watcher continuity across session replacement (#3498)

* fix(pi): rearm watcher after session replacement

* no-mistakes(review): Queue actionable closes across Pi session replacement

* no-mistakes(review): Stop replacement arm when handoff persistence fails

* no-mistakes(review): Preserve actionable wakes through branch and late child races

* no-mistakes(review): Surface late handoff failures without crashing Pi

* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens

* no-mistakes(review): Retry stale deliveries and release settled claims

* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup

* no-mistakes(review): Deduplicate persistent handoff cleanup alerts

* no-mistakes(review): Acknowledge watcher follow-ups only when consumed

* no-mistakes(review): Persist idle follow-ups until agent consumption

* no-mistakes(review): Preserve pending outcomes when handoff persistence fails

* no-mistakes(review): Arm replacement before awaiting prior delivery settlement

* no-mistakes(review): Adopt pending handoffs after lock reclamation

* no-mistakes(review): Prevent stale generations from adopting replacement handoffs

* no-mistakes(review): Scope replacement handoffs by watcher state

* no-mistakes(document): Clarify replacement handoff documentation

* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks

* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes

* no-mistakes(document): Document watcher-owned replacement handoffs

* no-mistakes(document): Verify replacement handoff documentation

* test(pi): cover watcher-owned branch fallback

* no-mistakes(document): Refresh watcher-owned fallback documentation

* fix(bin): resurface task statuses missed by wake handling (#3495)

* fix(bin): resurface terminal statuses lost after branch handling

* test(watch): canonicalize process-event fixture homes

* no-mistakes(review): Index branch outcomes by causal status position

* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses

* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics

* no-mistakes(review): Keep unclassifiable oversized statuses silent

* no-mistakes(document): Document lost-wake outcome backstop

* no-mistakes(document): Update outcome backstop documentation

* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally

* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes

* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift

* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state

* fix(bin): collect follow-up results from remote work homes (#3503)

* fix(bin): deliver typed terminal results from remote work homes

A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.

The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.

A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.

This is the emit-side counterpart of the retire/clear fix in #3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.

* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes

* no-mistakes(review): Fail collection when remote outbox is unreadable

* no-mistakes(review): Surface reassigned remote routes during empty collection

* no-mistakes(review): Fail remote collection on invalid registrations

* no-mistakes(review): Reject unsafe registration entries during remote collection

* no-mistakes(review): Restore healthy empty remote collection behavior

* no-mistakes(review): Skip remote collection for delivered registrations

* no-mistakes(review): Skip delivered registrations before route validation

* no-mistakes(document): Document remote follow-up collection semantics

* fix(bin): exclude secondmates from home-summary validity (#3504)

* fix(bin): exclude secondmates from home-summary child inventory

kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed.

* no-mistakes(review): Cover terminal secondmate in-flight exclusion

* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds

* fix(bin): self-heal outcome indexes on first drain (#3509)

* fix(bin): self-heal status-outcome indexes on every drain

Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault.

* no-mistakes(review): Guard held-lock initialization and fail marker writes

* no-mistakes(document): Document cross-harness outcome-index self-healing

* fix(bearings): keep active children underway during captain holds (#3505)

* fix(bearings): keep active children underway beside a captain hold

Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work.

* no-mistakes(review): Preserve Underway repos and disclose child truncation

* no-mistakes(review): Fall back to task project for Underway repos

* no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean

* fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513)

* fix(pi): settle watcher delivery on Pi accepting the follow-up

A follow-up queued while main is streaming joins the running run without
ever raising before_agent_start, so waiting on that event before clearing
the successor pipeline (#3498) stalled every later actionable close: no
successor started, no wake was delivered or offered to the branch, and the
turn-end guard woke main to re-arm by hand after every close.

The pipeline now settles once Pi accepts the follow-up. Consumption is
observed at before_agent_start for an idle main and at the user
message_start for a streaming main, and decides only what a replacement
session (/new, /resume, /fork, reload) replays. An exhausted restoration
delivers its typed failure without launching an arm past the retry bound,
which the stall had hidden. The replacement-coordinator map is typed so the
strict no-emit typecheck passes again.

Tests: the doubles no longer raise before_agent_start for a streaming send,
a portable regression drives two actionable closes while main streams and
proves the successor chain plus consumption-scoped replay, and a
credential-free real-SDK probe pins Pi's event contract for both the
streaming and the idle follow-up.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(pi): retry a verified successor that fails during wake delivery

A verified successor can exit while the wake it was started for is still
being delivered, most plausibly during a branch turn that holds the
settlement for minutes. Its failure close arrived while the pipeline's
single-flight guard was set, so the close handler skipped the retry, and
the pipeline's end no longer launched an arm, which left the live
generation with no watcher and no retry timer.

The close handler now records that failure when the child had reported
readiness and was not retired by the restoration itself, and the pipeline
runs the ordinary bounded, lock-checked retry for it once the delivery
settles. A restoration started for a later pending supersedes it, and an
exhausted restoration still hands repair to main without a further arm.

The regression holds a branch settlement open while the verified
successor exits with a failure and proves one retry watcher starts after
the settlement releases, none while it is held.

Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a

* fix(bin): bound repeat stale wakes for parked workers (#3532)

* fix(bin): bound repeat stale wakes for a parked but live worker

A worker parked on a declared wait - `paused:` for an external or pipeline
wait, or a verified `captain-held` transfer - kept waking firstmate far inside
FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one
captain-held worker and dozens across a day on a pipeline wait, and reported
upstream as four wakes in 75 minutes against a 3600s window.

pause_state_class deliberately answers `none` for a still-live agent even under
a declared wait, so a worker genuinely waiting on a decision is never silenced.
That classification is correct and is left alone; it routes every parked but
live worker through surface_nonterminal_stale on first sight of each distinct
stale hash, and an idle parked pane still churns its hash on a clock or a token
counter without changing what is being waited on.

Two places let that churn re-alarm:

- surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was
  declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should
  have suppressed it. The throttle was never read on this path and was advanced
  by the wake it should have prevented.
- The hash-change path cleared that throttle through clear_pause_tracking
  whenever the classification came back `none`, so each tick also bought the same
  declared wait a fresh window. Fixing only the first site changes nothing.

Read the throttle before anything is queued and advance it only on a wake that
really fires, and on the hash-change path reset only the per-hash bookkeeping
while the declaration still stands, via a clear_stale_hash_tracking split so
neither half of clear_pause_tracking is duplicated. The throttle is keyed to the
declaration, not to the pane.

First sight still wakes, so an inconclusive state is still inspected, and the
window's end still re-surfaces once, so a forgotten wait cannot rot invisibly -
noise traded for a bounded cadence, never for silence. The wake identity stays
the plain `stale: <win>` the away-mode handoff depends on.

Tests cover both observed forms and were confirmed to fail against three
deliberate breaks: each site reverted on its own, and a re-surface that never
fires again.

* fix(document): Clarify declared-wait wake cadence documentation

* fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor

* fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed

* fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567)

* fix(turnend): accept the away-mode daemon as the supervision owner

While state/.afk exists the away-mode daemon owns supervision and runs
bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon
starts its replacement. The turn-end guard tested for a live watcher
process holding the watch lock at that instant, so a turn boundary that
landed in the hand-off blocked with "TURN WOULD END BLIND" while
supervision was completely healthy, costing a full handling turn each
time.

Reproduced with the real daemon wrapping the real watcher and the real
guard sampling the same home: 6 of 40 samples blocked, every one of them
with the daemon alive and the beacon 2-3 seconds old, and a new watcher
pid on each cycle. After the fix the same reproduction blocks 0 of 40,
and killing the daemon and its watcher (away mode still on, beacon still
fresh) blocks again.

The guard now accepts a live, identity-matched daemon holding this home
as proof of supervision while away mode is active. The identity match is
the same discipline the watcher lock uses, so a recycled pid or a lock
left by a killed daemon proves nothing. The fresh-beacon half of the
predicate is unchanged: a daemon that stops restarting its watcher still
blocks once the beacon passes grace, a home with no supervisor blocks
exactly as before, and with away mode off the strict watcher predicate is
untouched.

The predicate reads only durable state, so it behaves identically for
every primary harness and runtime backend.

* no-mistakes(document): clarify away-mode daemon supervision proof and test coverage

* no-mistakes(document): generalize stale turn-end predicate summary in architecture.md

* fix(backlog): omit --file from row probes for non-markdown backends (#3582)

* fix(backlog): omit markdown file for beads probes

* no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes

* no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change)

* fix(bin): classify progress updates on requested work as routine (#3589)

The supervision branch's verdict rule escalated every outcome that
answered a captain request, so "the work started" and "still working"
notes reached the captain with nothing to look at. The rule now keeps a
finished result of requested work captain-facing, even when healthy, and
treats start or still-working updates that bring no new artifact,
finding, or decision as routine. The captain list for review-ready PRs,
ask-user findings, exhausted blockers, credentials, and destructive or
security-sensitive cases is unchanged, as are the unsolicited-routine,
silent-fleet-review, and doubt-chooses-captain rules.

The fm_branch_report tool description and the two docs that restated the
old unconditional rule now point at the prompt's "Verdict: routine or
captain" section as the one owner instead of carrying a second copy.

* fix(bin): preserve captain calls during teardown (#3595)

* fix(bin): never close a captain call during cleanup

A scout that held its own work item for the captain, which is what
captain-hold-lifecycle prefers ("hold the work item the question gates"),
was closed by bin/fm-teardown.sh's automatic backlog transition. The
completion gate passed, cleanup ran, and the captain's question moved to
Done with no recorded answer: the one thing the policy says must never
happen. `tasks-axi done` closes a held row silently, and nothing in
teardown asked whether the row was the captain's own call.

bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when
the task is still an open captain call, 1 when it is not, 2 when that
cannot be established. It reads the row through the transition library's
backend-aware probe, so it addresses the same backlog teardown does; the
script's other commands now address the configured data directory the
same way instead of FM_HOME, which also fixes captain holds in a home
with a relocated data directory.

Teardown asks `open` before any destructive step and refuses on 2. On 0
only the close changes: after cleanup and still under the task's own
lock, the row gets one "Deliverable of the finished work" line at the end
of its body and returns to Queued through `tasks-axi reopen`, keeping its
hold, so it lands in Captain's Call instead of reading as work under way.
--force does not lift this: it authorizes discarding unlanded work, never
the captain's question. The deliverable goes into the body because
`tasks-axi update --report` rewrites the title of a row that is not Done.

The crash window reuses the pending-close record teardown already stages:
a `mode=retain` line makes the existing replay record the deliverable and
reopen instead of closing, with the same validator, stale-generation
check, cleanup-incomplete marking, and non-blocking bootstrap lock as an
ordinary close. A retained row the captain answered first simply retires
the record. No parallel record type, recovery command, or second bootstrap
loop is introduced.

Regressions run the real executables: the captain-held scout survives
cleanup queued, held, with its deliverable and on the board, only
`answer` closes it, --force keeps it open, and an ordinary scout still
closes with its report; an interrupted cleanup leaves the row untouched
and the next session start retains it; a relocated backlog keeps the
retention in its one configured file; and a ship row whose hold cannot be
read refuses cleanup before anything destructive.

Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np

* no-mistakes(review): Serialize captain holds and fix backend-aware listing

* no-mistakes(document): Update captain-call retention documentation

* no-mistakes(document): Fix relocated captain-hold backlog diagnostics

* fix(bin): deliver secondmate outcomes to the parent channel (#3592)

* fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts

A secondmate's captain-facing outcomes could miss: the mate model addressed
the captain in its own unread chat instead of appending to the parent
channel, and a PR-ready report, a finding, a decision, a blocker, and a
failure all depended on that one remembered append. Make delivery
structural, so the parent channel never depends on the model:

- bin/fm-parent-channel-lib.sh is the one owner of channel resolution and
  exact-line append-once; the merge outcome path and the inactive-outcome
  scan now publish through it instead of two private copies.
- bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every
  watcher poll in a secondmate home: a direct child's whole terminal done or
  failed line is delivered at once with its note, recorded PR, mode, merge
  posture, and scout report pointer, keyed and receipted so it is delivered
  once, and the inactive path yields to it. `report <task-id>` runs the same
  delivery for a caller holding the child's meta lock.
- bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at
  registration.
- bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id
  and resolution-record count, with no new persisted state.
- bin/fm-teardown.sh delivers the child's final line before removing its
  record and refuses, retaining every record, while the channel cannot be
  written.
- The charter opens with the parent-channel rule and confines the mate's own
  appends to judgement; AGENTS.md carries the carve-out at the persona
  address rule and the escalation list.

docs/secondmate-parent-channel.md records the design and its coverage, and
docs/verification/secondmate-parent-channel.md records the live run with real
tmux panes and both real watchers delivering every line with no model.
Supersedes #3569.

* no-mistakes(review): Fix parent outcome retries and reconciliation locking

* no-mistakes(review): Prevent busy children from starving ledger delivery

* no-mistakes(review): Correct ledger metadata and hold occurrence handling

* no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons

* no-mistakes(review): Close ledger races and preserve teardown records

* no-mistakes(document): Correct parent-channel receipt and scanner documentation

* no-mistakes(lint): Quote done arguments for ShellCheck compliance

* no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks

* no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure

* fix(bin): sync remote second mates to primary commit (#3599)

* fix(bin): sync remote second-mate homes to the parent primary commit

Session start and remote launch pointed a remote second-mate home at whatever
Firstmate copy its own host kept, so a home that had already advanced past that
copy refused as a non-fast-forward and every other home stopped at the host's
older commit while the primary ran ahead.

The parent now resolves ITS primary default-branch commit with the existing
helper and hands that commit to the host on both paths. Because a remote home
is a standalone clone, the host imports that one commit before advancing -
already present, else from that host's Firstmate copy without moving it, else
from the home's own origin - and then runs the SAME ff_target guards a local
home gets, so dirty, diverged, feature-branch, and unresolvable targets skip
untouched and the ancestry rules keep one owner. An unimportable target now
names /updatefirstmate instead of failing opaquely, and a host still running an
older Firstmate copy is reported the same way rather than echoing a bare
refusal.

The host-local launch leg no longer re-runs its own secondmate sync, so the
spawn it drives cannot re-target that host's copy after the parent has already
converged the home.

/updatefirstmate is unchanged: it still refreshes the remote code root from that
host's origin and then syncs the home to that refreshed copy, which is what the
sync call with no target commit means.

* no-mistakes(document): Document primary-targeted remote secondmate synchronization

* fix(bin): separate captain intent from firstmate specs (#3597)

* fix(bin): split brief task into captain intent and firstmate spec

Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs.

* fix(bin): stop task-subsection copies at the next heading

Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body.

* no-mistakes(review): Validate brief content and preserve nested specifications

* no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies

* no-mistakes(review): Ignore fenced subsection headings during brief validation

* no-mistakes(review): Preserve captain intent across scout promotion

* no-mistakes(review): Enforce safe intent boundaries for legacy promotions

* no-mistakes(review): Allow marked legacy intent and reject empty promotions

* no-mistakes(review): Scope task parsing and overlay legacy intent contracts

* no-mistakes(review): Overlay current intent contract for all no-mistakes spawns

* no-mistakes(review): Preserve later captain clarifications in intent overlays

* no-mistakes(document): Document brief intent enforcement and ownership

* no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed

* no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint

* fix: start a fresh supervision branch for every main session (#3600)

* fix(pi): start a new supervision branch conversation per main session

The supervision branch reopened one recorded conversation forever, so
every main session start reloaded the current generated prompt and then
weeks of accumulated thread, where a superseded rule could still outweigh
today's.

The branch conversation is now scoped to one main session: the session
generation owns the recorded conversation, so a cold start, /new,
/resume, /fork, or a reload always builds a new one, while a rebuild
inside one session (a model or effort change) still continues that
session's own conversation.

The dialog mirror re-anchors with it. Its durable cursor records what the
previous branch conversation received, so a /resume or reload - which
keeps main's own session file - would otherwise leave the new branch
blind to dialog main itself still has. The reset is bounded by the
current main session, and the cursor keeps advancing incrementally within
it. The durable outcome store and its processed marker are untouched, so
unacknowledged captain-facing outcomes still re-present on the new main
session.

* no-mistakes(document): Document fresh Pi supervision conversations

* no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks

* feat: restart second mates after instruction updates (#3614)

* feat(update): restart second mates whose instructions changed

/updatefirstmate pulled new bytes onto disk and then asked each advanced
second mate to re-read them. A running agent holds AGENTS.md and every
loaded skill frozen from launch and no verified harness offers a reload,
so that steer could not reach a loaded skill at all and left the mate
holding two contradictory copies of its own job description.

An eligible mate is now restarted instead, in the same home and endpoint,
through the existing transactional relaunch. The restart is gated on the
mate first writing down the open work it holds only in conversation - the
open-record half of /stow, never its memory sweeps - so an unregistered
captain call is flushed before the conversation is spent. Anything that
leaves the reload unprovable falls back to the old re-read message and is
reported as exactly that, never as a clean reload.

Remote mates take the same path: fm-remote-secondmate-control.sh gains a
relaunch verb whose host-local leg runs that same control plane, since the
mate is an ordinary local secondmate from its host's point of view. The
primary resolves the profile and passes it explicitly, because
config/secondmate-harness is not inherited and the file on that host
belongs to a different home.

fm-update.sh now splits its advanced live mates into a restart set and a
nudge residual, and both sets require a changed instruction surface, which
also closes the over-nudge against the session-start sweep. Restart is
stricter still: a bin/-only advance reloads itself on the next call, so it
never costs a conversation.

Colocated tests cover the gating, the persist-then-restart order, the
task-subset persist request, each unsafe fallback, the remote hop, and the
remote sync's new instruction-surface report.

* no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting

* no-mistakes(review): Parallelize relaunches and classify replacement incarnations

* no-mistakes(review): Gate restart actions on live agent state

* no-mistakes(review): Handle failed restart workers without hanging

* no-mistakes(review): Nudge legacy remotes and preserve persist recovery

* no-mistakes(review): Document one-time secondmate restart rollout

* no-mistakes(review): Honor arrived replies and refresh remote profiles

* no-mistakes(review): Revert remote parent profile reconciliation

* no-mistakes(review): Reset remote profile defaults and honor published results

* no-mistakes(review): Preserve fallback nudges for unverifiable secondmates

* no-mistakes(document): Document second-mate restart update flow

* no-mistakes(lint): Fix ShellCheck warnings in restart scripts

* perf: accelerate local validation with bounded concurrency (#3644)

* perf(tests): route gate verification through the bounded concurrent runner

Local validation was the pipeline's dominant cost: across 67 recorded
no-mistakes agent sessions on this repo, 99.3% of command execution was
`bash tests/*.test.sh`, run strictly one script at a time, and 2% of those
calls were killed by an agent-guessed timeout and paid for twice.

Three changes, each measured:

- `.no-mistakes.yaml` pins `commands.test` to
  `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner
  already owns changed-file selection, bounded concurrency, the refusal of
  unproven scripts, and a generous automatic per-script bound, so the gate's
  baseline is neither a serial chain nor a guessed timeout. It stays
  intent-targeted - the Test step still runs its evidence agent on top - and
  excludes the live-Herdr family the required Herdr lane owns.

- `bin/fm-test-run.sh` gives a plain list of script paths the same bounded
  automatic scheduler and automatic bound that `--changed` gets. Naming several
  subjects is how a verification round asks for exactly those scripts. The
  curated selections are untouched: `--lane` still composes CI shards whose
  serial lane must stay serial, `--family` is what the required Herdr lane runs,
  and `--all` stays a deliberate complete regression.

- `pr-forge` is admitted to the concurrent-safe family registry on two
  consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those,
  and records `secondmate` and `session-bootstrap` as refused with the exact
  script and reason each failed on, so the refusals are actionable rather than
  silent.

Measured on this host, 0 failures on both sides:

  verification round, 4 scripts   448s chained -> 231s through the runner (-48%)
  pr-forge family                 409.2s at 1 worker -> 237.9s at 4 (1.72x)
  watcher-wake-lock family        1311.1s at 1 worker -> 539.3s at 4 (2.43x)

A fourth lever was implemented and then removed because the measurement
refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made
`fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s
unchanged, back to back. Those sleeps are not overhead added to the clock -
they are how a test waits for a subject moving on fm-watch.sh's own one-second
cadence - so sampling less often only delays detection. It also broke
`fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a
settled condition. CONTRIBUTING.md records that result so the experiment is not
repeated.

* no-mistakes(review): Separate concurrent runs by isolation proof family

* no-mistakes(review): Limit automatic timeouts to changed-file validation

* no-mistakes(document): Clarify validation concurrency documentation

* fix: copy PR URLs from durable records (#3648)

* fix: copy PR URLs from records or abstain, never assemble them

Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.

Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:

- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
  or abstain" section requires a URL to be copied verbatim from a durable
  record (the done: PR <url> status line, pr= metadata, or the backlog note),
  forbids assembling owner, repository, host, or number from memory, and has
  the branch report only the identifier it actually holds when no record names
  the URL yet, leaving the PR check unarmed until the worker's ready line
  arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for
  main in place of the bare full-URL mandate.

- Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full
  https:// URL wherever a PR is mentioned - status line, terminal, or summary -
  never a bare "PR 108", so the link is in view as early as the number is.

- bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that
  the task's own done lines contradict, printing both spellings; a log naming
  no URL still records the argument as before. fm_pr_status_ready_urls in
  bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches
  bin/fm-pr-merge.sh, so nothing merges under a contradicted URL.

Tests cover the offline refusal with zero side effects, the recorded spelling
being accepted, markdown-wrapped and punctuated URLs, working lines not
counting, the merge wrapper propagation, a self-hosted merge request with no
forge call, the prompt carrying the rule, and the brief carrying the worker
rule.

* no-mistakes(review): Remove stale PR URL enforcement

* no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`

* fix(bin): disable Claude feedback drafts for fleet launches (#3661)

* fix(bin): disable Claude's feedback-draft flow for fleet-launched agents

Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched
Claude crewmate and secondmate, so /bug and /feedback never queue or
submit a bug report on the captain's behalf. feedbackDrafts is the
documented settings key (Claude Code changelog 2.1.247); the
per-launch CLI flag never touches the captain's global settings.json.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts

* no-mistakes(document): Fix Claude feedback documentation formatting

* fix(bin): layer both feedback-draft controls for defense in depth

The prior --settings-only fix can be overridden by a managed Claude
settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0
alongside --settings '{"feedbackDrafts":"off"}': either control alone
disables the SendFeedback tool, so a managed override of one still
leaves the other in force.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(document): Document Claude feedback-draft suppression ownership

* feat(tests): run three more validation families concurrently (#3662)

* perf(tests): admit three more families to concurrent validation

The three families that `docs/fm-test-isolation-proof.md` recorded as refused
were not refused for concurrency. Each blocker was a test that decided a
property by wall clock, or a script filed where it cannot run. Fixing those
three things admits all three families and recovers 28.6 minutes of local
validation with no assertion removed or weakened.

- `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the
  handoff, sleeping a fixed second, then delegating the move to the real
  binary. Nothing ever killed the fake, so on a host slow enough for the case's
  next assertions to take longer than a second, the orphan woke and completed
  the very move the case requires left undone, and recovery then failed with
  `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs
  during the injected crash showed exactly that, the item moving one second
  after the crash. All four crash injections in the file now go through a new
  `fm_fake_crash_injector` shim that signals the target and returns only once
  it is observably gone, and the pre-move fake never delegates the move at all.

- `tests/fm-session-start.test.sh` proved the startup digest does not block on
  a slow current-state read by timing the whole digest against a fixed
  eight-second sleep, which a loaded host exceeds without the property being
  violated. It now holds that read open until the case releases it and asserts,
  the moment the digest returns, that the read has not finished. A digest that
  waited would wait indefinitely rather than for an interval a slow host can
  out-run, so the assertion is stronger than the bound it replaces. Its scan
  budget moves to the maximum, because the old value left two seconds of margin
  over the fixed sleep and measured the host rather than the deadline that
  `tests/fm-inactive-reconcile.test.sh` owns.

- `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all,
  which put it in the portable serial lane, where Linux CI gate-skips it: that
  real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and
  the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its
  opt-in variable and moves to `live-harness-optin`.

The 28 remaining ungrouped scripts become an enumerated `standalone` family
instead of admitting `unclassified` itself. `unclassified` is the family map's
`*)` arm, so admitting it would silently grant concurrency to every test added
afterwards, which is exactly the population with no proof. A new test still
lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh`
covers that split behaviorally.

Each family passes two consecutive four-worker proofs with zero failures. On
the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap`
756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s
recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock
against 121 minutes of summed script time.

* no-mistakes(document): Refresh concurrent validation and shard documentation

* no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2

* feat: structure no-mistakes ask-user escalations (#3670)

* feat(brief): structure no-mistakes ask-user escalation as event + snapshot file

Crewmates escalating a no-mistakes ask-user gate now report one status
event naming every finding id plus a snapshot file holding the gate's
axi finding records verbatim (id, severity, file, line, description,
authority), using the same shape even for a single finding. The status
line never paraphrases. The format is defined once in fm-dod-lib.sh and
rendered into both the scout and ship rule 6 in fm-brief.sh, so a
promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets
the identical contract as a freshly-spawned no-mistakes ship worker.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei

* no-mistakes(review): Preserve ask-user escalation output contract

* no-mistakes(review): Align escalation format test expectation

* no-mistakes(review): Scope ask-user escalation instructions correctly

* no-mistakes(review): Remove ask-user from generic decision rules

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix(bin): require self-sufficient no-mistakes intent (#3671)

* fix(bin): require a self-sufficient no-mistakes intent

A no-mistakes worker's --intent is only as useful as the string it
passes. PR #3604 shipped with an intent that was only "do 1, 2, 3, 7
from the report": the real contract lived in a private scout report and
never reached --intent, so nobody holding that string plus the codebase
could have derived the specification.

This is pure instruction at the contract's one owner; no spawn-side or
promotion-side check is added.

- bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now
  states that the --intent string must be self-sufficient (the string
  plus the codebase reconstructs roughly the same specification) and
  tells the worker to write the substance of any report, decision, or
  PR the captain's intent refers to into --intent rather than the
  pointer, while Firstmate build instructions and the worker's own
  decisions still stay out. The spawn-time overlay points back at that
  rule so its "supersedes" wording cannot cancel it, and the header's
  owner statement carries the rule.
- AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to
  include the substance of referenced material when filling
  ## Captain's intent, and section 11 points at the owner of the rule.
- tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the
  rendered brief and launch contract carry the rule.

Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim

* no-mistakes(document): Replace incident-specific intent test commentary

* fix: accelerate local Bearings snapshot composition (#3499)

* Speed local fleet snapshot composition

* no-mistakes(review): Stabilize task inventory during concurrent snapshot composition

* no-mistakes(document): Document local snapshot observation concurrency

* no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks

* no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks

* no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks

* no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky

* no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks

* fix(snapshot): keep live observations generation-coherent

* no-mistakes(review): Keep secondmate observations generation-bound without copying reports

* no-mistakes(document): Document generation-coherent snapshot observations

* test(bearings): measure local read overlap instead of wall-clock budget

The large-local-snapshot regression asserted that a whole snapshot
composed in under five seconds. That bound measures how loaded the host
is, not whether the per-task reads actually overlap, so it failed
intermittently on a contended machine: one run in six on a box at load
16-20, landing exactly on the five second boundary.

Time a serialized run and a concurrent run of the same workload instead
and require the concurrent one to save at least two seconds. Both runs
pay the same composition overhead, so the difference isolates the
overlap this change delivers. Five one-second reads serialize into five
seconds and overlap into about one, and re-serializing the reads
collapses the saving to roughly zero, so the assertion still fails
loudly if the concurrency regresses.

Also bump the pinned Bearings test count to 48, since rebasing onto the
current default branch picked up its captain-hold test.

* no-mistakes(review): Restore JSON-derived decision flags

* no-mistakes(review): Unify status-derived snapshot observations

* no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass

* fix: prevent stale supervision wake loops (#3672)

* fix(bin): stop the supervision branch's stale-ack and ghost-report loops

Clean-slate implementation of the four authorized recommendations from the
supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal
form, superseding PR #3604:

- fm_branch_report refuses a task the wake being handled never named. The
  extension fixes the reportable task set from the eligible rows before each
  prompt (signal and stale rows resolve to their tasks, a heartbeat allows any
  task with a live record, fleet is always allowed), so a report typed from
  memory about a task whose records teardown already removed is never stored
  or delivered.
- An acknowledgement that consumes nothing says "nothing was acknowledged
  through N" and prints the exact --ack-through / --recovery-generation
  command for the current presented wake, instead of "re-run the drain",
  which re-fed the same stale acknowledgement in a loop.
- bin/fm-guard.sh no longer tells the branch actor to drain queued wakes
  while it is handling them; it names the granted rows instead.
- Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and
  descendants; the index rebuild and the append-side index write both skip a
  task with neither a live record nor a status log, so the branch's report of
  a teardown it just performed is stored without recreating the index.

No new locking, no spawn-generation binding, and no retired-task refusal: the
branch can still report the outcome of a task it just tore down, and the
teardown test now proves that path end to end.

* fix(bin): narrow the branch report scope and guard silence to the minimal form

Apply the four review decisions on the clean-slate branch:

- A signal or stale prompt may report only the tasks its own rows resolve
  to; fleet is refused there too. A heartbeat review is not scoped by task
  at all, so the extension no longer tracks live task records and refuses
  nothing by task id during a fleet review.
- The outcome-index rebuild no longer skips retired tasks; the append-side
  skip alone keeps a torn-down task's index from being recreated.
- bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor
  instead of printing a replacement note.

* no-mistakes(document): Align supervision docs with scoped wake handling

* fix(bin): avoid fleet snapshot argument limits (#3677)

* Fix fleet snapshot large JSON transport

* no-mistakes(review): Captain: file-back fleet snapshot transport safely

* no-mistakes(review): Captain: file-back parent summary aggregation

* no-mistakes(ci): Rebased the PR's three commits onto f4d7875824ecc5e274b4bb896f10c1e1f207b7e4 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor

* fix(bin): attribute active runs with unfetched pipeline heads (#3681)

* fix(bin): rec…
NewAiCoder added a commit to NewAiCoder/firstmate that referenced this pull request Sep 26, 2026
…rt (kunchenguid#28)

* fix(bin): preserve subshell lock ownership on Bash 3.2 (#3789)

* fix: distinguish subshell wake-lock owners on stock Bash

Restore distinct process ownership for issue #3743 using the existing PID helper, consistently across lock publication, reclaim, release, role checks, and bounded handoff.

The existing wake-queue regression fails on pristine upstream Bash 3.2 with rc=13. The complete suite now passes on Bash 3.2.57 and Bash 5.3.15, with added coverage for ownership when BASHPID is unset. Canonical lint and stock-Bash syntax checks pass.

* no-mistakes(document): Correct lock grace-period documentation

* no-mistakes(ci): Captain, fixed all 14 SC2031 false positives with nine ShellCheck source-boundary annotations across three tests. Full CI-mode lint and the complete wake-queue suite on stock Bash 3.2 passed. Runtime behavior is unchanged

* fix(bin): resolve captain holds and legacy teardowns on non-markdown backends (#3782)

* fix(bin): close legacy records on the Beads backend honestly

Two pre-Beads reads blocked honest closure of leftover records:

1. fm-captain-hold.sh complete/verify resolved attested legacy hold ids
   only against the live backend and the pre-collapse derived identity, so
   a home whose holds fm-hold-migration rehomed under fm- ids failed with
   an empty-name absence message (the resolve failure was swallowed by the
   command substitution feeding verify_hold_durable). Resolution now falls
   back, on the Beads backend only, to the legacy id under the configured
   beads prefix and to the row whose notes carry the exact marker line
   'migrated from data/backlog.md id <legacy id>'; every refusal names the
   id it could not resolve, and the markdown path is unchanged.

2. fm-teardown.sh refused any record without spawn_gen forever. A record
   that predates the field can now be torn down with an explicit
   --legacy-record flag once the recovery-grade endpoint classifier
   confirms the recorded endpoint dead or agent-less; the accepted
   incarnation is stamped into the record right before its close marker
   binds to it and named in the teardown line. Refusals leave the record
   byte-identical, the unlanded-work refusal is not relaxed, and a corrupt
   (multi-valued) spawn_gen is never accepted.

The companion repair this branch carries (follow-up commit) is the
backend-gated --file and markdown-file requirement in the mutate path and
lifecycle gates: fm_backlog_mutate passed --file and required the markdown
backlog file regardless of the resolved backend, and the transition gate
plus row probe required that file before any backend work, so a home on a
non-markdown backend could neither gate, probe, nor close its rows.

Behavior tests: self-contained beads fixtures over a scratch bd graph
(self-skipping on markdown-only tasks-axi installs), legacy meta fixtures
for every teardown gate, and the relocated markdown backlog coverage stays
green.

* no-mistakes(review): fix(review): report migrated-hold scan refusals and guard legacy spawn_gen stamp against newline-less records

* fix(backlog): address the configured backend for lifecycle writes

Completes the fm-backlog-transition-lib repair the first commit's message
claims: on this base fm_backlog_mutate passed --file and required the
markdown backlog file regardless of the resolved backend, and
fm_backlog_transition_applies plus fm_backlog_row_probe required that file
before any backend work, so a home on a non-markdown backend could neither
gate, probe, nor close its backlog rows. All three now gate the markdown
file on the resolved tasks-axi backend: markdown keeps exactly its explicit
<data>/backlog.md behavior, non-markdown homes address the backend their
own configuration selects with no markdown file requirement.
fm_backlog_row_show and fm_backlog_row_list already gated correctly and
are unchanged. docs/configuration.md owns the contract line.

Also extends the same backend gate to fm-captain-hold.sh's own mutation
wrapper - hold/add/update/answer/done append the markdown --file only when
the resolved backend is markdown, so a captain call on a Beads home reaches
the Beads store end to end - and applies the review round's two direct
remedies there: the [beads] graph path resolves against the backlog root
when relative (never the process CWD), and a failed bd graph read reports
bd's own trimmed stderr reason in the refusal.

Coverage: tests/fm-backlog-atomicity.test.sh gains a stub-driven Beads
completion case proving the transition gate applies, the row probe reads,
and done runs without any markdown file or --file override; the relocated
markdown backlog test stays green.

* no-mistakes(review): Document root-tasks.toml-only beads settings for migrated-hold resolution

* test(gotmp): stub fm_tasks_axi_backend so the fixture matches the backend-aware transition lib

The legacy-records change made fm-backlog-transition-lib.sh resolve the
configured backend via fm_tasks_axi_backend before the markdown-only skip.
The gotmp fixture's fm-tasks-axi-lib stub lacked that function, so the
markdown check fell through and teardown hit the incompatible-backend
error with unbound FM_TASKS_AXI_MIN under set -u. Stub the backend as
markdown and define the floor, restoring the intended no-backlog skip.

* fix(teardown): roll the legacy stamp back when the close marker fails

A legacy-record teardown stamps its accepted incarnation into the record
right before the close marker binds to it; when that marker write then
fails, the stamp survived, so a retried teardown sailed past the
dead-or-agent-less endpoint gate the stamp now proved unnecessary. The
failed marker write now truncates the record back to its exact pre-stamp
bytes (verified by size), restoring the byte-identical-refusal invariant;
when the rollback itself fails the operator is told to re-run with
--legacy-record after reconciling the endpoint.

Also completes the recorded review decision's coverage wording: the
beads stub test now drives the answer close end to end (update and done
through the gated wrapper), asserting no markdown file override reaches
either verb.

* fix(review): harden the legacy stamp rollback and resolve derived migrated ids

The legacy-record stamp rollback now uses perl (already in the teardown
curated PATH; truncate is not, and is absent on stock macOS), routes every
failure branch inside the stamp block through the same size-verified
rollback so the byte-identical-refusal invariant holds on those paths too,
and gains behavior coverage: an unrecordable close (an invalid pr= link)
fails the teardown, leaves the record byte-identical, keeps the backlog
row in flight, and a flag-less retry still refuses.

Migrated-hold resolution now probes the derived pre-collapse identity
(<origin>-decision-<entry>) alongside the raw entry - fm-hold-migration
recorded the DERIVED id in every migrated row's marker note - in both the
prefix and the migration-note forms, with the ambiguity refusal naming
every identity tried, plus behavior coverage for a bare decision key
resolved through its derived identity's marker.

Also aligns fm-backlog-transition-lib.sh's header ADDRESSING/SCOPE
paragraphs with the backend-gated contract, drops an unreachable FORCE
validity guard the parser rewrite left behind, and switches the new stub
fixture to the portable sed -i.bak idiom.

* no-mistakes(review): Name the configured backend in teardown's backlog reminder

* no-mistakes(review): Scan migration markers before the prefix guess

* no-mistakes(review): Document marker-first resolution and cover the prefix branch

* no-mistakes(document): Record prefix-attestation audit and marker-line forms

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-teardown.sh: a failed rollback of the synthetic legacy stamp let a retry bypass the dead-or-agent-less endpoint gate. Root cause: teardown minted `spawn_gen=legacy-<ts>-<pid>` into the task record before the close marker bound to it. When the close-marker write failed AND the rollback also failed, the record retained that token. On the next invocation `fm_backlog_meta_spawn_gen` succeeded, so `TEARDOWN_LEGACY_PENDING` stayed 0 and the endpoint gate was skipped entirely — even with `--legacy-record`. The script's own error text told the operator to "re-run teardown with --legacy-record", advice the code could not honor. Fix (bin/fm-teardown.sh): - A `legacy-*` spawn_gen is now recognized as a stamp this teardown path minted, never one a spawn published (fm-spawn.sh publishes `s<epoch>.<pid>.<random>`). Such a record still reads as the legacy record it is: it re-enters the endpoint gate, and a flag-less retry refuses naming `--legacy-record`. - Acceptance reuses the retained token instead of minting a second one; the append block is skipped when the record already carries it, so no duplicate spawn_gen is written. - The rollback attempt and its "could not be rolled back" message are guarded to runs that actually appended a stamp, so a run that appended nothing never claims a rollback it did not perform. - Usage header documents the retained-stamp rule. Test (tests/fm-teardown.test.sh): added `test_retained_legacy_stamp_still_faces_the_endpoint_gate`, an end-to-end reproduction — a `perl` stub that fails only the rollback's `truncate` (delegating every other perl call to the real interpreter) leaves the stamp behind, then the retry must still hit the gate, must not stamp a second incarnation, must not close the backlog row, and the flag-less retry must refuse. Verification: the new test fails against the pre-fix script on exactly the reported defect ("the retry skipped the dead-or-agent-less endpoint gate") and passes after. Full tests/fm-teardown.test.sh 80 ok / 0 failures / rc=0; tests/fm-backlog-atomicity.test.sh 80 ok / 0 failures / rc=0; bin/fm-lint.sh (pinned ShellCheck 0.11.0 + actionlint 1.7.12) clean

* fix: reduce local ShellCheck source-analysis cost (#3778)

* fix(lint): drop source following on the local changed-file gate

The local lint step was inlining library closures through --external-sources
and peaking above 8 GB on a single root. Keep full analysis in CI, on main,
and without a merge-base; exclude the four cross-file codes from the local
pass so those findings still land in CI.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Run local ShellCheck per root and document measurements

* no-mistakes(review): Correct local source-following telemetry

* no-mistakes(document): Clarify context-sensitive lint documentation

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(pi): route decision-owned wake batches to main (#3776)

* fix(pi): route needs-decision wakes and mixed batches wholly to main

Skip the supervision branch for every needs-decision status append, the
same way a check-kind wake already skips it. A coalesced signal/stale
trigger batch containing any needs-decision row is delivered wholly to
main, not split between the branch and a later main wake - the whole
batch, including any co-present routine rows for a different task,
travels together. Heartbeat and unread-status scans stay independent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* test(pi): cover distinct-file mixed batches and heartbeat independence

Add a regression using two distinct files (not the same status file
twice) in one coalesced trigger so a some-vs-every regression on the
file-list cross-reference cannot hide behind a degenerate same-key
case, and a heartbeat/needs-decision co-presence test proving a
needs-decision row neither vetoes nor rides along with an otherwise
eligible heartbeat scan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* no-mistakes(document): Clarify needs-decision and heartbeat routing

* no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect

* no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect

* no-mistakes(review): Route stale open decisions directly to main

* no-mistakes(review): Honor configured verbs in stale decision routing

* no-mistakes(review): Route second-mate escalations and configured decisions to main

* no-mistakes(review): Ignore trailing whitespace after captain holds

* no-mistakes(review): Cache stale decision classification per status file

* no-mistakes(review): Document unread decision precedence for later task wakes

* no-mistakes(review): Cache unchanged stale decisions across scope scans

* no-mistakes(review): Resolve decision aliases and reject symlinked statuses

* no-mistakes(review): Route surfaced captain-held signals directly to main

* no-mistakes(document): Document decision-owned main routing

* no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034)

* no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* feat(bin): add opt-in worker launch environment allowlist (#3802)

* feat(spawn): add an opt-in worker environment allowlist

Honor a home-local launch-env-allowlist at the shared worker command
boundary and inherit it into secondmate homes. Preserve the existing
launch behavior when the file is absent. Keep the operational environment
and explicit launch assignments, and account for filtered Muse credentials.

Refs https://github.com/kunchenguid/firstmate/issues/3742

Verification:
- Red on origin/main 1820316b66ac2c68e244dd04a02512859ee8c1f4:
  the new enabled-allowlist regression observed synthetic-unrelated in
  the worker; the absent-file control passed.
- Green: fm-test-run.sh on fm-spawn-dispatch-profile, fm-muse-harness,
  and fm-trace-context-spawn; all three passed without skips.
- Synthetic emitted-command probes ran through sh, stock Bash, and zsh.
- Canonical lint, documentation audience checks, and stock Bash syntax
  checks passed.

* no-mistakes(review): Reject inaccessible launch environment configuration

* no-mistakes(review): Preserve inherited allowlists on source inspection errors

* no-mistakes(document): Clarify worker environment grants and inheritance documentation

* no-mistakes(lint): Fix inheritance test ShellCheck source boundary

* feat(bin): add rovo crewmate/scout adapter with home-path file access (#3575)

* feat(bin): add rovo as a verified crewmate/scout worker harness

Wire the Atlassian Rovo CLI (202609.1.2) into the TUI-under-tmux/herdr
adapter contract: detection with marker-precedence ordering, one-shot
positional launch with --startup-receipt readiness polling instead of
composer scraping, model/effort flags, a screen-scrape busy fallback
scoped like grok's, and crew/scout-only lifecycle control that refuses
secondmate launches. Ships with a portable regression suite, a live PTY
guard against the real binary, a per-harness reference doc, and a dated
verification record covering the silent OAuth refresh, the interrupt-ack
divergence from the originating scout report, and the still-open
composer-ghost and tmux/herdr pane-liveness gaps.

* no-mistakes(review): revert rovo launch to positional brief, drop startup-receipt

* no-mistakes(test): rewire rovo adapter to kimi-style launch-then-send shape

* no-mistakes(document): add rovo to stale worker-harness enumerations in docs

* docs(verification): close the rovo herdr-liveness gap with live isolated-lab evidence

Placement, launch-then-send, and busy/idle rendering are now verified live
in an isolated non-default Herdr lab session (bin/fm-herdr-lab.sh), driven
directly through fm-spawn.sh's/fm-backend.sh's own shared primitives since
the cross-session launcher-identity guard refuses this task's own ambient
Herdr identity for a full fm-spawn.sh run.

fm_backend_agent_state reported dead for a live, responding rovo pane at
every point checked, because herdr's own agent-integration registry has no
rovo entry (herdr integration status), so herdr agent get returns
agent_not_found regardless of whether rovo is actually running. This is
recorded as a Herdr-side integration gap rather than a firstmate bug, left
unpatched to avoid a false-positive alive verdict for other idle shells.

Updates docs/verification/rovo.md's backend-liveness section and its two
cross-references (docs/verification/runtime-backends.md, docs/configuration.md)
accordingly.

* fix(bin): close rovo's failed-spawn leak and busy-scrape false idle

Greptile P1s on PR #3575: a failed rovo readiness/submission/delivery gate
exited without tearing down the just-created endpoint, leaving the launched
--yolo rovo process running as an orphaned agent outside task control.
Separately, the busy classifier's rendered-tail fallback returned definitive
idle whenever the "Rovo is thinking" marker scrolled out of the last 12
nonblank lines of a long turn, which could make supervision wrongly conclude
a still-working worker had gone idle.

fm-spawn.sh: rovo_spawn_fail now calls rovo_endpoint_cleanup, which kills the
created endpoint (tmux/herdr/zellij/cmux) via the same generic fm_backend_kill
dispatch fm-spawn.sh's own orca-abort path already uses; orca's worktree and
terminal remain owned by the separate ORCA_ABORT_CLEANUP trap.

fm-busy-lib.sh: the rovo classifier arm now reports "unknown rovo-regex"
instead of "idle rovo-regex" when the marker is absent, matching how muse and
cursor already express "can't tell" for their own fallbacks. The positive
busy match is unchanged.

Extends tests/fm-rovo-harness.test.sh: the readiness and delivery failure
tests now assert the endpoint is torn down (and the success test asserts it
is not), and a new test drives the busy marker out of the tail window to
confirm the verdict is unknown, never idle. bin/fm-lint.sh is clean on both
changed files.

* test(rovo): align spawn fixture with the launch-brief validation contract

Upstream main now requires a brief's ## Captain's intent and
## Firstmate spec subsections (or a nonempty legacy # Task body)
before spawn, and rewrites ship+no-mistakes briefs into
launch-brief.md. Update the rovo harness fixture and pointer
assertions to match, mirroring the kimi harness fixture.

* no-mistakes(review): align rovo.md delivery-gate note with live herdr evidence

* fix: prevent stale supervision wake loops (#3672)

* fix(bin): stop the supervision branch's stale-ack and ghost-report loops

Clean-slate implementation of the four authorized recommendations from the
supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal
form, superseding PR #3604:

- fm_branch_report refuses a task the wake being handled never named. The
  extension fixes the reportable task set from the eligible rows before each
  prompt (signal and stale rows resolve to their tasks, a heartbeat allows any
  task with a live record, fleet is always allowed), so a report typed from
  memory about a task whose records teardown already removed is never stored
  or delivered.
- An acknowledgement that consumes nothing says "nothing was acknowledged
  through N" and prints the exact --ack-through / --recovery-generation
  command for the current presented wake, instead of "re-run the drain",
  which re-fed the same stale acknowledgement in a loop.
- bin/fm-guard.sh no longer tells the branch actor to drain queued wakes
  while it is handling them; it names the granted rows instead.
- Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and
  descendants; the index rebuild and the append-side index write both skip a
  task with neither a live record nor a status log, so the branch's report of
  a teardown it just performed is stored without recreating the index.

No new locking, no spawn-generation binding, and no retired-task refusal: the
branch can still report the outcome of a task it just tore down, and the
teardown test now proves that path end to end.

* fix(bin): narrow the branch report scope and guard silence to the minimal form

Apply the four review decisions on the clean-slate branch:

- A signal or stale prompt may report only the tasks its own rows resolve
  to; fleet is refused there too. A heartbeat review is not scoped by task
  at all, so the extension no longer tracks live task records and refuses
  nothing by task id during a fleet review.
- The outcome-index rebuild no longer skips retired tasks; the append-side
  skip alone keeps a torn-down task's index from being recreated.
- bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor
  instead of printing a replacement note.

* no-mistakes(document): Align supervision docs with scoped wake handling

* fix(bin): grant rovo the per-task home paths its standard crewmate flow needs

rovo confines every file-tool operation to its worktree by default, and its
bash tool independently refuses the same external paths regardless of any
grant (confirmed live), so a rovo worker could not read its own brief or
steering messages or write its status/report - all of which live in the
firstmate home outside the worktree - without hand-feeding it. Grant
toolPermissions.allowedExternalPaths for exactly the task's brief directory,
steering inbox, and status file at launch time via --config-override,
merged with agent.efficiencyLevel into one JSON object since that flag is
single-value and silently discards a second occurrence.

Extends the live PTY guard to prove, against the real binary, that the
grant lets rovo read an external brief and append to an external status
file, and that the same flow is blocked without the grant.

* no-mistakes(document): align rovo reference Effort row with merged single --config-override

* no-mistakes(document): document rovo file-access grant in harness reference

---------

Co-authored-by: PUNEET PATWARI <ppatwari@atlassian.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>

* fix(bin): prevent false pipeline blocks after drive timeouts (#3813)

* fix(bin): read a crew's pipeline-death claim against the live run

A crew's no-mistakes drive call blocks until the next gate or outcome,
routinely far longer than its harness lets one command live, so the call
gets killed or times out while the daemon runs the fix round on in the
background. Crews read that as daemon death and block on it, and firstmate
had nothing that contradicted them.

Rule 7 of every generated brief now says a drive-call error or a harness
command timeout is not a daemon error, requires `no-mistakes daemon status`
plus `no-mistakes axi status` before a pipeline `blocked:`, and reserves
that report for a refused socket or a run record failed with a daemon
error. The no-mistakes definition of done adds the harness command limit
and the background-and-poll shape that fits inside it.

fm-crew-state gains one classification case: a `blocked:` line blaming the
daemon, a timeout, or unreachability, while the run is running or fixing
AND the pipeline reports fresh activity, now reads as superseded because
the run is alive. Recency comes from the client's own `quiet` marker on
active_steps.last_activity rather than a threshold invented here, and
positive evidence is required, so a run record that outlives a genuinely
dead daemon keeps the plain reading.

stuck-crewmate-recovery gains the inverse-of-a-dead-endpoint playbook:
firstmate reads both statuses itself, steers a reattach, never restarts the
shared daemon on a crew's claim, and escalates only a refused socket.

Nothing here depends on an unshipped no-mistakes capability.

* no-mistakes(review): Prioritize daemon socket failure and narrow unreachable matching

* no-mistakes(review): Honor socket refusal across coarse status and crew guidance

* no-mistakes(test): Replace flaky settle timing assertion with pane-read count

* no-mistakes(document): Document daemon timeout recovery contract

* no-mistakes(ci): Fixed daemon socket failures being suppressed by terminal attributed runs. Positive refused/missing socket evidence now remains blocked regardless of run status. Added a behavioral regression test for terminal failed runs. Verified with fm-crew-state tests, project ShellCheck lint, and git diff checks

* fix(bin): scope the worker role contract for ship and scout launches (#3797)

* fix(brief): scope Firstmate workers to their launch contract

* no-mistakes(document): Clarify supervisor scope and worker contract ownership

* no-mistakes(ci): Removed the heading-based bypass so every ship/scout launch receives the current worker-role contract. Added a regression that failed before the fix and passes afterward. Dispatch, brief, and delivery suites, focused ShellCheck, and git diff --check all passed

* no-mistakes(review): make launch overlay sole owner of worker role contract

* no-mistakes(review): narrow heading test dimension and fix publish error wording

* no-mistakes(review): gate role supersession, fix render guard, drop AGENTS twin

* no-mistakes(document): align architecture AGENTS.md scope and spawn launch-brief header

* fix(bin): stop ringing steering doorbells into dead panes (#3823)

* fix(bin): stop ringing steering doorbells into dead panes

The steering-inbox doorbell was a plain sentence plus Enter typed into a
worker's pane, and the watcher re-rang it on the assumption that a ring is
free. In a pane whose agent has exited that line is a shell command, and the
re-ring ladder kept typing it into a shell that can never acknowledge it.

- Prefix the doorbell with the shell no-op `: ` so a bare shell executes
  nothing while a live worker still reads the same self-describing line.
  `#` is not used because interactive zsh does not treat it as a comment by
  default and the claude harness binds it to memory mode.
- fm_task_inbox_ring skips the pane (return 3) when the backend positively
  classifies the agent as dead; missing, ambiguous, unreadable, and unverified
  endpoints still ring so a blind classifier never starves a live worker.
- The watcher caps the ladder for a dead pane: one stale wake for recovery,
  no ring, no ladder walk, and the durable record stays for
  stuck-crewmate-recovery. fm-send and the remote steer leg report the skip.

Tests cover the no-op in real shells, the dead/live/unclassifiable ring
verdicts, and the single-surfacing watcher path.

* no-mistakes(review): Quote doorbell paths against shell injection

* no-mistakes(review): Reject terminal-control paths before ringing

* no-mistakes(review): Document accepted partial doorbell delivery race

* no-mistakes(review): Skip unavailable endpoints before busy-state handling

* no-mistakes(test): Respect shell startup PATH in environment allowlist test

* no-mistakes(test): Fix doorbell test fixtures for endpoint liveness

* no-mistakes(test): Prioritize confirmed restarts and clean shell test syntax

* no-mistakes(document): Document dead and missing doorbell recovery

* no-mistakes(ci): Fixed the persistence-reply timeout race by rechecking for a correlated reply immediately before falling back to a nudge. Added a deterministic regression covering replies arriving between the preliminary resolution pass and timeout handling. Verified with the targeted restart suite, project lint, coverage guard, bash syntax checks, and diff checks

* fix(bin): wait out transient primary-checkout reads in the spawn worktree poll (#3834)

* fix(spawn): keep the worktree poll from adopting the repository primary

After `treehouse get` is sent, the worktree-discovery poll reads the pane's
foreground-process cwd. While treehouse is still fetching and checking a slot
out, the foreground process is treehouse itself and it reports the repository's
PRIMARY checkout as its cwd for several seconds. The poll accepted any path
that merely differed from the spawning project, so from a linked spawning home
- whose project is itself a worktree of that repository - it adopted the
primary, and the isolation guard then refused a launch whose slot treehouse
went on to create normally.

Screen every candidate with the isolation guard's own conditions, extracted as
spawn_worktree_isolated, so a read the guard would reject stays a transient the
poll keeps waiting through. The two-consecutive-reads rule and the guard as
final backstop are unchanged; a pane that never reaches an isolated worktree
still fails at the existing 60s deadline, now naming the last path it reported.

The already-settled timing assertion counted whole-spawn wall time against a
5s budget and failed on unmodified HEAD on slower machines; it now counts pane
reads, which is what "one confirming read, not an extra cycle" actually means.

* fix(spawn): say which path the worktree wait rejected, and why

Screening every discovery-poll candidate means a host that never reaches an
isolated worktree spends the whole 60s window before refusing. That wait is
deliberate - separating a transient from a terminal misconfiguration needs
machinery this path does not want - so the refusal explains itself instead:
the isolation check records why a candidate failed, and the deadline names the
last path seen together with that reason. Message and diagnostics only; the
poll's control flow is unchanged.

Two suites asserted the guard's wording on paths the poll now rejects rather
than adopts, so their refusal arrives from the deadline instead: realign
fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also
asserting the stated reason, and the second the metadata absence it was
missing) and the herdr projection e2e's forced non-worktree cwd.

* no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test

* no-mistakes(document): document spawn poll isolation screen in fm-spawn header

* test(spawn): make the non-git isolation case non-git anywhere

The refusal-reason assertion for a path outside any repository assumed TMPDIR
is not inside a git repository. Where it is, git walks up from the temporary
directory, finds that repository, and the spawn reports the subdirectory cause
instead - so the case passed or failed on a property of the host rather than on
the behaviour under test.

Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES,
which git documents as not chdir-ing up into a listed directory while looking
for a repository. Git never excludes the directory being searched, so the
ceiling is the parent of the path handed to the spawn.

The assertions pin which cause fired rather than the sentence that explains it,
leaving the operator wording free to improve.

* no-mistakes(document): point spawn poll comment at the isolation screen's comparison

* no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure

* fix: support stock macOS Bash 3.2 paths (#3732)

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): read orphaned green ci monitor and daemon-down failed record as not failed (#3846)

* fix(bin): read an orphaned green ci monitor as held-for-merge, not failed

A no-mistakes run held for a captain merge decision keeps its ci step
polling until merged or closed; when the shared daemon restarts under
that poll, the run is recorded failed although every substantive step
completed and GitHub reports the PR green. A monitor whose only
remaining job is to observe a human decision must not convert the
absence of that decision into a failure verdict.

fm-crew-state.sh now reclassifies a terminal failed run as done
(held-for-merge), surfacing the run's PR URL, when the steps table
shows every step completed except exactly ci failed and the ci log's
last recognized marker reads checks green. A genuinely red check, an
unreadable ci log, or a second failed step keeps the failure.

* no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed

* no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification

* fix(spawn): verify preserved backlog state after interrupted spawn delivery (#3852)

* fix(bin): read preserved spawn state back before the interrupted exit claims it

The deferred-signal exit path asserted the paired task record and
In-flight backlog state were preserved without reading either back,
exactly when a reader is least able to check (fm-yi4j evidence,
2026-09-05). The commit's exit status alone has been observed to agree
with a row that did not actually move.

The exit path now re-reads the record and the row under the same
per-task lock as the commit, repairs a row the commit believed it
moved, and phrases the error as exactly what was verified or attempted
- verified preserved, repaired and verified, or an explicit
preservation-could-not-be-verified with the reason and hand-closeout
instruction. Two behavior tests drive a lying tasks-axi start through a
real interrupted spawn and assert the printed claim and the real
backlog state agree.

* no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask

* no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean

* docs: correct runtime-backend maturity labels for Herdr (#3821)

* docs: correct stale tmux/herdr backend maturity claims

Herdr now has 21 test files, its own required CI job (tests-herdr)
that installs a pinned build and hard-fails on "skip: herdr not
found", while tmux has 3 test files and is only required as a
dependency of the portable-serial e2e lane. zellij, orca, and cmux
still have no CI lane at all. AGENTS.md and docs/herdr-backend.md
still called Herdr merely "experimental" alongside those three,
misleading every session and reader about actual coverage.

Update AGENTS.md's config/backend entry, the opening lines of
docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend
section of docs/configuration.md, and the matching claims in
docs/architecture.md, CONTRIBUTING.md, and README.md so they agree
and distinguish tmux (default), herdr (own required CI lane, largest
suite, Windows still spike-only), and zellij/orca/cmux (still
experimental, no CI lane). No behavior, selection order, or
dispatch logic changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj

* no-mistakes(review): docs: fix stale herdr label and CI-lane wording

* no-mistakes(document): docs: align tmux adapter label in scripts.md

* no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page

* no-mistakes(review): docs: drop windows claim, align contributing backend wording

* no-mistakes(review): docs: trim duplicated CI claim from herdr opening line

* no-mistakes(review): docs: drop unguarded largest-test-suite superlative

* no-mistakes(review): docs: restore tmux verified label and README experimental scope

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: restore intent-targeted no-mistakes validation (#3865)

* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR #823 added and PR #1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: verify Treehouse slot ownership before teardown (#3837)

* fix(bin): verify pool-slot ownership before returning a worktree slot

Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.

Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.

* no-mistakes(review): Protect slots across cloned Firstmate homes

* no-mistakes(test): Gate teardown locking on genuine Treehouse slots

* no-mistakes(test): Clarify pooled descendant slot gating

* no-mistakes(test): Synchronize watcher re-arm test on process exit

* no-mistakes(test): Wait for watcher cleanup before timeout escalation

* no-mistakes(document): Document pool-slot ownership safeguards

* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion

* fix(bin): resolve relative origins from repository root

* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout

* feat(bin): add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary (#3867)

* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary

Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.

Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test: prove the omp guard continuation through a guard spy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* fix(spawn): clear the gemini marker at the omp launch boundary

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): force the guard stage by freezing the watcher and clear lint findings

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): reap the live lab by path and record omp's rpc shutdown as a note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces

Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin

* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs

* no-mistakes(review): omp: pin config-model validation with a test, trim overlay

* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family

* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes

* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux

* no-mistakes(document): docs: add omp subagent-guard row, fix live test header

* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π|󰵗|pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)

* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|󰵗)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix(pi): invoke Bash helpers correctly on native Windows (#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>

* test(bin): pin teardown outcomes for squash-merged rebased branches (#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>

* fix(bin): keep supervision armed for registered custom checks (#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes #3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner

* fix(bin): resolve Treehouse locks for remote secondmate homes (#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership

* fix(bearings): repair board listening and decision reconciliation (#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable…
ktapa added a commit to ktapa/firstmate that referenced this pull request Sep 27, 2026
…ork's four changes (#5)

* fix(bin): prevent false pipeline blocks after drive timeouts (#3813)

* fix(bin): read a crew's pipeline-death claim against the live run

A crew's no-mistakes drive call blocks until the next gate or outcome,
routinely far longer than its harness lets one command live, so the call
gets killed or times out while the daemon runs the fix round on in the
background. Crews read that as daemon death and block on it, and firstmate
had nothing that contradicted them.

Rule 7 of every generated brief now says a drive-call error or a harness
command timeout is not a daemon error, requires `no-mistakes daemon status`
plus `no-mistakes axi status` before a pipeline `blocked:`, and reserves
that report for a refused socket or a run record failed with a daemon
error. The no-mistakes definition of done adds the harness command limit
and the background-and-poll shape that fits inside it.

fm-crew-state gains one classification case: a `blocked:` line blaming the
daemon, a timeout, or unreachability, while the run is running or fixing
AND the pipeline reports fresh activity, now reads as superseded because
the run is alive. Recency comes from the client's own `quiet` marker on
active_steps.last_activity rather than a threshold invented here, and
positive evidence is required, so a run record that outlives a genuinely
dead daemon keeps the plain reading.

stuck-crewmate-recovery gains the inverse-of-a-dead-endpoint playbook:
firstmate reads both statuses itself, steers a reattach, never restarts the
shared daemon on a crew's claim, and escalates only a refused socket.

Nothing here depends on an unshipped no-mistakes capability.

* no-mistakes(review): Prioritize daemon socket failure and narrow unreachable matching

* no-mistakes(review): Honor socket refusal across coarse status and crew guidance

* no-mistakes(test): Replace flaky settle timing assertion with pane-read count

* no-mistakes(document): Document daemon timeout recovery contract

* no-mistakes(ci): Fixed daemon socket failures being suppressed by terminal attributed runs. Positive refused/missing socket evidence now remains blocked regardless of run status. Added a behavioral regression test for terminal failed runs. Verified with fm-crew-state tests, project ShellCheck lint, and git diff checks

* fix(bin): scope the worker role contract for ship and scout launches (#3797)

* fix(brief): scope Firstmate workers to their launch contract

* no-mistakes(document): Clarify supervisor scope and worker contract ownership

* no-mistakes(ci): Removed the heading-based bypass so every ship/scout launch receives the current worker-role contract. Added a regression that failed before the fix and passes afterward. Dispatch, brief, and delivery suites, focused ShellCheck, and git diff --check all passed

* no-mistakes(review): make launch overlay sole owner of worker role contract

* no-mistakes(review): narrow heading test dimension and fix publish error wording

* no-mistakes(review): gate role supersession, fix render guard, drop AGENTS twin

* no-mistakes(document): align architecture AGENTS.md scope and spawn launch-brief header

* fix(bin): stop ringing steering doorbells into dead panes (#3823)

* fix(bin): stop ringing steering doorbells into dead panes

The steering-inbox doorbell was a plain sentence plus Enter typed into a
worker's pane, and the watcher re-rang it on the assumption that a ring is
free. In a pane whose agent has exited that line is a shell command, and the
re-ring ladder kept typing it into a shell that can never acknowledge it.

- Prefix the doorbell with the shell no-op `: ` so a bare shell executes
  nothing while a live worker still reads the same self-describing line.
  `#` is not used because interactive zsh does not treat it as a comment by
  default and the claude harness binds it to memory mode.
- fm_task_inbox_ring skips the pane (return 3) when the backend positively
  classifies the agent as dead; missing, ambiguous, unreadable, and unverified
  endpoints still ring so a blind classifier never starves a live worker.
- The watcher caps the ladder for a dead pane: one stale wake for recovery,
  no ring, no ladder walk, and the durable record stays for
  stuck-crewmate-recovery. fm-send and the remote steer leg report the skip.

Tests cover the no-op in real shells, the dead/live/unclassifiable ring
verdicts, and the single-surfacing watcher path.

* no-mistakes(review): Quote doorbell paths against shell injection

* no-mistakes(review): Reject terminal-control paths before ringing

* no-mistakes(review): Document accepted partial doorbell delivery race

* no-mistakes(review): Skip unavailable endpoints before busy-state handling

* no-mistakes(test): Respect shell startup PATH in environment allowlist test

* no-mistakes(test): Fix doorbell test fixtures for endpoint liveness

* no-mistakes(test): Prioritize confirmed restarts and clean shell test syntax

* no-mistakes(document): Document dead and missing doorbell recovery

* no-mistakes(ci): Fixed the persistence-reply timeout race by rechecking for a correlated reply immediately before falling back to a nudge. Added a deterministic regression covering replies arriving between the preliminary resolution pass and timeout handling. Verified with the targeted restart suite, project lint, coverage guard, bash syntax checks, and diff checks

* fix(bin): wait out transient primary-checkout reads in the spawn worktree poll (#3834)

* fix(spawn): keep the worktree poll from adopting the repository primary

After `treehouse get` is sent, the worktree-discovery poll reads the pane's
foreground-process cwd. While treehouse is still fetching and checking a slot
out, the foreground process is treehouse itself and it reports the repository's
PRIMARY checkout as its cwd for several seconds. The poll accepted any path
that merely differed from the spawning project, so from a linked spawning home
- whose project is itself a worktree of that repository - it adopted the
primary, and the isolation guard then refused a launch whose slot treehouse
went on to create normally.

Screen every candidate with the isolation guard's own conditions, extracted as
spawn_worktree_isolated, so a read the guard would reject stays a transient the
poll keeps waiting through. The two-consecutive-reads rule and the guard as
final backstop are unchanged; a pane that never reaches an isolated worktree
still fails at the existing 60s deadline, now naming the last path it reported.

The already-settled timing assertion counted whole-spawn wall time against a
5s budget and failed on unmodified HEAD on slower machines; it now counts pane
reads, which is what "one confirming read, not an extra cycle" actually means.

* fix(spawn): say which path the worktree wait rejected, and why

Screening every discovery-poll candidate means a host that never reaches an
isolated worktree spends the whole 60s window before refusing. That wait is
deliberate - separating a transient from a terminal misconfiguration needs
machinery this path does not want - so the refusal explains itself instead:
the isolation check records why a candidate failed, and the deadline names the
last path seen together with that reason. Message and diagnostics only; the
poll's control flow is unchanged.

Two suites asserted the guard's wording on paths the poll now rejects rather
than adopts, so their refusal arrives from the deadline instead: realign
fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also
asserting the stated reason, and the second the metadata absence it was
missing) and the herdr projection e2e's forced non-worktree cwd.

* no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test

* no-mistakes(document): document spawn poll isolation screen in fm-spawn header

* test(spawn): make the non-git isolation case non-git anywhere

The refusal-reason assertion for a path outside any repository assumed TMPDIR
is not inside a git repository. Where it is, git walks up from the temporary
directory, finds that repository, and the spawn reports the subdirectory cause
instead - so the case passed or failed on a property of the host rather than on
the behaviour under test.

Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES,
which git documents as not chdir-ing up into a listed directory while looking
for a repository. Git never excludes the directory being searched, so the
ceiling is the parent of the path handed to the spawn.

The assertions pin which cause fired rather than the sentence that explains it,
leaving the operator wording free to improve.

* no-mistakes(document): point spawn poll comment at the isolation screen's comparison

* no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure

* fix: support stock macOS Bash 3.2 paths (#3732)

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): read orphaned green ci monitor and daemon-down failed record as not failed (#3846)

* fix(bin): read an orphaned green ci monitor as held-for-merge, not failed

A no-mistakes run held for a captain merge decision keeps its ci step
polling until merged or closed; when the shared daemon restarts under
that poll, the run is recorded failed although every substantive step
completed and GitHub reports the PR green. A monitor whose only
remaining job is to observe a human decision must not convert the
absence of that decision into a failure verdict.

fm-crew-state.sh now reclassifies a terminal failed run as done
(held-for-merge), surfacing the run's PR URL, when the steps table
shows every step completed except exactly ci failed and the ci log's
last recognized marker reads checks green. A genuinely red check, an
unreadable ci log, or a second failed step keeps the failure.

* no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed

* no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification

* fix(spawn): verify preserved backlog state after interrupted spawn delivery (#3852)

* fix(bin): read preserved spawn state back before the interrupted exit claims it

The deferred-signal exit path asserted the paired task record and
In-flight backlog state were preserved without reading either back,
exactly when a reader is least able to check (fm-yi4j evidence,
2026-09-05). The commit's exit status alone has been observed to agree
with a row that did not actually move.

The exit path now re-reads the record and the row under the same
per-task lock as the commit, repairs a row the commit believed it
moved, and phrases the error as exactly what was verified or attempted
- verified preserved, repaired and verified, or an explicit
preservation-could-not-be-verified with the reason and hand-closeout
instruction. Two behavior tests drive a lying tasks-axi start through a
real interrupted spawn and assert the printed claim and the real
backlog state agree.

* no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask

* no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean

* docs: correct runtime-backend maturity labels for Herdr (#3821)

* docs: correct stale tmux/herdr backend maturity claims

Herdr now has 21 test files, its own required CI job (tests-herdr)
that installs a pinned build and hard-fails on "skip: herdr not
found", while tmux has 3 test files and is only required as a
dependency of the portable-serial e2e lane. zellij, orca, and cmux
still have no CI lane at all. AGENTS.md and docs/herdr-backend.md
still called Herdr merely "experimental" alongside those three,
misleading every session and reader about actual coverage.

Update AGENTS.md's config/backend entry, the opening lines of
docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend
section of docs/configuration.md, and the matching claims in
docs/architecture.md, CONTRIBUTING.md, and README.md so they agree
and distinguish tmux (default), herdr (own required CI lane, largest
suite, Windows still spike-only), and zellij/orca/cmux (still
experimental, no CI lane). No behavior, selection order, or
dispatch logic changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj

* no-mistakes(review): docs: fix stale herdr label and CI-lane wording

* no-mistakes(document): docs: align tmux adapter label in scripts.md

* no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page

* no-mistakes(review): docs: drop windows claim, align contributing backend wording

* no-mistakes(review): docs: trim duplicated CI claim from herdr opening line

* no-mistakes(review): docs: drop unguarded largest-test-suite superlative

* no-mistakes(review): docs: restore tmux verified label and README experimental scope

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: restore intent-targeted no-mistakes validation (#3865)

* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR #823 added and PR #1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: verify Treehouse slot ownership before teardown (#3837)

* fix(bin): verify pool-slot ownership before returning a worktree slot

Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.

Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.

* no-mistakes(review): Protect slots across cloned Firstmate homes

* no-mistakes(test): Gate teardown locking on genuine Treehouse slots

* no-mistakes(test): Clarify pooled descendant slot gating

* no-mistakes(test): Synchronize watcher re-arm test on process exit

* no-mistakes(test): Wait for watcher cleanup before timeout escalation

* no-mistakes(document): Document pool-slot ownership safeguards

* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion

* fix(bin): resolve relative origins from repository root

* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout

* feat(bin): add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary (#3867)

* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary

Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.

Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test: prove the omp guard continuation through a guard spy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* fix(spawn): clear the gemini marker at the omp launch boundary

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): force the guard stage by freezing the watcher and clear lint findings

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): reap the live lab by path and record omp's rpc shutdown as a note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces

Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin

* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs

* no-mistakes(review): omp: pin config-model validation with a test, trim overlay

* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family

* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes

* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux

* no-mistakes(document): docs: add omp subagent-guard row, fix live test header

* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π|󰵗|pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)

* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|󰵗)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix(pi): invoke Bash helpers correctly on native Windows (#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>

* test(bin): pin teardown outcomes for squash-merged rebased branches (#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>

* fix(bin): keep supervision armed for registered custom checks (#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes #3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner

* fix(bin): resolve Treehouse locks for remote secondmate homes (#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership

* fix(bearings): repair board listening and decision reconciliation (#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable state - refuses
rather than acting.

Regression coverage fails without each fix, and pins every leak path: a bare
reconcile, a standalone close or note with no pending request, an any-channel
reconcile, an annotated selection from a freeform card, and a
generation-skewed authorization. An opt-in guard re-proves the lavish-axi
shapes and the reopen against the installed tool.

* fix(bin): quote the done comparison in the reconcile intake

shellcheck SC1010 reads the bare word as the loop keyword. The failed run
never reached its lint step, so this shipped in the recovered content.

* no-mistakes(review): Publish reconciled parent resolution before request retirement

* no-mistakes(review): Clarify committed cleanup and reconcile reservation scope

* no-mistakes(review): Preserve remote cards and legacy answer compatibility

* no-mistakes(test): Separate live claim release from stale reclamation

* no-mistakes(test): Allow terminal self-retirement during active capture

* no-mistakes(document): Document Bearings repair contracts

* no-mistakes(ci): Stabilized the failing Herdr presentation E2E by serializing test-harness Treehouse allocator calls, preventing concurrent recovery spawns from claiming the same pool slot while preserving Herdr concurrency coverage. Verified with the full E2E suite on Herdr 0.8.2, bash syntax checks, ShellCheck, and git diff checks

* fix(bin): allow pooled spawns without a git origin (#3885)

* fix(bin): skip pooled-worktree freshness fetch when no origin is configured

An origin-less local-only project has nothing remote to be stale
against, so fm-spawn's freshen_spawn_worktree_base refused to launch
crews for it. Detect a missing origin remote and skip the fetch
freshness gate entirely; an existing-but-unreachable origin keeps
refusing as before.

* no-mistakes(review): Preserve pool safety for absent and unusable origins

* no-mistakes(review): Refuse empty origin configurations during pooled spawn

* no-mistakes(review): Detect empty origin sections across config includes

* no-mistakes(review): Honor globbed includes when detecting origin configuration

* no-mistakes(review): Document conservative conditional include handling

* no-mistakes(review): Use Git-resolved config files for origin detection

* no-mistakes(review): Document included empty-origin detection boundary

* no-mistakes(document): Document originless pooled spawn behavior

* feat(tests): run live harness guards by default when available (#3889)

* feat(tests): run live harness guards by default where the harness is installed

The 24 live-harness guards each opened with their own env check, so on the
machine that has every harness - the one the product and its validation
actually run on - all of them skipped and passed. Fourteen had never been run
by the pipeline at all.

tests/lib.sh gains fm_live_gate as the single owner of that decision: a guard
that spends no model tokens runs wherever its tools are installed, a guard that
submits prompts stays opt-in, an absent tool is a named capability skip, and a
guard's own variable or FM_LIVE forces it on (turning an absent tool into a
failure) or off. Every live guard now opens with it, which also carries the
test-suite gate-refusal bypass into the guards that never sourced the shared
helpers and were therefore refused whenever a gate agent ran them.

bin/fm-test-run.sh records what a skip means: the family's expected class is
live-capability rather than a bare env opt-in, and each gate skip's reason is
logged and written to the timing artifact, so a lane can say which tool this
host could not exercise.

Only the token-free guards flip to default-on: composer-matrix, the harness
liveness drift guard, and the Herdr version floor. cursor-primary submits three
prompts, so it stays opt-in.

Running the drift guard unasked immediately found a real defect it existed to
catch: it resolved the harness through a generic `command -v cursor`, which on
a machine that also has the Cursor editor finds the editor launcher rather than
cursor-agent. That binary exits at once, leaving a bare shell in the pane and a
liveness-drift failure no classifier change could fix. It now asks
fm_cursor_resolve_binary first, the same verified owner fm-spawn uses.

CI installs the public Pi package in the portable serial lane and fails on its
skip token, so the Pi extension tests stop passing silently against a package
that is not there. No secret is added.

Verified on macOS 26.5.2 arm64: the drift guard runs with no variable set and
classifies 8 installed harnesses alive; the Herdr version-floor guard runs by
default and checks 4 real releases; every live guard refuses together under
FM_LIVE=0.

* fix(tests): keep the composer-matrix guard opt-in

Running it unasked is red on a healthy machine for reasons no code change here
removes: a harness that has not trusted this checkout sits on its own trust
dialog, which the guard treats as an unreadable composer and correctly fails.
The opencode 1.18.29 and grok 1.0.13 composer drift it also surfaced reproduces
identically on main and is filed as separate work.

So this token-free guard stays opt-in with the reason stated in its header, and
the coding guidelines record the narrow exception: a guard whose verdict
depends on host state that installing its tools does not establish may stay
opt-in, because one that is permanently red is one the fleet learns to ignore.
The other two token-free guards keep running by default.

* no-mistakes(review): Wire bearings guard and remove composer exception policy

* no-mistakes(review): Run Pi responsiveness guard by default

* no-mistakes(review): Gate AFK Pi Herdr through authoritative family sweep

* no-mistakes(review): Sanitize live gate test environments

* no-mistakes(document): Document default-on live guard behavior

* no-mistakes(ci): Fixed CI by installing the Pi package in portable-parallel-1, where fm-pi-primary-types.test.sh runs, and enforcing its package-missing gate skip there. Verified with fm-lint.sh, workflow actionlint, coverage partition checks, lane membership, and git diff checks

* no-mistakes(ci): Fixed CI’s Pi typecheck skip enforcement by giving npm, tsc, and Pi-package capability skips a shared prefix and configuring both relevant CI lanes to fail on that prefix. Verified missing tsc emits the expected skip, missing Pi package becomes a runner failure, and actionlint, ShellCheck, and git diff checks pass

* fix(bin): refuse test runs in the primary checkout when a task marker is set (#3891)

* fix(bin): refuse the behavior suite in the repository primary checkout

A task worker's isolated worktree placement is verified exactly once, when
its task starts, and nothing re-checks it afterwards. A worker that later
changes directory into the repository's primary checkout runs its Git
commands, and this branch-switching suite, against the one checkout every
linked worktree resolves against and every landing merges into. A run that
dies mid-suite can leave that checkout on a stray branch.

bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task
worker and the runner resolves to the primary checkout, every executing mode
exits non-zero before selecting a suite, with one line naming the primary
path and pointing at the assigned task worktree. The predicate is the one
bin/fm-spawn.sh already uses for launch placement: the working tree's own
git dir is the repository's common git dir, which separates the primary from
every linked worktree even when their top levels differ. A run with no
FM_TASK_ID set is unchanged, and so are the inspection modes, which execute
nothing. When git resolves neither directory - a non-repository fixture, a
detached copy - nothing proves this is the primary, so the run proceeds.

bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID
into the pane shell on the same pre-launch channel as GOTMPDIR, and the name
joins the sanitized launch environment allowlist so an isolated launch keeps
it.

* no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT

* no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal

---------

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): bound stale alarms for backlog captain holds (#3842)

* fix(bin): bound a stale alarm with the backlog hold, not only the status line

A legitimate wait has two records and the stale alarm reads only one.
`status_is_paused_or_captain_held` takes a status line, so it sees a wait the
worker declared. It cannot see the wait firstmate records when it hands work to
the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and
leaves the status log alone, so a delivered task keeps `done: PR ...` as its
last line for the whole time the captain is deciding.

Both stale branches were blind to it, and each churned a new pane hash back into
its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale
branch, while a held task whose last line is `working:` reaches
`surface_nonterminal_stale` and fails its declared-wait test.

Consult that second record where the watcher is about to alarm, through
`bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and
bound the alarm on the shared `.paused-resurfaced-<key>` marker and
`PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first
sight still alarms, the window's end alarms once more, and a held crew that goes
genuinely silent still escalates through the wedge timer.

Only an established open captain call bounds anything: an unreadable backlog, an
absent or incompatible tasks-axi, a row this home does not carry, and every task
with no hold keep alarming exactly as before. The backlog hold is deliberately
not recorded as a declared pause, because the loop-top reconciliation and
`pause_state_class` both read the status line and would clear a flag that line
does not support.

Extends the fix in #3443, which closed the forms of this loop that the status
line itself can express.

* fix(bin): identify the captain call a stale alarm is bounded by

Three gaps in the bound added by the previous commit, all in how the throttle is
scoped and where the backlog is consulted.

The scope carried only the status-log signature. A task can be held, answered
with `--release`, and re-held as a genuinely different captain call without any
status append, so the second call inherited the first one's marker and its first
sight was absorbed - the one thing this bound must never do. The task id is not
the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the
call's own lifecycle - its hold-set stamp and the number of recorded answers -
on an exit 0 and only then, leaving the silent predicate every existing caller
reads unchanged. The throttle scope now carries that identity.

The terminal path recorded the throttle before publishing the durable wake. A
failed append exits the watcher with nothing queued, and the next sighting then
read that fresh marker and absorbed the retry, turning a delayed alarm into a
lost one. Recording moves behind the append, as the non-terminal path already
had it, and the comment claiming the marker could not outlive its wake is gone
because it was false.

The backlog was consulted only on a new terminal pane hash. A captain call can
open after a hash was absorbed as provably working, changing neither the pane nor
the status log, so nothing re-read the backlog and the wedge timer kept firing
possible-wedge alarms through a legitimate wait. That timer now consults the call
at its own alarm boundary and takes the same bounded cadence - and only at that
boundary, so an ordinary repeat poll under the bound stays the local-only read it
was.

Regression coverage for each, all driving churn through one watcher process
rather than relaunching per pane change: relaunch cost dominated the earlier
shape, and an absorbing watcher stays in its poll loop across churn in production
anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer
with no open captain call still escalates as a possible wedge.

* fix(review): Compose stale throttles with captain-call lifecycle identity

* fix(review): Preserve bounded same-hash captain-call resurfacing

* revert(bin): narrow the captain-hold stale bound to its observed defect

Lifts the lifecycle-identity and cadence-ownership work back out, leaving the
change at the shape that matches the defect actually observed: the stale alarm
did not consult the backlog captain hold, on either stale branch.

Reviewing the wider version surfaced a series of adjacent gaps in the watcher's
alarm state machine - a call opening after the first alarm, marker invalidation
at the hold lifecycle boundary, and which deadline a terminal timer represents.
They are real, but fixing them turns a small extension into a state-machine
change to the alarm path, which is a different review on a subsystem that is
being actively reworked. They are named as known limitations rather than carried
here, and none of them is load-bearing for what remains: the bound does strictly
less than the reverted version, leaves the wedge path escalating on
STALE_ESCALATE_SECS exactly as before, and introduces no silence that the
existing terminal-alarm path did not already have.

Kept from the reverted work is the record-after-append ordering, because that is
a defect in the code being shipped rather than an adjacent one: recording the
cadence marker before publishing the durable wake let a failed append lose an
alarm outright instead of delaying it.

History is preserved: the earlier commits stay on the branch and this removal
sits on top of them.

* fix(review): Document secondmate captain-hold scope boundary

* fix(document): Document captain-hold stale alarm scope

* fix(bin): bind the stale throttle to the captain call, not the status log

The throttle this change introduces was scoped to the task's status-log
signature. Answering a call with `--release` and holding the task again creates a
genuinely different captain call without necessarily appending to that log, so
the second call inherited the first one's marker and its first sight was
absorbed.

That is the one alarm this bound must never swallow. A delivery announced twice
is noise; a decision waiting on the captain that is never surfaced is invisible,
because nobody asks for what they do not know to ask for.

Measured rather than assumed, on the same fixture - a delivered task held for the
captain, released, and re-held with no status append, driven through bin/fm-watch.sh:

  base c499f84    call-1 first=ALARM  call-1 churn=ALARM     new call first sight=ALARM
  before this fix call-1 first=ALARM  call-1 churn=absorbed  new call first sight=absorbed
  after           call-1 first=ALARM  call-1 churn=absorbed  new call first sight=ALARM

Base never suppresses the new call, so the suppression came from this change and
closing it completes the fix rather than widening it.

`bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle -
its hold-set stamp and count of recorded answers - on an exit 0 and only then, so
the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope
carries that identity beside the status signature.

The sibling case was measured too and is NOT included: on the status-declared
path, where the last line is `captain-held:`, base already absorbs a re-held
call's first sight. That behaviour predates this change and stays documented as a
known limitation rather than repaired here.

* fix(document): Document captain-call throttle lifecycle scope

* fix(ci): isolate the Herdr restart fixtures from a claimed worktree

The Herdr behaviour test intermittently reused a local worktree still claimed
by an earlier fixture after a restart.
The restart scenarios now use an isolated Treehouse project.

The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as
well.

* feat(pi): resolve extension-registered providers in the supervision branch (#3871)

* Let the supervision branch resolve extension-registered providers

The isolated branch ModelRuntime cannot see providers an extension
registered into main's runtime at run time, so a pin on pi-devin-auth's
devin/swe-1-7 (or an unpinned branch following a main session on devin)
failed with "unavailable to the isolated branch runtime".

Capture main's ModelRegistry alongside mainModel and copy each
extension-registered provider config into the branch runtime at
model-resolution time. The config carries the provider's own streamSimple
and oauth wiring by reference, so the custom gRPC transport reaches the
branch unchanged instead of being reimplemented. The /supervision-model
picker uses the same copy so those models are offered.

Update configuration.md and pi-supervision-branch.md, which previously
stated extension-registered providers were not offered.

* no-mistakes(document): docs: own devin provider carve-out in branch architecture doc

* no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered

* ci: retrigger flaky Herdr/serial-1 lanes

* no-mistakes(document): docs already cover branch extension-provider copy

* fix(bin): gate secondmate wake-loop stall alerts on real queue no-progress (#3943)

* fix(watch): detect stalled secondmate queue progress

* no-mistakes(review): gate secondmate stall on active turns and progress episodes

* no-mistakes(document): align secondmate wake-stall docs with progress-episode detector

* no-mistakes(document): clarify active-turn gate in wake-stall config docs

* no-mistakes(ci): Fixed the one real defect behind the failing checks. ROOT CAUSE (Greptile P1, real code defect in this PR): `secondmate_wake_stall_tick` in bin/fm-watch.sh reset the no-progress timer only when the oldest actionable queue sequence INCREASED (`[ "$seq" -gt "$observed_seq" ]`). When a secondmate is retired and reprovisioned under the same task ID, its fresh home's queue sequence restarts BELOW the recorded position, so the comparison is false, no reset happens, and the new queue inherits the retired generation's already-expired idle interval — emitting a false `secondmate wake-loop stalled` on its very first observation. That is precisely the false-alarm class the user intent requires this PR to remove. FIX (smallest, removal-first): bin/fm-watch.sh:749 now resets when the drain position MOVES AT ALL (`-ne` instead of `-gt`). Draining moves it up, reprovisioning moves it down; neither is a continued no-progress episode. The asymmetric `-gt` branch is removed rather than special-cased or hardened. Updated the function header comment plus the two doc sentences in docs/architecture.md and docs/configuration.md that stated the old advance-only semantics. REGRESSION TEST: added `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` to tests/fm-wake-queue.test.sh (registered in the invocation list). It drives the real watcher through retired generation (seq 9) -> reprovision (seq 3, later clock) -> freeze, asserting observable wake-queue output, no source-text inspection. VERIFICATION: - Fails before / passes after: with the fix reverted the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted`; with the fix it passes, and its third leg confirms the restarted generation still escalates on a genuine freeze (row=3 idle=2s), so the fix does not merely mute the alarm. - `bash tests/fm-wake-queue.test.sh`: exit 0, 38/38 pass, all five secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs (34152740610, 34152740604) both ended with conclusion `action_required` — workflow approval pending, not a test/build failure. Separately, `bin/fm-test-run.sh --check-coverage` exits 1 in this environment, but I confirmed by stashing my changes that it fails identically on the unmodified base tree (locale-related `comm: input is not in sorted order`); it is pre-existing and this change adds no new test file for the partition to account for. Changes are left uncommitted in the worktree

* no-mistakes(ci): Fixed the one real code defect behind the failing checks. ROOT CAUSE (Greptile P1, second round, on the head commit e1304e6): `secondmate_wake_stall_tick` in bin/fm-watch.sh identified the queue's drain position by the sequence number ALONE. The previous round changed the comparison from `-gt` to `-ne`, which handles a reprovisioned queue that restarts BELOW the recorded position, but not one that restarts ON it. A mate retired and reprovisioned under the same task id gets a fresh home whose wake-queue sequence counter restarts at 1 — and the retained parent progress marker very plausibly holds a low sequence too (a queue frozen on its first row records seq 1). Equal sequence ⇒ no reset ⇒ the brand-new queue inherits the retired generation's long-expired idle interval and emits a false `secondmate wake-loop stalled` on its very first observation. That is exactly the false-alarm class this PR exists to remove. FIX (smallest, removal-first): the file already defines the identity of a queue row once, as `row_key="$epoch-$seq"` (used for stall receipts, the stall marker, and the notify key). The progress marker's separate, weaker seq-only identity is removed: `row_key` is now computed once right after the row is parsed, …
wonder-media-fleet Bot pushed a commit to wonder-media/firstmate that referenced this pull request Sep 27, 2026
* fix(bin): resolve captain holds and legacy teardowns on non-markdown backends (#3782)

* fix(bin): close legacy records on the Beads backend honestly

Two pre-Beads reads blocked honest closure of leftover records:

1. fm-captain-hold.sh complete/verify resolved attested legacy hold ids
   only against the live backend and the pre-collapse derived identity, so
   a home whose holds fm-hold-migration rehomed under fm- ids failed with
   an empty-name absence message (the resolve failure was swallowed by the
   command substitution feeding verify_hold_durable). Resolution now falls
   back, on the Beads backend only, to the legacy id under the configured
   beads prefix and to the row whose notes carry the exact marker line
   'migrated from data/backlog.md id <legacy id>'; every refusal names the
   id it could not resolve, and the markdown path is unchanged.

2. fm-teardown.sh refused any record without spawn_gen forever. A record
   that predates the field can now be torn down with an explicit
   --legacy-record flag once the recovery-grade endpoint classifier
   confirms the recorded endpoint dead or agent-less; the accepted
   incarnation is stamped into the record right before its close marker
   binds to it and named in the teardown line. Refusals leave the record
   byte-identical, the unlanded-work refusal is not relaxed, and a corrupt
   (multi-valued) spawn_gen is never accepted.

The companion repair this branch carries (follow-up commit) is the
backend-gated --file and markdown-file requirement in the mutate path and
lifecycle gates: fm_backlog_mutate passed --file and required the markdown
backlog file regardless of the resolved backend, and the transition gate
plus row probe required that file before any backend work, so a home on a
non-markdown backend could neither gate, probe, nor close its rows.

Behavior tests: self-contained beads fixtures over a scratch bd graph
(self-skipping on markdown-only tasks-axi installs), legacy meta fixtures
for every teardown gate, and the relocated markdown backlog coverage stays
green.

* no-mistakes(review): fix(review): report migrated-hold scan refusals and guard legacy spawn_gen stamp against newline-less records

* fix(backlog): address the configured backend for lifecycle writes

Completes the fm-backlog-transition-lib repair the first commit's message
claims: on this base fm_backlog_mutate passed --file and required the
markdown backlog file regardless of the resolved backend, and
fm_backlog_transition_applies plus fm_backlog_row_probe required that file
before any backend work, so a home on a non-markdown backend could neither
gate, probe, nor close its backlog rows. All three now gate the markdown
file on the resolved tasks-axi backend: markdown keeps exactly its explicit
<data>/backlog.md behavior, non-markdown homes address the backend their
own configuration selects with no markdown file requirement.
fm_backlog_row_show and fm_backlog_row_list already gated correctly and
are unchanged. docs/configuration.md owns the contract line.

Also extends the same backend gate to fm-captain-hold.sh's own mutation
wrapper - hold/add/update/answer/done append the markdown --file only when
the resolved backend is markdown, so a captain call on a Beads home reaches
the Beads store end to end - and applies the review round's two direct
remedies there: the [beads] graph path resolves against the backlog root
when relative (never the process CWD), and a failed bd graph read reports
bd's own trimmed stderr reason in the refusal.

Coverage: tests/fm-backlog-atomicity.test.sh gains a stub-driven Beads
completion case proving the transition gate applies, the row probe reads,
and done runs without any markdown file or --file override; the relocated
markdown backlog test stays green.

* no-mistakes(review): Document root-tasks.toml-only beads settings for migrated-hold resolution

* test(gotmp): stub fm_tasks_axi_backend so the fixture matches the backend-aware transition lib

The legacy-records change made fm-backlog-transition-lib.sh resolve the
configured backend via fm_tasks_axi_backend before the markdown-only skip.
The gotmp fixture's fm-tasks-axi-lib stub lacked that function, so the
markdown check fell through and teardown hit the incompatible-backend
error with unbound FM_TASKS_AXI_MIN under set -u. Stub the backend as
markdown and define the floor, restoring the intended no-backlog skip.

* fix(teardown): roll the legacy stamp back when the close marker fails

A legacy-record teardown stamps its accepted incarnation into the record
right before the close marker binds to it; when that marker write then
fails, the stamp survived, so a retried teardown sailed past the
dead-or-agent-less endpoint gate the stamp now proved unnecessary. The
failed marker write now truncates the record back to its exact pre-stamp
bytes (verified by size), restoring the byte-identical-refusal invariant;
when the rollback itself fails the operator is told to re-run with
--legacy-record after reconciling the endpoint.

Also completes the recorded review decision's coverage wording: the
beads stub test now drives the answer close end to end (update and done
through the gated wrapper), asserting no markdown file override reaches
either verb.

* fix(review): harden the legacy stamp rollback and resolve derived migrated ids

The legacy-record stamp rollback now uses perl (already in the teardown
curated PATH; truncate is not, and is absent on stock macOS), routes every
failure branch inside the stamp block through the same size-verified
rollback so the byte-identical-refusal invariant holds on those paths too,
and gains behavior coverage: an unrecordable close (an invalid pr= link)
fails the teardown, leaves the record byte-identical, keeps the backlog
row in flight, and a flag-less retry still refuses.

Migrated-hold resolution now probes the derived pre-collapse identity
(<origin>-decision-<entry>) alongside the raw entry - fm-hold-migration
recorded the DERIVED id in every migrated row's marker note - in both the
prefix and the migration-note forms, with the ambiguity refusal naming
every identity tried, plus behavior coverage for a bare decision key
resolved through its derived identity's marker.

Also aligns fm-backlog-transition-lib.sh's header ADDRESSING/SCOPE
paragraphs with the backend-gated contract, drops an unreachable FORCE
validity guard the parser rewrite left behind, and switches the new stub
fixture to the portable sed -i.bak idiom.

* no-mistakes(review): Name the configured backend in teardown's backlog reminder

* no-mistakes(review): Scan migration markers before the prefix guess

* no-mistakes(review): Document marker-first resolution and cover the prefix branch

* no-mistakes(document): Record prefix-attestation audit and marker-line forms

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-teardown.sh: a failed rollback of the synthetic legacy stamp let a retry bypass the dead-or-agent-less endpoint gate. Root cause: teardown minted `spawn_gen=legacy-<ts>-<pid>` into the task record before the close marker bound to it. When the close-marker write failed AND the rollback also failed, the record retained that token. On the next invocation `fm_backlog_meta_spawn_gen` succeeded, so `TEARDOWN_LEGACY_PENDING` stayed 0 and the endpoint gate was skipped entirely — even with `--legacy-record`. The script's own error text told the operator to "re-run teardown with --legacy-record", advice the code could not honor. Fix (bin/fm-teardown.sh): - A `legacy-*` spawn_gen is now recognized as a stamp this teardown path minted, never one a spawn published (fm-spawn.sh publishes `s<epoch>.<pid>.<random>`). Such a record still reads as the legacy record it is: it re-enters the endpoint gate, and a flag-less retry refuses naming `--legacy-record`. - Acceptance reuses the retained token instead of minting a second one; the append block is skipped when the record already carries it, so no duplicate spawn_gen is written. - The rollback attempt and its "could not be rolled back" message are guarded to runs that actually appended a stamp, so a run that appended nothing never claims a rollback it did not perform. - Usage header documents the retained-stamp rule. Test (tests/fm-teardown.test.sh): added `test_retained_legacy_stamp_still_faces_the_endpoint_gate`, an end-to-end reproduction — a `perl` stub that fails only the rollback's `truncate` (delegating every other perl call to the real interpreter) leaves the stamp behind, then the retry must still hit the gate, must not stamp a second incarnation, must not close the backlog row, and the flag-less retry must refuse. Verification: the new test fails against the pre-fix script on exactly the reported defect ("the retry skipped the dead-or-agent-less endpoint gate") and passes after. Full tests/fm-teardown.test.sh 80 ok / 0 failures / rc=0; tests/fm-backlog-atomicity.test.sh 80 ok / 0 failures / rc=0; bin/fm-lint.sh (pinned ShellCheck 0.11.0 + actionlint 1.7.12) clean

* fix: reduce local ShellCheck source-analysis cost (#3778)

* fix(lint): drop source following on the local changed-file gate

The local lint step was inlining library closures through --external-sources
and peaking above 8 GB on a single root. Keep full analysis in CI, on main,
and without a merge-base; exclude the four cross-file codes from the local
pass so those findings still land in CI.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Run local ShellCheck per root and document measurements

* no-mistakes(review): Correct local source-following telemetry

* no-mistakes(document): Clarify context-sensitive lint documentation

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(pi): route decision-owned wake batches to main (#3776)

* fix(pi): route needs-decision wakes and mixed batches wholly to main

Skip the supervision branch for every needs-decision status append, the
same way a check-kind wake already skips it. A coalesced signal/stale
trigger batch containing any needs-decision row is delivered wholly to
main, not split between the branch and a later main wake - the whole
batch, including any co-present routine rows for a different task,
travels together. Heartbeat and unread-status scans stay independent.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* test(pi): cover distinct-file mixed batches and heartbeat independence

Add a regression using two distinct files (not the same status file
twice) in one coalesced trigger so a some-vs-every regression on the
file-list cross-reference cannot hide behind a degenerate same-key
case, and a heartbeat/needs-decision co-presence test proving a
needs-decision row neither vetoes nor rides along with an otherwise
eligible heartbeat scan.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013LB2CeerSMCZN4oNsVfLeE

* no-mistakes(document): Clarify needs-decision and heartbeat routing

* no-mistakes(ci): Fixed captain-held stale reminders so they bypass supervision and wake main, while unrelated unread rows and heartbeats remain independent. Added routing regressions and updated documentation. Pi watcher tests, strict TypeScript checks, lint, and diff checks pass. The no-mistakes attestation failure was pipeline-state related, not a source defect

* no-mistakes(ci): Fixed CI lint by narrowly suppressing false-positive SC2031 diagnostics where background PIDs are captured immediately in the same shell. Verified with `CI=true bin/fm-lint.sh` and `git diff --check`. The no-mistakes attestation failure is pipeline-state-related (`test` was skipped), not a source defect

* no-mistakes(review): Route stale open decisions directly to main

* no-mistakes(review): Honor configured verbs in stale decision routing

* no-mistakes(review): Route second-mate escalations and configured decisions to main

* no-mistakes(review): Ignore trailing whitespace after captain holds

* no-mistakes(review): Cache stale decision classification per status file

* no-mistakes(review): Document unread decision precedence for later task wakes

* no-mistakes(review): Cache unchanged stale decisions across scope scans

* no-mistakes(review): Resolve decision aliases and reject symlinked statuses

* no-mistakes(review): Route surfaced captain-held signals directly to main

* no-mistakes(document): Document decision-owned main routing

* no-mistakes(ci): Fixed captain-held spans to remain actionable while crew working evidence is positive, ensuring the watcher delivers their main-only marker. Updated the executable regression test to cover this case. Verified with the full fm-watch-triage suite, bash syntax checks, and git diff checks. Shellcheck reported only pre-existing test harness warnings (SC1091/SC2034)

* no-mistakes(ci): Fixed the CI regression: captain-held transfers now retain their established non-actionable stale classification while the signal-routing side-band still surfaces them main-only. Verified with tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh, bash syntax checks, and git diff --check

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* feat(bin): add opt-in worker launch environment allowlist (#3802)

* feat(spawn): add an opt-in worker environment allowlist

Honor a home-local launch-env-allowlist at the shared worker command
boundary and inherit it into secondmate homes. Preserve the existing
launch behavior when the file is absent. Keep the operational environment
and explicit launch assignments, and account for filtered Muse credentials.

Refs https://github.com/kunchenguid/firstmate/issues/3742

Verification:
- Red on origin/main 1820316b66ac2c68e244dd04a02512859ee8c1f4:
  the new enabled-allowlist regression observed synthetic-unrelated in
  the worker; the absent-file control passed.
- Green: fm-test-run.sh on fm-spawn-dispatch-profile, fm-muse-harness,
  and fm-trace-context-spawn; all three passed without skips.
- Synthetic emitted-command probes ran through sh, stock Bash, and zsh.
- Canonical lint, documentation audience checks, and stock Bash syntax
  checks passed.

* no-mistakes(review): Reject inaccessible launch environment configuration

* no-mistakes(review): Preserve inherited allowlists on source inspection errors

* no-mistakes(document): Clarify worker environment grants and inheritance documentation

* no-mistakes(lint): Fix inheritance test ShellCheck source boundary

* feat(bin): add rovo crewmate/scout adapter with home-path file access (#3575)

* feat(bin): add rovo as a verified crewmate/scout worker harness

Wire the Atlassian Rovo CLI (202609.1.2) into the TUI-under-tmux/herdr
adapter contract: detection with marker-precedence ordering, one-shot
positional launch with --startup-receipt readiness polling instead of
composer scraping, model/effort flags, a screen-scrape busy fallback
scoped like grok's, and crew/scout-only lifecycle control that refuses
secondmate launches. Ships with a portable regression suite, a live PTY
guard against the real binary, a per-harness reference doc, and a dated
verification record covering the silent OAuth refresh, the interrupt-ack
divergence from the originating scout report, and the still-open
composer-ghost and tmux/herdr pane-liveness gaps.

* no-mistakes(review): revert rovo launch to positional brief, drop startup-receipt

* no-mistakes(test): rewire rovo adapter to kimi-style launch-then-send shape

* no-mistakes(document): add rovo to stale worker-harness enumerations in docs

* docs(verification): close the rovo herdr-liveness gap with live isolated-lab evidence

Placement, launch-then-send, and busy/idle rendering are now verified live
in an isolated non-default Herdr lab session (bin/fm-herdr-lab.sh), driven
directly through fm-spawn.sh's/fm-backend.sh's own shared primitives since
the cross-session launcher-identity guard refuses this task's own ambient
Herdr identity for a full fm-spawn.sh run.

fm_backend_agent_state reported dead for a live, responding rovo pane at
every point checked, because herdr's own agent-integration registry has no
rovo entry (herdr integration status), so herdr agent get returns
agent_not_found regardless of whether rovo is actually running. This is
recorded as a Herdr-side integration gap rather than a firstmate bug, left
unpatched to avoid a false-positive alive verdict for other idle shells.

Updates docs/verification/rovo.md's backend-liveness section and its two
cross-references (docs/verification/runtime-backends.md, docs/configuration.md)
accordingly.

* fix(bin): close rovo's failed-spawn leak and busy-scrape false idle

Greptile P1s on PR #3575: a failed rovo readiness/submission/delivery gate
exited without tearing down the just-created endpoint, leaving the launched
--yolo rovo process running as an orphaned agent outside task control.
Separately, the busy classifier's rendered-tail fallback returned definitive
idle whenever the "Rovo is thinking" marker scrolled out of the last 12
nonblank lines of a long turn, which could make supervision wrongly conclude
a still-working worker had gone idle.

fm-spawn.sh: rovo_spawn_fail now calls rovo_endpoint_cleanup, which kills the
created endpoint (tmux/herdr/zellij/cmux) via the same generic fm_backend_kill
dispatch fm-spawn.sh's own orca-abort path already uses; orca's worktree and
terminal remain owned by the separate ORCA_ABORT_CLEANUP trap.

fm-busy-lib.sh: the rovo classifier arm now reports "unknown rovo-regex"
instead of "idle rovo-regex" when the marker is absent, matching how muse and
cursor already express "can't tell" for their own fallbacks. The positive
busy match is unchanged.

Extends tests/fm-rovo-harness.test.sh: the readiness and delivery failure
tests now assert the endpoint is torn down (and the success test asserts it
is not), and a new test drives the busy marker out of the tail window to
confirm the verdict is unknown, never idle. bin/fm-lint.sh is clean on both
changed files.

* test(rovo): align spawn fixture with the launch-brief validation contract

Upstream main now requires a brief's ## Captain's intent and
## Firstmate spec subsections (or a nonempty legacy # Task body)
before spawn, and rewrites ship+no-mistakes briefs into
launch-brief.md. Update the rovo harness fixture and pointer
assertions to match, mirroring the kimi harness fixture.

* no-mistakes(review): align rovo.md delivery-gate note with live herdr evidence

* fix: prevent stale supervision wake loops (#3672)

* fix(bin): stop the supervision branch's stale-ack and ghost-report loops

Clean-slate implementation of the four authorized recommendations from the
supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal
form, superseding PR #3604:

- fm_branch_report refuses a task the wake being handled never named. The
  extension fixes the reportable task set from the eligible rows before each
  prompt (signal and stale rows resolve to their tasks, a heartbeat allows any
  task with a live record, fleet is always allowed), so a report typed from
  memory about a task whose records teardown already removed is never stored
  or delivered.
- An acknowledgement that consumes nothing says "nothing was acknowledged
  through N" and prints the exact --ack-through / --recovery-generation
  command for the current presented wake, instead of "re-run the drain",
  which re-fed the same stale acknowledgement in a loop.
- bin/fm-guard.sh no longer tells the branch actor to drain queued wakes
  while it is handling them; it names the granted rows instead.
- Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and
  descendants; the index rebuild and the append-side index write both skip a
  task with neither a live record nor a status log, so the branch's report of
  a teardown it just performed is stored without recreating the index.

No new locking, no spawn-generation binding, and no retired-task refusal: the
branch can still report the outcome of a task it just tore down, and the
teardown test now proves that path end to end.

* fix(bin): narrow the branch report scope and guard silence to the minimal form

Apply the four review decisions on the clean-slate branch:

- A signal or stale prompt may report only the tasks its own rows resolve
  to; fleet is refused there too. A heartbeat review is not scoped by task
  at all, so the extension no longer tracks live task records and refuses
  nothing by task id during a fleet review.
- The outcome-index rebuild no longer skips retired tasks; the append-side
  skip alone keeps a torn-down task's index from being recreated.
- bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor
  instead of printing a replacement note.

* no-mistakes(document): Align supervision docs with scoped wake handling

* fix(bin): grant rovo the per-task home paths its standard crewmate flow needs

rovo confines every file-tool operation to its worktree by default, and its
bash tool independently refuses the same external paths regardless of any
grant (confirmed live), so a rovo worker could not read its own brief or
steering messages or write its status/report - all of which live in the
firstmate home outside the worktree - without hand-feeding it. Grant
toolPermissions.allowedExternalPaths for exactly the task's brief directory,
steering inbox, and status file at launch time via --config-override,
merged with agent.efficiencyLevel into one JSON object since that flag is
single-value and silently discards a second occurrence.

Extends the live PTY guard to prove, against the real binary, that the
grant lets rovo read an external brief and append to an external status
file, and that the same flow is blocked without the grant.

* no-mistakes(document): align rovo reference Effort row with merged single --config-override

* no-mistakes(document): document rovo file-access grant in harness reference

---------

Co-authored-by: PUNEET PATWARI <ppatwari@atlassian.com>
Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>

* fix(bin): prevent false pipeline blocks after drive timeouts (#3813)

* fix(bin): read a crew's pipeline-death claim against the live run

A crew's no-mistakes drive call blocks until the next gate or outcome,
routinely far longer than its harness lets one command live, so the call
gets killed or times out while the daemon runs the fix round on in the
background. Crews read that as daemon death and block on it, and firstmate
had nothing that contradicted them.

Rule 7 of every generated brief now says a drive-call error or a harness
command timeout is not a daemon error, requires `no-mistakes daemon status`
plus `no-mistakes axi status` before a pipeline `blocked:`, and reserves
that report for a refused socket or a run record failed with a daemon
error. The no-mistakes definition of done adds the harness command limit
and the background-and-poll shape that fits inside it.

fm-crew-state gains one classification case: a `blocked:` line blaming the
daemon, a timeout, or unreachability, while the run is running or fixing
AND the pipeline reports fresh activity, now reads as superseded because
the run is alive. Recency comes from the client's own `quiet` marker on
active_steps.last_activity rather than a threshold invented here, and
positive evidence is required, so a run record that outlives a genuinely
dead daemon keeps the plain reading.

stuck-crewmate-recovery gains the inverse-of-a-dead-endpoint playbook:
firstmate reads both statuses itself, steers a reattach, never restarts the
shared daemon on a crew's claim, and escalates only a refused socket.

Nothing here depends on an unshipped no-mistakes capability.

* no-mistakes(review): Prioritize daemon socket failure and narrow unreachable matching

* no-mistakes(review): Honor socket refusal across coarse status and crew guidance

* no-mistakes(test): Replace flaky settle timing assertion with pane-read count

* no-mistakes(document): Document daemon timeout recovery contract

* no-mistakes(ci): Fixed daemon socket failures being suppressed by terminal attributed runs. Positive refused/missing socket evidence now remains blocked regardless of run status. Added a behavioral regression test for terminal failed runs. Verified with fm-crew-state tests, project ShellCheck lint, and git diff checks

* fix(bin): scope the worker role contract for ship and scout launches (#3797)

* fix(brief): scope Firstmate workers to their launch contract

* no-mistakes(document): Clarify supervisor scope and worker contract ownership

* no-mistakes(ci): Removed the heading-based bypass so every ship/scout launch receives the current worker-role contract. Added a regression that failed before the fix and passes afterward. Dispatch, brief, and delivery suites, focused ShellCheck, and git diff --check all passed

* no-mistakes(review): make launch overlay sole owner of worker role contract

* no-mistakes(review): narrow heading test dimension and fix publish error wording

* no-mistakes(review): gate role supersession, fix render guard, drop AGENTS twin

* no-mistakes(document): align architecture AGENTS.md scope and spawn launch-brief header

* fix(bin): stop ringing steering doorbells into dead panes (#3823)

* fix(bin): stop ringing steering doorbells into dead panes

The steering-inbox doorbell was a plain sentence plus Enter typed into a
worker's pane, and the watcher re-rang it on the assumption that a ring is
free. In a pane whose agent has exited that line is a shell command, and the
re-ring ladder kept typing it into a shell that can never acknowledge it.

- Prefix the doorbell with the shell no-op `: ` so a bare shell executes
  nothing while a live worker still reads the same self-describing line.
  `#` is not used because interactive zsh does not treat it as a comment by
  default and the claude harness binds it to memory mode.
- fm_task_inbox_ring skips the pane (return 3) when the backend positively
  classifies the agent as dead; missing, ambiguous, unreadable, and unverified
  endpoints still ring so a blind classifier never starves a live worker.
- The watcher caps the ladder for a dead pane: one stale wake for recovery,
  no ring, no ladder walk, and the durable record stays for
  stuck-crewmate-recovery. fm-send and the remote steer leg report the skip.

Tests cover the no-op in real shells, the dead/live/unclassifiable ring
verdicts, and the single-surfacing watcher path.

* no-mistakes(review): Quote doorbell paths against shell injection

* no-mistakes(review): Reject terminal-control paths before ringing

* no-mistakes(review): Document accepted partial doorbell delivery race

* no-mistakes(review): Skip unavailable endpoints before busy-state handling

* no-mistakes(test): Respect shell startup PATH in environment allowlist test

* no-mistakes(test): Fix doorbell test fixtures for endpoint liveness

* no-mistakes(test): Prioritize confirmed restarts and clean shell test syntax

* no-mistakes(document): Document dead and missing doorbell recovery

* no-mistakes(ci): Fixed the persistence-reply timeout race by rechecking for a correlated reply immediately before falling back to a nudge. Added a deterministic regression covering replies arriving between the preliminary resolution pass and timeout handling. Verified with the targeted restart suite, project lint, coverage guard, bash syntax checks, and diff checks

* fix(bin): wait out transient primary-checkout reads in the spawn worktree poll (#3834)

* fix(spawn): keep the worktree poll from adopting the repository primary

After `treehouse get` is sent, the worktree-discovery poll reads the pane's
foreground-process cwd. While treehouse is still fetching and checking a slot
out, the foreground process is treehouse itself and it reports the repository's
PRIMARY checkout as its cwd for several seconds. The poll accepted any path
that merely differed from the spawning project, so from a linked spawning home
- whose project is itself a worktree of that repository - it adopted the
primary, and the isolation guard then refused a launch whose slot treehouse
went on to create normally.

Screen every candidate with the isolation guard's own conditions, extracted as
spawn_worktree_isolated, so a read the guard would reject stays a transient the
poll keeps waiting through. The two-consecutive-reads rule and the guard as
final backstop are unchanged; a pane that never reaches an isolated worktree
still fails at the existing 60s deadline, now naming the last path it reported.

The already-settled timing assertion counted whole-spawn wall time against a
5s budget and failed on unmodified HEAD on slower machines; it now counts pane
reads, which is what "one confirming read, not an extra cycle" actually means.

* fix(spawn): say which path the worktree wait rejected, and why

Screening every discovery-poll candidate means a host that never reaches an
isolated worktree spends the whole 60s window before refusing. That wait is
deliberate - separating a transient from a terminal misconfiguration needs
machinery this path does not want - so the refusal explains itself instead:
the isolation check records why a candidate failed, and the deadline names the
last path seen together with that reason. Message and diagnostics only; the
poll's control flow is unchanged.

Two suites asserted the guard's wording on paths the poll now rejects rather
than adopts, so their refusal arrives from the deadline instead: realign
fm-tangle-guard's non-git and subdirectory-of-primary cases (each now also
asserting the stated reason, and the second the metadata absence it was
missing) and the herdr projection e2e's forced non-worktree cwd.

* no-mistakes(review): stub poll sleep in tangle-guard spawn isolation test

* no-mistakes(document): document spawn poll isolation screen in fm-spawn header

* test(spawn): make the non-git isolation case non-git anywhere

The refusal-reason assertion for a path outside any repository assumed TMPDIR
is not inside a git repository. Where it is, git walks up from the temporary
directory, finds that repository, and the spawn reports the subdirectory cause
instead - so the case passed or failed on a property of the host rather than on
the behaviour under test.

Build the path under a directory the test then names in GIT_CEILING_DIRECTORIES,
which git documents as not chdir-ing up into a listed directory while looking
for a repository. Git never excludes the directory being searched, so the
ceiling is the parent of the path handed to the spawn.

The assertions pin which cause fired rather than the sentence that explains it,
leaving the operator wording free to improve.

* no-mistakes(document): point spawn poll comment at the isolation screen's comparison

* no-mistakes(ci): Fixed the "Behavior portable serial 2" failure in tests/fm-tangle-guard.test.sh ("non-worktree spawn did not say why the path was rejected (missing: 'not inside a git worktree')"). Root cause, in this PR's code: bin/fm-spawn.sh's spawn_worktree_isolated resolved the git toplevel with `wt_top_real=$(cd "$SPAWN_WT_TOP" ...)`. For a path in no repository, `git rev-parse --show-toplevel` yields empty, and `cd ""` is a SUCCESSFUL no-op on bash before 5.3 (CI's ubuntu-latest ships bash 5.2). The empty toplevel therefore resolved to fm-spawn's own cwd — the CI checkout — so the poll reported "it is a subdirectory of worktree root '/home/runner/work/firstmate/firstmate'" instead of the correct "it is not inside a git worktree". Dev machines with bash 5.3 fail `cd ""`, which is why the suite passed locally and only failed on CI; it is a genuine shell-portability defect in the reason vocabulary this change added, not a test-environment artifact. Fix (smallest root-cause change, 1 line + comment, bin/fm-spawn.sh:2168-2173): guard the empty value so it never reaches `cd` — if [ -n "$SPAWN_WT_TOP" ] && ! wt_top_real=$(cd "$SPAWN_WT_TOP" 2>/dev/null && pwd -P); then No change to the poll's timing or deadline behavior (respecting the recorded refusal-latency and spawn-wt-reason-vocabulary decisions), no new tests, no other files touched. Verification: - Reproduced the exact CI failure locally by putting bash 3.2 (same `cd ""` semantics as CI's 5.2) first on PATH: fails before the fix with the identical message shape, passes after. - tests/fm-tangle-guard.test.sh passes under both bash 3.2 and bash 5.3. - tests/fm-spawn-worktree-settle.test.sh and tests/fm-spawn-pool-base-freshen.test.sh pass; shellcheck -x bin/fm-spawn.sh clean. - bin/fm-test-run.sh --changed: 46 suites completed, every FM_TEST_END exit=0, 1076 passing assertions, 0 "not ok" (including fm-tangle-guard, fm-control-relaunch, fm-lint). The run ended on my own 900s wall-clock cap (rc=124), not on any test failure

* fix: support stock macOS Bash 3.2 paths (#3732)

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): read orphaned green ci monitor and daemon-down failed record as not failed (#3846)

* fix(bin): read an orphaned green ci monitor as held-for-merge, not failed

A no-mistakes run held for a captain merge decision keeps its ci step
polling until merged or closed; when the shared daemon restarts under
that poll, the run is recorded failed although every substantive step
completed and GitHub reports the PR green. A monitor whose only
remaining job is to observe a human decision must not convert the
absence of that decision into a failure verdict.

fm-crew-state.sh now reclassifies a terminal failed run as done
(held-for-merge), surfacing the run's PR URL, when the steps table
shows every step completed except exactly ci failed and the ci log's
last recognized marker reads checks green. A genuinely red check, an
unreadable ci log, or a second failed step keeps the failure.

* no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed

* no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification

* fix(spawn): verify preserved backlog state after interrupted spawn delivery (#3852)

* fix(bin): read preserved spawn state back before the interrupted exit claims it

The deferred-signal exit path asserted the paired task record and
In-flight backlog state were preserved without reading either back,
exactly when a reader is least able to check (fm-yi4j evidence,
2026-09-05). The commit's exit status alone has been observed to agree
with a row that did not actually move.

The exit path now re-reads the record and the row under the same
per-task lock as the commit, repairs a row the commit believed it
moved, and phrases the error as exactly what was verified or attempted
- verified preserved, repaired and verified, or an explicit
preservation-could-not-be-verified with the reason and hand-closeout
instruction. Two behavior tests drive a lying tasks-axi start through a
real interrupted spawn and assert the printed claim and the real
backlog state agree.

* no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask

* no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean

* docs: correct runtime-backend maturity labels for Herdr (#3821)

* docs: correct stale tmux/herdr backend maturity claims

Herdr now has 21 test files, its own required CI job (tests-herdr)
that installs a pinned build and hard-fails on "skip: herdr not
found", while tmux has 3 test files and is only required as a
dependency of the portable-serial e2e lane. zellij, orca, and cmux
still have no CI lane at all. AGENTS.md and docs/herdr-backend.md
still called Herdr merely "experimental" alongside those three,
misleading every session and reader about actual coverage.

Update AGENTS.md's config/backend entry, the opening lines of
docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend
section of docs/configuration.md, and the matching claims in
docs/architecture.md, CONTRIBUTING.md, and README.md so they agree
and distinguish tmux (default), herdr (own required CI lane, largest
suite, Windows still spike-only), and zellij/orca/cmux (still
experimental, no CI lane). No behavior, selection order, or
dispatch logic changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj

* no-mistakes(review): docs: fix stale herdr label and CI-lane wording

* no-mistakes(document): docs: align tmux adapter label in scripts.md

* no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page

* no-mistakes(review): docs: drop windows claim, align contributing backend wording

* no-mistakes(review): docs: trim duplicated CI claim from herdr opening line

* no-mistakes(review): docs: drop unguarded largest-test-suite superlative

* no-mistakes(review): docs: restore tmux verified label and README experimental scope

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: restore intent-targeted no-mistakes validation (#3865)

* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR #823 added and PR #1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: verify Treehouse slot ownership before teardown (#3837)

* fix(bin): verify pool-slot ownership before returning a worktree slot

Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.

Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.

* no-mistakes(review): Protect slots across cloned Firstmate homes

* no-mistakes(test): Gate teardown locking on genuine Treehouse slots

* no-mistakes(test): Clarify pooled descendant slot gating

* no-mistakes(test): Synchronize watcher re-arm test on process exit

* no-mistakes(test): Wait for watcher cleanup before timeout escalation

* no-mistakes(document): Document pool-slot ownership safeguards

* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion

* fix(bin): resolve relative origins from repository root

* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout

* feat(bin): add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary (#3867)

* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary

Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.

Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test: prove the omp guard continuation through a guard spy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* fix(spawn): clear the gemini marker at the omp launch boundary

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): force the guard stage by freezing the watcher and clear lint findings

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): reap the live lab by path and record omp's rpc shutdown as a note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces

Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin

* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs

* no-mistakes(review): omp: pin config-model validation with a test, trim overlay

* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family

* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes

* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux

* no-mistakes(document): docs: add omp subagent-guard row, fix live test header

* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π|󰵗|pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)

* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|󰵗)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix(pi): invoke Bash helpers correctly on native Windows (#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>

* test(bin): pin teardown outcomes for squash-merged rebased branches (#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>

* fix(bin): keep supervision armed for registered custom checks (#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes #3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner

* fix(bin): resolve Treehouse locks for remote secondmate homes (#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership

* fix(bearings): repair board listening and decision reconciliation (#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable state - refuses
rather than acting.

Regression coverage fails without each fix, and pins every leak path: a bare
reconcile, a standalone close or note with no pending request, an any-channel
reconcile, an annotated selection from a freeform card, and a
generation-skewed authorization. An opt-in guard re-proves the lavish-axi
shapes and the reopen against the installed tool.

* fix(bin): quote the done comparison in the reconcile intake

shellcheck SC1010 reads the bare word as the loop keyword. The failed run
never reached its lint step, so this shipped in the recovered content.

* no-mistakes(review): Publish reconciled parent resolution before request retirement

* no-mistakes(review): Clarify committed cleanup and reconcile reservation scope

* no-mistakes(review): Preserve remote cards and legacy answer compatibility

* no-mistakes(test): Separate live claim relea…
Hozzy02 added a commit to Hozzy02/firstmate that referenced this pull request Sep 27, 2026
* fix: support stock macOS Bash 3.2 paths (#3732)

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): read orphaned green ci monitor and daemon-down failed record as not failed (#3846)

* fix(bin): read an orphaned green ci monitor as held-for-merge, not failed

A no-mistakes run held for a captain merge decision keeps its ci step
polling until merged or closed; when the shared daemon restarts under
that poll, the run is recorded failed although every substantive step
completed and GitHub reports the PR green. A monitor whose only
remaining job is to observe a human decision must not convert the
absence of that decision into a failure verdict.

fm-crew-state.sh now reclassifies a terminal failed run as done
(held-for-merge), surfacing the run's PR URL, when the steps table
shows every step completed except exactly ci failed and the ci log's
last recognized marker reads checks green. A genuinely red check, an
unreadable ci log, or a second failed step keeps the failure.

* no-mistakes(review): Read daemon-down coarse failed ledger as unknown, not failed

* no-mistakes(document): docs: align AGENTS.md failed-verdict guidance with crew-state reclassification

* fix(spawn): verify preserved backlog state after interrupted spawn delivery (#3852)

* fix(bin): read preserved spawn state back before the interrupted exit claims it

The deferred-signal exit path asserted the paired task record and
In-flight backlog state were preserved without reading either back,
exactly when a reader is least able to check (fm-yi4j evidence,
2026-09-05). The commit's exit status alone has been observed to agree
with a row that did not actually move.

The exit path now re-reads the record and the row under the same
per-task lock as the commit, repairs a row the commit believed it
moved, and phrases the error as exactly what was verified or attempted
- verified preserved, repaired and verified, or an explicit
preservation-could-not-be-verified with the reason and hand-closeout
instruction. Two behavior tests drive a lying tasks-axi start through a
real interrupted spawn and assert the printed claim and the real
backlog state agree.

* no-mistakes(test): Fix calm suite for Pi 0.85 and pin test umask

* no-mistakes(document): Document interrupted-spawn preservation claim in backlog gate owner

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-spawn.sh's deferred-signal exit path: during preservation verification, the no-op HUP/INT/TERM re-trap combined with an unresponsive `tasks-axi show`/`start` (bash cannot run traps while a foreground child runs) held the per-task meta lock - and every lifecycle operation waiting on it - indefinitely. Root-cause fix: bound every tasks-axi invocation made under the lock. bin/fm-backlog-transition-lib.sh gains fm_tasks_axi, an exec-based wrapper (GNU timeout, gtimeout fallback) used by fm_backlog_row_show and fm_backlog_mutate that preserves the exact process placement of the plain tasks-axi call; bin/fm-spawn.sh sets FM_TASKS_AXI_TIMEOUT (default 30s) at the commit point so both the commit and the read-back verification are bounded. A timed-out call fails through the existing error plumbing and probe/mutate name the timeout as the reason, so the interrupted exit path prints honest 'preservation could not be verified ... (reason)' wording - never intent phrased as outcome, matching the author's intent. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives a real interrupted spawn through a lying tasks-axi whose repair start never answers; it asserts the spawn exits promptly (self-bounded by an outer timeout), the attempted wording names the timeout, and the printed claim agrees with the real record/backlog state. Confirmed the test fails on the unfixed tree and passes with the fix. Verified: fm-backlog-atomicity (83 ok), fm-transition-lib, fm-backlog-handoff, fm-captain-hold, fm-teardown, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-spawn-batch, fm-spawn-dispatch-profile, fm-task-delivery, fm-control-relaunch all pass; bin/fm-lint.sh clean. The fm-bootstrap 'unsplit run lost its local diagnostic' failure reproduces on the pristine base commit and is unrelated to this change

* no-mistakes(ci): Fixed the Greptile P1 in bin/fm-backlog-transition-lib.sh: fm_tasks_axi bounded tasks-axi only through GNU timeout/gtimeout and fell through to an unbounded exec on hosts with neither (stock macOS), so an unresponsive call could hold the per-task meta lock forever during interrupted-spawn verification. Root-cause fix: the bound now has no unbounded path. GNU timeout is preferred, gtimeout next, then a small perl watchdog (fork + waitpid WNOHANG polling at 50ms, TERM on expiry, one bound of grace, then KILL, exit 124 so the callers' existing timeout plumbing reports it; exit statuses and output pass through unchanged). Polling was chosen over alarm+die to avoid perl's platform-dependent syscall-restart semantics. When a bound is requested but no bounding mechanism exists, the call fails closed (exit 127 with a diagnostic) rather than running unbounded, so the interrupted exit path prints honest attempted wording, never intent as outcome. The unbounded exec remains only for the no-bound plain-call case. Added three behavior tests in tests/fm-backlog-atomicity.test.sh driving fm_tasks_axi through a PATH with no timeout binary: a hanging stub must exit 124 within the bound (verified to fail on the pre-fix code), a failing stub's status/output must pass through, and a tool-less PATH must fail closed with the diagnostic. Verified: fm-backlog-atomicity 86/86 ok, fm-transition-lib, fm-backlog-handoff, fm-teardown, fm-spawn-batch, fm-task-delivery, fm-fleet-snapshot-view, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile all pass; bin/fm-lint.sh clean

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-backlog-transition-lib.sh: fm_tasks_axi's GNU timeout and gtimeout paths sent TERM at the bound but had no kill-after, so a tasks-axi that ignores SIGTERM kept the bounded call - and the per-task meta lock - held indefinitely during interrupted-spawn verification. Root-cause fix: both GNU execs now carry -k "$bound" (TERM at the bound, KILL after one further bound of grace), giving every bounded path the same forced-termination contract the perl watchdog already had. Because GNU timeout exits 137 (128+SIGKILL) when the kill-after fires - versus 124 for a TERM expiry - the probe/mutate timeout detection now goes through a new fm_tasks_axi_timeout_expired helper that treats 124 and 137 alike, so the interrupted exit path still names the timeout as the reason; the helper keeps the bound check in one place. Added a behavior test in tests/fm-backlog-atomicity.test.sh that drives fm_tasks_axi through a real GNU timeout with a tasks-axi stub that traps and ignores TERM (the ignored disposition survives exec into sleep) and asserts a bound-expiry status plus completion within bound+grace; on the pre-fix code the suite hangs until killed, confirming the reproduction. Verified: fm-backlog-atomicity 87/87 ok, fm-transition-lib, fm-backlog-handoff, fm-spawn-batch, fm-task-delivery, fm-teardown, fm-secondmate-reconcile, fm-control-relaunch, fm-spawn-dispatch-profile, fm-captain-hold-lifecycle all pass; bin/fm-lint.sh clean

* docs: correct runtime-backend maturity labels for Herdr (#3821)

* docs: correct stale tmux/herdr backend maturity claims

Herdr now has 21 test files, its own required CI job (tests-herdr)
that installs a pinned build and hard-fails on "skip: herdr not
found", while tmux has 3 test files and is only required as a
dependency of the portable-serial e2e lane. zellij, orca, and cmux
still have no CI lane at all. AGENTS.md and docs/herdr-backend.md
still called Herdr merely "experimental" alongside those three,
misleading every session and reader about actual coverage.

Update AGENTS.md's config/backend entry, the opening lines of
docs/herdr-backend.md and docs/tmux-backend.md, the runtime-backend
section of docs/configuration.md, and the matching claims in
docs/architecture.md, CONTRIBUTING.md, and README.md so they agree
and distinguish tmux (default), herdr (own required CI lane, largest
suite, Windows still spike-only), and zellij/orca/cmux (still
experimental, no CI lane). No behavior, selection order, or
dispatch logic changes.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GDYsWPEfwTuPNjQ2nGBCcj

* no-mistakes(review): docs: fix stale herdr label and CI-lane wording

* no-mistakes(document): docs: align tmux adapter label in scripts.md

* no-mistakes(review): docs: drop duplicated herdr CI claim from tmux page

* no-mistakes(review): docs: drop windows claim, align contributing backend wording

* no-mistakes(review): docs: trim duplicated CI claim from herdr opening line

* no-mistakes(review): docs: drop unguarded largest-test-suite superlative

* no-mistakes(review): docs: restore tmux verified label and README experimental scope

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: restore intent-targeted no-mistakes validation (#3865)

* fix: delete deterministic no-mistakes test baseline, restore intent-targeted Test

PR #3644 pinned commands.test to a fm-test-run.sh --changed walk of the
repository's 75-162 tests/*.test.sh scripts. no-mistakes runs commands.test
verbatim and unconditionally after every fix round, so that walk multiplied
by round count: measured at 32.7 minutes per validation versus 3.6 minutes
intent-targeted. Delete the pin and restore the 3.6-minute posture.

Add tests/fm-nm-test-contract.test.sh as a regression guard, parsing
.no-mistakes.yaml as YAML (ruby's bundled Psych, matching the parser
tests/fm-test-run.test.sh already uses for ci.yml) rather than grepping its
text, restoring in legal form what PR #823 added and PR #1282 removed.

Record the rule in docs/configuration.md's "Gate defaults" section (the
authoritative owner CONTRIBUTING.md already points at) and strengthen
CONTRIBUTING.md's existing local-Test guidance to state it plainly: never
configure commands.test to a deterministic test command, complete or partial.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01CSt9JvrQMVc4u3jPyUFCFC

* no-mistakes(review): Centralize no-mistakes test policy and narrow guard

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>

* fix: verify Treehouse slot ownership before teardown (#3837)

* fix(bin): verify pool-slot ownership before returning a worktree slot

Workers were killed when cleanup returned a Treehouse pool slot that a
different, live task had already taken. Teardown now proves the slot is
genuinely this task's before releasing it: it refuses when another task
record claims the same live worktree path, or when the endpoint's working
directory contradicts the recorded slot, and that refusal holds under
--force. Slot allocation, metadata publication, ownership verification,
and slot return are serialized across linked firstmate homes, and forced
secondmate cleanup verifies descendant slot ownership before returning
any child worktree.

Regression coverage drives the scripts with two task records naming one
slot path and asserts the live worker survives and its slot is not reset.

* no-mistakes(review): Protect slots across cloned Firstmate homes

* no-mistakes(test): Gate teardown locking on genuine Treehouse slots

* no-mistakes(test): Clarify pooled descendant slot gating

* no-mistakes(test): Synchronize watcher re-arm test on process exit

* no-mistakes(test): Wait for watcher cleanup before timeout escalation

* no-mistakes(document): Document pool-slot ownership safeguards

* no-mistakes(ci): Fixed all reported CI issues: normalized bare local Git origins to the same Treehouse project-lock identity as absolute clone origins; resolved ShellCheck SC1091 with explicit conditional sourcing; and taught concurrent Herdr teardown coverage to retry expected Treehouse lock contention. Added behavioral regression coverage for bare/absolute origin lock identity. Verified endpoint-safety tests, watcher tests, full CI lint, and the previously failing Herdr teardown assertion

* fix(bin): resolve relative origins from repository root

* no-mistakes(ci): Fixed teardown so an exact recorded endpoint may change cwd without falsely vetoing cleanup. Removed cwd-based ownership refusal while preserving cross-home record exclusivity and project locking. Updated behavioral coverage for both foreign slot ownership refusal and moved-cwd teardown success. Endpoint-safety, backend, watcher, checkpoint, and targeted lint checks pass. Real Herdr presentation E2E progressed successfully but exceeded the 600s local timeout

* feat(bin): add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary (#3867)

* feat: add verified omp (Oh My Pi) harness adapter for crew, secondmate, and primary

Add omp as a verified harness: anchored process-name detection with a
Firstmate-owned FM_OMP_HARNESS launch marker that needs real omp ancestry,
the fm-spawn launch template with foreign-marker clearing, the tracked
.omp/fm-worker-overlay.yml posture overlay, --auto-approve, --cwd, and
pre-launch model validation scoped to providers 'omp models --json' lists.
Workers get a state-resident busy-state extension keyed on agent_end
without willContinue (omp has no agent_settled). The primary gets two
tracked .omp/extensions: a turn-end guard that answers omp's blocking
session_stop hook by compelling one continuation per turn, with the
pre-tool seatbelts and Run-tier session-start delivery, and a watcher
extension ported from the Pi one with fm_watch_arm_omp. Control tables,
composer busy footers, omp's status row as a bare-composer boundary, the
extension supervision model with an omp-keyed ownership proof, the
session-start diagnostic, and the supervision protocol snippet follow.

Verified live on omp 18.1.11 with openai-codex/gpt-6-astra: a Herdr scout
through spawn, busy state, steer, interrupt, exit, and teardown, and the
isolated rpc primary lab through extension auto-discovery, digest
delivery, lock identity, watcher arm, successor and wake delivery, and
the compelled guard continuation.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test: prove the omp guard continuation through a guard spy

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* fix(spawn): clear the gemini marker at the omp launch boundary

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): force the guard stage by freezing the watcher and clear lint findings

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): reap the live lab by path and record omp's rpc shutdown as a note

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* test(omp): spawn a real secondmate for the discovery rule and classify the omp surfaces

Replace the template-extraction check with a genuine --secondmate launch
pinned to the fake tmux backend, assert the worker extension's handler set
through the executable rather than its bytes, classify the two new omp
surfaces in the documentation inventory, and record the Herdr worker
evidence in the runtime-backends verification doc.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_012AMyYaHU42Ltotn6fauPAh

* no-mistakes(review): omp: unverify remote routes, narrow busy regex, drop overlay approval pin

* no-mistakes(review): omp: validate config-pinned model, correct remote and marker docs

* no-mistakes(review): omp: pin config-model validation with a test, trim overlay

* no-mistakes(review): omp: sync guard evidence, drop dead param, map quota family

* no-mistakes(review): omp quota: refuse unmapped prefixes, match bare model scopes

* no-mistakes(document): docs: cover omp in cd-guard, quota, continuity, tmux

* no-mistakes(document): docs: add omp subagent-guard row, fix live test header

* no-mistakes(ci): Fixed both failing behavior shards and the Greptile P1 in bin/fm-composer-lib.sh. Root cause of "Behavior portable serial 1" and "Behavior portable parallel 2": the omp busy regex (FM_DELIVERY_OMP_BUSY_REGEX_DEFAULT) and omp status-row furniture regex (FM_COMPOSER_OMP_STATUS_RE_DEFAULT) used the bracket range [⠁-⣿]; BSD grep on macOS accepts it but GNU grep on Linux CI aborts with "Invalid collation character", failing every omp busy/furniture read (3 assertions across fm-omp-harness, fm-tmux-submit-busy, fm-composer-lib). Replaced the range with one shared explicit alternation FM_OMP_SPINNER_FRAMES_RE of omp 18.1.11's unicode-preset spinner frames (status set ⣾⣽⣻⢿⡿⣟⣯⣷ + activity set ⠋⠙⠹⠸⠼⠴⠦⠧⠇⠏, read from the installed binary), the same pattern the Kimi busy regex already uses in CI. For Greptile's finding (the harness-agnostic furniture rule's first alternative matched any 1–4-byte token + ' · ', so wrapped typed input like 'fix · tests' with the cursor on it regressed from pending to unknown; reproduced locally vs base), pinned that alternative to omp's identity cell (π|󰵗|pi, the icon.omp of each preset in the 18.1.11 binary). Tests: fm-composer-lib.test.sh asserts 'fix · tests' is not furniture, a status-set spinner row is furniture, and the wrapped composer screen reads pending under both locales (CAPS_TMUX cursor 3); fm-omp-harness.test.sh asserts a status-set frame reads busy. New negative cases fail against the pre-fix lib and pass after. Verified: fm-omp-harness, fm-tmux-submit-busy pass via bin/fm-test-run.sh; fm-composer-lib passes all cases except one pre-existing, unrelated local failure (Herdr half-block test uses printf '▀', unsupported by macOS bash 3.2; fails identically on a pristine HEAD export, passes on CI bash 5); shellcheck and bin/fm-lint.sh clean. Caveat: GNU grep is unavailable locally, so the Linux compile was not run directly; the fix uses only constructs already proven on CI's GNU grep (multibyte literal alternations, incl. under LC_ALL=C). Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh, tests/fm-omp-harness.test.sh. No docs needed changes (they describe the rule generically)

* no-mistakes(ci): Greptile Review: fixed. The omp status-row furniture regex FM_COMPOSER_OMP_STATUS_RE_DEFAULT in bin/fm-composer-lib.sh still accepted a literal `pi ·` opening, so wrapped composer input beginning with `pi ·` was truncated and misclassified. Read the installed omp 18.1.11 binary: the ascii preset's `icon.omp` is `pi` but its `sep.dot` separator is ` - ` (unicode/nerd use ` · `), so a real ascii status row never contains `pi ·` and that alternative could only ever match typed text. Removal-first fix: dropped `pi` from the identity alternation (now `(π|󰵗)`) and updated the comment to record why the ascii preset is excluded. Tests (tests/fm-composer-lib.test.sh): added a negative furniture case for 'pi · e · phi as the three constants' and a wrapped-screen assertion (CAPS_TMUX, cursor 3) that a continuation row opening `pi ·` reads pending in both locales; the new case fails against the unfixed lib and passes after. Verified: composer test with the half-block case skipped passes all 33 cases including the omp matrix; bin/fm-test-run.sh tests/fm-omp-harness.test.sh passes; shellcheck -x clean on both files; bin/fm-lint.sh clean. The full composer test via the runner fails locally only on the pre-existing half-block case (bash 3.2 printf cannot emit ▀; passes on CI bash 5), identical to before this change. Docs unchanged (they describe the rule generically and never mention the ascii identity cell). PR must be raised via no-mistakes: not caused by code. attestation.head_sha is cdddc60 while the PR head is cd51cf4 because the pipeline's ci-phase push moved the head; the outer executor's re-push will re-bind the attestation. No file change for that check. Files changed: bin/fm-composer-lib.sh, tests/fm-composer-lib.test.sh

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

* fix(pi): invoke Bash helpers correctly on native Windows (#3843)

* Fix Pi shell invocation on native Windows

* no-mistakes(document): Document Pi Windows Bash transport

* no-mistakes(ci): Captain, staged a narrow fix: register the Pi Windows regression for both extension paths, make Windows mode emulation non-failing, and enforce LF shell checkouts. Mapping and coverage checks pass; CI/Require no-mistakes were approval-gated externally

* no-mistakes(review): Cover async Windows branch-outcome Bash invocation

* no-mistakes(review): Preserve Cygwin checks and refresh Windows timing

* no-mistakes(test): Invoke OpenCode operational-input owner through Bash on Windows

* validation-fixture

* no-mistakes(document): Document Windows Bash helper invocation

* no-mistakes(ci): Fixed PR-caused changed-selection failure by removing the malformed tracked evidence artifact and allowing deleted, unconsumed source paths to retire cleanly while preserving fail-closed behavior for live unmapped paths. Added regression coverage. Verified native-Windows Pi shell-seam test passes and --changed selects the Windows regression

---------

Co-authored-by: test <test@example.invalid>

* test(bin): pin teardown outcomes for squash-merged rebased branches (#3870)

* fix(bin): recognise squash-merged rebased work as landed at teardown

A pipeline rebase can leave the local worktree on pre-rebase commits while
GitHub squash-merges the rebased head. The landed-work test then compared
those stale commits against a squashed main and refused cleanup of work
that had already landed.

When the forge reports the recorded PR merged and its merge commit is on
the default branch, treat a local branch that only repeats paths from the
pipeline push as stale rather than unlanded. If the forge is unreachable,
the same coverage check runs against a PR head whose content is already
on default. Extra local paths still refuse.

* fix(bin): drop unprovable squash-rebase landed-work coverage

Path-set coverage treated a diverged local branch as landed whenever it
touched the same files as the squash merge. That accepts the reviewer's
failing sequence: same path, different content, work discarded.

git cherry and merge-tree containment were already too strict on the real
rebase-fold case. No remaining check is both safe and permissive enough
to recognise a stale pre-rebase copy without also accepting unlanded
edits, so that case still refuses.

Keep the proofs that hold: a merged PR head that contains local work, or
a clean content-in-default tree match. Tests now refuse same-path
different content and extra unlanded commits, and still allow a local
branch that followed the pipeline rebase.

* no-mistakes(review): drop recorded-pr-head fallback and reverted-design leftovers

* no-mistakes(review): silence squash-merge stdout corrupting test PR head

* no-mistakes(review): make unlanded follow-up commit sole cause of refusal

* no-mistakes(document): correct stale squash-rebase fixture comments in teardown tests

* no-mistakes(ci): Split the three reported checks: - CI (run 34061098467) and Require no-mistakes (run 34061098460) both concluded `action_required` — approval-gated workflow runs that never executed a step. Not caused by this PR's code; no change can clear them. - Greptile Review was a genuine defect in the new tests: the three new refusal cases (tests/fm-teardown.test.sh) asserted only exit status 1 and a REFUSED line, so a teardown regression that destroyed the worktree, branch, and task record before reporting refusal would still pass. Fix (tests only): added one `assert_refusal_retained_task_state` helper and called it from `test_squash_merged_same_file_different_content_refuses`, `test_squash_merged_rebased_local_with_unlanded_commit_refuses`, and `test_squash_merged_stale_local_refuses_when_forge_unreachable`, each capturing the worktree HEAD before `run_teardown`. It pins that the refusal left the isolated copy on disk, the task branch still checked out at the same unlanded commit, and state/task-x1.meta intact. Verification: the four squash tests pass; a sensitivity probe ran the ALLOW fixture (teardown completes) and pointed the same helper at the outcome — it fires, because a completed teardown detaches/deletes the branch and removes the task record, proving the assertions discriminate. Full tests/fm-teardown.test.sh: 83 passing. bin/fm-lint.sh clean with pinned ShellCheck 0.11.0 + actionlint 1.7.12 (plus an explicit --external-sources pass on the changed file). bin/fm-test-run.sh --check-coverage ok. Caveat: test_herdr_flat_teardown_preflight_refuses_before_changes (mode missing-adapter) fails on this machine. Verified it fails identically on base commit f91a950 via `git archive`, so it is a pre-existing local environment difference untouched by this diff; skipped to run the rest of the suite, not modified

---------

Co-authored-by: Morten Gad <mogad@itm8.com>

* fix(bin): keep supervision armed for registered custom checks (#3860)

* fix(bin): keep supervision armed for registered custom checks

A custom check bound by bin/fm-check-register.sh only ever runs inside the
watcher's check sweep, but fm_supervision_status counted in-flight tasks, the
relay poll shim, and process-event sources as supervision need, and not
registered checks. Tearing down the last task therefore stopped every
home-level check silently until the next spawn.

Count a state/<id>.check.sh that carries its state/<id>.check-trust binding as
supervision need. The relay shim keeps its own trust path and task PR polls
carry no such binding and are torn down with their task, so neither arms a home
by accident. Presence of the binding is the whole test: the sweep validates the
bytes at execution time and wakes firstmate when it rejects one, which is the
outcome an idle home needs.

Closes #3856

* no-mistakes(review): name registered checks in turn-end block banner and doc invariant

* no-mistakes(review): narrow PR poll predicate test to what it proves

* no-mistakes(document): point Grok re-arm step at supervision-need owner

* fix(bin): resolve Treehouse locks for remote secondmate homes (#3883)

* fix(bin): resolve the shared Treehouse project lock inside remote secondmate homes

Every spawn and teardown inside a remote-seeded secondmate home refused,
because the project lock's anchor could not be resolved there.

fm_firstmate_root_home walks a home's parent bindings upward to find the
anchor the lock lives in, and treated a remote parent binding as an error.
A remote-seeded home's parent is on another machine, so that walk can never
succeed from there - and neither can the home's own local descendants, whose
chain terminates at the same record. Both fail closed on every Treehouse-backed
spawn and every pool-slot teardown.

A remote parent now terminates the walk at the home holding it, which is the
correct anchor: a lock taken on this filesystem is neither held nor observable
across that boundary, and that home is already the top of the local tree
teardown's collect_local_firstmate_states enumerates, since that walk skips
remote registry entries for the same reason. Mutual exclusion is unchanged -
every home reachable through local parent links still derives one identical
lock file per project, and an unreadable binding, an unsupported route, an
unreachable local parent, a cycle, and an over-deep chain all still refuse.
Origin-less local-only projects keep resolving through their worktree top.

Regression coverage pins the anchor for the main-home layout, a local
secondmate, a remote-seeded home, and its local child; drives teardown
end-to-end in a remote-seeded home; keeps the cross-home slot-ownership
refusal across that boundary; and proves two homes still serialize on the
one shared lock file.

* no-mistakes(document): Clarify machine-local Treehouse lock ownership

* fix(bearings): repair board listening and decision reconciliation (#3872)

* fix(bearings): repair the board's listening, card hygiene, and reconcile path

Three defects made the fleet board go quiet and then lie about what still
needs the captain.

Never arm a poll on a session that is not live. `lavish-axi <file>` exits 0
even when it refuses to reopen a session the captain ended from the browser,
reporting `status: user-ended` with the same session id, so the build's
exit-status check accepted a dead session, printed `already-armed`, and left
the board reading "not listening". The build now proves the session is live
from a fresh authoritative listing immediately before arming - not from the
establish call's status alone, which is already stale by then - reopens once
when it finds the session ended, and refuses rather than arming when it stays
ended. A reopen also replaces the pre-reopen source generation before
reporting success, so a runner on its way out cannot be mistaken for a
listener, and a board whose source is registered but unowned gets a
replacement started before the build returns.

Let a dead generation's ownership actually move. Reclaiming a claim ran its
capture-reservation cleanup first, and that cleanup re-verifies the recorded
state-root identity, so a claim naming a pid and a process group that were
both provably gone could not be cleared: reconcile reported a start while
nothing attached, and retire refused with "cannot release source ownership".
Reservation records are keyed by claim token and every replacement claims a
fresh one, so they are hygiene, not an ownership invariant. Reclamation now
additionally requires the owning process group to be absent independently,
which keeps a reused pid whose poll child still runs from ever reading as a
gone generation. A live owner and a crashed leader whose owned group survives
are still never reclaimed.

Stop carding decisions whose subject already landed. The build drops a
decision card whose work item or PR appears in the payload's own landed rows,
and one whose task is no longer an open captain call, naming each drop on
stderr. A task whose state cannot be established is kept, because a call
wrongly hidden is worse than a card wrongly shown.

Add the reconcile choice, and make it structurally incapable of closing a
call. Every decision card carries a standard `reconcile` option, injected by
the build rather than left to the composer. The board now emits the picked
option and any freeform note as separate structured fields instead of fusing
them, so a reconcile selection is not expressible as an answer value at all -
the defect that let `reconcile - <note>` reach the intake as an ordinary
answer. The adapter routes selections from that structured field, creation of
a reconcile request is bound to a verified board source rather than the shared
keyed-answer intake, and the intake still refuses the reserved value on every
channel. Each authorization is bound to the captain-hold generation that
produced the card, so an obsolete card cannot close a later call, and both
terminal outcomes require a pending request: `reconcile close` records the
evidence under its own `reconciled` mode so it never reads as the captain's
words, and `reconcile note` leaves the call open. Anything unprovable -
an unversioned row, a missing generation, an unreadable state - refuses
rather than acting.

Regression coverage fails without each fix, and pins every leak path: a bare
reconcile, a standalone close or note with no pending request, an any-channel
reconcile, an annotated selection from a freeform card, and a
generation-skewed authorization. An opt-in guard re-proves the lavish-axi
shapes and the reopen against the installed tool.

* fix(bin): quote the done comparison in the reconcile intake

shellcheck SC1010 reads the bare word as the loop keyword. The failed run
never reached its lint step, so this shipped in the recovered content.

* no-mistakes(review): Publish reconciled parent resolution before request retirement

* no-mistakes(review): Clarify committed cleanup and reconcile reservation scope

* no-mistakes(review): Preserve remote cards and legacy answer compatibility

* no-mistakes(test): Separate live claim release from stale reclamation

* no-mistakes(test): Allow terminal self-retirement during active capture

* no-mistakes(document): Document Bearings repair contracts

* no-mistakes(ci): Stabilized the failing Herdr presentation E2E by serializing test-harness Treehouse allocator calls, preventing concurrent recovery spawns from claiming the same pool slot while preserving Herdr concurrency coverage. Verified with the full E2E suite on Herdr 0.8.2, bash syntax checks, ShellCheck, and git diff checks

* fix(bin): allow pooled spawns without a git origin (#3885)

* fix(bin): skip pooled-worktree freshness fetch when no origin is configured

An origin-less local-only project has nothing remote to be stale
against, so fm-spawn's freshen_spawn_worktree_base refused to launch
crews for it. Detect a missing origin remote and skip the fetch
freshness gate entirely; an existing-but-unreachable origin keeps
refusing as before.

* no-mistakes(review): Preserve pool safety for absent and unusable origins

* no-mistakes(review): Refuse empty origin configurations during pooled spawn

* no-mistakes(review): Detect empty origin sections across config includes

* no-mistakes(review): Honor globbed includes when detecting origin configuration

* no-mistakes(review): Document conservative conditional include handling

* no-mistakes(review): Use Git-resolved config files for origin detection

* no-mistakes(review): Document included empty-origin detection boundary

* no-mistakes(document): Document originless pooled spawn behavior

* feat(tests): run live harness guards by default when available (#3889)

* feat(tests): run live harness guards by default where the harness is installed

The 24 live-harness guards each opened with their own env check, so on the
machine that has every harness - the one the product and its validation
actually run on - all of them skipped and passed. Fourteen had never been run
by the pipeline at all.

tests/lib.sh gains fm_live_gate as the single owner of that decision: a guard
that spends no model tokens runs wherever its tools are installed, a guard that
submits prompts stays opt-in, an absent tool is a named capability skip, and a
guard's own variable or FM_LIVE forces it on (turning an absent tool into a
failure) or off. Every live guard now opens with it, which also carries the
test-suite gate-refusal bypass into the guards that never sourced the shared
helpers and were therefore refused whenever a gate agent ran them.

bin/fm-test-run.sh records what a skip means: the family's expected class is
live-capability rather than a bare env opt-in, and each gate skip's reason is
logged and written to the timing artifact, so a lane can say which tool this
host could not exercise.

Only the token-free guards flip to default-on: composer-matrix, the harness
liveness drift guard, and the Herdr version floor. cursor-primary submits three
prompts, so it stays opt-in.

Running the drift guard unasked immediately found a real defect it existed to
catch: it resolved the harness through a generic `command -v cursor`, which on
a machine that also has the Cursor editor finds the editor launcher rather than
cursor-agent. That binary exits at once, leaving a bare shell in the pane and a
liveness-drift failure no classifier change could fix. It now asks
fm_cursor_resolve_binary first, the same verified owner fm-spawn uses.

CI installs the public Pi package in the portable serial lane and fails on its
skip token, so the Pi extension tests stop passing silently against a package
that is not there. No secret is added.

Verified on macOS 26.5.2 arm64: the drift guard runs with no variable set and
classifies 8 installed harnesses alive; the Herdr version-floor guard runs by
default and checks 4 real releases; every live guard refuses together under
FM_LIVE=0.

* fix(tests): keep the composer-matrix guard opt-in

Running it unasked is red on a healthy machine for reasons no code change here
removes: a harness that has not trusted this checkout sits on its own trust
dialog, which the guard treats as an unreadable composer and correctly fails.
The opencode 1.18.29 and grok 1.0.13 composer drift it also surfaced reproduces
identically on main and is filed as separate work.

So this token-free guard stays opt-in with the reason stated in its header, and
the coding guidelines record the narrow exception: a guard whose verdict
depends on host state that installing its tools does not establish may stay
opt-in, because one that is permanently red is one the fleet learns to ignore.
The other two token-free guards keep running by default.

* no-mistakes(review): Wire bearings guard and remove composer exception policy

* no-mistakes(review): Run Pi responsiveness guard by default

* no-mistakes(review): Gate AFK Pi Herdr through authoritative family sweep

* no-mistakes(review): Sanitize live gate test environments

* no-mistakes(document): Document default-on live guard behavior

* no-mistakes(ci): Fixed CI by installing the Pi package in portable-parallel-1, where fm-pi-primary-types.test.sh runs, and enforcing its package-missing gate skip there. Verified with fm-lint.sh, workflow actionlint, coverage partition checks, lane membership, and git diff checks

* no-mistakes(ci): Fixed CI’s Pi typecheck skip enforcement by giving npm, tsc, and Pi-package capability skips a shared prefix and configuring both relevant CI lanes to fail on that prefix. Verified missing tsc emits the expected skip, missing Pi package becomes a runner failure, and actionlint, ShellCheck, and git diff checks pass

* fix(bin): refuse test runs in the primary checkout when a task marker is set (#3891)

* fix(bin): refuse the behavior suite in the repository primary checkout

A task worker's isolated worktree placement is verified exactly once, when
its task starts, and nothing re-checks it afterwards. A worker that later
changes directory into the repository's primary checkout runs its Git
commands, and this branch-switching suite, against the one checkout every
linked worktree resolves against and every landing merges into. A run that
dies mid-suite can leave that checkout on a stray branch.

bin/fm-test-run.sh now refuses that case. When FM_TASK_ID marks a task
worker and the runner resolves to the primary checkout, every executing mode
exits non-zero before selecting a suite, with one line naming the primary
path and pointing at the assigned task worktree. The predicate is the one
bin/fm-spawn.sh already uses for launch placement: the working tree's own
git dir is the repository's common git dir, which separates the primary from
every linked worktree even when their top levels differ. A run with no
FM_TASK_ID set is unchanged, and so are the inspection modes, which execute
nothing. When git resolves neither directory - a non-repository fixture, a
detached copy - nothing proves this is the primary, so the run proceeds.

bin/fm-spawn.sh sets the marker: ship and scout launches export FM_TASK_ID
into the pane shell on the same pre-launch channel as GOTMPDIR, and the name
joins the sanitized launch environment allowlist so an isolated launch keeps
it.

* no-mistakes(review): clear inherited task marker in test lib; name resolved ROOT

* no-mistakes(document): docs: record FM_TASK_ID marker and runner placement refusal

---------

Co-authored-by: Talon Stark <talonstark@gmail.com>

* fix(bin): bound stale alarms for backlog captain holds (#3842)

* fix(bin): bound a stale alarm with the backlog hold, not only the status line

A legitimate wait has two records and the stale alarm reads only one.
`status_is_paused_or_captain_held` takes a status line, so it sees a wait the
worker declared. It cannot see the wait firstmate records when it hands work to
the captain: `bin/fm-captain-hold.sh hold` writes that into the backlog and
leaves the status log alone, so a delivered task keeps `done: PR ...` as its
last line for the whole time the captain is deciding.

Both stale branches were blind to it, and each churned a new pane hash back into
its own alarm: a `done:` line is captain-relevant and reaches the terminal-stale
branch, while a held task whose last line is `working:` reaches
`surface_nonterminal_stale` and fails its declared-wait test.

Consult that second record where the watcher is about to alarm, through
`bin/fm-captain-hold.sh open`, which already owns the predicate's semantics, and
bound the alarm on the shared `.paused-resurfaced-<key>` marker and
`PAUSE_RESURFACE_SECS` window the declared-wait absorb already uses. The first
sight still alarms, the window's end alarms once more, and a held crew that goes
genuinely silent still escalates through the wedge timer.

Only an established open captain call bounds anything: an unreadable backlog, an
absent or incompatible tasks-axi, a row this home does not carry, and every task
with no hold keep alarming exactly as before. The backlog hold is deliberately
not recorded as a declared pause, because the loop-top reconciliation and
`pause_state_class` both read the status line and would clear a flag that line
does not support.

Extends the fix in #3443, which closed the forms of this loop that the status
line itself can express.

* fix(bin): identify the captain call a stale alarm is bounded by

Three gaps in the bound added by the previous commit, all in how the throttle is
scoped and where the backlog is consulted.

The scope carried only the status-log signature. A task can be held, answered
with `--release`, and re-held as a genuinely different captain call without any
status append, so the second call inherited the first one's marker and its first
sight was absorbed - the one thing this bound must never do. The task id is not
the call: `bin/fm-captain-hold.sh open` gains `--identity`, which reports the
call's own lifecycle - its hold-set stamp and the number of recorded answers -
on an exit 0 and only then, leaving the silent predicate every existing caller
reads unchanged. The throttle scope now carries that identity.

The terminal path recorded the throttle before publishing the durable wake. A
failed append exits the watcher with nothing queued, and the next sighting then
read that fresh marker and absorbed the retry, turning a delayed alarm into a
lost one. Recording moves behind the append, as the non-terminal path already
had it, and the comment claiming the marker could not outlive its wake is gone
because it was false.

The backlog was consulted only on a new terminal pane hash. A captain call can
open after a hash was absorbed as provably working, changing neither the pane nor
the status log, so nothing re-read the backlog and the wedge timer kept firing
possible-wedge alarms through a legitimate wait. That timer now consults the call
at its own alarm boundary and takes the same bounded cadence - and only at that
boundary, so an ordinary repeat poll under the bound stays the local-only read it
was.

Regression coverage for each, all driving churn through one watcher process
rather than relaunching per pane change: relaunch cost dominated the earlier
shape, and an absorbing watcher stays in its poll loop across churn in production
anyway. An unheld task still alarms on every new hash, and an elapsed wedge timer
with no open captain call still escalates as a possible wedge.

* fix(review): Compose stale throttles with captain-call lifecycle identity

* fix(review): Preserve bounded same-hash captain-call resurfacing

* revert(bin): narrow the captain-hold stale bound to its observed defect

Lifts the lifecycle-identity and cadence-ownership work back out, leaving the
change at the shape that matches the defect actually observed: the stale alarm
did not consult the backlog captain hold, on either stale branch.

Reviewing the wider version surfaced a series of adjacent gaps in the watcher's
alarm state machine - a call opening after the first alarm, marker invalidation
at the hold lifecycle boundary, and which deadline a terminal timer represents.
They are real, but fixing them turns a small extension into a state-machine
change to the alarm path, which is a different review on a subsystem that is
being actively reworked. They are named as known limitations rather than carried
here, and none of them is load-bearing for what remains: the bound does strictly
less than the reverted version, leaves the wedge path escalating on
STALE_ESCALATE_SECS exactly as before, and introduces no silence that the
existing terminal-alarm path did not already have.

Kept from the reverted work is the record-after-append ordering, because that is
a defect in the code being shipped rather than an adjacent one: recording the
cadence marker before publishing the durable wake let a failed append lose an
alarm outright instead of delaying it.

History is preserved: the earlier commits stay on the branch and this removal
sits on top of them.

* fix(review): Document secondmate captain-hold scope boundary

* fix(document): Document captain-hold stale alarm scope

* fix(bin): bind the stale throttle to the captain call, not the status log

The throttle this change introduces was scoped to the task's status-log
signature. Answering a call with `--release` and holding the task again creates a
genuinely different captain call without necessarily appending to that log, so
the second call inherited the first one's marker and its first sight was
absorbed.

That is the one alarm this bound must never swallow. A delivery announced twice
is noise; a decision waiting on the captain that is never surfaced is invisible,
because nobody asks for what they do not know to ask for.

Measured rather than assumed, on the same fixture - a delivered task held for the
captain, released, and re-held with no status append, driven through bin/fm-watch.sh:

  base c499f84    call-1 first=ALARM  call-1 churn=ALARM     new call first sight=ALARM
  before this fix call-1 first=ALARM  call-1 churn=absorbed  new call first sight=absorbed
  after           call-1 first=ALARM  call-1 churn=absorbed  new call first sight=ALARM

Base never suppresses the new call, so the suppression came from this change and
closing it completes the fix rather than widening it.

`bin/fm-captain-hold.sh open` gains `--identity`, printing the call's lifecycle -
its hold-set stamp and count of recorded answers - on an exit 0 and only then, so
the silent predicate bin/fm-teardown.sh reads is untouched. The throttle scope
carries that identity beside the status signature.

The sibling case was measured too and is NOT included: on the status-declared
path, where the last line is `captain-held:`, base already absorbs a re-held
call's first sight. That behaviour predates this change and stays documented as a
known limitation rather than repaired here.

* fix(document): Document captain-call throttle lifecycle scope

* fix(ci): isolate the Herdr restart fixtures from a claimed worktree

The Herdr behaviour test intermittently reused a local worktree still claimed
by an earlier fixture after a restart.
The restart scenarios now use an isolated Treehouse project.

The full Herdr test passes on Herdr 0.8.2; bash -n and git diff --check pass as
well.

* feat(pi): resolve extension-registered providers in the supervision branch (#3871)

* Let the supervision branch resolve extension-registered providers

The isolated branch ModelRuntime cannot see providers an extension
registered into main's runtime at run time, so a pin on pi-devin-auth's
devin/swe-1-7 (or an unpinned branch following a main session on devin)
failed with "unavailable to the isolated branch runtime".

Capture main's ModelRegistry alongside mainModel and copy each
extension-registered provider config into the branch runtime at
model-resolution time. The config carries the provider's own streamSimple
and oauth wiring by reference, so the custom gRPC transport reaches the
branch unchanged instead of being reimplemented. The /supervision-model
picker uses the same copy so those models are offered.

Update configuration.md and pi-supervision-branch.md, which previously
stated extension-registered providers were not offered.

* no-mistakes(document): docs: own devin provider carve-out in branch architecture doc

* no-mistakes(ci): Fixed the Greptile P1 finding: the /supervision-model picker copied extension-registered providers into the branch ModelRuntime but checked hasConfiguredAuth without refreshing them, so providers with provisional post-registration auth were omitted from the picker while the pin-resolution path (which did refresh) accepted them. Root-cause fix in .pi/extensions/fm-branch-supervision.ts: moved the `refresh({ providers, allowNetwork: false })` call into `copyExtensionProviders` (now async, refreshing every provider it copied) and removed the duplicate per-provider refresh from `resolveBranchModel`. Both the picker and the resolution path now share one copy-and-refresh step, so hasConfiguredAuth is real in both. Regression coverage in tests/fm-pi-branch-extension.test.sh: the stubbed ModelRuntime now mirrors the real runtime by leaving a registered provider's auth pending until `refresh()` runs for it. With that stub, the existing extension-registered-provider case fails against the pre-fix extension (picker offers only anthropic/main-model) and passes with the fix. Verification: tests/fm-pi-branch-extension.test.sh passes (42 ok, no failures); tests/fm-branch-supervision.test.sh passes; tests/fm-pi-primary-types.test.sh skips locally because tsc is not installed (the refresh signature reused is the one the existing code already called). Intent constraints preserved: isolation flags untouched, carve-out still scoped to provider registration, graceful fallthrough when no providers are registered

* ci: retrigger flaky Herdr/serial-1 lanes

* no-mistakes(document): docs already cover branch extension-provider copy

* fix(bin): gate secondmate wake-loop stall alerts on real queue no-progress (#3943)

* fix(watch): detect stalled secondmate queue progress

* no-mistakes(review): gate secondmate stall on active turns and progress episodes

* no-mistakes(document): align secondmate wake-stall docs with progress-episode detector

* no-mistakes(document): clarify active-turn gate in wake-stall config docs

* no-mistakes(ci): Fixed the one real defect behind the failing checks. ROOT CAUSE (Greptile P1, real code defect in this PR): `secondmate_wake_stall_tick` in bin/fm-watch.sh reset the no-progress timer only when the oldest actionable queue sequence INCREASED (`[ "$seq" -gt "$observed_seq" ]`). When a secondmate is retired and reprovisioned under the same task ID, its fresh home's queue sequence restarts BELOW the recorded position, so the comparison is false, no reset happens, and the new queue inherits the retired generation's already-expired idle interval — emitting a false `secondmate wake-loop stalled` on its very first observation. That is precisely the false-alarm class the user intent requires this PR to remove. FIX (smallest, removal-first): bin/fm-watch.sh:749 now resets when the drain position MOVES AT ALL (`-ne` instead of `-gt`). Draining moves it up, reprovisioning moves it down; neither is a continued no-progress episode. The asymmetric `-gt` branch is removed rather than special-cased or hardened. Updated the function header comment plus the two doc sentences in docs/architecture.md and docs/configuration.md that stated the old advance-only semantics. REGRESSION TEST: added `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` to tests/fm-wake-queue.test.sh (registered in the invocation list). It drives the real watcher through retired generation (seq 9) -> reprovision (seq 3, later clock) -> freeze, asserting observable wake-queue output, no source-text inspection. VERIFICATION: - Fails before / passes after: with the fix reverted the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted`; with the fix it passes, and its third leg confirms the restarted generation still escalates on a genuine freeze (row=3 idle=2s), so the fix does not merely mute the alarm. - `bash tests/fm-wake-queue.test.sh`: exit 0, 38/38 pass, all five secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs (34152740610, 34152740604) both ended with conclusion `action_required` — workflow approval pending, not a test/build failure. Separately, `bin/fm-test-run.sh --check-coverage` exits 1 in this environment, but I confirmed by stashing my changes that it fails identically on the unmodified base tree (locale-related `comm: input is not in sorted order`); it is pre-existing and this change adds no new test file for the partition to account for. Changes are left uncommitted in the worktree

* no-mistakes(ci): Fixed the one real code defect behind the failing checks. ROOT CAUSE (Greptile P1, second round, on the head commit e1304e6): `secondmate_wake_stall_tick` in bin/fm-watch.sh identified the queue's drain position by the sequence number ALONE. The previous round changed the comparison from `-gt` to `-ne`, which handles a reprovisioned queue that restarts BELOW the recorded position, but not one that restarts ON it. A mate retired and reprovisioned under the same task id gets a fresh home whose wake-queue sequence counter restarts at 1 — and the retained parent progress marker very plausibly holds a low sequence too (a queue frozen on its first row records seq 1). Equal sequence ⇒ no reset ⇒ the brand-new queue inherits the retired generation's long-expired idle interval and emits a false `secondmate wake-loop stalled` on its very first observation. That is exactly the false-alarm class this PR exists to remove. FIX (smallest, removal-first): the file already defines the identity of a queue row once, as `row_key="$epoch-$seq"` (used for stall receipts, the stall marker, and the notify key). The progress marker's separate, weaker seq-only identity is removed: `row_key` is now computed once right after the row is parsed, stored in the progress marker, and compared with `!=`. Across generations the epoch differs (the new generation's rows are appended later), so no sequence collision can carry a stale interval; within a generation the key is stable exactly while the position does not move. bin/fm-wake-lib.sh's `fm_wake_secondmate_progress_marker_write` now takes `<oldest-row-key>` and validates it the same way the two neighbouring row-key writers do. Updated the function header comment and the two doc sentences (docs/architecture.md, docs/configuration.md) that described the old sequence-only semantics. REGRESSION TEST: `test_secondmate_reprovisioned_queue_starts_a_fresh_interval` in tests/fm-wake-queue.test.sh now drives the reported case — the reprovisioned generation restarts on the SAME sequence 9 (epoch 200) that the retired generation recorded (epoch 100), at a later clock — and asserts observable watcher output only. Its third leg still confirms the restarted generation escalates on a genuine freeze (row=9 idle=2s), so the fix does not merely mute the alarm. Three seeded progress markers in the symlink, crash-window and prefix-receipt tests were updated to the epoch-sequence form. VERIFICATION: - Fails before / passes after: with bin/fm-watch.sh and bin/fm-wake-lib.sh reverted to HEAD and the new test in place, the suite aborts on `not ok - a reprovisioned queue generation inherited the retired generation's idle interval and alerted` (exit 1); with the fix, `bash tests/fm-wake-queue.test.sh` exits 0 with 38/38 pass, all six secondmate cases green. - `bin/fm-lint.sh`: clean (ShellCheck 0.11.0, actionlint 1.7.12). - I also started `bin/fm-test-run.sh tests/fm-watch-checkpoint.test.sh tests/fm-watch-triage.test.sh tests/fm-watch-recovery-loop.test.sh` as a blast-radius check; it was still running when this phase had to return, so its result is not included. No other suite references the stall detector or the progress marker (grep over tests/ for `wake-loop stall|SECONDMATE_WAKE_STALL|secondmate-wake-progress` matches only fm-wake-queue.test.sh), and the changed lib function has exactly one caller. CHECKS NOT CAUSED BY THE CODE: the CI and "Require no-mistakes" runs on the head commit (34154091191, 34154091236, 34154091945) all ended with conclusion `action_required` — pending workflow approval, not a test/build failure. `bin/fm-test-run.sh --check-coverage` still exits 1 in this environment for the pre-existing locale reason recorded in the previous phase (`comm: input is not in sorted order` on the unmodified base tree); this change adds no new test file. Changes are left uncommitted in the worktree: bin/fm-watch.sh, bin/fm-wake-lib.sh, docs/architecture.md, docs/configuration.md, tests/fm-wake-queue.test.sh

---------

Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>

* test(calm): harden the export-DOM render step and record Pi 0.85.1 evidence (#3952)

* fix(tests): make the Calm export-DOM render step retry and report

The Calm suite's rendered-export-DOM assertion started breaking CI with a
bare "could not render calm-mode HTML export DOM", which read like a Pi
0.85 rendering change. It is not one. Calm's rendered rows are identical
across Pi 0.84.4, 0.85.0, and 0.85.1, and the CI break appeared in exactly
one of the thirteen most recent runs, all on the same Pi 0.85.1, with the
main runs immediately before and after it passing.

What actually failed is headless Chrome's start-up. The render step made a
single unattended attempt and discarded both Chrome's stderr and its exit
status, so the log held nothing to tell a Chrome crash apart from a real
change in Pi's export shape.

Rendering is a vendor-tool step; the DOM assertions that follow it are what
protect the Calm conversation boundary. So the step now retries a bounded
number of Chrome start-ups on a fresh profile, drops Chrome's background
network and /dev/shm dependencies without changing what a local file renders
to, and, when every attempt fails, reports the Chrome binary, its version,
the installed Pi version, each attempt's exit status, and Chrome's own
stderr. test_export_dom_render_guard pins that with real processes and no
browser: one clean render, one that only succeeds after a start-up failure,
and one that never renders and must report enough to diagnose itself.

The verification record adds the 0.85.1 evidence this contract is now
pinned to, the cross-version comparison run through isolated installs, and
the Pi 0.85.0 packaging gap - its dist/experimental/server.js statically
imports @earendil-works/pi-server, which 0.85.0 does not declare - that
made the contract look version-sensitive in the first place.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Y7rr2DHf51MjMvy7sMRauo

* no-mistakes(review): docs: attribute Pi 0.85 calm contract adaptation to renderer change

* no-mistakes(review): tests: drop inert chrome flags, report render timeouts

* no-mistakes(document): docs: fix stale Pi version facts and doc-lint link

---------

Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix(bin): make every counted wake queue row presentable or retired (#3950)

* fix(bin): make every counted wake queue row presentable or retired

A wake row could be counted as queued while no drain would ever present
it, leaving the operator told to "drain them before anything else" by a
command that printed nothing and offered no acknowledgement.

Two independent paths produced that state.
A row reserved by a live supervision-branch grant is excluded from a main
drain by design, but fm-guard.sh counted the whole queue, so main was
warned about rows only the branch could present, on every guarded command
for as long as the grant was held.
A row that lost its five appended fields or its numeric sequence can never
be claimed, presented, or named by an --ack-through cutoff, yet it still
counted as queued, wedging the queue permanently.

The guard now counts only the rows the calling actor can itself present or
retire, and a main drain retires unusable rows under the queue lock,
reporting them in bounded escaped form before removal so the evidence
survives for the separate row-generating defects. A retirement failure is
reported loudly and never suppresses unrelated consumable work. A main
drain whose remaining rows are all branch-held says so in one bounded line
instead of exiting silently. Grant row-list and owner-record reads move
into fm-wake-lib.sh so the drain, the grant publisher, and the guard share
one implementation.

Ownership is unchanged: a branch drain still touches nothing outside its
grant and never retires a row, and main still cannot present or acknowledge
an active grant.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GghGvsa4JDB1E5FuznX2i1

* no-mistakes(test): keep SIGTERM-safe arithmetic in wake queue retirement pass

* no-mistakes(review): add guard advisory for branch-held wake rows

* no-mistakes(document): document per-actor wake counting and unusable-row retirement

* no-mistakes(ci): Addressed the Greptile P1 on bin/fm-wake-lib.sh:1840 ("Unreadable queue suppresses alarms"). Root cause: fm_wake_actor_pending_count inferred "the queue could not be counted" only from awk's printed output (`case "$count" in ''|*[!0-9]*) count=1`). That relies on awk aborting before its END rule when the input cannot be opened. An awk that reaches END after a failed open prints `0`, which the fallback accepts as a genuine count; both actor counts then read zero and bin/fm-guard.sh emits neither the queued-wake warning nor the branch-held advisory for a queue nobody proved empty. Fix (bin/fm-wake-lib.sh:1828,1835): both counting awk invocations now set `count=''` on a non-zero awk exit status, so the existing "cannot be counted => report a pending row" fallback is driven by awk's exit status instead of an implementation-defined detail of what it printed. No new code path or behavior; the pre-existing fallback just becomes unconditional. Comment updated to state why. Regression test (tests/fm-wake-queue.test.sh: test_uncountable_queue_still_raises_the_pending_alarm, registered in the run list): runs the real bin/fm-guard.sh against a non-empty, unreadable queue with a PATH-injected awk emulating an END-running implementation (prints 0, exits 2; execs the real awk otherwise) and asserts "queued wakes pending" is still emitted; disconfirming half asserts the same fake awk over a readable, provably empty queue stays silent. Fails on the pre-fix library ("not ok - a queue that could not be counted silenced the queued-wake alarm"), passes after. Verified locally: bin/fm-test-run.sh tests/fm-wake-queue.test.sh -> 0 failed; tests/fm-guard-stale-banner.test.sh + tests/fm-watcher-lock.test.sh -> 0 failed; bin/fm-lint.sh (ShellCheck 0.11.0 + actionlint 1.7.12) clean. Caveat reported honestly: on the awks available/known here (mawk locally, plus gawk and BWK/macOS awk, all of which treat an unopenable input as fatal and skip END) the alarm was not actually suppressed - I reproduced the unreadable-queue case and the warning fired. The change removes the code's dependence on that awk detail rather than repairing an outage observed on this platform. Both intent constraints still hold: every counted row remains presentable or retirable, and no alarm-suppressing path was added

---------

Co-authored-by: Alex William <awilliam@v2202608403614505120.powersrv.de>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>

* fix(bin): use system stat for Darwin BSD formats (#3305)

* fix(bin): use /usr/bin/stat on Darwin to survive GNU stat shadowing

* fix(bin): extend /usr/bin/stat prefix to Darwin stat -f sites added on main

* test(bin): make fm-stat-shadowing skip visible on non-Darwin and isolate fm-watch state

* ci: re-trigger after Chrome headless timeout in calm HTML export test

* ci: re-trigger serial-5 after second Chrome headless timeout in calm HTML export test

* test(bin): skip PATH-based stat fault injection on Darwin where stat is /usr/bin/stat

* no-mistakes(document): Refresh stat and shard docs

* fix(spawn): carry attribution-off policy in every claude launch (#3945)

The captain's attribution policy (no Co-Authored-By trailer, no
Claude-Session link, no generated-with line) lives in Claude Code's `user`
settings scope. A spawned worker's settings sources are not guaranteed to
load that scope, so a launched worker could write attribution trailers into
its commits and PR bodies regardless of the captain's own configuration.

launch_template()'s claude case now carries the same policy
("attribution": {"commit": "…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant