Skip to content

feat(bin): sync the fork with upstream main and preserve the local features - #3

Merged
LeonidShamis merged 74 commits into
mainfrom
fm/firstmate-upstream-sync
Sep 6, 2026
Merged

LeonidShamis merged 74 commits into
mainfrom
fm/firstmate-upstream-sync

Conversation

@LeonidShamis

Copy link
Copy Markdown
Owner

Intent

Sync this fork's main with upstream by merging upstream's 67 new commits in (f66be0f..af8c4b6), resolving the conflicts, and keeping every local change working.

MERGE SHAPE IS THE POINT OF THE TASK. Use a real merge: do NOT rebase the local commits onto upstream and do NOT squash upstream's history away. Git must end up knowing upstream/main is an ancestor, or every future sync re-conflicts on the same content. The branch therefore carries a real merge commit with two parents (df933e8 + af8c4b6). If the pipeline's rebase step cannot carry that merge commit, that is a blocker to report with exact evidence, NOT something to fix by flattening the merge.

The three local commits are unique work; none exists upstream and none is superseded: c03bfbe (adopt beads backlog storage through tasks-axi), d2ccd7b (launch remote-less projects from their local default branch, PR #1), df933e8 (refuse teardown of a worktree reassigned to another task, PR #2). Never force-push, never rewrite main, never discard any of the three local commits, and never use a "discard commits and match upstream" style reset.

CONFLICT RESOLUTION RULE. Upstream is authoritative for its own 67 commits of evolution; the local side is authoritative for exactly the three features above. Where they touch the same lines, re-express the LOCAL feature on top of the UPSTREAM structure rather than reverting upstream's refactors. If upstream independently implemented the same guarantee, prefer upstream's implementation and delete the local duplicate, but say so explicitly and prove the guarantee still holds with the local regression tests.

Eight files conflicted and were resolved: bin/fm-spawn.sh, bin/fm-tasks-axi-lib.sh, bin/fm-teardown.sh, docs/architecture.md, docs/configuration.md, tests/fm-control-relaunch.test.sh, tests/fm-spawn-pool-base-freshen.test.sh, tests/fm-teardown.test.sh.

ONE LOCAL DUPLICATE WAS DELIBERATELY DELETED. Upstream's 353a8f0 (kunchenguid#3582) added fm_tasks_axi_backend_from_toml and fm_tasks_axi_backend, a strict superset of the fork's fm_tasks_axi_storage_backend: TOML-section aware (required, because .tasks.toml carries a [beads] section), honours a TASKS_AXI_BACKEND override, and falls back to ~/.tasks-axi/config.toml, mirroring the CLI's own precedence. The local duplicate was deleted and its five callers repointed (bin/fm-bootstrap.sh x2, bin/fm-backlog-handoff.sh x2, bin/fm-backlog-receive.sh x1). This is intentional, not an accidental deletion. Every TASKS_AXI_BACKEND=markdown pin in those scripts is a single-command env VAR=... cmd prefix, verified so it cannot leak into the gate that now honours that variable. The genuinely local fm_tasks_axi_has_beads_backend probe was KEPT and documented in the script header.

LOCAL BEHAVIOR THAT MUST STILL HOLD (each has regression tests in the tree; keep them passing and keep them meaningful):

  • bin/fm-spawn.sh remote-less base: a pooled worktree whose repository has NO remotes freshens against the local default branch (refs/heads/), resolved through the same default_branch() helper the other sync paths use, then the configured init.defaultBranch when that local branch exists. The currently checked-out branch is never taken on its own. An unresolvable default still refuses with the stale-base wording. A repository that has remotes, even none named origin, keeps the unchanged origin-fetch path and its exact messages.
  • bin/fm-spawn.sh relaunch and marker: spawn writes the .fm-task-owner marker; --relaunch refuses when recorded metadata carries worktree_returned=1, telling the operator the worktree was already returned and teardown must be rerun; the key is dropped only on a fresh non-relaunch spawn that genuinely acquires a new worktree.
  • bin/fm-teardown.sh ownership and convergence: teardown proves the recorded worktree is still this task's before reaping processes or returning it, refuses loudly otherwise without forcing or discarding, records worktree_returned so a rerun after a partial failure skips the worktree steps instead of repeating them, and in the Orca branch writes the returned record only after the removal actually succeeded. Upstream's 8f7b79c moved the run-conclusion and process-reap block after the backlog transition; that move was taken and the local WORKTREE_STEPS guard re-applied at the block's new location.
  • bin/fm-tasks-axi-lib.sh beads backend: the storage backend is read from the home's configuration rather than hardcoded, and fm_tasks_axi_has_beads_backend probes the installed tasks-axi build without needing bd or a database. Tracked .tasks.toml keeps backend = "beads" with the [beads] dir = "data" section. This home's live backlog runs on beads in Dolt server mode, so a regression here breaks the captain's backlog.
  • Docs: keep BOTH sides' content in docs/architecture.md and docs/configuration.md; those conflicts were adjacent additions, not disagreements.

TEST CHANGES ARE DELIBERATE AND ARE NOT WEAKENING. Upstream added suites that drive the real markdown backlog flow with the installed tasks-axi CLI. They read THIS repository's tracked .tasks.toml, which selects beads storage on this fork, so they addressed a beads store the fixture never created and lost the behavior they were written to prove. Four suites were fixed by pinning markdown storage where the suite's subject is not backend selection:

  • tests/fm-teardown.test.sh: case homes pinned with fm_test_markdown_tasks_toml, and the suite's own tasks-axi calls run from that home. 73 passing (clean upstream baseline 70; the 3 extra are the local ownership tests).
  • tests/fm-backlog-atomicity.test.sh: test_dispatch_moves_the_item_in_flight_in_the_same_run copies the tracked config precisely to prove the MARKDOWN addressing path, so it now seeds that config pinned to markdown. The suite also runs from its own temp root so its ~35 explicit --file helper calls stop inheriting a backend from the working directory. Deliberately NOT done by exporting TASKS_AXI_BACKEND: the suite's own header comment says that would outrank each case's .tasks.toml fixture and break its two beads cases.
  • tests/fm-captain-hold-lifecycle.test.sh and tests/fm-public-followup.test.sh: homes seeded by the real bin/fm-home-seed.sh inherit the tracked config because seeding clones this repository, while the beads database under gitignored data/ is not cloned. Those seeded homes are pinned the same way. pin_seeded_home owns the reason once in the public-followup suite.
  • tests/lib.sh: fm_test_markdown_tasks_toml's comment named four suites and had already rotted; it now states the rule, names the seeded-home case, and says which suites must NOT use the helper (those whose subject is backend selection and which write their own fixture).

VERIFICATION ALREADY PERFORMED on the final tree: bin/fm-lint.sh exit 0 (pinned ShellCheck 0.11.0, pinned actionlint 1.7.12); bin/fm-doc-audience-check.sh exit 0; bin/fm-test-run.sh --all = 180 suites, 19 failed, 27 gate-skipped; tasks-axi list against this home's beads database exit 0 (read-only, backlog untouched).

KNOWN-ENVIRONMENT FAILURES, NOT CAUSED BY THIS CHANGE AND NOT TO BE "FIXED". Each was proved against a clean git worktree add <tmp> upstream/main baseline with matching pass counts AND matching first failure: fm-bearings-board (2), fm-bearings-board-render (0), fm-bootstrap-network-parallel (0), fm-control-relaunch (3), fm-extension-binding (2), fm-procevent (0), fm-procevent-when (0), fm-procevent-quota (9), fm-remote-reply (0), fm-remote-secondmate-lifecycle-e2e (11), fm-remote-secondmate-parent-binding (1), fm-remote-secondmate-trace-context (0), fm-test-run (27), fm-tmux-agent-liveness (0), fm-watch-triage (79). Host causes: lsof absent (teardown leaked-process-reap), ruby absent, umask 0002 failing the process-event private-directory check (the whole procevent family plus bearings board render), no remote host configured (remote-* suites), and one 2-second timing bound (control-relaunch). fm-public-followup exits 1 on a fixture cleanup "rm: Directory not empty" with ZERO assertion failures, identically on baseline. Never weaken a guard or a test to make any of these pass.

OUT OF SCOPE, REPORTED NOT FIXED: bin/fm-home-seed.sh seeds a secondmate home by cloning this repository, so a seeded home inherits beads storage with no database, and upstream's af8c4b6 captain-hold read now makes that home's teardown refuse. Combined with c03bfbe already refusing secondmate backlog handoff on beads homes, a beads primary cannot currently stand up or hand work to a secondmate. That is fork configuration behavior for the captain to decide, deliberately not changed inside a sync task.

Do not touch data/, state/, config/, projects/, or .no-mistakes/ in any firstmate home. This change is firstmate's own shared tracked material, so firstmate-coding-guidelines applies: one sentence per line in tracked Markdown, plain dashes never em dashes, bin/*.sh must pass bin/fm-lint.sh, tests colocated in tests/ named .test.sh following tests/lib.sh conventions, tests must exercise behavior through the executable and never assert on implementation source bytes, each script header is the single owner of that script's contract, one owner per contract with cross-references rather than restatements, and never add an agent name as a commit co-author.

What Changed

  • Merged upstream's 67 commits (f66be0f..af8c4b6) into the fork as a real two-parent merge (df933e8 + af8c4b6), resolving eight conflicted files (bin/fm-spawn.sh, bin/fm-tasks-axi-lib.sh, bin/fm-teardown.sh, docs/architecture.md, docs/configuration.md, tests/fm-control-relaunch.test.sh, tests/fm-spawn-pool-base-freshen.test.sh, tests/fm-teardown.test.sh) by re-expressing the three local features on top of upstream's structure, including re-applying the teardown WORKTREE_STEPS guard at the new position upstream's 8f7b79c moved the run-conclusion and process-reap block to, and keeping both sides' additions in the two conflicted docs.
  • Deleted the local fm_tasks_axi_storage_backend in favour of upstream's TOML-section-aware fm_tasks_axi_backend / fm_tasks_axi_backend_from_toml and repointed its five callers in bin/fm-bootstrap.sh, bin/fm-backlog-handoff.sh, and bin/fm-backlog-receive.sh; kept the local fm_tasks_axi_has_beads_backend probe and documented it in the script header, and added tests/fm-backlog-beads-storage-guard.test.sh pinning the beads refusal of all three secondmate backlog movers against an exported and an empty TASKS_AXI_BACKEND.
  • Pinned markdown backlog storage in the upstream suites whose subject is not backend selection (fm-teardown, fm-backlog-atomicity, fm-captain-hold-lifecycle, fm-public-followup, fm-control-relaunch), so they stop addressing a beads store their fixtures never create; rewrote fm_test_markdown_tasks_toml's comment in tests/lib.sh to state the rule and the seeded-home case instead of a rotted suite list, and documented the beads storage model plus the seeded-secondmate-home limitation in docs/configuration.md.

Risk Assessment

⚠️ Medium: The fix round is small, correct, and empirically verified against the real bd and tasks-axi binaries, but it lands on top of a 67-commit upstream merge with hand-resolved conflicts across eight files touching backlog storage gates, spawn base resolution, and teardown convergence, whose full behavior cannot be confirmed without the pipeline's test step.

Testing

I exercised the round-2 review fixes through the real executables rather than only through unit assertions. For the beads storage gates I captured an operator transcript on a home whose own .tasks.toml declares beads while the shell exports TASKS_AXI_BACKEND=markdown: before the fix the handoff succeeded and moved the work item out of the regenerated beads mirror into the secondmate, after the fix it refuses with the storage message and both backlogs are intact; the new tests/fm-backlog-beads-storage-guard.test.sh reproduces that failure when the guards are reverted and passes on the committed tree, covering all three movers and the empty-variable case. For the tracked [beads] path I drove bin/fm-captain-hold.sh verify against a real bd 1.2.2 graph on a home carrying the tracked .tasks.toml verbatim: before the fix the migrated-hold scan refused with a configuration complaint and could never run on this fork, after it resolves the graph and reports the unknown hold, and deleting the path line makes the new captain-hold case fail with exactly the old message. The bootstrap HOME pin was proved by running the suite under a simulated developer HOME whose user-level tasks-axi config selects beads: pinned it passes 29/29, unpinned it fails on a spurious missing-bd line. The local feature suites all pass (control-relaunch 54, spawn-pool-base-freshen, handoff 25, remote handoff 12, atomicity 80, captain-hold 32 under umask 077), and I re-confirmed the merge shape upstream/main is an ancestor, the merge commit has both parents, 67 upstream commits, all three local commits reachable. The only non-zero exits are the two the intent already documents as host-caused and identical on the upstream baseline: fm-teardown's leaked-process-reap with lsof absent, and fm-public-followup's fixture cleanup with zero assertion failures. This change is shell/CLI only with no rendered surface, so the reviewer-visible evidence is before/after terminal transcripts rather than screenshots. Worktree left clean.

Evidence: Operator transcript: beads home + ambient TASKS_AXI_BACKEND=markdown, handoff before/after the fix

Source: Operator transcript: beads home + ambient TASKS_AXI_BACKEND=markdown, handoff before/after the fix

BEFORE: $ bin/fm-backlog-handoff.sh design guarded-item handed off 1 item(s) to design: guarded-item -> captain/data/backlog.md (the beads mirror) now has an EMPTY Queued section AFTER: $ bin/fm-backlog-handoff.sh design guarded-item error: secondmate backlog handoff requires markdown backlog storage; this home's .tasks.toml selects beads (data/backlog.md is a mirror). Keep handoff homes on backend = "markdown". -> the item is still in the mirror and the secondmate backlog is untouched

Operator transcript: secondmate backlog handoff on a home whose own .tasks.toml declares beads,
run from a shell that exports TASKS_AXI_BACKEND=markdown (the value every mover pins on its own
tasks-axi calls). The gate must answer on the home's declared storage, not the environment.

################ BEFORE FIX (guard reads fm_tasks_axi_backend, which honours the environment) ################
$ cat captain/.tasks.toml
backend = "beads"

[beads]
dir = "data"
path = "data/.beads"
$ export TASKS_AXI_BACKEND=markdown        # operator's ambient shell
$ bin/fm-backlog-handoff.sh design guarded-item
handed off 1 item(s) to design: guarded-item
  into /tmp/fm-guard-demo.EUrFtl/design-mate/data/backlog.md
error: handed off work to secondmate design, but no live receiver endpoint is recorded; the destination backlog is durable and the receiver was not woken
  (exit 1)

$ cat captain/data/backlog.md      # the beads mirror
## In flight

## Queued
## Done
$ cat design-mate/data/backlog.md  # the secondmate's backlog
## In flight

## Queued
- [ ] guarded-item - routed nowhere (repo: alpha)

## Done

>>> RESULT: the item WAS MOVED OUT of the beads mirror into the secondmate. Mirror and beads database now diverge.

################ AFTER FIX (guard reads fm_tasks_axi_backend_from_toml on the home's own file) ################
$ cat captain/.tasks.toml
backend = "beads"

[beads]
dir = "data"
path = "data/.beads"
$ export TASKS_AXI_BACKEND=markdown        # operator's ambient shell
$ bin/fm-backlog-handoff.sh design guarded-item
error: secondmate backlog handoff requires markdown backlog storage; this home's .tasks.toml selects beads (data/backlog.md is a mirror). Keep handoff homes on backend = "markdown".
  (exit 1)

$ cat captain/data/backlog.md      # the beads mirror
## In flight

## Queued
- [ ] guarded-item - routed nowhere (repo: alpha)

## Done
$ cat design-mate/data/backlog.md  # the secondmate's backlog
## In flight

## Queued

## Done

>>> RESULT: the move was refused; the item is still in the beads mirror and the secondmate backlog is untouched.
Evidence: Operator transcript: fm-captain-hold.sh verify on the tracked .tasks.toml against a real bd graph

Source: Operator transcript: fm-captain-hold.sh verify on the tracked .tasks.toml against a real bd graph

BEFORE ([beads] carries only dir): fm-captain-hold: the beads backend carries no graph path in <home>/.tasks.toml, so a migrated hold cannot be found AFTER (path = "data/.beads" added): fm-captain-hold: no captain-held task absent-tracked-key and no migrated hold for it in this home's configured backlog (data directory <home>/data); the nearest legacy identity ... also resolves to nothing

Operator transcript: bin/fm-captain-hold.sh verify on a home carrying this repository's tracked
.tasks.toml verbatim, against a real bd 1.2.2 graph, for a decision key the graph does not carry.

################ BEFORE FIX (tracked [beads] section carries only dir) ################
$ cat $FM_HOME/.tasks.toml     # copied verbatim from the tracked repo file
backend = "beads"

[markdown]
path = "data/backlog.md"
archive = "data/done-archive.md"
done_keep = 10

[beads]
dir = "data"

$ bin/fm-captain-hold.sh verify tracked-config-scout
fm-captain-hold: the beads backend carries no graph path in /tmp/fm-hold-demo.KQXoFA/home/.tasks.toml, so a migrated hold cannot be found
  (exit 1)

################ AFTER FIX (tracked [beads] section adds path = "data/.beads") ################
$ cat $FM_HOME/.tasks.toml     # copied verbatim from the tracked repo file
backend = "beads"

[markdown]
path = "data/backlog.md"
archive = "data/done-archive.md"
done_keep = 10

[beads]
dir = "data"
path = "data/.beads"

$ bin/fm-captain-hold.sh verify tracked-config-scout
fm-captain-hold: no captain-held task absent-tracked-key and no migrated hold for it in this home's configured backlog (data directory /tmp/fm-hold-demo.AUJwJu/home/data); the nearest legacy identity tracked-config-scout-decision-absent-tracked-key also resolves to nothing
  (exit 1)
Evidence: Bootstrap suite under a hostile user-level tasks-axi config

Source: Bootstrap suite under a hostile user-level tasks-axi config

BEFORE: HOME=<beads host config> ./tests/fm-bootstrap.test.sh -> exit 1 not ok - treehouse --lease support is accepted silently: expected silence, got: MISSING: bd (install: npm install -g @beads/bd) AFTER: HOME=<beads host config> ./tests/fm-bootstrap.test.sh -> exit 0, 29 ok, 0 not ok

Host-config hermeticity of tests/fm-bootstrap.test.sh.
Most of its cases give their home no .tasks.toml, so fm_tasks_axi_backend falls through to
$HOME/.tasks-axi/config.toml. Simulated hostile developer host:

$ cat /tmp/fm-hostile-home/.tasks-axi/config.toml
backend = "beads"

[beads]
dir = "data"
path = "data/.beads"

\### BEFORE FIX (suite at f9064c5, HOME=/tmp/fm-hostile-home)
$ HOME=/tmp/fm-hostile-home ./tests/fm-bootstrap.test.sh   -> exit 1
not ok - treehouse --lease support is accepted silently: expected silence, got: MISSING: bd (install: npm install -g @beads/bd)

\### AFTER FIX (suite at HEAD, same hostile HOME)
$ HOME=/tmp/fm-hostile-home ./tests/fm-bootstrap.test.sh   -> exit 0
29 assertions passed, no failures; last one:
ok - bootstrap validates crew-dispatch.json and reports malformed or unverified configs
Evidence: New beads storage guard suite, red before the fix and green after

Source: New beads storage guard suite, red before the fix and green after

guards reverted -> not ok - the markdown refusal must name the storage requirement (got: handed off 1 item(s) to design: guarded-item ...) as committed -> ok - local handoff refuses beads storage whatever TASKS_AXI_BACKEND says ok - remote outbox staging refuses beads storage whatever TASKS_AXI_BACKEND says ok - receive refuses beads storage whatever TASKS_AXI_BACKEND says ok - a declared markdown home is unaffected by an ambient beads override

$ bash tests/fm-backlog-beads-storage-guard.test.sh   # with the guard fix reverted
not ok - the markdown refusal must name the storage requirement (got: handed off 1 item(s) to design: guarded-item
  into <tmp>/local-markdown-sub/data/backlog.md
error: handed off work to secondmate design, but no live receiver endpoint is recorded; the destination backlog is durable and the receiver was not woken) (missing: 'requires markdown backlog storage')
--- output ---
handed off 1 item(s) to design: guarded-item
  into <tmp>/local-markdown-sub/data/backlog.md
error: handed off work to secondmate design, but no live receiver endpoint is recorded; the destination backlog is durable and the receiver was not woken
[exit 1]

$ bash tests/fm-backlog-beads-storage-guard.test.sh   # on the change as committed
ok - local handoff refuses beads storage whatever TASKS_AXI_BACKEND says
ok - remote outbox staging refuses beads storage whatever TASKS_AXI_BACKEND says
ok - receive refuses beads storage whatever TASKS_AXI_BACKEND says
ok - a declared markdown home is unaffected by an ambient beads override
[exit 0]
Evidence: control-relaunch markdown pin, before/after

Source: control-relaunch markdown pin, before/after

tests/fm-control-relaunch.test.sh - a fifth upstream suite that drives the real markdown backlog
through the installed tasks-axi CLI and was still inheriting this fork's tracked beads config.
Fixed here with the same pin the change already applies to four other suites.

\### BEFORE (HEAD as committed, run from the repo root whose .tasks.toml selects beads)
$ ./tests/fm-control-relaunch.test.sh   -> exit 1 after 50 assertions
not ok - a relaunch must not re-run a transition the row already reflects
warning: /tmp/fm-control-relaunch.t5pnP6/reverify-16026/home/data/rl40/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
error: task rl40 has no backlog item in this home, so dispatching it would leave a worker no record owns; add it first (tasks-axi add rl40 '<title>' --kind ship) and re-run
error: the replacement agent for rl40 could not be launched on claude
error: rl40's agent was stopped but the replacement did not launch; no agent is running, and its work plus the recorded progress note are preserved at /tmp/fm-control-relaunch.t5pnP6/reverify-16026/wt: expected exit 0, got 1

\### Cause, isolated: same test, only TASKS_AXI_BACKEND=markdown added
$ TASKS_AXI_BACKEND=markdown <same single test>
ok - relaunch re-reads the backlog item instead of blindly re-running the transition

\### AFTER (case home pinned with fm_test_markdown_tasks_toml, suite's own tasks-axi calls run from it)
$ ./tests/fm-control-relaunch.test.sh   -> exit 0, 54 assertions
ok - relaunch re-reads the backlog item instead of blindly re-running the transition
ok - relaunch heals an item that drifted out of In flight while the task stayed live
ok - fm-spawn --relaunch: refuses to adopt a worktree an earlier teardown already returned
ok - fm-control relaunch: a returned worktree refuses before the agent is touched

Upstream baseline (git archive af8c4b6, markdown .tasks.toml) for comparison: exit 0, 52 assertions.
54 - 52 = the two local relaunch/marker tests from df933e8.
Evidence: Targeted suite results on the merged tree, with upstream baseline comparison

Source: Targeted suite results on the merged tree, with upstream baseline comparison

Targeted suite results on the merged tree (host umask pinned to 077 so the process-event
private-directory check can pass; the host's default 0002 fails it identically on the
upstream baseline).

SUITE                                      FORK (HEAD)                  UPSTREAM BASELINE af8c4b6
tests/fm-backlog-beads-storage-guard       exit 0, 4 assertions (new)   n/a - new local suite
tests/fm-captain-hold-lifecycle            exit 0, 32 assertions        not run (fork-only case added)
tests/fm-control-relaunch                  exit 0, 54 assertions        exit 0, 52 assertions
tests/fm-spawn-pool-base-freshen           exit 0, 16 assertions        not run
tests/fm-teardown                          exit 1, 73 ok, 1 host fail   exit 1, 70 ok, SAME host fail
tests/fm-backlog-atomicity                 exit 0, 80 assertions        not run
tests/fm-bootstrap                         exit 0, 29 assertions        not run
tests/fm-public-followup                   exit 1, 73 ok, 0 assert fail fixture rm cleanup only

fm-teardown: the single failure on both trees is 'leaked-process-reap: leaked worktree process
survived teardown'. lsof is not installed on this host and bin/fm-teardown.sh resolves worktree-
rooted processes with 'lsof -a -d cwd'. The three fork-only assertions are exactly the local
worktree-ownership feature from df933e8:
  ok - an ordinary teardown still reaps and returns the worktree that is its own
  ok - a rerun after a post-return failure refuses loudly instead of tearing down a reassigned worktree
  ok - a rerun after the records are reconciled completes without repeating the worktree steps

fm-public-followup: 73 assertions pass and zero fail; the run exits 1 only on the fixture
teardown's "rm: cannot remove ...: Directory not empty", matching the author's own report.

Nine further suites that call the real tasks-axi CLI without pinning a .tasks.toml were swept
for the same beads-inheritance hazard; all nine pass on the merged tree:
  tests/fm-cd-pretool-check.test.sh exit=0
  tests/fm-secondmate-liveness.test.sh exit=0
  tests/fm-on.test.sh exit=0
  tests/fm-startup-memory-budget.test.sh exit=0
  tests/fm-shared-captain-inheritance.test.sh exit=0
  tests/fm-remote-doctor.test.sh exit=0
  tests/fm-secondmate-sync.test.sh exit=0
  tests/fm-secondmate-harness.test.sh exit=0
  tests/fm-session-start.test.sh exit=0
  FM_TEST_SUMMARY total=9 failed=0 skipped_gate=0 duration_ms=323926

Re-run independently on the final tree (HEAD = 1706a0e), same host, same results:
  tests/fm-backlog-beads-storage-guard  exit 0, 4 assertions
  tests/fm-backlog-handoff              exit 0, 25 assertions
  tests/fm-remote-backlog-handoff       exit 0, 12 assertions
  tests/fm-control-relaunch             exit 0, 54 assertions
  tests/fm-backlog-atomicity            exit 0, 80 assertions
  tests/fm-bootstrap                    exit 0, 29 assertions
  tests/fm-captain-hold-lifecycle       exit 0, 32 assertions (umask 077)
  tests/fm-public-followup              exit 1, 73 ok / 0 assertion failures (fixture rm cleanup only)
  tests/fm-spawn-pool-base-freshen      exit 0, 22 "ok -" lines
  tests/fm-teardown                     exit 1, 73 ok / 1 failure: leaked-process-reap
`command -v lsof` on this host returns nothing, which is the documented cause of that
single teardown failure and reproduces identically on the upstream baseline.
Evidence: Merge shape verification

Source: Merge shape verification

$ git merge-base --is-ancestor af8c4b6 HEAD -> 0 (upstream/main IS an ancestor) $ git log -1 --format='%h %p %s' ebafb88 -> ebafb88 parents: df933e8 af8c4b6 $ git rev-list --count f66be0f..af8c4b6 -> 67 reachable c03bfbe / d2ccd7b / df933e8

Merge shape: git must know upstream/main is an ancestor, and all three local commits must survive.

$ git merge-base --is-ancestor af8c4b6 HEAD; echo $?
0 (0 = upstream/main af8c4b6 IS an ancestor of HEAD)

$ git log -1 --format='%h %p %s' ebafb88
ebafb88  parents: df933e8 af8c4b6
  Merge upstream/main into the fork, preserving the three local features

$ git rev-list --count f66be0f..af8c4b6   # upstream commits merged in
67

The three local feature commits, each still reachable from HEAD:
  reachable c03bfbe feat: adopt beads backlog storage through tasks-axi
  reachable d2ccd7b fix(bin): launch remote-less projects from their local default branch (#1)
  reachable df933e8 fix(bin): refuse teardown of a worktree reassigned to another task (#2)

$ git log --oneline --first-parent df933e8..HEAD
db78839 no-mistakes(review): fix beads storage gates and tracked graph path
f9064c5 test: make the markdown-pin helper own its rule instead of a suite list
8bc450c test: pin seeded secondmate homes to markdown backlog storage
45b8f4e test: keep upstream's markdown backlog suites off this fork's beads storage
ebafb88 Merge upstream/main into the fork, preserving the three local features
- Outcome: 🔧 2 issues found → auto-fixed ✅ across 2 runs (46m36s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 3 infos
  • ⚠️ .tasks.toml:8 - The tracked [beads] section carries only dir = &#34;data&#34;, but upstream's af8c4b6 (now merged) reads the beads graph from [beads] path|binary|prefix only: captain_beads_toml_entries (bin/fm-captain-hold.sh:475) matches exactly those three keys, and captain_migration_scan_load refuses with rc=2 and "the beads backend carries no graph path in <root>/.tasks.toml" (bin/fm-captain-hold.sh:520) when path is absent. Concrete path on this fork's beads home: any hold key that no longer resolves exactly reaches resolve_entry -> resolve_migrated_entry -> rc=2; command_verify then exits 1 and bin/fm-teardown.sh:2913 refuses the scout with "has not passed the captain-call completion gate", while answer reports skipped: &lt;key&gt; (migrated-hold scan refused: ...). The scan can never run on this fork, so upstream's migrated-hold feature is inert here and the operator-facing reason is wrong (it says the graph path is missing, not that the id is unknown). The merged docs/configuration.md:102 documents this [beads] contract while docs/configuration.md:91 documents a tracked config that does not satisfy it. The intent's out-of-scope note attributes the seeded-home teardown refusal to the missing database, but source evidence shows the refusal fires earlier, on the missing key, and applies to the captain's own live home too. Fix is either adding path/binary/prefix to the tracked [beads] section or teaching captain_beads_toml_entries to accept dir - both are configuration/product calls.
  • ⚠️ bin/fm-backlog-handoff.sh:857 - Repointing the beads guards from the deleted fm_tasks_axi_storage_backend (which read $FM_HOME/.tasks.toml only) to fm_tasks_axi_backend makes them honour the environment: bin/fm-tasks-axi-lib.sh:143 returns $TASKS_AXI_BACKEND whenever the variable is set. Concrete sequence: on a home whose .tasks.toml selects beads, an operator shell with export TASKS_AXI_BACKEND=markdown (or even TASKS_AXI_BACKEND=, since the check is +x, not -n) makes the guards at bin/fm-backlog-handoff.sh:631, :857 and bin/fm-backlog-receive.sh:106 all evaluate non-beads and pass; the pinned env TASKS_AXI_BACKEND=markdown tasks-axi mv ... --file &#34;$MAIN_BACKLOG&#34; --to &#34;$outbox&#34; then moves item blocks out of data/backlog.md, which on beads storage is a regenerated mirror. The next tasks-axi render rewrites the mirror from the beads database and the moved items reappear in the parent while the secondmate also holds them - exactly the divergence the guard exists to prevent, and the pre-merge guard could not be overridden this way. The intent's verification covers in-script leakage (correct - all four pins are single-command env prefixes) but not an inherited exported value. Consider reading the home's declared storage via fm_tasks_axi_backend_from_toml &#34;$FM_HOME/.tasks.toml&#34; for these safety gates, or refusing outright when TASKS_AXI_BACKEND is set on a home whose file selects beads.
  • ℹ️ tests/fm-bootstrap.test.sh:477 - fm_tasks_axi_backend adds a third precedence step the deleted local function did not have: when the home has no .tasks.toml it falls back to $HOME/.tasks-axi/config.toml (bin/fm-tasks-axi-lib.sh:151). Only 5 of this suite's 37 bootstrap invocations write a .tasks.toml into their case home, and none of them pin HOME, so on a host whose user-level tasks-axi config selects beads the other 32 runs now emit MISSING: bd (install: npm install -g @beads/bd) and every [ -z &#34;$out&#34; ] / exact-output assertion in the suite fails for a reason unrelated to what it tests. Pinning HOME to a throwaway directory in the run helper (or seeding a markdown .tasks.toml in make_fake_toolchain's home) removes the host dependence.
  • ℹ️ tests/fm-public-followup.test.sh:1863 - pin_seeded_home is called on $child in test_retire_refuses_unbound_existing_secondmate, but that directory is built by a bare mkdir -p &#34;$home/unbound-mate&#34; rather than by bin/fm-home-seed.sh, so the helper's stated reason (tests/fm-public-followup.test.sh:132, "bin/fm-home-seed.sh clones this repository, so a seeded secondmate home inherits the tracked backlog storage") does not apply there. It is not harmful - without the pin that directory has no .tasks.toml at all and would fall through to the user-level config - so keeping it is defensible; noting it only because the helper name and comment now cover a call site they do not describe.

🔧 Fix: fix beads storage gates and tracked graph path
3 infos still open:

  • ℹ️ bin/fm-backlog-handoff.sh:857 - The three gates now read only $FM_HOME/.tasks.toml, so they no longer see the $HOME/.tasks-axi/config.toml step that fm_tasks_axi_backend (bin/fm-tasks-axi-lib.sh:154) consults. Concrete residual path: a home whose .tasks.toml exists but carries no root-level backend key makes fm_tasks_axi_backend_from_toml return rc=1 with empty output, the gate reads non-beads and passes, while tasks-axi itself would fall through to a user-level config selecting beads - in which case data/backlog.md really is a regenerated mirror and the pinned env TASKS_AXI_BACKEND=markdown tasks-axi mv would move blocks out of it. Reachability is low: every firstmate and seeded secondmate home is a clone of this repository and so carries the tracked .tasks.toml with an explicit backend. Noting it only because it is the one precedence step the gates deliberately dropped; this is exactly the pre-merge guarantee the author asked for, so no change is implied.
  • ℹ️ tests/fm-backlog-beads-storage-guard.test.sh:128 - make_beads_receiver computes bytes and the SHA-256 of the delivered outbox and its comment says the fixture is built "so receipt is refused only by storage", but neither value is ever compared: bin/fm-backlog-receive.sh validates them against $PARENT_REAL/.mate.upload-generation (bin/fm-backlog-receive.sh:120-125), and the fixture never writes that file. Without the storage gate this fixture would die at "delivered outbox generation is unavailable or unsafe" rather than proceeding, so the stated "only by storage" property does not hold. The test still discriminates correctly, because assert_contains &#34;$out&#34; &#39;requires markdown backlog storage&#39; would fail on that other refusal. Either drop the unused byte/digest computation and reword the comment, or write the generation file so the fixture matches its claim.
  • ℹ️ tests/fm-backlog-beads-storage-guard.test.sh:13 - The new suite's header states "This suite owns that contract for all three entry points; the surrounding handoff behavior lives in tests/fm-backlog-handoff.test.sh", but tests/fm-backlog-handoff.test.sh:64 test_handoff_refuses_beads_backlog_storage still asserts the same local-handoff beads refusal with the same 'requires markdown backlog storage' message - it is not "surrounding handoff behavior", it is the contract this file claims to own, and it duplicates this suite's unset ambient case. Under the repository's "one owner per contract with cross-references rather than restatements" rule the cross-reference is inaccurate. Either fold the older case into this suite or reword the header to say the local-only refusal is co-owned there.
🔧 **Test** - 2 issues found → auto-fixed ✅
  • ⚠️ tests/fm-control-relaunch.test.sh:129 - tests/fm-control-relaunch.test.sh (one of the eight conflicted files) was still inheriting this fork's tracked beads .tasks.toml, so its own tasks-axi add/start --file ... calls addressed a beads store the fixture never created. test_relaunch_reverifies_an_already_in_flight_item_instead_of_rewriting_it failed with "task rl40 has no backlog item in this home" (exit 1 after 50 assertions), while the clean upstream af8c4b6 baseline passed 52/52. Isolating the single test with only TASKS_AXI_BACKEND=markdown added made it pass, confirming the cause. Fixed in the working tree with the same pin the change already applies to four other suites: fm_test_markdown_tasks_toml on the case home in new_case, and seed_backlog/backlog_state now run their tasks-axi calls from that home. Suite now exits 0 with 54 assertions (upstream's 52 plus the two local relaunch/marker tests from df933e8). The fix is uncommitted in tests/fm-control-relaunch.test.sh.
  • ℹ️ Two suites cannot reach a clean exit on this host for reasons proven identical on a clean git archive af8c4b6 baseline, and per the intent were not "fixed": tests/fm-teardown.test.sh fails only on leaked-process-reap because lsof is not installed and bin/fm-teardown.sh resolves worktree-rooted processes via lsof -a -d cwd (fork 73 ok, baseline 70 ok, same single failure); tests/fm-captain-hold-lifecycle.test.sh fails at "process-event state root is not a private directory" under the host's default umask 0002 (fork 13 ok, baseline 13 ok, same failure) and passes fully 32/32 when re-run under umask 077. tests/fm-public-followup.test.sh passes 73 assertions with zero failures and exits 1 only on its fixture teardown's "rm: Directory not empty".
  • git merge-base --is-ancestor af8c4b6 HEAD plus git log -1 --format=&#39;%h %p&#39; ebafb88 and reachability of c03bfbe/d2ccd7b/df933e8 - merge shape and local-commit survival
  • Manual operator transcript: bin/fm-backlog-handoff.sh design guarded-item on a beads-declaring home with TASKS_AXI_BACKEND=markdown exported, run against both the committed tree and a pre-fix copy with the guards reverted to fm_tasks_axi_backend
  • Manual operator transcript: bin/fm-captain-hold.sh verify tracked-config-scout on a home carrying the tracked .tasks.toml verbatim against a real bd init graph (bd 1.2.2), run with and without the added path = &#34;data/.beads&#34;
  • bin/fm-test-run.sh tests/fm-backlog-beads-storage-guard.test.sh - new suite, exit 0, 4 assertions
  • Pre-fix reproduction of the new suite: same suite run against a copy of HEAD with the three guards reverted to fm_tasks_axi_backend - fails, and the handoff actually moves the item out of the beads mirror
  • HOME=/tmp/fm-hostile-home ./tests/fm-bootstrap.test.sh with a user-level ~/.tasks-axi/config.toml selecting beads, against the suite at f9064c5 (exit 1, "MISSING: bd") and at HEAD (exit 0, 29 assertions)
  • umask 077; ./tests/fm-captain-hold-lifecycle.test.sh - exit 0, 32 assertions, including ok - the tracked .tasks.toml gives the migrated-hold scan a reachable beads graph
  • ./tests/fm-control-relaunch.test.sh before the fix (exit 1, 50 ok) and after (exit 0, 54 ok), plus the same suite on a git archive af8c4b6 baseline (exit 0, 52 ok)
  • Cause isolation: TASKS_AXI_BACKEND=markdown applied to test_relaunch_reverifies_an_already_in_flight_item_instead_of_rewriting_it alone, which flips it from failing to passing
  • ./tests/fm-teardown.test.sh on HEAD (exit 1, 73 ok) vs git archive af8c4b6 baseline (exit 1, 70 ok) - same single leaked-process-reap failure, 3-assertion delta is exactly the local worktree-ownership tests
  • bin/fm-test-run.sh tests/fm-spawn-pool-base-freshen.test.sh - exit 0, remote-less default-branch freshen behavior
  • umask 077; ./tests/fm-backlog-atomicity.test.sh - exit 0, 80 assertions
  • umask 077; ./tests/fm-public-followup.test.sh - 73 assertions pass, zero failures, exit 1 only on the fixture rm: Directory not empty
  • bin/fm-test-run.sh over the nine remaining suites that invoke the real tasks-axi CLI without pinning a .tasks.toml (fm-cd-pretool-check, fm-on, fm-remote-doctor, fm-secondmate-harness, fm-secondmate-liveness, fm-secondmate-sync, fm-session-start, fm-shared-captain-inheritance, fm-startup-memory-budget) - total=9 failed=0

🔧 Fix: pin control-relaunch case homes to markdown backlog storage
✅ Re-checked - no issues remain.

  • bash tests/fm-backlog-beads-storage-guard.test.sh - exit 0, 4 assertions (new suite covering fm-backlog-handoff.sh:631, :857 and fm-backlog-receive.sh:106 under TASKS_AXI_BACKEND unset/markdown/empty, plus the inverse markdown-home case)
  • Red/green proof of the guard fix: reverted fm_tasks_axi_backend_from_toml back to fm_tasks_axi_backend in bin/fm-backlog-handoff.sh and bin/fm-backlog-receive.sh, re-ran the suite (fails: the item is handed out of the beads mirror), then restored the tree
  • Operator CLI demo on a beads-declared home with export TASKS_AXI_BACKEND=markdown: bin/fm-backlog-handoff.sh design ship-the-thing before/after the fix, capturing both backlog.md files each time
  • (umask 077; bash tests/fm-captain-hold-lifecycle.test.sh) - exit 0, 32 assertions, including the new test_tracked_tasks_toml_reaches_a_real_beads_graph running for real against bd 1.2.2
  • Red/green proof of the tracked config fix: deleted path = &#34;data/.beads&#34; from .tasks.toml, re-ran the captain-hold suite (fails with "the beads backend carries no graph path"), then restored .tasks.toml
  • Operator CLI demo: bin/fm-captain-hold.sh verify &lt;scout&gt; on a home carrying the tracked .tasks.toml verbatim over a real bd init graph, before/after the path line
  • bash tests/fm-bootstrap.test.sh - exit 0, 29 assertions; and HOME=&lt;hostile home with .tasks-axi/config.toml selecting beads&gt; bash tests/fm-bootstrap.test.sh - exit 0, 29 ok / 0 not ok, versus the same run with the HOME pin removed which fails on MISSING: bd
  • bash tests/fm-control-relaunch.test.sh - exit 0, 54 assertions
  • bash tests/fm-backlog-handoff.test.sh - exit 0, 25 assertions
  • bash tests/fm-remote-backlog-handoff.test.sh - exit 0, 12 assertions
  • bash tests/fm-backlog-atomicity.test.sh - exit 0, 80 assertions
  • bash tests/fm-spawn-pool-base-freshen.test.sh - exit 0
  • bash tests/fm-public-followup.test.sh - 73 assertions pass, 0 assertion failures; exit 1 only from the fixture teardown's rm: Directory not empty
  • bash tests/fm-teardown.test.sh - 73 ok, single leaked-process-reap failure; confirmed command -v lsof returns nothing on this host, the documented cause
  • Merge shape: git merge-base --is-ancestor af8c4b6 HEAD (ancestor), git log -1 --format=&#39;%h %p&#39; ebafb88 (two parents df933e8 + af8c4b6), git rev-list --count f66be0f..af8c4b6 = 67, and each of c03bfbe / d2ccd7b / df933e8 reachable from HEAD
  • git status --porcelain after every temporary revert - worktree left clean, no transient artifacts
⚠️ **Document** - 2 infos
  • ℹ️ docs/captain-hold-lifecycle.md:126 - Judgment call, deliberately not applied. This change adds test_tracked_tasks_toml_reaches_a_real_beads_graph to tests/fm-captain-hold-lifecycle.test.sh, whose "Verification record" section enumerates what that suite proves. I recorded the new guarantee beside its own invariant in docs/configuration.md (the owner of this fork's tracked .tasks.toml contract) rather than extending the upstream-owned enumeration, because the fact is fork configuration rather than captain-hold migration behavior, and a fork-local clause in that paragraph would re-conflict on every future upstream sync. The enumeration is now non-exhaustive but states nothing false. If the captain would rather have every case of that suite listed in one place, extending the migration-family sentence there is the alternative.
  • ℹ️ docs/configuration.md:113 - The intent reports the seeded-secondmate-on-beads limitation as out of scope and deliberately unfixed, so I documented the current behavior and its remedy rather than changing anything. Verified directly: a directory carrying the tracked .tasks.toml with no database fails tasks-axi list with beads backend: no beads database found (rc=2), and non-forced scout teardown then refuses at bin/fm-teardown.sh:2913 on the captain-hold verification it cannot complete. The product decision (whether a beads primary should seed markdown secondmate homes automatically) remains the captain's and is untouched here.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 30 commits August 29, 2026 10:34
…nchenguid#3273)

* fix(bin): keep home-summary publication bounded and off the watcher beat

A home whose tasks had accumulated ordinary status history could not publish
state/home-summary.json at all, and every attempt starved the watcher's
liveness beacon while it failed silently.

The producer's per-task open-decision fold spent tens of milliseconds per
status line on a bash 3.2 global bracket-class substitution used only as a
blank-line guard. On a real home that made the whole ledger producer take
minutes, so publication burned its full FM_HOME_SUMMARY_TIMEOUT on every
attempt and never completed. Replace that guard with an equivalent case glob
in the one fold owner, which both the whole-file and cursor-backed folds use.

Bound each per-task current-state read in the snapshot with
FM_SNAPSHOT_CREW_STATE_TIMEOUT. For a remote secondmate that read crosses ssh,
whose dead-peer detection deliberately never kills a slow-but-alive remote
command, so nothing else bounded it.

Detach the watcher's two publication triggers from the poll loop. The loop
owns the beacon that fm-guard.sh reads as proof supervision is alive, and an
inline publication put up to a full publication deadline between two beacon
touches. A single in-flight publication is tracked so a slow one cannot
accumulate clones.

Report a repeatedly failing publication at session start. Publication stays
deliberately non-fatal to its caller, so the existing bounded home-local
failure record is now surfaced as a HOME_SUMMARY bootstrap line once the
ledger is absent or stale and failures have been recorded since.

* no-mistakes(review): Preserve home-summary failure attempt ordering

* no-mistakes(review): Enforce durable home-summary single-flight and ordering

* no-mistakes(review): Derive failure ordering from publication boundaries

* no-mistakes(review): Restore best-effort failure logging and publication scoping

* no-mistakes(review): Make ordering regression sensitive to one failure

* no-mistakes(document): Correct HOME_SUMMARY diagnostic guidance
…henguid#3268)

* fix(supervision): classify the appended status span, not the last line

An actionable project update could be classified as routine and absorbed, so
a worker that raised a decision, hit a blocker, failed, or finished stalled
silently with the captain never told.

Trigger, mask, symptom. A worker appends a captain-relevant event
(`needs-decision`, `blocked`, `failed`, `done`). Any later routine append -
a `working:` progress note - lands before the supervisor classifies the
batch; the watcher's 30s signal-grace linger exists precisely to coalesce a
status write with the same turn's turn-end, so this window is ordinary
rather than rare. Both supervisors then asked "is the LAST line
captain-relevant?", read the routine line, and absorbed the wake. The
`.seen-*` suppressor advanced either way, so nothing ever re-read the event.
When the crew was also provably working, the no-verb fallback absorbed it
too, which is why the event disappeared completely instead of surfacing late.

Reproduced end to end against a real watcher before any change: with the
trailing `working:` append the watcher never exits and the wake queue stays
empty; with that one line removed - the smallest counterfactual - the same
`needs-decision` surfaces and queues. The away-mode daemon's `classify_signal`
returns `self|routine signal` for a `blocked:` event under the same mask,
which is the worse case because no captain is present to notice.

The proven path was already in the tree: `status_open_decisions` fixed this
exact masking for the durable decision fold, and its header states the rule -
reading an append-only event log last-event-wins cannot represent an earlier
event that a later unrelated line moved past. The classification path was
never migrated to that read model. That is the earliest divergence, and the
fix is to migrate it rather than to special-case the symptom.

`status_span_first_actionable` in bin/fm-classify-lib.sh is the new single
owner: it reads the bytes at or after a caller-supplied position and returns
the first still-live captain-relevant event. Each supervisor supplies its own
position, because the always-on watcher and the away-mode daemon classify the
same stream independently and must not share one cursor: the watcher reads
the size already recorded in its `.seen-*` signature (no new state) and its
`.hb-surfaced-<task>` backstop marker, and the daemon its
`.subsuper-seen-status-<task>` marker. Those two markers held the escalated
line and now hold the escalated-through byte offset, which also removes a
second defect in the same code - content dedup silently swallowed a genuinely
new event whose text repeated an older one. An absent, malformed, or
past-the-end position reads the whole log, so uncertainty surfaces events
rather than losing them, and a marker an older build wrote as a status line
reads that way too. Status logs are only ever appended to, including across a
reused task id, so a recorded position keeps its meaning.

A `needs-decision`/`blocked` event in the span is retired only when the
whole-file fold proves its key closed; `status_open_decisions` stays the sole
owner of that rule, so same-key reopening and reserved-key namespaces need no
second implementation here. Every other captain-relevant event is terminal
and always actionable.

Both backstops now walk every status log instead of only those whose last
line looks captain-relevant, because the event a backstop most needs to catch
is exactly one a later append has moved past. That leaves
`scan_captain_relevant_statuses` with no callers, and it is removed rather
than left as a working copy of the defective read model.

Regression coverage exercises the classifier and both supervisors through
their own interfaces: the masked decision, the captain-reported
release/install completion followed by cleanup chatter, and the away-mode
blocker all surface; a routine append after an already-classified event stays
absorbed, so the fix does not convert ordinary progress into wakes; and the
heartbeat backstop catches a masked event the per-wake path missed. The
end-to-end watcher tests drive a real fm-watch.sh with the crew reported as
provably working, which is the configuration that made the original stall
silent.

Two further claims in the supplied RCA are deliberately not patched here.
"Repeated operational recoveries produced all-clear replies despite known
actions" is downstream of this same cause, not an independent contributor: an
all-clear reply is the documented response when the specific event needs no
action, so a classification that wrongly reported "no action" produces it, and
correcting the classification removes it. "The project was subjected to
validation requirements outside its accepted path" is delivery-mode selection,
which AGENTS.md section 7 owns; no code changed here touches it, so it is out
of scope.

Harness and backend axes were inspected rather than assumed: nothing in this
path reads a vendor-emitted signal. The status log's format and append
protocol are Firstmate's own and identical for every harness, and no runtime
backend reads or writes `.status` files (`bin/backends/*` contain no reference
to them). The surrounding triage's only backend touchpoints - pane capture and
the authoritative crew-state read - are unchanged. No live-harness guard
applies and no per-harness verification record changes.

Verified with `bin/fm-lint.sh`, `bin/fm-doc-audience-check.sh`, and
`bin/fm-test-run.sh --changed --base origin/main`.

* no-mistakes(review): Prevent status races and surface classification failures

* no-mistakes(review): Surface unreadable signals and preserve AFK endpoints

* no-mistakes(review): Route stale wakes through captured span verdicts

* no-mistakes(review): Retire supervision offsets with reused task state

* no-mistakes(review): Bind status offsets and preserve live decision origins

* no-mistakes(review): Strengthen status identity with verified birth time

* no-mistakes(review): Skip turn-end markers during status classification

* no-mistakes(review): Preserve status presentation with platform-strength identities

* no-mistakes(review): Retain failed wakes and advance routine checkpoints

* no-mistakes(review): Surface all events and retain unreadable wakes

* no-mistakes(review): Treat absent status logs as successful empty spans

* no-mistakes(review): Bound repeated classification failures with durable receipts

* revert(supervision): drop the failure-receipt and durable-retry machinery

Captain-authorized revert to the minimal fix. Review rounds added a durable
failure-receipt store and wake-retention-on-failure to bound repeated
classification failures. That machinery grew larger than the fix it protected
and kept producing its own defects: an unreadable log still looped forever
because the always-on watcher never consulted the receipt, and the receipt was
persisted before its diagnostic was durably queued, so a crash in between
swallowed the alarm outright. Those two defects go away with the code that
contained them rather than being repaired.

Removed: the failure-receipt path, fingerprint, record and clear helpers and
their retirement bookkeeping; the retention of a durable wake when
classification fails; and the error-propagation plumbing in both supervisors
that existed only to drive them.

Kept, because it is the accepted fix rather than the declined machinery: span
classification of the events appended since a supervisor last looked, in both
supervisors and both backstops; reporting every actionable event in a span and
committing a position only through what was reported; naming the live opening of
a reopened decision; treating an absent log as ordinary and an unreadable one as
worth reporting; the non-.status filter; and the platform-strength identity that
guards a position commit without failing a read.

Replacement behavior for a log that cannot be classified: report it once, do NOT
advance the classification position so the content is classified from where it
stopped once readable, and DO advance the wake signature so the report is
bounded to one per distinct file state. Reporting and reading are different acts:
telling the captain about a log is not the same as having read it, and only the
latter may move a classification position.

The residual risk is explicit and accepted: there is no guaranteed automatic
retry inside a crash-mid-read window, and the locked session-start replay of the
durable queue covers it. That rationale is recorded at mark_escalated_seen so a
future reader does not reintroduce the retry as a "missing" guarantee.

Also fixes lint failures that arrived with the review-fix commits and were never
caught because the run never reached its lint step: an unfollowable conditional
source directive, a second unquoted-expansion site left after a call was split
across lines, cleanup of the file being read inside its own read loop (restructured
to one post-loop teardown rather than three in-loop copies), stub functions in
tests that are invoked indirectly, and a test local left unused when its
assignment was replaced by a helper. bin/fm-lint.sh passes on the default branch,
so these were introduced here.

Verified with `bin/fm-lint.sh`, the end-to-end masked-decision and away-mode
reproductions, and `bin/fm-test-run.sh` over the supervision, wake-queue,
wake-drain, watch-arm and inactive-reconcile suites (6 scripts, 0 failures).

* no-mistakes(review): Correct classification failure contract documentation

* no-mistakes(review): Bound unreadable status reports without skipping classification

* no-mistakes(review): Preserve escalation markers when buffering fails

* no-mistakes(review): Detect permission recovery without advancing classification

* no-mistakes(document): Document status span classification contract

* no-mistakes(ci): Fixed CI failures by lazily loading classification helpers in fm-wake-lib, preserving minimal recovery/remote fixtures; added a public current-status marker helper and updated behavioral fixtures to use the v2 marker contract; resolved ShellCheck variable collisions in fm-control and fm-public-followup-lib. Verified fm-lint, bash syntax, fm-control, public-followup, wake-queue, send-resolve-key, captain-hold, pending-reply, remote-reply, remote-backlog-handoff, turnend-guard, and Claude autoarm tests. The Pi branch suite reached a separate local stock-render mismatch under Node 24; its CI-reported missing-classifier failure path is fixed

* no-mistakes(review): Escalate blockers while preserving declared-wait cadence

* no-mistakes(review): Clarify actionable events override wait self-handling

* no-mistakes(review): Surface rejected decisions and dangling status links

* no-mistakes(document): Document reserved-key reconciliation classification

* no-mistakes(ci): Fixed the flaky portable serial CI test by modeling the retained staging directory as genuinely owned by a live process and aging both fixtures deterministically. This removes scheduler-timing dependence while verifying the worker reaps abandoned staging and preserves live staging. Verified with fm-remote-transport-lanes.test.sh, bin/fm-lint.sh, bash syntax, and git diff --check

* no-mistakes(document): Correct away-mode classification documentation
…#3289)

* docs: split harness adapter operations reference

* no-mistakes(review): Fix harness adapter routing and ownership contracts

* no-mistakes(review): Prune duplicate harness adapter ownership prose

* no-mistakes(review): Fix default effort routing and Grok max semantics

* no-mistakes(review): Remove source-only routing test and duplicate semantics

* no-mistakes(review): Add local harness adapter instruction evaluation

* no-mistakes(review): Fix harness evaluation gating and change mapping

* no-mistakes(test): Captain, require explicit harness instruction evaluator model

* no-mistakes(document): Fix harness adapter documentation references
* test(fixtures): share fake-toolchain and spawn-world builders

Future tests can start from tests/fixtures.sh instead of copying stubs, and a
no-mistakes version-floor bump is one constant rather than a multi-file edit.

Migrated this round: fm-busy-adapter-wiring, fm-spawn-pool-base-freshen,
fm-grok-harness, fm-tangle-guard, fm-gate-refuse, fm-spawn-dispatch-profile.
Left for opportunistic migration: remaining make_spawn_fakebin copies
(trace-context, kimi, muse, backend), the make_stubs send cluster, and the
fake no-mistakes version banners in bootstrap/session-start/secondmate suites.
Did not touch tests/fm-pr-check-security.test.sh.

* no-mistakes(review): Prevent fake SSH test from blocking on stdin

* no-mistakes(document): Clarify shared fixture documentation

* no-mistakes(ci): Fixed the flaky watcher triage test by extending its startup-sensitive timer-repair wait from 3s to 10s, matching existing loaded-runner budgets. Verified with the full tests/fm-watch-triage.test.sh suite, bash syntax validation, and git diff checks

* no-mistakes(ci): Fixed portable serial shard 4 by updating the inactive-reconcile fixture to prime status through the public fm_wake_status_mark_current API, ensuring classifier helpers load correctly and preventing the idle watcher from exiting. Verified the test three consecutive times, ran fm-test-fixtures, ShellCheck, bash syntax checks, and git diff checks. The outer no-mistakes executor can now bind a fresh attestation to the new head

* no-mistakes(ci): Added behavioral coverage proving the shared spawn tmux fixture defaults an unset FM_FAKE_PANE_PATH to empty. Verified the fixture suite, ShellCheck, syntax/diff checks, and all six migrated test suites; all passed. The outer executor can now bind a fresh no-mistakes attestation to the updated head
* feat(bin): retire completed PR-check migration machinery

Every registered home already carried both completion markers, and no
installer still creates pre-migration checks. Remove the one-time migrate
script, its bootstrap/watch/teardown/docs surface, and migration-path tests
without weakening live check-trust or PR-poll authentication.

* no-mistakes(review): Restore live PR-check security coverage

* no-mistakes(document): Refresh retired PR-check documentation

* no-mistakes(ci): Fixed both failing CI checks. Updated inactive-reconcile setup to use the public status-marking interface, preventing false watcher exits. Made remote-job shutdown deterministic by stopping the complete worker tree before tampering. Verified both affected test suites, repeated inactive reconciliation, shell syntax, and git diff checks
…3247)

* feat(extensions): bind trusted external process-event adapters

* no-mistakes(review): Enforce owner and remote-home conformance

* no-mistakes(review): Enforce serialized remote extension package lifecycle

* no-mistakes(review): Enforce identity-conditional extension retirement

* no-mistakes(review): Serialize extension retirement and recover crash cuts

* no-mistakes(review): Unify retirement worker and lifecycle lock ownership

* no-mistakes(review): Harden extension lifecycle retirement serialization

* no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries

* no-mistakes(document): Clarify built-in-only captain answer routing

* no-mistakes(lint): Captain: fix extension binding ShellCheck findings

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Use isolated UID mapping for owner conformance

* no-mistakes(review): Captain: remove forbidden CI ownership wrapper

* no-mistakes(review): Serialize extension binding publication

* no-mistakes(review): Document ordinary CI owner-fixture exclusion

* no-mistakes(review): Quarantine orphaned handshake descendants

* no-mistakes(test): Fix orphan attribution

* no-mistakes(test): Harden process tracker baseline

* no-mistakes(test): Harden detached descendant attribution

* no-mistakes(test): Use exact invocation-group cleanup

* no-mistakes(test): Bound remote conformance transport crossings

* no-mistakes(test): Parallelize isolated extension conformance tests

* no-mistakes(test): Lifecycle suite still exceeds deadline

* feat(extensions): bind trusted external process-event adapters

* no-mistakes(review): Enforce owner and remote-home conformance

* no-mistakes(review): Enforce serialized remote extension package lifecycle

* no-mistakes(review): Enforce identity-conditional extension retirement

* no-mistakes(review): Serialize extension retirement and recover crash cuts

* no-mistakes(review): Unify retirement worker and lifecycle lock ownership

* no-mistakes(review): Harden extension lifecycle retirement serialization

* no-mistakes(review): Unify extension registration and overridden-state lifecycle boundaries

* no-mistakes(document): Clarify built-in-only captain answer routing

* no-mistakes(lint): Captain: fix extension binding ShellCheck findings

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Use isolated UID mapping for owner conformance

* no-mistakes(review): Captain: remove forbidden CI ownership wrapper

* no-mistakes(review): Serialize extension binding publication

* no-mistakes(review): Document ordinary CI owner-fixture exclusion

* no-mistakes(review): Quarantine orphaned handshake descendants

* no-mistakes(test): Fix orphan attribution

* no-mistakes(test): Harden process tracker baseline

* no-mistakes(test): Harden detached descendant attribution

* no-mistakes(test): Use exact invocation-group cleanup

* no-mistakes(test): Bound remote conformance transport crossings

* no-mistakes(test): Parallelize isolated extension conformance tests

* no-mistakes(test): Lifecycle suite still exceeds deadline

* no-mistakes(review): Split extension conformance and forward remote transfer input

* no-mistakes(review): Forward malformed remote payloads through fm-on

* no-mistakes(review): Bound extension coordinator failure cleanup

* no-mistakes(test): Skip repeated orphan sweep in coordinator children

* no-mistakes(test): Queue isolated extension sections through bounded workers

* no-mistakes(test): Bound extension coordinator lane cleanup

* no-mistakes(test): Split remote lifecycle coordinator sections

* no-mistakes(test): Coordinator probes pass; aggregate deadline remains

* no-mistakes(test): Launch extension sections concurrently

* no-mistakes(test): Fix coordinator marker publication

* no-mistakes(test): Stabilize extension binding coordinator timing

* no-mistakes(lint): Fix extension binding ShellCheck warnings

* fix(extensions): prove invocation cleanup before retirement

* no-mistakes(review): Harden process-event inbox confinement

* no-mistakes(review): Preserve legacy capture parity

* no-mistakes(review): Protect external registry staging

* no-mistakes(test): Stabilize bounded extension conformance aggregate

* no-mistakes(document): Document external evidence confinement

* no-mistakes(ci): CI phase fixed. The failure was a flaky fixture in `tests/fm-remote-transport-lanes.test.sh`: its “fresh/in-use” staging directory had no live owner identity, so the real worker correctly reaped it once the 1-second age boundary elapsed on slower CI. The fixture now records the active test shell’s exact PID/start identity and cleans those records before removal. Verified: `bash tests/fm-remote-transport-lanes.test.sh` exits 0 with all checks passing; `git diff --check` passes. Provider check retrieval was also retried successfully, resolving the selected manual CI finding. Changed file: `tests/fm-remote-transport-lanes.test.sh`

* no-mistakes(review): Harden extension staging and lifecycle reservation

* no-mistakes(review): Harden external staging and lifecycle reservations

* no-mistakes(review): Wire capture helper into remote conformance

* no-mistakes(review): Pin external capture handoff and signal failures

* no-mistakes(review): Bind pinned capture authority to inherited descriptor

* no-mistakes(review): Harden descriptor-bound capture authority

* no-mistakes(review): Harden core capture reservation authority

* no-mistakes(review): Harden capture reservation boundaries

* no-mistakes(review): Harden capture reservations and cleanup

* no-mistakes(review): Harden capture handoff and reservation cleanup

* no-mistakes(review): Bind capture handoff to claim descriptors

* no-mistakes(review): Release lifecycle locks after host crashes

* no-mistakes(review): Pin reservation recovery to recorded state roots

* no-mistakes(review): Reject control bytes in claim state roots

* no-mistakes(test): Stabilize extension capture descriptor handoff

* no-mistakes(document): Document extension capture authority boundary

* no-mistakes(lint): Fix ShellCheck extension binding warnings

* no-mistakes(ci): CI phase result: fixed `bin/fm-procevent.sh` by initializing the shared `capture_state` sentinel for built-in adapters under `set -u`. This prevents normal built-in captures from aborting before publication. Verified: `bash -n bin/fm-procevent.sh` and `git diff --check` pass. The focused process-event suite was run locally but stopped earlier at a local detached-runner claim failure (`reconcile never claimed the registered source`), before the CI-reported post-capture path; CI evidence confirms the fixed unset-variable failure affected the failing remote, board, watcher, and process-event checks

* no-mistakes(document): Correct extension namespace creation timing

* no-mistakes(lint): Initialize capture locals for ShellCheck
* fix(bin): deliver the real definition of done to a promoted scout, and ban --yes

A promoted scout used to receive a free-form placeholder instead of the
mode-specific Definition of done a briefed ship worker gets, so it never
saw the ask-user escalation rule or the --yes prohibition. That gap is the
concrete reason one incident's worker drove validation with --yes and
answered its own ask-user findings.

- Add bin/fm-dod-lib.sh as the single owner of a ship task's mode-specific
  Definition of done, rendered by both bin/fm-brief.sh and bin/fm-promote.sh
  so the two contracts cannot drift.
- bin/fm-promote.sh now writes data/<id>/ship-instructions.md carrying the
  scratch inventory, clean base, ship branch, and that Definition of done,
  and prints the fm-send.sh command that delivers it.
- State the --yes ban as a prohibition rather than a preference, without
  claiming an enforcement the tool does not provide.
- Cover both through the real promotion and brief paths in
  tests/fm-task-delivery.test.sh and tests/fm-brief.test.sh.

* no-mistakes(review): Publish promotion instructions before committing task state

* no-mistakes(review): Supersede conflicting scout delivery rules after promotion

* no-mistakes(review): Reject invalid promotion instruction destinations

* no-mistakes(document): Align documentation with promotion delivery contracts

* no-mistakes(ci): Fixed both CI findings. Promoted workers now receive an explicit worktree-isolation check before branch creation, with instructions to stop and escalate if they are in the primary checkout. Updated behavioral coverage to verify the delivered promotion payload, and aligned the ask-user authority test with the new fleet-wide --yes prohibition. Verified with bin/fm-lint.sh, tests/fm-brief.test.sh, tests/fm-ask-user-authority.test.sh, tests/fm-task-delivery.test.sh, and git diff --check

* no-mistakes(ci): Made tests/fm-ask-user-authority.test.sh executable so the modified colocated behavioral test runs directly like the surrounding test suite. Verified bin/fm-lint.sh, fm-brief, ask-user-authority, and task-delivery tests; all pass. git diff --check is clean

* no-mistakes(ci): Strengthened tests/fm-task-delivery.test.sh to behaviorally verify that real promotion and brief generation deliver byte-identical Definition-of-done blocks for all three modes. Verified tests/fm-task-delivery.test.sh, tests/fm-brief.test.sh, bin/fm-lint.sh, and git diff --check. The outer pipeline can now commit and attest the updated head

* no-mistakes(ci): Fixed promotion isolation instructions so any checkout other than the launched disposable worktree requires escalation, including another non-primary worktree. Updated behavioral coverage against the delivered promotion payload. Verified fm-task-delivery, fm-brief, fm-ask-user-authority, full fm-lint/ShellCheck, workflow lint, and git diff checks
)

* fix(bin): present complete Lavish board feedback as structured output

Give the Lavish adapter a read-only presentation so a handler sees every
annotation and the session-ending tag=message as its own field, instead of
grepping a truncated raw capture.

* no-mistakes(review): Preserve unquoted messages and prioritize captain prose

* no-mistakes(document): Document structured Lavish result reads

* no-mistakes(ci): Fixed Lavish `read` completeness: rows missing declared fields are excluded from presented items, counted as malformed, and force `complete: no`. Added behavioral regression coverage through the adapter interface. `bin/fm-lint.sh`, syntax checks, and focused valid/malformed read checks passed. The portable-serial failure was an unrelated secondmate cooldown timing flake
* fix(records): pair backlog transitions with the record that moves

Dispatch and completion each moved a task's physical record and its
backlog row as two independently timed steps, so a crash or a forgotten
follow-up could leave the two disagreeing: a record with no in-flight
row, an in-flight row with no owner, or a finished task still shown in
flight.

Fold each backlog transition into the script that performs the physical
change, under the per-task lock it already holds and before it reports
success. Dispatch moves the item to In flight after publishing the task
record and fails loudly, removing its provisional record, when that
transition cannot land. Completion records an authoritative close and
performs it before removing the record, so an interrupted cleanup can be
finished later, and its closing message now confirms what already
happened rather than instructing a future step.

Add a same-home reconciliation sweep to session start so a home that was
interrupted mid-transition settles its own books on restart, replaying a
recorded close and restoring an in-flight row it already owns a worker
for. It never reads or writes another home; the fleet snapshot and the
cross-home nudge stay as backstops.

Close records are validated before they are trusted: the file is read as
raw bytes and rejected outright when it carries a NUL or other control
byte, every field must be well formed and non-duplicated, the id must
match the record it was found under, the data location must resolve
inside this home, and each close argument must carry a permitted,
well-formed value. Writer and reader share one validator so a record
this home publishes always remains replayable, independent of locale.

Homes configured for a manual backlog, and homes with no backlog at all,
stay exempt and are unaffected.

* no-mistakes(review): Remove stale bootstrap migration helper invocation

* no-mistakes(review): Preserve pending closes and narrow signal deferral

* no-mistakes(review): Record close before destructive teardown

* no-mistakes(review): Refuse pending closes before creating resources

* no-mistakes(review): Guard relaunches and preserve cleanup warnings

* no-mistakes(review): Reject symlinked records and clarify cleanup guidance

* no-mistakes(review): Align dispatch eligibility and protect close replay

* no-mistakes(review): Unify exact task incarnation parsing

* no-mistakes(review): Render resolved configured backlog path

* no-mistakes(review): Harden transition path boundaries against symlinks

* no-mistakes(review): Validate lifecycle state before resource actions

* no-mistakes(review): Enforce transition tooling and continuous state locks

* no-mistakes(review): Consolidate same-home lifecycle file boundaries

* no-mistakes(review): Enforce canonical lifecycle containment and tooling contracts

* no-mistakes(review): Reject final-component lifecycle record symlinks

* no-mistakes(document): Document lifecycle record path boundaries

* no-mistakes(lint): Quote literal done tokens in atomicity tests

* no-mistakes(ci): Fixed all PR-caused CI failures: bootstrap now treats an absent state directory as an empty fresh home while retaining unsafe-state checks; nested remote secondmate retirement accepts records already removed with the retired home; teardown fixtures now provide valid data/manual-backend configuration; and the manual reminder assertion checks the configured absolute backlog path. Verified the reported tests, remote lifecycle E2E, backlog atomicity suite, Bash syntax, diff checks, and ShellCheck. The documented pre-existing captain-hold failure was intentionally untouched

* no-mistakes(ci): Fixed Behavior portable serial 3 by adding `od` to the teardown test’s lsof-free PATH fixture. The new close-record validator legitimately requires `od`; its omission caused teardown to fail before process-group cleanup and stall the shard. Verified the full `tests/fm-teardown.test.sh` suite passes, plus Bash syntax, ShellCheck, and `git diff --check`

* no-mistakes(ci): Fixed close replay to durably retain incomplete-cleanup evidence before removing task metadata. Subsequent retries now emit the reconciliation warning even after a backlog probe or close failure. Updated the behavioral regression and verified the full atomicity suite under stock macOS Bash 3.2, plus shellcheck and diff checks

* fix(records): validate record bytes without an uncurated tool

The byte validation added for close records and directory paths shelled
out to od. The spawn and teardown lifecycle runs under a curated command
set that deliberately excludes it, so on any restricted PATH the check
could not run, the data directory read as unresolvable, and dispatch and
cleanup refused - wedging the lifecycle rather than protecting it.

An earlier attempt made the failing test pass by adding od to that
curated set. That fixed the test to agree with the defect and quietly
widened the contract the fixture exists to pin, so it is reverted here.

Inspect the bytes with perl instead, which is already in the curated set
and already used in this repo for the same portability reason. The
emitted values are identical to od's, so the rejection semantics are
unchanged: NUL and other control bytes are still refused, legitimate
paths containing spaces or non-ASCII characters still round-trip, and
the check stays independent of the process locale.

The restricted-PATH teardown case now passes because the validator no
longer needs od, not because the fixture was loosened.

* no-mistakes(review): Enforce dispatch eligibility and atomic remote record publication

* no-mistakes(document): Document dispatch eligibility and cleanup alerts
…3342)

* fix: publish promote and Relay meta rewrites through contained replace

Bare mv still rewrote live task records in place, so a symlink meta could
be followed to a target outside state/. Route those field rewrites through
the shared publisher and drop the unused library aliases.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Refuse dangling symlinks during X metadata clear

* no-mistakes(review): Refuse unsafe metadata before follow-up and promotion side effects

* no-mistakes(review): Exercise dangling symlink refusal through clear helper

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…d#2877)

* fix(watch): absorb a turn-end whose pane churned since the previous poll

The watcher's "absorb a benign turn-end when the crew is provably working"
triage was structurally unreachable for any harness whose semantic busy state
has no verified source. crew_absorb_class only reports working for an actively
running no-mistakes step or an exact busy verdict, and bin/fm-crew-state.sh can
only answer unknown for such an adapter, so codex crewmates surfaced a signal
wake at every turn boundary with nothing to act on - a full supervisor drain,
inspect and acknowledge turn per worker turn, scaling with the number of workers
in flight and drowning the wakes that matter in identical noise.

Widen the proof rather than bound the wake rate. A wake carrying only bare
turn-ended markers is now also benign when the task's pane content changed since
the previous poll, compared against the same state/.hash-* marker the staleness
backbone already records and already trusts as liveness. That evidence claims no
harness semantics, so it fabricates no busy verdict an adapter has not earned,
and it needs no adapter cooperation.

Absorb stays evidence-driven in both directions. A wake naming any status file
keeps the strict proof, every captain-relevant verb still surfaces immediately,
and an unresolvable task, a missing prior hash, a failed or empty capture, or an
unchanged pane all surface exactly as before. The absorb defers rather than
swallows: a crew that has stopped renders nothing further, so its now-static pane
surfaces through the staleness backbone within a poll or two. Bounding the
surfacing rate instead would have suppressed genuinely stopped workers.

The derivation lives with the .hash-* marker format in bin/fm-watch.sh, which
owns it, and costs one bounded capture reached only for a no-verb turn-end whose
crew is not already provably working.

* no-mistakes(review): Captain, guard pane-churn absorption from collisions and secondmates

* no-mistakes(review): Captain, make watcher marker identities injective

* no-mistakes(review): Captain, isolate ambiguous legacy markers and restore Herdr sourcing

* no-mistakes(review): Captain, localize pane-churn collision guard

* no-mistakes(review): Captain, reject malformed pane-churn hashes

* no-mistakes(document): Document pane-churn turn-end evidence

* no-mistakes: apply CI fixes

* fix(watch): gate and bound the pane-churn turn-end absorb

Make the pane-churn form of positive work evidence opt-in per home and
bound how long it may defer one endpoint's bare turn-ends.

Absorbing a bare turn-end on pane churn is now reached only when the home
creates config/turnend-churn-absorb. The other two proofs read a verdict
the harness itself vouches for, while this one infers execution from
rendered bytes, so widening the absorb is a home's choice rather than a
default every fleet inherits. With the flag absent the predicate returns
on its first line and triage is unchanged.

Churn and pane staleness read the same pane, so neither can be the
other's only backstop. A pane that renders continuously never presents
the two consecutive identical hashes the staleness backbone needs, so an
unbounded churn absorb left a worker that had genuinely stopped behind
such a renderer with no path to surface at all. One endpoint's turn-ends
may now ride churn evidence for at most FM_TURNEND_CHURN_ABSORB_SECS,
tracked in state/.churn-since-*, after which the wake surfaces and the
window restarts. The bound is evaluated before any .stale- state is
touched, so a wake that surfaces there leaves the staleness backbone's
own classification alone.

Covers both with behavioral tests: the same churning fixture that absorbs
with the flag surfaces and queues without it, and a spent deferral window
surfaces and restarts. The four existing safety guards now run with the
flag enabled so they keep proving their specific guard.

* no-mistakes(review): Fail closed on invalid churn deferral state

* no-mistakes(review): Validate persisted churn deadlines before arithmetic

* no-mistakes(review): Make churn deadlines transactional and bounds safe

* no-mistakes(review): Compose turn-end evidence per task from one snapshot

* no-mistakes(review): Restore strict turn-end fallback guards

* no-mistakes(document): Clarify pane-churn supervision documentation

* no-mistakes(lint): Fix watcher arithmetic lint issues

* no-mistakes: apply CI fixes

* no-mistakes(document): Clarify pane-churn fail-closed documentation

* fix(bin): prioritize active pipeline-owned crew runs (kunchenguid#3194)

* fix(bin): bind the live pipeline-owned run instead of a superseded failed row

fm-crew-state.sh bound a superseded FAILED no-mistakes run to a task instead
of the LIVE replacement run: the live run's pipeline-owned lane head is not a
git object in the task worktree, so head-equality attribution rejected it and
the coarse runs-list fallback silently continued past the RUNNING row onto an
older failed row whose head equalled the stale worktree HEAD. The home summary
then flipped invalid and Bearings hid the home's live work (F10).

Attribution precedence now follows the daemon's own identity:
- An ACTIVE run for the task's branch binds without head equality while
  branch_sync.state is pipeline_owned (fm_nm_run_is_pipeline_owned_active);
  the pipeline owning the branch is itself the attribution.
- A genuinely failed run with no later run on the branch still reports failed
  through the unchanged head-equality path - real failures are not hidden.
- In the coarse runs scan, an unresolvable head is unknown attribution and
  stops the scan (fm_nm_head_resolvable) instead of falling through to an
  older row; a resolvable-but-mismatched head keeps the historical
  reused-branch skip.

The exemption never applies to a terminal run and requires pipeline_owned
specifically, both pinned by negative-control tests. Fixture shape verified
against the live incident run's real axi status output.

* no-mistakes(document): Updated run-attribution documentation ownership

* no-mistakes(review): Captain, make watcher marker identities injective

* no-mistakes(review): Captain, localize pane-churn collision guard

* no-mistakes(review): Compose turn-end evidence per task from one snapshot

* no-mistakes(review): Restore strict turn-end fallback guards

* no-mistakes(document): Align pane-churn watcher documentation

* no-mistakes(ci): Captain, fixed the flaky cooldown boundary test by freezing its executable clock. The failure reproduced before the fix and passed five consecutive full-suite runs afterward. Extended ShellCheck passed; full lint stopped because actionlint 1.7.12 is not installed

---------

Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
* fix(bin): add a safe owner for custom-check retirement

Agents were improvising rm of check files with unset STATE/ID, which wedges
headless panes. Unregister validates the id and state directory first.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Refuse explicitly empty custom-check state overrides

* no-mistakes(document): Document custom-check retirement safety contract

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…o dedicated scripts (kunchenguid#3221)

* Add quota exhaustion detection and safe fallback helpers

- bin/fm-procevent-quota.sh: generic procevent adapter that arms a
  recurring quota-axi --json poll and wakes firstmate when a tracked
  provider's effectivePercentRemaining drops below a threshold or its
  runway.status becomes exhausted_now.
- bin/fm-quota-choose.sh: worker-side helper that picks the first ranked
  harness:model candidate with positive effectivePercentRemaining.
- AGENTS.md and .agents/skills/quota-array-dispatch/SKILL.md: document
  the new helpers and the mid-task quota-exhaustion wake path.
- tests/fm-quota-choose.test.sh: unit tests with a mocked quota-axi JSON
  source.

* no-mistakes(review): Fix quota polling and scope bounds

* no-mistakes(review): Enforce safe default quota selection

* no-mistakes(review): Handle decimal quota values safely

* no-mistakes(review): Fail closed on invalid quota inputs

* no-mistakes(review): Reject empty quota candidate segments

* no-mistakes(review): Harden quota parsing and timeout ownership

* no-mistakes(review): Reuse captured quota snapshots consistently

* no-mistakes(review): Match quota using explicit candidate providers

* no-mistakes(review): Centralize fail-closed quota schema validation

* no-mistakes(review): Reject out-of-range quota percentages

* no-mistakes(review): Validate quota runway status enum

* no-mistakes(review): Tighten quota scope and status contracts

* no-mistakes(review): Preserve unknown quota and exact product bounds

* no-mistakes(review): Preserve provider-level unknown quota

* no-mistakes(review): Reuse canonical verified harness validation

* no-mistakes(document): Document mid-task quota handling

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(docs): restore default routing contract, keep quota helper optional

Restore the AGENTS.md section 4 always-loaded routing paragraph the PR
had deleted, so the standing TOON-first intake, spendPriority ranker,
every-candidate accounting, and load-trigger contract stay exactly as
before this PR. The mid-task quota wake is optional and must not alter
default routing.

Restore the quota-array-dispatch skill ownership line to section 4 as
the always-loaded intake boundary owner; keep the worker-side helper
section as an addition only, without rewiring ownership or load
triggers to section 13.

* fix(bin): use harness-keyed quota matching in optional helper

Revert fm-quota-choose.sh from harness:provider:model tuples back to
harness:model candidates with harness-keyed provider matching, per the
resolved ask-user finding. The helper is optional; authoritative
multi-provider routing (provider discovery from the harness catalog and
quota matching by that explicit provider) stays owned by AGENTS.md
section 4 and the quota-array-dispatch skill intake procedure, not the
helper.

Document the multi-provider limitation in the helper header and the
quota-array-dispatch skill: the helper maps each harness to one primary
provider family only, so a candidate whose established provider differs
from that primary family is checked against the wrong quota row. Use it
only when the brief fixed the candidate order and every candidate's
provider is the harness's primary family.

The helper still consumes one already-captured default-TOON or JSON
snapshot via stdin or --snapshot and never calls quota-axi itself, so
it selects from the same quota state as the intake.

* no-mistakes(review): Fix Muse quota mapping and helper contract docs

* no-mistakes(review): Reject known-empty quotas and map quota tests explicitly

* no-mistakes(review): Preserve unmeasured candidates and enforce snapshot reuse

* no-mistakes(review): Fix quota retirement and dependent regression coverage

* no-mistakes(review): Accept zero-row quota TOON snapshots

* no-mistakes(review): Enforce quota semantics status consistency

* no-mistakes(review): Veto dispatch on any exhausted applicable scope

* no-mistakes(review): Record exhausted quota scope in wake details

* no-mistakes(review): Fix quota help and control dependency coverage

* no-mistakes(review): Decode quoted TOON fields and document quota wakes

* no-mistakes(review): Validate zero-row TOON and map timeout coverage

* no-mistakes(review): Reject multi-value JSON and malformed TOON envelopes

* no-mistakes(review): Validate complete nonzero TOON envelopes

* no-mistakes(review): Accept producer-shaped quota TOON envelopes

* no-mistakes(review): Support empty quota arrays and validate counted rows

* no-mistakes(review): Harden TOON completion, scopes, and quoted fields

* no-mistakes(review): Preserve unknown-headroom exhaustion and reject trailing fields

* no-mistakes(review): Allow unknown headroom under known semantics

* no-mistakes(review): Reject noncanonical quota identities

* no-mistakes(review): Preserve empty quota polling and validate attention identities

* no-mistakes(review): Reject noncanonical provider watches

* no-mistakes(review): Validate all candidates before quota selection

* no-mistakes(document): Correct quota helper safety documentation

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes
* fix(bin): keep typed Lavish comments when an element is also annotated

read preferred element text over prompt, so an annotate-and-comment
item dropped the captain's words. Surface prompt as its own field.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Filter non-comment prompts from Lavish reader output

* no-mistakes(document): Clarify Lavish comment presentation contract

* no-mistakes(ci): Fixed Lavish reader comment provenance: non-choice prompts are now emitted even when identical to element text. Added observable regression coverage for identical selector+comment input while retaining pure annotation/message coverage. Reader cases, bash syntax, and diff checks pass. Full fm-procevent suite stops earlier at unrelated “reconcile never claimed” setup failure

* no-mistakes(ci): Fixed duplicate pure-annotation prompts by emitting `prompt:` only when it differs from captured element text. Updated behavioral coverage for selector+comment, pure annotation, and pure message cases. Focused reader regressions, syntax checks, and diff checks pass. Full suite remains blocked by the pre-existing “reconcile never claimed the registered source” failure

* fix(bin): always emit Lavish comments and use real annotation fixtures

Stop inferring comment provenance from prompt==text. Real pure
annotations have no prompt, so always-emit does not duplicate.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…uid#3420)

* Fix public-followup register crashing on empty lock arrays under bash 3.2.

bash 3.2 with set -u treats "${arr[@]}" on an empty array as unbound, so the first register in a fresh home aborted before taking the registry lock.
The empty-lock regression also runs under the existing stock macOS Bash CI lane so pre-fix code would fail there.

* no-mistakes(document): Document stock Bash registration coverage

* no-mistakes(ci): Pinned the stock macOS Bash CI lane to tasks-axi@0.2.5, eliminating dependency drift. Verified workflow YAML parsing, git diff checks, and the focused regression under /bin/bash 3.2.57 with tasks-axi 0.2.5

* no-mistakes(ci): Fixed the flaky portable CI test: it treated exited zombie processes as live because `kill -0` succeeds for zombies. The watcher and descendant assertions now check process state and regard zombies as exited. Verified `tests/fm-pr-check-security.test.sh`, ShellCheck, `git diff --check`, and the focused Bash public-followup regression
* fix(herdr): isolate server launch environment

* no-mistakes(review): Clear inherited supervision model from Herdr launches

* no-mistakes(document): Document Herdr server launch environment isolation
* fix: surface inbound Relay attachments to the responding agent

A Discord support thread's screenshots were never seen by the agent
handling the mention. The relay delivered them and the poll stashed
them: the reporter's images arrived on the `thread_starter` entry of
`in_reply_to_chain` while the mention's own media list was empty. The
gap was in the responder's playbook, which enumerated a fixed field
list (`request_id`, `text`, `in_reply_to`, `in_reply_to_chain`) and so
made every other field, attachments included, invisible.

Fix it where the gap is, in prose:

- Read the complete payload object rather than a fixed field list, so
  media and later relay fields are never skipped again.
- Fetch and view attached media with the agent's own tools, on the
  mention and on every chain entry, and call out the common shape where
  only the thread starter carries the screenshots.
- Restrict those fetches to known-good platform media hosts over https
  (Discord: cdn.discordapp.com, media.discordapp.net,
  images-ext-1.discordapp.net, images-ext-2.discordapp.net; X:
  pbs.twimg.com, video.twimg.com), report a blocked host instead of
  working around it, and treat everything fetched as untrusted public
  input on the same terms as the surrounding thread text.

The poll stays out of it and downloads nothing, so no third-party bytes
are pulled on the polling path.

The new test pins the contract the playbook depends on: a mention in the
incident's shape, with an empty top-level media list and screenshots on
the thread starter, must reach the inbox with the payload intact and its
media URLs unfetched.

* no-mistakes(review): Preserve media authority and enforce poll-only fetching

* no-mistakes(document): Clarify Relay attachment safety prose
)

* Defer inactive startup reconciliation

* no-mistakes(review): Queue deferred inactive reconciliation diagnostics durably

* no-mistakes(review): Require worker phases to cover startup requests

* no-mistakes(review): Make diagnostic wakes safely acknowledgeable

* no-mistakes(document): Document deferred startup phase coverage
* fix: bound status presentation lock waits

* no-mistakes(review): Distinguish malformed presentation locks from live contention

* no-mistakes(review): Bound no-ack drain queue lock acquisition

* no-mistakes(document): Document bounded presentation-lock drain behavior

* no-mistakes(lint): Annotate bounded lock output global

* no-mistakes(ci): Added deterministic regression coverage for successful bounded-lock acquisition after live contention, verifying helper-to-caller PID ownership handoff and caller release. Verified with bash syntax checks, git diff checks, and the full fm-wake-queue test suite
* fix(relay): close a public loop whose work lives in a remote secondmate home

A public-followup loop bound to a REMOTE secondmate could never be closed.
`clear_public_followup_link` (bin/fm-public-followup.sh:701) required an
absolute recorded `work_home_path` for a `secondmate:*` work home, but a remote
route has no local path on this machine, so registration records that field
empty (bin/fm-public-followup.sh:291). Every close ran that clear first, so
`retire` died with "could not clear the legacy X link ... retained for
reconciliation" forever, and `deliver` posted the public reply and then stranded
the loop at `posted`. `--force` never covered that step.

The clear now goes to the remote home over that route's SSH transport, running
`fm-x-followup.sh --clear <work-id>` through `bin/fm-on.sh`. The route is decided
from `data/secondmates.md` before any local path is consulted, so a same-named
local directory can never stand in for a remote home, and registrations already
on disk retire without needing a new field. `fm-on.sh` passes ssh's status
through, so 255 stays the established "delivered but completion unknown" result
this codebase already reconciles: the close is refused, the registration and the
remote link are left exactly as they were, and the message names the unknown
completion instead of claiming a definite failure.

Local secondmate and `main` work homes are untouched, and `--force` still
governs only the unresolved-obligation refusal.

Three regression cases drive a remote route end to end, faking only the ssh
binary at the FM_SSH_BIN seam and then running the real remote entrypoint
against a local checkout, so the clear that must reach the remote home actually
happens there.

* no-mistakes(review): Guard remote link clears by request identity

* no-mistakes(review): Fail guarded clears on unreadable remote state

* no-mistakes(review): Reject guarded clears on non-writable remote state

* no-mistakes(review): Allow no-link retirement in non-writable remote state

* no-mistakes(document): Correct public-followup verification guarantee count

* no-mistakes(ci): Fixed the guarded link-clear race by ensuring absence is decided under the metadata lock whenever publication is possible. Added a behavioral concurrency regression test. Verified with fm-x-mode and fm-public-followup suites, Bash syntax checks, diff checks, and bin/fm-lint.sh

* no-mistakes(ci): Fixed the guarded link-clear race by refusing an unlocked absence decision when a publisher already owns the metadata lock in a non-writable directory. Added a behavioral concurrency regression test. Verified with fm-x-mode, fm-public-followup, syntax/diff checks, and fm-lint

* no-mistakes(ci): Fixed the guarded-clear race by refusing all guarded clears when the metadata parent is non-writable, including apparent link absence. Added a behavioral regression with a publisher waiting to create the lock, updated remote-retirement expectations and verification docs. Passed fm-x-mode, fm-public-followup, fm-lint, documentation audience, Bash syntax, and diff checks

* fix(relay): bound the guarded remote link clear so it refuses instead of hanging

The guarded clear checks that the remote state directory is writable before
taking the metadata lock, but that check cannot close the window: the parent can
turn non-writable between the check and lock creation, and a lock held by a live
holder is indistinguishable from that at the acquire. `fm_lock_acquire_wait` is
an unbounded `while ! try; do sleep 0.1; done`, so either case retried forever
and `deliver` or `retire` wedged with nothing reported, instead of returning the
retained-for-reconciliation refusal the guard exists to produce. This path runs
unattended over the secondmate transport, where a wedge is worse than either
outcome the guard defines.

The guarded clear now acquires through `fm_lock_acquire_wait_bounded`
(FMX_LINK_CLEAR_LOCK_TIMEOUT, default 10 seconds) and refuses on timeout through
the existing failure path. Unguarded local callers keep the ordinary unbounded
wait, so local behavior is unchanged.

The bounded primitive's header no longer claims presentation-only scope, since
this is a second authorized caller; nothing else in the shared lock
infrastructure changed.

The regression holds the metadata lock with a genuinely live process while
leaving the state directory writable, so the refusal can only come from the
bound and never from the writability precondition. Against the unbounded wait it
does not terminate at all; with the bound it refuses, retains the registration,
writes no receipt, and leaves the remote link untouched.

* no-mistakes(review): Harden lock-timeout regression with independent deadline

* no-mistakes(review): Restore no-op guarded clears on read-only state

* no-mistakes(document): Clarify remote public-followup cleanup contract
)

* fix(bin): resolve process-event state roots before validating them

The process-event module validated the caller's spelling of a home's state
root instead of the directory it operates on: it required the supplied path
to equal its own lexical normalization, which rejects any path reached
through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks,
so an operator home under either could never claim a source. Reconcile still
reported the runner started, while the detached runner died writing "cannot
claim source" to the discarded stderr, and the source silently never fired.

Resolve the state root to its physical directory once, then apply the
existing private-directory validation to that resolved directory and derive
every path, recorded claim identity, and later confinement check from it.
This keeps the confinement contract for the directory actually operated on
rather than only for callers that already spelled it physically, and removes
the window where an ancestor symlink could be repointed between check and
use. Homes already spelled physically behave identically.

This was the single cause of both deterministic macOS failures in
tests/fm-procevent.test.sh ("reconcile never claimed the registered source")
and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not
produce an outcome"). The new case pins the behavior with an explicit
symlinked-ancestor home, so it fails without the fix on any platform rather
than only where the temp root happens to be a symlink.

* fix(bin): pin the external capture staging boundary to its physical path

The extension capture path pinned its registry staging boundary by comparing
`pwd -P` against the caller-spelled registry directory, so a home reached
through a symlinked ancestor still refused to start an extension-backed
source after the state root itself resolved correctly. That left such a home
half working: built-in sources ran while external ones failed.

The staging preparer now prints the physical registry directory it validated,
matching the inbox and reservation preparers beside it, and the start path
pins on that returned path. The new end-to-end case drives the shipped
file-signal package from a symlinked home spelling.

* no-mistakes(review): Propagate canonical process-event state roots

* no-mistakes(review): Propagate canonical state to process-event adapters

* no-mistakes(document): Document physical process-event state roots
…kunchenguid#3312)

* fix(pi): persist captain outcomes visibly

* no-mistakes(review): Recover captain outcomes after cold-start lock acquisition

* no-mistakes(document): Document cold-start captain-outcome recovery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* no-mistakes: apply CI fixes

* fix(pi): process captain outcomes through a sequence-keyed turn

PR kunchenguid#3312 made every captain-facing supervision outcome a durable, exact-once
visible transcript entry with the read cursor advancing only after that entry
exists. That is the display half of the delivery contract. Left alone it turns
a probabilistic silent loss into a deterministic one: the captain sees an
anchor line, and firstmate never acts, because nothing opens a turn and
nothing records whether main ever processed the outcome.

The 2026-08-31 timeline showed the two shapes this must survive on the
previous hidden-turn path: seven delivered decision outcomes each answered by
an empty assistant message (cursor advanced, no retry, unanswered for close
to three hours), and two answered by an unrelated prior reply. Both happened
because delivery advanced the cursor at enqueue and accepted whatever the
next assistant message was.

Add the processing half on top of the persistence half:

- bin/fm-branch-outcome.sh keeps a processed marker separate from the read
  cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It
  only advances through an explicit sequence-bound acknowledgement, never
  past the read cursor and never backwards; an absent marker reads as zero
  and `processed-init` migrates delivered history once so an upgraded home
  is not re-presented its past.
- After the visible entry for a captain outcome exists, the extension hands
  every still-unprocessed captain row to main as one hidden, typed
  `fm-branch-process` request listing each `[seq N] task: summary`, opening
  exactly one main turn. Main closes it only by calling the new
  `fm_branch_processed` tool with the highest sequence listed. An unrelated,
  empty, or paraphrased answer leaves the sequence open, and the same request
  is presented again at the end of the next main run and at session start.
  The first two presentations of a sequence set open a turn of their own;
  after that the request rides the captain's next prompt so an ignored
  request cannot loop, and a session replacement resets that budget.
  Routine outcomes stay turn-free.
- The regressions cover exactly those incident shapes against the real store
  scripts: an empty answer and an unrelated prior answer neither advance the
  marker nor stop re-presentation, the acknowledgement is refused beyond the
  read cursor and outside lock ownership, a partial acknowledgement keeps the
  newer sequence open, and kunchenguid#3312's own assertions now forbid an unkeyed turn
  rather than any turn. The store suite pins the marker's bounds and the
  migration; the real-SDK guard for appendEntry persistence and model
  exclusion is unchanged.

Docs move the protocol from "no model turn" to "one sequence-keyed processing
turn closed only by its acknowledgement", and the verification record carries
the dated run against Pi 0.84.4.

* no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements

* no-mistakes(review): Harden outcome state validation and request pacing

* no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores

* no-mistakes(review): Validate canonical mark-read cursor state

* no-mistakes(review): Guard cursor advancement against corrupt processed state

* no-mistakes(review): Bind acknowledgements to active processing requests

* no-mistakes(review): Reset pacing when processing sequence membership changes

* no-mistakes(review): Enforce silent outcome invariants at storage boundary

* no-mistakes(document): Document hardened captain outcome processing contracts

---------

Co-authored-by: kunchenguid <kun@kunchenguid.com>
…3481)

* feat: bound Bearings remote ledger collection

* no-mistakes(review): Clarify default remote-ledger collection behavior

* no-mistakes(review): Detach reconcile delivery from watcher loop

* no-mistakes(review): Enforce bounded snapshot and request captures

* no-mistakes(review): Bound legacy summary capture before parsing

* no-mistakes(review): Bound primary remote ledger captures

* no-mistakes(document): Correct snapshot and reconcile documentation

* no-mistakes(lint): Fix ShellCheck quoting in bounded collector

* no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks

* test: await reconcile request retirement

* no-mistakes(review): Avoid empty reconcile queue process churn

* no-mistakes(review): Read ledger summaries from immutable snapshots

* no-mistakes(review): Reject multi-document home ledger streams

* no-mistakes(review): Coalesce durable reconcile requests per target

* no-mistakes(review): Unify reconcile keys and reject snapshot streams

* no-mistakes(review): Key reconcile requests by stable target ID

* no-mistakes(document): Document per-target reconcile request coalescing

* no-mistakes(lint): Remove unused snapshot summary file variable

* no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass

* no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks
* fix(ci): rebalance the portable serial shards on measured durations

The "Behavior portable serial 3" shard ran 17-20 minutes against its
20-minute job cap and intermittently timed out seconds after a passing
test, on branches and on main alike.

Shards are packed longest-processing-time from per-script duration hints,
and those hints were last measured on 2026-08-21 at 116 scripts. The lane
has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had
no hint at all and fell back to the 20 s default, and several existing
hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured,
fm-public-followup 36 s vs 197 s). The partition therefore looked
perfectly balanced in hint space, 734.6 s per shard, while really running
11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the
tests asserted, stayed normal throughout and hid it.

Refresh the hints from the timing artifacts of three green runs, taking
the slowest measurement of each script so the balance holds on a slow
runner, and split the lane across five shards instead of four. Replayed
against those runs' real per-script durations the worst shard is now
12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's
wall clock drops from ~20 to ~12.5 minutes.

Bound the drift that caused this rather than relying on the hints being
refreshed by hand: the coverage guard now reports the unmeasured share as
serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT,
which leaves room for newly added tests while making a stale table fail
the guard instead of silently pushing one shard into its cap.

No test changes what it asserts and no test stops running; only the
partition across shards changes.

* no-mistakes(document): Clarify conservative shard timing aggregate
…uid#3491)

* fix(pi): fall back after settled branch errors

* no-mistakes(review): Detect provider errors across prompt compaction

* no-mistakes(review): Preserve in-flight branch state across selection changes
* fix(pi): recover supervision branch after cooldown

* no-mistakes(review): Defer branch recovery until prompt settlement

* no-mistakes(document): Clarify supervision cooldown recovery contract
* refactor: remove legacy remote summary reads

* no-mistakes(document): Document ledger-only snapshot reads

* no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass

* no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean

* no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux
…henguid#3498)

* fix(pi): rearm watcher after session replacement

* no-mistakes(review): Queue actionable closes across Pi session replacement

* no-mistakes(review): Stop replacement arm when handoff persistence fails

* no-mistakes(review): Preserve actionable wakes through branch and late child races

* no-mistakes(review): Surface late handoff failures without crashing Pi

* no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens

* no-mistakes(review): Retry stale deliveries and release settled claims

* no-mistakes(review): Distinguish branch settlement and retry handoff cleanup

* no-mistakes(review): Deduplicate persistent handoff cleanup alerts

* no-mistakes(review): Acknowledge watcher follow-ups only when consumed

* no-mistakes(review): Persist idle follow-ups until agent consumption

* no-mistakes(review): Preserve pending outcomes when handoff persistence fails

* no-mistakes(review): Arm replacement before awaiting prior delivery settlement

* no-mistakes(review): Adopt pending handoffs after lock reclamation

* no-mistakes(review): Prevent stale generations from adopting replacement handoffs

* no-mistakes(review): Scope replacement handoffs by watcher state

* no-mistakes(document): Clarify replacement handoff documentation

* no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks

* no-mistakes(review): Update branch settlement tests and preserve chunked outcomes

* no-mistakes(document): Document watcher-owned replacement handoffs

* no-mistakes(document): Verify replacement handoff documentation

* test(pi): cover watcher-owned branch fallback

* no-mistakes(document): Refresh watcher-owned fallback documentation
…d#3495)

* fix(bin): resurface terminal statuses lost after branch handling

* test(watch): canonicalize process-event fixture homes

* no-mistakes(review): Index branch outcomes by causal status position

* no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses

* no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics

* no-mistakes(review): Keep unclassifiable oversized statuses silent

* no-mistakes(document): Document lost-wake outcome backstop

* no-mistakes(document): Update outcome backstop documentation

* no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally

* no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes

* no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift

* no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state
…id#3503)

* fix(bin): deliver typed terminal results from remote work homes

A public commitment whose work is bound to a REMOTE secondmate home could
never receive its typed terminal result. `fm-public-followup.sh brief`
printed an emit command carrying this home's own absolute path and this
checkout's own script path, neither of which exists on the machine the
worker runs on, so the worker had nothing it could write to that the
owning home would ever read - and `consume` kept finding nothing while
the promise stayed open.

The brief is now route-aware: for a remote work home it prints that
route's own code root and home with `--stage-in`, so the typed event is
staged in the home where the work actually runs, and the closing
paragraph names the owning home as the one on the other machine instead
of pointing at the path above it. The owning home collects those staged
results over the same SSH route it reaches that secondmate on, because
the transport only runs outbound: `consume` pulls them into its own
inbox and reconciles them exactly as it reconciles a local report.
Collection is non-destructive until the result is durably held, so a
dropped connection cannot lose a terminal result, and a route that could
not be reached is named in `consume`'s output with the promise left open
rather than reported as an empty inbox.

A local work home is untouched: the brief still prints `--home` with this
home and this checkout's script, and the event still lands directly in
this home's typed terminal-result inbox.

This is the emit-side counterpart of the retire/clear fix in kunchenguid#3479 and
reuses the remote-route resolution that landed with it. Reconciling a
loop bound to a remote route now reaches that route, so the existing
remote cases drive `consume` through the same faked transport their
other steps already use.

* no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes

* no-mistakes(review): Fail collection when remote outbox is unreadable

* no-mistakes(review): Surface reassigned remote routes during empty collection

* no-mistakes(review): Fail remote collection on invalid registrations

* no-mistakes(review): Reject unsafe registration entries during remote collection

* no-mistakes(review): Restore healthy empty remote collection behavior

* no-mistakes(review): Skip remote collection for delivered registrations

* no-mistakes(review): Skip delivered registrations before route validation

* no-mistakes(document): Document remote follow-up collection semantics
kunchenguid and others added 29 commits September 3, 2026 12:00
* fix: copy PR URLs from records or abstain, never assemble them

Supervision reported a plausible but dead PR link three times because its
prompt demanded a full https:// URL at a moment when only a PR number was
observable, so the model assembled an owner/repository from memory, and the PR
check then accepted that URL and wrote it into the task record, after which the
model kept defending its own tool-endorsed guess over the worker's real link.

Three changes close that chain without any live forge lookup, so private
forges are treated exactly like public ones:

- bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy
  or abstain" section requires a URL to be copied verbatim from a durable
  record (the done: PR <url> status line, pr= metadata, or the backlog note),
  forbids assembling owner, repository, host, or number from memory, and has
  the branch report only the identifier it actually holds when no record names
  the URL yet, leaving the PR check unarmed until the worker's ready line
  arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for
  main in place of the bare full-URL mandate.

- Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full
  https:// URL wherever a PR is mentioned - status line, terminal, or summary -
  never a bare "PR 108", so the link is in view as early as the number is.

- bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that
  the task's own done lines contradict, printing both spellings; a log naming
  no URL still records the argument as before. fm_pr_status_ready_urls in
  bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches
  bin/fm-pr-merge.sh, so nothing merges under a contradicted URL.

Tests cover the offline refusal with zero side effects, the recorded spelling
being accepted, markdown-wrapped and punctuated URLs, working lines not
counting, the merge wrapper propagation, a self-hosted merge request with no
forge call, the prompt carrying the rule, and the brief carrying the worker
rule.

* no-mistakes(review): Remove stale PR URL enforcement

* no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check`
…uid#3661)

* fix(bin): disable Claude's feedback-draft flow for fleet-launched agents

Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched
Claude crewmate and secondmate, so /bug and /feedback never queue or
submit a bug report on the captain's behalf. feedbackDrafts is the
documented settings key (Claude Code changelog 2.1.247); the
per-launch CLI flag never touches the captain's global settings.json.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts

* no-mistakes(document): Fix Claude feedback documentation formatting

* fix(bin): layer both feedback-draft controls for defense in depth

The prior --settings-only fix can be overridden by a managed Claude
settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0
alongside --settings '{"feedbackDrafts":"off"}': either control alone
disables the SendFeedback tool, so a managed override of one still
leaves the other in force.

Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3

* no-mistakes(document): Document Claude feedback-draft suppression ownership
…guid#3662)

* perf(tests): admit three more families to concurrent validation

The three families that `docs/fm-test-isolation-proof.md` recorded as refused
were not refused for concurrency. Each blocker was a test that decided a
property by wall clock, or a script filed where it cannot run. Fixing those
three things admits all three families and recovers 28.6 minutes of local
validation with no assertion removed or weakened.

- `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the
  handoff, sleeping a fixed second, then delegating the move to the real
  binary. Nothing ever killed the fake, so on a host slow enough for the case's
  next assertions to take longer than a second, the orphan woke and completed
  the very move the case requires left undone, and recovery then failed with
  `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs
  during the injected crash showed exactly that, the item moving one second
  after the crash. All four crash injections in the file now go through a new
  `fm_fake_crash_injector` shim that signals the target and returns only once
  it is observably gone, and the pre-move fake never delegates the move at all.

- `tests/fm-session-start.test.sh` proved the startup digest does not block on
  a slow current-state read by timing the whole digest against a fixed
  eight-second sleep, which a loaded host exceeds without the property being
  violated. It now holds that read open until the case releases it and asserts,
  the moment the digest returns, that the read has not finished. A digest that
  waited would wait indefinitely rather than for an interval a slow host can
  out-run, so the assertion is stronger than the bound it replaces. Its scan
  budget moves to the maximum, because the old value left two seconds of margin
  over the fixed sleep and measured the host rather than the deadline that
  `tests/fm-inactive-reconcile.test.sh` owns.

- `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all,
  which put it in the portable serial lane, where Linux CI gate-skips it: that
  real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and
  the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its
  opt-in variable and moves to `live-harness-optin`.

The 28 remaining ungrouped scripts become an enumerated `standalone` family
instead of admitting `unclassified` itself. `unclassified` is the family map's
`*)` arm, so admitting it would silently grant concurrency to every test added
afterwards, which is exactly the population with no proof. A new test still
lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh`
covers that split behaviorally.

Each family passes two consecutive four-worker proofs with zero failures. On
the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap`
756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s
recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock
against 121 minutes of summed script time.

* no-mistakes(document): Refresh concurrent validation and shard documentation

* no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2
* feat(brief): structure no-mistakes ask-user escalation as event + snapshot file

Crewmates escalating a no-mistakes ask-user gate now report one status
event naming every finding id plus a snapshot file holding the gate's
axi finding records verbatim (id, severity, file, line, description,
authority), using the same shape even for a single finding. The status
line never paraphrases. The format is defined once in fm-dod-lib.sh and
rendered into both the scout and ship rule 6 in fm-brief.sh, so a
promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets
the identical contract as a freshly-spawned no-mistakes ship worker.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei

* no-mistakes(review): Preserve ask-user escalation output contract

* no-mistakes(review): Align escalation format test expectation

* no-mistakes(review): Scope ask-user escalation instructions correctly

* no-mistakes(review): Remove ask-user from generic decision rules

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* fix(bin): require a self-sufficient no-mistakes intent

A no-mistakes worker's --intent is only as useful as the string it
passes. PR kunchenguid#3604 shipped with an intent that was only "do 1, 2, 3, 7
from the report": the real contract lived in a private scout report and
never reached --intent, so nobody holding that string plus the codebase
could have derived the specification.

This is pure instruction at the contract's one owner; no spawn-side or
promotion-side check is added.

- bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now
  states that the --intent string must be self-sufficient (the string
  plus the codebase reconstructs roughly the same specification) and
  tells the worker to write the substance of any report, decision, or
  PR the captain's intent refers to into --intent rather than the
  pointer, while Firstmate build instructions and the worker's own
  decisions still stay out. The spawn-time overlay points back at that
  rule so its "supersedes" wording cannot cancel it, and the header's
  owner statement carries the rule.
- AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to
  include the substance of referenced material when filling
  ## Captain's intent, and section 11 points at the owner of the rule.
- tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the
  rendered brief and launch contract carry the rule.

Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim

* no-mistakes(document): Replace incident-specific intent test commentary
* Speed local fleet snapshot composition

* no-mistakes(review): Stabilize task inventory during concurrent snapshot composition

* no-mistakes(document): Document local snapshot observation concurrency

* no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks

* no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks

* no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks

* no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky

* no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks

* fix(snapshot): keep live observations generation-coherent

* no-mistakes(review): Keep secondmate observations generation-bound without copying reports

* no-mistakes(document): Document generation-coherent snapshot observations

* test(bearings): measure local read overlap instead of wall-clock budget

The large-local-snapshot regression asserted that a whole snapshot
composed in under five seconds. That bound measures how loaded the host
is, not whether the per-task reads actually overlap, so it failed
intermittently on a contended machine: one run in six on a box at load
16-20, landing exactly on the five second boundary.

Time a serialized run and a concurrent run of the same workload instead
and require the concurrent one to save at least two seconds. Both runs
pay the same composition overhead, so the difference isolates the
overlap this change delivers. Five one-second reads serialize into five
seconds and overlap into about one, and re-serializing the reads
collapses the saving to roughly zero, so the assertion still fails
loudly if the concurrency regresses.

Also bump the pinned Bearings test count to 48, since rebasing onto the
current default branch picked up its captain-hold test.

* no-mistakes(review): Restore JSON-derived decision flags

* no-mistakes(review): Unify status-derived snapshot observations

* no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass
* fix(bin): stop the supervision branch's stale-ack and ghost-report loops

Clean-slate implementation of the four authorized recommendations from the
supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal
form, superseding PR kunchenguid#3604:

- fm_branch_report refuses a task the wake being handled never named. The
  extension fixes the reportable task set from the eligible rows before each
  prompt (signal and stale rows resolve to their tasks, a heartbeat allows any
  task with a live record, fleet is always allowed), so a report typed from
  memory about a task whose records teardown already removed is never stored
  or delivered.
- An acknowledgement that consumes nothing says "nothing was acknowledged
  through N" and prints the exact --ack-through / --recovery-generation
  command for the current presented wake, instead of "re-run the drain",
  which re-fed the same stale acknowledgement in a loop.
- bin/fm-guard.sh no longer tells the branch actor to drain queued wakes
  while it is handling them; it names the granted rows instead.
- Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and
  descendants; the index rebuild and the append-side index write both skip a
  task with neither a live record nor a status log, so the branch's report of
  a teardown it just performed is stored without recreating the index.

No new locking, no spawn-generation binding, and no retired-task refusal: the
branch can still report the outcome of a task it just tore down, and the
teardown test now proves that path end to end.

* fix(bin): narrow the branch report scope and guard silence to the minimal form

Apply the four review decisions on the clean-slate branch:

- A signal or stale prompt may report only the tasks its own rows resolve
  to; fleet is refused there too. A heartbeat review is not scoped by task
  at all, so the extension no longer tracks live task records and refuses
  nothing by task id during a fleet review.
- The outcome-index rebuild no longer skips retired tasks; the append-side
  skip alone keeps a torn-down task's index from being recreated.
- bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor
  instead of printing a replacement note.

* no-mistakes(document): Align supervision docs with scoped wake handling
* Fix fleet snapshot large JSON transport

* no-mistakes(review): Captain: file-back fleet snapshot transport safely

* no-mistakes(review): Captain: file-back parent summary aggregation

* no-mistakes(ci): Rebased the PR's three commits onto f4d7875 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor
…nguid#3681)

* fix(bin): recognize active pipeline fix rounds with unfetched run heads

A no-mistakes fix round advances the run head beyond the submitted head,
and the pipeline commits in its own checkout, so the task copy never
receives the new commit object. fm-crew-state's strict head rule rejected
the active row, the coarse runs-list scan skipped it and matched the
older failed row at the submitted head, and an active validation read as
failed (observed on model-routing-benchmark-hardening: active head
ac61c64 vs task copy at fb47636d).

fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns
runs-ledger attribution: the branch's newest row alone decides, and a
newest row whose head cannot resolve locally is recognized only as a
provable pipeline-owned continuation - active (running) and anchored by
the immediately older row for the same branch having ended at exactly
this worktree's HEAD. The reader keeps the axi TOON as full detail for
that proven same-branch run. Unanchored, ancestor-anchored, and terminal
unresolvable rows stay unattributed, so branch-name coincidence and other
tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact
prior semantics for teardown (verified by the full teardown suite).

Tests: reproduction regression for the unfetched active fix head (reads
working via full run-step detail), coarse-path continuation when axi
answers another branch, and negative controls for the unanchored active
row and the unresolvable terminal row with the historical fallback
preserved.

Ported onto upstream/main f4d7875, where kunchenguid#3194 independently added the
branch_sync custody exemption on the full axi-status path: both mechanisms
now coexist, each owning one surface (TOON custody on the full path, the
runs ledger on the coarse path). The port deletes the superseded coarse
scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers
(fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the
exemption comment's "the one exemption" phrasing now that a second
complementary exemption exists, and points the stale
FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree
(judge follow-up #1). The parent coarse-guard test's fixture is the
ledger-anchored continuation shape, so its expectation flips to the fixed
behavior (working via run-step, never the older failed row); a new
mismatched-anchor coarse negative control preserves that guard's original
no-anchor protection (pane answers, never the older row).

* no-mistakes(document): Clarify pipeline attribution documentation
…uid#3663)

* fix(bin): pre-register claude workspace trust for task worktrees

A claude crewmate launched into a fresh task worktree met Claude Code's
interactive workspace-trust dialog before it ever read its brief, and firstmate
could not answer it: the key plane carries only Enter, Escape, and C-c with no
arrow navigation, and the dialog's selection starts on "No, exit", so the
documented Enter recipe ended the session instead of accepting it. Two workers
wedged this way and were unblocked only by hand-seeding the trust store per
path.

--dangerously-skip-permissions does not cover that gate. `claude --help`
records the dialog as skipped only in non-interactive mode, through -p or a
non-TTY stdout, and a crewmate pane is interactive, so there is no launch flag
to reach for.

fm-spawn now pre-registers the worktree through bin/fm-claude-trust.sh in the
existing claude branch, before the project settings that the same gate would
otherwise block, and refuses the spawn when that write fails rather than
launching a worker that would wedge.

The scope test is the safety property and is structural rather than a path
policy: the path must be a linked git worktree, sharing the spawning project's
common dir, whose top level is exactly the resolved argument. Git is the ground
truth, so the argument is never trusted on its own word, and a primary
checkout, an unrelated repo, a worktree subdirectory, a plain directory, and a
home directory are each refused rather than warned about or skipped. A
treehouse or orca path prefix was deliberately avoided because treehouse's root
is configurable, which would make a prefix both wrong and a new policy surface.
One structural test covers both worktree providers.

tests/fm-claude-trust.test.sh pins both halves, including a case where HOME is
itself a valid linked worktree so the home guard is proven load-bearing rather
than passing vacuously, plus the spawn-level proof that a claude spawn trusts
its worktree and launches with the brief pointed at the same store.

The adapter reference no longer tells a firstmate to press Enter on that
dialog, and the shared trust reference now names every harness surface: which
harnesses gate, which suppress at launch, which dodge the gate, which now
pre-registers, and that a claude secondmate is excluded by design.

The spawn fixture runs each spawn against a throwaway HOME so the suite cannot
write the developer's real store, isolating through HOME rather than
CLAUDE_CONFIG_DIR because the spawn forwards a set CLAUDE_CONFIG_DIR onto the
launch command that launch-shape assertions read.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v

* fix(bin): create the staged trust store exclusively

The staged store was written to a predictable pid-based path with a plain
write, which follows a symlink. Where the Claude config directory is writable
by another local account, that account could pre-create the path as a symlink
and redirect the write into another file the launching user owns.

The staged name now carries random bytes and is created with an exclusive
"wx" open, so an existing path is refused outright instead of followed. The
happy-path test also asserts no staged store survives the rename.

The durability comment now states the residual window plainly: the readback
proves the entry landed, not that it survives, because a vendor session that
rewrites the whole store afterwards can still drop it and no lock closes that
window when the writer is Claude itself. The worker then meets the dialog and
stalls, which reaches firstmate as the ordinary stale wake rather than as
silent success.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v

* no-mistakes(review): neutralise CDPATH in claude trust scope guard

* no-mistakes(review): sandbox HOME in spawn tests, drop out-of-scope artifacts

* no-mistakes(review): refuse unresolvable git dir, compact store, fix secondmate doc

* no-mistakes(review): clear git env overrides, resolve symlinked store target

* no-mistakes(review): degrade without node, fix Pi gate claim, record trust proof

* no-mistakes(review): refuse without node, pin CLAUDE_CONFIG_DIR in spawn tests

* no-mistakes(review): refuse relative config dir and concurrent store modification

* no-mistakes(review): correct orca worktree claim, clean staged store on failure

* no-mistakes(review): restore pretty-printed store, correct trust dialog docs

* no-mistakes(review): arm trust gate before busy state to avoid orphans

* no-mistakes(document): record claude trust pre-registration in its owner docs

* no-mistakes(document): note orca limit for claude trust pre-registration

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-spawn.sh by moving the Claude trust gate earlier rather than adding cleanup machinery. Diagnosis: Greptile reported that when Claude trust registration fails on tmux/Zellij/cmux/non-projected Herdr, the exit runs after the backend endpoint and /tmp/fm-<id> were created, and the abort trap cleans neither. The endpoint half is pre-existing, deliberate architecture — the two refusals immediately above the gate (the 60s `treehouse get` timeout at fm-spawn.sh:2550 and `validate_spawn_worktree` at :2487) also exit with the endpoint live and direct the operator with "inspect window $T"; spawn_abort_cleanup only reclaims orca endpoints (already covered via ORCA_ABORT_CLEANUP) and herdr projections. The temp-root half was genuinely introduced by this PR: the gate was placed beside the busy-state arm, ~30 lines after `mkdir -p "$TASK_TMP/gotmp"`, and fm-teardown can only find that root through `tasktmp=` in a meta record a refused spawn never publishes. Root-cause fix (smallest correct change, no new subsystem): - bin/fm-spawn.sh — moved the `claude*` trust gate from inside the busy-arm block up to the first point $WT is known, immediately after the `freshen_spawn_worktree_base` block and before TASK_TMP creation, the STATE setup, and the relaunch `clear_relaunch_harness_wiring` retirement. A refusal now leaves no temp root, no retired relaunch wiring, and no busy record; only the endpoint remains, in the same class as the two refusals just above it. - bin/fm-spawn.sh — the refusal message now ends with "inspect window $T", matching the existing convention so control/teardown can identify the endpoint. $T is set for every backend on the non-secondmate path. - bin/fm-spawn.sh:196 — header note corrected from "before any state is armed" to "before any per-task state exists". - tests/fm-claude-trust.test.sh — the existing refused-spawn test's own comment claimed "before any task state exists" but only asserted busy state. Renamed to test_refused_spawn_leaves_no_task_state and added an assertion that /tmp/fm-<id> is absent, with the task id suffixed by the test process pid so the assertion reads only this run's path (a stale /tmp/fm-refusedspawn from the fixed-id version was in fact present on this box). No assertions on implementation source bytes. Verification run locally: - The new assertion fails against the pre-fix bin/fm-spawn.sh ("not ok - a refused spawn stranded a temp root no teardown can find") and passes after — a real before/after regression proof. - tests/fm-claude-trust.test.sh: 20/20 ok. - tests/fm-backend.test.sh, fm-backend-orca, fm-control-relaunch, fm-spawn-dispatch-profile, fm-trace-context-spawn, fm-gotmp: all pass. - tests/fm-backlog-atomicity.test.sh: rc=0, 79 assertions ok. - bin/fm-lint.sh (repo's single lint owner, pinned ShellCheck 0.11.0 + actionlint 1.7.12): clean. - No /tmp/fm-refusedspawn* leftovers after the runs. Scope respected: no trust subsystem, no policy layer, no config surface, no endpoint-cleanup mechanism added; the change is an ordering move plus one error-message clause and the test that pins it. Adapter references and docs made no ordering claim, so none needed updating. Changes are left uncommitted in the worktree for the outer executor

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* feat(update): restart every live second mate after a successful update

/updatefirstmate only restarted a second mate when that pass advanced its
AGENTS.md or .agents/skills. An already-current home was skipped entirely, a
bin/-only advance was steered instead, and a remote host that could not report
its instruction diff was downgraded to a re-read. A running agent also freezes
its launch-time wiring - turn-end hooks, harness flags, per-harness feature
switches - and none of that is derivable from a file diff, so an unchanged
tracked surface is not evidence the agent is already on the current behavior.

Restart is now unconditional on a successful update of that home. Every live
second mate the pass leaves on the target commit is restarted, whether it
advanced or was already there.

The safety contract is unchanged: open records are persisted before the agent is
replaced, nothing is forced, stashed, or discarded, a home the pass had to skip
is not restarted at all, and a mate whose runtime cannot prove a restart keeps
the honest re-read path and is never reported as reloaded.

bin/fm-ff-lib.sh gains a settled-state hook that fires for a home left at the
base whether it advanced or was already there, and never for a skipped one; the
instruction-gated hook the session-start convergence sweep uses is untouched.

Regressions: fm-update pins the already-current mate into the restart set and
the unprovable one into the nudge set, and fm-secondmate-restart drives both
real commands end to end - an already-current home is named, persisted, and
genuinely replaced with its checkout untouched, while the unprovable one keeps
its running agent.

* no-mistakes(document): Document unconditional secondmate restarts
…3696)

* fix(bin): close reserved pending-reply keys via fm-send --resolve-key

fm-send wrote answered: notes that the reserved-key fold ignores, so
operator closes exited 0 while OPEN DECISIONS kept the decision open.
Speak the owning library's close vocabulary on that path, and refuse
when a reserved close cannot take effect.

* no-mistakes(review): Safely quote manual decision-close recovery commands

* no-mistakes(review): Reject unclosable overlong decision keys before sending

* no-mistakes(review): Remove contract suffix from open decisions hint

* no-mistakes(document): Document resolve-key line-cap refusal
* fix(bin): stop false missed-reply escalations for same-basename self-home answers

A healthy secondmate that wrote corr= to its own state/<id>.status never matched the parent channel, so recovery confirmed and the record escalated as pending-reply-missed. Make the report helper resolve the parent channel itself, skip parent-replies.status as wrong-home, put a readable sighting path on the missed line, and restatement-copy only that same-basename self-home file onto the parent channel.

* no-mistakes(review): Resolve late replies before recovery escalation

* no-mistakes(review): Tighten reply routing and regression coverage

* no-mistakes(review): Preserve reply paths and require explicit home

* no-mistakes(review): Encode wrong-home paths before persistence

* no-mistakes(document): Document corrected secondmate reply routing

* no-mistakes(lint): Fix pending-reply ShellCheck warnings
* feat(harness): verify gemini as a crewmate runtime adapter

Adds Gemini CLI as a fourth dispatch target alongside claude, codex, and
grok, scoped to crewmate and scout work only. Every axis was proven against
gemini-cli 0.58.0 rather than inferred; docs/verification/runtime-backends.md
carries the dated evidence and names what stayed unverified.

Busy state is semantic, not rendered: BeforeAgent opens a turn and AfterAgent
and SessionEnd close it. AfterAgent also fires on a manual interrupt, so a
cancelled turn closes its own record.

Three findings shaped the wiring rather than a config line:

- --skip-trust and GEMINI_CLI_TRUST_WORKSPACE=true are presented by the CLI
  as equivalents and are not. A controlled A/B showed --skip-trust leaves
  project configuration unloaded, so workspace skills never load.
- The worktree's .gemini/settings.json is the PROJECT's committed settings
  file, unlike claude's settings.local.json. Firstmate's hooks therefore go
  to a firstmate-owned state/<id>.gemini-settings.json reached through
  GEMINI_CLI_SYSTEM_SETTINGS_PATH, which also works untrusted and merges
  with a project's own hooks instead of replacing them.
- The shipped CLI is a node bundle whose live process reports comm as
  MainThread, so ancestry cannot see it. GEMINI_CLI=1 is load-bearing and is
  tested before an inherited CLAUDECODE, and pane liveness identifies gemini
  from the script argument through the new bin/fm-gemini-lib.sh.

Gemini is refused for secondmates: it has no primary supervision protocol.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L

* test: clear gemini's marker in launch and detection expectations

Every non-gemini launch now clears GEMINI_CLI the way it already clears
cursor's markers, so the two tests that pin the exact launch prefix are
updated to match. The harness-detection tests that scrub foreign markers
before probing ancestry scrub GEMINI_CLI too, so running the suite from
inside a gemini session cannot produce a false verdict.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L

* docs: classify the gemini harness reference

The documentation inventory is the single classification owner for maintained
prose surfaces, and every surface must appear in it exactly once. The new
harness reference is agent-runtime, matching its siblings.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L

* no-mistakes(review): Narrow Gemini ancestry detection

* no-mistakes(review): Restrict Gemini hooks to canonical launches

* no-mistakes(document): Document Gemini adapter support boundaries

* no-mistakes(ci): Fixed Gemini process identity when interpreter or script paths contain whitespace. Tmux liveness now uses NUL-delimited /proc argv on Linux, with the existing flattened ps fallback elsewhere. Added a real-process regression test. Verified with the Gemini harness test suite, full fm-lint, ShellCheck, and git diff --check. The CI and Require no-mistakes runs were action_required/attestation outcomes rather than code failures

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
…uid#3704)

* conclude parked runs the pipeline advanced past the task copy

A no-mistakes fix round commits in the daemon's own gate-repo clone, so a
run parked at a gate can carry a head whose object the task copy never
received. Teardown's strict object-local identity rule then declined to
conclude the run, and cleanup left it parked forever holding a fleet slot
(observed 2026-09-03; the same masking condition PR 3681 fixed on the
read path, now closing the teardown half its scope boundary deferred).

task_status_is_own_parked_run now falls back - only when the reported
head resolves to no local object - to the one shared runs-ledger
attribution rule fm_nm_runs_status_for_worktree (bin/fm-nm-run-lib.sh),
whose anchored continuation proof binds the branch's newest active row
to this worktree's exact submitted head. Foreign branches, stale
history, terminal rows, ancestor-only anchors, diverged newer rows, and
ambiguous multi-row shapes all still refuse, and runs that are actively
running, fixing, or in CI remain untouched: only the parked-at-a-gate
determination ever reaches the abort. No sqlite access, no fetches into
another task copy, no custody changes, no duplicated matching logic.

* tighten the parked-run ledger fallback and pin both judge corrections

The teardown ledger fallback now authorizes concluding this task's parked
run only when the shared runs-ledger rule's proved answer is the explicitly
active word (running): a terminal newest row - even anchored at exactly the
worktree's head - is finished history and never an abort authorization.
The read path may classify the same owner's answer; teardown's abort must
never fire for a run that already ended.

Two bounded pre-validation corrections from the implementation review:
- a fetched-object counterfactual pins the strict-rule path: a pipeline fix
  head fetched into the task copy aborts through object-local identity
  alone, with an empty ledger and a proof the runs query never fired;
- a negative fixture pins the tightened boundary: an unresolvable reported
  head with a terminal newest same-branch row anchored at the worktree head
  engages the ledger fallback and still refuses, so the refusal is the
  terminal-word boundary and not an earlier guard.

* no-mistakes(review): Bind teardown ledger fallback to validated run heads

* no-mistakes(review): Restore validated advanced-head ledger continuation

* no-mistakes(review): Reject invalid ledger dates and terminal statuses

* no-mistakes(document): Document teardown ledger scan limit
* Keep parked and aged undated captain holds off live Captain's Call.

Bearings was treating undated parked-style holds as live calls; mark those phrasings deferred and project holds older than a configurable 14-day since date as Charted Next gates instead.

* no-mistakes(review): Bound parked marker matching to lexical tokens

* no-mistakes(review): Age undated holds from durable hold-set dates

* no-mistakes(review): Reset re-held timestamps and scan full bodies

* no-mistakes(review): Preserve timestamp precision and prioritize parked suppression

* no-mistakes(document): Document undated captain-hold aging

* no-mistakes(ci): Fixed stock Bash CI test-count expectations (16 snapshot, 45 Bearings). Prevented fresh holds on old tasks from aging via stale `since` dates by aging only stamped holds. Added behavioral regressions and verified both suites plus Bash 3.2 parsing

* no-mistakes(ci): account for rebased snapshot regression

* no-mistakes(review): Restore legacy hold aging and mandate wrapper

* no-mistakes(review): Restrict hold stamps to canonical leading lines

* no-mistakes(review): Exclude historical answers and deduplicate revealed holds

* no-mistakes(document): Correct captain-hold projection documentation

* no-mistakes(ci): Rebased onto 8988af2 and resolved Bearings conflicts. Fixed the hold timestamp race by persisting and verifying the timestamp before publishing the captain hold; failures now leave the task unheld. Added behavioral coverage for ordering and failure handling. Preserved the required parked-phrase projection behavior. Relevant snapshot, Bearings, lifecycle, syntax, and ShellCheck validations pass

* no-mistakes(review): Bound current prose before historical resolutions

* no-mistakes(review): Preserve hold age across interrupted answers

* no-mistakes(review): Preserve leading hold stamps until answer closure

* no-mistakes(review): Normalize answer bodies on matching retries

* no-mistakes(review): Document concurrent re-hold age-basis limitation

* no-mistakes(document): Refresh captain hold lifecycle documentation

* no-mistakes(ci): Fixed both CI failures. Updated the macOS Bash snapshot expectation from 45 to 46 Bearings tests. Narrowed parked-style deferral matching to explicit hold-reason prefixes while preserving legacy explicit markers and preventing contextual prose from hiding active decisions. Added behavioral regression coverage. Verified with stock Bash 3.2: 17 fleet snapshot tests and 46 Bearings tests pass; full lint and workflow validation also pass

* no-mistakes(ci): Fixed Greptile’s P1 finding by restricting parked-style deferral phrases to complete hold-reason markers. Contextual reasons beginning with “not urgent,” “queued opportunity,” or “captain-gated” now remain visible decisions. Added behavioral coverage through the real fleet and Bearings snapshot paths and updated documentation. Verified both snapshot suites under Bash 3.2 (17 fleet tests and 46 Bearings tests), syntax checks, and git diff checks. The no-mistakes attestation failure is external/stale and requires the outer pipeline to refresh it for the new head

* no-mistakes(ci): Fixed parked-style undated captain holds disappearing from the default Bearings board. They now project to Charted Next with omitted[] disclosure, while --all-decisions reveals them and removes the safety gate. Added behavioral coverage for the reported “not urgent” case and aligned documentation. Verified fm-bearings-snapshot, fleet snapshot view, and captain-hold lifecycle tests; shellcheck, bash syntax, and git diff checks pass

* no-mistakes(test): Stabilize concurrency budget and provision timeout tests

* no-mistakes(document): Correct captain hold documentation details

* no-mistakes(ci): Fixed hold-reason parsing so commas in contextual reasons are preserved and do not incorrectly defer live Captain's Call decisions. Added end-to-end fleet/Bearings regression coverage. Reworked the flaky Herdr timeout test to assert observable late-launch behavior rather than process-ID liveness. Verified both snapshot suites, Herdr test 5 consecutive times, shell syntax, shellcheck, and git diff checks

* Restore the Herdr lab timeout test to its main version.

The stabilization rounds reworked tests/fm-herdr-lab.test.sh while chasing a
load-induced flake, replacing the fake server's wall-clock delay with a
SIGSTOP'd process and asserting that the blocked process is gone after a
timed-out provision. A stopped process does not die from SIGTERM, so that
assertion fails on Linux and the portable parallel shard stayed red.

That test is unrelated to the undated captain-hold projection this branch
delivers and was identical to main before these rounds, so restore main's
version exactly. It still proves that a timed-out provision cancels its late
launch before teardown.

* no-mistakes(review): Preserve metadata-like prose in captain hold reasons

* no-mistakes(review): Resurface due dated captain holds

* no-mistakes(review): Distinguish parked holds from explicit deferrals

* no-mistakes(review): Invalidate legacy secondmate summary caches

* no-mistakes(review): Keep blocked deferred holds in Charted Next

* no-mistakes(review): Count blocked deferred holds in omission disclosure

* no-mistakes(document): Correct captain-hold projection documentation

* no-mistakes(ci): Fixed both CI failures. Updated the macOS Bearings test count to 51. Preserved the v1 summary schema for compatibility while rejecting hold-bearing summaries missing the new aging fields, preventing stale caches from restoring noisy calls. Verified fleet snapshot, Bearings snapshot (51 tests), home-summary refresh, secondmate reconciliation, Bash 3.2 parsing, and diff checks

* no-mistakes(ci): Fixed Greptile’s valid finding: `--all-decisions` now reveals deferred/aged captain holds even when blocked, for both main and secondmate homes, and removes their duplicate Charted Next gates. Added behavioral regression coverage and updated documentation. The prose-classifier finding was not applied because exact complete-phrase matching is explicitly required by the author intent; contextual wording remains live. Verified with Bearings and fleet snapshot tests, `bin/fm-lint.sh`, Bash syntax checking, and `git diff --check`

* no-mistakes(ci): Fixed the actionable-state bug in Bearings: an arrived parked-style hold is live only when it is not explicitly non-actionable, so blocked due holds remain gated by default and are revealed by --all-decisions. Added behavioral regression coverage for that case. Preserved complete-reason parked-style classification as required by the author intent. Verified with tests/fm-bearings-snapshot.test.sh, bin/fm-lint.sh, and git diff --check

* Show why a revealed captain hold is deferred.

Under --all-decisions a deferred hold is revealed and its Charted Next gate
is removed, but the revealed row carried only the bare hold reason. A
date-deferred or blocked hold therefore read exactly like a genuine live
decision, because the until date, the age, and the blocking work only ever
appeared on the gate row that the reveal replaces.

Annotate a row that is revealed because it is deferred with the same
vocabulary the gate uses - until <date>, held <n>d, and the blocking work -
so the expanded view reads as deferred-but-shown. A genuinely live call is
left unannotated, and the default board is unchanged.

* Classify captain holds from structured fields alone.

Bucket membership was decided by several independent expressions, and two of
them matched hold reason or body prose. That produced a recurring class of
defects: holds that fell through every bucket and vanished from the board, and
live decisions silently suppressed because their wording happened to contain a
marker word - a reason of "non-deferred release choice" matched DEFERRED and
disappeared.

Replace all of it with one total classifier over structured fields only:
hold_kind, state, hold_until, unresolved_blocker_ids, and the machine-written
hold-set timestamp. Every captain hold gets exactly one hold_bucket - blocked,
dated, aged, or live - so no hold can fall through and none can match two.
captain_actionable is exactly the live bucket, and the --all-decisions reveal
is a property of the bucket rather than a second filter.

No hold reason or body prose is matched anywhere in the projection, so wording
can no longer hide, reveal, or reclassify a decision. A hold that is superseded
or no longer required is closed through the hold lifecycle instead of lingering
as an open hold flagged by a keyword.

* no-mistakes(review): Preserve working captain holds across bucket surfaces

* no-mistakes(review): Reject pre-classifier secondmate summary caches

* no-mistakes(review): Preserve complete live hold summaries

* no-mistakes(review): Clarify working hold decision bucket semantics

* no-mistakes(review): Reveal bounded remote holds and preserve blocker notes

* no-mistakes(review): Make blocker overflow explicit in hold summaries

* no-mistakes(document): Correct captain-hold projection documentation

* no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 17 to 18 tests. Verified the suite under Bash 3.2.57: all 18 tests pass. `git diff --check` also passes

* no-mistakes(ci): Updated the stock macOS Bash CI expectation from 51 to 53 Bearings tests. Verified all 53 pass under Bash 3.2.57; git diff --check passes
* fix(pi): deliver supervision outcomes off Pi's render thread

The supervision branch runs inside the captain's own Pi process, and Pi
runs extensions, their tools, and their event handlers on the single
JavaScript thread that also draws the TUI and reads the keyboard. Every
delivered outcome ran roughly five bash script invocations plus several
`ps` calls through spawnSync on that thread, so the TUI could not repaint
or echo a keystroke for the whole chain - the subsecond freeze the
captain saw every time a routine or captain-facing outcome arrived.

Convert the delivery path's subprocess calls to an awaited spawn behind a
serializing queue. lib/fm-async-exec.ts is the single owner of the
awaited-spawn replacement and returns the same capture shape and failure
verdicts spawnSync returned. Awaiting yields the thread, so what the
single thread used to guarantee for free is now an explicit queue: every
delivery, acknowledgement, and turn-boundary reconciliation runs as one
unit of it, preserving the durable append before anything visible, one
delivery at a time in sequence order, the read cursor advanced before the
next reader sees a row, and one ownership activation per generation.
Cancellation is preserved by the generation and lock-ownership rechecks
the awaits are placed around.

Two reads stay synchronous because Pi's own API is synchronous there, not
as an optimization: its bash spawn hook is typed as a plain function, and
the watcher reads offer.accepted the moment its dispatch event returns, so
a session that does not own the fleet lock must still refuse a wake
without waiting. Both walk the lock's process ancestry in full every time,
never cached, because reparenting and pid reuse can invalidate a
remembered chain and that answer decides ownership rather than hinting at
it. The store scripts and their durability contracts are unchanged.

Measured through the real fm_branch_report tool and real bin/ scripts with
a 1 ms interval timer, the largest block of the JS thread falls from 273
to 2.0 ms for a routine outcome, 286 to 2.0 ms for a captain outcome, and
134 to 1.9 ms for main's acknowledgement, against a 1.3-2.2 ms idle floor.
In a real Pi 0.82.0 TUI the worst keystroke echo while two outcomes arrive
falls from 676.9 ms to 36.8 ms, against a 22.6 ms extension-free floor.

Regressions: a delivery must leave the event loop running (zero timer
ticks before this change, in 250 ms), interleaved reports stay ordered and
exactly once, a session replaced mid-delivery neither loses nor duplicates
an outcome, and a failing store script surfaces without losing or doubling
one. The real-TUI half is an opt-in live guard that types into an isolated
Pi pane while outcomes are delivered and fails if echo leaves the class of
the same machine's own floor.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013bzoWyr2EcJGBKuoUjVRSp

* no-mistakes(review): Revalidate ownership and bound asynchronous subprocess output

* no-mistakes(ci): Fixed CI defects: routine outcomes now persist a sequence-keyed delivery receipt before awaiting cursor advancement, preventing duplicate delivery after mark-read failure. Corrected the session-replacement test to exercise an actual asynchronous ps ancestry lookup. Targeted behavioral tests, strict Pi typecheck, ShellCheck, and diff checks pass. The full extension test remains locally blocked by an unrelated stock-render assertion under the installed Pi runtime. The no-mistakes attestation failure is external pipeline state (test was previously skipped), not a source defect

* fix(pi): keep the declined routine receipt out and skip the renderer case below its Pi floor

Four follow-ups on the same branch, plus one revert.

Revert the routine-delivery receipt a CI auto-fix round added. It introduced a
new persisted `fm-branch-routine-delivery` entry, written into the captain's
transcript for every routine note, to deduplicate a note whose cursor write
failed. That is a change to the delivery contract, which this task is not
authorized to make: the approved work is the asynchronous conversion with the
existing durability contract preserved. The ownership re-read and output
bounding from the review round are kept - both are genuine asynchronous
correctness, not contract changes - as is that round's use of a real parent pid
so the replacement regression traverses an actual ps subprocess.

Record the routine gap instead of closing it. A routine note is a plain
message with no sequence-keyed record, so a mark-read failure after delivery
makes the next reconciliation send it once more; a captain row cannot
duplicate that way because its visible entry is found by store sequence. That
asymmetry predates moving delivery off the render thread. It is now stated at
the call site and in the delivery-contract docs, tracked as
fm-pi-routine-delivery-idempotency-followup-r1, and pinned by a regression
that proves the routine note is re-delivered exactly once more and never
again, the captain entry stays single, and the store keeps both rows.

Give the stock-renderer case a Pi version floor. It compares the extension's
renderers against Pi's stock rendering, so its verdict only means anything
against the contract those renderers target: since 0.84.4 the stock renderer
no longer supplies an implicit reset at multiline boundaries and the extension
emits that reset itself, so an older installed Pi differs legitimately. It now
names the installed version and the floor and skips, while a package whose
version cannot be read at all still fails.

Make the responsiveness regression's second signal a fraction rather than a
millisecond budget. A loaded machine that deschedules the process inflates an
absolute stall budget into a false failure, but it inflates the delivery's own
wall time too, so requiring the worst stall to be a minority of that wall time
holds under load. Synchronous delivery sits near 1.0 there whatever the load,
and the tick-count signal still reads zero on it.

Replace the test-family mapping for the Pi extension libraries with per-script
targeting. Routing them to whole families - or leaving them unmapped, which
widens through the reference scan to each referencing suite's entire family -
selected dozens of suites with nothing to do with Pi and pulled an unrelated
flake into the run. The changed-file selection drops from 112 scripts to 61.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013bzoWyr2EcJGBKuoUjVRSp

* no-mistakes(document): Clarify asynchronous execution documentation

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…nguid#3763)

* fix(memory): honor explicit project maintenance guidance

* no-mistakes(test): Blocked by pre-existing Bash and Muse fixture failures

* no-mistakes(test): Remove accidentally tracked test attribution report

* no-mistakes(ci): Restricted the marker to the exact first line, preventing fenced examples from suppressing governance, and corrected the documentation. Regression failed before the fix; all 18 helper tests, focused ShellCheck, documentation validation, and diff checks pass. CI and Require no-mistakes report action_required with zero jobs executed; those external checks remain unresolved
…nguid#3783)

* fix(spawn): refuse repository primary from linked spawning homes

Compare the resolved task git directory with the spawning repository's
common git directory before refreshing a fresh copy or relaunching a task.
This protects the primary even when the spawning project is a linked home.
Keep pooled copies accepted and preserve recorded work on relaunch.

Fixes kunchenguid#3741.

Verification for the pipeline PR body:
- Red on origin/main 1820316 with the new
  regression and unchanged production code: bin/fm-test-run.sh
  tests/fm-spawn-pool-base-freshen.test.sh exited 1 with
  "linked spawning home accepted primary as a disposable copy".
- Green after the guard: the complete pool-base-freshen and control-relaunch
  suites passed through bin/fm-test-run.sh, covering primary and symlink
  refusal before fetch/reset, spawning-directory refusal, scout acceptance,
  and committed plus unfinished work preserved during linked-home relaunch.
- The worktree-settle suite passed on pristine main and the final branch.
  An earlier loaded-host run exceeded its five-second assertion (6s);
  the final retry passed without changing code or the assertion.
- Test fixture commits ran with GIT_CONFIG_COUNT=1,
  GIT_CONFIG_KEY_0=commit.gpgsign, GIT_CONFIG_VALUE_0=false.
- bin/fm-lint.sh and /bin/bash -n for all three changed scripts passed.

The upstream cwd-selection cause remains outside this change.

* no-mistakes(document): Clarify spawn isolation ownership and relaunch preservation
…kunchenguid#3785)

* fix: read a failed herdr CLI as unreachable, not a gone backend target

The no-run fallback in bin/fm-crew-state.sh collapsed every failed pane
capture into 'backend target gone', which downstream consumers treat as
positive death evidence - so a herdr CLI that errors or stalls under load
briefly scored dozens of live claims dead on a busy box. Only a successful
herdr answer proving the pane absent (fm_backend_agent_state's 'missing',
backed by pane get answering pane_not_found) may now read as gone; every
other verdict reports 'backend unreachable' with the endpoint state, which
is never positive death evidence. Adds a behavior test: an always-failing
fake herdr reads unknown/unreachable, never gone.

* test: pin the herdr suite's ambient home to a marker-free fixture

FM_HOME defaults to the suite's own root when unset, and any secondmate-
marked checkout (every treehouse crew home carries .fm-secondmate-home)
flips the default workspace label to 2ndmate-*, so the ambiguous-label
placement test found zero firstmate matches and fell into the create path
instead of refusing (expected exit 3, got 1) - deterministically green in
CI, deterministically red from a crew home. Export a marker-free ambient
FM_HOME fixture; per-test FM_HOME prefixes still override it.

* fix: classify herdr endpoint answers instead of every non-missing verdict

Review decision (firstmate, 2026-09-05): a failed pane capture is not
itself evidence of death, but neither is every non-missing classifier
verdict a failed answer. missing (pane get answered pane_not_found) and
dead (pane present, agent_not_found husk) keep gone-class text so a
stale-claim sweep may still reclaim them; an alive answer falls through
to the normal busy/state flow instead of being discarded when only the
heavy 200-line scrollback read failed; only when the cheap pane get /
agent get calls themselves fail to answer does the line read
'backend unreachable'. Adds the two missing cases: alive with a failed
scrollback read stays live, and a husk pane still reads gone.

* no-mistakes(review): route tmux through agent-state classifier; drop test stall

* no-mistakes(review): narrow inaccurate tmux socket and alive-arm fallback comments

* no-mistakes(document): document classifier-backed endpoint verdicts in crew-state contract
)

* fix: distinguish subshell wake-lock owners on stock Bash

Restore distinct process ownership for issue kunchenguid#3743 using the existing PID helper, consistently across lock publication, reclaim, release, role checks, and bounded handoff.

The existing wake-queue regression fails on pristine upstream Bash 3.2 with rc=13. The complete suite now passes on Bash 3.2.57 and Bash 5.3.15, with added coverage for ownership when BASHPID is unset. Canonical lint and stock-Bash syntax checks pass.

* no-mistakes(document): Correct lock grace-period documentation

* no-mistakes(ci): Captain, fixed all 14 SC2031 false positives with nine ShellCheck source-boundary annotations across three tests. Full CI-mode lint and the complete wake-queue suite on stock Bash 3.2 passed. Runtime behavior is unchanged
…backends (kunchenguid#3782)

* fix(bin): close legacy records on the Beads backend honestly

Two pre-Beads reads blocked honest closure of leftover records:

1. fm-captain-hold.sh complete/verify resolved attested legacy hold ids
   only against the live backend and the pre-collapse derived identity, so
   a home whose holds fm-hold-migration rehomed under fm- ids failed with
   an empty-name absence message (the resolve failure was swallowed by the
   command substitution feeding verify_hold_durable). Resolution now falls
   back, on the Beads backend only, to the legacy id under the configured
   beads prefix and to the row whose notes carry the exact marker line
   'migrated from data/backlog.md id <legacy id>'; every refusal names the
   id it could not resolve, and the markdown path is unchanged.

2. fm-teardown.sh refused any record without spawn_gen forever. A record
   that predates the field can now be torn down with an explicit
   --legacy-record flag once the recovery-grade endpoint classifier
   confirms the recorded endpoint dead or agent-less; the accepted
   incarnation is stamped into the record right before its close marker
   binds to it and named in the teardown line. Refusals leave the record
   byte-identical, the unlanded-work refusal is not relaxed, and a corrupt
   (multi-valued) spawn_gen is never accepted.

The companion repair this branch carries (follow-up commit) is the
backend-gated --file and markdown-file requirement in the mutate path and
lifecycle gates: fm_backlog_mutate passed --file and required the markdown
backlog file regardless of the resolved backend, and the transition gate
plus row probe required that file before any backend work, so a home on a
non-markdown backend could neither gate, probe, nor close its rows.

Behavior tests: self-contained beads fixtures over a scratch bd graph
(self-skipping on markdown-only tasks-axi installs), legacy meta fixtures
for every teardown gate, and the relocated markdown backlog coverage stays
green.

* no-mistakes(review): fix(review): report migrated-hold scan refusals and guard legacy spawn_gen stamp against newline-less records

* fix(backlog): address the configured backend for lifecycle writes

Completes the fm-backlog-transition-lib repair the first commit's message
claims: on this base fm_backlog_mutate passed --file and required the
markdown backlog file regardless of the resolved backend, and
fm_backlog_transition_applies plus fm_backlog_row_probe required that file
before any backend work, so a home on a non-markdown backend could neither
gate, probe, nor close its backlog rows. All three now gate the markdown
file on the resolved tasks-axi backend: markdown keeps exactly its explicit
<data>/backlog.md behavior, non-markdown homes address the backend their
own configuration selects with no markdown file requirement.
fm_backlog_row_show and fm_backlog_row_list already gated correctly and
are unchanged. docs/configuration.md owns the contract line.

Also extends the same backend gate to fm-captain-hold.sh's own mutation
wrapper - hold/add/update/answer/done append the markdown --file only when
the resolved backend is markdown, so a captain call on a Beads home reaches
the Beads store end to end - and applies the review round's two direct
remedies there: the [beads] graph path resolves against the backlog root
when relative (never the process CWD), and a failed bd graph read reports
bd's own trimmed stderr reason in the refusal.

Coverage: tests/fm-backlog-atomicity.test.sh gains a stub-driven Beads
completion case proving the transition gate applies, the row probe reads,
and done runs without any markdown file or --file override; the relocated
markdown backlog test stays green.

* no-mistakes(review): Document root-tasks.toml-only beads settings for migrated-hold resolution

* test(gotmp): stub fm_tasks_axi_backend so the fixture matches the backend-aware transition lib

The legacy-records change made fm-backlog-transition-lib.sh resolve the
configured backend via fm_tasks_axi_backend before the markdown-only skip.
The gotmp fixture's fm-tasks-axi-lib stub lacked that function, so the
markdown check fell through and teardown hit the incompatible-backend
error with unbound FM_TASKS_AXI_MIN under set -u. Stub the backend as
markdown and define the floor, restoring the intended no-backlog skip.

* fix(teardown): roll the legacy stamp back when the close marker fails

A legacy-record teardown stamps its accepted incarnation into the record
right before the close marker binds to it; when that marker write then
fails, the stamp survived, so a retried teardown sailed past the
dead-or-agent-less endpoint gate the stamp now proved unnecessary. The
failed marker write now truncates the record back to its exact pre-stamp
bytes (verified by size), restoring the byte-identical-refusal invariant;
when the rollback itself fails the operator is told to re-run with
--legacy-record after reconciling the endpoint.

Also completes the recorded review decision's coverage wording: the
beads stub test now drives the answer close end to end (update and done
through the gated wrapper), asserting no markdown file override reaches
either verb.

* fix(review): harden the legacy stamp rollback and resolve derived migrated ids

The legacy-record stamp rollback now uses perl (already in the teardown
curated PATH; truncate is not, and is absent on stock macOS), routes every
failure branch inside the stamp block through the same size-verified
rollback so the byte-identical-refusal invariant holds on those paths too,
and gains behavior coverage: an unrecordable close (an invalid pr= link)
fails the teardown, leaves the record byte-identical, keeps the backlog
row in flight, and a flag-less retry still refuses.

Migrated-hold resolution now probes the derived pre-collapse identity
(<origin>-decision-<entry>) alongside the raw entry - fm-hold-migration
recorded the DERIVED id in every migrated row's marker note - in both the
prefix and the migration-note forms, with the ambiguity refusal naming
every identity tried, plus behavior coverage for a bare decision key
resolved through its derived identity's marker.

Also aligns fm-backlog-transition-lib.sh's header ADDRESSING/SCOPE
paragraphs with the backend-gated contract, drops an unreachable FORCE
validity guard the parser rewrite left behind, and switches the new stub
fixture to the portable sed -i.bak idiom.

* no-mistakes(review): Name the configured backend in teardown's backlog reminder

* no-mistakes(review): Scan migration markers before the prefix guess

* no-mistakes(review): Document marker-first resolution and cover the prefix branch

* no-mistakes(document): Record prefix-attestation audit and marker-line forms

* no-mistakes(ci): Fixed the Greptile P1 on bin/fm-teardown.sh: a failed rollback of the synthetic legacy stamp let a retry bypass the dead-or-agent-less endpoint gate. Root cause: teardown minted `spawn_gen=legacy-<ts>-<pid>` into the task record before the close marker bound to it. When the close-marker write failed AND the rollback also failed, the record retained that token. On the next invocation `fm_backlog_meta_spawn_gen` succeeded, so `TEARDOWN_LEGACY_PENDING` stayed 0 and the endpoint gate was skipped entirely — even with `--legacy-record`. The script's own error text told the operator to "re-run teardown with --legacy-record", advice the code could not honor. Fix (bin/fm-teardown.sh): - A `legacy-*` spawn_gen is now recognized as a stamp this teardown path minted, never one a spawn published (fm-spawn.sh publishes `s<epoch>.<pid>.<random>`). Such a record still reads as the legacy record it is: it re-enters the endpoint gate, and a flag-less retry refuses naming `--legacy-record`. - Acceptance reuses the retained token instead of minting a second one; the append block is skipped when the record already carries it, so no duplicate spawn_gen is written. - The rollback attempt and its "could not be rolled back" message are guarded to runs that actually appended a stamp, so a run that appended nothing never claims a rollback it did not perform. - Usage header documents the retained-stamp rule. Test (tests/fm-teardown.test.sh): added `test_retained_legacy_stamp_still_faces_the_endpoint_gate`, an end-to-end reproduction — a `perl` stub that fails only the rollback's `truncate` (delegating every other perl call to the real interpreter) leaves the stamp behind, then the retry must still hit the gate, must not stamp a second incarnation, must not close the backlog row, and the flag-less retry must refuse. Verification: the new test fails against the pre-fix script on exactly the reported defect ("the retry skipped the dead-or-agent-less endpoint gate") and passes after. Full tests/fm-teardown.test.sh 80 ok / 0 failures / rc=0; tests/fm-backlog-atomicity.test.sh 80 ok / 0 failures / rc=0; bin/fm-lint.sh (pinned ShellCheck 0.11.0 + actionlint 1.7.12) clean
Merges 67 upstream commits (f66be0f..af8c4b6) into the fork's main, keeping
upstream ancestry intact so future syncs do not re-conflict on the same
content. Eight files conflicted and were resolved by re-expressing the local
feature on top of upstream's structure.

bin/fm-tasks-axi-lib.sh: upstream's 353a8f0 independently implemented the
backend resolver the local beads commit had added as
fm_tasks_axi_storage_backend, and does it better - it is TOML-section aware,
honours a TASKS_AXI_BACKEND override, and falls back to the user's own
tasks-axi config, mirroring the CLI's real precedence. The local duplicate is
deleted and its five callers (bootstrap, backlog handoff, backlog receive) now
use upstream's fm_tasks_axi_backend. The genuinely local
fm_tasks_axi_has_beads_backend probe is kept and documented in the header.
All TASKS_AXI_BACKEND=markdown pins in those callers are single-command `env`
prefixes, so they never leak into the gates that read the resolver.

bin/fm-spawn.sh: kept upstream's linked-home isolation rewrite and its
relaunch/base-freshen wording, and re-expressed the remote-less base paragraph
on top of it. The remote-less code path, the .fm-task-owner marker write, and
the worktree_returned relaunch refusal all survive unchanged.

bin/fm-teardown.sh: upstream's 8f7b79c moved the run-conclusion and
process-reap block after the backlog transition. Took that move and re-applied
the local WORKTREE_STEPS guard at the block's new location, so an already
returned worktree still reaps only the task's own temp root. Ownership is
still proven before anything is killed or returned.

docs/architecture.md, docs/configuration.md: adjacent additions on both sides;
kept both, folded into upstream's pointer-style prose and its new
backlog-transition section.

tests/fm-spawn-pool-base-freshen.test.sh: adopted upstream's shared
fm_test_run_spawn fixture and its new linked-home case, keeping the local
GIT_CONFIG_GLOBAL/GIT_CONFIG_SYSTEM isolation as a prefix on that call, since
the shared helper pins HOME but not git's global or system config. The
local-only case fixture now writes its brief through upstream's
fm_test_spawn_brief, which 28fb5ac made the required format.

tests/fm-control-relaunch.test.sh, tests/fm-teardown.test.sh: pure additions
on both sides; kept every case from both.

tests/fm-teardown.test.sh additionally pins each case home to markdown
storage through fm_test_markdown_tasks_toml and runs its own tasks-axi calls
from that home. Upstream's new backlog-closing cases drive the real markdown
flow with the installed CLI, which otherwise resolved this repository's own
.tasks.toml and refused on the beads backend - the same pin the other
markdown-flow suites already carry.
…torage

Two suites upstream added since the fork point drive the real markdown backlog
flow with the installed tasks-axi. Both read this repository's own tracked
.tasks.toml, which selects beads storage here, so they addressed a beads store
with no database and refused before reaching what they assert.

fm-backlog-atomicity: test_dispatch_moves_the_item_in_flight_in_the_same_run
copies the tracked config precisely to prove the MARKDOWN addressing path, so
it now seeds that config pinned to markdown through fm_test_markdown_tasks_toml
rather than following the tracked default. The suite also runs from its own
temp root, so the roughly thirty-five explicit `--file` helper calls stop
inheriting a backend from the working directory; ROOT and TMP_ROOT are already
absolute and the scripts under test address their own home explicitly. This is
deliberately not done by exporting TASKS_AXI_BACKEND, which the suite's own
header warns would outrank each case's .tasks.toml fixture and break its two
beads cases.

fm-captain-hold-lifecycle: test_secondmate_home_publishes_holds_and_answers
copied the tracked config into a synthetic secondmate home; it now uses the
same markdown pin. This was the last raw copy of the tracked config in tests/.

Both suites keep exercising exactly the behavior they were written for, and
each now reaches the same point as a clean upstream/main worktree on this host.
bin/fm-home-seed.sh seeds a secondmate home by cloning this repository, so the
seeded home inherits the tracked .tasks.toml - beads storage on this fork -
while the beads database itself lives under the gitignored data/ and is not
cloned. Teardown in such a home now correctly refuses, because upstream's
captain-hold read cannot tell whether the item is still held for the captain,
and nine cases in this suite lost the behavior they were written to prove.

Every case here drives markdown backlogs, so a seeded home is now pinned the
same way make_home already pins the homes it builds. pin_seeded_home carries
the reason once and is called wherever the suite resolves a seeded child's real
path. The suite goes from 14 to 73 passing cases, matching a clean upstream/main
worktree on this host exactly.

This changes tests only. That a seeded secondmate home inherits beads storage
without a database is real behavior of the fork's tracked configuration and is
reported separately rather than papered over here.
The helper's comment named four suites, and that list was already wrong: the
teardown, backlog-atomicity and captain-hold suites need the same pin, and an
enumerated list rots every time a suite is added. State the rule instead, name
the seeded-home case that is easy to miss because bin/fm-home-seed.sh clones
this repository, and say which suites must NOT use the helper - the ones whose
subject is backend selection and which write their own fixture.
@LeonidShamis
LeonidShamis merged commit 7f19868 into main Sep 6, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.