fix(bin): attribute active runs with unfetched pipeline heads - #3681
kunchenguid merged 2 commits into
Conversation
A no-mistakes fix round advances the run head beyond the submitted head, and the pipeline commits in its own checkout, so the task copy never receives the new commit object. fm-crew-state's strict head rule rejected the active row, the coarse runs-list scan skipped it and matched the older failed row at the submitted head, and an active validation read as failed (observed on model-routing-benchmark-hardening: active head ac61c64 vs task copy at fb47636d). fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns runs-ledger attribution: the branch's newest row alone decides, and a newest row whose head cannot resolve locally is recognized only as a provable pipeline-owned continuation - active (running) and anchored by the immediately older row for the same branch having ended at exactly this worktree's HEAD. The reader keeps the axi TOON as full detail for that proven same-branch run. Unanchored, ancestor-anchored, and terminal unresolvable rows stay unattributed, so branch-name coincidence and other tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact prior semantics for teardown (verified by the full teardown suite). Tests: reproduction regression for the unfetched active fix head (reads working via full run-step detail), coarse-path continuation when axi answers another branch, and negative controls for the unanchored active row and the unresolvable terminal row with the historical fallback preserved. Ported onto upstream/main f4d7875, where kunchenguid#3194 independently added the branch_sync custody exemption on the full axi-status path: both mechanisms now coexist, each owning one surface (TOON custody on the full path, the runs ledger on the coarse path). The port deletes the superseded coarse scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers (fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the exemption comment's "the one exemption" phrasing now that a second complementary exemption exists, and points the stale FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree (judge follow-up #1). The parent coarse-guard test's fixture is the ledger-anchored continuation shape, so its expectation flips to the fixed behavior (working via run-step, never the older failed row); a new mismatched-anchor coarse negative control preserves that guard's original no-anchor protection (pane answers, never the older row).
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
Confidence Score: 5/5The PR appears safe to merge with no concrete changed-code defect identified. The ledger helper checks only the newest same-branch row, requires an unresolved row to be active, and attributes it only when the immediately older same-branch row resolves exactly to the worktree head; the added tests exercise the principal positive and negative paths. Reviews (1): Last reviewed commit: "no-mistakes(document): Clarify pipeline ..." | Re-trigger Greptile |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: c5b9aff15c
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| fi | ||
| return 0 | ||
| fi | ||
| if [ -n "$(fm_nm_resolve_commit "$wt" "$sha")" ]; then |
There was a problem hiding this comment.
Reject ambiguous run-head abbreviations before anchoring
Captain, when the ledger's abbreviated SHA is ambiguous in the task repository, fm_nm_resolve_commit returns empty just as it does for an absent object, so this branch treats the row as an unfetched active head. If the next same-branch row resolves to the worktree HEAD, the helper returns running and fm-crew-state.sh makes that potentially unrelated run authoritative, masking the worker's actual parked or failed state. Distinguish an absent object from an ambiguous prefix and reject the latter before applying the continuation anchor.
AGENTS.md reference: AGENTS.md:L367-L368
Useful? React with 👍 / 👎.
| # A branch-matching answer the strict rule rejected is this branch's | ||
| # own current run once the ledger proves the pipeline-owned | ||
| # continuation, so its axi TOON is the authoritative run detail | ||
| # (RUN_SOURCE stays full); only a foreign-branch answer leaves | ||
| # coarse status-word detail. | ||
| [ "$run_branch" = "$CREW_BRANCH" ] || RUN_SOURCE=coarse |
There was a problem hiding this comment.
Keep coarse status when full-run identity was rejected
When axi status returns a same-branch run whose head fails attribution, and the subsequent runs-ledger query finds a newer matching row, this branch-name check leaves RUN_SOURCE=full. The authoritative path then ignores COARSE_STATUS and reports the stale RUN_OUT gate, status, or outcome; for example, a newly running row can still be rendered as the older run's failure. Branch equality does not prove that two separately sampled records are the same run, so retain coarse mode unless the ledger row is explicitly tied to the full record.
AGENTS.md reference: AGENTS.md:L367-L368
Useful? React with 👍 / 👎.
| Any direct or remaining historical annotation prints every status line unread at the presentation cursor instead of replaying only the latest line. | ||
| `bin/fm-crew-state.sh <id>` is the cheap current-state read for an actionable heartbeat review: it attributes an active or terminal no-mistakes run under the shared run-attribution contract, then keeps that run-step authoritative even if the pane has closed. | ||
| [`bin/fm-nm-run-lib.sh`](../bin/fm-nm-run-lib.sh)'s header owns the exact branch, head, pipeline-custody, and newest-first attribution rules. | ||
| A run head the task copy cannot resolve locally is attributed only when the pipeline's own runs ledger proves it is an active continuation of the submitted head, so a pipeline fix round never reads as an older failed run. |
There was a problem hiding this comment.
Document the custody exemption alongside ledger anchoring
This new statement says every locally unresolvable run head is attributed only through the runs ledger, but fm-crew-state.sh lines 417-419 also attributes a same-branch active run directly when branch_sync.state=pipeline_owned, without querying that ledger. Because this architecture document owns supervision mechanisms, the absolute wording gives maintainers a false safety invariant; qualify it as the non-custody fallback or explicitly describe both attribution paths.
AGENTS.md reference: AGENTS.md:L400-L400
Useful? React with 👍 / 👎.
|
Speaking as Kun's firstmate: this is waiting on the author (CI in progress after fork approval). First look on HEAD Attestation: MATCH (body binds tip). Workflow approvals this pass: yes — approved first-time fork runs CI Contract-class: restore. An active no-mistakes fix round whose pipeline head object the task copy never fetched was misread as an older failed same-branch row. VISION.md per-rule
Eligible for auto-merge as restore once CI |
|
Speaking as Kun's firstmate: this is merged. Thank you @npayette84 — really appreciate you taking the time on this. |
* fix: start a fresh supervision branch for every main session (kunchenguid#3600) * fix(pi): start a new supervision branch conversation per main session The supervision branch reopened one recorded conversation forever, so every main session start reloaded the current generated prompt and then weeks of accumulated thread, where a superseded rule could still outweigh today's. The branch conversation is now scoped to one main session: the session generation owns the recorded conversation, so a cold start, /new, /resume, /fork, or a reload always builds a new one, while a rebuild inside one session (a model or effort change) still continues that session's own conversation. The dialog mirror re-anchors with it. Its durable cursor records what the previous branch conversation received, so a /resume or reload - which keeps main's own session file - would otherwise leave the new branch blind to dialog main itself still has. The reset is bounded by the current main session, and the cursor keeps advancing incrementally within it. The durable outcome store and its processed marker are untouched, so unacknowledged captain-facing outcomes still re-present on the new main session. * no-mistakes(document): Document fresh Pi supervision conversations * no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks * feat: restart second mates after instruction updates (kunchenguid#3614) * feat(update): restart second mates whose instructions changed /updatefirstmate pulled new bytes onto disk and then asked each advanced second mate to re-read them. A running agent holds AGENTS.md and every loaded skill frozen from launch and no verified harness offers a reload, so that steer could not reach a loaded skill at all and left the mate holding two contradictory copies of its own job description. An eligible mate is now restarted instead, in the same home and endpoint, through the existing transactional relaunch. The restart is gated on the mate first writing down the open work it holds only in conversation - the open-record half of /stow, never its memory sweeps - so an unregistered captain call is flushed before the conversation is spent. Anything that leaves the reload unprovable falls back to the old re-read message and is reported as exactly that, never as a clean reload. Remote mates take the same path: fm-remote-secondmate-control.sh gains a relaunch verb whose host-local leg runs that same control plane, since the mate is an ordinary local secondmate from its host's point of view. The primary resolves the profile and passes it explicitly, because config/secondmate-harness is not inherited and the file on that host belongs to a different home. fm-update.sh now splits its advanced live mates into a restart set and a nudge residual, and both sets require a changed instruction surface, which also closes the over-nudge against the session-start sweep. Restart is stricter still: a bin/-only advance reloads itself on the next call, so it never costs a conversation. Colocated tests cover the gating, the persist-then-restart order, the task-subset persist request, each unsafe fallback, the remote hop, and the remote sync's new instruction-surface report. * no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting * no-mistakes(review): Parallelize relaunches and classify replacement incarnations * no-mistakes(review): Gate restart actions on live agent state * no-mistakes(review): Handle failed restart workers without hanging * no-mistakes(review): Nudge legacy remotes and preserve persist recovery * no-mistakes(review): Document one-time secondmate restart rollout * no-mistakes(review): Honor arrived replies and refresh remote profiles * no-mistakes(review): Revert remote parent profile reconciliation * no-mistakes(review): Reset remote profile defaults and honor published results * no-mistakes(review): Preserve fallback nudges for unverifiable secondmates * no-mistakes(document): Document second-mate restart update flow * no-mistakes(lint): Fix ShellCheck warnings in restart scripts * perf: accelerate local validation with bounded concurrency (kunchenguid#3644) * perf(tests): route gate verification through the bounded concurrent runner Local validation was the pipeline's dominant cost: across 67 recorded no-mistakes agent sessions on this repo, 99.3% of command execution was `bash tests/*.test.sh`, run strictly one script at a time, and 2% of those calls were killed by an agent-guessed timeout and paid for twice. Three changes, each measured: - `.no-mistakes.yaml` pins `commands.test` to `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner already owns changed-file selection, bounded concurrency, the refusal of unproven scripts, and a generous automatic per-script bound, so the gate's baseline is neither a serial chain nor a guessed timeout. It stays intent-targeted - the Test step still runs its evidence agent on top - and excludes the live-Herdr family the required Herdr lane owns. - `bin/fm-test-run.sh` gives a plain list of script paths the same bounded automatic scheduler and automatic bound that `--changed` gets. Naming several subjects is how a verification round asks for exactly those scripts. The curated selections are untouched: `--lane` still composes CI shards whose serial lane must stay serial, `--family` is what the required Herdr lane runs, and `--all` stays a deliberate complete regression. - `pr-forge` is admitted to the concurrent-safe family registry on two consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those, and records `secondmate` and `session-bootstrap` as refused with the exact script and reason each failed on, so the refusals are actionable rather than silent. Measured on this host, 0 failures on both sides: verification round, 4 scripts 448s chained -> 231s through the runner (-48%) pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x) watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x) A fourth lever was implemented and then removed because the measurement refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made `fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s unchanged, back to back. Those sleeps are not overhead added to the clock - they are how a test waits for a subject moving on fm-watch.sh's own one-second cadence - so sampling less often only delays detection. It also broke `fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a settled condition. CONTRIBUTING.md records that result so the experiment is not repeated. * no-mistakes(review): Separate concurrent runs by isolation proof family * no-mistakes(review): Limit automatic timeouts to changed-file validation * no-mistakes(document): Clarify validation concurrency documentation * fix: copy PR URLs from durable records (kunchenguid#3648) * fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check` * fix(bin): disable Claude feedback drafts for fleet launches (kunchenguid#3661) * fix(bin): disable Claude's feedback-draft flow for fleet-launched agents Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched Claude crewmate and secondmate, so /bug and /feedback never queue or submit a bug report on the captain's behalf. feedbackDrafts is the documented settings key (Claude Code changelog 2.1.247); the per-launch CLI flag never touches the captain's global settings.json. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts * no-mistakes(document): Fix Claude feedback documentation formatting * fix(bin): layer both feedback-draft controls for defense in depth The prior --settings-only fix can be overridden by a managed Claude settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0 alongside --settings '{"feedbackDrafts":"off"}': either control alone disables the SendFeedback tool, so a managed override of one still leaves the other in force. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(document): Document Claude feedback-draft suppression ownership * feat(tests): run three more validation families concurrently (kunchenguid#3662) * perf(tests): admit three more families to concurrent validation The three families that `docs/fm-test-isolation-proof.md` recorded as refused were not refused for concurrency. Each blocker was a test that decided a property by wall clock, or a script filed where it cannot run. Fixing those three things admits all three families and recovers 28.6 minutes of local validation with no assertion removed or weakened. - `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the handoff, sleeping a fixed second, then delegating the move to the real binary. Nothing ever killed the fake, so on a host slow enough for the case's next assertions to take longer than a second, the orphan woke and completed the very move the case requires left undone, and recovery then failed with `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs during the injected crash showed exactly that, the item moving one second after the crash. All four crash injections in the file now go through a new `fm_fake_crash_injector` shim that signals the target and returns only once it is observably gone, and the pre-move fake never delegates the move at all. - `tests/fm-session-start.test.sh` proved the startup digest does not block on a slow current-state read by timing the whole digest against a fixed eight-second sleep, which a loaded host exceeds without the property being violated. It now holds that read open until the case releases it and asserts, the moment the digest returns, that the read has not finished. A digest that waited would wait indefinitely rather than for an interval a slow host can out-run, so the assertion is stronger than the bound it replaces. Its scan budget moves to the maximum, because the old value left two seconds of margin over the fixed sleep and measured the host rather than the deadline that `tests/fm-inactive-reconcile.test.sh` owns. - `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all, which put it in the portable serial lane, where Linux CI gate-skips it: that real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its opt-in variable and moves to `live-harness-optin`. The 28 remaining ungrouped scripts become an enumerated `standalone` family instead of admitting `unclassified` itself. `unclassified` is the family map's `*)` arm, so admitting it would silently grant concurrency to every test added afterwards, which is exactly the population with no proof. A new test still lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh` covers that split behaviorally. Each family passes two consecutive four-worker proofs with zero failures. On the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap` 756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock against 121 minutes of summed script time. * no-mistakes(document): Refresh concurrent validation and shard documentation * no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2 * feat: structure no-mistakes ask-user escalations (kunchenguid#3670) * feat(brief): structure no-mistakes ask-user escalation as event + snapshot file Crewmates escalating a no-mistakes ask-user gate now report one status event naming every finding id plus a snapshot file holding the gate's axi finding records verbatim (id, severity, file, line, description, authority), using the same shape even for a single finding. The status line never paraphrases. The format is defined once in fm-dod-lib.sh and rendered into both the scout and ship rule 6 in fm-brief.sh, so a promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets the identical contract as a freshly-spawned no-mistakes ship worker. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei * no-mistakes(review): Preserve ask-user escalation output contract * no-mistakes(review): Align escalation format test expectation * no-mistakes(review): Scope ask-user escalation instructions correctly * no-mistakes(review): Remove ask-user from generic decision rules --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix(bin): require self-sufficient no-mistakes intent (kunchenguid#3671) * fix(bin): require a self-sufficient no-mistakes intent A no-mistakes worker's --intent is only as useful as the string it passes. PR kunchenguid#3604 shipped with an intent that was only "do 1, 2, 3, 7 from the report": the real contract lived in a private scout report and never reached --intent, so nobody holding that string plus the codebase could have derived the specification. This is pure instruction at the contract's one owner; no spawn-side or promotion-side check is added. - bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now states that the --intent string must be self-sufficient (the string plus the codebase reconstructs roughly the same specification) and tells the worker to write the substance of any report, decision, or PR the captain's intent refers to into --intent rather than the pointer, while Firstmate build instructions and the worker's own decisions still stay out. The spawn-time overlay points back at that rule so its "supersedes" wording cannot cancel it, and the header's owner statement carries the rule. - AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to include the substance of referenced material when filling ## Captain's intent, and section 11 points at the owner of the rule. - tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the rendered brief and launch contract carry the rule. Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim * no-mistakes(document): Replace incident-specific intent test commentary * fix: accelerate local Bearings snapshot composition (kunchenguid#3499) * Speed local fleet snapshot composition * no-mistakes(review): Stabilize task inventory during concurrent snapshot composition * no-mistakes(document): Document local snapshot observation concurrency * no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks * no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks * no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky * no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks * fix(snapshot): keep live observations generation-coherent * no-mistakes(review): Keep secondmate observations generation-bound without copying reports * no-mistakes(document): Document generation-coherent snapshot observations * test(bearings): measure local read overlap instead of wall-clock budget The large-local-snapshot regression asserted that a whole snapshot composed in under five seconds. That bound measures how loaded the host is, not whether the per-task reads actually overlap, so it failed intermittently on a contended machine: one run in six on a box at load 16-20, landing exactly on the five second boundary. Time a serialized run and a concurrent run of the same workload instead and require the concurrent one to save at least two seconds. Both runs pay the same composition overhead, so the difference isolates the overlap this change delivers. Five one-second reads serialize into five seconds and overlap into about one, and re-serializing the reads collapses the saving to roughly zero, so the assertion still fails loudly if the concurrency regresses. Also bump the pinned Bearings test count to 48, since rebasing onto the current default branch picked up its captain-hold test. * no-mistakes(review): Restore JSON-derived decision flags * no-mistakes(review): Unify status-derived snapshot observations * no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass * fix: prevent stale supervision wake loops (kunchenguid#3672) * fix(bin): stop the supervision branch's stale-ack and ghost-report loops Clean-slate implementation of the four authorized recommendations from the supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal form, superseding PR kunchenguid#3604: - fm_branch_report refuses a task the wake being handled never named. The extension fixes the reportable task set from the eligible rows before each prompt (signal and stale rows resolve to their tasks, a heartbeat allows any task with a live record, fleet is always allowed), so a report typed from memory about a task whose records teardown already removed is never stored or delivered. - An acknowledgement that consumes nothing says "nothing was acknowledged through N" and prints the exact --ack-through / --recovery-generation command for the current presented wake, instead of "re-run the drain", which re-fed the same stale acknowledgement in a loop. - bin/fm-guard.sh no longer tells the branch actor to drain queued wakes while it is handling them; it names the granted rows instead. - Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and descendants; the index rebuild and the append-side index write both skip a task with neither a live record nor a status log, so the branch's report of a teardown it just performed is stored without recreating the index. No new locking, no spawn-generation binding, and no retired-task refusal: the branch can still report the outcome of a task it just tore down, and the teardown test now proves that path end to end. * fix(bin): narrow the branch report scope and guard silence to the minimal form Apply the four review decisions on the clean-slate branch: - A signal or stale prompt may report only the tasks its own rows resolve to; fleet is refused there too. A heartbeat review is not scoped by task at all, so the extension no longer tracks live task records and refuses nothing by task id during a fleet review. - The outcome-index rebuild no longer skips retired tasks; the append-side skip alone keeps a torn-down task's index from being recreated. - bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor instead of printing a replacement note. * no-mistakes(document): Align supervision docs with scoped wake handling * fix(bin): avoid fleet snapshot argument limits (kunchenguid#3677) * Fix fleet snapshot large JSON transport * no-mistakes(review): Captain: file-back fleet snapshot transport safely * no-mistakes(review): Captain: file-back parent summary aggregation * no-mistakes(ci): Rebased the PR's three commits onto f4d7875 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor * fix(bin): attribute active runs with unfetched pipeline heads (kunchenguid#3681) * fix(bin): recognize active pipeline fix rounds with unfetched run heads A no-mistakes fix round advances the run head beyond the submitted head, and the pipeline commits in its own checkout, so the task copy never receives the new commit object. fm-crew-state's strict head rule rejected the active row, the coarse runs-list scan skipped it and matched the older failed row at the submitted head, and an active validation read as failed (observed on model-routing-benchmark-hardening: active head ac61c64 vs task copy at fb47636d). fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns runs-ledger attribution: the branch's newest row alone decides, and a newest row whose head cannot resolve locally is recognized only as a provable pipeline-owned continuation - active (running) and anchored by the immediately older row for the same branch having ended at exactly this worktree's HEAD. The reader keeps the axi TOON as full detail for that proven same-branch run. Unanchored, ancestor-anchored, and terminal unresolvable rows stay unattributed, so branch-name coincidence and other tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact prior semantics for teardown (verified by the full teardown suite). Tests: reproduction regression for the unfetched active fix head (reads working via full run-step detail), coarse-path continuation when axi answers another branch, and negative controls for the unanchored active row and the unresolvable terminal row with the historical fallback preserved. Ported onto upstream/main f4d7875, where kunchenguid#3194 independently added the branch_sync custody exemption on the full axi-status path: both mechanisms now coexist, each owning one surface (TOON custody on the full path, the runs ledger on the coarse path). The port deletes the superseded coarse scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers (fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the exemption comment's "the one exemption" phrasing now that a second complementary exemption exists, and points the stale FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree (judge follow-up #1). The parent coarse-guard test's fixture is the ledger-anchored continuation shape, so its expectation flips to the fixed behavior (working via run-step, never the older failed row); a new mismatched-anchor coarse negative control preserves that guard's original no-anchor protection (pane answers, never the older row). * no-mistakes(document): Clarify pipeline attribution documentation * fix(bin): pre-register claude workspace trust at spawn time (kunchenguid#3663) * fix(bin): pre-register claude workspace trust for task worktrees A claude crewmate launched into a fresh task worktree met Claude Code's interactive workspace-trust dialog before it ever read its brief, and firstmate could not answer it: the key plane carries only Enter, Escape, and C-c with no arrow navigation, and the dialog's selection starts on "No, exit", so the documented Enter recipe ended the session instead of accepting it. Two workers wedged this way and were unblocked only by hand-seeding the trust store per path. --dangerously-skip-permissions does not cover that gate. `claude --help` records the dialog as skipped only in non-interactive mode, through -p or a non-TTY stdout, and a crewmate pane is interactive, so there is no launch flag to reach for. fm-spawn now pre-registers the worktree through bin/fm-claude-trust.sh in the existing claude branch, before the project settings that the same gate would otherwise block, and refuses the spawn when that write fails rather than launching a worker that would wedge. The scope test is the safety property and is structural rather than a path policy: the path must be a linked git worktree, sharing the spawning project's common dir, whose top level is exactly the resolved argument. Git is the ground truth, so the argument is never trusted on its own word, and a primary checkout, an unrelated repo, a worktree subdirectory, a plain directory, and a home directory are each refused rather than warned about or skipped. A treehouse or orca path prefix was deliberately avoided because treehouse's root is configurable, which would make a prefix both wrong and a new policy surface. One structural test covers both worktree providers. tests/fm-claude-trust.test.sh pins both halves, including a case where HOME is itself a valid linked worktree so the home guard is proven load-bearing rather than passing vacuously, plus the spawn-level proof that a claude spawn trusts its worktree and launches with the brief pointed at the same store. The adapter reference no longer tells a firstmate to press Enter on that dialog, and the shared trust reference now names every harness surface: which harnesses gate, which suppress at launch, which dodge the gate, which now pre-registers, and that a claude secondmate is excluded by design. The spawn fixture runs each spawn against a throwaway HOME so the suite cannot write the developer's real store, isolating through HOME rather than CLAUDE_CONFIG_DIR because the spawn forwards a set CLAUDE_CONFIG_DIR onto the launch command that launch-shape assertions read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * fix(bin): create the staged trust store exclusively The staged store was written to a predictable pid-based path with a plain write, which follows a symlink. Where the Claude config directory is writable by another local account, that account could pre-create the path as a symlink and redirect the write into another file the launching user owns. The staged name now carries random bytes and is created with an exclusive "wx" open, so an existing path is refused outright instead of followed. The happy-path test also asserts no staged store survives the rename. The durability comment now states the residual window plainly: the readback proves the entry landed, not that it survives, because a vendor session that rewrites the whole store afterwards can still drop it and no lock closes that window when the writer is Claude itself. The worker then meets the dialog and stalls, which reaches firstmate as the ordinary stale wake rather than as silent success. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * no-mistakes(review): neutralise CDPATH in claude trust scope guard * no-mistakes(review): sandbox HOME in spawn tests, drop out-of-scope artifacts * no-mistakes(review): refuse unresolvable git dir, compact store, fix secondmate doc * no-mistakes(review): clear git env overrides, resolve symlinked store target * no-mistakes(review): degrade without node, fix Pi gate claim, record trust proof * no-mistakes(review): refuse without node, pin CLAUDE_CONFIG_DIR in spawn tests * no-mistakes(review): refuse relative config dir and concurrent store modification * no-mistakes(review): correct orca worktree claim, clean staged store on failure * no-mistakes(review): restore pretty-printed store, correct trust dialog docs * no-mistakes(review): arm trust gate before busy state to avoid orphans * no-mistakes(document): record claude trust pre-registration in its owner docs * no-mistakes(document): note orca limit for claude trust pre-registration * no-mistakes(ci): Fixed the Greptile P1 on bin/fm-spawn.sh by moving the Claude trust gate earlier rather than adding cleanup machinery. Diagnosis: Greptile reported that when Claude trust registration fails on tmux/Zellij/cmux/non-projected Herdr, the exit runs after the backend endpoint and /tmp/fm-<id> were created, and the abort trap cleans neither. The endpoint half is pre-existing, deliberate architecture — the two refusals immediately above the gate (the 60s `treehouse get` timeout at fm-spawn.sh:2550 and `validate_spawn_worktree` at :2487) also exit with the endpoint live and direct the operator with "inspect window $T"; spawn_abort_cleanup only reclaims orca endpoints (already covered via ORCA_ABORT_CLEANUP) and herdr projections. The temp-root half was genuinely introduced by this PR: the gate was placed beside the busy-state arm, ~30 lines after `mkdir -p "$TASK_TMP/gotmp"`, and fm-teardown can only find that root through `tasktmp=` in a meta record a refused spawn never publishes. Root-cause fix (smallest correct change, no new subsystem): - bin/fm-spawn.sh — moved the `claude*` trust gate from inside the busy-arm block up to the first point $WT is known, immediately after the `freshen_spawn_worktree_base` block and before TASK_TMP creation, the STATE setup, and the relaunch `clear_relaunch_harness_wiring` retirement. A refusal now leaves no temp root, no retired relaunch wiring, and no busy record; only the endpoint remains, in the same class as the two refusals just above it. - bin/fm-spawn.sh — the refusal message now ends with "inspect window $T", matching the existing convention so control/teardown can identify the endpoint. $T is set for every backend on the non-secondmate path. - bin/fm-spawn.sh:196 — header note corrected from "before any state is armed" to "before any per-task state exists". - tests/fm-claude-trust.test.sh — the existing refused-spawn test's own comment claimed "before any task state exists" but only asserted busy state. Renamed to test_refused_spawn_leaves_no_task_state and added an assertion that /tmp/fm-<id> is absent, with the task id suffixed by the test process pid so the assertion reads only this run's path (a stale /tmp/fm-refusedspawn from the fixed-id version was in fact present on this box). No assertions on implementation source bytes. Verification run locally: - The new assertion fails against the pre-fix bin/fm-spawn.sh ("not ok - a refused spawn stranded a temp root no teardown can find") and passes after — a real before/after regression proof. - tests/fm-claude-trust.test.sh: 20/20 ok. - tests/fm-backend.test.sh, fm-backend-orca, fm-control-relaunch, fm-spawn-dispatch-profile, fm-trace-context-spawn, fm-gotmp: all pass. - tests/fm-backlog-atomicity.test.sh: rc=0, 79 assertions ok. - bin/fm-lint.sh (repo's single lint owner, pinned ShellCheck 0.11.0 + actionlint 1.7.12): clean. - No /tmp/fm-refusedspawn* leftovers after the runs. Scope respected: no trust subsystem, no policy layer, no config surface, no endpoint-cleanup mechanism added; the change is an ordering move plus one error-message clause and the test that pins it. Adapter references and docs made no ordering claim, so none needed updating. Changes are left uncommitted in the worktree for the outer executor --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix: restart every live second mate after updates (kunchenguid#3690) * feat(update): restart every live second mate after a successful update /updatefirstmate only restarted a second mate when that pass advanced its AGENTS.md or .agents/skills. An already-current home was skipped entirely, a bin/-only advance was steered instead, and a remote host that could not report its instruction diff was downgraded to a re-read. A running agent also freezes its launch-time wiring - turn-end hooks, harness flags, per-harness feature switches - and none of that is derivable from a file diff, so an unchanged tracked surface is not evidence the agent is already on the current behavior. Restart is now unconditional on a successful update of that home. Every live second mate the pass leaves on the target commit is restarted, whether it advanced or was already there. The safety contract is unchanged: open records are persisted before the agent is replaced, nothing is forced, stashed, or discarded, a home the pass had to skip is not restarted at all, and a mate whose runtime cannot prove a restart keeps the honest re-read path and is never reported as reloaded. bin/fm-ff-lib.sh gains a settled-state hook that fires for a home left at the base whether it advanced or was already there, and never for a skipped one; the instruction-gated hook the session-start convergence sweep uses is untouched. Regressions: fm-update pins the already-current mate into the restart set and the unprovable one into the nudge set, and fm-secondmate-restart drives both real commands end to end - an already-current home is named, persisted, and genuinely replaced with its checkout untouched, while the unprovable one keeps its running agent. * no-mistakes(document): Document unconditional secondmate restarts * fix(bin): close pending-reply decisions via resolve-key (kunchenguid#3696) * fix(bin): close reserved pending-reply keys via fm-send --resolve-key fm-send wrote answered: notes that the reserved-key fold ignores, so operator closes exited 0 while OPEN DECISIONS kept the decision open. Speak the owning library's close vocabulary on that path, and refuse when a reserved close cannot take effect. * no-mistakes(review): Safely quote manual decision-close recovery commands * no-mistakes(review): Reject unclosable overlong decision keys before sending * no-mistakes(review): Remove contract suffix from open decisions hint * no-mistakes(document): Document resolve-key line-cap refusal * fix(bin): prevent false missed-reply escalations (kunchenguid#3697) * fix(bin): stop false missed-reply escalations for same-basename self-home answers A healthy secondmate that wrote corr= to its own state/<id>.status never matched the parent channel, so recovery confirmed and the record escalated as pending-reply-missed. Make the report helper resolve the parent channel itself, skip parent-replies.status as wrong-home, put a readable sighting path on the missed line, and restatement-copy only that same-basename self-home file onto the parent channel. * no-mistakes(review): Resolve late replies before recovery escalation * no-mistakes(review): Tighten reply routing and regression coverage * no-mistakes(review): Preserve reply paths and require explicit home * no-mistakes(review): Encode wrong-home paths before persistence * no-mistakes(document): Document corrected secondmate reply routing * no-mistakes(lint): Fix pending-reply ShellCheck warnings * feat: add verified Gemini crewmate runtime (kunchenguid#3695) * feat(harness): verify gemini as a crewmate runtime adapter Adds Gemini CLI as a fourth dispatch target alongside claude, codex, and grok, scoped to crewmate and scout work only. Every axis was proven against gemini-cli 0.58.0 rather than inferred; docs/verification/runtime-backends.md carries the dated evidence and names what stayed unverified. Busy state is semantic, not rendered: BeforeAgent opens a turn and AfterAgent and SessionEnd close it. AfterAgent also fires on a manual interrupt, so a cancelled turn closes its own record. Three findings shaped the wiring rather than a config line: - --skip-trust and GEMINI_CLI_TRUST_WORKSPACE=true are presented by the CLI as equivalents and are not. A controlled A/B showed --skip-trust leaves project configuration unloaded, so workspace skills never load. - The worktree's .gemini/settings.json is the PROJECT's committed settings file, unlike claude's settings.local.json. Firstmate's hooks therefore go to a firstmate-owned state/<id>.gemini-settings.json reached through GEMINI_CLI_SYSTEM_SETTINGS_PATH, which also works untrusted and merges with a project's own hooks instead of replacing them. - The shipped CLI is a node bundle whose live process reports comm as MainThread, so ancestry cannot see it. GEMINI_CLI=1 is load-bearing and is tested before an inherited CLAUDECODE, and pane liveness identifies gemini from the script argument through the new bin/fm-gemini-lib.sh. Gemini is refused for secondmates: it has no primary supervision protocol. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * test: clear gemini's marker in launch and detection expectations Every non-gemini launch now clears GEMINI_CLI the way it already clears cursor's markers, so the two tests that pin the exact launch prefix are updated to match. The harness-detection tests that scrub foreign markers before probing ancestry scrub GEMINI_CLI too, so running the suite from inside a gemini session cannot produce a false verdict. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * docs: classify the gemini harness reference The documentation inventory is the single classification owner for maintained prose surfaces, and every surface must appear in it exactly once. The new harness reference is agent-runtime, matching its siblings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * no-mistakes(review): Narrow Gemini ancestry detection * no-mistakes(review): Restrict Gemini hooks to canonical launches * no-mistakes(document): Document Gemini adapter support boundaries * no-mistakes(ci): Fixed Gemini process identity when interpreter or script paths contain whitespace. Tmux liveness now uses NUL-delimited /proc argv on Linux, with the existing flattened ps fallback elsewhere. Added a real-process regression test. Verified with the Gemini harness test suite, full fm-lint, ShellCheck, and git diff --check. The CI and Require no-mistakes runs were action_required/attestation outcomes rather than code failures --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(teardown): conclude parked runs advanced past task copy (kunchenguid#3704) * conclude parked runs the pipeline advanced past the task copy A no-mistakes fix round commits in the daemon's own gate-repo clone, so a run parked at a gate can carry a head whose object the task copy never received. Teardown's strict object-local identity rule then declined to conclude the run, and cleanup left it parked forever holding a fleet slot (observed 2026-09-03; the same masking condition PR 3681 fixed on the read path, now closing the teardown half its scope boundary deferred). task_status_is_own_parked_run now falls back - only when the reported head resolves to no local object - to the one shared runs-ledger attribution rule fm_nm_runs_status_for_worktree (bin/fm-nm-run-lib.sh), whose anchored continuation proof binds the branch's newest active row to this worktree's exact submitted head. Foreign branches, stale history, terminal rows, ancestor-only anchors, diverged newer rows, and ambiguous multi-row shapes all still refuse, and runs that are actively running, fixing, or in CI remain untouched: only the parked-at-a-gate determination ever reaches the abort. No sqlite access, no fetches into another task copy, no custody changes, no duplicated matching logic. * tighten the parked-run ledger fallback and pin both judge corrections The teardown ledger fallback now authorizes concluding this task's parked run only when the shared runs-ledger rule's proved answer is the explicitly active word (running): a terminal newest row - even anchored at exactly the worktree's head - is finished history and never an abort authorization. The read path may classify the same owner's answer; teardown's abort must never fire for a run that already ended. Two bounded pre-validation corrections from the implementation review: - a fetched-object counterfactual pins the strict-rule path: a pipeline fix head fetched into the task copy aborts through object-local identity alone, with an empty ledger and a proof the runs query never fired; - a negative fixture pins the tightened boundary: an unresolvable reported head with a terminal newest same-branch row anchored at the worktree head engages the ledger fallback and still refuses, so the refusal is the terminal-word boundary and not an earlier guard. * no-mistakes(review): Bind teardown ledger fallback to validated run heads * no-mistakes(review): Restore validated advanced-head ledger continuation * no-mistakes(review): Reject invalid ledger dates and terminal statuses * no-mistakes(document): Document teardown ledger scan limit * feat(bin): show requested vs effective model in Herdr agent view Track spawn-config requested_model separately from runtime-verified effective_model, probe Claude/Pi transcripts for exact API ids, push compact display metadata to Herdr, and preserve verified models across relaunch/compaction hooks without inferring aliases as truth. * fix(bin): keep re-probing effective model after first exact reading fm-model-sync.sh only probed for the runtime-verified effective model while it was still pending/UNKNOWN, so a session that later switched models (manual switch, provider fallback) kept displaying the first verified model forever and never appended a fallback-history entry. Probe unconditionally instead; fm_model_record_effective already no-ops when the probed value is unchanged, so this stays cheap. Addresses the Greptile P1 finding on PR kunchenguid#3705's fm-model-sync.sh. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm * fix(bin): distinguish Cursor Grok, direct xAI Grok, and Anthropic Claude in the display Kapitänskorrektur: harness alone conflated Cursor-hosted Grok models (cursor-grok-4.6-*) and direct xAI Grok models (xai/grok-4.6) under one generic label, and displayed Anthropic Claude without naming the provider. Add fm_model_source_label, pattern-matched on the verified exact model id, so the compact display always reads Cursor · Grok, xAI · Grok, or Anthropic · Claude with the exact model id appended. Falls back to the existing harness label for every other model. No routing change: this only affects display strings. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm * fix(bin): wire model-sync into the Pi extension's turn lifecycle fm-model-sync.sh was only invoked from Claude's SessionStart/ UserPromptSubmit/Stop hooks; the Pi harness's own extension (state/<id>.pi-ext.ts) never called it, so a Pi-hosted session (e.g. a pi/xai-grok crewmate) never refreshed its effective model after the first probe and Herdr kept showing the stale value with no fallback-history entry. Call fm-model-sync.sh from the same agent_start/turn_end boundaries Pi already uses for busy-state and the turn-end notification touch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm * fix(bin): serialize fm-model-sync.sh's meta read-probe-write Overlapping lifecycle events (Pi's agent_start/turn_end, Claude's SessionStart/UserPromptSubmit/Stop) can invoke fm-model-sync.sh concurrently for the same task. The unlocked read-probe-write let interleaved runs revert a newer effective model, mismatch its source, or duplicate a model-history entry. Serialize the critical section through the same per-task meta lock fm-spawn.sh already uses (fm_meta_lock_path + fm_lock_acquire_wait/fm_lock_release), released before every exit path. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: Arthur Haro <38157909+haroarthur@users.noreply.github.com> Co-authored-by: Nicolas Payette <nicolas.payette@specira.ai> Co-authored-by: Jon Roosevelt <rooseveltadvisors@gmail.com> Co-authored-by: att430 <41454889+att430@users.noreply.github.com> Co-authored-by: Valentino-Sole <171032438+Valentino-Sole@users.noreply.github.com>
* fix: start a fresh supervision branch for every main session (kunchenguid#3600) * fix(pi): start a new supervision branch conversation per main session The supervision branch reopened one recorded conversation forever, so every main session start reloaded the current generated prompt and then weeks of accumulated thread, where a superseded rule could still outweigh today's. The branch conversation is now scoped to one main session: the session generation owns the recorded conversation, so a cold start, /new, /resume, /fork, or a reload always builds a new one, while a rebuild inside one session (a model or effort change) still continues that session's own conversation. The dialog mirror re-anchors with it. Its durable cursor records what the previous branch conversation received, so a /resume or reload - which keeps main's own session file - would otherwise leave the new branch blind to dialog main itself still has. The reset is bounded by the current main session, and the cursor keeps advancing incrementally within it. The durable outcome store and its processed marker are untouched, so unacknowledged captain-facing outcomes still re-present on the new main session. * no-mistakes(document): Document fresh Pi supervision conversations * no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks * feat: restart second mates after instruction updates (kunchenguid#3614) * feat(update): restart second mates whose instructions changed /updatefirstmate pulled new bytes onto disk and then asked each advanced second mate to re-read them. A running agent holds AGENTS.md and every loaded skill frozen from launch and no verified harness offers a reload, so that steer could not reach a loaded skill at all and left the mate holding two contradictory copies of its own job description. An eligible mate is now restarted instead, in the same home and endpoint, through the existing transactional relaunch. The restart is gated on the mate first writing down the open work it holds only in conversation - the open-record half of /stow, never its memory sweeps - so an unregistered captain call is flushed before the conversation is spent. Anything that leaves the reload unprovable falls back to the old re-read message and is reported as exactly that, never as a clean reload. Remote mates take the same path: fm-remote-secondmate-control.sh gains a relaunch verb whose host-local leg runs that same control plane, since the mate is an ordinary local secondmate from its host's point of view. The primary resolves the profile and passes it explicitly, because config/secondmate-harness is not inherited and the file on that host belongs to a different home. fm-update.sh now splits its advanced live mates into a restart set and a nudge residual, and both sets require a changed instruction surface, which also closes the over-nudge against the session-start sweep. Restart is stricter still: a bin/-only advance reloads itself on the next call, so it never costs a conversation. Colocated tests cover the gating, the persist-then-restart order, the task-subset persist request, each unsafe fallback, the remote hop, and the remote sync's new instruction-surface report. * no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting * no-mistakes(review): Parallelize relaunches and classify replacement incarnations * no-mistakes(review): Gate restart actions on live agent state * no-mistakes(review): Handle failed restart workers without hanging * no-mistakes(review): Nudge legacy remotes and preserve persist recovery * no-mistakes(review): Document one-time secondmate restart rollout * no-mistakes(review): Honor arrived replies and refresh remote profiles * no-mistakes(review): Revert remote parent profile reconciliation * no-mistakes(review): Reset remote profile defaults and honor published results * no-mistakes(review): Preserve fallback nudges for unverifiable secondmates * no-mistakes(document): Document second-mate restart update flow * no-mistakes(lint): Fix ShellCheck warnings in restart scripts * perf: accelerate local validation with bounded concurrency (kunchenguid#3644) * perf(tests): route gate verification through the bounded concurrent runner Local validation was the pipeline's dominant cost: across 67 recorded no-mistakes agent sessions on this repo, 99.3% of command execution was `bash tests/*.test.sh`, run strictly one script at a time, and 2% of those calls were killed by an agent-guessed timeout and paid for twice. Three changes, each measured: - `.no-mistakes.yaml` pins `commands.test` to `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner already owns changed-file selection, bounded concurrency, the refusal of unproven scripts, and a generous automatic per-script bound, so the gate's baseline is neither a serial chain nor a guessed timeout. It stays intent-targeted - the Test step still runs its evidence agent on top - and excludes the live-Herdr family the required Herdr lane owns. - `bin/fm-test-run.sh` gives a plain list of script paths the same bounded automatic scheduler and automatic bound that `--changed` gets. Naming several subjects is how a verification round asks for exactly those scripts. The curated selections are untouched: `--lane` still composes CI shards whose serial lane must stay serial, `--family` is what the required Herdr lane runs, and `--all` stays a deliberate complete regression. - `pr-forge` is admitted to the concurrent-safe family registry on two consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those, and records `secondmate` and `session-bootstrap` as refused with the exact script and reason each failed on, so the refusals are actionable rather than silent. Measured on this host, 0 failures on both sides: verification round, 4 scripts 448s chained -> 231s through the runner (-48%) pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x) watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x) A fourth lever was implemented and then removed because the measurement refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made `fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s unchanged, back to back. Those sleeps are not overhead added to the clock - they are how a test waits for a subject moving on fm-watch.sh's own one-second cadence - so sampling less often only delays detection. It also broke `fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a settled condition. CONTRIBUTING.md records that result so the experiment is not repeated. * no-mistakes(review): Separate concurrent runs by isolation proof family * no-mistakes(review): Limit automatic timeouts to changed-file validation * no-mistakes(document): Clarify validation concurrency documentation * fix: copy PR URLs from durable records (kunchenguid#3648) * fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check` * fix(bin): disable Claude feedback drafts for fleet launches (kunchenguid#3661) * fix(bin): disable Claude's feedback-draft flow for fleet-launched agents Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched Claude crewmate and secondmate, so /bug and /feedback never queue or submit a bug report on the captain's behalf. feedbackDrafts is the documented settings key (Claude Code changelog 2.1.247); the per-launch CLI flag never touches the captain's global settings.json. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts * no-mistakes(document): Fix Claude feedback documentation formatting * fix(bin): layer both feedback-draft controls for defense in depth The prior --settings-only fix can be overridden by a managed Claude settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0 alongside --settings '{"feedbackDrafts":"off"}': either control alone disables the SendFeedback tool, so a managed override of one still leaves the other in force. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(document): Document Claude feedback-draft suppression ownership * feat(tests): run three more validation families concurrently (kunchenguid#3662) * perf(tests): admit three more families to concurrent validation The three families that `docs/fm-test-isolation-proof.md` recorded as refused were not refused for concurrency. Each blocker was a test that decided a property by wall clock, or a script filed where it cannot run. Fixing those three things admits all three families and recovers 28.6 minutes of local validation with no assertion removed or weakened. - `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the handoff, sleeping a fixed second, then delegating the move to the real binary. Nothing ever killed the fake, so on a host slow enough for the case's next assertions to take longer than a second, the orphan woke and completed the very move the case requires left undone, and recovery then failed with `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs during the injected crash showed exactly that, the item moving one second after the crash. All four crash injections in the file now go through a new `fm_fake_crash_injector` shim that signals the target and returns only once it is observably gone, and the pre-move fake never delegates the move at all. - `tests/fm-session-start.test.sh` proved the startup digest does not block on a slow current-state read by timing the whole digest against a fixed eight-second sleep, which a loaded host exceeds without the property being violated. It now holds that read open until the case releases it and asserts, the moment the digest returns, that the read has not finished. A digest that waited would wait indefinitely rather than for an interval a slow host can out-run, so the assertion is stronger than the bound it replaces. Its scan budget moves to the maximum, because the old value left two seconds of margin over the fixed sleep and measured the host rather than the deadline that `tests/fm-inactive-reconcile.test.sh` owns. - `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all, which put it in the portable serial lane, where Linux CI gate-skips it: that real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its opt-in variable and moves to `live-harness-optin`. The 28 remaining ungrouped scripts become an enumerated `standalone` family instead of admitting `unclassified` itself. `unclassified` is the family map's `*)` arm, so admitting it would silently grant concurrency to every test added afterwards, which is exactly the population with no proof. A new test still lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh` covers that split behaviorally. Each family passes two consecutive four-worker proofs with zero failures. On the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap` 756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock against 121 minutes of summed script time. * no-mistakes(document): Refresh concurrent validation and shard documentation * no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2 * feat: structure no-mistakes ask-user escalations (kunchenguid#3670) * feat(brief): structure no-mistakes ask-user escalation as event + snapshot file Crewmates escalating a no-mistakes ask-user gate now report one status event naming every finding id plus a snapshot file holding the gate's axi finding records verbatim (id, severity, file, line, description, authority), using the same shape even for a single finding. The status line never paraphrases. The format is defined once in fm-dod-lib.sh and rendered into both the scout and ship rule 6 in fm-brief.sh, so a promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets the identical contract as a freshly-spawned no-mistakes ship worker. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei * no-mistakes(review): Preserve ask-user escalation output contract * no-mistakes(review): Align escalation format test expectation * no-mistakes(review): Scope ask-user escalation instructions correctly * no-mistakes(review): Remove ask-user from generic decision rules --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix(bin): require self-sufficient no-mistakes intent (kunchenguid#3671) * fix(bin): require a self-sufficient no-mistakes intent A no-mistakes worker's --intent is only as useful as the string it passes. PR kunchenguid#3604 shipped with an intent that was only "do 1, 2, 3, 7 from the report": the real contract lived in a private scout report and never reached --intent, so nobody holding that string plus the codebase could have derived the specification. This is pure instruction at the contract's one owner; no spawn-side or promotion-side check is added. - bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now states that the --intent string must be self-sufficient (the string plus the codebase reconstructs roughly the same specification) and tells the worker to write the substance of any report, decision, or PR the captain's intent refers to into --intent rather than the pointer, while Firstmate build instructions and the worker's own decisions still stay out. The spawn-time overlay points back at that rule so its "supersedes" wording cannot cancel it, and the header's owner statement carries the rule. - AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to include the substance of referenced material when filling ## Captain's intent, and section 11 points at the owner of the rule. - tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the rendered brief and launch contract carry the rule. Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim * no-mistakes(document): Replace incident-specific intent test commentary * fix: accelerate local Bearings snapshot composition (kunchenguid#3499) * Speed local fleet snapshot composition * no-mistakes(review): Stabilize task inventory during concurrent snapshot composition * no-mistakes(document): Document local snapshot observation concurrency * no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks * no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks * no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky * no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks * fix(snapshot): keep live observations generation-coherent * no-mistakes(review): Keep secondmate observations generation-bound without copying reports * no-mistakes(document): Document generation-coherent snapshot observations * test(bearings): measure local read overlap instead of wall-clock budget The large-local-snapshot regression asserted that a whole snapshot composed in under five seconds. That bound measures how loaded the host is, not whether the per-task reads actually overlap, so it failed intermittently on a contended machine: one run in six on a box at load 16-20, landing exactly on the five second boundary. Time a serialized run and a concurrent run of the same workload instead and require the concurrent one to save at least two seconds. Both runs pay the same composition overhead, so the difference isolates the overlap this change delivers. Five one-second reads serialize into five seconds and overlap into about one, and re-serializing the reads collapses the saving to roughly zero, so the assertion still fails loudly if the concurrency regresses. Also bump the pinned Bearings test count to 48, since rebasing onto the current default branch picked up its captain-hold test. * no-mistakes(review): Restore JSON-derived decision flags * no-mistakes(review): Unify status-derived snapshot observations * no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass * fix: prevent stale supervision wake loops (kunchenguid#3672) * fix(bin): stop the supervision branch's stale-ack and ghost-report loops Clean-slate implementation of the four authorized recommendations from the supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal form, superseding PR kunchenguid#3604: - fm_branch_report refuses a task the wake being handled never named. The extension fixes the reportable task set from the eligible rows before each prompt (signal and stale rows resolve to their tasks, a heartbeat allows any task with a live record, fleet is always allowed), so a report typed from memory about a task whose records teardown already removed is never stored or delivered. - An acknowledgement that consumes nothing says "nothing was acknowledged through N" and prints the exact --ack-through / --recovery-generation command for the current presented wake, instead of "re-run the drain", which re-fed the same stale acknowledgement in a loop. - bin/fm-guard.sh no longer tells the branch actor to drain queued wakes while it is handling them; it names the granted rows instead. - Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and descendants; the index rebuild and the append-side index write both skip a task with neither a live record nor a status log, so the branch's report of a teardown it just performed is stored without recreating the index. No new locking, no spawn-generation binding, and no retired-task refusal: the branch can still report the outcome of a task it just tore down, and the teardown test now proves that path end to end. * fix(bin): narrow the branch report scope and guard silence to the minimal form Apply the four review decisions on the clean-slate branch: - A signal or stale prompt may report only the tasks its own rows resolve to; fleet is refused there too. A heartbeat review is not scoped by task at all, so the extension no longer tracks live task records and refuses nothing by task id during a fleet review. - The outcome-index rebuild no longer skips retired tasks; the append-side skip alone keeps a torn-down task's index from being recreated. - bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor instead of printing a replacement note. * no-mistakes(document): Align supervision docs with scoped wake handling * fix(bin): avoid fleet snapshot argument limits (kunchenguid#3677) * Fix fleet snapshot large JSON transport * no-mistakes(review): Captain: file-back fleet snapshot transport safely * no-mistakes(review): Captain: file-back parent summary aggregation * no-mistakes(ci): Rebased the PR's three commits onto f4d7875 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor * fix(bin): attribute active runs with unfetched pipeline heads (kunchenguid#3681) * fix(bin): recognize active pipeline fix rounds with unfetched run heads A no-mistakes fix round advances the run head beyond the submitted head, and the pipeline commits in its own checkout, so the task copy never receives the new commit object. fm-crew-state's strict head rule rejected the active row, the coarse runs-list scan skipped it and matched the older failed row at the submitted head, and an active validation read as failed (observed on model-routing-benchmark-hardening: active head ac61c64 vs task copy at fb47636d). fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns runs-ledger attribution: the branch's newest row alone decides, and a newest row whose head cannot resolve locally is recognized only as a provable pipeline-owned continuation - active (running) and anchored by the immediately older row for the same branch having ended at exactly this worktree's HEAD. The reader keeps the axi TOON as full detail for that proven same-branch run. Unanchored, ancestor-anchored, and terminal unresolvable rows stay unattributed, so branch-name coincidence and other tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact prior semantics for teardown (verified by the full teardown suite). Tests: reproduction regression for the unfetched active fix head (reads working via full run-step detail), coarse-path continuation when axi answers another branch, and negative controls for the unanchored active row and the unresolvable terminal row with the historical fallback preserved. Ported onto upstream/main f4d7875, where kunchenguid#3194 independently added the branch_sync custody exemption on the full axi-status path: both mechanisms now coexist, each owning one surface (TOON custody on the full path, the runs ledger on the coarse path). The port deletes the superseded coarse scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers (fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the exemption comment's "the one exemption" phrasing now that a second complementary exemption exists, and points the stale FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree (judge follow-up #1). The parent coarse-guard test's fixture is the ledger-anchored continuation shape, so its expectation flips to the fixed behavior (working via run-step, never the older failed row); a new mismatched-anchor coarse negative control preserves that guard's original no-anchor protection (pane answers, never the older row). * no-mistakes(document): Clarify pipeline attribution documentation * fix(bin): pre-register claude workspace trust at spawn time (kunchenguid#3663) * fix(bin): pre-register claude workspace trust for task worktrees A claude crewmate launched into a fresh task worktree met Claude Code's interactive workspace-trust dialog before it ever read its brief, and firstmate could not answer it: the key plane carries only Enter, Escape, and C-c with no arrow navigation, and the dialog's selection starts on "No, exit", so the documented Enter recipe ended the session instead of accepting it. Two workers wedged this way and were unblocked only by hand-seeding the trust store per path. --dangerously-skip-permissions does not cover that gate. `claude --help` records the dialog as skipped only in non-interactive mode, through -p or a non-TTY stdout, and a crewmate pane is interactive, so there is no launch flag to reach for. fm-spawn now pre-registers the worktree through bin/fm-claude-trust.sh in the existing claude branch, before the project settings that the same gate would otherwise block, and refuses the spawn when that write fails rather than launching a worker that would wedge. The scope test is the safety property and is structural rather than a path policy: the path must be a linked git worktree, sharing the spawning project's common dir, whose top level is exactly the resolved argument. Git is the ground truth, so the argument is never trusted on its own word, and a primary checkout, an unrelated repo, a worktree subdirectory, a plain directory, and a home directory are each refused rather than warned about or skipped. A treehouse or orca path prefix was deliberately avoided because treehouse's root is configurable, which would make a prefix both wrong and a new policy surface. One structural test covers both worktree providers. tests/fm-claude-trust.test.sh pins both halves, including a case where HOME is itself a valid linked worktree so the home guard is proven load-bearing rather than passing vacuously, plus the spawn-level proof that a claude spawn trusts its worktree and launches with the brief pointed at the same store. The adapter reference no longer tells a firstmate to press Enter on that dialog, and the shared trust reference now names every harness surface: which harnesses gate, which suppress at launch, which dodge the gate, which now pre-registers, and that a claude secondmate is excluded by design. The spawn fixture runs each spawn against a throwaway HOME so the suite cannot write the developer's real store, isolating through HOME rather than CLAUDE_CONFIG_DIR because the spawn forwards a set CLAUDE_CONFIG_DIR onto the launch command that launch-shape assertions read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * fix(bin): create the staged trust store exclusively The staged store was written to a predictable pid-based path with a plain write, which follows a symlink. Where the Claude config directory is writable by another local account, that account could pre-create the path as a symlink and redirect the write into another file the launching user owns. The staged name now carries random bytes and is created with an exclusive "wx" open, so an existing path is refused outright instead of followed. The happy-path test also asserts no staged store survives the rename. The durability comment now states the residual window plainly: the readback proves the entry landed, not that it survives, because a vendor session that rewrites the whole store afterwards can still drop it and no lock closes that window when the writer is Claude itself. The worker then meets the dialog and stalls, which reaches firstmate as the ordinary stale wake rather than as silent success. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * no-mistakes(review): neutralise CDPATH in claude trust scope guard * no-mistakes(review): sandbox HOME in spawn tests, drop out-of-scope artifacts * no-mistakes(review): refuse unresolvable git dir, compact store, fix secondmate doc * no-mistakes(review): clear git env overrides, resolve symlinked store target * no-mistakes(review): degrade without node, fix Pi gate claim, record trust proof * no-mistakes(review): refuse without node, pin CLAUDE_CONFIG_DIR in spawn tests * no-mistakes(review): refuse relative config dir and concurrent store modification * no-mistakes(review): correct orca worktree claim, clean staged store on failure * no-mistakes(review): restore pretty-printed store, correct trust dialog docs * no-mistakes(review): arm trust gate before busy state to avoid orphans * no-mistakes(document): record claude trust pre-registration in its owner docs * no-mistakes(document): note orca limit for claude trust pre-registration * no-mistakes(ci): Fixed the Greptile P1 on bin/fm-spawn.sh by moving the Claude trust gate earlier rather than adding cleanup machinery. Diagnosis: Greptile reported that when Claude trust registration fails on tmux/Zellij/cmux/non-projected Herdr, the exit runs after the backend endpoint and /tmp/fm-<id> were created, and the abort trap cleans neither. The endpoint half is pre-existing, deliberate architecture — the two refusals immediately above the gate (the 60s `treehouse get` timeout at fm-spawn.sh:2550 and `validate_spawn_worktree` at :2487) also exit with the endpoint live and direct the operator with "inspect window $T"; spawn_abort_cleanup only reclaims orca endpoints (already covered via ORCA_ABORT_CLEANUP) and herdr projections. The temp-root half was genuinely introduced by this PR: the gate was placed beside the busy-state arm, ~30 lines after `mkdir -p "$TASK_TMP/gotmp"`, and fm-teardown can only find that root through `tasktmp=` in a meta record a refused spawn never publishes. Root-cause fix (smallest correct change, no new subsystem): - bin/fm-spawn.sh — moved the `claude*` trust gate from inside the busy-arm block up to the first point $WT is known, immediately after the `freshen_spawn_worktree_base` block and before TASK_TMP creation, the STATE setup, and the relaunch `clear_relaunch_harness_wiring` retirement. A refusal now leaves no temp root, no retired relaunch wiring, and no busy record; only the endpoint remains, in the same class as the two refusals just above it. - bin/fm-spawn.sh — the refusal message now ends with "inspect window $T", matching the existing convention so control/teardown can identify the endpoint. $T is set for every backend on the non-secondmate path. - bin/fm-spawn.sh:196 — header note corrected from "before any state is armed" to "before any per-task state exists". - tests/fm-claude-trust.test.sh — the existing refused-spawn test's own comment claimed "before any task state exists" but only asserted busy state. Renamed to test_refused_spawn_leaves_no_task_state and added an assertion that /tmp/fm-<id> is absent, with the task id suffixed by the test process pid so the assertion reads only this run's path (a stale /tmp/fm-refusedspawn from the fixed-id version was in fact present on this box). No assertions on implementation source bytes. Verification run locally: - The new assertion fails against the pre-fix bin/fm-spawn.sh ("not ok - a refused spawn stranded a temp root no teardown can find") and passes after — a real before/after regression proof. - tests/fm-claude-trust.test.sh: 20/20 ok. - tests/fm-backend.test.sh, fm-backend-orca, fm-control-relaunch, fm-spawn-dispatch-profile, fm-trace-context-spawn, fm-gotmp: all pass. - tests/fm-backlog-atomicity.test.sh: rc=0, 79 assertions ok. - bin/fm-lint.sh (repo's single lint owner, pinned ShellCheck 0.11.0 + actionlint 1.7.12): clean. - No /tmp/fm-refusedspawn* leftovers after the runs. Scope respected: no trust subsystem, no policy layer, no config surface, no endpoint-cleanup mechanism added; the change is an ordering move plus one error-message clause and the test that pins it. Adapter references and docs made no ordering claim, so none needed updating. Changes are left uncommitted in the worktree for the outer executor --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix: restart every live second mate after updates (kunchenguid#3690) * feat(update): restart every live second mate after a successful update /updatefirstmate only restarted a second mate when that pass advanced its AGENTS.md or .agents/skills. An already-current home was skipped entirely, a bin/-only advance was steered instead, and a remote host that could not report its instruction diff was downgraded to a re-read. A running agent also freezes its launch-time wiring - turn-end hooks, harness flags, per-harness feature switches - and none of that is derivable from a file diff, so an unchanged tracked surface is not evidence the agent is already on the current behavior. Restart is now unconditional on a successful update of that home. Every live second mate the pass leaves on the target commit is restarted, whether it advanced or was already there. The safety contract is unchanged: open records are persisted before the agent is replaced, nothing is forced, stashed, or discarded, a home the pass had to skip is not restarted at all, and a mate whose runtime cannot prove a restart keeps the honest re-read path and is never reported as reloaded. bin/fm-ff-lib.sh gains a settled-state hook that fires for a home left at the base whether it advanced or was already there, and never for a skipped one; the instruction-gated hook the session-start convergence sweep uses is untouched. Regressions: fm-update pins the already-current mate into the restart set and the unprovable one into the nudge set, and fm-secondmate-restart drives both real commands end to end - an already-current home is named, persisted, and genuinely replaced with its checkout untouched, while the unprovable one keeps its running agent. * no-mistakes(document): Document unconditional secondmate restarts * fix(bin): close pending-reply decisions via resolve-key (kunchenguid#3696) * fix(bin): close reserved pending-reply keys via fm-send --resolve-key fm-send wrote answered: notes that the reserved-key fold ignores, so operator closes exited 0 while OPEN DECISIONS kept the decision open. Speak the owning library's close vocabulary on that path, and refuse when a reserved close cannot take effect. * no-mistakes(review): Safely quote manual decision-close recovery commands * no-mistakes(review): Reject unclosable overlong decision keys before sending * no-mistakes(review): Remove contract suffix from open decisions hint * no-mistakes(document): Document resolve-key line-cap refusal * fix(bin): prevent false missed-reply escalations (kunchenguid#3697) * fix(bin): stop false missed-reply escalations for same-basename self-home answers A healthy secondmate that wrote corr= to its own state/<id>.status never matched the parent channel, so recovery confirmed and the record escalated as pending-reply-missed. Make the report helper resolve the parent channel itself, skip parent-replies.status as wrong-home, put a readable sighting path on the missed line, and restatement-copy only that same-basename self-home file onto the parent channel. * no-mistakes(review): Resolve late replies before recovery escalation * no-mistakes(review): Tighten reply routing and regression coverage * no-mistakes(review): Preserve reply paths and require explicit home * no-mistakes(review): Encode wrong-home paths before persistence * no-mistakes(document): Document corrected secondmate reply routing * no-mistakes(lint): Fix pending-reply ShellCheck warnings * feat: add verified Gemini crewmate runtime (kunchenguid#3695) * feat(harness): verify gemini as a crewmate runtime adapter Adds Gemini CLI as a fourth dispatch target alongside claude, codex, and grok, scoped to crewmate and scout work only. Every axis was proven against gemini-cli 0.58.0 rather than inferred; docs/verification/runtime-backends.md carries the dated evidence and names what stayed unverified. Busy state is semantic, not rendered: BeforeAgent opens a turn and AfterAgent and SessionEnd close it. AfterAgent also fires on a manual interrupt, so a cancelled turn closes its own record. Three findings shaped the wiring rather than a config line: - --skip-trust and GEMINI_CLI_TRUST_WORKSPACE=true are presented by the CLI as equivalents and are not. A controlled A/B showed --skip-trust leaves project configuration unloaded, so workspace skills never load. - The worktree's .gemini/settings.json is the PROJECT's committed settings file, unlike claude's settings.local.json. Firstmate's hooks therefore go to a firstmate-owned state/<id>.gemini-settings.json reached through GEMINI_CLI_SYSTEM_SETTINGS_PATH, which also works untrusted and merges with a project's own hooks instead of replacing them. - The shipped CLI is a node bundle whose live process reports comm as MainThread, so ancestry cannot see it. GEMINI_CLI=1 is load-bearing and is tested before an inherited CLAUDECODE, and pane liveness identifies gemini from the script argument through the new bin/fm-gemini-lib.sh. Gemini is refused for secondmates: it has no primary supervision protocol. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * test: clear gemini's marker in launch and detection expectations Every non-gemini launch now clears GEMINI_CLI the way it already clears cursor's markers, so the two tests that pin the exact launch prefix are updated to match. The harness-detection tests that scrub foreign markers before probing ancestry scrub GEMINI_CLI too, so running the suite from inside a gemini session cannot produce a false verdict. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * docs: classify the gemini harness reference The documentation inventory is the single classification owner for maintained prose surfaces, and every surface must appear in it exactly once. The new harness reference is agent-runtime, matching its siblings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * no-mistakes(review): Narrow Gemini ancestry detection * no-mistakes(review): Restrict Gemini hooks to canonical launches * no-mistakes(document): Document Gemini adapter support boundaries * no-mistakes(ci): Fixed Gemini process identity when interpreter or script paths contain whitespace. Tmux liveness now uses NUL-delimited /proc argv on Linux, with the existing flattened ps fallback elsewhere. Added a real-process regression test. Verified with the Gemini harness test suite, full fm-lint, ShellCheck, and git diff --check. The CI and Require no-mistakes runs were action_required/attestation outcomes rather than code failures --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(teardown): conclude parked runs advanced past task copy (kunchenguid#3704) * conclude parked runs the pipeline advanced past the task copy A no-mistakes fix round commits in the daemon's own gate-repo clone, so a run parked at a gate can carry a head whose object the task copy never received. Teardown's strict object-local identity rule then declined to conclude the run, and cleanup left it parked forever holding a fleet slot (observed 2026-09-03; the same masking condition PR 3681 fixed on the read path, now closing the teardown half its scope boundary deferred). task_status_is_own_parked_run now falls back - only when the reported head resolves to no local object - to the one shared runs-ledger attribution rule fm_nm_runs_status_for_worktree (bin/fm-nm-run-lib.sh), whose anchored continuation proof binds the branch's newest active row to this worktree's exact submitted head. Foreign branches, stale history, terminal rows, ancestor-only anchors, diverged newer rows, and ambiguous multi-row shapes all still refuse, and runs that are actively running, fixing, or in CI remain untouched: only the parked-at-a-gate determination ever reaches the abort. No sqlite access, no fetches into another task copy, no custody changes, no duplicated matching logic. * tighten the parked-run ledger fallback and pin both judge corrections The teardown ledger fallback now authorizes concluding this task's parked run only when the shared runs-ledger rule's proved answer is the explicitly active word (running): a terminal newest row - even anchored at exactly the worktree's head - is finished history and never an abort authorization. The read path may classify the same owner's answer; teardown's abort must never fire for a run that already ended. Two bounded pre-validation corrections from the implementation review: - a fetched-object counterfactual pins the strict-rule path: a pipeline fix head fetched into the task copy aborts through object-local identity alone, with an empty ledger and a proof the runs query never fired; - a negative fixture pins the tightened boundary: an unresolvable reported head with a terminal newest same-branch row anchored at the worktree head engages the ledger fallback and still refuses, so the refusal is the terminal-word boundary and not an earlier guard. * no-mistakes(review): Bind teardown ledger fallback to validated run heads * no-mistakes(review): Restore validated advanced-head ledger continuation * no-mistakes(review): Reject invalid ledger dates and terminal statuses * no-mistakes(document): Document teardown ledger scan limit * feat(bin): show requested vs effective model in Herdr agent view Track spawn-config requested_model separately from runtime-verified effective_model, probe Claude/Pi transcripts for exact API ids, push compact display metadata to Herdr, and preserve verified models across relaunch/compaction hooks without inferring aliases as truth. * fix(bin): keep re-probing effective model after first exact reading fm-model-sync.sh only probed for the runtime-verified effective model while it was still pending/UNKNOWN, so a session that later switched models (manual switch, provider fallback) kept displaying the first verified model forever and never appended a fallback-history entry. Probe unconditionally instead; fm_model_record_effective already no-ops when the probed value is unchanged, so this stays cheap. Addresses the Greptile P1 finding on PR kunchenguid#3705's fm-model-sync.sh. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm * fix(bin): distinguish Cursor Grok, direct xAI Grok, and Anthropic Claude in the display Kapitänskorrektur: harness alone conflated Cursor-hosted Grok models (cursor-grok-4.6-*) and direct xAI Grok models (xai/grok-4.6) under one generic label, and displayed Anthropic Claude without naming the provider. Add fm_model_source_label, pattern-matched on the verified exact model id, so the compact display always reads Cursor · Grok, xAI · Grok, or Anthropic · Claude with the exact model id appended. Falls back to the existing harness label for every other model. No routing change: this only affects display strings. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm * fix(bin): wire model-sync into the Pi extension's turn lifecycle fm-model-sync.sh was only invoked from Claude's SessionStart/ UserPromptSubmit/Stop hooks; the Pi harness's own extension (state/<id>.pi-ext.ts) never called it, so a Pi-hosted session (e.g. a pi/xai-grok crewmate) never refreshed its effective model after the first probe and Herdr kept showing the stale value with no fallback-history entry. Call fm-model-sync.sh from the same agent_start/turn_end boundaries Pi already uses for busy-state and the turn-end notification touch. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm * fix(bin): serialize fm-model-sync.sh's meta read-probe-write Overlapping lifecycle events (Pi's agent_start/turn_end, Claude's SessionStart/UserPromptSubmit/Stop) can invoke fm-model-sync.sh concurrently for the same task. The unlocked read-probe-write let interleaved runs revert a newer effective model, mismatch its source, or duplicate a model-history entry. Serialize the critical section through the same per-task meta lock fm-spawn.sh already uses (fm_meta_lock_path + fm_lock_acquire_wait/fm_lock_release), released before every exit path. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BwfFjeYQcz9cZ3vEZohmpm --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> Co-authored-by: Arthur Haro <38157909+haroarthur@users.noreply.github.com> Co-authored-by: Nicolas Payette <nicolas.payette@specira.ai> Co-authored-by: Jon Roosevelt <rooseveltadvisors@gmail.com> Co-authored-by: att430 <41454889+att430@users.noreply.github.com> Co-authored-by: Valentino-Sole <171032438+Valentino-Sole@users.noreply.github.com>
…nguid#3681) * fix(bin): recognize active pipeline fix rounds with unfetched run heads A no-mistakes fix round advances the run head beyond the submitted head, and the pipeline commits in its own checkout, so the task copy never receives the new commit object. fm-crew-state's strict head rule rejected the active row, the coarse runs-list scan skipped it and matched the older failed row at the submitted head, and an active validation read as failed (observed on model-routing-benchmark-hardening: active head ac61c64 vs task copy at fb47636d). fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns runs-ledger attribution: the branch's newest row alone decides, and a newest row whose head cannot resolve locally is recognized only as a provable pipeline-owned continuation - active (running) and anchored by the immediately older row for the same branch having ended at exactly this worktree's HEAD. The reader keeps the axi TOON as full detail for that proven same-branch run. Unanchored, ancestor-anchored, and terminal unresolvable rows stay unattributed, so branch-name coincidence and other tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact prior semantics for teardown (verified by the full teardown suite). Tests: reproduction regression for the unfetched active fix head (reads working via full run-step detail), coarse-path continuation when axi answers another branch, and negative controls for the unanchored active row and the unresolvable terminal row with the historical fallback preserved. Ported onto upstream/main f4d7875, where kunchenguid#3194 independently added the branch_sync custody exemption on the full axi-status path: both mechanisms now coexist, each owning one surface (TOON custody on the full path, the runs ledger on the coarse path). The port deletes the superseded coarse scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers (fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the exemption comment's "the one exemption" phrasing now that a second complementary exemption exists, and points the stale FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree (judge follow-up #1). The parent coarse-guard test's fixture is the ledger-anchored continuation shape, so its expectation flips to the fixed behavior (working via run-step, never the older failed row); a new mismatched-anchor coarse negative control preserves that guard's original no-anchor protection (pane answers, never the older row). * no-mistakes(document): Clarify pipeline attribution documentation
#133) * fix(bin): accept converged synchronized runs in validation binding A correctly attributed no-mistakes run that passed its required checks stays active only to monitor its open PR. In that state `axi status --run` still reports status running with no outcome and no branch_sync block, while `axi sync --check` reports the branch converged: state synchronized, safety already_synchronized, relation equal, and the pipeline block naming the same run, its submitted head, and its pushed-back current head. Both completion paths required pipeline_owned, so the bound run could never bind or seal its own advance (observed on fm-omp-orchestrate-opt-in run 01M25YRZ7NP7HVZ33K1XC94AYN, PR #131: submitted 5d52705, pushed c096921, checks green, still monitoring). Add fm_nm_run_branch_ownership to bin/fm-nm-run-lib.sh as the one owner of active-run branch evidence, used by both --bind-run and --complete. It keeps pipeline_owned acceptance unchanged and additionally accepts a synchronized state only on the full sync readout: same run id, submitted_head resolving to the validated head, current_head and the reported local head resolving to the run's observed head, relation equal, and safety already_synchronized. Checks-green readiness stays a separate requirement, so a merely synchronized-but-not-passed run still refuses. Upstream kunchenguid/firstmate has no fix to port: fm-receipt-check.sh is fork-local and upstream's diverged fm-nm-run-lib.sh has no synchronized handling. Adjacent upstream repairs (kunchenguid#2881, kunchenguid#3681, kunchenguid#3704, kunchenguid#3846) cover other read paths; open upstream PR kunchenguid#2799 confirms the checks-green merge-wait state but targets the supervision classifier, not completion. Extend tests/fm-receipt-check.test.sh with the converged bind+complete path and negative controls for foreign run ids, mismatched submitted heads, incomplete convergence fields, stale local heads, and synchronized-with-pending-CI completion. * no-mistakes(document): Correct run-owned branch wording in verification evidence
…nguid#3681) * fix(bin): recognize active pipeline fix rounds with unfetched run heads A no-mistakes fix round advances the run head beyond the submitted head, and the pipeline commits in its own checkout, so the task copy never receives the new commit object. fm-crew-state's strict head rule rejected the active row, the coarse runs-list scan skipped it and matched the older failed row at the submitted head, and an active validation read as failed (observed on model-routing-benchmark-hardening: active head ac61c64 vs task copy at fb47636d). fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns runs-ledger attribution: the branch's newest row alone decides, and a newest row whose head cannot resolve locally is recognized only as a provable pipeline-owned continuation - active (running) and anchored by the immediately older row for the same branch having ended at exactly this worktree's HEAD. The reader keeps the axi TOON as full detail for that proven same-branch run. Unanchored, ancestor-anchored, and terminal unresolvable rows stay unattributed, so branch-name coincidence and other tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact prior semantics for teardown (verified by the full teardown suite). Tests: reproduction regression for the unfetched active fix head (reads working via full run-step detail), coarse-path continuation when axi answers another branch, and negative controls for the unanchored active row and the unresolvable terminal row with the historical fallback preserved. Ported onto upstream/main f4d7875, where kunchenguid#3194 independently added the branch_sync custody exemption on the full axi-status path: both mechanisms now coexist, each owning one surface (TOON custody on the full path, the runs ledger on the coarse path). The port deletes the superseded coarse scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers (fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the exemption comment's "the one exemption" phrasing now that a second complementary exemption exists, and points the stale FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree (judge follow-up #1). The parent coarse-guard test's fixture is the ledger-anchored continuation shape, so its expectation flips to the fixed behavior (working via run-step, never the older failed row); a new mismatched-anchor coarse negative control preserves that guard's original no-anchor protection (pane answers, never the older row). * no-mistakes(document): Clarify pipeline attribution documentation
…nguid#3681) * fix(bin): recognize active pipeline fix rounds with unfetched run heads A no-mistakes fix round advances the run head beyond the submitted head, and the pipeline commits in its own checkout, so the task copy never receives the new commit object. fm-crew-state's strict head rule rejected the active row, the coarse runs-list scan skipped it and matched the older failed row at the submitted head, and an active validation read as failed (observed on model-routing-benchmark-hardening: active head ac61c64 vs task copy at fb47636d). fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns runs-ledger attribution: the branch's newest row alone decides, and a newest row whose head cannot resolve locally is recognized only as a provable pipeline-owned continuation - active (running) and anchored by the immediately older row for the same branch having ended at exactly this worktree's HEAD. The reader keeps the axi TOON as full detail for that proven same-branch run. Unanchored, ancestor-anchored, and terminal unresolvable rows stay unattributed, so branch-name coincidence and other tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact prior semantics for teardown (verified by the full teardown suite). Tests: reproduction regression for the unfetched active fix head (reads working via full run-step detail), coarse-path continuation when axi answers another branch, and negative controls for the unanchored active row and the unresolvable terminal row with the historical fallback preserved. Ported onto upstream/main f4d7875, where kunchenguid#3194 independently added the branch_sync custody exemption on the full axi-status path: both mechanisms now coexist, each owning one surface (TOON custody on the full path, the runs ledger on the coarse path). The port deletes the superseded coarse scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers (fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the exemption comment's "the one exemption" phrasing now that a second complementary exemption exists, and points the stale FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree (judge follow-up #1). The parent coarse-guard test's fixture is the ledger-anchored continuation shape, so its expectation flips to the fixed behavior (working via run-step, never the older failed row); a new mismatched-anchor coarse negative control preserves that guard's original no-anchor protection (pane answers, never the older row). * no-mistakes(document): Clarify pipeline attribution documentation
…nguid#3681) * fix(bin): recognize active pipeline fix rounds with unfetched run heads A no-mistakes fix round advances the run head beyond the submitted head, and the pipeline commits in its own checkout, so the task copy never receives the new commit object. fm-crew-state's strict head rule rejected the active row, the coarse runs-list scan skipped it and matched the older failed row at the submitted head, and an active validation read as failed (observed on model-routing-benchmark-hardening: active head ac61c64 vs task copy at fb47636d). fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns runs-ledger attribution: the branch's newest row alone decides, and a newest row whose head cannot resolve locally is recognized only as a provable pipeline-owned continuation - active (running) and anchored by the immediately older row for the same branch having ended at exactly this worktree's HEAD. The reader keeps the axi TOON as full detail for that proven same-branch run. Unanchored, ancestor-anchored, and terminal unresolvable rows stay unattributed, so branch-name coincidence and other tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact prior semantics for teardown (verified by the full teardown suite). Tests: reproduction regression for the unfetched active fix head (reads working via full run-step detail), coarse-path continuation when axi answers another branch, and negative controls for the unanchored active row and the unresolvable terminal row with the historical fallback preserved. Ported onto upstream/main f4d7875, where kunchenguid#3194 independently added the branch_sync custody exemption on the full axi-status path: both mechanisms now coexist, each owning one surface (TOON custody on the full path, the runs ledger on the coarse path). The port deletes the superseded coarse scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers (fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the exemption comment's "the one exemption" phrasing now that a second complementary exemption exists, and points the stale FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree (judge follow-up #1). The parent coarse-guard test's fixture is the ledger-anchored continuation shape, so its expectation flips to the fixed behavior (working via run-step, never the older failed row); a new mismatched-anchor coarse negative control preserves that guard's original no-anchor protection (pane answers, never the older row). * no-mistakes(document): Clarify pipeline attribution documentation
…g handling (#6) * feat: restart second mates after instruction updates (#3614) * feat(update): restart second mates whose instructions changed /updatefirstmate pulled new bytes onto disk and then asked each advanced second mate to re-read them. A running agent holds AGENTS.md and every loaded skill frozen from launch and no verified harness offers a reload, so that steer could not reach a loaded skill at all and left the mate holding two contradictory copies of its own job description. An eligible mate is now restarted instead, in the same home and endpoint, through the existing transactional relaunch. The restart is gated on the mate first writing down the open work it holds only in conversation - the open-record half of /stow, never its memory sweeps - so an unregistered captain call is flushed before the conversation is spent. Anything that leaves the reload unprovable falls back to the old re-read message and is reported as exactly that, never as a clean reload. Remote mates take the same path: fm-remote-secondmate-control.sh gains a relaunch verb whose host-local leg runs that same control plane, since the mate is an ordinary local secondmate from its host's point of view. The primary resolves the profile and passes it explicitly, because config/secondmate-harness is not inherited and the file on that host belongs to a different home. fm-update.sh now splits its advanced live mates into a restart set and a nudge residual, and both sets require a changed instruction surface, which also closes the over-nudge against the session-start sweep. Restart is stricter still: a bin/-only advance reloads itself on the next call, so it never costs a conversation. Colocated tests cover the gating, the persist-then-restart order, the task-subset persist request, each unsafe fallback, the remote hop, and the remote sync's new instruction-surface report. * no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting * no-mistakes(review): Parallelize relaunches and classify replacement incarnations * no-mistakes(review): Gate restart actions on live agent state * no-mistakes(review): Handle failed restart workers without hanging * no-mistakes(review): Nudge legacy remotes and preserve persist recovery * no-mistakes(review): Document one-time secondmate restart rollout * no-mistakes(review): Honor arrived replies and refresh remote profiles * no-mistakes(review): Revert remote parent profile reconciliation * no-mistakes(review): Reset remote profile defaults and honor published results * no-mistakes(review): Preserve fallback nudges for unverifiable secondmates * no-mistakes(document): Document second-mate restart update flow * no-mistakes(lint): Fix ShellCheck warnings in restart scripts * perf: accelerate local validation with bounded concurrency (#3644) * perf(tests): route gate verification through the bounded concurrent runner Local validation was the pipeline's dominant cost: across 67 recorded no-mistakes agent sessions on this repo, 99.3% of command execution was `bash tests/*.test.sh`, run strictly one script at a time, and 2% of those calls were killed by an agent-guessed timeout and paid for twice. Three changes, each measured: - `.no-mistakes.yaml` pins `commands.test` to `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner already owns changed-file selection, bounded concurrency, the refusal of unproven scripts, and a generous automatic per-script bound, so the gate's baseline is neither a serial chain nor a guessed timeout. It stays intent-targeted - the Test step still runs its evidence agent on top - and excludes the live-Herdr family the required Herdr lane owns. - `bin/fm-test-run.sh` gives a plain list of script paths the same bounded automatic scheduler and automatic bound that `--changed` gets. Naming several subjects is how a verification round asks for exactly those scripts. The curated selections are untouched: `--lane` still composes CI shards whose serial lane must stay serial, `--family` is what the required Herdr lane runs, and `--all` stays a deliberate complete regression. - `pr-forge` is admitted to the concurrent-safe family registry on two consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those, and records `secondmate` and `session-bootstrap` as refused with the exact script and reason each failed on, so the refusals are actionable rather than silent. Measured on this host, 0 failures on both sides: verification round, 4 scripts 448s chained -> 231s through the runner (-48%) pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x) watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x) A fourth lever was implemented and then removed because the measurement refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made `fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s unchanged, back to back. Those sleeps are not overhead added to the clock - they are how a test waits for a subject moving on fm-watch.sh's own one-second cadence - so sampling less often only delays detection. It also broke `fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a settled condition. CONTRIBUTING.md records that result so the experiment is not repeated. * no-mistakes(review): Separate concurrent runs by isolation proof family * no-mistakes(review): Limit automatic timeouts to changed-file validation * no-mistakes(document): Clarify validation concurrency documentation * fix: copy PR URLs from durable records (#3648) * fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check` * fix(bin): disable Claude feedback drafts for fleet launches (#3661) * fix(bin): disable Claude's feedback-draft flow for fleet-launched agents Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched Claude crewmate and secondmate, so /bug and /feedback never queue or submit a bug report on the captain's behalf. feedbackDrafts is the documented settings key (Claude Code changelog 2.1.247); the per-launch CLI flag never touches the captain's global settings.json. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts * no-mistakes(document): Fix Claude feedback documentation formatting * fix(bin): layer both feedback-draft controls for defense in depth The prior --settings-only fix can be overridden by a managed Claude settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0 alongside --settings '{"feedbackDrafts":"off"}': either control alone disables the SendFeedback tool, so a managed override of one still leaves the other in force. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(document): Document Claude feedback-draft suppression ownership * feat(tests): run three more validation families concurrently (#3662) * perf(tests): admit three more families to concurrent validation The three families that `docs/fm-test-isolation-proof.md` recorded as refused were not refused for concurrency. Each blocker was a test that decided a property by wall clock, or a script filed where it cannot run. Fixing those three things admits all three families and recovers 28.6 minutes of local validation with no assertion removed or weakened. - `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the handoff, sleeping a fixed second, then delegating the move to the real binary. Nothing ever killed the fake, so on a host slow enough for the case's next assertions to take longer than a second, the orphan woke and completed the very move the case requires left undone, and recovery then failed with `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs during the injected crash showed exactly that, the item moving one second after the crash. All four crash injections in the file now go through a new `fm_fake_crash_injector` shim that signals the target and returns only once it is observably gone, and the pre-move fake never delegates the move at all. - `tests/fm-session-start.test.sh` proved the startup digest does not block on a slow current-state read by timing the whole digest against a fixed eight-second sleep, which a loaded host exceeds without the property being violated. It now holds that read open until the case releases it and asserts, the moment the digest returns, that the read has not finished. A digest that waited would wait indefinitely rather than for an interval a slow host can out-run, so the assertion is stronger than the bound it replaces. Its scan budget moves to the maximum, because the old value left two seconds of margin over the fixed sleep and measured the host rather than the deadline that `tests/fm-inactive-reconcile.test.sh` owns. - `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all, which put it in the portable serial lane, where Linux CI gate-skips it: that real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its opt-in variable and moves to `live-harness-optin`. The 28 remaining ungrouped scripts become an enumerated `standalone` family instead of admitting `unclassified` itself. `unclassified` is the family map's `*)` arm, so admitting it would silently grant concurrency to every test added afterwards, which is exactly the population with no proof. A new test still lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh` covers that split behaviorally. Each family passes two consecutive four-worker proofs with zero failures. On the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap` 756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock against 121 minutes of summed script time. * no-mistakes(document): Refresh concurrent validation and shard documentation * no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2 * feat: structure no-mistakes ask-user escalations (#3670) * feat(brief): structure no-mistakes ask-user escalation as event + snapshot file Crewmates escalating a no-mistakes ask-user gate now report one status event naming every finding id plus a snapshot file holding the gate's axi finding records verbatim (id, severity, file, line, description, authority), using the same shape even for a single finding. The status line never paraphrases. The format is defined once in fm-dod-lib.sh and rendered into both the scout and ship rule 6 in fm-brief.sh, so a promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets the identical contract as a freshly-spawned no-mistakes ship worker. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei * no-mistakes(review): Preserve ask-user escalation output contract * no-mistakes(review): Align escalation format test expectation * no-mistakes(review): Scope ask-user escalation instructions correctly * no-mistakes(review): Remove ask-user from generic decision rules --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix(bin): require self-sufficient no-mistakes intent (#3671) * fix(bin): require a self-sufficient no-mistakes intent A no-mistakes worker's --intent is only as useful as the string it passes. PR #3604 shipped with an intent that was only "do 1, 2, 3, 7 from the report": the real contract lived in a private scout report and never reached --intent, so nobody holding that string plus the codebase could have derived the specification. This is pure instruction at the contract's one owner; no spawn-side or promotion-side check is added. - bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now states that the --intent string must be self-sufficient (the string plus the codebase reconstructs roughly the same specification) and tells the worker to write the substance of any report, decision, or PR the captain's intent refers to into --intent rather than the pointer, while Firstmate build instructions and the worker's own decisions still stay out. The spawn-time overlay points back at that rule so its "supersedes" wording cannot cancel it, and the header's owner statement carries the rule. - AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to include the substance of referenced material when filling ## Captain's intent, and section 11 points at the owner of the rule. - tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the rendered brief and launch contract carry the rule. Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim * no-mistakes(document): Replace incident-specific intent test commentary * fix: accelerate local Bearings snapshot composition (#3499) * Speed local fleet snapshot composition * no-mistakes(review): Stabilize task inventory during concurrent snapshot composition * no-mistakes(document): Document local snapshot observation concurrency * no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks * no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks * no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky * no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks * fix(snapshot): keep live observations generation-coherent * no-mistakes(review): Keep secondmate observations generation-bound without copying reports * no-mistakes(document): Document generation-coherent snapshot observations * test(bearings): measure local read overlap instead of wall-clock budget The large-local-snapshot regression asserted that a whole snapshot composed in under five seconds. That bound measures how loaded the host is, not whether the per-task reads actually overlap, so it failed intermittently on a contended machine: one run in six on a box at load 16-20, landing exactly on the five second boundary. Time a serialized run and a concurrent run of the same workload instead and require the concurrent one to save at least two seconds. Both runs pay the same composition overhead, so the difference isolates the overlap this change delivers. Five one-second reads serialize into five seconds and overlap into about one, and re-serializing the reads collapses the saving to roughly zero, so the assertion still fails loudly if the concurrency regresses. Also bump the pinned Bearings test count to 48, since rebasing onto the current default branch picked up its captain-hold test. * no-mistakes(review): Restore JSON-derived decision flags * no-mistakes(review): Unify status-derived snapshot observations * no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass * fix: prevent stale supervision wake loops (#3672) * fix(bin): stop the supervision branch's stale-ack and ghost-report loops Clean-slate implementation of the four authorized recommendations from the supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal form, superseding PR #3604: - fm_branch_report refuses a task the wake being handled never named. The extension fixes the reportable task set from the eligible rows before each prompt (signal and stale rows resolve to their tasks, a heartbeat allows any task with a live record, fleet is always allowed), so a report typed from memory about a task whose records teardown already removed is never stored or delivered. - An acknowledgement that consumes nothing says "nothing was acknowledged through N" and prints the exact --ack-through / --recovery-generation command for the current presented wake, instead of "re-run the drain", which re-fed the same stale acknowledgement in a loop. - bin/fm-guard.sh no longer tells the branch actor to drain queued wakes while it is handling them; it names the granted rows instead. - Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and descendants; the index rebuild and the append-side index write both skip a task with neither a live record nor a status log, so the branch's report of a teardown it just performed is stored without recreating the index. No new locking, no spawn-generation binding, and no retired-task refusal: the branch can still report the outcome of a task it just tore down, and the teardown test now proves that path end to end. * fix(bin): narrow the branch report scope and guard silence to the minimal form Apply the four review decisions on the clean-slate branch: - A signal or stale prompt may report only the tasks its own rows resolve to; fleet is refused there too. A heartbeat review is not scoped by task at all, so the extension no longer tracks live task records and refuses nothing by task id during a fleet review. - The outcome-index rebuild no longer skips retired tasks; the append-side skip alone keeps a torn-down task's index from being recreated. - bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor instead of printing a replacement note. * no-mistakes(document): Align supervision docs with scoped wake handling * fix(bin): avoid fleet snapshot argument limits (#3677) * Fix fleet snapshot large JSON transport * no-mistakes(review): Captain: file-back fleet snapshot transport safely * no-mistakes(review): Captain: file-back parent summary aggregation * no-mistakes(ci): Rebased the PR's three commits onto f4d7875824ecc5e274b4bb896f10c1e1f207b7e4 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor * fix(bin): attribute active runs with unfetched pipeline heads (#3681) * fix(bin): recognize active pipeline fix rounds with unfetched run heads A no-mistakes fix round advances the run head beyond the submitted head, and the pipeline commits in its own checkout, so the task copy never receives the new commit object. fm-crew-state's strict head rule rejected the active row, the coarse runs-list scan skipped it and matched the older failed row at the submitted head, and an active validation read as failed (observed on model-routing-benchmark-hardening: active head ac61c64b vs task copy at fb47636d). fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh now owns runs-ledger attribution: the branch's newest row alone decides, and a newest row whose head cannot resolve locally is recognized only as a provable pipeline-owned continuation - active (running) and anchored by the immediately older row for the same branch having ended at exactly this worktree's HEAD. The reader keeps the axi TOON as full detail for that proven same-branch run. Unanchored, ancestor-anchored, and terminal unresolvable rows stay unattributed, so branch-name coincidence and other tasks' runs never match, and fm_nm_head_matches_worktree keeps its exact prior semantics for teardown (verified by the full teardown suite). Tests: reproduction regression for the unfetched active fix head (reads working via full run-step detail), coarse-path continuation when axi answers another branch, and negative controls for the unanchored active row and the unresolvable terminal row with the historical fallback preserved. Ported onto upstream/main f4d78758, where #3194 independently added the branch_sync custody exemption on the full axi-status path: both mechanisms now coexist, each owning one surface (TOON custody on the full path, the runs ledger on the coarse path). The port deletes the superseded coarse scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers (fm_nm_head_resolvable, nm_coarse_head_matches_worktree), renames the exemption comment's "the one exemption" phrasing now that a second complementary exemption exists, and points the stale FM_CREW_STATE_RUNS_LIMIT comment at fm_nm_runs_status_for_worktree (judge follow-up #1). The parent coarse-guard test's fixture is the ledger-anchored continuation shape, so its expectation flips to the fixed behavior (working via run-step, never the older failed row); a new mismatched-anchor coarse negative control preserves that guard's original no-anchor protection (pane answers, never the older row). * no-mistakes(document): Clarify pipeline attribution documentation * fix(bin): pre-register claude workspace trust at spawn time (#3663) * fix(bin): pre-register claude workspace trust for task worktrees A claude crewmate launched into a fresh task worktree met Claude Code's interactive workspace-trust dialog before it ever read its brief, and firstmate could not answer it: the key plane carries only Enter, Escape, and C-c with no arrow navigation, and the dialog's selection starts on "No, exit", so the documented Enter recipe ended the session instead of accepting it. Two workers wedged this way and were unblocked only by hand-seeding the trust store per path. --dangerously-skip-permissions does not cover that gate. `claude --help` records the dialog as skipped only in non-interactive mode, through -p or a non-TTY stdout, and a crewmate pane is interactive, so there is no launch flag to reach for. fm-spawn now pre-registers the worktree through bin/fm-claude-trust.sh in the existing claude branch, before the project settings that the same gate would otherwise block, and refuses the spawn when that write fails rather than launching a worker that would wedge. The scope test is the safety property and is structural rather than a path policy: the path must be a linked git worktree, sharing the spawning project's common dir, whose top level is exactly the resolved argument. Git is the ground truth, so the argument is never trusted on its own word, and a primary checkout, an unrelated repo, a worktree subdirectory, a plain directory, and a home directory are each refused rather than warned about or skipped. A treehouse or orca path prefix was deliberately avoided because treehouse's root is configurable, which would make a prefix both wrong and a new policy surface. One structural test covers both worktree providers. tests/fm-claude-trust.test.sh pins both halves, including a case where HOME is itself a valid linked worktree so the home guard is proven load-bearing rather than passing vacuously, plus the spawn-level proof that a claude spawn trusts its worktree and launches with the brief pointed at the same store. The adapter reference no longer tells a firstmate to press Enter on that dialog, and the shared trust reference now names every harness surface: which harnesses gate, which suppress at launch, which dodge the gate, which now pre-registers, and that a claude secondmate is excluded by design. The spawn fixture runs each spawn against a throwaway HOME so the suite cannot write the developer's real store, isolating through HOME rather than CLAUDE_CONFIG_DIR because the spawn forwards a set CLAUDE_CONFIG_DIR onto the launch command that launch-shape assertions read. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * fix(bin): create the staged trust store exclusively The staged store was written to a predictable pid-based path with a plain write, which follows a symlink. Where the Claude config directory is writable by another local account, that account could pre-create the path as a symlink and redirect the write into another file the launching user owns. The staged name now carries random bytes and is created with an exclusive "wx" open, so an existing path is refused outright instead of followed. The happy-path test also asserts no staged store survives the rename. The durability comment now states the residual window plainly: the readback proves the entry landed, not that it survives, because a vendor session that rewrites the whole store afterwards can still drop it and no lock closes that window when the writer is Claude itself. The worker then meets the dialog and stalls, which reaches firstmate as the ordinary stale wake rather than as silent success. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HNEN2GLnew27HFyfi4ms4v * no-mistakes(review): neutralise CDPATH in claude trust scope guard * no-mistakes(review): sandbox HOME in spawn tests, drop out-of-scope artifacts * no-mistakes(review): refuse unresolvable git dir, compact store, fix secondmate doc * no-mistakes(review): clear git env overrides, resolve symlinked store target * no-mistakes(review): degrade without node, fix Pi gate claim, record trust proof * no-mistakes(review): refuse without node, pin CLAUDE_CONFIG_DIR in spawn tests * no-mistakes(review): refuse relative config dir and concurrent store modification * no-mistakes(review): correct orca worktree claim, clean staged store on failure * no-mistakes(review): restore pretty-printed store, correct trust dialog docs * no-mistakes(review): arm trust gate before busy state to avoid orphans * no-mistakes(document): record claude trust pre-registration in its owner docs * no-mistakes(document): note orca limit for claude trust pre-registration * no-mistakes(ci): Fixed the Greptile P1 on bin/fm-spawn.sh by moving the Claude trust gate earlier rather than adding cleanup machinery. Diagnosis: Greptile reported that when Claude trust registration fails on tmux/Zellij/cmux/non-projected Herdr, the exit runs after the backend endpoint and /tmp/fm-<id> were created, and the abort trap cleans neither. The endpoint half is pre-existing, deliberate architecture — the two refusals immediately above the gate (the 60s `treehouse get` timeout at fm-spawn.sh:2550 and `validate_spawn_worktree` at :2487) also exit with the endpoint live and direct the operator with "inspect window $T"; spawn_abort_cleanup only reclaims orca endpoints (already covered via ORCA_ABORT_CLEANUP) and herdr projections. The temp-root half was genuinely introduced by this PR: the gate was placed beside the busy-state arm, ~30 lines after `mkdir -p "$TASK_TMP/gotmp"`, and fm-teardown can only find that root through `tasktmp=` in a meta record a refused spawn never publishes. Root-cause fix (smallest correct change, no new subsystem): - bin/fm-spawn.sh — moved the `claude*` trust gate from inside the busy-arm block up to the first point $WT is known, immediately after the `freshen_spawn_worktree_base` block and before TASK_TMP creation, the STATE setup, and the relaunch `clear_relaunch_harness_wiring` retirement. A refusal now leaves no temp root, no retired relaunch wiring, and no busy record; only the endpoint remains, in the same class as the two refusals just above it. - bin/fm-spawn.sh — the refusal message now ends with "inspect window $T", matching the existing convention so control/teardown can identify the endpoint. $T is set for every backend on the non-secondmate path. - bin/fm-spawn.sh:196 — header note corrected from "before any state is armed" to "before any per-task state exists". - tests/fm-claude-trust.test.sh — the existing refused-spawn test's own comment claimed "before any task state exists" but only asserted busy state. Renamed to test_refused_spawn_leaves_no_task_state and added an assertion that /tmp/fm-<id> is absent, with the task id suffixed by the test process pid so the assertion reads only this run's path (a stale /tmp/fm-refusedspawn from the fixed-id version was in fact present on this box). No assertions on implementation source bytes. Verification run locally: - The new assertion fails against the pre-fix bin/fm-spawn.sh ("not ok - a refused spawn stranded a temp root no teardown can find") and passes after — a real before/after regression proof. - tests/fm-claude-trust.test.sh: 20/20 ok. - tests/fm-backend.test.sh, fm-backend-orca, fm-control-relaunch, fm-spawn-dispatch-profile, fm-trace-context-spawn, fm-gotmp: all pass. - tests/fm-backlog-atomicity.test.sh: rc=0, 79 assertions ok. - bin/fm-lint.sh (repo's single lint owner, pinned ShellCheck 0.11.0 + actionlint 1.7.12): clean. - No /tmp/fm-refusedspawn* leftovers after the runs. Scope respected: no trust subsystem, no policy layer, no config surface, no endpoint-cleanup mechanism added; the change is an ordering move plus one error-message clause and the test that pins it. Adapter references and docs made no ordering claim, so none needed updating. Changes are left uncommitted in the worktree for the outer executor --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix: restart every live second mate after updates (#3690) * feat(update): restart every live second mate after a successful update /updatefirstmate only restarted a second mate when that pass advanced its AGENTS.md or .agents/skills. An already-current home was skipped entirely, a bin/-only advance was steered instead, and a remote host that could not report its instruction diff was downgraded to a re-read. A running agent also freezes its launch-time wiring - turn-end hooks, harness flags, per-harness feature switches - and none of that is derivable from a file diff, so an unchanged tracked surface is not evidence the agent is already on the current behavior. Restart is now unconditional on a successful update of that home. Every live second mate the pass leaves on the target commit is restarted, whether it advanced or was already there. The safety contract is unchanged: open records are persisted before the agent is replaced, nothing is forced, stashed, or discarded, a home the pass had to skip is not restarted at all, and a mate whose runtime cannot prove a restart keeps the honest re-read path and is never reported as reloaded. bin/fm-ff-lib.sh gains a settled-state hook that fires for a home left at the base whether it advanced or was already there, and never for a skipped one; the instruction-gated hook the session-start convergence sweep uses is untouched. Regressions: fm-update pins the already-current mate into the restart set and the unprovable one into the nudge set, and fm-secondmate-restart drives both real commands end to end - an already-current home is named, persisted, and genuinely replaced with its checkout untouched, while the unprovable one keeps its running agent. * no-mistakes(document): Document unconditional secondmate restarts * fix(bin): close pending-reply decisions via resolve-key (#3696) * fix(bin): close reserved pending-reply keys via fm-send --resolve-key fm-send wrote answered: notes that the reserved-key fold ignores, so operator closes exited 0 while OPEN DECISIONS kept the decision open. Speak the owning library's close vocabulary on that path, and refuse when a reserved close cannot take effect. * no-mistakes(review): Safely quote manual decision-close recovery commands * no-mistakes(review): Reject unclosable overlong decision keys before sending * no-mistakes(review): Remove contract suffix from open decisions hint * no-mistakes(document): Document resolve-key line-cap refusal * fix(bin): prevent false missed-reply escalations (#3697) * fix(bin): stop false missed-reply escalations for same-basename self-home answers A healthy secondmate that wrote corr= to its own state/<id>.status never matched the parent channel, so recovery confirmed and the record escalated as pending-reply-missed. Make the report helper resolve the parent channel itself, skip parent-replies.status as wrong-home, put a readable sighting path on the missed line, and restatement-copy only that same-basename self-home file onto the parent channel. * no-mistakes(review): Resolve late replies before recovery escalation * no-mistakes(review): Tighten reply routing and regression coverage * no-mistakes(review): Preserve reply paths and require explicit home * no-mistakes(review): Encode wrong-home paths before persistence * no-mistakes(document): Document corrected secondmate reply routing * no-mistakes(lint): Fix pending-reply ShellCheck warnings * feat: add verified Gemini crewmate runtime (#3695) * feat(harness): verify gemini as a crewmate runtime adapter Adds Gemini CLI as a fourth dispatch target alongside claude, codex, and grok, scoped to crewmate and scout work only. Every axis was proven against gemini-cli 0.58.0 rather than inferred; docs/verification/runtime-backends.md carries the dated evidence and names what stayed unverified. Busy state is semantic, not rendered: BeforeAgent opens a turn and AfterAgent and SessionEnd close it. AfterAgent also fires on a manual interrupt, so a cancelled turn closes its own record. Three findings shaped the wiring rather than a config line: - --skip-trust and GEMINI_CLI_TRUST_WORKSPACE=true are presented by the CLI as equivalents and are not. A controlled A/B showed --skip-trust leaves project configuration unloaded, so workspace skills never load. - The worktree's .gemini/settings.json is the PROJECT's committed settings file, unlike claude's settings.local.json. Firstmate's hooks therefore go to a firstmate-owned state/<id>.gemini-settings.json reached through GEMINI_CLI_SYSTEM_SETTINGS_PATH, which also works untrusted and merges with a project's own hooks instead of replacing them. - The shipped CLI is a node bundle whose live process reports comm as MainThread, so ancestry cannot see it. GEMINI_CLI=1 is load-bearing and is tested before an inherited CLAUDECODE, and pane liveness identifies gemini from the script argument through the new bin/fm-gemini-lib.sh. Gemini is refused for secondmates: it has no primary supervision protocol. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * test: clear gemini's marker in launch and detection expectations Every non-gemini launch now clears GEMINI_CLI the way it already clears cursor's markers, so the two tests that pin the exact launch prefix are updated to match. The harness-detection tests that scrub foreign markers before probing ancestry scrub GEMINI_CLI too, so running the suite from inside a gemini session cannot produce a false verdict. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * docs: classify the gemini harness reference The documentation inventory is the single classification owner for maintained prose surfaces, and every surface must appear in it exactly once. The new harness reference is agent-runtime, matching its siblings. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01GYdQkKfSEQrFUxtcTXZ66L * no-mistakes(review): Narrow Gemini ancestry detection * no-mistakes(review): Restrict Gemini hooks to canonical launches * no-mistakes(document): Document Gemini adapter support boundaries * no-mistakes(ci): Fixed Gemini process identity when interpreter or script paths contain whitespace. Tmux liveness now uses NUL-delimited /proc argv on Linux, with the existing flattened ps fallback elsewhere. Added a real-process regression test. Verified with the Gemini harness test suite, full fm-lint, ShellCheck, and git diff --check. The CI and Require no-mistakes runs were action_required/attestation outcomes rather than code failures --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(teardown): conclude parked runs advanced past task copy (#3704) * conclude parked runs the pipeline advanced past the task copy A no-mistakes fix round commits in the daemon's own gate-repo clone, so a run parked at a gate can carry a head whose object the task copy never received. Teardown's strict object-local identity rule then declined to conclude the run, and cleanup left it parked forever holding a fleet slot (observed 2026-09-03; the same masking condition PR 3681 fixed on the read path, now closing the teardown half its scope boundary deferred). task_status_is_own_parked_run now falls back - only when the reported head resolves to no local object - to the one shared runs-ledger attribution rule fm_nm_runs_status_for_worktree (bin/fm-nm-run-lib.sh), whose anchored continuation proof binds the branch's newest active row to this worktree's exact submitted head. Foreign branches, stale history, terminal rows, ancestor-only anchors, diverged newer rows, and ambiguous multi-row shapes all still refuse, and runs that are actively running, fixing, or in CI remain untouched: only the parked-at-a-gate determination ever reaches the abort. No sqlite access, no fetches into another task copy, no custody changes, no duplicated matching logic. * tighten the parked-run ledger fallback and pin both judge corrections The teardown ledger fallback now authorizes concluding this task's parked run only when the shared runs-ledger rule's proved answer is the explicitly active word (running): a terminal newest row - even anchored at exactly the worktree's head - is finished history and never an abort authorization. The read path may classify the same owner's answer; teardown's abort must never fire for a run that already ended. Two bounded pre-validation corrections from the implementation review: - a fetched-object counterfactual pins the strict-rule path: a pipeline fix head fetched into the task copy aborts through object-local identity alone, with an empty ledger and a proof the runs query never fired; - a negative fixture pins the tightened boundary: an unresolvable reported head with a terminal newest same-branch row anchored at the worktree head engages the ledger fallback and still refuses, so the refusal is the terminal-word boundary and not an earlier guard. * no-mistakes(review): Bind teardown ledger fallback to validated run heads * no-mistakes(review): Restore validated advanced-head ledger continuation * no-mistakes(review): Reject invalid ledger dates and terminal statuses * no-mistakes(document): Document teardown ledger scan limit * fix: classify captain holds from structured state (#3508) * Keep parked and aged undated captain holds off live Captain's Call. Bearings was treating undated parked-style holds as live calls; mark those phrasings deferred and project holds older than a configurable 14-day since date as Charted Next gates instead. * no-mistakes(review): Bound parked marker matching to lexical tokens * no-mistakes(review): Age undated holds from durable hold-set dates * no-mistakes(review): Reset re-held timestamps and scan full bodies * no-mistakes(review): Preserve timestamp precision and prioritize parked suppression * no-mistakes(document): Document undated captain-hold aging * no-mistakes(ci): Fixed stock Bash CI test-count expectations (16 snapshot, 45 Bearings). Prevented fresh holds on old tasks from aging via stale `since` dates by aging only stamped holds. Added behavioral regressions and verified both suites plus Bash 3.2 parsing * no-mistakes(ci): account for rebased snapshot regression * no-mistakes(review): Restore legacy hold aging and mandate wrapper * no-mistakes(review): Restrict hold stamps to canonical leading lines * no-mistakes(review): Exclude historical answers and deduplicate revealed holds * no-mistakes(document): Correct captain-hold projection documentation * no-mistakes(ci): Rebased onto 8988af2 and resolved Bearings conflicts. Fixed the hold timestamp race by persisting and verifying the timestamp before publishing the captain hold; failures now leave the task unheld. Added behavioral coverage for ordering and failure handling. Preserved the required parked-phrase projection behavior. Relevant snapshot, Bearings, lifecycle, syntax, and ShellCheck validations pass * no-mistakes(review): Bound current prose before historical resolutions * no-mistakes(review): Preserve hold age across interrupted answers * no-mistakes(review): Preserve leading hold stamps until answer closure * no-mistakes(review): Normalize answer bodies on matching retries * no-mistakes(review): Document concurrent re-hold age-basis limitation * no-mistakes(document): Refresh captain hold lifecycle documentation * no-mistakes(ci): Fixed both CI failures. Updated the macOS Bash snapshot expectation from 45 to 46 Bearings tests. Narrowed parked-style deferral matching to explicit hold-reason prefixes while preserving legacy explicit markers and preventing contextual prose from hiding active decisions. Added behavioral regression coverage. Verified with stock Bash 3.2: 17 fleet snapshot tests and 46 Bearings tests pass; full lint and workflow validation also pass * no-mistakes(ci): Fixed Greptile’s P1 finding by restricting parked-style deferral phrases to complete hold-reason markers. Contextual reasons beginning with “not urgent,” “queued opportunity,” or “captain-gated” now remain visible decisions. Added behavioral coverage through the real fleet and Bearings snapshot paths and updated documentation. Verified both snapshot suites under Bash 3.2 (17 fleet tests and 46 Bearings tests), syntax checks, and git diff checks. The no-mistakes attestation failure is external/stale and requires the outer pipeline to refresh it for the new head * no-mistakes(ci): Fixed parked-style undated captain holds disappearing from the default Bearings board. They now project to Charted Next with omitted[] disclosure, while --all-decisions reveals them and removes the safety gate. Added behavioral coverage for the reported “not urgent” case and aligned documentation. Verified fm-bearings-snapshot, fleet snapshot view, and captain-hold lifecycle tests; shellcheck, bash syntax, and git diff checks pass * no-mistakes(test): Stabilize concurrency budget and provision timeout tests * no-mistakes(document): Correct captain hold documentation details * no-mistakes(ci): Fixed hold-reason parsing so commas in contextual reasons are preserved and do not incorrectly defer live Captain's Call decisions. Added end-to-end fleet/Bearings regression coverage. Reworked the flaky Herdr timeout test to assert observable late-launch behavior rather than process-ID liveness. Verified both snapshot suites, Herdr test 5 consecutive times, shell syntax, shellcheck, and git diff checks * Restore the Herdr lab timeout test to its main version. The stabilization rounds reworked tests/fm-herdr-lab.test.sh while chasing a load-induced flake, replacing the fake server's wall-clock delay with a SIGSTOP'd process and asserting that the blocked process is gone after a timed-out provision. A stopped process does not die from SIGTERM, so that assertion fails on Linux and the portable parallel shard stayed red. That test is unrelated to the undated captain-hold projection this branch delivers and was identical to main before these rounds, so restore main's version exactly. It still proves that a timed-out provision cancels its late launch before teardown. * no-mistakes(review): Preserve metadata-like prose in captain hold reasons * no-mistakes(review): Resurface due dated captain holds * no-mistakes(review): Distinguish parked holds from explicit deferrals * no-mistakes(review): Invalidate legacy secondmate summary caches * no-mistakes(review): Keep blocked deferred holds in Charted Next * no-mistakes(review): Count blocked deferred holds in omission disclosure * no-mistakes(document): Correct captain-hold projection documentation * no-mistakes(ci): Fixed both CI failures. Updated the macOS Bearings test count to 51. Preserved the v1 summary schema for compatibility while rejecting hold-bearing summaries missing the new aging fields, preventing stale caches from restoring noisy calls. Verified fleet snapshot, Bearings snapshot (51 tests), home-summary refresh, secondmate reconciliation, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed Greptile’s valid finding: `--all-decisions` now reveals deferred/aged captain holds even when blocked, for both main and secondmate homes, and removes their duplicate Charted Next gates. Added behavioral regression coverage and updated documentation. The prose-classifier finding was not applied because exact complete-phrase matching is explicitly required by the author intent; contextual wording remains live. Verified with Bearings and fleet snapshot tests, `bin/fm-lint.sh`, Bash syntax checking, and `git diff --check` * no-mistakes(ci): Fixed the actionable-state bug in Bearings: an arrived parked-style hold is live only when it is not explicitly non-actionable, so blocked due holds remain gated by default and are revealed by --all-decisions. Added behavioral regression coverage for that case. Preserved complete-reason parked-style classification as required by the author intent. Verified with tests/fm-bearings-snapshot.test.sh, bin/fm-lint.sh, and git diff --check * Show why a revealed captain hold is deferred. Under --all-decisions a deferred hold is revealed and its Charted Next gate is removed, but the revealed row carried only the bare hold reason. A date-deferred or blocked hold therefore read exactly like a genuine live decision, because the until date, the age, and the blocking work only ever appeared on the gate row that the reveal replaces. Annotate a row that is revealed because it is deferred with the same vocabulary the gate uses - until <date>, held <n>d, and the blocking work - so the expanded view reads as deferred-but-shown. A genuinely live call is left unannotated, and the default board is unchanged. * Classify captain holds from structured fields alone. Bucket membership was decided by several independent expressions, and two of them matched hold reason or body prose. That produced a recurring class of defects: holds that fell through every bucket and vanished from the board, and live decisions silently suppressed because their wording happened to contain a marker word - a reason of "non-deferred release choice" matched DEFERRED and disappeared. Replace all of it with one total classifier over structured fields only: hold_kind, state, hold_until, unresolved_blocker_ids, and the machine-written hold-set timestamp. Every captain hold gets exactly one hold_bucket - blocked, dated, aged, or live - so no hold can fall through and none can match two. captain_actionable is exactly the live bucket, and the --all-decisions reveal is a property of the bucket rather than a second filter. No hold reason or body prose is matched anywhere in the projection, so wording can no longer hide, reveal, or reclassify a decision. A hold that is superseded or no longer required is closed through the hold lifecycle instead of lingering as an open hold flagged by a keyword. * no-mistakes(review): Preserve working captain holds across bucket surfaces * no-mistakes(review): Reject pre-classifier secondmate summary caches * no-mistakes(review): Preserve complete live hold summaries * no-mistakes(review): Clarify working hold decision bucket semantics * no-mistakes(review): Reveal bounded remote holds and preserve blocker notes * no-mistakes(review): Make blocker overflow explicit in hold summaries * no-mistakes(document): Correct captain-hold projection documentation * no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 17 to 18 tests. Verified the suite under Bash 3.2.57: all 18 tests pass. `git diff --check` also passes * no-mistakes(ci): Updated the stock macOS Bash CI expectation from 51 to 53 Bearings tests. Verified all 53 pass under Bash 3.2.57; git diff --check passes * fix(pi): keep supervision outcome delivery responsive (#3767) * fix(pi): deliver supervision outcomes off Pi's render thread The supervision branch runs inside the captain's own Pi process, and Pi runs extensions, their tools, and their event handlers on the single JavaScript thread that also draws the TUI and reads the keyboard. Every delivered outcome ran roughly five bash script invocations plus several `ps` calls through spawnSync on that thread, so the TUI could not repaint or echo a keystroke for the whole chain - the subsecond freeze the captain saw every time a routine or captain-facing outcome arrived. Convert the delivery path's subprocess calls to an awaited spawn behind a serializing queue. lib/fm-async-exec.ts is the single owner of the awaited-spawn replacement and returns the same capture shape and failure verdicts spawnSync returned. Awaiting yields the thread, so what the single thread used to guarantee for free is now an explicit queue: every delivery, acknowledgement, and turn-boundary reconciliation runs as one unit of it, preserving the durable append before anything visible, one delivery at a time in sequence order, the read cursor advanced before the next reader sees a row, and one ownership activation per generation. Cancellation is preserved by the generation and lock-ownership rechecks the awaits are placed around. Two reads stay synchronous because Pi's own API is synchronous there, not as an optimization: its bash spawn hook is typed as a plain function, and the watcher reads offer.accepted the moment its dispatch event returns, so a session that does not own the fleet lock must still refuse a wake without waiting. Both walk the lock's process ancestry in full every time, never cached, because reparenting and pid reuse can invalidate a remembered chain and that answer decides ownership rather than hinting at it. The store scripts and their durability contracts are unchanged. Measured through the real fm_branch_report tool and real bin/ scripts with a 1 ms interval timer, the largest block of the JS thread falls from 273 to 2.0 ms for a routine outcome, 286 to 2.0 ms for a captain outcome, and 134 to 1.9 ms for main's acknowledgement, against a 1.3-2.2 ms idle floor. In a real Pi 0.82.0 TUI the worst keystroke echo while two outcomes arrive falls from 676.9 ms to 36.8 ms, against a 22.6 ms extension-free floor. Regressions: a delivery must leave the event loop running (zero timer ticks before this change, in 250 ms), interleaved reports stay ordered and exactly once, a session replaced mid-delivery neither loses nor duplicates an outcome, and a failing store script surfaces without losing or doubling one. The real-TUI half is an opt-in live guard that types into an isolated Pi pane while outcomes are delivered and fails if echo leaves the class of the same machine's own floor. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013bzoWyr2EcJGBKuoUjVRSp * no-mistakes(review): Revalidate ownership and bound asynchronous subprocess output * no-mistakes(ci): Fixed CI defects: routine outcomes now persist a sequence-keyed delivery receipt before awaiting cursor advancement, preventing duplicate delivery after mark-read failure. Corrected the session-replacement test to exercise an actual asynchronous ps ancestry lookup. Targeted behavioral tests, strict Pi typecheck, ShellCheck, and diff checks pass. The full extension test remains locally blocked by an unrelated stock-render assertion under the installed Pi runtime. The no-mistakes attestation failure is external pipeline state (test was previously skipped), not a source defect * fix(pi): keep the declined routine receipt out and skip the renderer case below its Pi floor Four follow-ups on the same branch, plus one revert. Revert the routine-delivery receipt a CI auto-fix round added. It introduced a new persisted `fm-branch-routine-delivery` entry, written into the captain's transcript for every routine note, to deduplicate a note whose cursor write failed. That is a change to the delivery contract, which this task is not authorized to make: the approved work is the asynchronous conversion with the existing durability contract preserved. The ownership re-read and output bounding from the review round are kept - both are genuine asynchronous correctness, not contract changes - as is that round's use of a real parent pid so the replacement regression traverses an actual ps subprocess. Record the routine gap instead of closing it. A routine note is a plain message with no sequence-keyed record, so a mark-read failure after delivery makes the next reconciliation send it once more; a captain row cannot duplicate that way because its visible entry is found by store sequence. That asymmetry predates moving delivery off the render thread. It is now stated at the call site and in the delivery-contract docs, tracked as fm-pi-routine-delivery-idempotency-followup-r1, and pinned by a regression that proves the routine note is re-delivered exactly once more and never again, the captain entry stays single, and the store keeps both rows. Give the stock-renderer case a Pi version floor. It compares the extension's renderers against Pi's stock rendering, so its verdict only means anything against the contract those renderers target: since 0.84.4 the stock renderer no longer supplies an implicit reset at multiline boundaries and the extension emits that reset itself, so an older installed Pi differs legitimately. It now names the installed version and the floor and skips, while a package whose version cannot be read at all still fails. Make the responsiveness regression's second signal a fraction rather than a millisecond budget. A loaded machine that deschedules the process inflates an absolute stall budget into a false failure, but it inflates the delivery's own wall time too, so requiring the worst stall to be a minority of that wall time holds under load. Synchronous delivery sits near 1.0 there whatever the load, and the tick-count signal still reads zero on it. Replace the test-family mapping for the Pi extension libraries with per-script targeting. Routing them to whole families - or leaving them unmapped, which widens through the reference scan to each referencing suite's entire family - selected dozens of suites with nothing to do with Pi and pulled an unrelated flake into the run. The changed-file selection drops from 112 scripts to 61. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013bzoWyr2EcJGBKuoUjVRSp * no-mistakes(document): Clarify asynchronous execution documentation --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com> * fix: avoid duplicate AGENTS.md governance for marked projects (#3763) * fix(memory): honor explicit project maintenance guidance * no-mistakes(test): Blocked by pre-existing Bash and Muse fixture failures * no-mistakes(test): Remove accidentally tracked test attribution report * no-mistakes(ci): Restricted the marker to the exact first line, preventing fenced examples from suppressing governance, and corrected the documentation. Regression failed before the fix; all 18 helper tests, focused ShellCheck, documentation validation, and diff checks pass. CI and Require no-mistakes report action_required with zero jobs executed; those external checks remain unresolved * fix: protect primary checkout when spawning from linked homes (#3783) * fix(spawn): refuse repository primary from linked spawning homes Compare the resolved task git directory with the spawning repository's common git directory before refreshing a fresh copy or relaunching a task. This protects the primary even when the spawning project is a linked home. Keep pooled copies accepted and preserve recorded work on relaunch. Fixes #3741. Verification for the pipeline PR body: - Red on origin/main 1820316b66ac2c68e244dd04a02512859ee8c1f4 with the new regression and unchanged production code: bin/fm-test-run.sh tests/fm-spawn-pool-base-freshen.test.sh exited 1 with "linked spawning home accepted primary as a disposable copy". - Green after the guard: the complete pool-base-freshen and control-relaunch suites passed through bin/fm-test-run.sh, covering primary and symlink refusal before fetch/reset, spawning-directory refusal, scout acceptance, and committed plus unfinished work preserved during linked-home relaunch. - The worktree-settle suite passed on pristine main and the final branch. An earlier loaded-host run exceeded its five-second assertion (6s); the final retry passed without changing code or the assertion. - Test fixture commits ran with GIT_CONFIG_COUNT=1, GIT_CONFIG_KEY_0=commit.gpgsign, GIT_CONFIG_VALUE_0=false. - bin/fm-lint.sh and /bin/bash -n for all three changed scripts passed. The upstream cwd-selection cause remains outside this change. * no-mistakes(document): Clarify spawn isolation ownership and relaunch preservation * fix(bin): stop reading an unanswered backend probe as a dead endpoint (#3785) * fix: read a failed herdr CLI as unreachable, not a gone backend target The no-run fallback in bin/fm-crew-state.sh collapsed every failed pane capture into 'backend target gone', which downstream consumers treat as positive death evidence - so a herdr CLI that errors or stalls under load briefly scored dozens of live claims dead on a busy box. Only a successful herdr answer proving the pane absent (fm_backend_agent_state's 'missing', backed by pane get answering pane_not_found) may now read as gone; every other verdict reports 'backend unreachable' with the endpoint state, which is never positive death evidence. Adds a behavior test: an always-failing fake herdr reads unknown/unreachable, never gone. * test: pin the herdr suite's ambient home to a marker-free fixture FM_HOME defaults to the suite's own root when unset, and any secondmate- marked checkout (every treehouse crew home carries .fm-secondmate-home) flips the default workspace label to 2ndmate-*, so the ambiguous-label placement test found zero firstmate matches and fell into the create path instead of refusing (expected exit 3, got 1) - deterministically green in CI, deterministically red from a crew home. Export a marker-free ambient FM_HOME fixture; per-test FM_HOME prefixes still override it. * fix: classify herdr endpoint answers instead of every non-missing verdict Review decision (firstmate, 2026-09-05): a failed pane capture is not itself evidence of death, but neither is every non-missing classifier verdict a failed answer. missing (pane get answered pane_not_found) and dead (pane present, agent_not_found husk) keep gone-class text so a stale-claim sweep may still reclaim them; an alive answer falls through to the normal busy/state flow instead of being discarded when only the heavy 200-line scrollback read failed; only when the cheap pane get / agent get calls themselves fail to answer does the line read 'backend unreachable'. Adds the two missing cases: alive with a failed scrollback read stays live, and a husk pane still reads gone. * no-mistakes(review): route tmux through agent-state classifier; drop test stall * no-mistakes(review): narrow inaccurate tmux socket and alive-arm fallback comments * no-mistakes(document): document classifier-backed endpoint verdicts in crew-state contract * fix(bin): preserve subshell lock ownership on Bash 3.2 (#3789) * fix: distinguish subshell wake-lock owners on stock Bash Restore distinct process ownership for issue #3743 using the existing PID helper, consistently across lock publication, reclaim, release, role checks, and bounded handoff. The existing wake-queue regression fails on pristine upstream Bash 3.2 with rc=13. The complete suite now passes on Bash 3.2.57 and Bash 5.3.15, with added coverage for ownership when BASHPID is unset. Canonical lint and stock-Bash syntax checks pass. * no-mistakes(document): Correct lock grace-period documentation * no-mistakes(ci): Captain, fixed all 14 SC2031 false positives with nine ShellCheck source-boundary annotations across three tests. Full CI-mode lint and the complete wake-queue suite on stock Bash 3.2 passed. Runtime behavior is unchanged * fix(bin): resolve captain holds and legacy teardowns on non-markdown backends (#3782) * fix(bin): close legacy records on the Beads backend honestly Two pre-Beads reads blocked honest closure of leftover records: 1. fm-captain-hold.sh complete/verify resolved attested legacy hold ids only against the live backend and the pre-collapse derived identity, so a home whose holds fm-hold-migration rehomed under fm- ids failed with an empty-name absence message (the resolve failure was swallowed by the command substitution feeding verify_hold_durable). Resolution now falls …
…ts resolved (#19) * fix(bin): support process events under symlinked homes (#3484) * fix(bin): resolve process-event state roots before validating them The process-event module validated the caller's spelling of a home's state root instead of the directory it operates on: it required the supplied path to equal its own lexical normalization, which rejects any path reached through a symlinked ancestor. On macOS both /tmp and $TMPDIR are symlinks, so an operator home under either could never claim a source. Reconcile still reported the runner started, while the detached runner died writing "cannot claim source" to the discarded stderr, and the source silently never fired. Resolve the state root to its physical directory once, then apply the existing private-directory validation to that resolved directory and derive every path, recorded claim identity, and later confinement check from it. This keeps the confinement contract for the directory actually operated on rather than only for callers that already spelled it physically, and removes the window where an ancestor symlink could be repointed between check and use. Homes already spelled physically behave identically. This was the single cause of both deterministic macOS failures in tests/fm-procevent.test.sh ("reconcile never claimed the registered source") and tests/fm-procevent-when.test.sh ("the winning concurrent arm did not produce an outcome"). The new case pins the behavior with an explicit symlinked-ancestor home, so it fails without the fix on any platform rather than only where the temp root happens to be a symlink. * fix(bin): pin the external capture staging boundary to its physical path The extension capture path pinned its registry staging boundary by comparing `pwd -P` against the caller-spelled registry directory, so a home reached through a symlinked ancestor still refused to start an extension-backed source after the state root itself resolved correctly. That left such a home half working: built-in sources ran while external ones failed. The staging preparer now prints the physical registry directory it validated, matching the inbox and reservation preparers beside it, and the start path pins on that returned path. The new end-to-end case drives the shipped file-signal package from a symlinked home spelling. * no-mistakes(review): Propagate canonical process-event state roots * no-mistakes(review): Propagate canonical state to process-event adapters * no-mistakes(document): Document physical process-event state roots * fix(pi): deliver captain outcomes as deterministic transcript entries (#3312) * fix(pi): persist captain outcomes visibly * no-mistakes(review): Recover captain outcomes after cold-start lock acquisition * no-mistakes(document): Document cold-start captain-outcome recovery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Prove immediate Pi captain-outcome transcript delivery * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix(pi): process captain outcomes through a sequence-keyed turn PR #3312 made every captain-facing supervision outcome a durable, exact-once visible transcript entry with the read cursor advancing only after that entry exists. That is the display half of the delivery contract. Left alone it turns a probabilistic silent loss into a deterministic one: the captain sees an anchor line, and firstmate never acts, because nothing opens a turn and nothing records whether main ever processed the outcome. The 2026-08-31 timeline showed the two shapes this must survive on the previous hidden-turn path: seven delivered decision outcomes each answered by an empty assistant message (cursor advanced, no retry, unanswered for close to three hours), and two answered by an unrelated prior reply. Both happened because delivery advanced the cursor at enqueue and accepted whatever the next assistant message was. Add the processing half on top of the persistence half: - bin/fm-branch-outcome.sh keeps a processed marker separate from the read cursor (`unprocessed`, `mark-processed --through`, `processed-init`). It only advances through an explicit sequence-bound acknowledgement, never past the read cursor and never backwards; an absent marker reads as zero and `processed-init` migrates delivered history once so an upgraded home is not re-presented its past. - After the visible entry for a captain outcome exists, the extension hands every still-unprocessed captain row to main as one hidden, typed `fm-branch-process` request listing each `[seq N] task: summary`, opening exactly one main turn. Main closes it only by calling the new `fm_branch_processed` tool with the highest sequence listed. An unrelated, empty, or paraphrased answer leaves the sequence open, and the same request is presented again at the end of the next main run and at session start. The first two presentations of a sequence set open a turn of their own; after that the request rides the captain's next prompt so an ignored request cannot loop, and a session replacement resets that budget. Routine outcomes stay turn-free. - The regressions cover exactly those incident shapes against the real store scripts: an empty answer and an unrelated prior answer neither advance the marker nor stop re-presentation, the acknowledgement is refused beyond the read cursor and outside lock ownership, a partial acknowledgement keeps the newer sequence open, and #3312's own assertions now forbid an unkeyed turn rather than any turn. The store suite pins the marker's bounds and the migration; the real-SDK guard for appendEntry persistence and model exclusion is unchanged. Docs move the protocol from "no model turn" to "one sequence-keyed processing turn closed only by its acknowledgement", and the verification record carries the dated run against Pi 0.84.4. * no-mistakes(review): Harden outcome listing and sequence-bound acknowledgements * no-mistakes(review): Harden outcome state validation and request pacing * no-mistakes(review): Reject unsafe sidecars and unterminated outcome stores * no-mistakes(review): Validate canonical mark-read cursor state * no-mistakes(review): Guard cursor advancement against corrupt processed state * no-mistakes(review): Bind acknowledgements to active processing requests * no-mistakes(review): Reset pacing when processing sequence membership changes * no-mistakes(review): Enforce silent outcome invariants at storage boundary * no-mistakes(document): Document hardened captain outcome processing contracts --------- Co-authored-by: kunchenguid <kun@kunchenguid.com> * feat: add bounded concurrent Bearings ledger collection (#3481) * feat: bound Bearings remote ledger collection * no-mistakes(review): Clarify default remote-ledger collection behavior * no-mistakes(review): Detach reconcile delivery from watcher loop * no-mistakes(review): Enforce bounded snapshot and request captures * no-mistakes(review): Bound legacy summary capture before parsing * no-mistakes(review): Bound primary remote ledger captures * no-mistakes(document): Correct snapshot and reconcile documentation * no-mistakes(lint): Fix ShellCheck quoting in bounded collector * no-mistakes(ci): Fixed all three CI failures: updated the macOS Bearings assertion to 44 tests, made the home-summary test deterministic and aligned with default ledger consumption, and increased the asynchronous reconcile retirement wait for loaded CI. Verified both focused suites, all 44 Bearings tests, ShellCheck, actionlint, Bash parsing, and git diff checks * test: await reconcile request retirement * no-mistakes(review): Avoid empty reconcile queue process churn * no-mistakes(review): Read ledger summaries from immutable snapshots * no-mistakes(review): Reject multi-document home ledger streams * no-mistakes(review): Coalesce durable reconcile requests per target * no-mistakes(review): Unify reconcile keys and reject snapshot streams * no-mistakes(review): Key reconcile requests by stable target ID * no-mistakes(document): Document per-target reconcile request coalescing * no-mistakes(lint): Remove unused snapshot summary file variable * no-mistakes(ci): Adjusted the concurrent collector regression’s end-to-end timing ceiling to account for stock macOS process/jq overhead outside the three-second remote collection budget, while remaining below the 15-second serial-read floor. Verified with stock /bin/bash 3.2: all 44 Bearings tests pass; bash syntax and git diff checks pass * no-mistakes(ci): Fixed legacy summary validation to require exactly one top-level JSON document and added behavioral regression coverage. Stabilized CI by conditionally waiting longer for durable reconcile delivery and synchronously stopping the fm-on worker tree before fixture cleanup. Removed a redundant flaky healthy-path timing assertion; the wedged-reader test still proves concurrent bounded collection. Verified fm-bearings-snapshot, fm-secondmate-reconcile, and fm-on tests, plus project ShellCheck, bash syntax, and git diff checks * ci: rebalance portable serial test shards (#3489) * fix(ci): rebalance the portable serial shards on measured durations The "Behavior portable serial 3" shard ran 17-20 minutes against its 20-minute job cap and intermittently timed out seconds after a passing test, on branches and on main alike. Shards are packed longest-processing-time from per-script duration hints, and those hints were last measured on 2026-08-21 at 116 scripts. The lane has since grown to 139 scripts and from ~42 to ~63 minutes: 17 scripts had no hint at all and fell back to the 20 s default, and several existing hints were low by 2-5x (fm-watch-triage 142 s hinted vs 263 s measured, fm-public-followup 36 s vs 197 s). The partition therefore looked perfectly balanced in hint space, 734.6 s per shard, while really running 11.5, 13.6, 18.8 and 16.5 minutes. Script-count balance, which is what the tests asserted, stayed normal throughout and hid it. Refresh the hints from the timing artifacts of three green runs, taking the slowest measurement of each script so the balance holds on a slow runner, and split the lane across five shards instead of four. Replayed against those runs' real per-script durations the worst shard is now 12.54 minutes, 63% of the unchanged 20-minute cap, and the serial lane's wall clock drops from ~20 to ~12.5 minutes. Bound the drift that caused this rather than relying on the hints being refreshed by hand: the coverage guard now reports the unmeasured share as serial_unhinted= and refuses past PORTABLE_SERIAL_MAX_UNHINTED_PERCENT, which leaves room for newly added tests while making a stale table fail the guard instead of silently pushing one shard into its cap. No test changes what it asserts and no test stops running; only the partition across shards changes. * no-mistakes(document): Clarify conservative shard timing aggregate * fix(pi): fall back on incomplete supervision branch prompts (#3491) * fix(pi): fall back after settled branch errors * no-mistakes(review): Detect provider errors across prompt compaction * no-mistakes(review): Preserve in-flight branch state across selection changes * fix(pi): re-probe supervision branch after cooldown (#3497) * fix(pi): recover supervision branch after cooldown * no-mistakes(review): Defer branch recovery until prompt settlement * no-mistakes(document): Clarify supervision cooldown recovery contract * fix(bin): remove legacy remote snapshot reads (#3501) * refactor: remove legacy remote summary reads * no-mistakes(document): Document ledger-only snapshot reads * no-mistakes(ci): Fixed the snapshot test fixture so ledger refreshes use the same fake executable PATH as the snapshot consumer. This preserves observable endpoint freshness after removing legacy summary computation. Verified stock Bash parsing and all 44 Bearings tests pass under /bin/bash; git diff checks pass * no-mistakes(ci): Fixed the CI-only snapshot fixture failure by ensuring the bounded-ledger refresh uses its fake tmux backend. This removes host tmux availability as a source of nondeterminism. Verified all 44 Bearings tests pass, Bash syntax passes, and git diff checks are clean * no-mistakes(ci): Fixed CI nondeterminism in the Bearings fixture: all local ledger refreshes now use the fixture’s fake tmux backend when available, instead of depending on host tmux state. Verified stock /bin/bash syntax, git diff checks, and all 44 Bearings tests with a deliberately failing host tmux * fix(pi): preserve watcher continuity across session replacement (#3498) * fix(pi): rearm watcher after session replacement * no-mistakes(review): Queue actionable closes across Pi session replacement * no-mistakes(review): Stop replacement arm when handoff persistence fails * no-mistakes(review): Preserve actionable wakes through branch and late child races * no-mistakes(review): Surface late handoff failures without crashing Pi * no-mistakes(review): Coordinate replacement delivery settlement and unique handoff tokens * no-mistakes(review): Retry stale deliveries and release settled claims * no-mistakes(review): Distinguish branch settlement and retry handoff cleanup * no-mistakes(review): Deduplicate persistent handoff cleanup alerts * no-mistakes(review): Acknowledge watcher follow-ups only when consumed * no-mistakes(review): Persist idle follow-ups until agent consumption * no-mistakes(review): Preserve pending outcomes when handoff persistence fails * no-mistakes(review): Arm replacement before awaiting prior delivery settlement * no-mistakes(review): Adopt pending handoffs after lock reclamation * no-mistakes(review): Prevent stale generations from adopting replacement handoffs * no-mistakes(review): Scope replacement handoffs by watcher state * no-mistakes(document): Clarify replacement handoff documentation * no-mistakes(ci): Fixed the failing branch-extension tests to model the new settlement-promise contract. Failure cases now assert that delivery ownership returns to the watcher instead of expecting direct extension fallback. Verified the updated branch suite, Pi watcher suite, shell syntax, and diff checks * no-mistakes(review): Update branch settlement tests and preserve chunked outcomes * no-mistakes(document): Document watcher-owned replacement handoffs * no-mistakes(document): Verify replacement handoff documentation * test(pi): cover watcher-owned branch fallback * no-mistakes(document): Refresh watcher-owned fallback documentation * fix(bin): resurface task statuses missed by wake handling (#3495) * fix(bin): resurface terminal statuses lost after branch handling * test(watch): canonicalize process-event fixture homes * no-mistakes(review): Index branch outcomes by causal status position * no-mistakes(review): Recover outcome indexes and deduplicate resurfaced statuses * no-mistakes(review): Handle legacy ambiguity and oversized status diagnostics * no-mistakes(review): Keep unclassifiable oversized statuses silent * no-mistakes(document): Document lost-wake outcome backstop * no-mistakes(document): Update outcome backstop documentation * no-mistakes(ci): Fixed CI regressions in wake-drain: parseable reserved-key decisions can no longer bypass the durable decision-fold guard, and status output is prepared and receipt-committed before presentation to prevent repeated one-shot outcomes after later failures. Added a behavioral regression for receipt commit failure and retry. Targeted backstop, correlation-token, decision-cursor, open-decision, unread-status, syntax, and diff checks pass locally. Shard-4 failures appeared unrelated/flaky; the network-parallel test passed locally * no-mistakes(ci): Fixed the Greptile P1 data-loss issue by committing presentation receipts only after prepared output reaches stdout. Added behavioral coverage proving output failure leaves the backstop retryable and receipt failure may duplicate but never lose a presentation. Relevant wake-drain suites and syntax/diff checks pass. The shard-4 Pi extension failure is unrelated to this PR and did not warrant changes * no-mistakes(ci): Stabilized tests/fm-bootstrap-network-parallel.test.sh by replacing scheduler-sensitive equal-sleep timing with bounded synchronization between mocked fetch and remote probes. This preserves detection of real serialization while avoiding false failures under CI load. Verified with five consecutive test runs, bash syntax validation, ShellCheck, and git diff checks. The separate Pi stock-rendering failure reproduces locally but is unrelated environment/version drift * no-mistakes(ci): Fixed Behavior portable serial 4 by adding fm-classify-lib.sh and fm-timeout-lib.sh to the broken-root Pi test fixture; fm-branch-outcome.sh now depends on them. Verified the full Pi branch-extension suite with real-Pi checks skipped, the wake-drain outcome-backstop suite, Bash syntax, and git diff checks. Greptile findings are already addressed at HEAD; the no-mistakes attestation failure is external head-SHA state * fix(bin): collect follow-up results from remote work homes (#3503) * fix(bin): deliver typed terminal results from remote work homes A public commitment whose work is bound to a REMOTE secondmate home could never receive its typed terminal result. `fm-public-followup.sh brief` printed an emit command carrying this home's own absolute path and this checkout's own script path, neither of which exists on the machine the worker runs on, so the worker had nothing it could write to that the owning home would ever read - and `consume` kept finding nothing while the promise stayed open. The brief is now route-aware: for a remote work home it prints that route's own code root and home with `--stage-in`, so the typed event is staged in the home where the work actually runs, and the closing paragraph names the owning home as the one on the other machine instead of pointing at the path above it. The owning home collects those staged results over the same SSH route it reaches that secondmate on, because the transport only runs outbound: `consume` pulls them into its own inbox and reconciles them exactly as it reconciles a local report. Collection is non-destructive until the result is durably held, so a dropped connection cannot lose a terminal result, and a route that could not be reached is named in `consume`'s output with the promise left open rather than reported as an empty inbox. A local work home is untouched: the brief still prints `--home` with this home and this checkout's script, and the event still lands directly in this home's typed terminal-result inbox. This is the emit-side counterpart of the retire/clear fix in #3479 and reuses the remote-route resolution that landed with it. Reconciling a loop bound to a remote route now reaches that route, so the existing remote cases drive `consume` through the same faked transport their other steps already use. * no-mistakes(review): Fail loudly on unresolved routes and invalid staging homes * no-mistakes(review): Fail collection when remote outbox is unreadable * no-mistakes(review): Surface reassigned remote routes during empty collection * no-mistakes(review): Fail remote collection on invalid registrations * no-mistakes(review): Reject unsafe registration entries during remote collection * no-mistakes(review): Restore healthy empty remote collection behavior * no-mistakes(review): Skip remote collection for delivered registrations * no-mistakes(review): Skip delivered registrations before route validation * no-mistakes(document): Document remote follow-up collection semantics * fix(bin): exclude secondmates from home-summary validity (#3504) * fix(bin): exclude secondmates from home-summary child inventory kind=secondmate meta records never have backlog rows, so counting them in unowned_children or terminal_in_flight made a clean main home look invalid once earlier ledger checks passed. * no-mistakes(review): Cover terminal secondmate in-flight exclusion * no-mistakes(ci): Updated the stock macOS Bash CI snapshot expectation from 15 to 16 tests. Verified all 16 snapshot/fleet-view tests pass under Bash 3.2.57 and `git diff --check` succeeds * fix(bin): self-heal outcome indexes on first drain (#3509) * fix(bin): self-heal status-outcome indexes on every drain Missing ready markers were skipping the lost-wake backstop on non-Pi homes because only the Pi branch ran processed-init. Drain now rebuilds those indexes under the outcome lock and fails closed only on a real store fault. * no-mistakes(review): Guard held-lock initialization and fail marker writes * no-mistakes(document): Document cross-harness outcome-index self-healing * fix(bearings): keep active children underway during captain holds (#3505) * fix(bearings): keep active children underway beside a captain hold Project each readable home's active children into Underway independently of the home-level captain-decision classification so a hold no longer hides live work. * no-mistakes(review): Preserve Underway repos and disclose child truncation * no-mistakes(review): Fall back to task project for Underway repos * no-mistakes(ci): Updated the stock macOS Bash CI assertion from 44 to 45 Bearings tests, matching the newly added behavioral regression. Verified all 45 tests pass under /bin/bash, Bash syntax checks pass, and git diff validation is clean * fix(pi): settle watcher delivery on Pi accepting the follow-up (#3513) * fix(pi): settle watcher delivery on Pi accepting the follow-up A follow-up queued while main is streaming joins the running run without ever raising before_agent_start, so waiting on that event before clearing the successor pipeline (#3498) stalled every later actionable close: no successor started, no wake was delivered or offered to the branch, and the turn-end guard woke main to re-arm by hand after every close. The pipeline now settles once Pi accepts the follow-up. Consumption is observed at before_agent_start for an idle main and at the user message_start for a streaming main, and decides only what a replacement session (/new, /resume, /fork, reload) replays. An exhausted restoration delivers its typed failure without launching an arm past the retry bound, which the stall had hidden. The replacement-coordinator map is typed so the strict no-emit typecheck passes again. Tests: the doubles no longer raise before_agent_start for a streaming send, a portable regression drives two actionable closes while main streams and proves the successor chain plus consumption-scoped replay, and a credential-free real-SDK probe pins Pi's event contract for both the streaming and the idle follow-up. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(pi): retry a verified successor that fails during wake delivery A verified successor can exit while the wake it was started for is still being delivered, most plausibly during a branch turn that holds the settlement for minutes. Its failure close arrived while the pipeline's single-flight guard was set, so the close handler skipped the retry, and the pipeline's end no longer launched an arm, which left the live generation with no watcher and no retry timer. The close handler now records that failure when the child had reported readiness and was not retired by the restoration itself, and the pipeline runs the ordinary bounded, lock-checked retry for it once the delivery settles. A restoration started for a later pending supersedes it, and an exhausted restoration still hands repair to main without a further arm. The regression holds a branch settlement open while the verified successor exits with a failure and proves one retry watcher starts after the settlement releases, none while it is held. Claude-Session: https://claude.ai/code/session_01QJjTsUvKkWAwLGNoncaZ3a * fix(bin): bound repeat stale wakes for parked workers (#3532) * fix(bin): bound repeat stale wakes for a parked but live worker A worker parked on a declared wait - `paused:` for an external or pipeline wait, or a verified `captain-held` transfer - kept waking firstmate far inside FM_PAUSE_RESURFACE_SECS. Observed as five consecutive alarms on one captain-held worker and dozens across a day on a pipeline wait, and reported upstream as four wakes in 75 minutes against a 3600s window. pause_state_class deliberately answers `none` for a still-live agent even under a declared wait, so a worker genuinely waiting on a decision is never silenced. That classification is correct and is left alone; it routes every parked but live worker through surface_nonterminal_stale on first sight of each distinct stale hash, and an idle parked pane still churns its hash on a clock or a token counter without changing what is being waited on. Two places let that churn re-alarm: - surface_nonterminal_stale queued the wake BEFORE consulting whether a wait was declared, then wrote `.paused-resurfaced-<key>` - the very throttle that should have suppressed it. The throttle was never read on this path and was advanced by the wake it should have prevented. - The hash-change path cleared that throttle through clear_pause_tracking whenever the classification came back `none`, so each tick also bought the same declared wait a fresh window. Fixing only the first site changes nothing. Read the throttle before anything is queued and advance it only on a wake that really fires, and on the hash-change path reset only the per-hash bookkeeping while the declaration still stands, via a clear_stale_hash_tracking split so neither half of clear_pause_tracking is duplicated. The throttle is keyed to the declaration, not to the pane. First sight still wakes, so an inconclusive state is still inspected, and the window's end still re-surfaces once, so a forgotten wait cannot rot invisibly - noise traded for a bounded cadence, never for silence. The wake identity stays the plain `stale: <win>` the away-mode handoff depends on. Tests cover both observed forms and were confirmed to fail against three deliberate breaks: each site reverted on its own, and a re-surface that never fires again. * fix(document): Clarify declared-wait wake cadence documentation * fix(ci): Captain, fixed the stale-throttle inheritance: cadence markers now bind to the current wait declaration, so replacement paused and captain-held waits each emit their first plain `stale:` wake. Added behavioral coverage for both forms. Bite proof failed as expected when identity matching was removed, then passed after restoration. Full watcher triage suite, `bin/fm-lint.sh`, syntax checks, and diff checks pass. Changes remain uncommitted for the outer executor * fix(ci): Captain, fixed the confirmed Greptile finding. `resurface_absorbed` now applies a throttle only when its stored declaration scope matches the current wait, so replacement `paused:` and `captain-held` waits surface immediately without changing classification. Added executable coverage for both absorbed forms. Bite proof failed before the fix at the intended assertion; afterward the full watcher triage suite, `bin/fm-lint.sh`, shell syntax checks, and `git diff --check` passed * fix(bin): accept the away-mode daemon as the turn-end supervision owner (#3567) * fix(turnend): accept the away-mode daemon as the supervision owner While state/.afk exists the away-mode daemon owns supervision and runs bin/fm-watch.sh one-shot: the watcher exits on every wake and the daemon starts its replacement. The turn-end guard tested for a live watcher process holding the watch lock at that instant, so a turn boundary that landed in the hand-off blocked with "TURN WOULD END BLIND" while supervision was completely healthy, costing a full handling turn each time. Reproduced with the real daemon wrapping the real watcher and the real guard sampling the same home: 6 of 40 samples blocked, every one of them with the daemon alive and the beacon 2-3 seconds old, and a new watcher pid on each cycle. After the fix the same reproduction blocks 0 of 40, and killing the daemon and its watcher (away mode still on, beacon still fresh) blocks again. The guard now accepts a live, identity-matched daemon holding this home as proof of supervision while away mode is active. The identity match is the same discipline the watcher lock uses, so a recycled pid or a lock left by a killed daemon proves nothing. The fresh-beacon half of the predicate is unchanged: a daemon that stops restarting its watcher still blocks once the beacon passes grace, a home with no supervisor blocks exactly as before, and with away mode off the strict watcher predicate is untouched. The predicate reads only durable state, so it behaves identically for every primary harness and runtime backend. * no-mistakes(document): clarify away-mode daemon supervision proof and test coverage * no-mistakes(document): generalize stale turn-end predicate summary in architecture.md * fix(backlog): omit --file from row probes for non-markdown backends (#3582) * fix(backlog): omit markdown file for beads probes * no-mistakes(document): Narrow backlog addressing doc to mutations for backend-aware probes * no-mistakes(ci): Fixed the Greptile P2 review comment (the only failing check) on tests/fm-backlog-atomicity.test.sh. The comment correctly noted that an exported TASKS_AXI_BACKEND environment variable would inherit into the spawned scripts and, because fm_tasks_axi_backend gives it top precedence, override each test case's .tasks.toml backend fixture — making the backend-specific argv assertions fail for environmental reasons. Fix: unset TASKS_AXI_BACKEND in the test harness right after sourcing tests/lib.sh, with a comment explaining why, so every case deterministically exercises its declared backend (4 lines added; no production code touched). Verified: reproduced the leak before the fix (TASKS_AXI_BACKEND=beads made the markdown dispatch case fail with 'beads show failed', exactly the reported failure mode); after the fix the full suite passes (0 failures, exit 0) both with and without TASKS_AXI_BACKEND=beads exported. The added lines are shellcheck-clean (the only shellcheck note, SC1091 on the lib.sh source line, pre-exists this change) * fix(bin): classify progress updates on requested work as routine (#3589) The supervision branch's verdict rule escalated every outcome that answered a captain request, so "the work started" and "still working" notes reached the captain with nothing to look at. The rule now keeps a finished result of requested work captain-facing, even when healthy, and treats start or still-working updates that bring no new artifact, finding, or decision as routine. The captain list for review-ready PRs, ask-user findings, exhausted blockers, credentials, and destructive or security-sensitive cases is unchanged, as are the unsolicited-routine, silent-fleet-review, and doubt-chooses-captain rules. The fm_branch_report tool description and the two docs that restated the old unconditional rule now point at the prompt's "Verdict: routine or captain" section as the one owner instead of carrying a second copy. * fix(bin): preserve captain calls during teardown (#3595) * fix(bin): never close a captain call during cleanup A scout that held its own work item for the captain, which is what captain-hold-lifecycle prefers ("hold the work item the question gates"), was closed by bin/fm-teardown.sh's automatic backlog transition. The completion gate passed, cleanup ran, and the captain's question moved to Done with no recorded answer: the one thing the policy says must never happen. `tasks-axi done` closes a held row silently, and nothing in teardown asked whether the row was the captain's own call. bin/fm-captain-hold.sh gains the read-only `open` predicate: exit 0 when the task is still an open captain call, 1 when it is not, 2 when that cannot be established. It reads the row through the transition library's backend-aware probe, so it addresses the same backlog teardown does; the script's other commands now address the configured data directory the same way instead of FM_HOME, which also fixes captain holds in a home with a relocated data directory. Teardown asks `open` before any destructive step and refuses on 2. On 0 only the close changes: after cleanup and still under the task's own lock, the row gets one "Deliverable of the finished work" line at the end of its body and returns to Queued through `tasks-axi reopen`, keeping its hold, so it lands in Captain's Call instead of reading as work under way. --force does not lift this: it authorizes discarding unlanded work, never the captain's question. The deliverable goes into the body because `tasks-axi update --report` rewrites the title of a row that is not Done. The crash window reuses the pending-close record teardown already stages: a `mode=retain` line makes the existing replay record the deliverable and reopen instead of closing, with the same validator, stale-generation check, cleanup-incomplete marking, and non-blocking bootstrap lock as an ordinary close. A retained row the captain answered first simply retires the record. No parallel record type, recovery command, or second bootstrap loop is introduced. Regressions run the real executables: the captain-held scout survives cleanup queued, held, with its deliverable and on the board, only `answer` closes it, --force keeps it open, and an ordinary scout still closes with its report; an interrupted cleanup leaves the row untouched and the next session start retains it; a relocated backlog keeps the retention in its one configured file; and a ship row whose hold cannot be read refuses cleanup before anything destructive. Claude-Session: https://claude.ai/code/session_01FqdTiHCwTqrAQrz8K2y4Np * no-mistakes(review): Serialize captain holds and fix backend-aware listing * no-mistakes(document): Update captain-call retention documentation * no-mistakes(document): Fix relocated captain-hold backlog diagnostics * fix(bin): deliver secondmate outcomes to the parent channel (#3592) * fix(bin): deliver every secondmate outcome on the parent channel from the recording scripts A secondmate's captain-facing outcomes could miss: the mate model addressed the captain in its own unread chat instead of appending to the parent channel, and a PR-ready report, a finding, a decision, a blocker, and a failure all depended on that one remembered append. Make delivery structural, so the parent channel never depends on the model: - bin/fm-parent-channel-lib.sh is the one owner of channel resolution and exact-line append-once; the merge outcome path and the inactive-outcome scan now publish through it instead of two private copies. - bin/fm-inactive-reconcile.sh gains a ledger-first path that runs on every watcher poll in a secondmate home: a direct child's whole terminal done or failed line is delivered at once with its note, recorded PR, mode, merge posture, and scout report pointer, keyed and receipted so it is delivered once, and the inactive path yields to it. `report <task-id>` runs the same delivery for a caller holding the child's meta lock. - bin/fm-pr-check.sh publishes the PR-ready line with the canonical URL at registration. - bin/fm-captain-hold.sh publishes a hold and its answer, keyed by task id and resolution-record count, with no new persisted state. - bin/fm-teardown.sh delivers the child's final line before removing its record and refuses, retaining every record, while the channel cannot be written. - The charter opens with the parent-channel rule and confines the mate's own appends to judgement; AGENTS.md carries the carve-out at the persona address rule and the escalation list. docs/secondmate-parent-channel.md records the design and its coverage, and docs/verification/secondmate-parent-channel.md records the live run with real tmux panes and both real watchers delivering every line with no model. Supersedes #3569. * no-mistakes(review): Fix parent outcome retries and reconciliation locking * no-mistakes(review): Prevent busy children from starving ledger delivery * no-mistakes(review): Correct ledger metadata and hold occurrence handling * no-mistakes(review): Disambiguate ledger outcomes and normalize hold reasons * no-mistakes(review): Close ledger races and preserve teardown records * no-mistakes(document): Correct parent-channel receipt and scanner documentation * no-mistakes(lint): Quote done arguments for ShellCheck compliance * no-mistakes(ci): Fixed both CI failures. Updated GOTMP teardown fixtures for the new final-outcome reporter and isolated them from host tmux state. Updated the PR security assertion to distinguish the accepted PR-ready line from duplicate merge outcomes. Verified with both failing test suites, bash syntax checks, and git diff checks * no-mistakes(ci): Fixed Greptile’s duplicate-delivery race in bin/fm-inactive-reconcile.sh. Ledger events now claim matching already-delivered inactive receipts using the prior status fingerprint, preventing duplicate parent reports while preserving later same-state completions. Added behavioral regression coverage. Verified inactive-reconcile tests, project lint, documentation audience checks, syntax, and diff checks. Teardown tests passed relevant cases before the documented pre-existing herdr-preflight-missing-adapter failure * fix(bin): sync remote second mates to primary commit (#3599) * fix(bin): sync remote second-mate homes to the parent primary commit Session start and remote launch pointed a remote second-mate home at whatever Firstmate copy its own host kept, so a home that had already advanced past that copy refused as a non-fast-forward and every other home stopped at the host's older commit while the primary ran ahead. The parent now resolves ITS primary default-branch commit with the existing helper and hands that commit to the host on both paths. Because a remote home is a standalone clone, the host imports that one commit before advancing - already present, else from that host's Firstmate copy without moving it, else from the home's own origin - and then runs the SAME ff_target guards a local home gets, so dirty, diverged, feature-branch, and unresolvable targets skip untouched and the ancestry rules keep one owner. An unimportable target now names /updatefirstmate instead of failing opaquely, and a host still running an older Firstmate copy is reported the same way rather than echoing a bare refusal. The host-local launch leg no longer re-runs its own secondmate sync, so the spawn it drives cannot re-target that host's copy after the parent has already converged the home. /updatefirstmate is unchanged: it still refreshes the remote code root from that host's origin and then syncs the home to that refreshed copy, which is what the sync call with no target commit means. * no-mistakes(document): Document primary-targeted remote secondmate synchronization * fix(bin): separate captain intent from firstmate specs (#3597) * fix(bin): split brief task into captain intent and firstmate spec Keep no-mistakes --intent as the captain's ask plus later captain words, not the build spec or worker tradeoffs. * fix(bin): stop task-subsection copies at the next heading Promotion was swallowing the scout Setup contract into Firstmate spec, and pre-subsection briefs lost their # Task body. * no-mistakes(review): Validate brief content and preserve nested specifications * no-mistakes(review): Scope placeholder validation to scaffold-only subsection bodies * no-mistakes(review): Ignore fenced subsection headings during brief validation * no-mistakes(review): Preserve captain intent across scout promotion * no-mistakes(review): Enforce safe intent boundaries for legacy promotions * no-mistakes(review): Allow marked legacy intent and reject empty promotions * no-mistakes(review): Scope task parsing and overlay legacy intent contracts * no-mistakes(review): Overlay current intent contract for all no-mistakes spawns * no-mistakes(review): Preserve later captain clarifications in intent overlays * no-mistakes(document): Document brief intent enforcement and ownership * no-mistakes(ci): Updated spawn-related test fixtures to use valid Captain intent and Firstmate spec subsections, corrected launch-path expectations to launch-brief.md, and resolved ShellCheck quoting findings. Verified with fm-lint.sh and 15 affected behavior tests, including real Herdr tests; all passed * no-mistakes(ci): Updated stale spawn/promotion fixtures in the Muse, Orca, secondmate-harness, and public-followup suites to provide valid Captain's intent and Firstmate spec subsections. Verified full Orca and secondmate-harness suites, targeted public-followup promotion behavior, Bash syntax, diff checks, and fm-lint * fix: start a fresh supervision branch for every main session (#3600) * fix(pi): start a new supervision branch conversation per main session The supervision branch reopened one recorded conversation forever, so every main session start reloaded the current generated prompt and then weeks of accumulated thread, where a superseded rule could still outweigh today's. The branch conversation is now scoped to one main session: the session generation owns the recorded conversation, so a cold start, /new, /resume, /fork, or a reload always builds a new one, while a rebuild inside one session (a model or effort change) still continues that session's own conversation. The dialog mirror re-anchors with it. Its durable cursor records what the previous branch conversation received, so a /resume or reload - which keeps main's own session file - would otherwise leave the new branch blind to dialog main itself still has. The reset is bounded by the current main session, and the cursor keeps advancing incrementally within it. The durable outcome store and its processed marker are untouched, so unacknowledged captain-facing outcomes still re-present on the new main session. * no-mistakes(document): Document fresh Pi supervision conversations * no-mistakes(ci): Fixed the flaky concurrent inbox failure. Lock acquisition now retries when a competing lock disappears between a failed claim and inspection. Added a behavioral regression covering that race. Verified the full inbox test four times, project lint, and git diff checks * feat: restart second mates after instruction updates (#3614) * feat(update): restart second mates whose instructions changed /updatefirstmate pulled new bytes onto disk and then asked each advanced second mate to re-read them. A running agent holds AGENTS.md and every loaded skill frozen from launch and no verified harness offers a reload, so that steer could not reach a loaded skill at all and left the mate holding two contradictory copies of its own job description. An eligible mate is now restarted instead, in the same home and endpoint, through the existing transactional relaunch. The restart is gated on the mate first writing down the open work it holds only in conversation - the open-record half of /stow, never its memory sweeps - so an unregistered captain call is flushed before the conversation is spent. Anything that leaves the reload unprovable falls back to the old re-read message and is reported as exactly that, never as a clean reload. Remote mates take the same path: fm-remote-secondmate-control.sh gains a relaunch verb whose host-local leg runs that same control plane, since the mate is an ordinary local secondmate from its host's point of view. The primary resolves the profile and passes it explicitly, because config/secondmate-harness is not inherited and the file on that host belongs to a different home. fm-update.sh now splits its advanced live mates into a restart set and a nudge residual, and both sets require a changed instruction surface, which also closes the over-nudge against the session-start sweep. Restart is stricter still: a bin/-only advance reloads itself on the next call, so it never costs a conversation. Colocated tests cover the gating, the persist-then-restart order, the task-subset persist request, each unsafe fallback, the remote hop, and the remote sync's new instruction-surface report. * no-mistakes(review): Fix restart correlation, concurrent waits, and lifecycle reporting * no-mistakes(review): Parallelize relaunches and classify replacement incarnations * no-mistakes(review): Gate restart actions on live agent state * no-mistakes(review): Handle failed restart workers without hanging * no-mistakes(review): Nudge legacy remotes and preserve persist recovery * no-mistakes(review): Document one-time secondmate restart rollout * no-mistakes(review): Honor arrived replies and refresh remote profiles * no-mistakes(review): Revert remote parent profile reconciliation * no-mistakes(review): Reset remote profile defaults and honor published results * no-mistakes(review): Preserve fallback nudges for unverifiable secondmates * no-mistakes(document): Document second-mate restart update flow * no-mistakes(lint): Fix ShellCheck warnings in restart scripts * perf: accelerate local validation with bounded concurrency (#3644) * perf(tests): route gate verification through the bounded concurrent runner Local validation was the pipeline's dominant cost: across 67 recorded no-mistakes agent sessions on this repo, 99.3% of command execution was `bash tests/*.test.sh`, run strictly one script at a time, and 2% of those calls were killed by an agent-guessed timeout and paid for twice. Three changes, each measured: - `.no-mistakes.yaml` pins `commands.test` to `bin/fm-test-run.sh --changed --exclude-family real-herdr-gated`. The runner already owns changed-file selection, bounded concurrency, the refusal of unproven scripts, and a generous automatic per-script bound, so the gate's baseline is neither a serial chain nor a guessed timeout. It stays intent-targeted - the Test step still runs its evidence agent on top - and excludes the live-Herdr family the required Herdr lane owns. - `bin/fm-test-run.sh` gives a plain list of script paths the same bounded automatic scheduler and automatic bound that `--changed` gets. Naming several subjects is how a verification round asks for exactly those scripts. The curated selections are untouched: `--lane` still composes CI shards whose serial lane must stay serial, `--family` is what the required Herdr lane runs, and `--all` stays a deliberate complete regression. - `pr-forge` is admitted to the concurrent-safe family registry on two consecutive clean proofs. `docs/fm-test-isolation-proof.md` records those, and records `secondmate` and `session-bootstrap` as refused with the exact script and reason each failed on, so the refusals are actionable rather than silent. Measured on this host, 0 failures on both sides: verification round, 4 scripts 448s chained -> 231s through the runner (-48%) pr-forge family 409.2s at 1 worker -> 237.9s at 4 (1.72x) watcher-wake-lock family 1311.1s at 1 worker -> 539.3s at 4 (2.43x) A fourth lever was implemented and then removed because the measurement refused it: raising the bounded-wait sample interval from 0.1s to 0.5s made `fm-watch-triage.test.sh` slower, 435s and 440s against 390s and 393s unchanged, back to back. Those sleeps are not overhead added to the clock - they are how a test waits for a subject moving on fm-watch.sh's own one-second cadence - so sampling less often only delays detection. It also broke `fm-watcher-lock.test.sh`, which catches a transient rather than waiting for a settled condition. CONTRIBUTING.md records that result so the experiment is not repeated. * no-mistakes(review): Separate concurrent runs by isolation proof family * no-mistakes(review): Limit automatic timeouts to changed-file validation * no-mistakes(document): Clarify validation concurrency documentation * fix: copy PR URLs from durable records (#3648) * fix: copy PR URLs from records or abstain, never assemble them Supervision reported a plausible but dead PR link three times because its prompt demanded a full https:// URL at a moment when only a PR number was observable, so the model assembled an owner/repository from memory, and the PR check then accepted that URL and wrote it into the task record, after which the model kept defending its own tool-endorsed guess over the worker's real link. Three changes close that chain without any live forge lookup, so private forges are treated exactly like public ones: - bin/fm-branch-prompt.sh no longer mandates a URL. Its new "PR identity: copy or abstain" section requires a URL to be copied verbatim from a durable record (the done: PR <url> status line, pr= metadata, or the backlog note), forbids assembling owner, repository, host, or number from memory, and has the branch report only the identifier it actually holds when no record names the URL yet, leaving the PR check unarmed until the worker's ready line arrives. AGENTS.md section 7 and 9 carry the same copy-or-abstain rule for main in place of the bare full-URL mandate. - Worker briefs (bin/fm-brief.sh, ship and scout rules) require the full https:// URL wherever a PR is mentioned - status line, terminal, or summary - never a bare "PR 108", so the link is in view as early as the number is. - bin/fm-pr-check.sh refuses, offline and before any side effect, a URL that the task's own done lines contradict, printing both spellings; a log naming no URL still records the argument as before. fm_pr_status_ready_urls in bin/fm-pr-lib.sh owns reading those lines. The refusal also reaches bin/fm-pr-merge.sh, so nothing merges under a contradicted URL. Tests cover the offline refusal with zero side effects, the recorded spelling being accepted, markdown-wrapped and punctuated URLs, working lines not counting, the merge wrapper propagation, a self-hosted merge request with no forge call, the prompt carrying the rule, and the brief carrying the worker rule. * no-mistakes(review): Remove stale PR URL enforcement * no-mistakes(ci): Removed backlog notes as an accepted PR identity source. PR URLs may now be copied only from the task’s `done: PR <url>` status or canonical `pr=` metadata; otherwise supervision reports only the known identifier and leaves PR checking unarmed. Updated related guidance/docs and verified with branch-supervision tests, brief tests, ShellCheck, and `git diff --check` * fix(bin): disable Claude feedback drafts for fleet launches (#3661) * fix(bin): disable Claude's feedback-draft flow for fleet-launched agents Scope --settings '{"feedbackDrafts":"off"}' to every Firstmate-launched Claude crewmate and secondmate, so /bug and /feedback never queue or submit a bug report on the captain's behalf. feedbackDrafts is the documented settings key (Claude Code changelog 2.1.247); the per-launch CLI flag never touches the captain's global settings.json. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(review): Prevent managed settings from re-enabling Claude feedback drafts * no-mistakes(document): Fix Claude feedback documentation formatting * fix(bin): layer both feedback-draft controls for defense in depth The prior --settings-only fix can be overridden by a managed Claude settings policy (feedbackDrafts precedence). Keep CLAUDE_CODE_SEND_FEEDBACK=0 alongside --settings '{"feedbackDrafts":"off"}': either control alone disables the SendFeedback tool, so a managed override of one still leaves the other in force. Claude-Session: https://claude.ai/code/session_01XYAXXzr4oZx9NjZb1veeE3 * no-mistakes(document): Document Claude feedback-draft suppression ownership * feat(tests): run three more validation families concurrently (#3662) * perf(tests): admit three more families to concurrent validation The three families that `docs/fm-test-isolation-proof.md` recorded as refused were not refused for concurrency. Each blocker was a test that decided a property by wall clock, or a script filed where it cannot run. Fixing those three things admits all three families and recovers 28.6 minutes of local validation with no assertion removed or weakened. - `tests/fm-backlog-handoff.test.sh` injected its pre-move crash by killing the handoff, sleeping a fixed second, then delegating the move to the real binary. Nothing ever killed the fake, so on a host slow enough for the case's next assertions to take longer than a second, the orphan woke and completed the very move the case requires left undone, and recovery then failed with `Task "pre-move-crash" not found in this backlog`. Watching the two backlogs during the injected crash showed exactly that, the item moving one second after the crash. All four crash injections in the file now go through a new `fm_fake_crash_injector` shim that signals the target and returns only once it is observably gone, and the pre-move fake never delegates the move at all. - `tests/fm-session-start.test.sh` proved the startup digest does not block on a slow current-state read by timing the whole digest against a fixed eight-second sleep, which a loaded host exceeds without the property being violated. It now holds that read open until the case releases it and asserts, the moment the digest returns, that the read has not finished. A digest that waited would wait indefinitely rather than for an interval a slow host can out-run, so the assertion is stronger than the bound it replaces. Its scan budget moves to the maximum, because the old value left two seconds of margin over the fixed sleep and measured the host rather than the deadline that `tests/fm-inactive-reconcile.test.sh` owns. - `fm-backend-herdr-focus-flash-e2e` was filed in the family map's catch-all, which put it in the portable serial lane, where Linux CI gate-skips it: that real-Herdr regression was running nowhere. It moves to `real-herdr-gated` and the required Herdr lane. `fm-claude-stop-autoarm-live-e2e` gate-skips on its opt-in variable and moves to `live-harness-optin`. The 28 remaining ungrouped scripts become an enumerated `standalone` family instead of admitting `unclassified` itself. `unclassified` is the family map's `*)` arm, so admitting it would silently grant concurrency to every test added afterwards, which is exactly the population with no proof. A new test still lands in `unclassified` and stays serial, and `tests/fm-test-run.test.sh` covers that split behaviorally. Each family passes two consecutive four-worker proofs with zero failures. On the production runner, `secondmate` goes 1233.1s to 453.4s, `session-bootstrap` 756.4s to 286.4s, and `standalone` 724.6s to 261.1s: 2.71x overall and 1713.2s recovered. The whole suite runs 177 scripts in 52.6 minutes of wall clock against 121 minutes of summed script time. * no-mistakes(document): Refresh concurrent validation and shard documentation * no-mistakes(ci): Fixed the real-Herdr focus-flash E2E race exposed by reclassification. Part C now starts its persistent child atomically via `pane run` and verifies stable child identity through Herdr’s public `process-info` interface, avoiding the racy send-text/send-keys sequence and platform-specific `ps` matching. Verified with bash syntax checking, ShellCheck, git diff checks, and the complete E2E test on Herdr 0.8.2 * feat: structure no-mistakes ask-user escalations (#3670) * feat(brief): structure no-mistakes ask-user escalation as event + snapshot file Crewmates escalating a no-mistakes ask-user gate now report one status event naming every finding id plus a snapshot file holding the gate's axi finding records verbatim (id, severity, file, line, description, authority), using the same shape even for a single finding. The status line never paraphrases. The format is defined once in fm-dod-lib.sh and rendered into both the scout and ship rule 6 in fm-brief.sh, so a promoted scout - whose rule 6 fm-promote.sh preserves unchanged - gets the identical contract as a freshly-spawned no-mistakes ship worker. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01PpiWaDerbYavTLPPtEjQei * no-mistakes(review): Preserve ask-user escalation output contract * no-mistakes(review): Align escalation format test expectation * no-mistakes(review): Scope ask-user escalation instructions correctly * no-mistakes(review): Remove ask-user from generic decision rules --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * fix(bin): require self-sufficient no-mistakes intent (#3671) * fix(bin): require a self-sufficient no-mistakes intent A no-mistakes worker's --intent is only as useful as the string it passes. PR #3604 shipped with an intent that was only "do 1, 2, 3, 7 from the report": the real contract lived in a private scout report and never reached --intent, so nobody holding that string plus the codebase could have derived the specification. This is pure instruction at the contract's one owner; no spawn-side or promotion-side check is added. - bin/fm-dod-lib.sh: the generated no-mistakes Definition of done now states that the --intent string must be self-sufficient (the string plus the codebase reconstructs roughly the same specification) and tells the worker to write the substance of any report, decision, or PR the captain's intent refers to into --intent rather than the pointer, while Firstmate build instructions and the worker's own decisions still stay out. The spawn-time overlay points back at that rule so its "supersedes" wording cannot cancel it, and the header's owner statement carries the rule. - AGENTS.md section 11 and bin/fm-brief.sh's header ask Firstmate to include the substance of referenced material when filling ## Captain's intent, and section 11 points at the owner of the rule. - tests/fm-brief.test.sh and tests/fm-task-delivery.test.sh assert the rendered brief and launch contract carry the rule. Claude-Session: https://claude.ai/code/session_01YMhEe42q7BAAoN6RxNuzim * no-mistakes(document): Replace incident-specific intent test commentary * fix: accelerate local Bearings snapshot composition (#3499) * Speed local fleet snapshot composition * no-mistakes(review): Stabilize task inventory during concurrent snapshot composition * no-mistakes(document): Document local snapshot observation concurrency * no-mistakes(ci): Fixed CI failures by making empty task manifests compatible with stock macOS Bash 3.2, snapshotting task metadata before concurrent observations to prevent generation drift, strengthening the behavioral race regression, and updating the stock-Bash Bearings test count to 45. Verified fleet snapshot tests (15), Bearings tests (45), workflow lint tests, project lint, Bash 3.2 parsing, and diff checks * no-mistakes(ci): Fixed the Linux CI failure caused by passing large backlog/task JSON through jq command-line arguments, which exceeded the per-argument size limit. Both inventory projections now stream large JSON inputs through stdin. Verified with fm-bearings-snapshot.test.sh (45 tests), fm-fleet-snapshot-view.test.sh (15 tests), Bash syntax, and git diff checks * no-mistakes(ci): Fixed concurrent task teardown during metadata capture: vanished metadata is now omitted while genuine copy failures remain fatal. Added a deterministic public Bearings regression test and updated CI’s expected test count. Verified with the full Bearings suite, workflow-lint suite, Bash syntax checks, and git diff checks * no-mistakes(ci): Fixed PR-caused CI and review issues: streamed large fleet JSON through jq stdin to avoid Linux argument limits, kept crew-state reads bound to captured metadata generations, and strengthened the behavioral race test. Bearings (46 tests), fleet snapshot (15 tests), crew-state, backend, lint, Bash syntax, and diff checks pass locally. Serial shard 5’s unrelated task-inbox segmentation fault appears infrastructural/flaky * no-mistakes(ci): Fixed endpoint-state generation crossing by validating captured spawn_gen before and after local endpoint probes, falling back to exact metadata identity for legacy tasks. Stale probe results now become unknown instead of false unhealthy state. Added a behavioral relaunch-race regression test. Verified the full Bearings snapshot suite, shellcheck, bash syntax, and git diff checks * fix(snapshot): keep live observations generation-coherent * no-mistakes(review): Keep secondmate observations generation-bound without copying reports * no-mistakes(document): Document generation-coherent snapshot observations * test(bearings): measure local read overlap instead of wall-clock budget The large-local-snapshot regression asserted that a whole snapshot composed in under five seconds. That bound measures how loaded the host is, not whether the per-task reads actually overlap, so it failed intermittently on a contended machine: one run in six on a box at load 16-20, landing exactly on the five second boundary. Time a serialized run and a concurrent run of the same workload instead and require the concurrent one to save at least two seconds. Both runs pay the same composition overhead, so the difference isolates the overlap this change delivers. Five one-second reads serialize into five seconds and overlap into about one, and re-serializing the reads collapses the saving to roughly zero, so the assertion still fails loudly if the concurrency regresses. Also bump the pinned Bearings test count to 48, since rebasing onto the current default branch picked up its captain-hold test. * no-mistakes(review): Restore JSON-derived decision flags * no-mistakes(review): Unify status-derived snapshot observations * no-mistakes(ci): Updated the stock macOS Bash CI check’s Bearings test count from 48 to 49. Verified the full Bearings suite passes and emits exactly 49 TAP successes; git diff checks pass * fix: prevent stale supervision wake loops (#3672) * fix(bin): stop the supervision branch's stale-ack and ghost-report loops Clean-slate implementation of the four authorized recommendations from the supervision-ghost-retrigger analysis (items 1, 2, 3, and 7), in their minimal form, superseding PR #3604: - fm_branch_report refuses a task the wake being handled never named. The extension fixes the reportable task set from the eligible rows before each prompt (signal and stale rows resolve to their tasks, a heartbeat allows any task with a live record, fleet is always allowed), so a report typed from memory about a task whose records teardown already removed is never stored or delivered. - An acknowledgement that consumes nothing says "nothing was acknowledged through N" and prints the exact --ack-through / --recovery-generation command for the current presented wake, instead of "re-run the drain", which re-fed the same stale acknowledgement in a loop. - bin/fm-guard.sh no longer tells the branch actor to drain queued wakes while it is handling them; it names the granted rows instead. - Teardown removes state/.<task>.branch-outcome-index for ordinary tasks and descendants; the index rebuild and the append-side index write both skip a task with neither a live record nor a status log, so the branch's report of a teardown it just performed is stored without recreating the index. No new locking, no spawn-generation binding, and no retired-task refusal: the branch can still report the outcome of a task it just tore down, and the teardown test now proves that path end to end. * fix(bin): narrow the branch report scope and guard silence to the minimal form Apply the four review decisions on the clean-slate branch: - A signal or stale prompt may report only the tasks its own rows resolve to; fleet is refused there too. A heartbeat review is not scoped by task at all, so the extension no longer tracks live task records and refuses nothing by task id during a fleet review. - The outcome-index rebuild no longer skips retired tasks; the append-side skip alone keeps a torn-down task's index from being recreated. - bin/fm-guard.sh keeps the queued-wakes warning silent for the branch actor instead of printing a replacement note. * no-mistakes(document): Align supervision docs with scoped wake handling * fix(bin): avoid fleet snapshot argument limits (#3677) * Fix fleet snapshot large JSON transport * no-mistakes(review): Captain: file-back fleet snapshot transport safely * no-mistakes(review): Captain: file-back parent summary aggregation * no-mistakes(ci): Rebased the PR's three commits onto f4d7875824ecc5e274b4bb896f10c1e1f207b7e4 and resolved the fleet snapshot conflict while preserving the base's task-observation lifecycle. Fixed Greptile's valid finding by recursively removing the private mktemp transport directory, so future transport files cannot cause cleanup to fail. Verified with tests/fm-home-summary-refresh.test.sh, bin/fm-lint.sh, git diff --check, and ancestry checks. All passed; the fix remains as an uncommitted worktree change for the outer executor * fix(bin): attribute active runs with unfetched pipeline heads (#3681) * fix(bin): rec…
Intent
Ship the selected correction for fm-crew-state misreporting an active validation fix round as an older failed run when the pipeline-owned head is absent from the task copy's object store.
The independent judge selected candidate B (exact commit f61dbcb0 from fm/fm-crew-state-pipeline-fix-head-visibility-b) as implementation-ready and rejected candidate A's sqlite-dependent design for incomplete paths, stale-row attribution, and duplicated ownership. The selected commit was cherry-picked onto upstream/main (f4d7875) and port-resolved: the parent's independent branch_sync custody exemption (#3194) coexists, each mechanism owning one surface (TOON custody on the full axi-status path, the runs ledger on the coarse path); the superseded coarse scan-and-skip (nm_runs_status_for_branch) and its now caller-less helpers (fm_nm_head_resolvable, nm_coarse_head_matches_worktree) are deleted; the parent coarse-guard test whose fixture is the ledger-anchored continuation shape now asserts the fixed behavior, and a new mismatched-anchor coarse negative control preserves that guard's original no-anchor protection.
Required behavior (candidate B's conservative rule, preserved exactly): fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh is the single owner of runs-ledger attribution; the branch's newest same-branch ledger row alone decides; an unresolved active head is attributed only when the immediately older same-branch row resolves exactly to the task copy head; unrelated, ambiguous, stale historical, terminal-unfetched, and mismatched rows must never produce a false verdict. fm_nm_head_matches_worktree keeps its exact prior semantics so teardown behavior is byte-for-byte unchanged. Read-only: the reader never fetches into or mutates another task copy and never moves branch custody. The judge's one-line correctness cleanup is included: the stale FM_CREW_STATE_RUNS_LIMIT comment in bin/fm-crew-state.sh now points at fm_nm_runs_status_for_worktree in bin/fm-nm-run-lib.sh.
Scope boundary: the separately filed teardown continuation task owns the parked-run cleanup extension; do not broaden this change into teardown.
Validation evidence already gathered on the final head: tests/fm-crew-state.test.sh 63/63 pass; tests/fm-inactive-reconcile.test.sh 28/28 pass; bin/fm-lint.sh clean; tests/fm-backend-orca.test.sh and tests/fm-watch-triage.test.sh pass. The reproduction differential proved on the original candidate base (12f6613) that an active fix round at an unfetched head read "state: failed . source: run-step . run failed" from the stale row, and reads "state: working . source: run-step" on this branch, with the fetched-object counterfactual attributing on every revision. Known pre-existing environmental failures, proven identical on the untouched base revisions: tests/fm-teardown.test.sh reports herdr-preflight-missing-adapter (macOS stock bash 3.2.57 herdr preflight flow), and tests/fm-wake-queue.test.sh fails at the subshell-hold assertion (rc=13).
What Changed
fm_nm_runs_status_for_worktree, using only the newest same-branch row.Risk Assessment
✅ Low: The shared ledger attribution now handles the unfetched active-head sequence conservatively while preserving strict teardown head matching and rejecting stale, terminal, and mismatched rows.
Testing
Completed 1 recorded test check.
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ **Review** - passed
✅ No issues found.
bin/fm-test-run.sh --changed --exclude-family real-herdr-gated✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.