Skip to content

Merge upstream Firstmate through b182d0f908b7 (28 changes) - #278

Merged
HelloWorldSungin merged 32 commits into
mainfrom
fm/fm-upstream-sync-2026-09-14-tail
Sep 14, 2026
Merged

HelloWorldSungin merged 32 commits into
mainfrom
fm/fm-upstream-sync-2026-09-14-tail

Conversation

@HelloWorldSungin

@HelloWorldSungin HelloWorldSungin commented Sep 14, 2026 •

Copy link
Copy Markdown
Owner

Merge the next contiguous upstream prefix through kunchenguid/firstmate@b182d0f908b78d08c7ccb8dce3775bdca8c5d657, retaining the fork contracts and full upstream parentage.

Exactly 28 incoming first-parent changes follow merge base 4768e98d469b7216569acf6e7a5cc696ecd15c2e.
The upstream merge commit’s first parent is the exact fork base d4ea58736c481f5acd924466ab7514fa9960de21; the second parent is the pinned upstream endpoint.
The later upstream change e3bd750cab2d58e33d5148a8e2cf7b2e7428014b is excluded.
This branch contains one upstream merge, with no cherry-picks, rebase, squash, parked-branch intake, or default-branch rewrite.

Exact pushed head: 66da6fb840582c7a2c256d6341dba8768938ed87.
The original merge 66173794ee745a30da64a216255fdcc1c46a04ca is followed by three ordinary correction commits for reviewed-sync policy/fixtures, bounded CI execution, and the completed-poll retirement fixture.
Final live forge observation: open against main, mergeable: true, mergeable_state: unstable because the expected policy check remains red.
All 16 substantive checks pass on this exact pushed head; only PR must be raised via no-mistakes is red.
The complete successful run is https://github.com/HelloWorldSungin/firstmate/actions/runs/34881222435.

Validation

Both pristine parents ran the complete suite and canonical lint in separate isolated temporary linked worktrees before the merge commit was created.
Their original failures remain separately attributed below.

Tree Complete sweep Canonical bin/fm-lint.sh
Fork parent 249 scripts, 1 initial failures, 34 explicit gated skips PASS
Pinned upstream parent 205 scripts, 9 initial failures, 27 explicit gated skips PASS
Merged tree before policy follow-ups 258 scripts, 10 initial failures, 35 explicit gated skips PASS
Final pushed head, complete hosted lanes 258 unique scripts, 0 failures, 38 explicit gated skips PASS locally and in hosted CI

Before the policy follow-up, reconciled complete-sweep coverage is 258 of 258 scripts, zero unresolved failures, 33 explicit gated skips.
This combines the complete merged sweep with the complete-script reruns listed below after fixture/dependency corrections; it does not represent the initial sweep as green.
Within that pre-policy reconciliation, production code, instruction contracts, and workflows did not change after the sweep started.
Subsequent policy and CI follow-ups have their own evidence below and complete final-head hosted coverage in the table above.
Canonical lint was rerun after the fixture changes with pinned ShellCheck 0.11.0 and actionlint 1.7.12.
Documentation audience checks and authenticated cross-system pointer checks passed.

Original tree / failing script Attribution and final evidence
Fork / tests/fm-secondmate-liveness.test.sh Unknown-pane fixture consulted the real nonexistent session server, yielding missing instead of unreadable; final server-state fixture correction and full suite pass.
Upstream / tests/fm-ci-workflow.test.sh Ruby installed but absent from baseline PATH.
Upstream / tests/fm-herdr-attached-viewer-live-e2e.test.sh Originally mandated shared helper predates viewer command. Task-specific authority approved the pinned upstream helper; isolated viewer rerun passed.
Upstream / tests/fm-nm-test-contract.test.sh Ruby installed but absent from baseline PATH.
Upstream / tests/fm-omp-harness.test.sh Leaked-marker fixture lacks a Claude ancestor; host Codex structural ancestry legitimately wins. Final fixture provides actual Claude ancestry.
Upstream / tests/fm-on.test.sh Pristine upstream fixture expects a fixed remote PATH that excludes this host Nix profile. Final fork fixture preserves host-aware portable PATH expectations.
Upstream / tests/fm-procevent-quota.test.sh Pristine upstream fixture inherits host umask 0002, producing a group-writable process-event root refused by the private-directory guard. Final fixture pins umask 022.
Upstream / tests/fm-remote-doctor.test.sh Pristine upstream repair fixture assumes no ambient harness and requires a new codex wrapper. This host already supplies a harness on the system/Nix profile path. The fork fixture isolates the user profile and tests both ambient-present and wrapper-required outcomes.
Upstream / tests/fm-secondmate-liveness.test.sh Same failure as fork baseline: unknown-pane fixture consults the real nonexistent session server and receives missing instead of expected unreadable. Final isolates server-state evidence and models stale-agent explicitly.
Upstream / tests/fm-test-run.test.sh The workflow-parser case requires Ruby, which was absent from baseline PATH although already installed. Final runner suite passes; pristine-parent full-script rerun exposes the installed Ruby.
Merge / tests/fm-agy-adapter.test.sh Synthetic Herdr response lacked structured process-info. Complete corrected script: PASS.
Merge / tests/fm-bootstrap.test.sh Optional Lavish fixtures used the upstream floor and agy schema fixtures selected tmux; preserved fork floor and Herdr-only restriction. Complete corrected script: PASS.
Merge / tests/fm-brief.test.sh Visual scout fixture expected upstream Lavish floor rather than retained installed-version floor. Complete corrected script: PASS.
Merge / tests/fm-captain-hold-lifecycle.test.sh Forge fixture reported MERGED before the merge action, bypassing the intended submission assertion; model OPEN then MERGED. Complete corrected script: PASS.
Merge / tests/fm-ci-workflow.test.sh Ruby was installed in Nix store but absent from runner PATH; focused rerun exposes existing Ruby 3.4.9 without installing or changing the test. Complete corrected script: PASS.
Merge / tests/fm-control-herdr-smoke.test.sh Reconciled stale-agent assertion ran before test agent exit; moved behind proven stale state and made cleanup idempotent. Complete corrected script: PASS.
Merge / tests/fm-crew-state.test.sh Synthetic process-info lacked matching terminal-wide process-table evidence. Complete corrected script: PASS.
Merge / tests/fm-herdr-attached-viewer-live-e2e.test.sh Pinned shared helper lacks incoming viewer subcommand; Firstmate approved exact pinned upstream helper in inbox005, safety audited and rerun after approval. Complete corrected script: PASS.
Merge / tests/fm-issue-writeback.test.sh Fixture assumed gh-axi submission and already-merged state; now models verified head/default tip, OPEN before submission and MERGED afterward, preserving independent tracker failures. Complete corrected script: PASS.
Merge / tests/fm-omp-harness.test.sh Host Codex ancestry overrode leaked Claude marker; fixture now provides actual named Claude ancestor for the claimed Claude worker scenario. Same original failure exists on pristine upstream. Complete corrected script: PASS.

The upstream CI-workflow, no-mistakes-config, and complete runner tests passed when existing Ruby 3.4.9 was exposed on PATH; the upstream viewer test passed with its own pinned helper after Firstmate relayed task-specific approval.
The merged viewer test was rerun after that approval, with exact helper source blobs verified against the endpoint and default-session/ownership tripwires preserved.
One earlier viewer rerun used that worktree helper before the exact-path exception was approved; this was disclosed to Firstmate and is not the authorization-bearing validation run.

Live supplementary coverage includes three consecutive complete agy launch/working/settle/steer/cleanup trials (24.444s, 20.245s, 19.678s), agy 1.2.2 signal/exit behavior, and Herdr 0.8.2 / Pi 0.85.1 stale-registration recovery.
Installed Claude, Codex, OpenCode, Pi, and Cursor process identities passed the live liveness guard.
The complete real GBrain 0.46.21.0 read-only sharing test passed after isolating child runtime homes and provider-key environment variables; its first opt-in attempt correctly refused a fixture that could see host runtime credentials.
The complete dashboard browser regression passed in a dedicated Chrome session at phone, 899/900 navigation-boundary, and desktop widths, including empty-page rejection, injected failure branches, and observation accounting. Its initial viewport setup failed before page observations; a ready-session rerun reached the 480-second default bound, and the complete explicit 1,200-second run passed in 494.453 seconds. The production default remains 480 seconds. Dark-theme phone and desktop screenshots were also visually inspected; no light-theme claim is made.
Absent vendors, optional services, and remaining explicit opt-in live tests remain disclosed gates; no OMP vendor run is claimed.

Ledger changes and parked exclusions

  • Reconcile agy through the fork trust/launch owners, retaining crew-only/Herdr-only/native-state and atomic-delivery contracts.
  • Record Firstmate-resolved attended quiet entry/exit without fabricated away authority or loss of actual away gates.
  • Record the terminal-wide negative agent proof while adopting upstream positive classification and stale-registration states.
  • Record the retained installed-version AXI floor policy alongside optional Lavish presentation.
  • Update timing evidence using parent maxima while retaining eight shards and fork bounds.
  • Clarify internal verified-head binding and preserve the standing landed-readback and queue-guidance decisions.

No parked branch entered: each parked branch’s exclusive commits has an empty intersection with the introduced upstream ancestry.

  • fm/fm-afk-injection-wedge at df4378e804c495a2886ff3f0afa715b9c09f279e.
  • fm/fm-crew-state-blind-during-fix-round at f0e80b3dae9c3ab1ed6199bcf45f2f98bb730291.
  • fm/fm-parked-decision-stale-noise at 6e0c8ff5b5b30c66b7fcd58a09d65451b382e5f3.
  • fm/fm-subagent-model-routing-guard at 45f2e55faf19a24742cfd2228c079fbc476aa98e.
  • fm/fm-vault-drift-check at 7ef747a8c9587fe8c341e13a66c7b6ed7ed7d9d0.

Delivery method

This is the explicitly approved direct-PR exception because the validation pipeline would rebase and linearize the merge-only branch.
The PR must be raised via no-mistakes policy check is expected to be red; every substantive check must pass.
Do not squash or rebase this branch.
Landing must use bin/fm-pr-merge.sh fm-upstream-sync-2026-09-14-tail https://github.com/HelloWorldSungin/firstmate/pull/278 -- --merge.

Reviewed sync policy follow-up

Firstmate relayed the explicit decision to preserve standing Firstmate landing authority for reviewed upstream-sync rounds and implement a narrowly scoped expected-policy exception.
The worker still never merges.
This ordinary follow-up commit preserves the original two-parent merge; it does not amend or rewrite the published merge.

github_verified_upstream_sync in bin/fm-pr-merge.sh owns the exact-head review declaration and the narrow completed-policy-failure exception.
Firstmate must review the final head and record its declaration before using the unchanged landing invocation bin/fm-pr-merge.sh fm-upstream-sync-2026-09-14-tail https://github.com/HelloWorldSungin/firstmate/pull/278 -- --merge.
This worker has not written that authority into live task metadata.

Behavioral cases exercise allowance attended, with an away grant, and after a reviewed linear fix-forward.
Refusal cases exercise pending/cancelled policy runs, policy status contexts, other failed/pending checks, stale or duplicated reviews, wrong task/repository/push/branch/head identities, wrong parents, extra merges, side-branch endpoints, absent away grants, and squash even when policy CI is green.
The existing complete merge suite also verifies generic attended waivers, live target/default-tip rereads, verified-head binding, task holds, away locking, and publishing boundaries.
The changed path reads Git/forge/task state and is independent of harness/runtime selection; backend/harness gates are unchanged.

The initial CI failure was a Lavish-dependent scout-brief golden fixture: Lavish was present locally but absent on the CI runner.
The fixture now pins its capable input while keeping the existing missing/old/current-version behavioral cases; no golden bytes or production version floor changed.
The earlier sibling parallel lane exceeded its 10-minute job cap, as confirmed by its check annotations; the initial fail-fast attribution was incorrect.
At the follow-up head, parallel shard 2 also exceeded the 10-minute cap and could not publish its timing artifact.
The workflow explicitly reserves cap and lane-count changes for a separate scope decision; no such change has been made.
The final changed files passed the complete merge suite (159.816s), task-hold suite (167.232s), issue-writeback suite (18.929s), and brief suite (17.452s), with zero failures or suite-level gate skips.
Full canonical CI=true bin/fm-lint.sh passed with pinned ShellCheck 0.11.0 extended analysis and actionlint 1.7.12 for all three workflows; scoped canonical lint also passed for the final executable/test roots.
Documentation audience and pointer checks passed.
The preceding run, https://github.com/HelloWorldSungin/firstmate/actions/runs/34873884448, failed parallel shard 2 at its job deadline and independently failed the remote trace-context fixture in serial shard 4.
Neither failure is waived by the sync policy exception.
An initial non-frozen probe reported a legacy merge-retry failure while sources were being edited; both subsequent complete frozen four-suite runs passed without a workaround for that case.
The preceding full merged sweep predates this follow-up; it is not claimed as new-code coverage, and old CI checks are not offered as validation of the new head.

Bounded CI correction

Firstmate authorized bounded attribution and the smallest supported correction, with no CI cap or lane-count change.
Completed logs show the earlier parallel shard 1 completed four passing scripts totaling 575.646 seconds before cancellation, and the next run's parallel shard 2 completed six passing scripts totaling 600.905 seconds before cancellation.
Both continued into another script; the logs support cumulative serial workload as the immediate deadline cause, not a demonstrated hung script or a general budget guarantee.
The existing admitted isolated pool now runs with two workers inside each of the same two parallel jobs.
All job and per-script bounds, lane counts, membership and runner logic remain unchanged.

Serial shard 4 failed during local clone with an object-copy ENOENT error.
A controlled real-Git clone/repack experiment reproduces that exact failure while source fsck remains clean; the no-repack control passes.
Natural local repetitions and the clean tracked-only fixture passed, which limits any claim of deterministic reproduction.
The fixture commit now waits for automatic maintenance before starting its clone, and its full remote route passes with the trace confirming non-detached maintenance.
The failed CI run did not capture maintenance process tracing: attribution of that particular instance to the demonstrated race is an inference from the fixture lifecycle and exact symptom.
Production clone behavior is unchanged.

Current-head local validation: complete portable isolation proof 24/24 passed, no gate skips, two workers on two CPUs, 434.026 seconds, with no surviving proof root or process retaining its TMPDIR.
The actual two-worker parallel lane invocations passed 12/12 in 219.363 seconds and 12/12 in 193.454 seconds, without gate skips.
These local measurements establish execution and isolation evidence, not hosted-runner speed.
A same-host, same-affinity one-worker counterfactual also passed all 12 lane-2 scripts in 290.852 seconds, compared with 193.454 seconds using two workers; it supports the effect of overlap while disproving any claim that local serial execution itself reliably exceeds the CI deadline.
The initial isolation attempt correctly failed its missing-TypeScript prerequisite; a task-private TypeScript 7.0.2 installation enabled the complete proof, with real Pi 0.85.1 and Ruby 3.4.9.
The complete trace-context and workflow suites passed, as did full canonical CI lint, final scoped fixture lint, documentation audience checks, pointer checks and diff whitespace checks.
The prior local 258-script sweep remains attributed to its earlier tree; complete hosted coverage of the final head is recorded above.

Process-event fixture follow-up

The bounded-CI head completed with 15 substantive jobs passing and one failure in serial shard 6: the deleted-artifact case in tests/fm-procevent.test.sh received the existing fail-closed runner-identity refusal during retirement.
Its exact fixture block is unchanged from the fork parent; the pinned upstream parent lacks this fork-specific case.
The complete unmodified local suite passed, but the isolated real-command case reproduced the same refusal naturally on trial seven after six passes.
Capturing a poll result precedes the runner's exit and claim release, so counting results is not evidence that retirement can safely prove a runner that is finishing.
The fixture now waits, with an asserted bound, for that completed runner to release its claim before checking retirement by source id.
Production identity guards, leaderless-group refusals, live-runner stopping cases and process cleanup behavior are unchanged.
All twelve isolated corrected repetitions passed, followed by the complete process-event suite in 182.381 seconds with no failure or gate skip.
The failed hosted run did not trace the exact disappearing-process transition; the causal attribution is supported by the natural reproduction and the capture-before-exit implementation, not an invented CI process trace.
The earlier parallel-job correction passes in hosted CI: both lanes ran all twelve scripts without gate skips in 464.365 and 385.901 seconds under unchanged caps.
The remote trace-context fixture also passed in serial shard 4 in 75.428 seconds.
Full canonical CI=true bin/fm-lint.sh passed after this fixture correction, including pinned extended ShellCheck and all three actionlint workflow checks.
Final-head hosted artifacts cover all 258 unique scripts with zero failures and 38 explicit gated skips; the opt-in gaps are disclosed and earlier supplementary live runs remain separately attributed.
The final parallel lane durations are 485.167 and 266.148 seconds, both without gate skips.
The exact final run is https://github.com/HelloWorldSungin/firstmate/actions/runs/34881222435; all 16 substantive jobs completed successfully, including hosted canonical lint.

Upstream applicability

Commit Upstream change Applicability and verification
869ae905779c kunchenguid/firstmate#4151 Applied with both-parent duration maxima and the fork's eight shards/480-second bounds; runner and CI workflow suites.
0a4e4a27cb7e kunchenguid/firstmate#4148 Applied delivered-PR attribution rules; public-followup, crew-state, inactive-reconcile, and teardown suites.
8a61b5054588 kunchenguid/firstmate#4212 Applied launch confirmation, dead-claim reclamation, and stranded-source wakes; procevent and watcher suites.
e0d269e07318 kunchenguid/firstmate#4191 Applied positive process classification and stale registration states with the fork's stricter terminal-wide absence proof; backend regression and real Pi guard.
31f062d43d25 kunchenguid/firstmate#4199 Applied live-head/check gates and explicit away grants while retaining fork target/tip, flag, no-rebase, and landed-readback contracts; complete merge suite.
9074f9d20d3d kunchenguid/firstmate#4221 Applied frozen-epoch reblock charging; turnend-guard tests and existing fork signal behavior retained.
ff5c7af9d765 kunchenguid/firstmate#4242 Applied guarded named-lab foreground viewer; lab unit and attached-viewer live suites.
dee156fe394a kunchenguid/firstmate#3766 Applied nonvisual fallback when Lavish is unavailable; bootstrap and brief suites.
eb0ea3ab862c kunchenguid/firstmate#3825 Applied global/system Git-config isolation and consolidated every synthetic runner constructor around its required helper; complete runner/fixture suites.
2c1017e5257f kunchenguid/firstmate#4239 Applied inherited Claude permission configuration through the fork's sole launch renderer; dispatch and inheritance suites.
7d14fc126aa8 kunchenguid/firstmate#4243 Applied reassigned Treehouse slot refusal without dropping fork exact branch cleanup or preserved records; endpoint-safety suite.
9e1e85e2fa15 kunchenguid/firstmate#4247 Applied scoped brief-filling instructions; retained fork worker-role, isolated-worktree, and definition-of-done contracts.
46ff99ccf9c7 kunchenguid/firstmate#4245 Applied durable underway names and filed-date ordering; snapshot and board-render suites.
ad14a8db06fd kunchenguid/firstmate#4248 Applied action-free catch-up Bearings warnings; snapshots may report while the ordinary-work return gate remains enforced.
a27646c4eae5 kunchenguid/firstmate#4223 Applied home-addressed tasks wrapper and duplicate code-root backlog detection; tasks, bootstrap, session-start, and hold suites.
92856fd4ffeb kunchenguid/firstmate#4262 Applied Claude secondmate home trust; retained explicit fork CLAUDE_CONFIG_DIR handling and shared launch construction.
c191eacd36a9 kunchenguid/firstmate#4258 Applied superseded-failure check selection; merge tests retain current failures, unfinished reruns, and separate legacy contexts.
de00521b9e67 kunchenguid/firstmate#4266 Applied identity-bound durable merge authority for poll outcomes; PR security suite covers returned-away grants, rebinding, and lifecycle locks.
83c63cc6aa79 kunchenguid/firstmate#4281 Applied superseded-run cancellation and job caps; fork shard count and substantive check surfaces preserved.
fa65b5df1074 kunchenguid/firstmate#4288 Applied optional tasks-axi gating for the away-record fixture in the fork's split watcher test owner.
fb19dd9a75f2 kunchenguid/firstmate#4027 Applied bounded per-item backlog reads and explicit omitted-row reporting; backlog-read-bound and session-start suites.
76d44055d810 kunchenguid/firstmate#4285 Applied synchronous away execution and locked grant checks; retained the fork's separate unconditional forwarded-auto refusal.
35761484e9f0 kunchenguid/firstmate#4200 Applied compatible agy discovery/model/readiness behavior through existing fork launch/trust owners; retained crew-only/Herdr-only, atomic submission, and cleanup custody; real vendor guard and three complete native trials.
b518a256e89c kunchenguid/firstmate#4337 Applied attended quiet; Firstmate resolved upstream's incompatible away-only entry/return gates without weakening away confirmation or adding authority; quiet lifecycle tests.
0962d4a02375 kunchenguid/firstmate#3578 Applied verified-ancestry precedence over retained markers; precedence suite and real installed-harness guard.
305dfffa9a53 kunchenguid/firstmate#3944 Applied external CLAUDE.md import trust and secondmate scope validation; Claude trust and secondmate tests.
ecfe071981b1 kunchenguid/firstmate#4355 Applied acknowledged watcher-down handling in the return gate; AFK return tests.
b182d0f908b7 kunchenguid/firstmate#4361 Applied exact when-source trust rebinding after self-update; procevent-when and update suites. This is the pinned final change.

Active divergence survival

Ledger entry Surviving behavior Regression evidence
Agy crew adapter Crew-only/Herdr-only and raw-launch refusals, owned trust cleanup, native busy identity, inbox atomic prompting, and distinct unverifiable delivery; three complete real native trials also passed. tests/fm-agy-adapter.test.sh
Attended quiet lifecycle Confirmed away entry still gates away mode; ordinary work preserves quiet, explicit quiet-off clears only quiet, and actual away records/catch-up gates cannot be consumed as quiet. tests/fm-afk-launch.test.sh, tests/fm-afk-return.test.sh
Herdr terminal-wide agent absence Suspended, backgrounded, reparented same-terminal, and escape-shell agents refuse absence; real Pi stays alive until exit and then becomes stale/dead without destructive husk authority. tests/fm-backend-herdr.test.sh, tests/fm-herdr-pi-stale-registration-live-e2e.test.sh
Installed AXI compatibility floors Lavish remains optional, but a below-floor installation cannot enable presentation instructions; the retained installed-version floor accepts 0.1.62 and newer compatible releases. tests/fm-bootstrap.test.sh, tests/fm-brief.test.sh
Pinned ShellCheck download retry budget Installer fixture exercises transient download recovery and exhaustion under the fork wall-time budget. tests/fm-lint.test.sh
LLM quota sidecar Reader preserves measured age and unknown/error distinctions without performing ranking or replacing quota-axi selection. tests/fm-quota-sidecar.test.sh
Scout completion gate reopened by a firstmate steer Typed success/unconfirmed/failure and durable inbox publication exercise reopening and the correct rollback direction. tests/fm-send-strict.test.sh
Watcher restart hand-over Restart obtains the next singleton lock within budget and proves the predecessor is gone. tests/fm-watcher-lock.test.sh
Watcher stop-signal disposition TERM exits 143 without parsed signal handlers; lock and caller-owned Herdr event-wait resources are released. tests/fm-backend-herdr.test.sh, tests/fm-watcher-lock.test.sh
Keyed decision repaint suppression Explicit keys suppress repeated repaint alarms; unkeyed blockers still alarm, dead agents recover, and acknowledgement preserves durable decisions. tests/fm-watch-triage.test.sh
Secondmate queue-stall semantic thresholds Separate idle/unknown thresholds, progress reset, reused endpoints, and bounded busy suppression remain exercised. tests/fm-wake-queue.test.sh
Run-progress wedge hold A demonstrably advancing validation run holds wedge escalation; stalled progress cannot hold it indefinitely. tests/fm-run-progress.test.sh, tests/fm-watch-triage.test.sh
Watcher live declared-wait routing and self-widening recheck cadence Live declared waits take paused routing and widen only after emitted reminders; daemon and watcher preserve handoff behavior. tests/fm-daemon.test.sh, tests/fm-watch-triage.test.sh
Herdr pre-Enter footer read on a native working baseline A working baseline cannot authorize typing past a newly rendered interactive footer. tests/fm-backend-herdr.test.sh
Remote job worker descendant reaping Worker shutdown reaps descendants, including orphan cases, without broad process-pattern kills. tests/fm-remote-job-orphan-reap.test.sh, tests/fm-remote-job.test.sh
Bounded remote job stdin capture Unclosed explicit stdin refuses within its bound; finite payloads remain byte-exact and queue deadlines start after capture. tests/fm-remote-job.test.sh
Default per-script bound on every test sweep Public runner exercises default arming, opt-out zero, timeout 124 versus refusal 125, and terminal-signal relay; fixture helper probes pass for every constructor. tests/fm-on.test.sh, tests/fm-test-run.test.sh, tests/fm-watch-triage-waits.test.sh, tests/fm-watch-triage.test.sh
Locale-independent test coverage comparisons Public coverage commands are exercised under installed locales with C-collation retained for inventory comparisons. tests/fm-test-run.test.sh
Queued wakes remain a supervision requirement Pending durable work keeps supervision required even when presentation ownership belongs to another actor. tests/fm-guard-stale-banner.test.sh, tests/fm-turnend-guard.test.sh, tests/fm-wake-queue.test.sh
No-mistakes run attribution Relation-table, branch identity, abandoned/degraded replay, pipeline custody, and strict teardown attribution cases remain covered. tests/fm-crew-state.test.sh, tests/fm-teardown.test.sh
Definition-of-done owner carries this fork's ready-to-validate handoff Generated worker briefs and relaunch instructions preserve the fork handoff and scoped self-contained intent. tests/fm-brief.test.sh, tests/fm-control-relaunch.test.sh, tests/fm-design-skills.test.sh, tests/fm-trigger-validation.test.sh
A preserving refusal withdraws the pending backlog close Refused teardown restores the pre-transition record and cannot leave a pending replay that closes preserved work. tests/fm-backlog-atomicity.test.sh
A watermark-capture failure keeps the published task record Spawn refusal retains the published task's metadata/endpoint and retry evidence after a failed watermark capture. tests/fm-spawn-dispatch-profile.test.sh
Fleet snapshot per-task timeout and abandoned child work Hung per-task state reads remain bounded and abandoned child work is reaped rather than blocking fleet projection. tests/fm-dashboard-backlog.test.sh, tests/fm-home-summary-refresh.test.sh
Pi and OMP away-mode supervision standby Extension interfaces retain flag-present standby, child retirement, pending-wake preservation, and one-cycle resumption; actual OMP binary is absent and no new vendor proof is claimed. tests/fm-afk-launch.test.sh, tests/fm-omp-harness.test.sh, tests/fm-pi-watch-extension.test.sh
Herdr presentation fixture ownership and cleanup Real named-lab presentation/recovery and cleanup suites retain source ownership, exact focus safety, task-specific journals, and failure-preserving cleanup. tests/fm-backend-herdr-presentation-e2e.test.sh, tests/fm-backend-herdr-recovery-e2e.test.sh, tests/fm-test-fixture-cleanup.test.sh
Fork-local no-mistakes compliance-gate event scope Gate fixtures preserve fork event scope and source selection without running the prohibited validation pipeline for this task. tests/fm-no-mistakes-required-gate.test.sh
Live pull-request body refresh before the compliance gate Gate fixtures exercise live-body refresh and best-effort warning fallback to the webhook body. tests/fm-no-mistakes-required-gate.test.sh
Merge-proof contract: one divergence accepted, one retired Complete merge suite preserves attended queue reporting and successful landed readback after command failure, alongside new away-only synchronous restrictions. tests/fm-pr-merge.test.sh
Upstream tracking mechanism Detector fixtures prove no-upstream inertness and read-only source handling; this round's exact ancestry and parked-commit intersections are separately recorded. The merge suite proves the reviewed-sync policy exception and its refusal boundaries through real Git graphs and mocked forge transport. tests/fm-upstream-status.test.sh, tests/fm-pr-merge.test.sh
Repository-local validation evidence Effective store_in_repo remains false; actual generated test artifacts resolve to the existing .no-mistakes ignore rule and none enter the staged tree. No pipeline was invoked, as required. Effective configuration and actual delivery/evidence observations described in this row.
Upstream-read-only posture in shared tracked docs Hard-rule regressions pass; both the inspected origin push URL and forge repository metadata resolve to the owned fork, with upstream used only for reads. tests/fm-agents-hard-rules.test.sh
Supervisor-only chat address Public spawn tests execute generated ship/scout launch commands and verify the worker role is delivered, preserving the supervisor/worker distinction. tests/fm-spawn-dispatch-profile.test.sh
GBrain per-home knowledge memory Capture/recall/bootstrap/brief/inheritance tests preserve per-home writes and absent-brain behavior; complete real GBrain 0.46.21.0 scope proof passed with disposable credential-isolated brains, denied main-brain writes, separate client cursors, and offline own-brain search. tests/fm-bootstrap.test.sh, tests/fm-brief.test.sh, tests/fm-gbrain-capture.test.sh, tests/fm-gbrain-readonly-e2e.test.sh, tests/fm-recall.test.sh, tests/fm-remote-secondmate-lifecycle-e2e.test.sh
Fleet dashboard and agent-event instrumentation HTTP/DOM/event tests retain read-only projection, opt-in instrumentation and hung-dashboard guards; complete real-browser proof passes phone, 899/900 boundary, and desktop rendering, empty-page rejection, injected failures and exact observation accounting in a dedicated session. tests/fm-bearings-snapshot.test.sh, tests/fm-dashboard-browser.test.sh, tests/fm-dashboard-events.test.sh, tests/fm-dashboard-gbrain-ui.test.sh, tests/fm-dashboard-gbrain.test.sh, tests/fm-dashboard.test.sh

Per-file contract review

Each maintained instruction, document, workflow, executable contract, and shared test helper changed by this round is accounted for below.

File Reconciliation reasoning
.agents/skills/afk/SKILL.md Adopt explicit task merge grants, synchronous away merges, and read-only catch-up Bearings; retain confirmed away records, return gates, and the fork's Pi/OMP daemon ownership.
.agents/skills/bearings/SKILL.md Adopt durable task names/filed ordering and the catch-up warning without turning warning rows into dispatchable work; retain fork snapshot and board contracts.
.agents/skills/bearings/assets/board-template.html Adopt the exact pinned upstream version for fix: identify underway tasks and sort charted work. No fork-specific implementation is displaced in this file.
.agents/skills/bootstrap-diagnostics/SKILL.md Add optional-presentation diagnostics and the home-addressed tasks wrapper while retaining fork GBrain, attribution, usage, and other existing diagnostic routes.
.agents/skills/firstmate-coding-guidelines/SKILL.md Update the quiet stub example; keep the fork's one-owner, audience, behavioral-test, lint, and no-agent-coauthor rules.
.agents/skills/fmx-respond/SKILL.md Route backlog creation through the home-addressed wrapper; preserve public-followup reconciliation and fork delivery boundaries.
.agents/skills/harness-adapters/SKILL.md Include upstream agy discovery and Claude permission configuration in the established routing artifact; retain agy crew-only/Herdr-only eligibility.
.agents/skills/harness-adapters/references/common/control-and-recovery.md Adopt the exact pinned upstream version for feat(bin): add Antigravity CLI (agy) as third worker/scout adapter; fix: pre-register Claude trust for secondmate homes. No fork-specific implementation is displaced in this file.
.agents/skills/harness-adapters/references/harness/agy.md Consolidate compatible upstream model/trust/readiness observations into the existing adapter; keep native Herdr authority, no raw-launch bypass, cleanup custody, and fork delivery verdicts.
.agents/skills/harness-adapters/references/harness/claude.md Adopt the exact pinned upstream version for fix(bin): pre-approve external CLAUDE.md import dialog for spawned workers; fix: pre-register Claude trust for secondmate homes; feat(bin): add config/claude-permission-mode to launch Claude workers in auto mode. No fork-specific implementation is displaced in this file.
.agents/skills/harness-adapters/references/harness/codex.md Adopt the exact pinned upstream version for fix(bin): let verified harness ancestry outrank retained markers. No fork-specific implementation is displaced in this file.
.agents/skills/harness-adapters/references/harness/cursor.md Adopt the exact pinned upstream version for fix(bin): let verified harness ancestry outrank retained markers. No fork-specific implementation is displaced in this file.
.agents/skills/harness-adapters/references/harness/kimi.md Adopt the exact pinned upstream version for fix(bin): let verified harness ancestry outrank retained markers. No fork-specific implementation is displaced in this file.
.agents/skills/harness-adapters/references/harness/muse.md Adopt the exact pinned upstream version for fix(bin): let verified harness ancestry outrank retained markers. No fork-specific implementation is displaced in this file.
.agents/skills/harness-adapters/references/harness/opencode.md Adopt the exact pinned upstream version for fix(bin): let verified harness ancestry outrank retained markers. No fork-specific implementation is displaced in this file.
.agents/skills/process-event-sources/SKILL.md Adopt launch-confirmation and stranded-claim handling while preserving durable captured results, handled acknowledgement, and the existing one-owner boundary.
.agents/skills/quiet/SKILL.md Firstmate resolved the upstream entry/return contradiction as attended quiet on the shared daemon, without an away record, extra merge authority, or implicit chat exit.
.agents/skills/stow/SKILL.md Use the home-addressed tasks wrapper without dropping fork GBrain capture or startup-memory maintenance.
.agents/skills/sync-upstream/SKILL.md Apply the explicit decision relayed by Firstmate: standing Firstmate review/landing authority for sync rounds, with workers still stopping at the PR; point to the merge guard as the one owner of the exact-head policy exception and require a fresh declaration after any new head.
.github/workflows/ci.yml Adopt superseded-run cancellation and bounded jobs; retain fork Node 22, eight serial shards, download retries, and the required Herdr lane. Run two already-admitted isolated workers inside each existing parallel job after complete two-worker isolation proof; caps, lane counts and membership remain unchanged.
AGENTS.md Integrate explicit grant/check-waiver ownership, scoped brief filling, optional visual tooling, tasks addressing, and quiet routing; preserve upstream-read-only policy, supervisor-only address, fork memory, and Pi/OMP standby.
CONTRIBUTING.md Expose the incoming configuration/workflow capabilities while keeping fork lint, documentation ownership, and contributor validation requirements.
README.md Add the upstream quiet entry point without moving detailed operational contracts into the public overview.
bin/backends/herdr.sh Adopt shared positive harness classification and stale-registration states; reuse the fork's terminal-wide negative proof, preserve atomic agy delivery and unknown verdicts, and retain focus/ownership/cleanup safety.
bin/backends/tmux.sh Adopt the exact pinned upstream version for fix(herdr): verify agent liveness at process level before trusting registration. No fork-specific implementation is displaced in this file.
bin/fm-afk-contract.sh Adopt the exact pinned upstream version for fix(merge): serialize away authority with synchronous merges; feat: add live-head merge gates and away task grants. No fork-specific implementation is displaced in this file.
bin/fm-afk-launch.sh Adopt quiet mode selection on the existing daemon; confirmed-away entry remains required outside attended quiet, and quiet refuses a real away return still needing reconciliation.
bin/fm-afk-return.sh Adopt acknowledged watcher-gap handling and mode-aware return checks; explicit quiet-off stops the shared daemon without fabricating or consuming an away return.
bin/fm-afk-start.sh Adopt the exact pinned upstream version for feat(afk): add quiet supervision mode for a present captain. No fork-specific implementation is displaced in this file.
bin/fm-agent-process-lib.sh Adopt the exact pinned upstream version for feat(bin): add Antigravity CLI (agy) as third worker/scout adapter; fix(herdr): verify agent liveness at process level before trusting registration. No fork-specific implementation is displaced in this file.
bin/fm-agy-trust-lib.sh Keep one locked mutation owner with created/preexisting custody and exact cleanup; add owned regular-file validation and external-write detection before atomic replacement.
bin/fm-agy-trust.sh Retain upstream linked-worktree scope and logical/resolved path validation; delegate settings writes to the fork's trust library instead of retaining a second writer.
bin/fm-backend.sh Document upstream process-level recovery while retaining fork caller-owned event-wait cleanup and native agy submission semantics.
bin/fm-backlog-handoff.sh Adopt the exact pinned upstream version for fix(bin): address the home's backlog from any directory and detect a forked code-root copy. No fork-specific implementation is displaced in this file.
bin/fm-backlog-transition-lib.sh Adopt the exact pinned upstream version for fix(backlog): bound per-item backlog row reads so a wedged backend cannot blind a session start. No fork-specific implementation is displaced in this file.
bin/fm-bearings-board.sh Adopt the exact pinned upstream version for fix: identify underway tasks and sort charted work. No fork-specific implementation is displaced in this file.
bin/fm-bearings-snapshot.sh Adopt catch-up warnings, durable labels, and filed ordering; preserve fork large-payload slurpfile use and dashboard read-only cache mode.
bin/fm-bootstrap.sh Adopt optional Lavish and code-root backlog detection; preserve fork GBrain/usage/attribution checks, required tool floors, and bounded startup behavior.
bin/fm-branch-prompt.sh Adopt the exact pinned upstream version for fix(bin): address the home's backlog from any directory and detect a forked code-root copy. No fork-specific implementation is displaced in this file.
bin/fm-brief.sh Gate visual scout instructions on compatible Lavish; preserve worker role, isolated-copy, brain recall, inbox, and fork definition-of-done instructions.
bin/fm-busy-lib.sh Adopt upstream agy recognition helpers while preserving native identity validation and refusing screen fallback for an unknown Herdr agy state.
bin/fm-captain-hold.sh Adopt the exact pinned upstream version for fix(backlog): bound per-item backlog row reads so a wedged backend cannot blind a session start. No fork-specific implementation is displaced in this file.
bin/fm-classify-lib.sh Own the read-only away/quiet flag interpretation so return guards need not initialize a wake queue; preserve all fork decision and run-progress classifiers.
bin/fm-claude-trust.sh Adopt the exact pinned upstream version for fix(bin): pre-approve external CLAUDE.md import dialog for spawned workers; fix: pre-register Claude trust for secondmate homes. No fork-specific implementation is displaced in this file.
bin/fm-composer-lib.sh Adopt the exact pinned upstream version for feat(bin): add Antigravity CLI (agy) as third worker/scout adapter. No fork-specific implementation is displaced in this file.
bin/fm-config-inherit-lib.sh Add Claude permission posture to the inherited set while retaining GBrain, project-board, and launch-environment validation.
bin/fm-control-lib.sh Use exact agy family recognition and the live-verified /quit command; retain crew-only control eligibility and native interrupt semantics.
bin/fm-crew-state.sh Explain stale registration as positive recovery evidence without changing fork no-mistakes attribution, abandoned-run, or unknown-state safety.
bin/fm-decision-hold.sh Adopt the exact pinned upstream version for fix(bin): address the home's backlog from any directory and detect a forked code-root copy. No fork-specific implementation is displaced in this file.
bin/fm-fleet-snapshot.sh Add durable names and bounded filed-date ordering; preserve per-task timeouts, child reaping, quiet allowance, and fork dashboard fields.
bin/fm-guard.sh Expose quiet-mode supervision while retaining the fork's queued-wake and stale-banner behavior.
bin/fm-harness.sh Adopt the exact pinned upstream version for fix(bin): let verified harness ancestry outrank retained markers; feat(bin): add Antigravity CLI (agy) as third worker/scout adapter. No fork-specific implementation is displaced in this file.
bin/fm-herdr-lab-viewer.py Adopt the exact pinned upstream version for feat(herdr): add guarded foreground viewer for live validation. No fork-specific implementation is displaced in this file.
bin/fm-herdr-lab.sh Adopt the exact pinned upstream version for feat(herdr): add guarded foreground viewer for live validation. No fork-specific implementation is displaced in this file.
bin/fm-inactive-reconcile.sh Adopt the exact pinned upstream version for fix(bin): stop claiming prose-mentioned PR URLs as a task's delivered PR. No fork-specific implementation is displaced in this file.
bin/fm-launch-lib.sh Keep the sole launch-template owner; render explicit Claude permission and agy binary bindings without rescanning substituted paths.
bin/fm-merge-authority-lib.sh Adopt the exact pinned upstream version for fix(bin): persist merge authority for poll-detected outcomes. No fork-specific implementation is displaced in this file.
bin/fm-merge-outcome-lib.sh Adopt the exact pinned upstream version for fix(bin): persist merge authority for poll-detected outcomes; feat: add live-head merge gates and away task grants. No fork-specific implementation is displaced in this file.
bin/fm-pr-merge.sh Combine verified-head/check gates and locked away authority with the fork's default-target/tip checks, forwarded-flag refusals, GitLab no-rebase boundary, and successful landed readback despite command failure. Add the Firstmate-authorized expected-policy exception only for reviewed direct-PR sync tasks, verified fork push identity, live branch/head, pinned base/endpoint and one upstream merge; preserve every other check/authority/hold/default-tip gate and require explicit merge even with green policy CI.
bin/fm-procevent-lib.sh Adopt the exact pinned upstream version for fix(procevent): confirm reconcile launches and reclaim provably dead claims instead of counting a dead drop as started. No fork-specific implementation is displaced in this file.
bin/fm-procevent-when.sh Adopt the exact pinned upstream version for fix(bin): rebind fm-procevent-when trust bindings after a self-update. No fork-specific implementation is displaced in this file.
bin/fm-procevent.sh Adopt the exact pinned upstream version for fix(procevent): confirm reconcile launches and reclaim provably dead claims instead of counting a dead drop as started. No fork-specific implementation is displaced in this file.
bin/fm-public-followup.sh Adopt the exact pinned upstream version for fix(bin): address the home's backlog from any directory and detect a forked code-root copy. No fork-specific implementation is displaced in this file.
bin/fm-send.sh Use home-addressed backlog reads and adopt the longer agy typed confirmation budget; retain scout completion reopening and native atomic inbox/typed delivery verdicts.
bin/fm-session-lock-lib.sh Adopt the exact pinned upstream version for fix(bin): let verified harness ancestry outrank retained markers. No fork-specific implementation is displaced in this file.
bin/fm-session-start.sh Adopt bounded home-addressed backlog pointers and explicit quiet guidance while retaining fork memory, attribution, lock, and startup ownership.
bin/fm-sessionstart-nudge.sh Adopt the exact pinned upstream version for fix(bin): let verified harness ancestry outrank retained markers. No fork-specific implementation is displaced in this file.
bin/fm-spawn.sh Combine upstream agy model/readiness and Claude trust/permission work with fork centralized launch rendering, exact trust custody, worker-role delivery, watermark refusal, and restricted agy scope.
bin/fm-supervision-instructions.sh Adopt the exact pinned upstream version for feat(afk): add quiet supervision mode for a present captain. No fork-specific implementation is displaced in this file.
bin/fm-tasks-axi.sh Adopt the exact pinned upstream version for fix(bin): address the home's backlog from any directory and detect a forked code-root copy. No fork-specific implementation is displaced in this file.
bin/fm-teardown.sh Adopt reassigned-slot refusal and authority-record cleanup; preserve exact branch reaping, agy trust removal, GBrain capture, usage/manifests, and preserving-refusal ordering.
bin/fm-test-run.sh Adopt Git-config isolation and refreshed duration packing using parent maxima; retain all-mode 480-second bounds, exit/signal behavior, C-locale coverage, eight shards, and fork-only suites. Update comments to distinguish concurrent wall time from serial hint sums; no additional runner logic, admission, bound or weight change.
bin/fm-turnend-guard.sh Adopt per-reblock charging for a frozen Claude auto-arm epoch while retaining fork pending-wake and safe stop behavior.
bin/fm-update.sh Adopt the exact pinned upstream version for fix(bin): rebind fm-procevent-when trust bindings after a self-update. No fork-specific implementation is displaced in this file.
bin/fm-wake-lib.sh Adopt quiet-aware shared supervision and backlog bounds; delegate mode interpretation to classify-lib while preserving fork wake ownership and queue guarantees.
bin/fm-watch.sh Adopt confirmed/stranded process-source headlines and identity-bound merge authority consumption; retain default signal disposition, restart handover, keyed suppression, declared-wait cadence, and run-progress holds.
bin/fm-x-link.sh Adopt the exact pinned upstream version for fix(bin): address the home's backlog from any directory and detect a forked code-root copy. No fork-specific implementation is displaced in this file.
docs/agent-control.md Describe incoming process-level recovery and verified control behavior without relaxing native agy or fork recovery refusals.
docs/architecture.md Integrate merge authority, source confirmation, quiet, and launch changes at their existing owners; preserve fork worker attribution, watcher, memory, and dashboard architecture.
docs/captain-hold-lifecycle.md Point task creation at the home-addressed wrapper without changing durable held-task completion or answer authority.
docs/cd-guard.md Adopt the exact pinned upstream version for fix(bin): address the home's backlog from any directory and detect a forked code-root copy. No fork-specific implementation is displaced in this file.
docs/configuration.md Add Claude permission, bounded backlog, quiet, and new upstream configuration fields; retain fork private layout, GBrain/sidecar/dashboard, and inherited-config schemas.
docs/documentation-audiences.json Classify incoming agy verification and quiet surfaces without removing fork documentation owners or duplicating an inventory key.
docs/fm-test-isolation-proof.md Add the complete 24-script two-worker proof with prerequisites present, two-CPU affinity and process/temp cleanup evidence; explicitly avoid a hosted timing guarantee.
docs/fm-test-portable-shards.md Record combined parent timing maxima and upstream job caps while keeping fork eight-shard and per-script timeout guarantees distinct from estimates. Distinguish historical serial timing inputs from current overlapping execution and point to the workflow owner.
docs/fork-divergence.md Record attended quiet, terminal-wide Herdr absence, agy owner reconciliation, current timing policy, and internal head binding; no parked branch is retired or adopted. Record the upstream-sync review/landing authority and exact-head exception by pointers to their owners. Record overlapping admitted scripts within existing jobs, without changing eight serial shards or any timeout, and point to the isolation proof.
docs/gitlab-merge-watch.md Adopt the exact pinned upstream version for feat: add live-head merge gates and away task grants. No fork-specific implementation is displaced in this file.
docs/herdr-backend.md Adopt stale-registration and guarded-viewer descriptions while preserving fork atomic agy delivery, focus safety, and strict absence/removal distinctions.
docs/scripts.md Index incoming commands at their authoritative headers; retain fork-only script and subsystem entries.
docs/sessionstart-nudge.md Adopt the exact pinned upstream version for fix(backlog): bound per-item backlog row reads so a wedged backend cannot blind a session start. No fork-specific implementation is displaced in this file.
docs/tmux-backend.md Point process recognition at upstream's shared classifier without changing fork isolation or supported-runtime boundaries.
docs/turnend-guard.md Explain frozen-epoch charging at the existing Claude guard owner while preserving fork queued-work and stop-lifecycle guarantees.
docs/verification/agy.md Retain upstream vendor observations while correcting fork-specific trust ownership, fail-closed registration, native-only busy authority, and measured startup liveness; link the current native trial evidence.
docs/verification/gbrain-readonly-share.md Retain historical versioned evidence and add the current complete live read-only scope regression, including isolated runtime credentials, client cursor separation, denied writes, and offline own-brain search.
docs/verification/muse.md Adopt the exact pinned upstream version for fix(bin): let verified harness ancestry outrank retained markers. No fork-specific implementation is displaced in this file.
docs/verification/process-event-sources.md Adopt launch-confirmation and stranded-claim evidence beside the fork's existing capture/acknowledgement verification.
docs/verification/rovo.md Adopt the exact pinned upstream version for fix(herdr): verify agent liveness at process level before trusting registration. No fork-specific implementation is displaced in this file.
docs/verification/runtime-backends.md Retain compatible parent observations, refresh real installed-harness and agy/Pi evidence, and explicitly distinguish stale recovery from destructive cleanup and absent harnesses from tested ones.
tests/git-config-helpers.sh Adopt the single owner for global/system Git-config isolation while preserving repository-local and explicit caller configuration.
tests/herdr-client-pair-fixture.sh Serve the process-info evidence now required to back a compatible client registration, preserving stale-client refusal cases.
tests/herdr-test-safety.sh Source the Git-config owner before real-lab fixture Git operations; retain named-session tripwires and inherited-pane isolation.
tests/lib.sh Adopt shared Git-config isolation, clear ambient tasks-backlog overrides, and provide scoped ancestry fixtures while preserving fork event/dashboard isolation.
tests/remote-herdr-fixture.sh Return the structured pane process view for remote lifecycle fixtures instead of a pre-classifier placeholder response.
tests/watch-triage-helpers.sh Keep the fork split-suite owner while adding upstream confirmed-launch, stranded-claim, and claim-release synchronization cases.

The recomputed overlap contains 85 paths, not 85 textual conflicts; 38 files initially had textual conflicts.
The excluded later change contributes no exclusive overlap path.
docs/trace-context.md is an overlap whose final content remains the fork version: agy trace coverage was already present there.

mremond and others added 30 commits September 10, 2026 23:55
…nchenguid#4151)

* ci: rebalance the portable parallel lanes on measured runner durations

Both portable parallel lanes are capped at 10 minutes. Lane 1 was cancelled at
that cap on every request raised on 2026-09-10 while lane 2 finished in about
3.5 minutes, so no request could go green.

CONTRACT CLASS: RESTORE.
The workflow already promises two duration-balanced lanes and the shard
documentation already claims a measured wall; this re-establishes both against
what the lanes now cost, and changes no lane count, no cap, and no scope of what
runs. The counter-argument, so nobody has to take that on trust: two pieces here
are genuinely new rather than restored, and either could be argued to make this
a NEW-behavior change. `--list-scheduled` now ranks a parallel lane on measured
durations where it previously handed every parallel script the serial default
weight and returned an alphabetical order; and `--check-coverage` gains three
reported fields. I classify the change RESTORE because both exist only to make
the already-promised property checkable, but they are named here rather than
folded into the restoration.

=== PART 1: THE TOTAL, AND HOW IT WAS OBTAINED ===

This section stands on its own. It establishes what the parallel set costs. It
derives no packing; Part 2 does that, from this number.

THE TOTAL: 828568 ms, about 13 min 49 s of serial work across the 24 scripts.
Lane 1 held 624299 ms of it and lane 2 held 204269 ms, a 3.06:1 split.

HOW IT WAS OBTAINED. The difficulty was that lane 1 had never finished, so its
duration did not exist as a recorded figure anywhere and no timing artifact was
expected for it. It turned out to be recoverable from the real lane without
estimating, by two routes, across six CI runs on 2026-09-10 (34459949083,
34460760299, 34462530836, 34462758357, 34466966385, 34470382458):

  - Run 34462758357's lane-1 job finished its suite 18 s BEFORE the wall and
    uploaded a complete fm-test-timing-portable-parallel-1 artifact carrying all
    11 scripts, FM_TEST_SUMMARY total=11 failed=0 duration_ms=598225. The
    upload step is if: always(), so the cancellation did not suppress it. This
    is one full, untruncated lane-1 measurement.
  - The five other lane-1 jobs were cancelled mid-suite, but each logs every
    script that had already finished as an FM_TEST_END duration_ms= marker.
    Those per-script records are complete measurements of completed scripts;
    only the script in flight at cancellation is lost, and it differs by run.

Lane 2 completed in all six runs, so its scripts come from the six uploaded
fm-test-timing-portable-parallel-2 artifacts.

Every one of the 24 scripts therefore carries at least one untruncated
measurement: 20 of them measured in all six runs, two in three or four runs, and
two (fm-brief, fm-transition-lib, the tail of lane 1) in the single complete run.
Each hint is the SLOWEST value that script reached, so the total is an upper
envelope rather than an average. NO FIGURE IN IT IS DERIVED FROM A TRUNCATED
LANE, and no lower bound was ever extrapolated into a total.

THE ENVIRONMENT, AND WHETHER IT TRANSFERS. Every hint is a serial run of the
real portable parallel lane on a GitHub ubuntu-latest runner, produced by the
lane's own CI job. It transfers because it is not a proxy for the lane; it is
the lane. Nothing in the total came from this machine or from any harness of
mine.

That mattered, and here is what it would have cost. A same-day macOS
cross-check of the same scripts ran 1.7x to 5.0x slower with the ratio varying
per script (fm-test-run 157420 ms against 92944 ms, fm-x-mode 67217 ms against
31870 ms, fm-composer-ghost 10521 ms against 2120 ms). Local timings therefore
do not scale the lane, they REORDER it, so a packing derived from them would
have balanced the wrong thing while looking clean.

WHAT IT REPLACES, which is the root cause. The lanes were packed from the
2026-08-20 concurrent isolation proof: 24 candidates across four LOCAL workers.
That record answers whether the candidates are isolation-safe, not how long a
SERIAL CI lane runs, so it was structurally incapable of representing lane wall
clock even when it was fresh. It was also never refreshed while the set grew
about 3.2x. Both the wrong instrument and the staleness are fixed here: the
hints now come from the lane itself and carry their run ids and date.

=== PART 2: THE SPLIT DERIVED FROM THAT TOTAL ===

Longest-processing-time assignment over those hints gives 414269 ms and
414299 ms, 30 ms apart, against 624299/204269 before.

tests/fm-pi-primary-types.test.sh stays in lane 1 because that is the job which
installs the Pi package, so ci.yml needs no step changes.

=== PART 3: DOES THE MARGIN SURVIVE MACHINE VARIANCE ===

Stated explicitly, because 6.90 min against a 10 min cap is 69% of cap before
any variance is applied, and the cap covers the whole job rather than the suite.

  worst lane, script time                         414299 ms   6.90 min
  job overhead, measured on the real lane             ~18 s   (see below)
  expected healthy job                            ~432300 ms  7.21 min
  x1.29 on the script time, plus overhead         ~552400 ms  9.21 min
  cap                                             600000 ms  10.00 min
  room left after the multiplication                ~47.6 s   7.9% of cap

The 1.29x is the runner variance measured today on the SIBLING SERIAL lane, as
supplied; it is not this lane's own figure. This lane family does have its own,
and it is tighter: the six full lane-2 sums today span 192939 ms to 203451 ms,
a spread of 1.054x. At that figure the worst lane lands near 7.58 min with about
2.4 min of room. I have used the LARGER, borrowed 1.29x for the verdict rather
than the tighter one this lane actually shows, and note that the hints are
already per-script maxima, so 1.29x on top is conservative twice over.

THE MARGIN SURVIVES THE MULTIPLICATION, so this proceeds rather than stopping.
The 18 s overhead is measured, not assumed: in run 34462758357 the lane-1 job
ran 10 min 16 s against a 598.2 s suite, and lane 2 ran 3 min 21 s against a
192.9 s suite, a ~10 s difference that matches lane 1's extra Pi package install.

The cap is unchanged, the lane count is unchanged, and nothing in the serial
lane, its shard count, its guard or its hint table is touched.

=== PART 4: THE RECORDED FACT ===

The workflow comment no longer restates the shard wall as a literal, which is
how "~1 min of serial sum" survived a 10x change without announcing it. It now
points at bin/fm-test-run.sh --check-coverage, which prints parallel_max_ms,
parallel_imbalance_ms and parallel_unhinted derived from the hint table, so the
current number is computed on demand. The shard documentation carries the dated
run ids, which route it was taken by, and the local cross-check that shows why
local numbers are not admissible as hints.

Two regressions pin what rotted: lane membership must be stored
longest-measured-first, and the lanes must be fully hinted and packed within 5%
of each other. Both were run against the old composition and both fail on it
(420030 ms imbalance against a 624299 ms worst lane). The ordering assertion they
replace named a specific script by hand and had itself gone stale.

=== PART 5: NAMED AND LEFT, OUTSIDE THIS REBALANCE ===

tests/fm-captain-hold-lifecycle.test.sh alone is 296481 ms, 36% of the whole
set, so it is the floor of any two-lane split: no repacking can put a lane below
it. After this rebalance the cap is about 1.45x the healthy lane where the
sibling serial lane keeps roughly 2x.

Nothing refuses a stale parallel hint the way PORTABLE_SERIAL_MAX_UNHINTED_PERCENT
bounds the serial lane. parallel_unhinted is reported, not enforced, which is
what let this drift for three weeks unnoticed.

* fix(review): Restrict parallel scheduling hints to portable parallel lanes

* fix(document): Clarify parallel lane scheduling and timing evidence
… PR (kunchenguid#4148)

pr_for_task fell back to scraping the whole status log with tail -1, so
any PR URL a worker ever mentioned in prose - including a scout citing
someone else's PR - became the task's delivered PR in the parent-channel
terminal report. Recorded meta pr= is now the only authoritative source,
the fallback scrape accepts only a preferred terminal line in a mode's
ready-signal shape (done: PR <url> or done: PR <url> checks green), and
a scout never carries pr= at all.
…claims instead of counting a dead drop as started (kunchenguid#4212)

* fix(procevent): stop a dead runner owning a source and reconcile reporting it

The captain answered ten calls on a bearings board, the board accepted
them, and nothing collected them. He had to answer all ten again in chat.
A surface that presents as armed while being a dead drop is worse than one
that visibly fails, because the answers looked recorded.

Two independent defects, reproduced together in an isolated home where
reconcile reports started=1 on every run while ownership never moves and
no runner ever attaches.

1. reconcile counted a launch it never verified. detach_runner is
   fire-and-forget and discards the child's stderr, so a runner that died
   before it could claim was counted exactly like one that is listening.
   Launches are now confirmed - the source observed owned, or its runner
   record moved - before being reported as started; the rest are reported
   as failed= with a non-zero exit. The runner-record clause is what keeps
   a fast-completing source from being reported as a failure when it
   finished between two polls. One bounded window covers a whole cycle's
   launches, so a home full of broken sources costs the same wait as one.

2. A claim whose whole generation is provably gone could be refused
   forever. Reclaiming it ran cleanups over that dead generation's own
   leftovers, and any failure vetoed the claim - permanently, because none
   of those conditions clears on its own. Every one of those leftovers is
   keyed by the dead generation's claim token and a replacement always
   claims a fresh one, so none can collide with what replaces it.
   fm_procevent_claim_capture_reservation_reclaim_locked already said this
   for the reservation record; the staging file and the shape check on the
   registry directory recorded to hold it now take the same rule. Removing
   the claim record itself stays a hard precondition: two owners is the one
   outcome worse than none.

Two smaller repairs to the same "registered is not listening" confusion:

- `list` reported OWNER=none for a source nothing can claim. A reused PID
  whose process group survives reaches that state through the stale branch
  rather than the leaderless one, so it read as an idle source waiting to
  be started - the reassuring answer this surface gave while a board
  collected nothing. It now reports the orphaned state it shares.
- reconcile relaunched into that same unclaimable state on every cycle,
  spawning a runner that could only die on the claim. docs/configuration.md
  already promised it preserves such a claim without starting a
  replacement; the code now does that and reports it as uncertain.

This is NOT a third instance of today's two lock-identity defects
(4e1bf9aa and its replayed predecessor). Those were wrong liveness
predicates: a reused PID read as a live holder, then an exec'd holder read
as dead. Here the predicate is right - the code correctly proves the owner
dead and refuses the claim anyway, on a condition unrelated to liveness.

Regression coverage, each failing on the parent commit for its own reason:
- tests/fm-procevent.test.sh: a source that cannot start is reported as
  failed rather than started; a dead generation whose leftovers cannot be
  tidied no longer keeps owning its source (the parent reports a start
  while nothing ever runs); the existing reused-PID fixture now also
  asserts the orphaned listing and that no doomed relaunch is reported.
- tests/fm-captain-hold-lifecycle.test.sh: a board answer reaches the
  keyed-answer intake through the runner end to end - durable capture, the
  wake, and the closed task carrying the captain's selection. This one
  passes on the parent, because that chain was never what broke.

fm-procevent 100, fm-bearings-board 18, fm-captain-hold-lifecycle 50,
fm-procevent-when 13 and fm-procevent-quota 18 pass; bin/fm-lint.sh and
bin/fm-doc-audience-check.sh clean. tests/fm-extension-binding.test.sh has
two failures identical on the parent commit (EACCES on package install in
this sandbox) and unrelated to this change.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016gxgshn5jkWJ3GEYWy7vTG

* no-mistakes(review): confirm reconcile launches on durable launch stamps

* no-mistakes(review): announce stranded sources and refuse bad confirm windows

* no-mistakes(review): announce leaderless strands, bound confirm window, fix recovery docs

* no-mistakes(review): announce unconfirmed launches once per episode, qualify start reclaim

* no-mistakes(review): nonce launch-failed keys, refuse bad window at arm

* no-mistakes(review): state only observed launch outcome, shorten episode nonce

* no-mistakes(test): assert launch-failed headline not re-delivered, allow recovery wake

* no-mistakes(document): docs: cover strand and launch-failure wakes in skill trigger and verification record

* no-mistakes(lint): restructure SC2015 chain into explicit if-block

* test(watch-triage): fix two timing-exposed defects the pipeline found

Both surfaced in the no-mistakes test step on this branch, each failing one
full run of tests/fm-watch-triage.test.sh; neither was accepted as a flake to
retry past.

1. The new launch-failed delivery test assumed an already-surfaced key never
   wakes the watcher again. That is false: a fresh watcher legitimately
   re-surfaces any unacknowledged queue row through its downtime-recovery
   path ("check: rearm-resurface"), so the assertion failed whenever a
   re-arm landed between its two checks. The pipeline's own fix tolerated any
   wake lacking the repeated key's headline; this tightens it to exactly one
   tolerated reason, by its exact line, with a failure message that names the
   expectation so a reworded path reads as "the tolerated recovery path
   changed" rather than as a mystery - and so nobody restores the strict
   silence check. The positive assertion (a fresh-suffix key is delivered
   under its own headline) is unchanged.

2. seed_captured_procevent_result retired its source in the gap between the
   runner publishing its wake and releasing its claim, so retire read the
   exiting runner's ownership as uncertain and refused ("cannot confirm
   runner identity"). The fixture and retire path pre-date this branch; the
   confirm window returns reconcile closer to the moment of capture, which
   made the gap easier to hit. The fixture now waits, bounded, for the claim
   release the publish promises, with the reason at the wait.

Verified on this head with tasks-axi on PATH: fm-watch-triage 113/113 with
no skips, fm-procevent 106/106, fm-captain-hold-lifecycle 50/50,
fm-watch-arm 15/15, fm-bearings-board 18/18, fm-procevent-when 13/13,
fm-procevent-quota 18/18; bin/fm-lint.sh and bin/fm-doc-audience-check.sh
exit 0. First attempt, no retries.

* no-mistakes(document): docs: route stranded and launch-failed wakes in skill handling

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…gistration (kunchenguid#4191)

* fix(herdr): verify agent registrations at process level before trusting them

Herdr keeps a Pi registration (`agent get` -> agent=pi, agent_status=idle)
after the Pi process has exited to a plain shell whenever a nested interactive
shell sits under the pane's top shell, which is the crew shape `treehouse get`
leaves behind. The pane classifier trusted that registration alone, so
`fm-control.sh <id> relaunch`, `fm-spawn.sh --relaunch`, and the crew-state
recovery read all treated a shell-only pane as a live agent and refused
recovery for as long as the record lived.

The Herdr adapter now reads `pane process-info` plus the real process table
through a shared harness-process classifier (bin/fm-agent-process-lib.sh,
moved verbatim out of the tmux adapter so both backends mean the same thing by
agent, shell, and other) before a registered agent counts as live. A
registration over a shell-only pane is the new explicit `stale-agent` pane
state, which the recovery-grade read maps to `dead`; husk detection, reclaim,
presentation recovery, and session cleanup keep refusing it, so recovery reuses
the pane and nothing gains close authority. A working record is verified the
same way before the native busy verdict reports busy, so the recovery
classifier never reports a shell-only pane as working. An unreadable process
view reads unknown, trusting neither the registration nor its absence.

Reproduced and measured on Herdr 0.9.0 with Pi 0.85.1 in an isolated lab; the
new default-on live guard tests/fm-herdr-pi-stale-registration-live-e2e.test.sh
exercises the real stale record, tests/fm-control-herdr-smoke.test.sh proves
exit and relaunch through the control plane, and the portable suites pin the
classifier over real processes.

Fixes kunchenguid#4115. Duplicates: kunchenguid#3639, kunchenguid#3487, kunchenguid#2908, kunchenguid#3545.

* no-mistakes(review): settle transient prompt helpers before trusting herdr process state

* no-mistakes(review): drop stray codegraph file; read spaced comm whole in descendant walk

* no-mistakes(review): untrack stray .codegraph/.gitignore

* no-mistakes(review): untrack codegraph file; make spaced-path walk test discriminating

* no-mistakes(review): untrack stray .codegraph/.gitignore

* no-mistakes(review): untrack stray .codegraph/.gitignore re-added by fix round

* no-mistakes(review): untrack stray .codegraph/.gitignore

* no-mistakes(review): untrack codegraph file, drop dead control case, record process-info floor

* no-mistakes(review): refuse stale-agent on fresh herdr spawn preflight

Documented non-goal: fresh-spawn, reclaim, and presentation-recovery auto-recovery for a stale-agent pane is a separate design change, out of scope here, to be proposed upstream as its own issue if wanted.

* no-mistakes(test): Fix herdr flake: don't misread transient empty foreground as unreadable

* no-mistakes(document): Add fm-agent-process-lib.sh to scripts inventory

* no-mistakes(fix): update remote herdr fixture to the real pane process-info shape

The shared remote-secondmate herdr fixture still returned the old flat
process-info body ({"result":{"process":{"name":...}}}). The process-level
liveness classifier added for kunchenguid#4115 requires the real
{"result":{"type":"pane_process_info","process_info":{...foreground_processes}}}
shape and treated the old body as unreadable, so an already-launched remote
endpoint's agent-state read failed and any relaunch attempt against it died
with "remote endpoint state is unreadable; refusing duplicate launch"
instead of reaching the state it was actually exercising
(tests/fm-remote-secondmate-parent-binding.test.sh,
tests/fm-remote-secondmate-lifecycle-e2e.test.sh).

* no-mistakes(review): test: add empty-foreground regression test for herdr flake fix

* no-mistakes(document): docs: register new stale-registration live-e2e test in herdr entry points
* Bind GitHub merges to a live green head and require an away-task grant.

A GitHub merge now re-reads the pull request and passes
--match-head-commit, so a red or moved head cannot land the way GitLab
already refused. While an away record exists, only yolo or a named
grant may merge, so hold-for-return cannot ship an ungated PR.

Co-authored-by: Cursor <cursoragent@cursor.com>

* no-mistakes(review): Harden away merge authorization and grant parsing

* no-mistakes(review): Restrict fallback outcomes to proved GitHub merges

* no-mistakes(document): Refresh merge safety documentation

* no-mistakes(ci): Fixed all three CI failures by updating legacy GitHub merge fixtures for live-head verification/direct gh merges and removing a process-event runner cleanup race. Verified fm-pr-check-security, fm-captain-hold-lifecycle, and fm-watch-triage pass locally; shell syntax and git diff checks also pass

* no-mistakes(document): Document attended red-check exception

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
…m epoch (kunchenguid#4221)

The --claude guard's re-block budget charged the auto-arm ledger epoch, not
the re-block: `budget_account_current_epoch` advanced the session count only
when `state/.claude-autoarm-epoch` named a different generation than the
previous accounting. The epoch advances only inside the auto-arm hook's
generation claim, so a hook kept inert before that claim - a session lock
held by a live harness outside its ancestry, a hook that never fires, or an
identity or write failure ahead of `fm_autoarm_claim_next` - left the ledger
frozen at its last outcome and the count frozen with it. Reproduced in a
fixture: twelve consecutive Stops re-blocked with the count at 0 and the
attended fail-open never fired, leaving only Claude's silent 8-block
override, the blind end the bounded alarm exists to prevent.

The budget now charges a re-block against an epoch the previous re-block
already charged, while still charging each epoch at most once per Stop so
the wait loop's repeated observations of one fresh terminal outcome and the
same invocation's block decision cannot double count. The advancing-epoch
progression is unchanged: three re-blocks, then one attended fail-open for a
verified failure episode, and a frozen epoch now follows the same shape.
Budget exhaustion without a verified failure still blocks, by the existing
contract, and positive watcher recovery still clears the whole episode.

Regression coverage drives the real auto-arm hook against a foreign session
lock holder, asserts the ledger itself stays frozen, and fails before the fix
in both the verified and unverified shapes; the existing unverified budget
test now proves its budget actually ran out.
…enguid#4242)

* feat(herdr): attach a real foreground viewer so the live-client teardown cases can be driven

PR kunchenguid#4131 gated the Herdr active-tab close refusal on a live foreground client
instead of the persisted `.focused` pointer, but only its two detached
scenarios could be validated live. Every pseudo-terminal the runner built
started at a zero-sized window grid, so Herdr registered no foreground client
and `terminal title clear` kept answering `no_foreground_client`, leaving the
four attached-client scenarios untested. That was a harness limit, not a
product one.

Add `fm-herdr-lab.sh viewer start|stop <session>`, backed by
`bin/fm-herdr-lab-viewer.py`. The launcher sets the pty window size on the
master fd BEFORE the fork, so the TUI cannot read the grid until it is already
non-zero, and scrubs the inherited `HERDR_*` variables so Herdr's nested-viewer
refusal does not fire when the helper runs inside one of its own panes. Attach
and detach are both confirmed against the session's own foreground-client
reason rather than assumed from a signal.

The viewer inherits the lab's isolation contract: it attaches only to a session
carrying this lab's ownership tripwire, never to `default`, and it signals only
the processes it recorded, so a client someone else attached is never touched.
Teardown now refuses while an owned viewer is still attached.

Turn the reproduction into the regression with
`tests/fm-herdr-attached-viewer-live-e2e.test.sh`, which drives kunchenguid#4131's
scenarios 3, 4, 5, and 7 live against real Herdr and asserts the close refusal
fires. Scenarios 4 and 5 need a focus change at one exact product boundary, so
a PATH shim performs the real `tab focus` when the close helper issues its
planning `pane get`. Removing either half of the recipe from the launcher makes
the guard fail with the same `no_foreground_client` symptom kunchenguid#4131 reported.

* test(herdr): fail loudly when an attached-viewer fixture cannot be created

The fixture helpers run inside command substitutions, where fail() exits only
the subshell and leaves the script running with empty ids. Return non-zero
instead and carry the message at each call site.

* fix(herdr): stop the viewer launcher's kill timer from raising on an exited child

The SIGALRM escalation called os.kill unguarded, so a viewer that exited
during the grace window turned an ordinary shutdown into a traceback inside
the signal handler.

* docs: list the lab viewer's pty engine in the bin toolbelt

* no-mistakes(review): Harden Herdr viewer ownership and live CI coverage

* no-mistakes(review): Validate viewer startup timeout and process ownership

* no-mistakes(document): Document Herdr viewer safety contracts

* no-mistakes(review): Fix viewer timeout to two seconds

* no-mistakes(review): Cancel timed-out viewers and fix PTY grid

* no-mistakes(review): Serialize viewer transitions and verify process parentage

* no-mistakes(review): Harden viewer ownership locks and deduplicate CI

* no-mistakes(review): Release interrupted locks and preserve viewer escalation

* no-mistakes(review): Remove viewer locks and cancel interrupted launches

* no-mistakes(review): Close viewer launch signal races

* no-mistakes(document): Document attached Herdr viewer regression
…unchenguid#3766)

* fix(bootstrap): allow nonvisual work without Lavish

* no-mistakes(review): Gate scout brief Lavish line on bootstrap version floor
…kunchenguid#3825)

* fix(tests): isolate fixture Git configuration from host preferences

Ignore global and system Git configuration in the shared test library,
which all four fixture helper entry points source. Keep local config,
command-line overrides and explicitly supplied test config usable without
changing the caller's environment or real project signing preferences.

Exercise global and system signing inputs through all four helpers, real
fixture and child commits, explicit signing overrides, unchanged input
files, and signing refusal outside fixture subprocesses.

Verification evidence for issue kunchenguid#3770:
On pristine upstream f09de8a, all 12 reported suites failed and each logged
"No secret key" using a private GIT_CONFIG_GLOBAL containing
commit.gpgsign=true and gpg.format=openpgp, GIT_CONFIG_NOSYSTEM=1, and an
empty private GNUPGHOME (GIT_CONFIG_COUNT and GIT_CONFIG_PARAMETERS unset).
With this change, all 12 pass in the identical environment through
bin/fm-test-run.sh --per-script-timeout-secs 900:
fm-backlog-atomicity, fm-bootstrap-network-parallel, fm-bootstrap,
fm-crew-state, fm-fleet-sync, fm-gate-refuse, fm-grok-harness,
fm-session-start, fm-sessionstart-nudge, fm-tangle-guard, fm-test-run,
and fm-update (all tests/<name>.test.sh).
The new fm-test-fixtures regression failed before the library change and
passes after it. Canonical bin/fm-lint.sh passes.

Additional verification exposed fm-teardown's
herdr-preflight-missing-adapter assertion on both this branch and an
unchanged f09de8a archive with signing neutralized. That pre-existing
failure needs separate disposition; it is not repaired or skipped here.
The separately owned Muse and composer fixture defects remain untouched.

Fixes kunchenguid#3770

* no-mistakes(review): Complete fixture Git isolation and scope config assertions

* no-mistakes(review): Share Git isolation across standalone fixture entry points

* no-mistakes(review): Map git-config helper changes to lib.sh dependents

* no-mistakes(review): Select fixture-isolation regression on runner change; halve config matrix

* no-mistakes(review): Scope fixture-isolation regression selection to the runner alone

* no-mistakes(document): Give fixture Git isolation helper its owning header

* no-mistakes(document): Record fixture Git-isolation coverage in fixtures suite header

* no-mistakes(review): Fix linked-worktree fixtures and remove redundant Git isolation

* no-mistakes(document): Correct stale runner-selection documentation

* no-mistakes(document): Clarify family antecedent in isolation-proof runner evidence
… in auto mode (kunchenguid#4239)

* feat(spawn): add config/claude-permission-mode to launch Claude workers in auto mode

Every Claude worker launched with --dangerously-skip-permissions, and a
captain who refuses bypass mode had no way to select Claude Code's
classifier-reviewed auto mode instead. A new one-token local config,
config/claude-permission-mode, selects the permission flag for every
Claude launch: absent or `bypass` keeps today's launch byte-for-byte,
`auto` swaps in --permission-mode auto, and any other value refuses the
spawn before any endpoint, worktree, or record exists and names the
accepted values.

fm-spawn resolves the file on every spawn and relaunch, threads the flag
through the Claude launch template for crewmates, scouts, and secondmates
alike, and records claude_permission_mode=auto in the task meta only
under auto so the default meta stays unchanged; a relaunch re-resolves
rather than preserving the line. The file is a captain-wide safety
preference, so it joins the inherited local material pushed into
secondmate homes.

The Claude adapter reference records the verified auto launch shape on
Claude Code 2.1.269 and that it never meets the once-per-machine bypass
confirmation dialog; docs/configuration.md owns the schema.

* no-mistakes(review): drop unread claude_permission_mode meta line and its assertions
… untouched (kunchenguid#4243)

* fix(teardown): refuse to return a Treehouse pool slot reassigned to another task

A pool slot is reused across tasks, so a finished task's worktree= line can name
a slot a different, live task now holds. Teardown already refused when a second
task record named the same live path, but that scan cannot prove the record it
is tearing down is the current owner: the task that took the slot next may leave
no record the scan can reach - its own worker may have exited and its record been
cleaned up, or it may live in a home this machine does not register. Teardown
then killed every process under the path, hard-reset it and returned it, and its
unlanded-work refusal never fired because it was inspecting a directory that no
longer belonged to the task being torn down (observed 2026-09-07).

Treehouse's own state file cannot answer the ownership question. It records a
slot's owner as a live process lease (owner_pid plus owner_started_at, with
`treehouse status` reporting in-use from the processes actually running under
the path), which names no task and is released by the very event that makes a
record stale - the worker exiting. An unleased slot therefore reads identical
whether it is still this task's or has since been handed on, and a slot whose
new holder has also exited but left uncommitted work reads as free. So the
identity source is Firstmate's own claim, not Treehouse's lease.

fm-spawn writes that claim - the task id - into the slot at the moment it takes
it, under the same project lock that allocates the slot, and fm-teardown drops it
only after the slot is genuinely returned. It lives at <pool>/<slot>/.fm-slot-owner,
a sibling of the repo checkout rather than a file inside it, so claiming a slot
can never dirty the copy the landed-work checks inspect. A claim naming another
task, or one that cannot be read, refuses; --force does not lift either refusal,
because --force authorizes discarding this task's unlanded work, never another
task's live work. A slot that cannot be claimed refuses the spawn instead.

An absent claim proceeds on exactly the record-scan protection it had before:
slots taken before claims existed, and slots already returned, carry none, and
refusing those would strand every task in flight across this change on no
evidence at all.

The refusal is deliberately all-or-nothing rather than partially completing the
task's own cleanup. state/<id>.meta is the only durable record naming the
worktree and endpoint, so removing it would destroy the evidence needed to
reconcile which record is wrong, and its removal is one step with the backlog
transition. Nothing is stranded: clearing the stale worktree= line leaves a
record with no slot to release, which then tears down normally, and the refusal
names that remedy.

Repairing the previous claimant's stale worktree= line at spawn time is left for
separate work. It would have the new owner write another task's record - the same
class of cross-task mutation this bug is - and would need that record's own meta
lock; with the claim in place teardown refuses on evidence rather than depending
on the stale pointer having been scrubbed. For the same reason the relaunch path
writes no claim: it holds no allocation lock, and a record whose worktree= is
already stale would stamp the wrong task's claim onto a live sibling's slot.

The regression reproduces the reuse sequence with only one discoverable record,
including a clean, fully landed ship copy torn down without --force - the shape
of the real incident, which the previous code returned to the pool - and fails
against the previous code; the existing two-record, cross-home, own-slot
and no-claim cases still pass unchanged.

This builds ON upstream b028e8b (kunchenguid#3837), which is already in this branch's base
(origin/main 40c50ea) and owns the record-exclusivity scan. Nothing here
replaces that scan; the claim is the positive proof it cannot supply.

Claude-Session: https://claude.ai/code/session_01JTBmuqKugaPUj7k9TXQwFS

* no-mistakes(review): teardown leaves reassigned slot; spawn abort drops claim

* no-mistakes(review): narrow Treehouse lease evidence; gate abort claim release on lock

* no-mistakes(review): pin spawn-side slot claim; narrow abort-release header

* no-mistakes(document): docs: point slot-claim rationale at fm-wake-lib owner
…nguid#4247)

The reviewer treats Captain's intent as acceptance criteria, so a widened ask there drives over-built work; the spec should carry only what the ask requires.
* feat(bearings): name the Underway rows and order Charted Next newest filed first

The fleet board's Underway rows led with the run status alone, so a scan told
the captain where a pipeline stood but never which task the row was, and
Charted Next rendered in backlog order rather than by when work was filed.

The snapshot now projects the durable task name onto every in_flight row - from
this home's backlog title, and from a secondmate home's own ledger for an active
child - and the durable filed date onto every gate. The board's Underway row
leads with that name and keeps the run status on its second line, and Charted
Next renders newest filed first, with rows carrying no comparable date keeping
their payload order after every dated row.

The payload validator requires an explicit name marker on every Underway row and
refuses a filed value that is not an ISO date, so the board can never sort on
garbage or invent a label.

* no-mistakes(review): Fix Bearings labels, bounds, and filed validation

* no-mistakes(review): Fix Bearings identifiers and eligible queue bounds

* no-mistakes(document): Document Bearings labels and newest-first bounds

* no-mistakes(ci): Updated the stock macOS Bash CI expectation from 56 to 59 Bearings tests. Verified the suite under /bin/bash 3.2: all 59 tests pass. git diff --check also passes
…nchenguid#4248)

* fix(bearings): report the away-return catch-up instead of refusing

A captain returning from away and asking for bearings got zero bytes and an
error: fm-bearings-snapshot.sh ran the away-return guard with `|| exit $?`
before reading any fleet state, so the mere existence of the catch-up gate
killed every bearings mode (and /ahoy with them).

Bearings now consults that guard rather than obeying it. fm-afk-return.sh
separates its two refusal branches by exit status, so an ACTIVE away window
still refuses exactly as before - the right answer there is to run the return
first - while return catch-up (exit 4) lets collection and projection proceed
and is disclosed as one action-free `(return-catchup)` gate row, following the
existing `(main-inventory)` precedent. It stays out of decisions_open: these
blockers are firstmate-actionable, not the captain's own call, and the per-task
blockers already project as their own Underway rows.

The guard's refusal text also stops promising a blocker list it cannot produce:
a gate retained for a lifecycle reason alone now names that retention reason,
and bearings carries the same reason in the gate row's title.

Reporting is not ordinary work. AGENTS.md already scopes the return hold to
work rather than reporting, so only the /afk and bearings skills needed the
correction.

* no-mistakes(document): Refresh away-return Bearings verification

* no-mistakes(review): Reserve catch-up gate outside Bearings truncation

* no-mistakes(review): Preserve filed dates in catch-up gate output

* no-mistakes(document): Document reserved catch-up gate projection
…forked code-root copy (kunchenguid#4223)

* fix(backlog): address the home's backlog from any directory and detect a forked code-root copy

A home outside the code root forks its queue: the tracked .tasks.toml names
data/backlog.md relative to tasks-axi's working directory, so a bare
tasks-axi call from the code root writes the code root's data/ while session
start, spawn, and teardown use $FM_HOME/data. Linking the code-root copy into
the home does not hold, because tasks-axi 0.2.4 writes by renaming a temp file
over its target and rename(2) replaces a symlink: add, start, hold, and done
from the code root each turn the link back into a regular file. The archive
path is resolved against the working directory too, even with --file.

bin/fm-tasks-axi.sh runs tasks-axi against this home's backlog from any
directory, using the lifecycle transitions' existing addressing (run from the
data directory's parent, pin <data>/backlog.md through TASKS_AXI_FILE). It
keeps relative --to/--*-file arguments meaning the caller's paths, and refuses
a caller --file, an unresolvable home, and a symlinked home backlog. The
fm-send hold lookup, fm-public-followup, and the fm-decision-hold shim, which
relied on cwd discovery, now go through it with an explicit FM_HOME and a
cleared data override, so they keep addressing exactly $FM_HOME/data and an
ambient TASKS_AXI_FILE cannot divert them; every agent-facing backlog command
names it instead of bare tasks-axi.

Bootstrap gains a detect-only BACKLOG_RECONCILE check, also run read-only:
when the home's data directory is not the code root's, a code-root
data/backlog.md or data/done-archive.md that is not the home's own file is
reported as a fork, with the merge procedure in bootstrap-diagnostics.

* test(teardown): assert the completion hint names bin/fm-tasks-axi.sh ready

The completion hint now points at the home-addressed command instead of a
bare tasks-axi call, so the dependency-cleared follow-up assertion checks for
that command.

* no-mistakes(test): clear ambient tasks-axi env in tests/lib.sh

* no-mistakes(document): drop bare tasks-axi example from cd-guard doc

* no-mistakes(lint): replace ls -A decoy listing with find for SC2012

* no-mistakes: apply CI fixes

* revert: keep the compliance gate unchanged; the synchronize race is filed separately
* fix(spawn): pre-register Claude workspace trust for secondmate homes

A claude --secondmate launch skipped workspace-trust registration
entirely, so a standalone-clone secondmate home (an explicit
~/fm-homes/<id> path) had no store entry and its pane wedged on the
"Is this a project you trust?" dialog before it read its charter.
The step was gated on the task kind rather than on the harness, so the
spawn's fail-closed guard had nothing to run against and reported a
launch that could never start work.

fm-claude-trust.sh gains a secondmate-home mode. A secondmate home is a
whole firstmate instance, produced either as a leased worktree or as a
standalone clone, so the linked-worktree test cannot decide it and the
seed is the evidence instead: the .fm-secondmate-home marker must be a
regular file this user owns naming exactly the id being spawned, the
home must hold AGENTS.md and bin/, and each operational directory must
resolve inside the home. That is the set fm-home-seed.sh writes and
fm-spawn.sh's own home validation re-checks, so nothing wider than a
home a secondmate spawn would launch into can earn home-level trust.
The worktree path is unchanged, and still refuses a home.

fm-spawn.sh now runs the registration for every claude launch and keeps
refusing the spawn when it fails, rather than launching an agent that
would wedge.

* no-mistakes(document): Correct Claude secondmate trust guidance
* fix(pr-merge): judge each required check by its current run

When the base branch advances, GitHub cancels a pull request's in-flight
run and re-triggers it. The cancelled run stays in statusCheckRollup
beside the passing re-run, so the rollup can hold several runs of one
check name at the same head while GitHub itself reports the pull request
CLEAN. github_checks_not_green judged every run independently, so that
superseded failure refused a genuinely mergeable pull request and pushed
the operator toward a needless --allow-red.

Group the rollup by the reported name and judge each check by its
current run. Supersession is proven, never assumed: a name leaves the red
set only when every one of its non-green runs is strictly older than one
of its green runs, dated by the forge's own settled timestamp - a check
run's completedAt once its status is COMPLETED, or a status context's
createdAt - and only in the whole-second UTC form GitHub emits, which is
the one spelling that orders correctly as plain text. A run with no such
timestamp is never superseded, so a still-running, queued or undated run
keeps its check red, and a name with no green run at all stays red. An
unnamed entry is grouped alone so two unrelated unnamed checks are never
treated as one.

Every comparison is one-directional: it can only clear a failure a later
success provably replaced, and never clears a check whose current run
failed, is pending, or is missing. No other guard moves - the pull
request must still be open, undrafted, mergeable, conflict-free and
head-bound, and --allow-red still waives exactly its named check with
every other check green.

Live reproduction: PR kunchenguid#4224 read CLEAN with an old FAILURE and a newer
SUCCESS for one check name and was refused; it now verifies, while
kunchenguid#4208 and kunchenguid#4210, whose latest runs failed, still refuse.

* no-mistakes(review): Use check-run start times for safe supersession

* no-mistakes(document): Clarify GitHub check-rollup documentation
…guid#4266)

* fix(merge): persist the merge authority on poll-detected merge outcomes

The merge ledger tags a merge with the authority that permitted it while the
away-posture record existed, but only the direct attended merge in
bin/fm-pr-merge.sh recorded it. A merge the forge queued, or one the merge
poll detected after the fact, published an untagged row, so exactly the
merges no agent watched were the least auditable.

bin/fm-merge-authority-lib.sh now owns that answer, read from the same
structured sources the merge gate already used: the task's recorded yolo
posture and the away-posture record's mechanical grant list, never prose.
bin/fm-pr-merge.sh keeps its own refusal wording and gates on that answer;
bin/fm-watch.sh only records it on the row its poll publishes, so reading the
authority never becomes a second path to a merge. An unresolved answer records
an untagged row rather than dropping the outcome or inventing an authority.

* no-mistakes(review): Persist canonical merge authority for queued poll outcomes

* no-mistakes(review): Harden merge authority persistence against lifecycle races

* no-mistakes(review): Serialize poll authority publication with teardown

* no-mistakes(document): Clarify persisted merge authority lifecycle

* no-mistakes(ci): Added targeted SC2034 suppressions for the two public result assignments in bin/fm-merge-authority-lib.sh. Verified successfully with `CI=true bin/fm-lint.sh`
…4281)

The 2026-09-12 Actions starvation incident found firstmate CI with no
concurrency deduplication, so every superseded PR head kept its full
13-job fan-out, and four jobs with no timeout at all.

Add per-PR supersession keyed on the PR number for pull_request events
and on the unique run id for push events, cancelling only pull_request
runs, so a new PR head replaces its own in-flight CI while every main
push keeps its own group and is never cancelled. Add hang tripwires to
the four previously unbounded jobs: 25 minutes for lint (measured at
14-16 minutes) and 5 minutes each for the coverage guard, the timing
aggregate, and the repo invariants. Measured lane bounds are unchanged.

tests/fm-ci-workflow.test.sh resolves the workflow's concurrency
expressions against simulated pull_request and push contexts and holds
every job's finite timeout.
…henguid#4288)

Every other make_hold_home caller in this file skips when tasks-axi is
absent; this test was the one unguarded call, so hosts without tasks-axi
hard-fail the fixture build instead of skipping.
…nnot blind a session start (kunchenguid#4027)

* fix(bin): bound each backlog row read so one wedged backend cannot blind a session start

bin/fm-bootstrap.sh's reconcile and close-replay sweeps read the backlog
backend once per item through fm_backlog_row_show, and that read was
unbounded. A single wedged `tasks-axi show` therefore consumed the whole
FM_SESSION_START_TIMEOUT and truncated the digest before the wake queue,
supervision instructions, fleet state, and context sections ever printed,
leaving the fleet unsupervised with no live watcher. The harm was a blind
startup, not a slow one.

Bound the read with the existing shared timeout primitive
(bin/fm-timeout-lib.sh), so a wedged backend degrades to a loud partial
reconcile: the sweep's existing BACKLOG_RECONCILE diagnostic names the item
it could not read and the loop continues to the next one. The first bound hit
also latches FM_BACKLOG_ROW_SHOW_WEDGED, so a sweep over many items pays one
bound rather than one per item and still names every item it skipped, which is
what keeps the digest whole on a home carrying a large fleet.

The bound holds regardless of any particular tasks-axi install, so it does not
depend on the 0.2.5 `show` hang being resolved separately.

* fix(bin): set the wedged-backend latch where it survives, and prove it

The latch added with the read bound was inert. fm_backlog_row_show runs inside
a command substitution in both of its status-capturing callers, so the subshell
read the inherited value correctly but its write died with the subshell. Every
item still paid a full bound and reported `exceeded`, never `skipped`, which
left the large-fleet case the latch existed to cover completely uncovered.

Move the write to the two callers that capture the read's status and own the
surviving shell, and leave fm_backlog_row_show reading the latch only. Correct
the comments that claimed an ownership the function never had.

The test that was supposed to cover this asserted only that the second read
finished under a generous ceiling, which is true whether or not the latch
works. Assert instead that a latched read is strictly faster than one bound and
that it reports its own item as skipped, so an inert latch fails the test.

* test: cover every item the wedged-backend latch skips

The latch assertion exercised a single skipped item, so "every skipped item is
still named" was inferred rather than tested. Probe three items instead and
assert each skipped one names itself and costs less than a bound.

Verified as a real guard by removing both latch writes: the suite then fails on
the first skipped item instead of passing.

* no-mistakes(review): distinguish backlog read-bound hits from absent rows

* no-mistakes(review): preserve read-bound status through the captain verify gates

* no-mistakes(review): Preserve backlog read-bound hits through resolve_entry and reconcile instead of spending them as absent rows

* no-mistakes(review): Preserve backlog read-bound 124 through migrated-prefix scan and remaining task_show call sites

* no-mistakes(document): Document bounded backlog row reads and FM_BACKLOG_ROW_TIMEOUT_SECS

* no-mistakes(ci): Fixed all four failing CI checks with one root-cause fix plus one test-heredity fix. (1) bin/fm-captain-hold.sh: task_show carries the row in TASK_SHOW_OUTPUT and emits no stdout, but four call sites still used the stale command-substitution convention show=$(task_show ...), leaving show empty: task_show_or_fail (every captain hold failed with 'did not retain its hold-set stamp' - broke fm-captain-hold-lifecycle in parallel 1 and fm-bearings-board in serial 3), resolve_migrated_entry (migrated-prefix resolution could never match), reconcile-requests (existing rows were refused as absent), and command_open --identity (printed a constant '#0' identity, so fm-watch-triage's re-held captain call inherited the previous call's silence in serial 1). This is also the Greptile P1. Fixed by invoking task_show in the current shell and reading show=$TASK_SHOW_OUTPUT, the convention the other eight call sites already use; read-bound hits still stop loudly by name. (2) tests/fm-backlog-read-bound.test.sh (serial 4, unclassified family): the new e2e half implicitly relied on the author's process tree containing a harness process so fm-lock.sh would grant the fleet lock; on CI runners the lock is refused, the reconcile sweep is skipped, and the final BACKLOG_RECONCILE assertion fails. Reproduced by simulating a CI ancestry via a ps shim, fixed by pinning the lock evidence with the established fake-ps harness fixture pattern from tests/fm-session-start.test.sh. Verified: shellcheck clean; parallel-1, serial-3, and serial-4 lanes fully green locally (failed=0); serial-1 lane green except fm-gemini-harness, which fails only under local Node v26 (comm=node-MainThread); CI's default Node 22 reports comm=node, the branch that test passes on, so it is not a CI failure

* no-mistakes(document): Verified bounded backlog read docs accurate across branch
…guid#4285)

* fix(merge): serialize the away-authority check with a synchronous merge

bin/fm-pr-merge.sh read the away-posture record for merge authority (the
per-task merge grant and the yolo/away-grant decision) and handed the merge to
the forge afterwards. An archive at the captain's return or a grant revoked by
a replacement record could land in between, so a merge could proceed on away
authority that no longer held.

The away record now carries a cross-subsystem lock, built on the existing
bounded lock primitive rather than a new lock format: the record-mutating
subcommands hold it across their mutation, and the merge holds it across both
its authority read and the forge command. Because a queued or auto merge
returns before the pull request lands, and would therefore outlive the lock,
an away merge is now refused whenever it could land asynchronously: a
requested --auto, a base branch whose merge-queue state does not prove an
immediate merge, and GitLab's asynchronous flags and configuration. What
remains permitted while away is the synchronous merge that lands inside the
lock.

This closes the common away-record/merge race against a live lock owner. It
does not make the merge atomic in every case, and two narrow races are
accepted and documented at their sites rather than hidden, both
confused-agent-grade in the sense bin/fm-lease-lib.sh already uses:

- A merge-queue rule change or a PR base change in the window between the
  queue-free preflight and the forge call can still enqueue the merge, which
  can then land after its grant lapses.
- Killing the lock-owning shell while its gh or glab child is still running
  lets stale-owner recovery reclaim the lock and the record be archived or
  replaced, after which the orphaned child can complete the merge on lapsed
  authority.

Closing either one needs landing verification or an ownership handoff, which
is deliberately out of scope here.

No existing gate is relaxed. The lock is taken after the live green-at-head
verify and the captain-hold check, the in-lock authority read is unchanged,
and a lock that cannot be taken refuses the merge rather than proceeding
unlocked. The away grant stays a structured field; no prose is parsed.

* no-mistakes(review): Fix GitHub rollup fixture base branch

* no-mistakes(document): Document atomic away-authority merge locking

* no-mistakes(ci): Updated two executable GitHub API fixtures to include the required baseRefName. Both previously failing test suites now pass: fm-captain-hold-lifecycle.test.sh and fm-pr-check-security.test.sh. git diff --check also passes
…unchenguid#4200)

* feat(agy): verify Antigravity CLI as third worker/scout adapter

Detection by anchored ancestry in fm-harness.sh (no marker of its own);
bootstrap harness and effort validation; launch template with model and
effort mapping plus reachable-catalog model validation; rendered-tail
busy fallback in fm-busy-lib.sh with delivery footer in fm-composer-lib.sh;
control mechanics with crewmate/scout-only refusal; tmux liveness naming;
router entry with concise adapter reference; dated verification record;
portable regression plus opt-in live drift guard.

Verified live on agy 1.2.0: supervised spawn, durable steering,
same-copy relaunch, and exit, with Herdr-native busy agreement.

* no-mistakes(review): bound agy model probe, gate trust dialog, narrow busy signature

* no-mistakes(review): pre-register agy workspace trust, make readiness gate strict

* no-mistakes(review): Close Orca terminal on gate failure; isolate live-guard HOME; tighten agy matching

* no-mistakes(document): Document agy adapter in stale harness enumerations

* no-mistakes(review): Clamp non-positive FM_AGY_MODELS_TIMEOUT to the default bound

* no-mistakes(document): Fix stale test-shard snapshots after agy lane additions

* no-mistakes(ci): Fixed ci-3 (tests/fm-agy-harness.test.sh:519). Root cause: the agy spawn fixture's default base PATH (/usr/bin:/bin:/usr/sbin:/sbin) omits node's directory, but the spawn drives the real bin/fm-agy-trust.sh (which hard-requires node to record trust) and the fixture's fake tmux trust lookup (node -e) under that PATH. On the ubuntu-latest CI runner node lives in the toolcache (/usr/local/bin), so trust pre-registration failed on portable serial 2; on typical Arch hosts node is in /usr/bin, masking the defect. Fix (smallest, following the existing tests/fm-kimi-harness.test.sh precedent of carrying the interpreter's resolved directory): resolve node from the invoking environment (failing the test with 'test needs node' if absent, as kimi does for python3) and prepend its directory to the fixture's default base PATH; the FM_TEST_BASE_PATH override contract is untouched. Verified locally: (1) pre-fix reproduction with a CI-shaped base PATH (system bins minus node) produced exactly the reported failure — 'node is required to record workspace trust and was not found on PATH' plus the fake tmux 'node: command not found'; (2) post-fix, all 29 tests in the file pass both with node available only via a leading non-standard dir in the base PATH (CI's shape) and with the default base PATH on this host. bash -n clean; ShellCheck is not installed in this worktree (previously recorded as environmental)

* no-mistakes(test): Give agy typed sends a longer submit-confirm budget

* no-mistakes(document): Document agy send budget, trust gate, and control coverage

* no-mistakes(document): Document agy busy fallback inventory and send-timing evidence
…uid#4337)

* feat(afk): add quiet supervision mode for a present captain

Adds a first-class quiet supervision mode alongside /afk for
kunchenguid#2356: the same away-mode daemon, injection,
busy/composer guards, classification policy, and reliability
properties, but the captain staying present and chatting no longer
exits it - only an explicit /quiet off does.

state/.afk's first line now declares its mode (away, the default, or
quiet); fm_afk_mode() in bin/fm-wake-lib.sh is the single reader,
falling back to away for missing/empty/unreadable/unrecognized
content (including the legacy bare-epoch-timestamp format written
before mode existed) so nothing regresses. fm_afk_flag_write()
preserves the on-disk mode on a bare refresh (no explicit mode given)
rather than defaulting to away, which is what keeps the daemon's own
redundant terminal-side re-write from silently resetting a captain's
quiet mode back to away underneath them.

New .agents/skills/quiet/SKILL.md is a thin wrapper cross-referencing
/afk for every shared mechanism, per the one-owner rule. AGENTS.md
gains the state/.afk table entry and section 8's exit-trigger line.
bin/fm-supervision-instructions.sh, bin/fm-session-start.sh, and
bin/fm-guard.sh's stale-watcher banner all become mode-aware so a
quiet-mode captain is never misdirected to /afk in captain-facing
text.

Closes kunchenguid#2356

* no-mistakes(review): Fix AFK epoch parsing and quiet-mode digest wording for two-line flag

* no-mistakes(document): Fix turnend-guard.md daemon-ownership contract for quiet mode

---------

Co-authored-by: NewAiCoder <claude@theinbtw.com>
Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…kunchenguid#3578)

* fix(bin): let verified harness ancestry outrank retained markers (#3)

* fix(bin): let a structural harness ancestor outrank a retained marker

bin/fm-harness.sh treated a verified environment marker as unconditionally
authoritative, so a Codex session started from an environment that had retained
CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned
supervision protocol to a Codex primary, and every turn end was blocked for
missing Claude recovery.

The defect is the precedence boundary, not any one harness. codex, opencode,
kimi, and muse publish no identity marker at all, so with markers winning
outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering
was a point patch on the same class of problem, and the launch-time marker
clearing only ever covered sessions fm-spawn started.

Markers and ancestry are now separate evidence layers that detect_own arbitrates:

- no ancestry match, or no marker: the single available layer answers, unchanged;
- same harness family: the marker's finer verdict stands, so a launch-selected
  pi-signed is not flattened to pi by an ancestry walk that can only see the
  shared launcher name;
- different harness with a structural (command-name) ancestor: ancestry wins,
  because only ancestry proves who owns the process tree;
- different harness with only a bare-interpreter script-path match: the marker
  wins, since a harness-shaped path in some node process's arguments is weaker
  evidence than a harness publishing its own identity.

The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude
worker nested under cursor either.

Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so
a real harness process can be asked what the walk makes of it.

tests/fm-harness-precedence.test.sh is the portable regression, built from real
renamed processes with no harness installed. Every case drives the two layers
apart and asserts each alone as well as the combination, so no case can pass
vacuously; it also pins Codex's real two-process install topology, since the fix
depends on the native binary being what a tool subprocess meets first. The
opt-in drift guard gains the matching live half: each installed harness's real
running process must still be identified by the ancestry walk, and it fails
naming the harness and version when a release changes that name.

Documentation follows the corrected contract in the script header, the
harness-adapters detection section, the codex, opencode, kimi, and cursor
references, and a dated verification record.

* fix(tests): drop the unused argument pass-through in the shim-topology helper

bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and
forwarded "$@", but every call site that varies the environment or passes the
ancestry subcommand invokes the shim entry point directly, so the helper is only
ever called with no arguments (ShellCheck SC2120/SC2119).

Behavior is unchanged: with no arguments "$@" expanded to nothing.

* fix(bin): examine the top of the process chain instead of assuming init

harness_ancestry stopped as soon as the next pid was 1, on the assumption that
pid 1 is always init and can never be a harness.
Inside a PID namespace that assumption inverts: the harness itself is pid 1, so
the walk never examined the one process that proves who owns the tree, reported
no ancestry at all, and handed the verdict straight back to a retained marker.

A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and
CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and
rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry
precedence boundary in place.
The same probe now resolves codex and renders the Codex foreground checkpoint.

A host's real pid 1 (init, systemd, launchd) matches no harness name, so
examining it costs one ps call and can introduce no false positive; the walk
still stops once that top process has been read, and a non-numeric or zero ppid
still ends it.

tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that
reports every process as bash with ppid 1 and pid 1 as the harness.
The case asserts the marker still answers alone when pid 1 is host-shaped, so it
cannot pass vacuously, and it fails against the previous stop condition.

* docs(verification): record the real-Codex retained-marker evidence

The existing record proved the precedence boundary with the portable regression
and recorded each installed harness's process name behind the ancestry walk, but
it had no evidence from a real Codex process actually holding a retained Claude
marker, which is the failure the boundary exists for.

Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`,
with the exact command and the decisive verdict and rendered protocol on each
side, and records the second boundary that shape exposed: the walk must examine
the top of the process chain, because inside a PID namespace the harness is pid 1.
Refreshes the portable regression's observed output for the case it gained.

* no-mistakes(review): blind ancestry in marker-pinned harness tests

* no-mistakes(review): blind ancestry in the Pi guard-routing test

* no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims

* no-mistakes(review): model the spawn-and-wait Codex shim topology

* no-mistakes(document): correct stale muse marker-clearing detection claims

* no-mistakes: apply CI fixes

* fix(bin): examine the top of the chain in the lock and nudge walks too

The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two
other harness-ancestry walks, on the exact topology the branch verified against
a real Codex process.

bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next
pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could
not find that harness at all and did not recognize its own session lock.
bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a
lock pid of 1, so the same session was told to run session start again on every
turn.

Both walks now compare the top process before stopping, matching the shape used
in bin/fm-harness.sh.
For the lock walk this is safe because fm_harness_process_matches rejects a
host's real pid 1.
For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged
`kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent
rather than acting on init.

Each walk gains one regression case. The lock case drives a deterministic process
table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds
nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace,
because the builtin `kill -0` gate cannot be reached through a fake ps, and it
first proves the same fixture nudges with no lock present; it skips explicitly
where unprivileged namespaces are unavailable.

* no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard

* fix(bin): verify the live harness guard at the strength the guarantee needs

The marker-versus-ancestry boundary this branch ships is a strength claim:
detect_own hands an args-strength verdict straight back to a retained foreign
marker, so a harness is only protected where the ancestry walk reaches it at
comm strength.

The installed-harness drift guard probed the pane process alone. Under an
interpreter shim the pane process IS the shim, whose own script path is args
strength, while the native binary that carries comm strength is its child. The
guard therefore observed args for Codex, passed, and would have kept passing if
a release stopped spawning that native child at all, while real sessions
silently regressed to the original bug.

fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane
process and every descendant of it, the vantage a tool subprocess actually
occupies. The guard now requires comm strength somewhere in that set and
requires every vantage to name the same harness.

This supersedes the preceding commit's in-guard leaf walk, which reached the
same vantage but left the logic inside the test file, where CI could not pin it
and nothing else could reuse it. A harness-dependent check needs both halves:
`tests/fm-harness-precedence.test.sh` now carries a portable case proving the
subtree probe reaches a strength the top-of-session probe cannot, mutation
checked twice, once against the pre-change script and once by disabling
descendant enumeration. The subtree walk also avoids depending on tty and
process-group semantics that differ between Linux and macOS.

Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code
2.1.257 reports [comm claude].

* no-mistakes(review): narrow drift guard to the upward vantage path

* no-mistakes(review): judge only comm-strength vantages in drift guard

* no-mistakes(document): drop duplicated rationale in detection precedence evidence

* no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion

* no-mistakes(document): drop branch-relative phrasing in detection precedence evidence

* no-mistakes(review): guard remaining empty positional expansions in fm-harness

* no-mistakes(document): scope cursor marker-ordering claim to the marker layer

* no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties

* no-mistakes(document): Document comm-strength descent tie-break

---------

* no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests

* no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript

* no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs

* no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…rkers (kunchenguid#3944)

Claude Code's external-imports check (hasClaudeMdExternalIncludesApproved)
reads only the canonical git-root project entry in ~/.claude.json, which its
own worktree-to-primary-checkout canonicalization means is never the task
worktree fm-claude-trust.sh registered. The trust dialog kept working
previously only because its check has an ancestor-walk fallback that happens
to reach the worktree entry; the external-imports check has no such
fallback.

Verified by disassembling the installed claude binary and reproducing in an
isolated three-way tmux launch: identical flags registered only at the
worktree key still showed the external-imports dialog, and registering them
at the primary checkout key suppressed both dialogs.

fm-claude-trust.sh now registers all three flags on both the worktree entry
and the primary-checkout entry in one atomic write, and refuses when the
<project> argument is not itself a primary checkout (its own write target
would then be wrong). Extends the harness-adapters Claude reference and the
trust test suite.

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
…nguid#4355)

The marker lifecycle (fm-wake-lib.sh _fm_recovery_marker_ack) leaves
state/.watcher-down behind in an acked:* state after a downtime episode
is handled. health_snapshot's presence check reported that as an open
gap on every later return, so a handled episode kept surfacing as a
false GAP forever.
…kunchenguid#4361)

* fix(update): rebind fm-procevent-when watches after a self-update

A self-update fast-forwards bin/ in place, changing an armed watch's
action executable bytes with no tampering involved. The watch's trust
binding was hashed at arm time, so the very next fire was refused as
not matching the registered binding and the watch died silently.

Add fm-procevent-when.sh rebind-all: it re-hashes and republishes the
trust binding for every watch whose action executable lives under
FM_ROOT, using the same spec/trust validation as an ordinary fire, and
leaves any watch whose action lives outside FM_ROOT untouched. Wire it
into fm-update.sh right after a successful fast-forward, for both the
primary home and any local secondmate home that advances.

* no-mistakes(review): Canonicalize FM_ROOT for rebind-all's containment check

* no-mistakes(document): Document fm-update.sh's automatic watch rebind and its verification evidence

* no-mistakes(lint): fix(tests): double-quote printf scripts to satisfy shellcheck SC2016

* no-mistakes(review): Reload trust binding from disk before firing to reach live pollers

* no-mistakes(review): Lock the fire-time trust reload against rebind_one's publish race

* no-mistakes(document): Document rebind-all's self-update guarantee and its two review-round test rows

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
Bring the next 28 first-parent changes after 4768e98 into the fork
with full upstream parentage and the standing fork contracts preserved.
Reconcile shared launch, trust, process-state, merge, and quiet owners,
and retain the pinned boundary and parked-branch exclusions.
Apply the explicit sync landing decision relayed by Firstmate. Preserve all other check and authority gates, and pin scout golden capability input for CI.
@HelloWorldSungin
HelloWorldSungin merged commit e08ff9d into main Sep 14, 2026
16 of 17 checks passed
@HelloWorldSungin
HelloWorldSungin deleted the fm/fm-upstream-sync-2026-09-14-tail branch September 14, 2026 18:52
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.