Skip to content

feat(bin): sync fork with upstream delivery contract, supervision, and CI sharding - #2

Merged
kaku-san merged 6 commits into
mainfrom
fm/firstmate-upstream-sync
Aug 3, 2026
Merged

kaku-san merged 6 commits into
mainfrom
fm/firstmate-upstream-sync

Conversation

@kaku-san

@kaku-san kaku-san commented Aug 3, 2026

Copy link
Copy Markdown
Owner

Intent

Synchronize the kaku-san/firstmate fork with the current kunchenguid/firstmate default branch after fetching and verifying origin/main=2bd6d0c1ba81a88baf279461ca06f61a33221ca6, upstream/main=4ee4a0a2790cfaa5e47b30fa462f16546f2ab5b6, and merge base=cd73e75e02a1c1e74811b00c5ee08ffae8a59e1e. Preserve the fork's existing parallel-first dispatch requirement and intent while integrating portable CI sharding, corrected session-lock and attached supervision, hardened Claude supervision recovery, and explicit per-task delivery contracts. Reconcile AGENTS.md deliberately so these requirements compose under one-owner and size discipline and retain every unique safety boundary; review semantic interactions beyond textual conflicts without unrelated customization or misplaced evidence prose. Do not force-push, rewrite history, drop upstream safety boundaries, push a default branch, merge the PR, touch the primary local copy, or touch its untracked bun-baseline.zip. Validate the complete integrated branch with the documented test runner, bin/fm-lint.sh, bin/fm-doc-audience-check.sh, and focused contract tests. Push only fm/firstmate-upstream-sync to kaku-san/firstmate and open a green PR against that fork's default branch; captain retains merge approval.

What Changed

  • Merged four upstream commits into the fork. bin/fm-spawn.sh, bin/fm-brief.sh, and bin/fm-promote.sh now require an explicit --mode/--yolo delivery contract on every ship task and refuse a spawn whose mode disagrees with the brief's recorded contract; bin/fm-session-lock-lib.sh, bin/fm-watch-arm.sh, and bin/fm-turnend-guard.sh correct session-lock ancestry and attached-watcher supervision; bin/fm-claude-stop-autoarm.sh gains failure-episode and attended-alarm records that harden auto-arm recovery. New coverage lands in tests/fm-task-delivery.test.sh, tests/fm-session-lock-ancestry.test.sh, and tests/fm-watch-arm.test.sh.
  • .github/workflows/ci.yml runs the portable serial lane as a 4-way matrix with fail-fast: false, deriving FM_SERIAL_LANE from strategy.job-total so a resized matrix is refused rather than silently leaving a shard unrun; bin/fm-test-run.sh owns shard membership, rejects an ofN that disagrees with it, and the coverage guard now proves the shards partition the serial lane exactly. The lane timeout drops from 20 to 15 minutes.
  • AGENTS.md was reconciled by hand to keep the fork's parallel-first dispatch clause intact alongside upstream's new intake rules for resolving delivery mode and yolo posture. README.md, docs/architecture.md, and the skill and contributor docs move from per-project "project modes" to per-task "delivery modes", and docs/fm-test-portable-shards.md records that the lane now holds 72 scripts against 69 weight hints, so the three newly added tests run on the conservative default weight until the hints are refreshed from a green CI run.

Risk Assessment

✅ Low: The merge is upstream-verbatim outside AGENTS.md (git diff 4ee4a0a..7bc4c0c touches only that file), the fork's parallel-first clause and its strengthened serialization wording are preserved intact, upstream's delivery-contract paragraph composes with them under disjoint ownership, both PreToolUse seatbelts and every upstream safety boundary remain unchanged, and the only issues found are a stale balance table and a message-less fail-closed exit that self-heals on the next Stop.

Testing

I exercised the four integrated upstream behaviors and the preserved fork requirement against the complete merged branch with the documented runner, bin/fm-test-run.sh. Fifteen targeted scripts covering the delivery contract, session-lock ancestry, Claude auto-arm recovery, portable CI sharding, attached watcher supervision, turn-end guard, spawn dispatch, and AGENTS.md doc discipline ran green except one case in tests/fm-watcher-lock.test.sh, which failed once under concurrent load and then passed four consecutive isolated reruns; that file is byte-identical to upstream/main, so the flake is pre-existing upstream rather than merge-induced, and it is reported as a warning for the captain. Beyond the automated suite I captured product-level evidence a reviewer can read directly: a CLI transcript of fm-spawn.sh refusing all four malformed delivery contracts with no task metadata written, a shard transcript proving the four portable serial shards exactly partition the 72-script lane and that a mismatched shard count is refused, a merge-reconciliation transcript showing AGENTS.md is the only file differing from upstream/main with all four fork lines intact and line accounting exactly additive, and a rendered HTML page plus screenshot of the merged section 7 intake block colour-coded by provenance so the fork's parallel-first clause and upstream's delivery-contract block can be seen composing under one owner. The UI-facing surface here is agent-instruction prose rather than an app screen, so the rendered-HTML screenshot is the reviewer-visible artifact for it. Two validations named in the intent, bin/fm-lint.sh and bin/fm-doc-audience-check.sh, were deliberately not run: this phase is barred from linters and static analysis, and the Lint phase owns them. During cleanup I killed two orphaned watcher processes the flaky case left pointing at this worktree (they were holding the runner's stdout pipe and stalling it) and confirmed the worktree is clean with no stray files or processes.

  • Evidence: Merged AGENTS.md section 7 intake, colour-coded by provenance (fork vs upstream vs base) (local file: /var/folders/hv/gd626zx51jb1_3rghgd97mp00000gn/T/no-mistakes-evidence/01KZ3ATDAK72A7T6ET2E48CYFM/agents-md-reconciled.png)
Evidence: Same reconciliation view as rendered HTML
<!doctype html><meta charset=utf-8><title>AGENTS.md section 7 intake - reconciled</title>
<style>
body{background:#0d1117;color:#e6edf3;font:15px/1.65 -apple-system,BlinkMacSystemFont,"Segoe UI",sans-serif;margin:0;padding:32px 40px}
h1{font-size:20px;margin:0 0 4px} h2{font-size:13px;font-weight:500;color:#8b949e;margin:0 0 24px}
ul.legend{list-style:none;padding:0;margin:0 0 28px;display:flex;flex-wrap:wrap;gap:8px 24px;font-size:12.5px;color:#c9d1d9}
ul.legend li{display:flex;align-items:center;gap:8px}
.dot{width:10px;height:10px;border-radius:3px;display:inline-block}
.doc{background:#161b22;border:1px solid #30363d;border-radius:10px;padding:22px 26px;max-width:1180px}
p.ln{margin:0 0 12px;padding-left:132px;text-indent:-132px}
.tag{display:inline-block;width:112px;margin-right:20px;padding:1px 0;border-radius:4px;color:#fff;
     font:600 10px/16px ui-monospace,SFMono-Regular,monospace;text-align:center;text-indent:0;vertical-align:2px}
p.base{color:#8b949e} p.fork{color:#e6edf3} p.upstream{color:#e6edf3} p.both{color:#e6edf3}
p.fork,p.upstream{background:#1c2128;border-left:3px solid;margin-left:-14px;padding:8px 12px 8px 146px;border-radius:0 6px 6px 0;text-indent:-132px}
p.fork{border-color:#1f6feb} p.upstream{border-color:#238636}
.blank{height:10px} code{background:#30363d;border-radius:4px;padding:1px 5px;font:12.5px ui-monospace,monospace}
footer{margin-top:22px;color:#8b949e;font-size:12px;max-width:1180px}
</style>
<h1>AGENTS.md &rsaquo; section 7 &rsaquo; task intake, as merged on fm/firstmate-upstream-sync</h1>
<h2>merge 7bc4c0c &nbsp;=&nbsp; fork 2bd6d0c &nbsp;+&nbsp; upstream 4ee4a0a &nbsp;over base cd73e75</h2>
<ul class=legend><li><span class="dot" style="background:#3a3f4b"></span>unchanged from merge base cd73e75</li><li><span class="dot" style="background:#1f6feb"></span>fork-only (kaku-san 2bd6d0c) - parallel-first dispatch</li><li><span class="dot" style="background:#238636"></span>upstream-only (kunchenguid 4ee4a0a) - per-task delivery contract</li><li><span class="dot" style="background:#8957e5"></span>present on both sides</li></ul>
<div class=doc><p class="ln base"><span class="tag" style="background:#3a3f4b">base</span>Before commissioning an investigation, consult existing reports and established evidence.</p><p class="ln base"><span class="tag" style="background:#3a3f4b">base</span>Classify the deliverable:</p><div class="blank"></div><p class="ln base"><span class="tag" style="background:#3a3f4b">base</span>- **Ship** is the default and produces a project change through the selected delivery mode; once implementation is authorized, dispatch a ship and keep any remaining bounded research inside it unless unresolved uncertainty could materially change whether or what to build.</p><p class="ln base"><span class="tag" style="background:#3a3f4b">base</span>- **Scout** produces knowledge in <code>data/&lt;id&gt;/report.md</code>, never a PR, and is appropriate for investigation, diagnosis, planning, reproduction, or audit work when the captain explicitly requests a separate knowledge or design deliverable or unresolved uncertainty could materially change whether or what to build.</p><div class="blank"></div><p class="ln base"><span class="tag" style="background:#3a3f4b">base</span>If established evidence already answers an informational question, relay it without a design-only scout; when implementation intent is unclear, answer and ask one concise implementation question when useful rather than dispatching speculative design work.</p><p class="ln base"><span class="tag" style="background:#3a3f4b">base</span>Never both present a likely-enough solution and launch a parallel design exercise that is not expected to change it.</p><p class="ln base"><span class="tag" style="background:#3a3f4b">base</span>A diagnostic request, report, recommendation, or implementation-ready finding is evidence, not authorization to change code.</p><p class="ln base"><span class="tag" style="background:#3a3f4b">base</span>Load <code>diagnostic-reasoning</code> before scoping a reported bug and before acting on a diagnostic report.</p><div class="blank"></div><p class="ln upstream"><span class="tag" style="background:#238636">upstream</span>Resolve every ship task&#x27;s concrete delivery mode and yolo posture at intake, and pass both explicitly to the brief, the spawn, and any scout promotion, which all refuse to guess.</p><p class="ln upstream"><span class="tag" style="background:#238636">upstream</span>A current explicit captain instruction wins; otherwise the project&#x27;s registry entry is the captain&#x27;s standing posture, and dropping below its rigor needs a reason you can state.</p><p class="ln upstream"><span class="tag" style="background:#238636">upstream</span>On a <code>no-mistakes-prod-only</code> project, classify the task&#x27;s surface: internal-only tooling, automation, contributor or operator process, and release or submission work ships <code>direct-PR</code>, while product-facing, mixed, and uncertain work ships <code>no-mistakes</code>; never infer internal-only from file location or project name.</p><p class="ln upstream"><span class="tag" style="background:#238636">upstream</span>An unregistered project or absent registry resolves to <code>no-mistakes</code> with yolo off, and the registration gap goes to the captain.</p><p class="ln upstream"><span class="tag" style="background:#238636">upstream</span>Record the resulting mode, yolo, and the one-line reason for any deviation in the backlog item note.</p><div class="blank"></div><p class="ln fork"><span class="tag" style="background:#1f6feb">fork</span>At every intake, and whenever long validation, infrastructure or platform work, an external wait, or a blocker appears, firstmate must identify independently valuable user-facing proof or delivery paths, dispatch immediately every such path already authorized by the captain&#x27;s original request or the accepted task criteria, and keep parallel-first decomposition to bounded independent outcomes rather than redundant planners, duplicate implementations, or competing product approaches or architectures.</p><p class="ln fork"><span class="tag" style="background:#1f6feb">fork</span>The Selected delivery path and approval authority subsection exclusively owns standing <code>yolo</code> authority, and this parallel-first clause does not broaden it.</p><p class="ln fork"><span class="tag" style="background:#1f6feb">fork</span>Identification may be silent and may find no material path, and identifying a path is never authorization to implement it, so raise an unauthorized path for a decision only when it is independently valuable to the requested outcome and could materially improve delivery or avoid meaningful delay or failure, never as speculative adjacent work.</p><p class="ln base"><span class="tag" style="background:#3a3f4b">base</span>Treat file or subsystem overlap as a risk signal rather than an automatic reason to wait, and dispatch isolated work immediately with no concurrency cap when each change can be independently implemented and validated and the selected delivery path can reconcile ordinary rebases or conflicts.</p><p class="ln fork"><span class="tag" style="background:#1f6feb">fork</span>Serialize only after naming a true semantic dependency, shared mutable external state, incompatible concurrent migration, or another concrete condition that makes independent progress or reconciliation unsafe; same-file editing, preference among bounded paths that all implement the same authorized outcome, or the existence of a longer integrated path is not itself a dependency, and genuine blockers remain durable.</p><p class="ln base"><span class="tag" style="background:#3a3f4b">base</span>Write the task-specific brief under section 11 before spawning.</p><div class="blank"></div></div>
<footer>Both edits landed in the same paragraph run. The fork's parallel-first clause survives verbatim, keeps its
own strengthened <em>Serialize only after naming&hellip;</em> sentence, and defers <code>yolo</code> authority to the
Selected delivery path subsection, so the upstream delivery-contract block that resolves each task's mode and yolo
posture composes under one owner rather than competing with it.</footer>
Evidence: Per-task delivery contract: fm-spawn.sh CLI refusals

$ fm-spawn.sh demo-1 <project> claude # contract never decided error: ship spawns require --mode <no-mistakes|direct-PR|local-only>; resolve it at intake from the captain's instruction and the project's registered posture in data/projects.md exit=1 $ fm-spawn.sh demo-1 <project> claude --mode fast-path --yolo off # mode outside the closed set error: --mode must be one of no-mistakes, direct-PR, local-only (got 'fast-path') exit=1 $ fm-spawn.sh demo-1 <project> claude --mode no-mistakes # yolo posture omitted error: ship spawns require --yolo <on|off>; it is this task's routine approval authority, not a project lookup exit=1 $ fm-spawn.sh demo-1 <project> claude --mode direct-PR --yolo off # brief records mode=no-mistakes error: delivery mismatch for demo-1: the brief says mode=no-mistakes but this spawn passed --mode direct-PR; correct the flag or re-scaffold the brief so the worker's instructions and the task record agree exit=1 # task metadata written by the refused spawns (must be none): state/: []

=== Explicit per-task delivery contract: fm-spawn.sh refuses to guess ===

$ fm-spawn.sh demo-1 <project> claude                                   # contract never decided
error: ship spawns require --mode <no-mistakes|direct-PR|local-only>; resolve it at intake from the captain's instruction and the project's registered posture in data/projects.md
exit=1

$ fm-spawn.sh demo-1 <project> claude --mode fast-path --yolo off       # mode outside the closed set
error: --mode must be one of no-mistakes, direct-PR, local-only (got 'fast-path')
exit=1

$ fm-spawn.sh demo-1 <project> claude --mode no-mistakes                # yolo posture omitted
error: ship spawns require --yolo <on|off>; it is this task's routine approval authority, not a project lookup
exit=1

$ fm-spawn.sh demo-1 <project> claude --mode direct-PR --yolo off       # brief records mode=no-mistakes
error: delivery mismatch for demo-1: the brief says mode=no-mistakes but this spawn passed --mode direct-PR; correct the flag or re-scaffold the brief so the worker's instructions and the task record agree
exit=1

# task metadata written by the refused spawns (must be none):
  state/: []
Evidence: Portable CI sharding: 4 shards exactly partition the serial lane

$ bin/fm-test-run.sh --check-coverage FM_TEST_COVERAGE ok total=107 parallel=24 serial=72 serial_shards=4 herdr=11 shard 1of4 -> 17 scripts shard 2of4 -> 18 shard 3of4 -> 16 shard 4of4 -> 21 whole lane: 72 union of shards: 72 duplicates in union: 0 PARTITION EXACT: union == whole lane, no missing, no duplicates $ bin/fm-test-run.sh --list --lane portable-serial-1of3 fm-test-run: lane 'portable-serial-1of3' asks for 3 portable serial shards but this runner is configured for 4 (see --list-lanes) exit=2

=== Portable CI serial sharding: 4 shards exactly partition the serial lane ===

$ bin/fm-test-run.sh --list-lanes
portable-parallel-1
portable-parallel-2
portable-serial
portable-serial-1of4
portable-serial-2of4
portable-serial-3of4
portable-serial-4of4
real-herdr-gated

$ bin/fm-test-run.sh --check-coverage
FM_TEST_COVERAGE ok total=107 parallel=24 serial=72 serial_shards=4 herdr=11

$ bin/fm-test-run.sh --list --lane portable-serial-1of4 | wc -l  -> 17 scripts
$ bin/fm-test-run.sh --list --lane portable-serial-2of4 | wc -l  -> 18 scripts
$ bin/fm-test-run.sh --list --lane portable-serial-3of4 | wc -l  -> 16 scripts
$ bin/fm-test-run.sh --list --lane portable-serial-4of4 | wc -l  -> 21 scripts

# union of the 4 shards vs the whole portable-serial lane:
  whole lane: 72   union of shards: 72   duplicates in union: 0
  PARTITION EXACT: union == whole lane, no missing, no duplicates

# a shard count that disagrees with the runner is refused:
$ bin/fm-test-run.sh --list --lane portable-serial-1of3
fm-test-run: lane 'portable-serial-1of3' asks for 3 portable serial shards but this runner is configured for 4 (see --list-lanes)
exit=2
Evidence: Merge reconciliation proof: parents, upstream-boundary sweep, fork-line containment, line accounting

actual merge parents of HEAD: 2bd6d0c1ba81a88baf279461ca06f61a33221ca6 4ee4a0a2790cfaa5e47b30fa462f16546f2ab5b6 # 1. Of the 56 files upstream changed, only AGENTS.md differs from upstream in the merge: differs from upstream: AGENTS.md total files differing from upstream/main: 1 (bin/, tests/, docs/, .github/ all byte-identical to upstream) # 3. The fork-only commit 2bd6d0c is fully contained in the merge (no fork line dropped): PRESENT At every intake, and whenever long validation, infrastructure or platform work, an external wait... PRESENT The Selected delivery path and approval authority subsection exclusively owns standing yolo au... PRESENT Identification may be silent and may find no material path, and identifying a path is never auth... PRESENT Serialize only after naming a true semantic dependency, shared mutable external state, incompati... # 4. Line accounting - the merge is exactly additive over the merge base: merge-base (cd73e75): AGENTS.md 60449 bytes, 536 lines fork-main (2bd6d0c): AGENTS.md 61593 bytes, 539 lines upstream-main (4ee4a0a): AGENTS.md 61576 bytes, 542 lines merged (7bc4c0c): AGENTS.md 62720 bytes, 545 lines 536 (base) + 3 (fork-only net) + 6 (upstream-only net) = 545 merged lines

=== Merge reconciliation proof: fork intent preserved, upstream boundaries intact ===

$ git rev-parse origin/main upstream/main merge-base  (per the sync brief)
  origin/main   = 2bd6d0c1ba81a88baf279461ca06f61a33221ca6
  upstream/main = 4ee4a0a2790cfaa5e47b30fa462f16546f2ab5b6
  merge-base    = cd73e75e02a1c1e74811b00c5ee08ffae8a59e1e
  actual merge parents of HEAD: 2bd6d0c1ba81a88baf279461ca06f61a33221ca6 4ee4a0a2790cfaa5e47b30fa462f16546f2ab5b6

# 1. Of the 56 files upstream changed, only AGENTS.md differs from upstream in the merge:
     differs from upstream: AGENTS.md
     total files differing from upstream/main: 1  (bin/, tests/, docs/, .github/ all byte-identical to upstream)

# 2. The only difference vs upstream is the fork parallel-first block:
     diff --git a/AGENTS.md b/AGENTS.md
     index 2adf742..f830b5f 100644
     --- a/AGENTS.md
     +++ b/AGENTS.md
     @@ -264,8 +264,11 @@ On a `no-mistakes-prod-only` project, classify the task's surface: internal-only
      An unregistered project or absent registry resolves to `no-mistakes` with yolo off, and the registration gap goes to the captain.
      Record the resulting mode, yolo, and the one-line reason for any deviation in the backlog item note.
      
     +At every intake, and whenever long validation, infrastructure or platform work, an external wait, or a blocker appears, firstmate must identify independently valuable user-facing proof or delivery paths, dispatch immediately every such path already authorized by the captain's original request or the accepted task criteria, and keep parallel-first decomposition to bounded independent outcomes rather than redundant planners, duplicate implementations, or competing product approaches or architectures.
     +The Selected delivery path and approval authority subsection exclusively owns standing `yolo` authority, and this parallel-first clause does not broaden it.
     +Identification may be silent and may find no material path, and identifying a path is never authorization to implement it, so raise an unauthorized path for a decision only when it is independently valuable to the requested outcome and could materially improve delivery or avoid meaningful delay or failure, never as speculative adjacent work.
      Treat file or subsystem overlap as a risk signal rather than an automatic reason to wait, and dispatch isolated work immediately with no concurrency cap when each change can be independently implemented and validated and the selected delivery path can reconcile ordinary rebases or conflicts.
     -Serialize only for a true semantic dependency, shared mutable external state, incompatible concurrent migration, or another concrete condition that makes independent progress or reconciliation unsafe; same-file editing alone is insufficient, and genuine blockers remain durable.
     +Serialize only after naming a true semantic dependency, shared mutable external state, incompatible concurrent migration, or another concrete condition that makes independent progress or reconciliation unsafe; same-file editing, preference among bounded paths that all implement the same authorized outcome, or the existence of a longer integrated path is not itself a dependency, and genuine blockers remain durable.
      Write the task-specific brief under section 11 before spawning.
      
      ### Dispatch and supervision handoff

# 3. The fork-only commit 2bd6d0c is fully contained in the merge (no fork line dropped):
     PRESENT  At every intake, and whenever long validation, infrastructure or platform work, an external wait...
     PRESENT  The Selected delivery path and approval authority subsection exclusively owns standing `yolo` au...
     PRESENT  Identification may be silent and may find no material path, and identifying a path is never auth...
     PRESENT  Serialize only after naming a true semantic dependency, shared mutable external state, incompati...

# 4. Line accounting - the merge is exactly additive over the merge base:
     536 (base) + 3 (fork-only net) + 6 (upstream-only net) = 545 merged lines
Evidence: Targeted behavior-test run log (15 scripts)
FM_TEST_BEGIN 2026-08-03T09:45:46Z tests/fm-task-delivery.test.sh family=pure-contract-unit expected_gate_skip=none
ok - fm-spawn: a ship spawn requires a valid explicit mode and yolo before anything is created
ok - fm-spawn: scout and secondmate spawns refuse ship delivery flags
ok - fm-spawn: the brief's recorded mode and the spawn's explicit mode must agree
ok - fm-spawn: a rigor downgrade against the registered posture is announced, never blocked
ok - fm-spawn: a scout spawn resolves no delivery posture from the registry
ok - fm-promote: promotion requires the delivery contract and records it exactly once
ok - fm-project-mode: the conditional policy is accepted, mapped for mechanical callers, and readable raw
# all fm-task-delivery tests passed
FM_TEST_END 2026-08-03T09:46:00Z tests/fm-task-delivery.test.sh exit=0 duration_ms=13279 gate_skip=false
FM_TEST_BEGIN 2026-08-03T09:46:00Z tests/fm-session-lock-ancestry.test.sh family=watcher-wake-lock expected_gate_skip=none
ok - session-lock: a version-named Claude Code session is identified from its install path and argv[0]
ok - session-lock: ordinary script paths under a harness directory are not harness processes
ok - session-lock: ownership stops at the first non-harness gap above the contiguous run
ok - session-lock: a live version-named session holding the lock is not mistaken for a stale owner
ok - session-lock e2e: a version-named session claims the home and arms supervision
ok - session-lock e2e: a session parented by a harness-named daemon claims the home and arms supervision
ok - session-lock e2e: a version-named session under a harness-named daemon keeps its own lock
FM_TEST_END 2026-08-03T09:46:12Z tests/fm-session-lock-ancestry.test.sh exit=0 duration_ms=11966 gate_skip=false
FM_TEST_BEGIN 2026-08-03T09:46:12Z tests/fm-claude-stop-autoarm.test.sh family=unclassified expected_gate_skip=none
ok - auto-arm: inert in a linked child worktree even when in-flight
ok - auto-arm: inert with no session lock
ok - auto-arm: a demonstrably dead recorded session owner is reclaimed through fm-lock.sh before arming
ok - auto-arm: inert without arm, rewake, or lock replacement when another live harness owns the home
ok - auto-arm: inert while AFK owns supervision
ok - auto-arm: stale-owner recovery leaves the AFK and supervision-need gates unchanged
ok - auto-arm: resolves the outermost pid of a nested contiguous claude ancestry (bg-spare chain)
ok - auto-arm: inert with nothing in flight and no X-mode need
ok - auto-arm: actionable close translates to exactly one exit-2 rewake with reason
ok - auto-arm: actionable close survives a healthy successor without duplicate delivery
ok - auto-arm: bounded failure verification emits one automatic-mechanism alarm
ok - auto-arm: consecutive failures keep Stop-owned retry without repeating notice
ok - auto-arm: unverified clean close exhausts retries and fails closed
ok - auto-arm: post-alarm actionable outcomes cannot continue or reset failure state
ok - auto-arm: benign cycle end with a live watcher and fresh beacon stays silent across the next cycle
ok - auto-arm: budget contention preserves the episode and forces a reset retry
ok - auto-arm: X-mode poll need arms the cycle even with no tasks in flight
ok - auto-arm: concurrent firings admit one owner and one rewake translation
ok - auto-arm: need vanishing mid-cycle closes without a rewake
ok - auto-arm: mid-cycle AFK hands triage to the daemon with no rewake
ok - auto-arm: active in a marked secondmate home
ok - fm-lock: shared session-lock lib preserves the status path
FM_TEST_END 2026-08-03T09:47:17Z tests/fm-claude-stop-autoarm.test.sh exit=0 duration_ms=64123 gate_skip=false
FM_TEST_BEGIN 2026-08-03T09:47:17Z tests/fm-test-run.test.sh family=pure-contract-unit expected_gate_skip=none
ok - exact suite coverage: --all lists every tests/*.test.sh once
ok - family selection returns a proper subset of the suite
ok - single-script selection lists exactly that path
ok - changed-file selection stays conservative (never silent full suite)
ok - changed selection covers dependents and fails closed for unmapped source
ok - empty changed selection emits deterministic text and JSON summaries
ok - timing markers and JSON artifact are valid
ok - aggregate exit reflects any script failure
ok - gate-skip accounting is honest and non-failing
ok - fail-on-gate-skip converts herdr-not-found into a hard failure
ok - exclude-family drops the named primary family after selection
ok - portable shard union, disjointness, and coverage guard hold
ok - portable serial shards are a deterministic disjoint cover of the serial lane
ok - portable serial shard lanes refuse mismatched, out-of-range, and countless names
ok - --jobs refuses non-proven / stateful selections
ok - jobs scheduler runs proven scripts; failure propagates; non-proven refused
ok - aggregate-json merges lane timing artifacts
FM_TEST_END 2026-08-03T09:49:21Z tests/fm-test-run.test.sh exit=0 duration_ms=124625 gate_skip=false
FM_TEST_BEGIN 2026-08-03T09:49:22Z tests/fm-watch-arm.test.sh family=watcher-wake-lock expected_gate_skip=none
ok - watch-arm: an attached arm reports the wake its cycle delivered instead of a false failure
ok - watch-arm: a delivered wake consumed by the handling turn still closes the attached arm cleanly
ok - watch-arm: a cycle that delivered no wake of its own still fails loudly
FM_TEST_END 2026-08-03T09:50:17Z tests/fm-watch-arm.test.sh exit=0 duration_ms=55361 gate_skip=false
FM_TEST_BEGIN 2026-08-03T09:50:17Z tests/fm-watcher-lock.test.sh family=watcher-wake-lock expected_gate_skip=none
ok - simultaneous watcher starts leave exactly one live process
ok - fm_pid_identity real ps fallback is locale-invariant
ok - fm_pid_identity is locale-invariant across LC_ALL/LC_TIME
ok - /proc process identity ignores simulated btime changes
ok - /proc process identity detects pid reuse
ok - MSYS /proc process identity regression skipped on non-Windows host
ok - killed watcher stale lock is reclaimed
ok - live watcher lock with stale heartbeat is actionable
ok - guard banner leads when down with pending wakes (repair-after-drain) and stays silent when live and fresh
ok - concurrent fm_lock_try_acquire yields exactly one winner
ok - dead-pid stale lock is reclaimed by a single acquirer
ok - concurrent stale-lock steal yields exactly one winner
ok - live steal mutex is not reclaimed
ok - live-held lock is not stolen
ok - empty mid-acquire lock keeps a minimum grace
ok - late original claimant cannot claim a recreated lock
ok - paused mid-acquire claimant backs off to active stealer
ok - watch restart refuses to signal a reused pid
not ok - restart did not attach to the verified healthy peer: watcher: started pid=78004 (beacon fresh)
FM_TEST_END 2026-08-03T10:01:02Z tests/fm-watcher-lock.test.sh exit=1 duration_ms=644623 gate_skip=false
FM_TEST_BEGIN 2026-08-03T10:01:02Z tests/fm-turnend-guard.test.sh family=watcher-wake-lock expected_gate_skip=none
ok - fm_supervision_unhealthy: false with no state/*.meta at all
ok - fm_supervision_unhealthy: true with in-flight task and no beacon ever
ok - fm_supervision_unhealthy: true with in-flight task and a beacon far outside the grace window
ok - fm_supervision_unhealthy: false with in-flight task and a fresh beacon
ok - fm_supervision_status: FM_SUP_QUEUE_PENDING tracks state/.wake-queue
ok - fm_supervision_needed: X-mode relay poll needs supervision
ok - fm_supervision_unhealthy: source-only home needs supervision
ok - fm-turnend-guard: silent no-op with nothing in flight
ok - fm-turnend-guard: blocks when a fresh beacon has no live watcher lock
ok - fm-turnend-guard: non-Claude path blocks a source-only home
ok - fm-turnend-guard: blocks on a dead watcher lock even when the beacon is fresh
ok - fm-turnend-guard: silent no-op with a live watcher lock and fresh beacon
ok - fm-turnend-guard: healthy non-Claude harness paths ignore Claude episode contention
ok - fm-turnend-guard: blocks on a live watcher lock with an ancient beacon
ok - fm-turnend-guard: blocks with the exact required reason in the primary when unhealthy
ok - fm-turnend-guard: blocks from active FM_HOME state, not only repo-root state
ok - f

... [9338 bytes truncated] ...


ok - documentation inventory classifies every maintained prose surface exactly once
ok - classification, setup routing, and maintained-prose scope fail safely
ok - required documentation owner pointers cannot silently disappear
ok - local links resolve while dates, versions, commands, and incident prose remain semantically reviewed
FM_TEST_END 2026-08-03T10:04:28Z tests/fm-documentation-audiences.test.sh exit=0 duration_ms=2394 gate_skip=false
FM_TEST_BEGIN 2026-08-03T10:04:28Z tests/fm-ensure-agents-md.test.sh family=pure-contract-unit expected_gate_skip=none
ok - fm-ensure-agents-md.sh: created AGENTS.md includes self-governance section
ok - fm-ensure-agents-md.sh: promoted CLAUDE.md includes self-governance section
ok - fm-ensure-agents-md.sh: newline-less promotion keeps a blank separator line
ok - fm-ensure-agents-md.sh: existing symlinked AGENTS.md gains the section idempotently
ok - fm-ensure-agents-md.sh: existing AGENTS.md without CLAUDE.md gains section and symlink
ok - fm-ensure-agents-md.sh: AGENTS.md that already has the section stays unchanged
ok - fm-ensure-agents-md.sh: CRLF AGENTS.md with the section stays unchanged
ok - fm-ensure-agents-md.sh: CRLF injection preserves line endings idempotently
ok - fm-ensure-agents-md.sh: refuses a case-variant lowercase agents.md (issue #389)
FM_TEST_END 2026-08-03T10:04:30Z tests/fm-ensure-agents-md.test.sh exit=0 duration_ms=1938 gate_skip=false
FM_TEST_BEGIN 2026-08-03T10:04:30Z tests/fm-supervision-instructions.test.sh family=pure-contract-unit expected_gate_skip=none
ok - renderer prints exactly the selected harness block
ok - renderer falls back to unknown.md for unverified harness names
ok - renderer includes read-only, afk, and effective x-mode current-state stanzas
ok - renderer repair-line mode is harness-aware and honors conditional state
ok - renderer preserves every harness ordinary-continuation and missing-cycle repair path
ok - pi-signed keeps its identity while sharing Pi's supervision protocol
ok - grok supervision is Claude-shaped background notify with passive Stop-hook backstop
ok - grok rendered command sources the effective x-mode config
ok - pi supervision snippet renders the effective extension path
FM_TEST_END 2026-08-03T10:04:32Z tests/fm-supervision-instructions.test.sh exit=0 duration_ms=1679 gate_skip=false
FM_TEST_BEGIN 2026-08-03T10:04:32Z tests/fm-guard-stale-banner.test.sh family=watcher-wake-lock expected_gate_skip=none
ok - fm-guard stale banner: first stale call prints the full actionable banner
ok - fm-guard stale banner: repeated same-episode calls print a concise reminder only
ok - fm-guard stale banner: a fresh beacon without a live watcher remains unhealthy
ok - fm-guard stale banner: X-mode polling without a live watcher remains unhealthy
ok - fm-guard stale banner: healthy recovery rearms the next stale episode
ok - fm-guard stale banner: concurrent same-episode calls claim exactly one full banner
ok - fm-guard stale banner: deduplication is isolated per FM_HOME
ok - fm-guard stale banner: queued-wake warning remains independent
ok - fm-guard stale banner: read-only before writable does not consume full banner
ok - fm-guard stale banner: read-only during episode observes without mutating marker
ok - fm-guard stale banner: healthy read-only does not clear marker
ok - fm-guard stale banner: read-only never mutates stale-banner state files
FM_TEST_END 2026-08-03T10:04:44Z tests/fm-guard-stale-banner.test.sh exit=0 duration_ms=11883 gate_skip=false
FM_TEST_BEGIN 2026-08-03T10:04:44Z tests/fm-wake-queue.test.sh family=watcher-wake-lock expected_gate_skip=none
ok - concurrent append plus drain preserves queue records
ok - signal written while no watcher runs is caught on next run
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no watcher has a fresh beacon (last beat: 4s ago, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  watcher supervision needs Stop-owned automatic recovery; inspect the hook registration and startup status before ending the turn.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
ok - stale wake is queued before suppressor state is advanced
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
●  WATCHER DOWN - SUPERVISION IS OFF
●  1 task(s) in flight, but no watcher has a fresh beacon (last beat: 3s ago, grace 300s).
●  Trust the emitted supervision protocol for this harness; do not use shell & for watcher repair.
●  This is a supervision warning only; the guarded operation WILL still run.
●  watcher supervision needs Stop-owned automatic recovery; inspect the hook registration and startup status before ending the turn.
●━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
ok - a not-provably-working stale wake is queued before its suppressor is advanced
ok - registered custom check output is queued before cadence suppression
ok - two atomic drains cannot consume the same records twice
ok - drain collapses obvious duplicate heartbeat and signal records
ok - drain asserts watcher liveness: warns on a lapse, stays silent for a live watcher with a fresh beacon
ok - structural signal enrichment is separate, deduped, home-local, and tier-zero for other wakes
ok - bounded reads and per-item/global caps fail open with explicit truncation and omission markers
ok - slow annotation releases the append lock and a deleted status file fails open
ok - interruptions restore before commitment and never replay after raw commitment
FM_TEST_END 2026-08-03T10:07:03Z tests/fm-wake-queue.test.sh exit=0 duration_ms=138893 gate_skip=false
FM_TEST_SUMMARY total=15 failed=1 skipped_gate=0 duration_ms=1277077
FM_TEST_SUMMARY_FAMILY family=backend-dispatch count=2 duration_ms=93524 failed=0
FM_TEST_SUMMARY_FAMILY family=pure-contract-unit count=6 duration_ms=152779 failed=0
FM_TEST_SUMMARY_FAMILY family=unclassified count=1 duration_ms=64123 failed=0
FM_TEST_SUMMARY_FAMILY family=watcher-wake-lock count=6 duration_ms=963571 failed=1
FM_TEST_SLOWEST rank=1 script=tests/fm-watcher-lock.test.sh duration_ms=644623
FM_TEST_SLOWEST rank=2 script=tests/fm-wake-queue.test.sh duration_ms=138893
FM_TEST_SLOWEST rank=3 script=tests/fm-test-run.test.sh duration_ms=124625
FM_TEST_SLOWEST rank=4 script=tests/fm-turnend-guard.test.sh duration_ms=100845
FM_TEST_SLOWEST rank=5 script=tests/fm-spawn-dispatch-profile.test.sh duration_ms=88966
FM_TEST_SLOWEST rank=6 script=tests/fm-claude-stop-autoarm.test.sh duration_ms=64123
FM_TEST_SLOWEST rank=7 script=tests/fm-watch-arm.test.sh duration_ms=55361
FM_TEST_SLOWEST rank=8 script=tests/fm-task-delivery.test.sh duration_ms=13279
FM_TEST_SLOWEST rank=9 script=tests/fm-session-lock-ancestry.test.sh duration_ms=11966
FM_TEST_SLOWEST rank=10 script=tests/fm-guard-stale-banner.test.sh duration_ms=11883
FM_TEST_SLOWEST rank=11 script=tests/fm-brief.test.sh duration_ms=8864
FM_TEST_SLOWEST rank=12 script=tests/fm-spawn-batch.test.sh duration_ms=4558
FM_TEST_SLOWEST rank=13 script=tests/fm-documentation-audiences.test.sh duration_ms=2394
FM_TEST_SLOWEST rank=14 script=tests/fm-ensure-agents-md.test.sh duration_ms=1938
FM_TEST_SLOWEST rank=15 script=tests/fm-supervision-instructions.test.sh duration_ms=1679
fm-test-run: wrote timing artifact: /var/folders/hv/gd626zx51jb1_3rghgd97mp00000gn/T/no-mistakes-evidence/01KZ3ATDAK72A7T6ET2E48CYFM/targeted-timing.json
- Evidence: Timing artifact for the targeted run (local file: /var/folders/hv/gd626zx51jb1_3rghgd97mp00000gn/T/no-mistakes-evidence/01KZ3ATDAK72A7T6ET2E48CYFM/targeted-timing.json)
Evidence: Clean isolated rerun of fm-watcher-lock confirming the flake

Source: Clean isolated rerun of fm-watcher-lock confirming the flake (local file: /var/folders/hv/gd626zx51jb1_3rghgd97mp00000gn/T/no-mistakes-evidence/01KZ3ATDAK72A7T6ET2E48CYFM/watcher-lock-rerun.log)

FM_TEST_END 2026-08-03T10:09:37Z tests/fm-watcher-lock.test.sh exit=0 duration_ms=138276 gate_skip=false
FM_TEST_SUMMARY total=1 failed=0 skipped_gate=0 duration_ms=138527
- Outcome: ⏭️ skipped across 1 run (27m56s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⏭️ **Review** - skipped
  • ℹ️ docs/fm-test-portable-shards.md:67 - The portable-serial shard table and the portable_serial_weight_hints block in bin/fm-test-run.sh were measured for a 69-script serial lane in the oldest of the four merged upstream commits (f5ab708 / PR perf: shard portable serial tests across CI runners kunchenguid/firstmate#1544). The three later upstream commits add tests/fm-session-lock-ancestry.test.sh, tests/fm-watch-arm.test.sh, and tests/fm-task-delivery.test.sh; none appears in list_portable_parallel_1/2 or the real-herdr-gated family, so all three derive into portable-serial with no hint and take PORTABLE_SERIAL_DEFAULT_WEIGHT_MS (20000). tests/fm-claude-stop-autoarm.test.sh (hinted 60521 ms) also grew by 186 lines and tests/fm-turnend-guard.test.sh (hinted 5986 ms) by 461 lines in the same window. Coverage is unaffected because list_portable_serial is derived and run_coverage_guard proves the shards partition it exactly; only the published 15/18/17/19 partition, the ~285.9 s per-shard estimate, and the '3x margin' rationale for cutting .github/workflows/ci.yml tests-portable-serial from 20 to 15 timeout-minutes are no longer evidenced. Worst-case rebalance still estimates well under the 15-minute cap. Refreshing hints from a green CI run on this branch would restore the doc's accuracy.
  • ℹ️ bin/fm-turnend-guard.sh:151 - On the healthy-watcher path in Claude mode, fm_failure_episode_reset &#34;$STATE&#34; &amp;&amp; exit 0 falls through to a bare exit 2, blocking the Stop with no stderr, so the model receives 'Stop hook feedback' with no reason or instruction. fm_failure_episode_reset acquires .turnend-claude-blocks.lock through fm_lock_try_acquire, which is single-attempt and non-blocking (bin/fm-wake-lib.sh), so any concurrent holder makes it return 1 even though the watcher is verified healthy. Both .claude/settings.json Stop hooks fire on the same event and the previous Stop's asyncRewake auto-arm stays attached across the idle period, so bin/fm-claude-stop-autoarm.sh:209 can be taking that same lock on its HEALTHY branch while a new Stop's guard runs; the loser blocks (guard) or rewakes with no banner (auto-arm). The window is microseconds and the next Stop self-heals, and the fail-closed choice matches docs/turnend-guard.md ('positive watcher recovery clears the block budget ... before either hook reports ordinary recovery'), but a blocking exit with no message gives the operator and the model nothing to act on. A one-line stderr naming the lost budget lock would close the diagnostic gap without changing the fail-closed semantics.
⏭️ **Test** - skipped
  • ⚠️ tests/fm-watcher-lock.test.sh:474 - tests/fm-watcher-lock.test.sh case watch restart attaches to a verified healthy peer is timing-flaky under concurrent machine load: it failed once during my run (not ok - restart did not attach to the verified healthy peer: watcher: started pid=78004 (beacon fresh)) while a Chrome screenshot render was running in parallel, then passed 4/4 in isolation. The case runs fm-watch-arm --restart with FM_ARM_CONFIRM_TIMEOUT=1 and FM_ARM_ATTACH_POLL=0.1, so under load the healthy-peer confirmation window can expire and the arm starts a fresh watcher instead of attaching. Two aggravating factors: (1) on failure the fail helper exits the script but leaves the started fm-watch.sh child alive, and that orphan holds fm-test-run.sh's stdout pipe, so the runner hung indefinitely (~10 min until I killed pid 78004) rather than moving to the next script; (2) the file is unchanged by this merge - it is byte-identical to upstream/main - so this is a pre-existing upstream flake, not merge-induced. Fixing it is outside a sync merge's scope, so the captain should decide whether to widen the confirmation budget / add child cleanup to the case now or file it upstream.
  • bin/fm-test-run.sh --json &lt;evidence&gt;/targeted-timing.json tests/fm-task-delivery.test.sh tests/fm-session-lock-ancestry.test.sh tests/fm-claude-stop-autoarm.test.sh tests/fm-test-run.test.sh tests/fm-watch-arm.test.sh tests/fm-watcher-lock.test.sh tests/fm-turnend-guard.test.sh tests/fm-brief.test.sh tests/fm-spawn-batch.test.sh tests/fm-spawn-dispatch-profile.test.sh tests/fm-documentation-audiences.test.sh tests/fm-ensure-agents-md.test.sh tests/fm-supervision-instructions.test.sh tests/fm-guard-stale-banner.test.sh tests/fm-wake-queue.test.sh - 15 scripts, 14 green on first pass
  • bash tests/fm-watcher-lock.test.sh x3 standalone plus bin/fm-test-run.sh tests/fm-watcher-lock.test.sh - 4/4 clean reruns (30 ok, 0 not ok each) isolating the healthy-peer flake
  • bin/fm-test-run.sh --check-coverage - FM_TEST_COVERAGE ok total=107 parallel=24 serial=72 serial_shards=4 herdr=11
  • bin/fm-test-run.sh --list-lanes and --list --lane portable-serial-{1,2,3,4}of4 diffed against --list --lane portable-serial - union equals the whole lane, zero duplicates, zero missing
  • bin/fm-test-run.sh --list --lane portable-serial-1of3 - runner refuses a shard count that disagrees with it (exit 2)
  • Manual CLI verification of the per-task delivery contract: bin/fm-spawn.sh &lt;id&gt; &lt;project&gt; claude with (a) no delivery flags, (b) --mode fast-path --yolo off, (c) --mode no-mistakes and no yolo, (d) --mode direct-PR --yolo off against a brief recording mode=no-mistakes - all four refused with distinct explanations and left state/ empty
  • git rev-list --parents -n1 7bc4c0c - merge parents match the brief's origin/main and upstream/main exactly
  • git diff 4ee4a0a 7bc4c0c plus a per-file git diff --quiet 4ee4a0a 7bc4c0c -- &lt;f&gt; sweep over all 56 upstream-changed files - AGENTS.md is the sole divergence
  • git show 2bd6d0c -- AGENTS.md | grep &#39;^+&#39; each line re-checked with grep -Fqx against merged AGENTS.md - all 4 fork lines PRESENT
  • git show &lt;sha&gt;:AGENTS.md | wc -c/-l at cd73e75 / 2bd6d0c / 4ee4a0a / 7bc4c0c - line accounting 536+3+6=545
  • Rendered the merged AGENTS.md section 7 intake block to HTML with per-line provenance and screenshotted it via chrome-devtools-axi open + screenshot at 1440x1620
⚠️ **Document** - 2 infos
  • ℹ️ docs/fm-test-portable-shards.md:79 - The portable-serial weight hints in bin/fm-test-run.sh cover 69 scripts while the lane now holds 72; the three tests this merge added (fm-session-lock-ancestry, fm-watch-arm, fm-task-delivery) run on the conservative PORTABLE_SERIAL_DEFAULT_WEIGHT_MS default. I documented the drift, but refreshing the hints needs per-shard timing artifacts from a green CI run and is a code change outside this docs phase. Follow-up: refresh the hints and the table from the next green run using the procedure already in docs/fm-test-portable-shards.md.
  • ℹ️ AGENTS.md:118 - This change adds state/.watch-deliveries.log (and its lock) as the watcher's terminal-delivery ledger, which is not listed in AGENTS.md section 2's state inventory. I deliberately did not add it: the pre-existing sibling ledger state/.watch-cycle-exits.log is also absent from that inventory, and docs/watcher-continuity.md already owns both, so adding a line would grow always-loaded guidance against the established precedent. Flagging as a judgment call in case the captain wants arm-layer ledgers enumerated there.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

kunchenguid and others added 6 commits August 2, 2026 14:42
* perf(ci): shard the portable serial behavior lane across runners

The Behavior portable serial job ran all 69 scripts of the serial
remainder on one runner. The measured serial sum on run 30725985757 was
1143762 ms (19m04s) against a 20-minute timeout, so the job intermittently
reached the cap and was cancelled with every step passing. Setup is only
about 7s, so the cost is entirely test wall time.

Split the lane into four separate-runner shards. Each shard is still
strictly serial, and separate runners mean no two of these stateful
scripts ever share a machine, so the split needs no concurrency isolation
proof. Assignment is longest-processing-time bin packing over measured
per-script duration hints, balancing every shard to 285941 ms (~4m46s) of
expected work, and the timeout tightens from 20 to 15 minutes.

bin/fm-test-run.sh owns the shard count and refuses a lane whose "ofN"
disagrees with it, while ci.yml derives the same count from
strategy.job-total rather than a literal, so changing it in either file
alone fails the lane loudly instead of leaving part of the required suite
unrun. --check-coverage additionally proves the shards are non-empty,
disjoint, and exactly equal to the serial lane. No test is weakened,
skipped, or removed.

Also replace the wall-clock sleeps in the --jobs scheduler test fixture
with an explicit signal handshake between the fixtures. The old
0.5s-versus-0.05s race failed on a loaded machine; the handshake passes
under sustained CPU saturation.

* no-mistakes(review): Correct portable serial shard balance evidence

* no-mistakes(document): Document portable serial shard evidence accurately
…henguid#1545)

* fix(bin): identify harness sessions by path and report delivered wakes

Two supervision faults, both reported by a contributor and both open on the
default branch.

Fault 1: the Stop auto-arm never claims the home. fm_harness_ancestry_pid()
matched only the basename of `ps -o comm=`, and Claude Code's native installer
names the per-session executable by its version (.../share/claude/versions/
2.1.220), so that basename identifies nothing. Three real failure shapes follow:
a version-named session is missed entirely and the hook exits 0 with the epoch
never written (unconditional on Linux, where procps reports the kernel exec name
and ignores argv[0]); a claude-named daemon that directly parents sessions wins
the outermost-contiguous-claude rule ahead of the session itself; and a session
that is both version-named and daemon-parented has its live lock reclaimed as
stale and rewritten to the shared daemon pid, corrupting the home's ownership
record.

Harness identity now also reads whole components of the executable path and of
argv[0], which is what both platforms still carry. Matching whole components
only keeps that widening safe: bin/fm-claude-stop-autoarm.sh and ~/.claude/hooks
scripts have no "claude" component. Ownership is then decided against the
session's whole contiguous harness ancestry rather than one chosen pid, which is
the honest form of the question the library already documents ("does the current
process descend from that same harness?"). That subsumes the outermost-pid rule
for Claude's nested bg-spare worker chain instead of reverting it, and lets a
daemon-parented session recognize its own lock. Lock acquisition still writes the
outermost pid of the run, the only pid that lives as long as the session.

Fault 2: an attached arm reports a delivered cycle as FAILED. The watcher prints
its one reason line to its own stdout, so only the arm that forked it can read
that line; an arm that attached observes nothing but a released lock and called a
completely successful cycle "cycle ended without an actionable reason". No
supervision event was lost - the durable queue held it - but every harness
protocol reads that line as "supervision is down" and directs a manual re-arm.

The arm now resolves an unobservable close against the durable wake queue, which
records every wake before the watcher prints it and whose sequence counter never
rewinds, not even across a drain. A cycle the queue proves delivered a wake
reports that wake and exits 0; a cycle whose records a handling turn already
drained reports the delivery without inventing a reason line; only a cycle that
delivered nothing is still the typed nonzero failure. Fixing it in the arm covers
codex, opencode, pi, grok and kimi, not just the Claude Stop path.

Regressions: tests/fm-session-lock-ancestry.test.sh pins both platforms' ps
semantics behind a deterministic process table and runs the real Stop auto-arm in
version-named, daemon-parented, and combined real process trees, each orphaned so
the walk cannot escape the fixture. tests/fm-watch-arm.test.sh drives a real
watcher and a real attached arm through a real wake. Every fault case fails on
the previous code.

* no-mistakes(review): Bind watcher delivery records to process identity

* no-mistakes(review): Return validated watcher identity atomically

* no-mistakes(review): Track watcher successors by PID and identity

* no-mistakes(document): Consolidate watcher arm-cycle documentation ownership
* fix(supervision): harden Claude auto-arm failure handling

* no-mistakes(review): Guarantee automatic retry after Claude auto-arm failures

* no-mistakes(review): Gate attended fail-open on verified supervision failure

* no-mistakes(document): Document Claude auto-arm retry and guard scope

* no-mistakes: apply CI fixes

* fix(supervision): make Claude fail-open progression monotonic

* no-mistakes(review): Preserve auto-arm failure episodes until verified watcher recovery

* no-mistakes(review): Linearize auto-arm failure progression across existing locks

* no-mistakes(review): Linearize positive recovery across shared failure episode lock

* no-mistakes(review): Scope Claude recovery contention to Claude guard mode

* no-mistakes(document): Align supervision auto-arm documentation

* no-mistakes(review): Preserve actionable wakes despite healthy successors

* no-mistakes(document): Refresh supervision auto-arm documentation
…d#1563)

* feat(bin): require an explicit ship delivery mode in fm-brief

A ship brief's definition of done was shaped by a silent per-project registry
lookup, so an adjusted brief and the task's recorded delivery could disagree and
no one had to decide anything per task.

fm-brief now requires --mode on ship scaffolds, validates it against the closed
set, refuses the conditional no-mistakes-prod-only registry policy as a task
mode, and records the choice as a fixed machine-readable "Delivery contract:
mode=<mode>" line that fm-spawn can check. --mode is refused on scout and
secondmate scaffolds, and --yolo is refused outright because the worker never
owns approval decisions.

* feat(bin): require an explicit ship delivery contract at spawn and promotion

fm-spawn resolved every ship and scout task's mode and yolo from the project
registry, so the delivery posture was never a per-task decision and could
contradict the brief the worker was about to follow.

fm-spawn now requires --mode and --yolo on ship spawns, validates both against
their closed sets, and reads the brief's recorded delivery contract line and
refuses a mismatch before any endpoint exists; a brief scaffolded before that
line existed warns once and launches on the flag. A batch carries one shared
contract that each pair still checks against its own brief. Scout and secondmate
spawns refuse the flags, and a scout now records no mode or yolo at all, which
teardown and the snapshot already tolerate. When the explicit mode carries less
rigor than the project's standing posture, a deviation notice is printed and the
spawn continues, so the registry stays advisory rather than an enforced default.

fm-promote requires the same two flags, because a scout carries no posture to
inherit, and writes them into the task record with the kind flip.

fm-project-mode keeps its one registry parser for the mechanical consumers that
have no task in hand, accepts the conditional no-mistakes-prod-only annotation
and maps it to its most rigorous leg for them, and grows --raw so the deviation
notice can tell a conditional policy apart from a flat mode.

* docs: record the explicit per-task delivery contract

AGENTS.md section 7 now owns how each ship task's mode and yolo are resolved at
intake, including the surface classification for a no-mistakes-prod-only project
and the unregistered-project fallback, and the project-management skill defines
that conditional policy as a registration-time posture with its defaults and
initialization consequences. The registry blurb, script table, and architecture
section follow: the registry records the captain's standing posture, and task
delivery is decided per task and passed explicitly.

* test: pass ship delivery flags per call site in the Herdr launcher e2e

The shared spawn helper also launches a secondmate, which refuses the flags, so
the contract belongs at each ship call site rather than inside the helper.

* test: pass the ship delivery contract in the secondmate suites

Both suites scaffold or spawn an ordinary ship task as the control case for a
secondmate assertion, so each needs the explicit contract the ship path now
requires.
@kaku-san
kaku-san merged commit 294d0e2 into main Aug 3, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants