Skip to content

fix(bin): correct false answered-decision and failed-run supervision reports - #3265

Open
tknguyen29032002 wants to merge 21 commits into
kunchenguid:mainfrom
tknguyen29032002:fm/fm-supervision-reliability
Open

tknguyen29032002 wants to merge 21 commits into
kunchenguid:mainfrom
tknguyen29032002:fm/fm-supervision-reliability

Conversation

@tknguyen29032002

@tknguyen29032002 tknguyen29032002 commented Aug 29, 2026 •

Copy link
Copy Markdown

Intent

Fix two supervision reliability defects in firstmate's own shared tracked material, both about how firstmate reads a worker's true state, treated as one coherent piece of work.

Defect 1 - a silently undelivered instruction reads as answered. The doorbell-skip fail-safe in bin/fm-send.sh (composer holds pending text) must stay exactly as-is; failing safe there is correct. The bug was that --resolve-key wrote the decision's closing "resolved [key=...]" line to the task status log at SEND time, so the OPEN DECISIONS fold and the session-start digest reported a decision as answered while the worker had never seen it. Fixed by parking the closure as a dot-prefixed sidecar beside the durable inbox record; it commits only when the worker acknowledges the record, either on the supervision poll or at teardown as a last commit while the status log is still live, naming any closure never acknowledged as undelivered. The typed and remote planes still close immediately since they prove delivery directly. The steering-inbox escalation ladder gained bounded causes beyond the spent delivery-attempt budget: a provably composer-blocked skip, an absolute unhandled-age bound that fires even while the pane reads busy, and orphaned-sidecar detection, quarantine and retry.

Defect 2 - a healthy waiting job reported as failed. bin/fm-crew-state.sh matched a stale or superseded no-mistakes run instead of the one belonging to the branch and current code, producing false "failed" wakes for healthy validating crews. Fixed by binding run attribution to the branch's newest run row or attributing nothing.

Constraints: do not weaken the doorbell-skip fail-safe; tests colocated and demonstrated to fail on reintroduction; shellcheck clean; where something cannot be made certain make it loud rather than quiet. The separately-filed fm-turnend-guard-readonly defect is explicitly out of scope.

Regression coverage is colocated in tests/fm-send-resolve-key.test.sh, tests/fm-task-inbox.test.sh, tests/fm-crew-state.test.sh, tests/fm-wake-drain-open-decisions.test.sh and tests/fm-teardown.test.sh, each confirmed to fail when only its own fix is reverted. An earlier review round in this same effort fixed two further real defects in the new code: the overdue escalation was not deduplicated and would have re-woken firstmate on every poll forever, and the orphan sidecar scan accepted a non-file and could starve the ladder.

IMPORTANT SCOPE RULE FOR THIS RUN: this branch must contain ZERO changes under .github/. A previous CI auto-fix attempt on this branch autonomously committed 4bdde97, which edited the required-no-mistakes workflow, added a new PR-body helper script under bin/, and edited CONTRIBUTING.md. That commit has been reverted in full by explicit decision and must NOT be reintroduced. Its diagnosis was correct - on a synchronize event the required-no-mistakes gate judges the event payload's PR body, which is a snapshot taken before the pr step writes the head-bound attestation, so that run fails while the later edited-event run on the same head passes - but changing the repo-wide verification surface from a fork pull request is out of scope here, was explicitly excluded by the maintainer's approval condition of no workflow writes, and is being tracked as its own separate piece of work. Do not edit, re-add, or work around anything under .github/, and do not modify any gate to make a check pass. If a check is red because of that known stale superseded synchronize run, that is expected and is for a maintainer to re-run; it is not a defect to fix here.

Also merged origin/main (commit c7fdef9) to clear a merge conflict, resolving one conflict in docs/scripts.md by keeping both sides' independent row edits.

Never add an agent name as a commit co-author. Do not merge the PR.

What Changed

  • --resolve-key no longer writes the closing resolved [key=...] line at send time on the local inbox plane. bin/fm-send.sh now parks the closure as a dot-prefixed sidecar beside the durable inbox record, and bin/fm-task-inbox-lib.sh commits it only once the worker acknowledges that record: on the watcher's supervision poll (bin/fm-watch.sh), or as a last commit at cleanup while the status log is still live (bin/fm-teardown.sh), which names any closure never acknowledged as undelivered. The typed and remote planes still close immediately, and bin/fm-wake-drain.sh marks an open decision whose answer is delivered but still unread.
  • The steering-inbox escalation ladder gained bounded causes beyond a spent delivery-attempt budget: a provably composer-blocked doorbell skip (FM_TASK_INBOX_BLOCKED_MAX), an absolute unhandled-age bound that fires even while the pane reads busy (FM_TASK_INBOX_UNHANDLED_MAX_SECS, default 900s), plus detection, quarantine under handled/orphaned/, and retry of orphaned sidecars. Each escalation is recorded once per record so a permanently failing closure cannot re-wake firstmate on every poll. The doorbell-skip fail-safe itself is unchanged.
  • bin/fm-crew-state.sh now binds no-mistakes run attribution to the branch's newest run row or attributes nothing, instead of walking past an unbindable newest row onto an older superseded one and reporting a healthy validating crew as failed; the now-unused fm_nm_head_resolvable helper is dropped from bin/fm-nm-run-lib.sh. Regression coverage is colocated in tests/fm-send-resolve-key.test.sh, tests/fm-task-inbox.test.sh, tests/fm-crew-state.test.sh, tests/fm-wake-drain-open-decisions.test.sh, and tests/fm-teardown.test.sh, with the docs, AGENTS.md, and skill files updated to describe acknowledgement-gated closure.

Risk Assessment

✅ Low: The change is large but well-bounded to its stated intent, satisfies every source-verifiable required constraint including the zero-.github scope rule, and carries colocated behavioral regression coverage for each fixed defect; nothing survived this pass except two mechanical documentation/test-selection findings that cannot change runtime behavior.

Testing

Ran the colocated regression suites the intent names, then demonstrated both defects end-to-end through the surfaces an operator actually reads, reproducing each on the base commit and showing it fixed on the target. For defect 1 that meant driving the real fm-send with a visibly occupied composer (so the doorbell fail-safe skips, unchanged) and reading fm-wake-drain's OPEN DECISIONS fold: base silently reports the decision answered, target keeps it open and marks the answer as delivered-but-unread until the worker's acknowledgement commits the closure. For defect 2, the real fm-crew-state over real git repos and real-shaped run rows reports a live validating crew as failed on base and as working on target, in both false-failure shapes. I also captured the three new escalation wakes, the teardown cleanup transcript naming an undelivered closure, and the last commit's changed-file selection newly picking up the captain-hold suite, and confirmed each fix's own suite fails when only that fix is reverted. Three suites fail on this machine - fm-teardown, fm-captain-hold-lifecycle and fm-gotmp - but all three fail identically on an unmodified origin/main checkout here, so none is a regression from this change. Root cause: stock macOS bash 3.2 kills the shell with status 0 on a failed dot-builtin under set -e even when it is guarded, so the preflight never reports the missing adapter. CI's authoritative lanes are ubuntu-latest and never run these three suites; the macos-stock-bash job deliberately runs only bash -n plus two snapshot suites. Tracked separately as fm-preexisting-test-failures. This change's own new cases inside those suites pass when run in isolation. The worktree is clean and the branch has no .github/ changes.

Evidence: Defect 1 before/after: an undelivered answer no longer reads as resolved (real fm-send + fm-wake-drain transcript)

Source: Defect 1 before/after: an undelivered answer no longer reads as resolved (real fm-send + fm-wake-drain transcript)

BASE c7fdef9: ----- 3. the task status log right after the answer was sent ----- 1 needs-decision [key=api-shape]: pick REST or RPC for the mapping endpoint 2 working: kept busy on an unrelated stream 3 resolved [key=api-shape]: answered: go with REST ----- 4. what the supervisor now reads (OPEN DECISIONS fold, fm-wake-drain.sh) ----- (empty) >>> VERDICT: decision reads as ANSWERED although the worker never saw it. THE DEFECT. TARGET fefdf60: ----- 3. the task status log right after the answer was sent ----- 1 needs-decision [key=api-shape]: pick REST or RPC for the mapping endpoint 2 working: kept busy on an unrelated stream ----- 4. what the supervisor now reads (OPEN DECISIONS fold, fm-wake-drain.sh) ----- OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line): t1 [key=api-shape] needs-decision: pick REST or RPC for the mapping endpoint (answer already delivered, still unread by the worker) >>> VERDICT: decision still OPEN and marked as delivered-but-unread. CORRECT. ----- 5/6/7. worker acknowledges -> commit_resolutions: ok -> 'resolved [key=api-shape]: answered: go with REST' appended -> fold empty >>> VERDICT: the decision closed exactly when the worker acknowledged it. CORRECT.

Defect 1 - a silently undelivered instruction reads as answered
Driven end to end through the real bin/fm-send.sh and bin/fm-wake-drain.sh.
The worker's composer visibly holds pending text, so the doorbell fail-safe skips - that part is unchanged and correct.

=========================================================================
 DEFECT 1 - the captain answers a decision the worker never sees
 tree under test: BASE  c7fdef9 (origin/main, before the fix)
=========================================================================

----- 1. the worker has an open captain decision; what a supervisor reads -----
OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line):
t1 [key=api-shape] needs-decision: pick REST or RPC for the mapping endpoint
OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'

----- 2. the captain answers it: fm-send.sh t1 --resolve-key api-shape "go with REST" -----
   (the worker's composer visibly holds pending text, so the
    doorbell fail-safe must skip the keystroke - that part is correct)

WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
fm-send: doorbell skipped (composer visibly holds pending text); the steer is durably recorded at /var/folders/mj/5x6plchd335fygk8y0j5r7mw0000gn/T//fmdemo1.ZLRSxJ/home/state/t1.inbox/001.msg and the watcher will re-ring
   exit=0
   keystrokes actually typed into the worker pane: 0 bytes
   durable inbox record on disk: 001.msg 
   worker acknowledgements (handled/): 0 - the worker has NOT seen it

----- 3. the task status log right after the answer was sent -----
     1	needs-decision [key=api-shape]: pick REST or RPC for the mapping endpoint
     2	working: kept busy on an unrelated stream
     3	resolved [key=api-shape]: answered: go with REST

----- 4. what the supervisor now reads (OPEN DECISIONS fold, fm-wake-drain.sh) -----

>>> VERDICT: decision reads as ANSWERED although the worker never saw it. THE DEFECT.

----- 5. the worker finally reads its inbox and acknowledges the record -----
   moved to handled/: 001.msg 
   the next supervision poll commits any parked closure:
   (this tree has no acknowledgement-gated closure to commit)

----- 6. the task status log after the acknowledgement -----
     1	needs-decision [key=api-shape]: pick REST or RPC for the mapping endpoint
     2	working: kept busy on an unrelated stream
     3	resolved [key=api-shape]: answered: go with REST

----- 7. what the supervisor reads once the worker has actually read the answer -----
(nothing - no open decisions)

>>> VERDICT: closed - but it had ALREADY read as closed back at step 4, before delivery.



=========================================================================
 DEFECT 1 - the captain answers a decision the worker never sees
 tree under test: TARGET  fefdf60 (fm/fm-supervision-reliability, after the fix)
=========================================================================

----- 1. the worker has an open captain decision; what a supervisor reads -----
OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line):
t1 [key=api-shape] needs-decision: pick REST or RPC for the mapping endpoint
OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'

----- 2. the captain answers it: fm-send.sh t1 --resolve-key api-shape "go with REST" -----
   (the worker's composer visibly holds pending text, so the
    doorbell fail-safe must skip the keystroke - that part is correct)

WARNING: watcher still down (same stale episode; last beat: never, grace 300s) - full banner already printed this episode.
fm-send: doorbell skipped (composer visibly holds pending text); the steer is durably recorded at /var/folders/mj/5x6plchd335fygk8y0j5r7mw0000gn/T//fmdemo1.6L8HDJ/home/state/t1.inbox/001.msg and the watcher will re-ring
   exit=0
   keystrokes actually typed into the worker pane: 0 bytes
   durable inbox record on disk: 001.msg 
   worker acknowledgements (handled/): 0 - the worker has NOT seen it

----- 3. the task status log right after the answer was sent -----
     1	needs-decision [key=api-shape]: pick REST or RPC for the mapping endpoint
     2	working: kept busy on an unrelated stream

----- 4. what the supervisor now reads (OPEN DECISIONS fold, fm-wake-drain.sh) -----
OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line):
t1 [key=api-shape] needs-decision: pick REST or RPC for the mapping endpoint (answer already delivered, still unread by the worker)
OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'

>>> VERDICT: decision still OPEN and marked as delivered-but-unread. CORRECT.

----- 5. the worker finally reads its inbox and acknowledges the record -----
   moved to handled/: 001.msg 
   the next supervision poll commits any parked closure:
api-shape
   commit_resolutions: ok

----- 6. the task status log after the acknowledgement -----
     1	needs-decision [key=api-shape]: pick REST or RPC for the mapping endpoint
     2	working: kept busy on an unrelated stream
     3	resolved [key=api-shape]: answered: go with REST

----- 7. what the supervisor reads once the worker has actually read the answer -----
(nothing - no open decisions)

>>> VERDICT: the decision closed exactly when the worker acknowledged it. CORRECT.
Evidence: Defect 2 before/after: a healthy validating crew no longer reads as failed (real fm-crew-state output)

Source: Defect 2 before/after: a healthy validating crew no longer reads as failed (real fm-crew-state output)

SCENARIO A - crew committed past the live run's head; older FAILED row binds at the worktree HEAD SCENARIO B - this branch's own live run at head 305b0969 that only the no-mistakes mirror holds BASE c7fdef9: $ bin/fm-crew-state.sh moved -> state: failed · source: run-step · run failed $ bin/fm-crew-state.sh mismatch -> state: failed · source: run-step · run failed >>> a live, busy, validating crew is reported FAILED off a superseded run. THE DEFECT. TARGET fefdf60: $ bin/fm-crew-state.sh moved -> state: working · source: pane · harness busy (fm-spawn) $ bin/fm-crew-state.sh mismatch -> state: working · source: pane · harness busy (fm-spawn) >>> the live crew reads as working, sourced from the pane. CORRECT.

Defect 2 - a healthy waiting job reported as failed
Driven end to end through the real bin/fm-crew-state.sh, over real throwaway git repos,
a fake no-mistakes CLI serving real-shaped run rows, and a genuinely busy pane.

=========================================================================
 DEFECT 2 - a healthy, actively validating crew reported as failed
 tree under test: BASE  c7fdef9 (origin/main, before the fix)
=========================================================================


#########################################################################
# SCENARIO A - the crew moved past the live run's head
#########################################################################

crew branch        : fm/switav-mapping-leaf-filter
crew worktree HEAD : df80c7e4
crew pane          : busy ("esc to interrupt") - the agent is working right now
crew status log    : working: validating

`no-mistakes axi status` (repo-wide answer; belongs to ANOTHER crew):
    run:
      id: "01OTHER"
      branch: fm/other-crew
      status: running
      head: "deadbee1"
      pr: ""
      findings: none

`no-mistakes runs --limit N` (newest first):
    running    fm/switav-mapping-leaf-filter 3c89ed2b  2026-08-27 19:02
    failed     fm/switav-mapping-leaf-filter df80c7e4  2026-08-27 17:17
    ^ row 1: this branch's LIVE run, pinned at 3c89ed2b - the crew has since
      committed on top, so this head resolves but is no longer current.
    ^ row 2: an OLDER, superseded FAILED run whose head is exactly the
      worktree's HEAD df80c7e4, so it binds perfectly. It is the trap.

--- what a captain/watcher actually reads: bin/fm-crew-state.sh moved ---
    state: failed · source: run-step · run failed

>>> VERDICT (scenario A): a live, busy, validating crew is reported FAILED off a superseded run. THE DEFECT.


#########################################################################
# SCENARIO B - this branch's own live run reports a head this worktree
#              cannot resolve (no-mistakes commits review fixes inside
#              its OWN repo mirror under ~/.no-mistakes/repos/<id>.git)
#########################################################################

crew branch        : fm/feat-mismatch
crew worktree HEAD : 3c89ed2b
crew pane          : busy ("esc to interrupt") - the agent is working right now
crew status log    : working: validating

`no-mistakes axi status` (this branch's OWN live run, head 305b0969
which does not exist in this worktree):
    run:
      id: "01LIVE"
      branch: fm/feat-mismatch
      status: running
      head: "305b0969"
      pr: ""
      findings: none

`no-mistakes runs --limit N`:
    failed     fm/feat-mismatch 3c89ed2b  2026-08-27 17:17
    ^ the only row is an OLDER FAILED run at the worktree's own HEAD.

--- what a captain/watcher actually reads: bin/fm-crew-state.sh mismatch ---
    state: failed · source: run-step · run failed

>>> VERDICT (scenario B): a live, busy, validating crew is reported FAILED off a superseded run. THE DEFECT.



=========================================================================
 DEFECT 2 - a healthy, actively validating crew reported as failed
 tree under test: TARGET  fefdf60 (fm/fm-supervision-reliability, after the fix)
=========================================================================


#########################################################################
# SCENARIO A - the crew moved past the live run's head
#########################################################################

crew branch        : fm/switav-mapping-leaf-filter
crew worktree HEAD : cca7d1d9
crew pane          : busy ("esc to interrupt") - the agent is working right now
crew status log    : working: validating

`no-mistakes axi status` (repo-wide answer; belongs to ANOTHER crew):
    run:
      id: "01OTHER"
      branch: fm/other-crew
      status: running
      head: "deadbee1"
      pr: ""
      findings: none

`no-mistakes runs --limit N` (newest first):
    running    fm/switav-mapping-leaf-filter 8acc7555  2026-08-27 19:02
    failed     fm/switav-mapping-leaf-filter cca7d1d9  2026-08-27 17:17
    ^ row 1: this branch's LIVE run, pinned at 8acc7555 - the crew has since
      committed on top, so this head resolves but is no longer current.
    ^ row 2: an OLDER, superseded FAILED run whose head is exactly the
      worktree's HEAD cca7d1d9, so it binds perfectly. It is the trap.

--- what a captain/watcher actually reads: bin/fm-crew-state.sh moved ---
    state: working · source: pane · harness busy (fm-spawn)

>>> VERDICT (scenario A): the live crew reads as working, sourced from the pane. CORRECT.


#########################################################################
# SCENARIO B - this branch's own live run reports a head this worktree
#              cannot resolve (no-mistakes commits review fixes inside
#              its OWN repo mirror under ~/.no-mistakes/repos/<id>.git)
#########################################################################

crew branch        : fm/feat-mismatch
crew worktree HEAD : 8acc7555
crew pane          : busy ("esc to interrupt") - the agent is working right now
crew status log    : working: validating

`no-mistakes axi status` (this branch's OWN live run, head 305b0969
which does not exist in this worktree):
    run:
      id: "01LIVE"
      branch: fm/feat-mismatch
      status: running
      head: "305b0969"
      pr: ""
      findings: none

`no-mistakes runs --limit N`:
    failed     fm/feat-mismatch 8acc7555  2026-08-27 17:17
    ^ the only row is an OLDER FAILED run at the worktree's own HEAD.

--- what a captain/watcher actually reads: bin/fm-crew-state.sh mismatch ---
    state: working · source: pane · harness busy (fm-spawn)

>>> VERDICT (scenario B): the live crew reads as working, sourced from the pane. CORRECT.
Evidence: The three new bounded escalation causes, as firstmate reads them (real fm-watch.sh -> wake queue)

Source: The three new bounded escalation causes, as firstmate reads them (real fm-watch.sh -> wake queue)

provably composer-blocked doorbell: stale: sess:fm-t1 (unread firstmate instruction: .../t1.inbox/001.msg cannot be delivered because the composer visibly holds pending text, so every doorbell is being skipped; clear the composer, then re-ring - it carries the answer to decision key(s) api-shape, which stay OPEN until the worker acknowledges the record) absolute unhandled bound, busy pane: stale: sess:fm-t1 (unread firstmate instruction: .../t1.inbox/001.msg has been unhandled for over 1s without an acknowledgement while the pane reads busy - the worker may be inside one long tool call; inspect it before treating it as stopped) orphaned closure (record removed instead of acknowledged): stale: sess:fm-t1 (steering-inbox contract violation: .../t1.inbox/.001.resolve is an answered decision's closure with no bound record in the inbox or handled/ - the worker likely removed its record instead of moving it into handled/, so the closure can never commit on its own; it is set aside under .../handled/orphaned/ and the answer it carries to decision key(s) api-shape - which the worker most likely did read - must be closed by hand)

Steering-inbox escalation ladder - the three new bounded causes, as firstmate reads them
Captured from the real bin/fm-watch.sh poll writing into the wake queue that
bin/fm-wake-drain.sh surfaces to the captain.

########## SUPERVISION WAKE - provably composer-blocked doorbell ##########
  1788062242	1	stale	sess:fm-t1	stale: sess:fm-t1 (unread firstmate instruction: /var/folders/mj/5x6plchd335fygk8y0j5r7mw0000gn/T/fm-task-inbox.QS4L8W/blocked-composer/state/t1.inbox/001.msg cannot be delivered because the composer visibly holds pending text, so every doorbell is being skipped; clear the composer, then re-ring - it carries the answer to decision key(s) api-shape, which stay OPEN until the worker acknowledges the record)
########## END ##########

ok - watcher: a blocked doorbell escalates by name and says which decision stays open

########## SUPERVISION WAKE - absolute unhandled bound, busy pane ##########
  1788062243	1	stale	sess:fm-t1	stale: sess:fm-t1 (unread firstmate instruction: /var/folders/mj/5x6plchd335fygk8y0j5r7mw0000gn/T/fm-task-inbox.QS4L8W/busy-overdue/state/t1.inbox/001.msg has been unhandled for over 1s without an acknowledgement while the pane reads busy - the worker may be inside one long tool call; inspect it before treating it as stopped)
########## END ##########

ok - watcher: a busy pane no longer holds an unread instruction in silence forever, and the bound speaks once

########## SUPERVISION WAKE - orphaned closure (record removed, not acknowledged) ##########
  1788062255	1	stale	sess:fm-t1	stale: sess:fm-t1 (steering-inbox contract violation: /var/folders/mj/5x6plchd335fygk8y0j5r7mw0000gn/T/fm-task-inbox.QS4L8W/orphan/state/t1.inbox/.001.resolve is an answered decision's closure with no bound record in the inbox or handled/ - the worker likely removed its record instead of moving it into handled/, so the closure can never commit on its own; it is set aside under /var/folders/mj/5x6plchd335fygk8y0j5r7mw0000gn/T/fm-task-inbox.QS4L8W/orphan/state/t1.inbox/handled/orphaned/ and the answer it carries to decision key(s) api-shape - which the worker most likely did read - must be closed by hand)
########## END ##########

ok - watcher: an orphaned closure is surfaced for hand closure by name, then set aside
Evidence: Teardown as the last commit point: acknowledged closure committed, unacknowledged one named undelivered

Source: Teardown as the last commit point: acknowledged closure committed, unacknowledged one named undelivered

$ bin/fm-teardown.sh task-x1 closed the acknowledged answer at cleanup: resolved [key=api-shape]: answered: go with REST undelivered at cleanup: task-x1 never acknowledged the answer to decision key(s) deploy-window (record 002.msg); that decision stays open teardown task-x1 complete (window firstmate:fm-task-x1, worktree ...) Backlog: task-x1 just finished. Run tasks-axi done task-x1 --note "local main", ... Backlog: include each 'closed the acknowledged answer at cleanup' line above in that done note - its closing line went with the task's status log, so the note is the only record of that answer after cleanup.

Teardown as the last commit point - what firstmate reads at cleanup
Real bin/fm-teardown.sh run: one closure the worker acknowledged, one it never did.

########## TEARDOWN TRANSCRIPT (what firstmate reads at cleanup) ##########
$ bin/fm-teardown.sh task-x1
  closed the acknowledged answer at cleanup: resolved [key=api-shape]: answered: go with REST
  undelivered at cleanup: task-x1 never acknowledged the answer to decision key(s) deploy-window (record 002.msg); that decision stays open
  teardown task-x1 complete (window firstmate:fm-task-x1, worktree /var/folders/mj/5x6plchd335fygk8y0j5r7mw0000gn/T//fm-teardown-tests.KCtRzy/inbox-closures/wt)
  Backlog: task-x1 just finished. Run tasks-axi done task-x1 --note "local main", then run tasks-axi ready for dependency-cleared candidates, check date gates, and dispatch only work whose blockers are gone and date is due.
  Backlog: include each 'closed the acknowledged answer at cleanup' line above in that done note - its closing line went with the task's status log, so the note is the only record of that answer after cleanup.
########## END TRANSCRIPT ##########

ok - teardown commits an acknowledged decision closure and names an undelivered one before removing the inbox
Evidence: Changed-file test selection before/after (fm-test-run.sh --list --changed over an edit confined to bin/fm-task-inbox-lib.sh)

Source: Changed-file test selection before/after (fm-test-run.sh --list --changed over an edit confined to bin/fm-task-inbox-lib.sh)

BASE c7fdef9 -> captain-hold suite selected? NO TARGET fefdf60 -> captain-hold suite selected? YES newly selected by the target that the base did not select: tests/fm-captain-hold-lifecycle.test.sh tests/fm-pr-check-security.test.sh tests/fm-pr-merge.test.sh tests/fm-review-diff.test.sh tests/fm-teardown.test.sh tests/fm-wake-drain-open-decisions.test.sh tests/fm-x-mode.test.sh

Changed-file test selection (fefdf60, the branch's last commit)
A change confined to bin/fm-task-inbox-lib.sh must now also select
tests/fm-captain-hold-lifecycle.test.sh, which gained a direct dependency on that library.
Driven through the runner's own selection surface: fm-test-run.sh --list --changed.

=== BASE  c7fdef9 (origin/main) ===
changed files:
  bin/fm-task-inbox-lib.sh

fm-test-run.sh --list --changed --base HEAD selects:

captain-hold suite selected? NO - a later edit that breaks the deferred closure would leave this suite unrun

  [53 suites listed, captain-hold not among them]

=== TARGET  fefdf60 (after the mapping fix) ===
changed files:
  bin/fm-task-inbox-lib.sh

fm-test-run.sh --list --changed --base HEAD selects:

captain-hold suite selected? YES

  [same 53 suites, plus the 6 newly mapped ones listed below]

newly selected by the target that the base did not select:
  tests/fm-captain-hold-lifecycle.test.sh
  tests/fm-pr-check-security.test.sh
  tests/fm-pr-merge.test.sh
  tests/fm-review-diff.test.sh
  tests/fm-teardown.test.sh
  tests/fm-wake-drain-open-decisions.test.sh
  tests/fm-x-mode.test.sh
Evidence: Regression coverage proof: each colocated suite fails when only its own fix is reverted

Source: Regression coverage proof: each colocated suite fails when only its own fix is reverted

Defect 1 revert (local inbox plane closes at ENQUEUE again): $ bash tests/fm-send-resolve-key.test.sh not ok - an unacknowledged answer must not read as resolved exit=1 Defect 2 revert (coarse list may walk past the newest unbindable row again): $ bash tests/fm-crew-state.test.sh not ok - a resolvable-but-mismatched newest run must not be resolved by an older run that binds (unexpected: 'state: failed') --- output --- state: failed · source: run-step · run failed

Regression coverage proof: each colocated suite fails when only its own fix is reverted.
Both reverts were applied to isolated copies of the target tree, never to the worktree.

=== Defect 1: revert bin/fm-send.sh's local-inbox plane back to closing at ENQUEUE ===
  -       fm_send_defer_resolved_keys "$INBOX_RECORD" "$RESOLVE_ANSWER_TEXT" || exit 1
  +       fm_send_close_resolved_keys "$RESOLVE_ANSWER_TEXT" || exit 1
  +       fm_send_feed_resolved_holds "$RESOLVE_ANSWER_TEXT" || exit 1

$ bash tests/fm-send-resolve-key.test.sh
  not ok - an unacknowledged answer must not read as resolved
  exit=1

=== Defect 2: revert bin/fm-crew-state.sh back to walking past an unbindable newest run ===
  (nm_runs_status_for_branch continues to the next row; the own-branch bind falls through
   to the coarse list again - both halves restored to the pre-fix shape)

$ bash tests/fm-crew-state.test.sh
  ok - cross-branch attribution picks the branch's most recent row
  ok - coarse run does not probe another branch's ci log
  ok - crew-state: an unbindable newest run never falls back to an older failed run
  not ok - a resolvable-but-mismatched newest run must not be resolved by an older run that binds (unexpected: 'state: failed')
  --- output ---
  state: failed · source: run-step · run failed
  exit=
Evidence: Targeted suite results and the pre-existing macOS bash-3.2 failures (with a reduced repro)

Source: Targeted suite results and the pre-existing macOS bash-3.2 failures (with a reduced repro)

FM_TEST_END tests/fm-task-inbox.test.sh exit=0 FM_TEST_END tests/fm-crew-state.test.sh exit=0 FM_TEST_END tests/fm-wake-drain-open-decisions.test.sh exit=0 FM_TEST_END tests/fm-send-resolve-key.test.sh exit=0 FM_TEST_END tests/fm-remote-transport-lanes.test.sh exit=0 FM_TEST_END tests/fm-teardown.test.sh exit=1 <- also fails on origin/main here FM_TEST_END tests/fm-captain-hold-lifecycle.test.sh exit=1 <- also fails on origin/main here tests/fm-gotmp.test.sh exit=1 <- also fails on origin/main here Reduced root cause (GNU bash 3.2.57, macOS /bin/bash): $ bash -c 'set -eu; src(){ . /nonexistent/x.sh || return 1; echo sourced-ok; }; req(){ echo "req entered"; if ! src; then echo "src FAILED"; return 1; fi; }; if req; then echo ok0; else echo nonzero; fi; echo "reached end"' req entered bash: line 1: /nonexistent/x.sh: No such file or directory (dies right there with status 0; on bash 4+/CI Ubuntu it prints "src FAILED" and continues)

Targeted suite results (worktree HEAD fefdf60), and the pre-existing local failures

Suites run (the ones the intent names as carrying this change's regression coverage):
  FM_TEST_END 2026-08-30T03:43:18Z tests/fm-task-inbox.test.sh exit=0 duration_ms=80902 gate_skip=false
  FM_TEST_END 2026-08-30T03:43:52Z tests/fm-crew-state.test.sh exit=0 duration_ms=34870 gate_skip=false
  FM_TEST_END 2026-08-30T03:44:03Z tests/fm-wake-drain-open-decisions.test.sh exit=0 duration_ms=10984 gate_skip=false
  FM_TEST_END 2026-08-30T03:44:55Z tests/fm-teardown.test.sh exit=1 duration_ms=51596 gate_skip=false
  FM_TEST_END 2026-08-30T03:45:39Z tests/fm-captain-hold-lifecycle.test.sh exit=1 duration_ms=43496 gate_skip=false
  FM_TEST_END tests/fm-send-resolve-key.test.sh exit=0 duration_ms=27019   (run separately)
  FM_TEST_END tests/fm-remote-transport-lanes.test.sh exit=0 duration_ms=73785

Two of those END lines report exit=1 (fm-teardown, fm-captain-hold-lifecycle), plus
tests/fm-gotmp.test.sh. All three fail IDENTICALLY on the base commit c7fdef9 (origin/main)
on this same machine, so none is a regression from this change:

  $ bash /tmp/nm-base-tree/tests/fm-teardown.test.sh   # origin/main, unmodified
    13:not ok - herdr-preflight-missing-adapter: teardown continued without its required preflight
  $ bash /tmp/nm-base-tree/tests/fm-captain-hold-lifecycle.test.sh
    aborts after the same 10 cases; test_bound_channel_answers_close_at_answer_time reports
    'not ok - the fixture channel captured no result to feed' (bin/fm-procevent.sh start
    captures nothing in this environment)
  $ bash /tmp/nm-base-tree/tests/fm-gotmp.test.sh
    not ok - teardown did not remove the tasktmp dir

Root cause of the fm-teardown one, reduced (macOS system bash 3.2, /bin/bash):
  under 'set -e' a failed '.' (source) builtin terminates the shell with status 0,
  even when guarded by '|| return 1' inside an 'if' condition, so
  teardown_herdr_require_prerequisites never gets to report the missing adapter:

  $ bash --version | head -1
    GNU bash, version 3.2.57(1)-release (arm64-apple-darwin25)
  $ bash -c 'set -eu
      src(){ . /nonexistent/x.sh || return 1; echo sourced-ok; }
      req(){ echo "req entered"; if ! src; then echo "src FAILED"; return 1; fi; }
      if req; then echo "req returned 0"; else echo "req returned nonzero"; fi
      echo "reached end"'
    req entered
    bash: line 1: /nonexistent/x.sh: No such file or directory
    (the script dies right after 'req entered'; on bash 4+/CI Ubuntu it continues)

The repository's authoritative CI lanes run on ubuntu-latest; .github/workflows/ci.yml's
macos-stock-bash job deliberately runs only 'bash -n' parsing plus two snapshot suites
under /bin/bash 3.2, not these.

The change's OWN new cases inside the two blocked suites were exercised in isolation and pass:
  tests/fm-teardown.test.sh  ok - teardown commits an acknowledged decision closure and names an undelivered one before removing the inbox
  tests/fm-teardown.test.sh  ok - teardown names a closure it cannot commit even after the watcher already surfaced that failure
  tests/fm-captain-hold-lifecycle.test.sh  ok - the chat channel feeds the same keyed-answer intake a captured review does
  tests/fm-gotmp.test.sh  ok - fm-teardown skips gracefully when tasktmp= is absent (backward compat)
  tests/fm-gotmp.test.sh  ok - fm-teardown skips gracefully when tasktmp= points to a nonexistent dir
Evidence: Reproduction script for the defect 1 transcript

Source: Reproduction script for the defect 1 transcript

#!/usr/bin/env bash
# End-to-end demonstration of Defect 1: "a silently undelivered instruction
# reads as answered".
#
# Drives the REAL bin/fm-send.sh and bin/fm-wake-drain.sh of the tree passed in
# $1, over a stubbed tmux whose composer visibly holds pending text - the exact
# condition that makes fm-send correctly skip the doorbell.
set -u
ROOT=$1
LABEL=$2
WORK=$(mktemp -d "${TMPDIR:-/tmp}/fmdemo1.XXXXXX")
HOME_DIR="$WORK/home"; mkdir -p "$HOME_DIR/state"
FB="$WORK/fakebin"; mkdir -p "$FB"

cat > "$FB/tmux" <<'SH'
#!/usr/bin/env bash
set -u
case "${1:-}" in
  send-keys)
    shift; literal=0
    while [ $# -gt 0 ]; do
      case "$1" in -t) shift 2 ;; -l) literal=1; shift ;; *) break ;; esac
    done
    [ "$literal" = 1 ] && printf '%s' "${1:-}" >> "$FM_SEND_LOG"
    exit 0 ;;
  display-message)
    for a in "$@"; do case "$a" in *cursor_y*) printf '1\n'; exit 0 ;; esac; done
    printf 'fakepane\n'; exit 0 ;;
  capture-pane)
    if [ "${FM_FAKE_TMUX_COMPOSER:-}" = pending ]; then
      printf '╭──────────────╮\n│ leftover txt │\n╰──────────────╯\n'
    else
      printf '╭────╮\n│    │\n╰────╯\n'
    fi
    exit 0 ;;
  list-windows) exit 0 ;;
esac
exit 0
SH
cat > "$FB/sleep" <<'SH'
#!/usr/bin/env bash
exit 0
SH
chmod +x "$FB/tmux" "$FB/sleep"

STATE="$HOME_DIR/state"
printf 'window=sess:fm-t1\nkind=ship\n' > "$STATE/t1.meta"
printf 'needs-decision [key=api-shape]: pick REST or RPC for the mapping endpoint\n' > "$STATE/t1.status"
printf 'working: kept busy on an unrelated stream\n' >> "$STATE/t1.status"

hr() { printf '\n----- %s -----\n' "$*"; }
drain() { FM_STATE_OVERRIDE="$STATE" "$ROOT/bin/fm-wake-drain.sh" 2>/dev/null; }

printf '=========================================================================\n'
printf ' DEFECT 1 - the captain answers a decision the worker never sees\n'
printf ' tree under test: %s\n' "$LABEL"
printf '=========================================================================\n'

hr '1. the worker has an open captain decision; what a supervisor reads'
drain

hr '2. the captain answers it: fm-send.sh t1 --resolve-key api-shape "go with REST"'
printf '   (the worker'\''s composer visibly holds pending text, so the\n'
printf '    doorbell fail-safe must skip the keystroke - that part is correct)\n\n'
: > "$WORK/send.log"
env PATH="$FB:$PATH" FM_FAKE_TMUX_COMPOSER=pending FM_GATE_REFUSE_BYPASS=1 \
  FM_ROOT_OVERRIDE="$HOME_DIR" FM_HOME="$HOME_DIR" FM_SEND_LOG="$WORK/send.log" FM_SEND_SETTLE=0 \
  "$ROOT/bin/fm-send.sh" t1 --resolve-key api-shape "go with REST" 2>&1
printf '   exit=%s\n' "$?"
printf '   keystrokes actually typed into the worker pane: %s bytes\n' "$(wc -c < "$WORK/send.log" | tr -d ' ')"
printf '   durable inbox record on disk: %s\n' "$(ls "$STATE"/t1.inbox/*.msg 2>/dev/null | xargs -n1 basename 2>/dev/null | tr '\n' ' ')"
printf '   worker acknowledgements (handled/): %s\n' "$(ls "$STATE"/t1.inbox/handled/*.msg 2>/dev/null | wc -l | tr -d ' ') - the worker has NOT seen it"

hr '3. the task status log right after the answer was sent'
nl -ba "$STATE/t1.status"

hr '4. what the supervisor now reads (OPEN DECISIONS fold, fm-wake-drain.sh)'
drain
CLOSED_EARLY=0
printf '\n>>> VERDICT: '
if drain | grep -qF '[key=api-shape]'; then
  if drain | grep -qF 'still unread by the worker'; then
    printf 'decision still OPEN and marked as delivered-but-unread. CORRECT.\n'
  else
    printf 'decision still open (no unread annotation).\n'
  fi
else
  CLOSED_EARLY=1
  printf 'decision reads as ANSWERED although the worker never saw it. THE DEFECT.\n'
fi

hr '5. the worker finally reads its inbox and acknowledges the record'
mkdir -p "$STATE/t1.inbox/handled"
for f in "$STATE"/t1.inbox/*.msg; do [ -e "$f" ] && mv "$f" "$STATE/t1.inbox/handled/"; done
printf '   moved to handled/: %s\n' "$(ls "$STATE"/t1.inbox/handled/*.msg | xargs -n1 basename | tr '\n' ' ')"
printf '   the next supervision poll commits any parked closure:\n'
FM_STATE_OVERRIDE="$STATE" bash -c '
  . "$1"
  if declare -F fm_task_inbox_commit_resolutions >/dev/null; then
    fm_task_inbox_commit_resolutions "$2" t1 "$2/t1.status" && echo "   commit_resolutions: ok"
  else
    echo "   (this tree has no acknowledgement-gated closure to commit)"
  fi
' _ "$ROOT/bin/fm-task-inbox-lib.sh" "$STATE"

hr '6. the task status log after the acknowledgement'
nl -ba "$STATE/t1.status"

hr '7. what the supervisor reads once the worker has actually read the answer'
out=$(drain)
if [ -n "$out" ]; then printf '%s\n' "$out"; else printf '(nothing - no open decisions)\n'; fi
printf '\n>>> VERDICT: '
if printf '%s' "$out" | grep -qF '[key=api-shape]'; then
  printf 'still open after acknowledgement - the answer never closed it.\n'
elif [ "$CLOSED_EARLY" = 1 ]; then
  printf 'closed - but it had ALREADY read as closed back at step 4, before delivery.\n'
else
  printf 'the decision closed exactly when the worker acknowledged it. CORRECT.\n'
fi
printf '\n'
Evidence: Reproduction script for the defect 2 transcript

Source: Reproduction script for the defect 2 transcript

#!/usr/bin/env bash
# End-to-end demonstration of Defect 2: "a healthy waiting job reported as
# failed".
#
# Drives the REAL bin/fm-crew-state.sh of the tree passed in $1 - the helper the
# watcher and the captain-facing fleet views read - over real throwaway git
# repos, a fake `no-mistakes` CLI serving real-shaped run rows, and a fake tmux
# pane that is genuinely busy.
set -u
ROOT=$1
LABEL=$2
WORK=$(mktemp -d "${TMPDIR:-/tmp}/fmdemo2.XXXXXX")
FB="$WORK/fakebin"; mkdir -p "$FB"

cat > "$FB/no-mistakes" <<'SH'
#!/usr/bin/env bash
set -u
case "${1:-}" in
  axi) shift
    case "${1:-}" in
      status) shift
        if [ "${1:-}" = --run ]; then printf '%s\n' "${FM_FAKE_AXI_STATUS_RUN:-}"
        else printf '%s\n' "${FM_FAKE_AXI_STATUS:-}"; fi ;;
      logs) printf '%s\n' "${FM_FAKE_CI_LOGS:-}" ;;
    esac ;;
  runs) printf '%s\n' "${FM_FAKE_RUNS_LIST:-}" ;;
esac
exit 0
SH
cat > "$FB/tmux" <<'SH'
#!/usr/bin/env bash
set -u
case "${1:-}" in
  display-message) printf '%%1\n' ;;
  capture-pane)
    if [ "${FM_FAKE_BUSY:-0}" = 1 ]; then printf 'reviewing the mapping leaf filter\nesc to interrupt\n'
    else printf 'all quiet\n> \n'; fi ;;
esac
exit 0
SH
chmod +x "$FB/no-mistakes" "$FB/tmux"
export FM_FAKE_BUSY=1 FM_FAKE_CI_LOGS="" FM_FAKE_AXI_STATUS_RUN=""

printf '=========================================================================\n'
printf ' DEFECT 2 - a healthy, actively validating crew reported as failed\n'
printf ' tree under test: %s\n' "$LABEL"
printf '=========================================================================\n'

# ---------------------------------------------------------------------------
# scenario <name> <branch>: builds a fresh crew worktree + state, busy pane.
setup() {
  local name=$1 branch=$2
  CASE="$WORK/$name"; STATE="$CASE/state"; WT="$CASE/wt"
  mkdir -p "$STATE" "$WT"
  git -C "$WT" init -q
  git -C "$WT" -c user.name=fmtest -c user.email=fmtest@example.invalid commit -q --allow-empty -m init
  git -C "$WT" checkout -q -b "$branch"
  printf 'window=fm:fm-%s\nworktree=%s\nkind=ship\nharness=claude\n' "$name" "$WT" > "$STATE/$name.meta"
  printf 'working: validating\n' > "$STATE/$name.status"
  FM_STATE_OVERRIDE="$STATE" "$ROOT/bin/fm-busy-event.sh" arm "$STATE" "$name" >/dev/null
}

report() {  # <id> <scenario title>
  local id=$1 title=$2 out
  printf '\n--- what a captain/watcher actually reads: bin/fm-crew-state.sh %s ---\n' "$id"
  out=$(PATH="$FB:$PATH" FM_STATE_OVERRIDE="$STATE" "$ROOT/bin/fm-crew-state.sh" "$id" 2>&1)
  printf '%s\n' "$out" | sed 's/^/    /'
  printf '\n>>> VERDICT (%s): ' "$title"
  if printf '%s' "$out" | grep -q 'state: failed'; then
    printf 'a live, busy, validating crew is reported FAILED off a superseded run. THE DEFECT.\n'
  elif printf '%s' "$out" | grep -q 'state: working'; then
    printf 'the live crew reads as working, sourced from the pane. CORRECT.\n'
  else
    printf 'unexpected verdict.\n'
  fi
}

# ===========================================================================
printf '\n\n#########################################################################\n'
printf '# SCENARIO A - the crew moved past the live run'\''s head\n'
printf '#########################################################################\n'
BR=fm/switav-mapping-leaf-filter
setup moved "$BR"
OLDER=$(git -C "$WT" rev-parse --short=8 HEAD)
git -C "$WT" -c user.name=fmtest -c user.email=fmtest@example.invalid commit -q --allow-empty -m "one more commit after the run started"
NOW=$(git -C "$WT" rev-parse --short=8 HEAD)
export FM_FAKE_AXI_STATUS="run:
  id: \"01OTHER\"
  branch: fm/other-crew
  status: running
  head: \"deadbee1\"
  pr: \"\"
  findings: none"
export FM_FAKE_RUNS_LIST="  running    $BR $OLDER  2026-08-27 19:02
  failed     $BR $NOW  2026-08-27 17:17"
printf '\ncrew branch        : %s\n' "$BR"
printf 'crew worktree HEAD : %s\n' "$NOW"
printf 'crew pane          : busy ("esc to interrupt") - the agent is working right now\n'
printf 'crew status log    : %s\n' "$(cat "$STATE/moved.status")"
printf '\n`no-mistakes axi status` (repo-wide answer; belongs to ANOTHER crew):\n'
printf '%s\n' "$FM_FAKE_AXI_STATUS" | sed 's/^/    /'
printf '\n`no-mistakes runs --limit N` (newest first):\n'
printf '%s\n' "$FM_FAKE_RUNS_LIST" | sed 's/^/  /'
printf '    ^ row 1: this branch'\''s LIVE run, pinned at %s - the crew has since\n' "$OLDER"
printf '      committed on top, so this head resolves but is no longer current.\n'
printf '    ^ row 2: an OLDER, superseded FAILED run whose head is exactly the\n'
printf '      worktree'\''s HEAD %s, so it binds perfectly. It is the trap.\n' "$NOW"
report moved "scenario A"

# ===========================================================================
printf '\n\n#########################################################################\n'
printf '# SCENARIO B - this branch'\''s own live run reports a head this worktree\n'
printf '#              cannot resolve (no-mistakes commits review fixes inside\n'
printf '#              its OWN repo mirror under ~/.no-mistakes/repos/<id>.git)\n'
printf '#########################################################################\n'
BR2=fm/feat-mismatch
setup mismatch "$BR2"
SHORT=$(git -C "$WT" rev-parse --short=8 HEAD)
export FM_FAKE_AXI_STATUS="run:
  id: \"01LIVE\"
  branch: $BR2
  status: running
  head: \"305b0969\"
  pr: \"\"
  findings: none"
export FM_FAKE_RUNS_LIST="  failed     $BR2 $SHORT  2026-08-27 17:17"
printf '\ncrew branch        : %s\n' "$BR2"
printf 'crew worktree HEAD : %s\n' "$SHORT"
printf 'crew pane          : busy ("esc to interrupt") - the agent is working right now\n'
printf 'crew status log    : %s\n' "$(cat "$STATE/mismatch.status")"
printf '\n`no-mistakes axi status` (this branch'\''s OWN live run, head 305b0969\n'
printf 'which does not exist in this worktree):\n'
printf '%s\n' "$FM_FAKE_AXI_STATUS" | sed 's/^/    /'
printf '\n`no-mistakes runs --limit N`:\n'
printf '%s\n' "$FM_FAKE_RUNS_LIST" | sed 's/^/  /'
printf '    ^ the only row is an OLDER FAILED run at the worktree'\''s own HEAD.\n'
report mismatch "scenario B"
printf '\n'

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 2 infos
  • ℹ️ bin/fm-test-run.sh:1201 - This change gives tests/fm-captain-hold-lifecycle.test.sh a brand-new direct dependency on bin/fm-task-inbox-lib.sh (the new run_ack_and_commit_resolutions helper sources the library and calls fm_task_inbox_commit_resolutions, and test_chat_channel_feeds_the_same_keyed_answer_intake now asserts state: queued before the ack and state: done after), but the bin/fm-task-inbox-lib.sh arm of families_for_changed_path - edited in this same change - was not extended to select it. That suite's family is pure-contract-unit, which this arm does not emit, and the families_for_unmapped_bin fallback does not apply because this path is explicitly mapped. Concretely: a later edit confined to fm_task_inbox_defer_resolution or fm_task_inbox_commit_resolutions that breaks the deferral (for example closing at enqueue again) would leave the selected suites green while fm-captain-hold-lifecycle.test.sh - the suite that pins that exact contract for the captain-hold channel - is never run. Fix is the same mechanical one already accepted for the wake-drain suite last round: add printf &#39;%s\n&#39; __script__:fm-captain-hold-lifecycle.test.sh beside the existing __script__ entry, rather than widening to the whole pure-contract-unit family.
  • ℹ️ bin/fm-teardown.sh:241 - teardown_inbox_closures derives its undelivered list from fm_task_inbox_pending_resolutions, which reports a sidecar as pending whenever its record is in neither the inbox root nor handled/. That is exactly the orphan case (the worker rm'd its record instead of moving it). If teardown runs before the watcher's next poll surfaces that orphan, the sidecar is still unsurfaced, so it lands in this loop and teardown prints undelivered at cleanup: &lt;id&gt; never acknowledged the answer to decision key(s) &lt;k&gt; (record 001.msg); that decision stays open - naming a record that no longer exists and telling firstmate the worker never saw an answer it almost certainly did read and act on. bin/fm-watch.sh's own orphan wake says the opposite for the identical state ("which the worker most likely did read"). Trace: write 001.msg + .001.resolve for key k, rm 001.msg, run teardown with no intervening watcher poll -> the message above. The two states are distinguishable here without new machinery: fm_task_inbox_resolution_record gives the bound record, and "still present in the inbox root" means genuinely undelivered while "absent from both root and handled/" means the orphan case. Flagging rather than patching because the wording is user-facing supervision text and the captain has stopped fix rounds on this machinery.

🔧 Fix: select captain-hold suite on task-inbox library changes
2 infos still open:

  • ℹ️ bin/fm-task-inbox-lib.sh:111 - The file header's commit contract says the idempotence check is against the status log's text - "a status key whose exact closing line is already in the status log is not appended again" - but the implementation deliberately does the opposite. _fm_task_inbox_resolution_is_committed (line 495) greps the per-sidecar ledger &lt;sidecar&gt;.committed, and both of the more detailed doc blocks in this same file explicitly forbid the status-log-text reading: _fm_task_inbox_resolution_committed_path's header (lines 484-489) says "Scoped to THIS sidecar, never to the status log's text ... a global 'does this exact line already exist anywhere in the status log' check cannot tell the two apart and silently orphans the reopened decision", and fm_task_inbox_commit_resolutions's own block (lines 668-670) repeats it. tests/fm-task-inbox.test.sh's test_reopened_key_with_identical_answer_closes_again exists precisely to fail if anyone implements what the header describes: a key legitimately reopened and re-answered with identical wording (routine for a short answer like "fix them all" on a repeated step key) would never close again. Since this file declares itself the ONE owner of the closure contract, a maintainer reading the header first is being pointed at the exact bug the test forbids. Fix is the same one-line kind already accepted for committed-ledger-comment-overstates: state that the dedupe is per-sidecar identity, not status-log text.
  • ℹ️ bin/fm-test-run.sh:1112 - This change adds new behavior to bin/fm-wake-drain.sh - the delivered-but-unread annotation in print_open_decisions_section (bin/fm-wake-drain.sh:270-289) - whose only coverage is the three cases added to tests/fm-wake-drain-open-decisions.test.sh. But bin/fm-wake-drain.sh matches the bin/fm-watch*|bin/fm-wake*|... arm here, which emits only watcher-wake-lock, and that family (family_for_basename lines 218-226) lists fm-wake-drain-unread-status.test.sh but not fm-wake-drain-open-decisions.test.sh; the latter is unmapped, so family_for_basename returns unclassified and it is never selected. Concretely: a later edit confined to bin/fm-wake-drain.sh that broke the annotation (for example dropping the unread_keys lookup, or capping the note away) would leave the selected suites green while the only suite pinning that behavior never runs. The same change already recognized this dependency in the other direction, adding __script__:fm-wake-drain-open-decisions.test.sh to the bin/fm-task-inbox-lib.sh arm at line 1203. Fix is the same mechanical one: add that __script__ entry to this arm too, rather than widening watcher-wake-lock.
⚠️ **Test** - 1 warning
  • ⚠️ tests/fm-teardown.test.sh:1696 - Three suites fail locally on this macOS machine, and all three fail IDENTICALLY on the base commit c7fdef9 (origin/main) with unmodified code, so none is a regression from this change: tests/fm-teardown.test.sh (herdr-preflight-missing-adapter: teardown continued without its required preflight), tests/fm-captain-hold-lifecycle.test.sh (aborts at test_bound_channel_answers_close_at_answer_time with the fixture channel captured no result to feed), and tests/fm-gotmp.test.sh (teardown did not remove the tasktmp dir). Root cause of the teardown one, reduced: macOS system /bin/bash is 3.2.57, where under set -e a failed . (source) builtin terminates the shell with status 0 even when guarded by || return 1 inside an if condition, so teardown_herdr_require_prerequisites never reaches its error path. The repository's authoritative CI lanes run on ubuntu-latest (bash 5); .github/workflows/ci.yml's macos-stock-bash job deliberately runs only bash -n parsing plus two snapshot suites under bash 3.2, not these. Because the suites abort at the first failure, this change's own new cases inside them never ran in the full-suite pass; I ran each in isolation and all pass. Flagging so you can decide whether the local bash-3.2 gap is worth its own piece of work - it is out of scope here (the change touches none of bin/fm-backend.sh, bin/backends/*, or the herdr preflight code path).
  • bash bin/fm-test-run.sh tests/fm-send-resolve-key.test.sh - 16/16 pass, including test_undelivered_answer_never_reads_as_resolved
  • bash bin/fm-test-run.sh tests/fm-task-inbox.test.sh tests/fm-crew-state.test.sh tests/fm-wake-drain-open-decisions.test.sh tests/fm-teardown.test.sh tests/fm-captain-hold-lifecycle.test.sh - task-inbox, crew-state and wake-drain-open-decisions all exit 0; teardown and captain-hold exit 1 on pre-existing local failures
  • bash bin/fm-test-run.sh tests/fm-gotmp.test.sh tests/fm-remote-transport-lanes.test.sh - remote-transport-lanes exit 0; gotmp exit 1 on a pre-existing local failure
  • Manual E2E, defect 1: real bin/fm-send.sh t1 --resolve-key api-shape &#39;go with REST&#39; under FM_FAKE_TMUX_COMPOSER=pending, then real bin/fm-wake-drain.sh, run against both the base tree (c7fdef9) and the worktree (fefdf60), asserting the status log, the inbox record, handled/, and the OPEN DECISIONS fold before and after the worker's acknowledgement
  • Manual E2E, defect 2: real bin/fm-crew-state.sh &lt;id&gt; over real throwaway git repos, a fake no-mistakes serving real-shaped axi status / runs --limit rows and a busy pane, for both false-failure shapes, run against base and target
  • Manual E2E, test selection: bash bin/fm-test-run.sh --list --changed --base HEAD over a one-line edit confined to bin/fm-task-inbox-lib.sh, in isolated git snapshots of base and target
  • Reversion check: reverted only fm_send_defer_resolved_keys back to fm_send_close_resolved_keys/fm_send_feed_resolved_holds in an isolated copy, then bash tests/fm-send-resolve-key.test.sh -> not ok - an unacknowledged answer must not read as resolved
  • Reversion check: reverted only nm_runs_status_for_branch and the own-branch bind in bin/fm-crew-state.sh in an isolated copy, then bash tests/fm-crew-state.test.sh -> not ok - a resolvable-but-mismatched newest run must not be resolved by an older run that binds (unexpected: &#39;state: failed&#39;)
  • Evidence capture: real bin/fm-watch.sh polls producing the wake-queue text for the composer-blocked, absolute-unhandled-bound and orphaned-closure escalation causes
  • Evidence capture: real bin/fm-teardown.sh task-x1 stdout with one acknowledged and one unacknowledged closure parked in the inbox
  • Pre-existing-failure isolation: bash /tmp/nm-base-tree/tests/fm-teardown.test.sh, .../fm-captain-hold-lifecycle.test.sh, .../fm-gotmp.test.sh at base commit c7fdef9 - identical failures
  • Isolated re-runs of the change's own new cases in the blocked suites: test_teardown_commits_acknowledged_closure_and_names_undelivered_one, test_teardown_names_a_quietly_retried_closure_it_cannot_commit, test_chat_channel_feeds_the_same_keyed_answer_intake, test_teardown_skips_gracefully_without_tasktmp, test_teardown_skips_gracefully_when_dir_missing - all pass
  • Scope check: git diff --name-only c7fdef9..fefdf60 | grep &#39;^\.github/&#39; returns nothing, and none of reverted 4bdde97's paths (.github/workflows/no-mistakes-required.yml, CONTRIBUTING.md, bin/fm-pr-body-settled.sh) appear in the net diff
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

tknguyen29032002 and others added 12 commits August 28, 2026 20:22
Two defects made a stopped worker read as a moving one.

A steer answering a keyed decision closed that decision at enqueue time,
but the local doorbell is deliberately skipped whenever the composer
holds pending text, so the answer could sit unread while every durable
record reported it answered. The closure is now parked beside its inbox
record and committed at the worker's acknowledgement, so the decision
stays open until it is actually read. The typed plane still closes on
confirmed submit and the remote plane on remote enqueue, where a
pending-reply expectation already backstops it.

fm-crew-state walked back to an older run for the same branch when the
newest one could not be bound to the worktree's code. A live run's
review commits live in no-mistakes' own repo mirror, so they often do
not resolve in the crew worktree, while the previous failed run's head
frequently is the worktree's HEAD and binds perfectly. Only a branch's
newest run may describe its current state; an unbindable newest run now
attributes nothing and falls through to the pane.

Regression tests for both fail on reintroduction.
Two defects made a stopped worker read as a moving one.

A steer answering a keyed decision closed that decision at enqueue time,
but the local doorbell is deliberately skipped whenever the composer
holds pending text, so the answer could sit unread while every durable
record reported it answered. The closure is now parked beside its inbox
record and committed at the worker's acknowledgement, so the decision
stays open until it is actually read. The typed plane still closes on
confirmed submit and the remote plane on remote enqueue, where a
pending-reply expectation already backstops it.

fm-crew-state walked back to an older run for the same branch when the
newest one could not be bound to the worktree's code. A live run's
review commits live in no-mistakes' own repo mirror, so they often do
not resolve in the crew worktree, while the previous failed run's head
frequently is the worktree's HEAD and binds perfectly. Only a branch's
newest run may describe its current state; an unbindable newest run now
attributes nothing and falls through to the pane.

Regression tests for both fail on reintroduction.
Two defects made a stopped worker read as a moving one.

A steer answering a keyed decision closed that decision at enqueue time,
but the local doorbell is deliberately skipped whenever the composer
holds pending text, so the answer could sit unread while every durable
record reported it answered. The closure is now parked beside its inbox
record and committed at the worker's acknowledgement, so the decision
stays open until it is actually read. The typed plane still closes on
confirmed submit and the remote plane on remote enqueue, where a
pending-reply expectation already backstops it. Recovery no longer waits
on a spent attempt budget alone: a proven composer-blocked skip and an
absolute unhandled-age bound each escalate on their own stated schedule,
so an unacknowledged instruction cannot sit silently behind a busy pane.

fm-crew-state walked back to an older run for the same branch when the
newest one could not be bound to the worktree's code. Rebasing onto main
picked up PR kunchenguid#3194's fix for most of this shape (an unresolvable newest
row now stops the coarse scan instead of walking onto an older terminal
row, and an active pipeline-owned run binds without head equality).
Reconciling against it found one gap it left open: when axi status
already answers for the crew's own branch but the head does not bind and
the run is not pipeline-owned-active, the code still fell through to the
coarse runs list, which can independently rediscover an older superseded
row for that same branch. Closed by never consulting the coarse list once
axi status has already answered for this branch - it can only re-find the
same run or an older, superseded one, never a better one.

Regression tests for both fail on reintroduction, confirmed by reverting
each fix in isolation.
Round 2 of the review's three ask-user findings, applied by hand after the
pipeline agent that was mid-fix died (recovered from the gate mirror's
last committed head, 7739dd9, per the documented recovery procedure).

The commit-idempotence check deduped a status key's closing line against
the WHOLE status log's text, so a decision legitimately reopened under the
same key and answered identically a second time was silently skipped as
already-closed and stayed open forever. Scoped to a per-sidecar committed-
identity marker instead, so idempotence protects only a retry of the same
sidecar's own earlier append.

A worker that rm's an acknowledged record instead of moving it into
handled/ (a contract violation, but not literally prevented) left its
sidecar bound to a record in neither place - a third state the ladder's
.msg-based scan can never see, so the closure never committed and never
escalated. The ladder now detects this and escalates once, independent of
pane busy state, naming the violation and the record.

Corrected the blocked-bound timing doc to what the code actually does
(~1 grace period plus one poll interval, not ~2 grace periods); the timing
itself is unchanged, since escalating on the very next poll after the
first skipped attempt is the intended loud-over-quiet behavior.

Regression tests for all three fail on reintroduction, confirmed by
reverting each fix in isolation.
An orphaned decision closure (its inbox record removed instead of moved to handled/) is now set aside under handled/orphaned/ once its escalation marker is durable, so it stops reading as a pending answer and cannot annotate a later, genuine reopening of the same key as already delivered. The watcher's orphan escalation names the orphan's own keys as something to close by hand and no longer appends the wait-for-acknowledgement suffix that contradicted it. A committed closure that cannot be filed under handled/ now respects the already-surfaced quiet flag like the other two commit failure paths, so a permanently unwritable handled/ surfaces once instead of re-waking firstmate on every poll. The inbox layout inventory documents the per-sidecar committed ledger, the orphan escalation marker, and the orphaned quarantine.
…eardown

- fm_task_inbox_commit_resolutions hands back exactly the closure names it
  failed on for the first time (.commit-failed), and
  fm_task_inbox_record_commit_escalated folds only those into
  .commit-escalated, so a closure acknowledged after the pass is never
  marked surfaced before its own first attempt.
- _fm_task_inbox_orphaned_sidecar skips names already in .orphan-escalated,
  so a surfaced orphan whose retirement failed cannot shadow a newer one.
- fm-teardown.sh runs the same closure-commit pass before removing the
  task inbox: an acknowledged answer lands its resolved line first, and an
  unacknowledged one is named as undelivered at cleanup without refusing.
- Regression tests for all three, each verified to fail on the pre-fix code.
@greptile-apps

greptile-apps Bot commented Aug 29, 2026 •

Copy link
Copy Markdown

Confidence Score: 5/5

The PR appears safe to merge.

No blocking failure remains.

Reviews (5): Last reviewed commit: "no-mistakes(document): name teardown and..." | Re-trigger Greptile

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 971b1d4419

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread bin/fm-task-inbox-lib.sh
fm_cap_line_var "resolved [key=$k]: answered: $note"
line=$FM_LINE_CAP_LINE
append_rc=0
fm_wake_status_append_self_announced "$state" "$status_file" "$line" || append_rc=$?

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Make the closure append and receipt crash-atomic

Captain, if the watcher exits after this status append succeeds but before the .committed receipt is written on line 707, the next pass treats the closure as uncommitted and appends it again. More seriously, if the worker reopens the same key during that interval, the retried stale resolved line appears after the new needs-decision and incorrectly closes the new decision; persist a transaction identity with the status event or otherwise make recovery able to recognize that this sidecar's append already landed.

Useful? React with 👍 / 👎.

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: first-time fork CI approved after a diff review (no .github/workflows writes). Runs 33231621276 (CI) and 33231808669 / 33231621299 (Require no-mistakes).

This replaces the specified local-inbox close-at-enqueue path with acknowledgement-gated close, and adds always-on escalation bounds (including a 900s unhandled-age fire while busy). That is a new unconfigured default, not a restore of a broken specified path. Pipeline tests were skipped; twelve known findings are parked for captain judgment. Not merge-ready. Not a captain-decision flag yet (CI still settling; author still owns attestation and the parked findings). Waiting on CI, then the author, not the captain.

VISION.md per-rule:

  • One captain, one interface: aligns (OPEN DECISIONS stays honest that an unread answer is still open) / cannot tell on the new busy-age escalation noise.
  • Authority is explicit and never inferred: aligns (no new merge/destructive grant).
  • Scripts own the mechanics, agents own the judgment: aligns (sidecar/ladder is script-owned).
  • A restart is a non-event: aligns (closures parked on disk until ack).
  • Delegation with a spine: aligns.
  • The fleet outlives any vendor: aligns.
  • Scope: aligns (command-layer send/watch/teardown).

contract-class: new-default (replaces specified close-at-enqueue; new always-on 900s busy escalation). No auto-merge.

…osure

fm-gotmp.test.sh's fake FM_HOME never symlinked fm-task-inbox-lib.sh (and
the fm-line-cap-lib.sh it sources), a new sibling fm-teardown.sh now sources
unconditionally to commit or name deferred decision closures at cleanup;
without it the fixture's teardown subprocess failed to source and exited
non-zero.

fm-captain-hold-lifecycle.test.sh's chat-channel test asserted the old
close-at-send-time contract on fm-send.sh's local steering-inbox plane. That
contract is what this branch's fix corrects: the closure now parks beside
the durable record and only commits once the worker acknowledges it. Updated
the test to assert the decision stays open immediately after send, simulate
the worker's acknowledgement and the watcher's commit pass, then assert the
close - matching the corrected provenance wording the deferred commit path
records.
…iability

# Conflicts:
#	docs/captain-hold-lifecycle.md
@tknguyen29032002 tknguyen29032002 changed the title fix(bin): gate decision closure on acknowledgement and fix stale run attribution fix(bin): gate keyed decision closure on acknowledgement, bind runs to the newest row Aug 30, 2026
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: newer-activity re-triage after HEAD moved. Fresh fork runs for HEAD 2fbf8327768af5002b4eef9b2204de93daee7a19 were action_required; after a diff review (workflow-zero; no .github/workflows writes; author not flagged) I approved CI 33285897752 and Require no-mistakes 33285897764 / 33285935869.

Attestation MATCH (head_sha 2fbf8327768af5002b4eef9b2204de93daee7a19 = HEAD; review/test/document completed in the raise). CI 33285897752 is in_progress. NM: 33285935869 SUCCESS (edited); 33285897764 FAILURE on synchronize with empty PR_* env (looks like an input race — live body still carries the attestation). Soft overlap only with #3159 (shared fm-teardown.sh / decision-hold surface; different problem). Not merge-ready. Waiting on CI, not the captain. Captain-flag only once CI + NM are green and this is otherwise ready except the new-default decision.

VISION.md per-rule (inspected HEAD vs main headers + bin/fm-send.sh, bin/fm-task-inbox-lib.sh, bin/fm-crew-state.sh, bin/fm-watch.sh, bin/fm-wake-drain.sh, bin/fm-teardown.sh):

  • One captain, one interface: aligns (OPEN DECISIONS keeps an unread answer open and names it) / cannot tell on always-on busy-pane overdue wakes as presentation load.
  • Authority is explicit and never inferred: aligns (no new merge/destructive grant; --resolve-key still captain-owned).
  • Scripts own the mechanics, agents own the judgment: aligns (sidecar/ladder/commit stay in scripts).
  • A restart is a non-event: aligns (.NNN.resolve + durable markers on disk until ack).
  • Delegation with a spine: aligns (close gated on acknowledgement contract).
  • The fleet outlives any vendor: aligns (no vendor/UI coupling).
  • Scope: aligns (command-layer send/watch/teardown/crew-state).

contract-class: new-default. Main's fm-send.sh header still specifies inbox-plane close at ENQUEUE time; this PR replaces that with acknowledgement-gated defer/commit and adds always-on bounded escalation (FM_TASK_INBOX_BLOCKED_MAX / FM_TASK_INBOX_UNHANDLED_MAX_SECS=900) plus newest-row-or-nothing run attribution. That is not a restore of a broken specified path. No auto-merge.

This reverts commit 4bdde97.

That commit was produced autonomously by the pipeline's CI auto-fix step. Its
diagnosis was correct - on a synchronize event the gate judges the event
payload's PR body, which is a snapshot taken before the pr step writes the
head-bound attestation into it - but the change does not belong on this branch:

- This branch's scope is two supervision reliability defects in bin/. A change
  to .github/workflows/no-mistakes-required.yml, a new bin/fm-pr-body-settled.sh,
  and CONTRIBUTING.md are none of them.
- It edits the verification surface that judges every pull request in this repo,
  from a fork pull request, which is a shape that needs deliberate maintainer
  review rather than arriving inside an unrelated change.
- The maintainer's approval on this PR was given explicitly on the basis of no
  .github/workflows writes.

The race it found is real and is being tracked as its own piece of work so it
gets reviewed on its own merits.
@tknguyen29032002 tknguyen29032002 changed the title fix(bin): gate keyed decision closure on acknowledgement, bind runs to the newest row fix(bin): correct false answered-decision and failed-run supervision reports Aug 30, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants