Skip to content

fix(bin): report wake drain presentation failures on stdout - #34

Merged
rub-a-dub-dub merged 7 commits into
mainfrom
fm/firstmate-drain-false-incomplete-label-recovery
Sep 27, 2026
Merged

rub-a-dub-dub merged 7 commits into
mainfrom
fm/firstmate-drain-false-incomplete-label-recovery

Conversation

@rub-a-dub-dub

Copy link
Copy Markdown
Owner

Intent

Captain's words, 2026-09-26: Can we dispatch the two Firstmate PRs on Codex? - authorizing this queued item. Earlier, at PR 28 on 2026-09-23, the captain said ship now, which deferred these cases here so that PR stopped growing one path per review round.

Redesign, as one piece, how bin/fm-wake-drain.sh reports its own failures. Known cases from the PR 28 review runs:

  1. notice-contradicts-already-printed-sections: when print_status_sections has computed and printed every section but the trailing status_commit_presentation_snapshot receipt write fails, the drain still prints STATUS PRESENTATION INCOMPLETE over correct sections.
  2. lock-failure-leaves-stdout-silent, the more important case: when fm_lock_acquire_wait_bounded fails with anything other than 124, print_status_presentation writes only to stderr and returns 1, leaving nothing on stdout below the wake rows - no sections and no incomplete notice.
  3. stale-nonempty-fold-set-printed-as-authoritative: a span-read failure on a task whose trusted open-decision set is non-empty still returns rc 0 with that stale set, so a resolved decision can print as open or a new one stay hidden.
  4. annotation-cursor-failure-drops-a-live-row-silently: the || continue in fm-wake-lib.sh's annotation pass drops a live task's annotation on a cursor-read hiccup with no incomplete notice.
    Goal: one consistent contract - every drain either presents the computed sections or states plainly on stdout what it could not compute and why, and never mislabels a complete presentation as incomplete.

What Changed

  • print_status_presentation in bin/fm-wake-drain.sh now routes every failure that leaves a drain with nothing to present — lock acquisition, fleet snapshot read, and the annotation pass — through a single stdout notice that names the failed operation, instead of the lock case writing only to stderr and returning 1; such a drain skips the acknowledge and commit passes so no presentation cursor advances. A separate print_status_receipt_failure_notice covers the case where all prepared section bytes reached stdout but status_commit_presentation_snapshot failed, so a complete presentation is no longer relabeled STATUS PRESENTATION INCOMPLETE, and print_status_sections passes a per-failure reason into the incomplete notice.
  • status_open_decisions_incremental in bin/fm-classify-lib.sh drops the trusted_open replay path: an ident, stat, size, or span-read failure now returns 1 without printing the persisted open-decision set (leaving it untouched on disk for a later recovery call), rather than returning rc 0 with a possibly stale non-empty set.
  • fm_wake_unread_events in bin/fm-wake-lib.sh returns 2 for a span it could not read versus 1 for a span with nothing unread, and fm_wake_print_annotations returns 1 when a live, readable status file's cursor read fails or its span read returns 2 — replacing the || continue that dropped a live row's annotation silently — while a missing, unreadable, or symlinked status path is still skipped without being called a failure. docs/architecture.md records the resulting contract, and tests/fm-wake-queue.test.sh plus tests/fm-wake-drain-open-decisions.test.sh add cases for the receipt-failure label, the stale fold set, and the annotation cursor/span failures holding every cursor.

Risk Assessment

⚠️ Medium: All four intent cases are implemented with behavioral, self-validating tests and the cursor/acknowledge invariants hold under tracing, but one notice can still contradict annotations already printed on a multi-task drain, and the new all-or-nothing rule widens the blast radius of a single transient per-task read failure to the whole fleet presentation (the captain's explicit design choice), so it is safe to merge with the wording addressed as a follow-up.

Testing

I read the diff to pin the four intent cases, then validated each one end-to-end by driving the real bin/fm-wake-drain.sh over crafted state fixtures with the same injected failure against both the base build (d32fe93) and the target build (8618dcd), capturing the captain-facing stdout as before/after CLI transcripts: a receipt-only failure now keeps its computed sections and reports a distinct RECEIPT FAILED notice instead of labelling them incomplete; a non-124 lock failure now speaks on stdout instead of stderr alone; a stale non-empty open-decision set is no longer printed as authoritative; and an annotation cursor failure yields exactly one generic notice, moves no presentation cursor, and the following drain replays the annotation that never printed. On the automated side I ran every test in tests/fm-wake-drain-open-decisions.test.sh, the new and adjacent annotation and presentation-lock tests in tests/fm-wake-queue.test.sh, and full clean passes of the three neighbouring drain suites plus the two classify suites that exercise the changed fold - all pass. I also ran the eight new or reworded tests against a base-commit checkout, where each fails with its own intended assertion, confirming they are genuine regression tests. One environmental problem surfaced and is documented rather than fixed: this host intermittently SIGSEGVs LC_ALL=C /usr/bin/stat (~5% of forks under the current load), which fails individual drains and makes whole-file suite runs flake at random tests - it reproduces identically at the base commit, so it predates the change; I worked around it by retrying per test.

Evidence: Intent case 1 - receipt-only failure no longer mislabels a complete presentation (base vs target drain stdout)

Source: Intent case 1 - receipt-only failure no longer mislabels a complete presentation (base vs target drain stdout)

Intent case 1: a receipt-only failure must not mislabel a complete presentation as incomplete
---------------------------------------------------------------------------------------------

print_status_sections printed every computed section; only the trailing status_commit_presentation_snapshot receipt write failed.

Same fixture, same injected failure, two builds of bin/fm-wake-drain.sh.

========== BEFORE  (base d32fe93) ==========
$ bin/fm-wake-drain.sh   # presentation receipt write fails (mv of .status-presentation-cursor)
OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line):
task [key=route] needs-decision: choose the presentation route
OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'
STATUS PRESENTATION INCOMPLETE: unread status, outcome backstop, OPEN DECISIONS, and record divergence could not be fully computed this drain (a status log or cursor read/write failed); do not read this drain's silence as nothing open or unread - retry on the next drain.

========== AFTER   (HEAD 8618dcd) ==========
$ bin/fm-wake-drain.sh   # presentation receipt write fails (mv of .status-presentation-cursor)
OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line):
task [key=route] needs-decision: choose the presentation route
OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'
STATUS PRESENTATION RECEIPT FAILED: the status sections above were fully computed and printed, but their presentation receipt could not be committed; they may repeat on the next drain.
Evidence: Intent case 2 - lock-acquire failure now reported on stdout (base vs target, stdout and stderr shown separately)

Source: Intent case 2 - lock-acquire failure now reported on stdout (base vs target, stdout and stderr shown separately)

BEFORE (base d32fe93) stdout: <wake row only> stderr: wake drain: status presentation lock could not be acquired safely AFTER (HEAD 8618dcd) stdout: <wake row> + STATUS PRESENTATION INCOMPLETE: status presentation lock could not be acquired safely; no status annotations or fleet-wide status sections were computed this drain, nothing was marked as seen and no presentation cursor advanced, so every unread status line is still unread.

Intent case 2: a non-124 presentation-lock failure must not leave stdout silent
-------------------------------------------------------------------------------

fm_lock_acquire_wait_bounded fails on a malformed lock. Before: stdout has nothing below the wake rows, only a stderr line. After: the captain reads the failure on stdout.

Same fixture, same injected failure, two builds of bin/fm-wake-drain.sh.

========== BEFORE  (base d32fe93) ==========
$ bin/fm-wake-drain.sh   # presentation lock is malformed, acquire fails (not 124 timeout)
--- stdout (what the captain reads) ---
1790537744	1	signal	task.status	signal: /private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-evidence.9w5VVJ/c2-base-1/state/task.status
--- stderr (diagnostics) ---
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 1 --recovery-generation 45841.1790537744.raKd1g
wake drain: status presentation lock could not be acquired safely

========== AFTER   (HEAD 8618dcd) ==========
$ bin/fm-wake-drain.sh   # presentation lock is malformed, acquire fails (not 124 timeout)
--- stdout (what the captain reads) ---
1790537746	1	signal	task.status	signal: /private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-evidence.9w5VVJ/c2-head-1/state/task.status
STATUS PRESENTATION INCOMPLETE: status presentation lock could not be acquired safely; no status annotations or fleet-wide status sections were computed this drain, nothing was marked as seen and no presentation cursor advanced, so every unread status line is still unread.
--- stderr (diagnostics) ---
WAKE_ACK_REQUIRED: after handling completes run bin/fm-wake-drain.sh --ack-through 1 --recovery-generation 46553.1790537747.GRUKmL
Evidence: Intent case 3 - a stale non-empty open-decision set is no longer printed as authoritative

Source: Intent case 3 - a stale non-empty open-decision set is no longer printed as authoritative

BEFORE: OPEN DECISIONS … firstmate-reconcile-fork-with-upstream [key=old-choice] needs-decision: choose the old route (already resolved in the log; the new [key=new-choice] decision is hidden) AFTER: STATUS PRESENTATION INCOMPLETE: unread status, outcome backstop, OPEN DECISIONS, and record divergence could not be fully computed this drain (a status log, cursor, or prepared-output write could not be completed); …

Intent case 3: a stale non-empty open-decision set must not print as authoritative
----------------------------------------------------------------------------------

The fold cursor is parked before a span that resolves [key=old-choice] and opens [key=new-choice], and that span read fails. Before: the stale set prints as OPEN DECISIONS - the resolved decision shows as open and the new one stays hidden. After: no decision is claimed and the drain says what it could not compute.

Same fixture, same injected failure, two builds of bin/fm-wake-drain.sh.

========== BEFORE  (base d32fe93) ==========
$ bin/fm-wake-drain.sh   # fold span read fails while the persisted open set is NON-empty and stale
#   (log now says: resolved [key=old-choice] + needs-decision [key=new-choice])
OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line):
firstmate-reconcile-fork-with-upstream [key=old-choice] needs-decision: choose the old route
OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'

========== AFTER   (HEAD 8618dcd) ==========
$ bin/fm-wake-drain.sh   # fold span read fails while the persisted open set is NON-empty and stale
#   (log now says: resolved [key=old-choice] + needs-decision [key=new-choice])
STATUS PRESENTATION INCOMPLETE: unread status, outcome backstop, OPEN DECISIONS, and record divergence could not be fully computed this drain (a status log, cursor, or prepared-output write could not be completed); do not read this drain's silence as nothing open or unread - retry on the next drain.
Evidence: Intent case 4 - annotation failure is never silent and no cursor advances (includes the next drain for both builds)

Source: Intent case 4 - annotation failure is never silent and no cursor advances (includes the next drain for both builds)

BEFORE: drain prints no notice and no annotation; the NEXT drain shows nothing - 'working: live row must not disappear silently' is lost. AFTER: drain prints one STATUS PRESENTATION INCOMPLETE: a supplemental status annotation could not be computed …; the NEXT drain prints 'wake annotation: unread wake-EVENT since last drain, not current state: task.status: working: live row must not disappear silently'.

Intent case 4: an annotation cursor failure must not silently drop a live row
-----------------------------------------------------------------------------

A live task annotation cursor read fails. Before: the annotation is dropped with no notice and the presentation cursor still advances, so the working: line is lost for good. After: one notice on stdout, no cursor advances, and the next drain replays the annotation.

Same fixture, same injected failure, two builds of bin/fm-wake-drain.sh.

========== BEFORE  (base d32fe93) ==========
# (host stat SIGSEGV flake: 2 attempt(s) discarded)
$ bin/fm-wake-drain.sh   # annotation presentation-cursor read fails for a live task
1790537768	1	signal	task.status	signal: /private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-evidence.9w5VVJ/c4-base-3/state/task.status
UNREAD STATUS (new since last drain, not re-printed after this presentation):
task note: the captain is still owed this one

$ bin/fm-wake-drain.sh   # the very next drain, no injected failure
1790537768	1	signal	task.status	signal: /private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-evidence.9w5VVJ/c4-base-3/state/task.status

========== AFTER   (HEAD 8618dcd) ==========
# (host stat SIGSEGV flake: 1 attempt(s) discarded)
$ bin/fm-wake-drain.sh   # annotation presentation-cursor read fails for a live task
1790537775	1	signal	task.status	signal: /private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-evidence.9w5VVJ/c4-head-2/state/task.status
STATUS PRESENTATION INCOMPLETE: a supplemental status annotation could not be computed; no status annotations or fleet-wide status sections were computed this drain, nothing was marked as seen and no presentation cursor advanced, so every unread status line is still unread.

$ bin/fm-wake-drain.sh   # the very next drain, no injected failure
1790537775	1	signal	task.status	signal: /private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-evidence.9w5VVJ/c4-head-2/state/task.status
wake annotation: unread wake-EVENT since last drain, not current state: task.status: working: live row must not disappear silently
wake annotation: latest wake-EVENT observed at drain, not current state: task.status: note: the captain is still owed this one
UNREAD STATUS (new since last drain, not re-printed after this presentation):
task note: the captain is still owed this one
Evidence: Targeted test results, per test, plus the same new tests failing at the base commit

Source: Targeted test results, per test, plus the same new tests failing at the base commit

Targeted test results - fm/firstmate-drain-false-incomplete-label-recovery
base d32fe93 -> target 8618dcd

Each test was driven through the real bin/fm-wake-drain.sh over crafted state fixtures.
This host intermittently SIGSEGVs 'LC_ALL=C /usr/bin/stat' (see host-flake-stat-sigsegv.txt),
which makes any single drain fail spuriously; attempt counts below are that host flake,
not the change. The same flake fails the same suites at the base commit.

== tests/fm-wake-drain-open-decisions.test.sh (all 16 tests, run per test) ==
PASS  test_no_open_decisions_prints_nothing                          (attempts: 1)
PASS  test_unrelated_task_read_failure_reports_incomplete_not_silent_empty (attempts: 1)
PASS  test_torn_down_task_wake_row_does_not_blank_the_sections       (attempts: 2)
PASS  test_partial_snapshot_does_not_truncate_the_cursor_manifest    (attempts: 1)
PASS  test_untrusted_fold_cursor_read_failure_is_not_a_silent_empty  (attempts: 1)
PASS  test_trusted_empty_fold_cursor_read_failure_is_not_a_silent_empty (attempts: 1)
PASS  test_trusted_nonempty_fold_cursor_read_failure_is_not_authoritative (attempts: 2)
PASS  test_receipt_failure_does_not_relabel_printed_sections_incomplete (attempts: 1)
PASS  test_status_symlink_is_not_followed                            (attempts: 1)
PASS  test_buried_decision_still_surfaces                            (attempts: 1)
PASS  test_over_long_decision_note_is_capped_with_a_marker           (attempts: 1)
PASS  test_explicit_resolution_closes_it                             (attempts: 1)
PASS  test_later_unrelated_terminal_line_does_not_close_it           (attempts: 1)
PASS  test_reserved_key_namespace_is_owned_by_its_library            (attempts: 1)
PASS  test_open_decision_surfaces_even_with_an_unrelated_queued_wake (attempts: 2)
PASS  test_buried_decision_surfaces_on_the_empty_queue_fast_path     (attempts: 1)

== tests/fm-wake-queue.test.sh (new + adjacent annotation/lock tests) ==
PASS  test_malformed_presentation_lock_reports_acquire_failure       (attempts: 1)
PASS  test_annotation_cursor_failure_is_reported_on_stdout           (attempts: 1)
PASS  test_annotation_span_read_failure_is_reported_and_retried      (attempts: 1)
PASS  test_unattributed_annotation_failure_holds_every_cursor        (attempts: 1)
PASS  test_annotation_failure_holds_sibling_cursors_too              (attempts: 1)
PASS  test_snapshot_failure_reports_the_uncomputed_annotations       (attempts: 1)
PASS  test_historical_annotation_skips_announced_status              (attempts: 1)
PASS  test_structural_signal_enrichment_preserves_raw_rows           (attempts: 1)
PASS  test_enrichment_preserves_all_unread_lines_and_status_file_failures (attempts: 1)
PASS  test_slow_annotation_does_not_block_append_and_deleted_file_fails_open (attempts: 1)
PASS  test_self_held_lock_reclaims_instead_of_deadlocking            (attempts: 1)
PASS  test_subshell_lock_ownership_without_bashpid                   (attempts: 1)
PASS  test_bounded_lock_handoff_after_contention                     (attempts: 1)
PASS  test_live_presentation_holder_is_deadlined_without_weakening_ack (attempts: 1)

== neighbouring suites over the same presentation machinery (whole file) ==
PASS tests/fm-wake-drain-unread-status.test.sh        (13 tests)
PASS tests/fm-wake-drain-outcome-backstop.test.sh     (18 tests)
PASS tests/fm-wake-drain-open-decisions-cursor.test.sh (7 tests)
PASS tests/fm-classify-decision-key.test.sh           (25 tests)
PASS tests/fm-classify-corr-token.test.sh             (12 tests)

== the same new tests run against the BASE commit (regression proof) ==
Each fails at d32fe93 with its own intended assertion, and passes at 8618dcd:

test_trusted_nonempty_fold_cursor_read_failure_is_not_authoritative
  -> OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'
test_receipt_failure_does_not_relabel_printed_sections_incomplete
  -> STATUS PRESENTATION INCOMPLETE: unread status, outcome backstop, OPEN DECISIONS, and record divergence could not be fully computed this drain (a status log or cursor read/write failed); do not read this drain's silence as nothing open or unread - retry on the next drain.
test_malformed_presentation_lock_reports_acquire_failure
  -> not ok - malformed presentation lock did not report its acquire failure on stdout
test_annotation_cursor_failure_is_reported_on_stdout
  -> not ok - the annotation cursor failure owed exactly one notice, got 0: 1790538074	1	signal	task.status	signal: /private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-wake-tests.1RnSpE/annotation-cursor-failure/state/task.status
test_annotation_span_read_failure_is_reported_and_retried
  -> not ok - the failed annotation span read remained silent on stdout: 1790538081	1	signal	task.status	signal: /private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-wake-tests.0BLoru/annotation-span-read-failure/state/task.status
test_unattributed_annotation_failure_holds_every_cursor
  -> not ok - the unattributed annotation failure remained silent on stdout: 1790538085	1	signal	task.status	signal: /private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-wake-tests.jEzOuL/unattributed-annotation-failure/state/task.status
test_annotation_failure_holds_sibling_cursors_too
  -> not ok - the sibling cursor failure owed exactly one notice, got 0: 1790538126	1	signal	alpha.status	signal: /private/var/folders/f_/hcy8g0q91p9fdyghn_9ypx5r0000gn/T/fm-wake-tests.eyx6I4/annotation-sibling-hold/state/alpha.status
test_snapshot_failure_reports_the_uncomputed_annotations
  -> (no stable result in 4 attempts)
test_snapshot_failure_reports_the_uncomputed_annotations
  -> not ok - the snapshot failure said nothing about the annotations it skipped
Evidence: Host flake write-up: /usr/bin/stat SIGSEGV reproducer and proof it predates the change

Source: Host flake write-up: /usr/bin/stat SIGSEGV reproducer and proof it predates the change

$ for i in $(seq 1 300); do bash -c 's=$(LC_ALL=C /usr/bin/stat -f "%z" /tmp/probe2.status 2>/dev/null); echo $?'; done | sort | uniq -c 284 0 16 139

Host flake, NOT a product defect: /usr/bin/stat SIGSEGVs under load on this machine
=====================================================================================

Darwin 27.0.0, load average ~14-25 (43 concurrent agent processes during this run).

Reproducer (no firstmate code involved) - stat prints the CORRECT value and then
exits 139 (128+SIGSEGV) in roughly 5-7% of forks:

$ for i in $(seq 1 300); do bash -c 's=$(LC_ALL=C /usr/bin/stat -f "%z" /tmp/probe2.status 2>/dev/null); echo $?'; done | sort | uniq -c
 284 0
  16 139

Every firstmate read that feeds the presentation snapshot goes through that call
(_fm_status_file_size / _fm_open_decisions_file_ident in bin/fm-classify-lib.sh), so a
segfaulting stat makes status_presentation_snapshot fail and the drain correctly print
'STATUS PRESENTATION INCOMPLETE: status snapshot could not be read' - the new contract
reporting a real (host-induced) read failure. Measured directly:

  status_presentation_snapshot over 4 static status files: 8 failures / 120 calls

The same flake fails the same suites at the BASE commit d32fe93 (4 runs of
tests/fm-wake-drain-open-decisions.test.sh at base: 3 random failures, 1 clean), so it
predates this change. It also breaks tests/lib.sh at source time, because
FM_TEST_OWNER_IDENTITY comes from 'LC_ALL=C ps' in fm_pid_identity and dies the same way.
- Outcome: ⚠️ 1 warning across 1 run (57m47s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 2 issues (1 warning, 1 info)
  • ⚠️ bin/fm-wake-drain.sh:414 - For a task the annotation pass held, the UNREAD STATUS header's promise is false. print_unread_status_section receives only $snapshot (bin/fm-wake-drain.sh:646 passes $annotation_held to print_status_sections but scan_unread_surface_snapshot at bin/fm-classify-lib.sh:1663 has no knowledge of it), so it prints the held task's unread note: lines under 'UNREAD STATUS (new since last drain, not re-printed after this presentation)' while status_acknowledge_presented_snapshot (bin/fm-classify-lib.sh:1417-1436) sets hold=true, skips the span scan, and forces endpoint=$offset - the cursor does not move, so the very next drain re-prints those same lines. The change's own fixture proves the sequence: test_annotation_cursor_failure_is_reported_on_stdout (tests/fm-wake-queue.test.sh:1868) primes the cursor, appends working: live row must not disappear silently plus note: the captain reads this one through the fleet-wide section, fails the annotation's first cat of .status-presentation-cursor (read at bin/fm-classify-lib.sh:1123), and then asserts the note: line IS on stdout (tests/fm-wake-queue.test.sh:1922) while asserting the span is retried on the next drain. So the note is printed under a 'not re-printed' header and is guaranteed to be re-printed. No data is lost - the direction is duplication - but a captain-facing section states a guarantee this drain then breaks, which is exactly the mislabeling the intent's 'never mislabels' goal targets. Needs the author's call because every remedy is a user-facing behavior choice: narrowest form is to pass the held list into print_unread_status_section/scan_unread_surface_snapshot and omit held tasks from that section for the drain (the STATUS ANNOTATION INCOMPLETE notice already tells the captain those bytes are coming back); alternatives are softening the header for any drain that held a cursor, or holding only the non-unread-surface part of the span.
  • ⚠️ bin/fm-wake-drain.sh:627 - The STATUS ANNOTATION INCOMPLETE notice promises 'every status line those annotations would have carried stays unread - retry on the next drain', but the held set is in-memory for one drain only (FM_WAKE_ANNOTATION_HELD at bin/fm-wake-lib.sh:2247, read at bin/fm-wake-drain.sh:628), and the retry depends on the wake row still being queued. Concrete ordinary sequence, using the change's own fixture shape: drain N holds task task (cursor read fails), prints the notice, and leaves the cursor at the previous offset; the drain then prints WAKE_ACK_REQUIRED on stderr, so the next thing the caller does is --ack-through &lt;seq&gt;, which removes the row and exits at bin/fm-wake-drain.sh:801 without ever calling print_status_presentation; the following drain takes the empty-queue path (bin/fm-wake-drain.sh:803-820) and calls print_status_presentation with no rows, so the annotation pass never runs, annotation_held is empty, and status_acknowledge_presented_snapshot applies the plain fleet rule - the span [held offset, endpoint) contains note: the captain reads this one through the fleet-wide section, which status_line_is_unread_surface accepts, so safe=true and the cursor jumps to the endpoint, past the working: line whose only surface was the annotation. That line is then never annotated and the promised retry never happens. The pre-existing fleet rule already accepts dropping routine lines once an unread-surface line in the same span is presented, so the defect introduced here is the notice's claim, not the drop. Classified ask-user because the smallest honest remedy is a choice between extending the change with durable state (persisting the held set across drains so the retry is real) and narrowing the notice's wording to what the drain actually guarantees - the remedy, not the defect, is what needs authorization.
  • ⚠️ bin/fm-wake-drain.sh:635 - The prior fix round introduced if [ &#34;$annotation_rc&#34; -eq 0 ] || [ -n &#34;$annotation_held&#34; ] to gate computing fully_presented. The || [ -n &#34;$annotation_held&#34; ] disjunct is the part no intent requirement needs, and it keeps a gap open: fm_wake_print_annotations has a failure return that names no task while held is already non-empty - printf &#39;%s\n&#39; &#34;$line&#34; || return 1 at bin/fm-wake-lib.sh:2407 aborts the manifest loop. With task A already in FM_WAKE_ANNOTATION_HELD (cursor or span read failure) and task B's annotation write then failing, annotation_held is non-empty so fully_presented is computed from the raw manifest as every direct task - including B, whose annotation write failed, and every task after B that the aborted loop never reached. status_acknowledge_presented_snapshot then sets safe=true for those tasks and advances their cursors past spans the annotation never printed, which is precisely the loss the hold mechanism exists to prevent and makes the notice's 'stays unread' claim false. Reachability is a double fault (a transient read failure plus a stdout write failure with stdout on a different filesystem than $STATE, since SIGPIPE kills the drain outright and an ENOSPC $STATE fails the earlier mktemp), so I am not claiming this is common. The point is that the narrower gate if [ &#34;$annotation_rc&#34; -eq 0 ] satisfies the intent on its own and closes the gap structurally: on any annotation failure no task is claimed fully presented, every cursor falls back to the unread-surface rule, and held tasks are still frozen by the hold branch. The only cost is that a fully annotated sibling task may replay its annotation on the next drain - duplication, not loss. Recommend removing the disjunct rather than adding a third repair to this machinery.

🔧 Fix: omit held unread lines, narrow annotation notice and gate
3 issues (2 warnings, 1 info) still open:

  • ⚠️ bin/fm-wake-drain.sh:647 - When the fleet snapshot read fails, the annotation pass is skipped entirely and no STATUS ANNOTATION INCOMPLETE notice is printed, so absent annotations read as "nothing unread" under the contract this change establishes. Concrete sequence: task T's presentation cursor is at O and working: rebasing onto upstream has been appended since O; a direct signal: row for T is queued. Concurrently task M is torn down, so _fm_open_decisions_file_ident fails mid-loop in status_presentation_snapshot (bin/fm-classify-lib.sh:1035) and it returns 1. The drain prints T's raw wake row, then only STATUS PRESENTATION INCOMPLETE: status snapshot could not be read. (bin/fm-wake-drain.sh:647), sets snapshot= and rc=1; the [ &#34;$rc&#34; -eq 0 ] &amp;&amp; [ -n &#34;$rows&#34; ] gate at bin/fm-wake-drain.sh:650 then skips fm_wake_print_annotations outright. Stdout carries no wake annotation: line and no STATUS ANNOTATION INCOMPLETE. This change made those two the complete set of outcomes for the annotation pass - the lock branch at bin/fm-wake-drain.sh:635 spells out "no status annotations or fleet-wide status sections were computed this drain", and the annotation branch at bin/fm-wake-drain.sh:655 names its own - so the snapshot branch is the one place that goes silent about annotations. The captain reads T as having nothing unread while a working: line whose only surface is that annotation went unpresented, which is the intent's "never mislabels" goal. Remedy is a user-facing wording/behavior choice: call print_annotation_incomplete_notice &#39;&#39; on this branch when $rows is non-empty, or extend the snapshot notice to state that annotations were not computed either.
  • ⚠️ bin/fm-wake-drain.sh:669 - The fix round's comments and its new test assert an invariant the code does not have. bin/fm-wake-drain.sh:669 states "Every task the annotation pass could not compute is held, so its presentation cursor does not advance" and bin/fm-wake-drain.sh:661 states "the cost is a replayed annotation, never a dropped span". Neither holds for the two failure paths that name no task: fm_wake_print_annotations's manifest awk (bin/fm-wake-lib.sh:2333, || return 1) and its per-line printf &#39;%s\n&#39; &#34;$line&#34; || return 1 (bin/fm-wake-lib.sh:2408). Concrete sequence: task T's span since its cursor holds working: rebasing onto upstream followed by note: answered the captain, with a direct signal row queued. The manifest awk fails - exactly the fault test_unattributed_annotation_failure_holds_every_cursor injects. FM_WAKE_ANNOTATION_HELD is empty, so annotation_held=&#39;&#39; and fully_presented=&#39;&#39;; print_status_sections still runs (bin/fm-wake-drain.sh:674) and status_acknowledge_presented_snapshot takes the hold=false branch (bin/fm-classify-lib.sh:1417), finds note: is an unread surface, sets safe=true, and commits T's cursor at the snapshot endpoint - past the working: line the annotation never printed. That line has no other surface and is dropped permanently. The new test only exercises a working:-only span, which the pre-existing fleet unread-surface rule holds on its own, so it cannot detect the mixed-span gap its pass message ("an annotation failure that names no task holds every presentation cursor") claims to cover; it is a valid regression test for the narrowed fully_presented gate, but not for the stated invariant. Needs the author's call because the two remedies differ in kind: correct the comments and the test's claim to what the fleet rule actually guarantees, or hold every manifest task when the failure names none - the latter extends the hold machinery rather than correcting it, and the prior round's chosen remedy (dropping the || [ -n &#34;$annotation_held&#34; ] disjunct) deliberately accepted the fleet-rule fallback.
  • ℹ️ bin/fm-wake-drain.sh:580 - The named-hold notice tells the reader each held task's status log "is readable at its path", but one of the two hold triggers is precisely a failed read of that log. A task enters FM_WAKE_ANNOTATION_HELD either from the cursor-read failure at bin/fm-wake-lib.sh:2369 (the status log itself is fine - only the shared .status-presentation-cursor manifest failed) or from the rc-2 span-read failure at bin/fm-wake-lib.sh:2392. rc 2 originates in fm_wake_unread_events's perl sysopen(... O_RDONLY | O_NOFOLLOW) or exit 1 (bin/fm-wake-lib.sh:2270), reached only after the pre-guard at bin/fm-wake-lib.sh:2364 already confirmed -f, -r and not-a-symlink, so rc 2 there means the file became unopenable or an I/O error hit between those two points. bin/fm-wake-drain.sh:580 then prints ... each status log is readable at its path: /state/&lt;task&gt;.status, pointing the captain at a file the drain just failed to read - in the one code path whose stated purpose is to never mislabel. Smallest honest remedy is wording: give the path without asserting readability ("their status logs are at: ..."), which reads correctly for the cursor-failure case too.

🔧 Fix: collapse presentation failures into one notice, hold every cursor
2 issues (1 warning, 1 info) still open:

  • ⚠️ bin/fm-wake-drain.sh:569 - The one generic notice can claim no annotations were computed while some are already on stdout above it. fm_wake_print_annotations prints each task's annotation lines as it walks the manifest (bin/fm-wake-lib.sh:2408) and only accumulates incomplete=1 for the task that failed (bin/fm-wake-lib.sh:2393), returning 1 at the end. Concrete sequence: tasks alpha and bravo both have primed presentation cursors and a queued direct signal: row. status_presentation_snapshot succeeds. alpha annotates cleanly, so wake annotation: unread wake-EVENT since last drain, not current state: alpha.status: working: ... is already on stdout. bravo.status is rotated/recreated (shorter) between the snapshot and bravo's span read, so fm_wake_unread_events's perl guard exit 1 unless ... $end &lt;= $size fails (bin/fm-wake-lib.sh:2284) -> rc 2 -> incomplete=1 -> return 1. print_status_presentation takes the reason branch at bin/fm-wake-drain.sh:640 and prints, directly below alpha's annotation: STATUS PRESENTATION INCOMPLETE: a supplemental status annotation could not be computed; no status annotations or fleet-wide status sections were computed this drain, nothing was marked as seen and no presentation cursor advanced, so every unread status line is still unread. The second and third clauses are true (print_status_sections is skipped, so no cursor moves), but "no status annotations ... were computed this drain" is contradicted by the visible line above it - the same mislabel class the intent forbids ("never mislabels a complete presentation as incomplete"), reintroduced at the annotation level rather than the section level. The failure mode the comment at bin/fm-wake-lib.sh:2387 names ("a file that disappears, rotates, or becomes unreadable after the snapshot") is exactly this path, so it is a real sequence, not a hypothetical. This needs the author's call because the single generic wording was the captain's own explicit instruction for this round; the narrowest honest remedy is wording only - keep one notice with no per-task naming, but state that SOME annotations could not be computed and that any annotation printed above may be partial, rather than that none were computed.
  • ℹ️ bin/fm-wake-drain.sh:336 - One computation failure inside the same presentation pass still resolves to silence with no notice: status_snapshot_latest_event &#34;$STATE/$task.status&#34; &#34;$endpoint&#34; &#34;$ident&#34; || continue. That helper returns 1 both for a genuine read failure (its _fm_status_read_span at bin/fm-classify-lib.sh:1078, or its inner perl at :1091) and for the documented benign deferrals (endpoint 0, size/ident changed mid-read, latest line crossing the 64 KiB bound), so the || continue cannot tell "could not read this task's latest event" from "not applicable this drain". The task is then absent from STATUS OUTCOME BACKSTOP with no STATUS OUTCOME BACKSTOP SKIPPED line, which under this change's contract reads as "computed, nothing uncovered". Nothing is lost: the task's backstop receipt is not advanced, so the next drain retries, and the benign deferral is by far the common case (any log that grew since the snapshot takes it). Recording this as an accepted tradeoff rather than a defect: separating a real I/O failure from the documented deferral would mean threading a distinct return code through status_snapshot_latest_event and its callers, which extends the change rather than correcting it, and the intent's four enumerated cases do not name this path.
⚠️ **Test** - 1 warning
  • ⚠️ bin/fm-classify-lib.sh:866 - Host-level flake, not a product defect: on this machine LC_ALL=C /usr/bin/stat prints the correct value and then exits 139 (SIGSEGV) in ~5% of forks under the current load (Darwin 27.0.0, load avg 14-25, 43 concurrent agent processes). Since _fm_status_file_size and _fm_open_decisions_file_ident (bin/fm-classify-lib.sh) run that call for every status file, status_presentation_snapshot fails ~7% of drains (measured: 8 failures / 120 calls over 4 static status files) and the drain correctly prints STATUS PRESENTATION INCOMPLETE: status snapshot could not be read. The result is that any whole-file run of tests/fm-wake-drain-open-decisions.test.sh or tests/fm-wake-queue.test.sh fails at a random test. This predates the change - 4 runs of the suite at base d32fe93 produced 3 random failures and 1 clean pass - and it also aborts tests/lib.sh at source time (FM_TEST_OWNER_IDENTITY comes from LC_ALL=C ps in fm_pid_identity), which surfaces as fm_test_tmproot: command not found. I worked around it by running every relevant test individually with retries; all pass. No code change is warranted here, but local full-suite runs will stay unreliable until the machine's load drops or the OS-level stat crash is resolved.
  • Before/after end-to-end drain transcripts for intent case 1 (fake mv failing the .status-presentation-cursor receipt write): base prints sections + STATUS PRESENTATION INCOMPLETE, target prints sections + STATUS PRESENTATION RECEIPT FAILED and no INCOMPLETE
  • Before/after transcripts for intent case 2 (malformed .status-presentation-lock, FM_STATUS_PRESENTATION_LOCK_TIMEOUT=1): base leaves stdout silent below the wake row with only a stderr line, target prints STATUS PRESENTATION INCOMPLETE: status presentation lock could not be acquired safely… on stdout
  • Before/after transcripts for intent case 3 (fold cursor rewound to a stale non-empty open set, FM_STATUS_SPAN_READER forced to exit 1): base prints the already-resolved [key=old-choice] under OPEN DECISIONS and hides [key=new-choice], target prints no decisions and states what it could not compute
  • Before/after transcripts for intent case 4 (fake cat failing the shared .status-presentation-cursor read once): base drops the annotation silently and the working: line is gone from the next drain, target prints exactly one notice and the next drain replays wake annotation: … working: live row must not disappear silently
  • bash tests/fm-wake-drain-open-decisions.test.sh - all 16 tests exercised individually (retried per test for the host stat flake), including the new test_trusted_nonempty_fold_cursor_read_failure_is_not_authoritative and test_receipt_failure_does_not_relabel_printed_sections_incomplete
  • bash tests/fm-wake-queue.test.sh - new tests test_annotation_cursor_failure_is_reported_on_stdout, test_annotation_span_read_failure_is_reported_and_retried, test_unattributed_annotation_failure_holds_every_cursor, test_annotation_failure_holds_sibling_cursors_too, test_snapshot_failure_reports_the_uncomputed_annotations, the reworded test_malformed_presentation_lock_reports_acquire_failure, plus adjacent test_historical_annotation_skips_announced_status, test_structural_signal_enrichment_preserves_raw_rows, test_enrichment_preserves_all_unread_lines_and_status_file_failures, test_slow_annotation_does_not_block_append_and_deleted_file_fails_open, test_self_held_lock_reclaims_instead_of_deadlocking, test_subshell_lock_ownership_without_bashpid, test_bounded_lock_handoff_after_contention, test_live_presentation_holder_is_deadlined_without_weakening_ack
  • bash tests/fm-wake-drain-unread-status.test.sh (13 tests), bash tests/fm-wake-drain-outcome-backstop.test.sh (18 tests), bash tests/fm-wake-drain-open-decisions-cursor.test.sh (7 tests) - whole-file clean passes over the same presentation machinery
  • bash tests/fm-classify-decision-key.test.sh (25 tests) and bash tests/fm-classify-corr-token.test.sh (12 tests) - whole-file clean passes over the changed status_open_decisions_incremental fold
  • Regression check: the same 8 new/updated tests run against a base-commit checkout of d32fe93 - each fails there with its own intended assertion (e.g. not ok - the annotation cursor failure owed exactly one notice, got 0) and passes at 8618dcd
  • Host-flake isolation: for i in $(seq 1 300); do bash -c &#39;s=$(LC_ALL=C /usr/bin/stat -f &#34;%z&#34; f 2&gt;/dev/null); echo $?&#39;; done | sort | uniq -c -> 16/300 exits 139 (SIGSEGV); status_presentation_snapshot over 4 static status files -> 8 failures / 120 calls; 4 base-commit runs of tests/fm-wake-drain-open-decisions.test.sh -> 3 random failures, 1 clean
⚠️ **Document** - 1 info
  • ℹ️ AGENTS.md:189 - AGENTS.md states twice (line 189 and line 425) that UNREAD STATUS lines "are not re-printed after that presentation", stated absolutely. This change makes the exception a named user-facing outcome: on a receipt-commit failure the drain now prints STATUS PRESENTATION RECEIPT FAILED, which says the fully printed sections "may repeat on the next drain". I deliberately did not hedge either AGENTS.md line: the possibility already existed at the base commit (sections printed, receipt commit failed), so this change did not make the claim newly wrong, the failure mode is duplication rather than loss, the drain announces it on stdout where it happens, and the instruction's load-bearing purpose - read those lines this turn, they are the only surface for them - is unchanged. Recording it as a judgment call in case the captain would rather AGENTS.md carry the caveat; doing so would cost two lines in the always-loaded surface that every session pays for.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@rub-a-dub-dub
rub-a-dub-dub merged commit c2d3d03 into main Sep 27, 2026
36 of 37 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant