Skip to content

fix(bin): hold the wedge alarm while a crew's validation run is progressing - #55

Merged
HelloWorldSungin merged 3 commits into
mainfrom
fm/fm-watch-escalates-during-validation
Aug 5, 2026
Merged

HelloWorldSungin merged 3 commits into
mainfrom
fm/fm-watch-escalates-during-validation

Conversation

@HelloWorldSungin

Copy link
Copy Markdown
Owner

Intent

GOAL (GitHub issue #53): stop firstmate's supervision raising a possible-wedge alarm for a worker that is parked on a validation run which is demonstrably progressing, WITHOUT weakening detection of one that is genuinely stuck. Observed 2026-08-05: one pane produced five consecutive 'stale: ... possible wedge, escalation N' escalations while bin/fm-crew-state.sh reported 'working - source: run-step - validating (running)' for the same task the whole time. The authoritative current-state reader already knew; the escalation path simply never consulted it.

THE HARD CONSTRAINT, AND THE BAR THIS CHANGE MUST BE JUDGED AGAINST. This NARROWS an alarm. The deliverable that earns it is the evidence of what STILL alarms, pinned by its own tests rather than by argument. All three of these must still escalate and each has a dedicated test: (1) a stranded step past the bound - a test step sat nineteen minutes on its opening line with a live but idle agent and was correctly aborted at thirty; (2) an orphaned step that never logged at all - firstmate has recorded a 'ci' step that orphaned and left a run reporting 'running' for days; (3) a worker wedged with no active run whatsoever, which is the alarm's original purpose. So 'status: running' is deliberately NOT accepted as evidence of health anywhere in this change - only observed PROGRESS is.

DELIBERATE DESIGN DECISIONS a reviewer reading only the diff would not know:

  1. The gate is a SUPPRESSOR THAT REQUIRES POSITIVE EVIDENCE, not a filter. bin/fm-run-progress.sh reports exactly three classes: progressing, stranded, none. Only 'progressing' may ever quiet an alarm. Every no-evidence shape - no metadata, torn-down worktree, no-mistakes absent, a bounded read that could not complete, a status answered for another branch, a run parked at a gate, no active_steps rows, an unparseable duration - collapses to 'none', which leaves the escalation path byte-identical to what it was before this change. This is why the change cannot weaken the alarm by accident: a bug in the reader produces 'none', which alarms.

  2. STRIPPING THE PIPELINE'S OWN 'quiet' MARKER RATHER THAN TRUSTING IT IS INTENTIONAL AND MUST SURVIVE REVIEW. no-mistakes prefixes last_activity with 'quiet' once its own step_quiet_warning (10m by default) elapses. That marker is a CLAIM about silence, not a MEASUREMENT of it, and no-mistakes' own guidance calls it 'a liveness clue, not permission to cancel'. Review and test steps routinely emit one opening line and then work silently for ten to eighteen minutes, so treating the marker as a wedge verdict would re-raise exactly the false alarms this change removes. The reader therefore strips the prefix and parses the AGE the field carries. tests/fm-run-progress.test.sh pins this directly (test_pipeline_quiet_marker_is_stripped_not_trusted): a step 14m quiet still reads 'progressing'. Do not 'simplify' this into trusting the marker.

  3. THE THRESHOLD IS A NAMED, DOCUMENTED, CONFIGURABLE VALUE AND THAT IS HALF THE POINT OF THE CHANGE. FM_RUN_STRANDED_SILENCE_SECS defaults to 1800 and its reasoning sits beside it in bin/fm-run-progress.sh, in docs/configuration.md, and in docs/architecture.md: 3x the pipeline's own 10m quiet warning, comfortably past the longest healthy silence observed (18m), and the point at which an operator actually judged a stranded step stranded and aborted it. An unexplained magic number here would be a defect in its own right. Do not inline it as a literal, do not drop the reasoning comment, and do not fold it into another module's constant.

  4. A CONFIDENTLY DEAD AGENT IS NEVER ABSORBED, however well its run is moving, and short-circuits before paying for the read. The pipeline executes its own steps, so a run can keep advancing with nobody left to answer its next gate - that is a wedge, and it is the one shape 'the run is fine' would otherwise hide. Pinned by test_dead_agent_escalates_even_while_its_run_progresses.

  5. THE READ IS AT THE ESCALATION POINT, NOT ON THE POLL PATH. bin/fm-watch.sh documents in its own comments that the poll path must stay cheap and must NOT re-read crew state. bin/fm-run-progress.sh shells out to a bounded 'no-mistakes axi status', so both supervisors consult it only where an alarm would otherwise be raised - at most one bounded call per STALE_ESCALATE_SECS window per pane, rather than one per FM_POLL tick. A hold RESTARTS the timer rather than clearing it, so the next look is a full window away instead of one poll away. The bounded cost of that: a crew that wedges immediately after a hold waits at most one extra window. A delay, never a loss.

  6. ONE POLICY OWNER FOR TWO SUPERVISORS. bin/fm-classify-lib.sh exists precisely so the always-on watcher and the away-mode daemon cannot drift apart on overlapping triage, so crew_wedge_progress lives there and both call it. I checked bin/fm-supervise-daemon.sh deliberately, as the task asked: away mode skips declared pauses in its stale loop, so it dodges the exact shape seen today, but it will still wedge-escalate a mid-run worker that never wrote a 'paused:' line - the same class of bug - so the same fix is applied there, through the same owner. Two new daemon tests pin the hold and what still escalates in away mode.

  7. Firstmate's own bin/fm-crew-state.sh was deliberately NOT extended. It answers 'is there an active run', which is a different question from 'is that run moving', and its single-line output contract ('state: X - source: Y - detail') is parsed by many consumers; widening its detail field would be the fragile choice.

VERIFICATION. The reader was proven against the real thing before any test was written: run read-only against a live mid-pipeline task on no-mistakes v1.41.2, it returned 'progress: progressing - test running, last activity 14m23s ago (silent 863s, bound 1800s)', and 'stranded' with the bound lowered. That evidence, the exact TOON shape parsed, and the refresh trigger are recorded in docs/verification/supervision.md, which is the classified maintainer-verification owner. tests/fm-run-progress.test.sh (17 cases) pins every class portably against that recorded shape, including quoted commas inside last_activity that a naive comma split would mis-align.

TEST-SAFETY DECISION worth noting: tests/wake-helpers.sh installs a hermetic fm-run-progress stub next to the existing wedge-alarm recorder, following that same established pattern, so no suite can reach a real pipeline from a fixture. The stub defaults to 'none', which is why every pre-existing wedge assertion in the watcher and daemon suites still asserts today's alarm unchanged.

SCOPE EXCLUSIONS the user explicitly ruled out of this task, tracked separately, and NOT to be fixed here: the tests/lib.sh fm_test_tmproot EXIT-trap leak; tests/fm-teardown.test.sh emitting no completion marker; and the open upstream-merge PR #44. If PR #44 lands under this branch, reconcile by MERGE, never rebase.

Delivery: firstmate ships this project no-mistakes; the PR body must carry 'Closes #53' and reference #53.

What Changed

  • Added bin/fm-run-progress.sh, a read-only reader that answers whether a crew's validation run is moving (a different question from fm-crew-state.sh's "is there an active run"). It parses the pipeline's own active_steps[].active_for / last_activity from a bounded no-mistakes axi status and reports exactly three classes - progressing, stranded, none - stripping the pipeline's quiet prefix and measuring the age it carries rather than trusting the marker, and separating slow from stranded at the named, documented FM_RUN_STRANDED_SILENCE_SECS (1800s, 3x the pipeline's own 10m quiet warning). Every no-evidence shape (no metadata, torn-down worktree, unbounded read, run parked at a gate, unparseable output) collapses to none.
  • Made crew_wedge_progress in bin/fm-classify-lib.sh the single owner of the wedge policy, and wired both supervisors to it at the escalation point only - never on the poll path, so the bounded read costs at most one call per escalate window per pane. bin/fm-watch.sh and bin/fm-supervise-daemon.sh now hold and restart the timer when the run is progressing, escalate byte-identically on every other answer, short-circuit a confidently dead agent before paying for the read, name the stopped step when a run is stranded, and cap consecutive holds (FM_RUN_PROGRESS_HOLD_MAX / RUN_PROGRESS_HOLD_MAX_DEFAULT, 15) past which the pane escalates anyway carrying the progress detail - a bounded delay that repeats as a check-in cadence, never a lost alarm.
  • Pinned the narrowing with evidence of what still alarms: new tests/fm-run-progress.test.sh covers the reader's parse/threshold/no-evidence contract, plus new watcher wedge-gate cases and away-mode daemon hold/escalate/cap cases; tests/wake-helpers.sh installs a hermetic fm-run-progress stub defaulting to none beside the existing wedge-alarm recorder, so every pre-existing wedge assertion still asserts today's alarm and no fixture can reach a real pipeline. Registered the new script in bin/fm-test-run.sh test selection and documented the gate, its threshold, and the recorded live-pipeline verification shape in docs/architecture.md, docs/configuration.md, docs/scripts.md, and docs/verification/supervision.md.

Closes #53 - #53

Risk Assessment

✅ Low: Every round-1 finding is fully addressed with correct fail-closed semantics, complete state-file reset coverage across both supervisors, and tests pinning the hold, the cap, the forced escalation's progress detail, the count reset, the repeating cadence, and the busy path; the only item left is a mechanical consolidation of one duplicated constant.

Testing

Ran the three suites that own this change (18 reader cases, 60 watcher-triage cases, 104 daemon cases) plus the test-routing and documentation suites that it also touches - all pass. Unit tests alone would not have shown the intent working, so I built an end-to-end harness wiring the real fm-watch.sh to the real fm-run-progress.sh behind a fake no-mistakes printing the exact axi status TOON recorded in docs/verification/supervision.md, with nothing stubbed between the pipeline's status output and the wake queue. Against base commit 5537370's bin/ tree that fixture reproduces the reported bug exactly - five consecutive possible wedge, escalation N wakes escalating to demand-deep-inspection while fm-crew-state reported working - source: run-step - validating (running); on this branch the identical fixture produces zero wakes and five held ... (progressing ...) triage lines. The bar the intent sets holds: a stranded test step, an orphaned ci step, a crew with no active run, and a dead agent whose run is progressing all still escalate through the real watcher, the first two now naming the step that stopped. I also verified the two decisions a reviewer could not confirm from the diff - the pipeline's quiet marker is stripped and its age parsed (14m quiet reads progressing; the same run reads stranded with the bound lowered), and the bounded read sits at the escalation point, measured at one axi status call across ~12 polls rather than one per tick. Away mode enforces the same policy through the shared owner, including the hold cap. No visual artifact: this is a shell supervision path with no rendered surface, so the end-user-visible output is the wake reason line and away-mode digest text, captured verbatim as CLI transcripts. One note on my own harness rather than the product: my first away-mode attempt showed the progressing crew escalating because tests/wake-helpers.sh exports FM_RUN_PROGRESS_BIN at source time and clobbered my override, so the daemon read the hermetic stub's none - the safety seam is genuinely unforgeable from a fixture, which is what it was built for. Worktree left clean; all scratch trees removed.

Evidence: Issue #53 reproduced on the base commit, then fixed on this branch (same fixture, real watcher + real reader)

=== 2. BEFORE (base commit 5537370): the reported bug === A crew parked on a validation run that is demonstrably progressing. fm-crew-state.sh, all five windows: state: working · source: run-step · validating (running) ------------------------------------------------------------------ window 1: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 1) window 2: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 2) window 3: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 3, demand-deep-inspection: same pane has wedge-escalated 3 times in a row - do not re-absorb on the run-step/pane state alone) window 4: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 4, demand-deep-inspection: ...) window 5: stale: test:fm-issue53 (idle 501s, possible wedge, escalation 5, demand-deep-inspection: ...) === 3. AFTER (this change): the same fixture, byte for byte === window 1: (no wake - watcher still supervising) window 2: (no wake - watcher still supervising) window 3: (no wake - watcher still supervising) window 4: (no wake - watcher still supervising) window 5: (no wake - watcher still supervising) wake queue delivered to the captain: (empty above = nothing surfaced) what the watcher logged instead (state/.watch-triage.log): held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 1/15): test:fm-issue53 ... hold 2/15 ... hold 3/15 ... hold 4/15 ... hold 5/15


=== 1. What the operator can now ask directly: is the run MOVING? ===

$ bin/fm-run-progress.sh reader
progress: progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s)
  # the healthy pane from issue #53: quiet for 14m, still moving

$ FM_RUN_STRANDED_SILENCE_SECS=60 bin/fm-run-progress.sh reader
progress: stranded · test running, last activity 14m23s ago (silent 863s, past the 60s bound)
  # same run, bound lowered - the class is a measurement, not a mood

$ bin/fm-run-progress.sh reader
progress: stranded · test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)
  # STILL ALARMS (1): a test step stranded on its opening line past the bound

$ bin/fm-run-progress.sh reader
progress: stranded · ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound)
  # STILL ALARMS (2): a ci step that orphaned and never logged at all

$ bin/fm-run-progress.sh reader
progress: none · no actively-executing step (run parked, terminal, or between steps)
  # STILL ALARMS (3): no executing step - the alarm's original subject

$ bin/fm-run-progress.sh reader   # with the worktree torn down
progress: none · worktree gone (torn down?)
  # every no-evidence shape collapses to none, which alarms


=== 2. BEFORE (base commit 5537370): the reported bug ===

A crew parked on a validation run that is demonstrably progressing.
fm-crew-state.sh, all five windows: state: working · source: run-step · validating (running)
------------------------------------------------------------------
window 1: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 1)
window 2: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 2)
window 3: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 3, demand-deep-inspection: same pane has wedge-escalated 3 times in a row - do not re-absorb on the run-step/pane state alone)
window 4: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 4, demand-deep-inspection: same pane has wedge-escalated 4 times in a row - do not re-absorb on the run-step/pane state alone)
window 5: stale: test:fm-issue53 (idle 501s, possible wedge, escalation 5, demand-deep-inspection: same pane has wedge-escalated 5 times in a row - do not re-absorb on the run-step/pane state alone)
------------------------------------------------------------------
wake queue delivered to the captain:
  1785906262	5	stale	test:fm-issue53	stale: test:fm-issue53 (idle 501s, possible wedge, escalation 5, demand-deep-inspection: same pane has wedge-escalated 5 times in a row - do not re-absorb on the run-step/pane state alone)


=== 3. AFTER (this change): the same fixture, byte for byte ===

------------------------------------------------------------------
window 1: (no wake - watcher still supervising)
window 2: (no wake - watcher still supervising)
window 3: (no wake - watcher still supervising)
window 4: (no wake - watcher still supervising)
window 5: (no wake - watcher still supervising)
------------------------------------------------------------------
wake queue delivered to the captain:
(empty above = nothing surfaced)

what the watcher logged instead (state/.watch-triage.log):
  [2026-08-05T05:04:22+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 1/15): test:fm-issue53
  [2026-08-05T05:04:26+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 2/15): test:fm-issue53
  [2026-08-05T05:04:30+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 3/15): test:fm-issue53
  [2026-08-05T05:04:34+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 4/15): test:fm-issue53
  [2026-08-05T05:04:39+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 5/15): test:fm-issue53


=== 4. AFTER: what STILL escalates ===

(1) a test step stranded past the bound, live but idle agent:
  reader says : progress: stranded · test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)
  watcher wake: stale: test:fm-stranded (idle 500s, possible wedge, escalation 1, validation run stranded: test running, last activity 31m0s ago (silent 1860s, past the 1800s bound))

(2) a ci step that orphaned and never logged, run still 'running' for days:
  reader says : progress: stranded · ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound)
  watcher wake: stale: test:fm-orphan (idle 500s, possible wedge, escalation 1, validation run stranded: ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound))

(3) a worker wedged with no active run whatsoever:
  reader says : progress: none · no actively-executing step (run parked, terminal, or between steps)
  watcher wake: stale: test:fm-norun (idle 500s, possible wedge, escalation 1)

(4) a confidently dead agent, however well its run is moving:
  reader says : progress: progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s)
  watcher wake: stale: test:fm-dead (idle 500s, possible wedge, escalation 1)


=== 5. AFTER: the hold is a bounded delay, never a lost alarm ===

FM_RUN_PROGRESS_HOLD_MAX=3, run reporting the same healthy verdict every window:
  window 1: held
  window 2: held
  window 3: held
  window 4: stale: test:fm-cap (idle 500s, possible wedge, escalation 1, validation run still progressing but this pane has been silent for 3 held windows: test running, last activity 14m23s ago (silent 863s, bound 1800s))
  window 5: held
Evidence: What still escalates: all four shapes the intent requires, through the real watcher

(1) a test step stranded past the bound, live but idle agent: reader says : progress: stranded · test running, last activity 31m0s ago (silent 1860s, past the 1800s bound) watcher wake: stale: test:fm-stranded (idle 500s, possible wedge, escalation 1, validation run stranded: test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)) (2) a ci step that orphaned and never logged, run still 'running' for days: reader says : progress: stranded · ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound) watcher wake: stale: test:fm-orphan (idle 500s, possible wedge, escalation 1, validation run stranded: ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound)) (3) a worker wedged with no active run whatsoever: reader says : progress: none · no actively-executing step (run parked, terminal, or between steps) watcher wake: stale: test:fm-norun (idle 500s, possible wedge, escalation 1) (4) a confidently dead agent, however well its run is moving: reader says : progress: progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s) watcher wake: stale: test:fm-dead (idle 500s, possible wedge, escalation 1) === 5. the hold is a bounded delay, never a lost alarm (FM_RUN_PROGRESS_HOLD_MAX=3) === window 1: held window 2: held window 3: held window 4: stale: test:fm-cap (idle 500s, possible wedge, escalation 1, validation run still progressing but this pane has been silent for 3 held windows: test running, last activity 14m23s ago (silent 863s, bound 1800s)) window 5: held


=== 1. What the operator can now ask directly: is the run MOVING? ===

$ bin/fm-run-progress.sh reader
progress: progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s)
  # the healthy pane from issue #53: quiet for 14m, still moving

$ FM_RUN_STRANDED_SILENCE_SECS=60 bin/fm-run-progress.sh reader
progress: stranded · test running, last activity 14m23s ago (silent 863s, past the 60s bound)
  # same run, bound lowered - the class is a measurement, not a mood

$ bin/fm-run-progress.sh reader
progress: stranded · test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)
  # STILL ALARMS (1): a test step stranded on its opening line past the bound

$ bin/fm-run-progress.sh reader
progress: stranded · ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound)
  # STILL ALARMS (2): a ci step that orphaned and never logged at all

$ bin/fm-run-progress.sh reader
progress: none · no actively-executing step (run parked, terminal, or between steps)
  # STILL ALARMS (3): no executing step - the alarm's original subject

$ bin/fm-run-progress.sh reader   # with the worktree torn down
progress: none · worktree gone (torn down?)
  # every no-evidence shape collapses to none, which alarms


=== 2. BEFORE (base commit 5537370): the reported bug ===

A crew parked on a validation run that is demonstrably progressing.
fm-crew-state.sh, all five windows: state: working · source: run-step · validating (running)
------------------------------------------------------------------
window 1: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 1)
window 2: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 2)
window 3: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 3, demand-deep-inspection: same pane has wedge-escalated 3 times in a row - do not re-absorb on the run-step/pane state alone)
window 4: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 4, demand-deep-inspection: same pane has wedge-escalated 4 times in a row - do not re-absorb on the run-step/pane state alone)
window 5: stale: test:fm-issue53 (idle 501s, possible wedge, escalation 5, demand-deep-inspection: same pane has wedge-escalated 5 times in a row - do not re-absorb on the run-step/pane state alone)
------------------------------------------------------------------
wake queue delivered to the captain:
  1785906262	5	stale	test:fm-issue53	stale: test:fm-issue53 (idle 501s, possible wedge, escalation 5, demand-deep-inspection: same pane has wedge-escalated 5 times in a row - do not re-absorb on the run-step/pane state alone)


=== 3. AFTER (this change): the same fixture, byte for byte ===

------------------------------------------------------------------
window 1: (no wake - watcher still supervising)
window 2: (no wake - watcher still supervising)
window 3: (no wake - watcher still supervising)
window 4: (no wake - watcher still supervising)
window 5: (no wake - watcher still supervising)
------------------------------------------------------------------
wake queue delivered to the captain:
(empty above = nothing surfaced)

what the watcher logged instead (state/.watch-triage.log):
  [2026-08-05T05:04:22+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 1/15): test:fm-issue53
  [2026-08-05T05:04:26+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 2/15): test:fm-issue53
  [2026-08-05T05:04:30+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 3/15): test:fm-issue53
  [2026-08-05T05:04:34+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 4/15): test:fm-issue53
  [2026-08-05T05:04:39+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 5/15): test:fm-issue53


=== 4. AFTER: what STILL escalates ===

(1) a test step stranded past the bound, live but idle agent:
  reader says : progress: stranded · test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)
  watcher wake: stale: test:fm-stranded (idle 500s, possible wedge, escalation 1, validation run stranded: test running, last activity 31m0s ago (silent 1860s, past the 1800s bound))

(2) a ci step that orphaned and never logged, run still 'running' for days:
  reader says : progress: stranded · ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound)
  watcher wake: stale: test:fm-orphan (idle 500s, possible wedge, escalation 1, validation run stranded: ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound))

(3) a worker wedged with no active run whatsoever:
  reader says : progress: none · no actively-executing step (run parked, terminal, or between steps)
  watcher wake: stale: test:fm-norun (idle 500s, possible wedge, escalation 1)

(4) a confidently dead agent, however well its run is moving:
  reader says : progress: progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s)
  watcher wake: stale: test:fm-dead (idle 500s, possible wedge, escalation 1)


=== 5. AFTER: the hold is a bounded delay, never a lost alarm ===

FM_RUN_PROGRESS_HOLD_MAX=3, run reporting the same healthy verdict every window:
  window 1: held
  window 2: held
  window 3: held
  window 4: stale: test:fm-cap (idle 500s, possible wedge, escalation 1, validation run still progressing but this pane has been silent for 3 held windows: test running, last activity 14m23s ago (silent 863s, bound 1800s))
  window 5: held
Evidence: The reader as an operator runs it: the quiet marker is measured, not trusted

$ bin/fm-run-progress.sh reader progress: progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s) # the healthy pane from issue #53: 'quiet 14m23s ago', still moving $ FM_RUN_STRANDED_SILENCE_SECS=60 bin/fm-run-progress.sh reader progress: stranded · test running, last activity 14m23s ago (silent 863s, past the 60s bound) # same run, bound lowered - the class is a measurement, not a mood $ bin/fm-run-progress.sh reader # with the worktree torn down progress: none · worktree gone (torn down?) # every no-evidence shape collapses to none, which alarms


=== 1. What the operator can now ask directly: is the run MOVING? ===

$ bin/fm-run-progress.sh reader
progress: progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s)
  # the healthy pane from issue #53: quiet for 14m, still moving

$ FM_RUN_STRANDED_SILENCE_SECS=60 bin/fm-run-progress.sh reader
progress: stranded · test running, last activity 14m23s ago (silent 863s, past the 60s bound)
  # same run, bound lowered - the class is a measurement, not a mood

$ bin/fm-run-progress.sh reader
progress: stranded · test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)
  # STILL ALARMS (1): a test step stranded on its opening line past the bound

$ bin/fm-run-progress.sh reader
progress: stranded · ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound)
  # STILL ALARMS (2): a ci step that orphaned and never logged at all

$ bin/fm-run-progress.sh reader
progress: none · no actively-executing step (run parked, terminal, or between steps)
  # STILL ALARMS (3): no executing step - the alarm's original subject

$ bin/fm-run-progress.sh reader   # with the worktree torn down
progress: none · worktree gone (torn down?)
  # every no-evidence shape collapses to none, which alarms


=== 2. BEFORE (base commit 5537370): the reported bug ===

A crew parked on a validation run that is demonstrably progressing.
fm-crew-state.sh, all five windows: state: working · source: run-step · validating (running)
------------------------------------------------------------------
window 1: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 1)
window 2: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 2)
window 3: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 3, demand-deep-inspection: same pane has wedge-escalated 3 times in a row - do not re-absorb on the run-step/pane state alone)
window 4: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 4, demand-deep-inspection: same pane has wedge-escalated 4 times in a row - do not re-absorb on the run-step/pane state alone)
window 5: stale: test:fm-issue53 (idle 501s, possible wedge, escalation 5, demand-deep-inspection: same pane has wedge-escalated 5 times in a row - do not re-absorb on the run-step/pane state alone)
------------------------------------------------------------------
wake queue delivered to the captain:
  1785906262	5	stale	test:fm-issue53	stale: test:fm-issue53 (idle 501s, possible wedge, escalation 5, demand-deep-inspection: same pane has wedge-escalated 5 times in a row - do not re-absorb on the run-step/pane state alone)


=== 3. AFTER (this change): the same fixture, byte for byte ===

------------------------------------------------------------------
window 1: (no wake - watcher still supervising)
window 2: (no wake - watcher still supervising)
window 3: (no wake - watcher still supervising)
window 4: (no wake - watcher still supervising)
window 5: (no wake - watcher still supervising)
------------------------------------------------------------------
wake queue delivered to the captain:
(empty above = nothing surfaced)

what the watcher logged instead (state/.watch-triage.log):
  [2026-08-05T05:04:22+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 1/15): test:fm-issue53
  [2026-08-05T05:04:26+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 2/15): test:fm-issue53
  [2026-08-05T05:04:30+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 3/15): test:fm-issue53
  [2026-08-05T05:04:34+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 4/15): test:fm-issue53
  [2026-08-05T05:04:39+0000] held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 5/15): test:fm-issue53


=== 4. AFTER: what STILL escalates ===

(1) a test step stranded past the bound, live but idle agent:
  reader says : progress: stranded · test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)
  watcher wake: stale: test:fm-stranded (idle 500s, possible wedge, escalation 1, validation run stranded: test running, last activity 31m0s ago (silent 1860s, past the 1800s bound))

(2) a ci step that orphaned and never logged, run still 'running' for days:
  reader says : progress: stranded · ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound)
  watcher wake: stale: test:fm-orphan (idle 500s, possible wedge, escalation 1, validation run stranded: ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound))

(3) a worker wedged with no active run whatsoever:
  reader says : progress: none · no actively-executing step (run parked, terminal, or between steps)
  watcher wake: stale: test:fm-norun (idle 500s, possible wedge, escalation 1)

(4) a confidently dead agent, however well its run is moving:
  reader says : progress: progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s)
  watcher wake: stale: test:fm-dead (idle 500s, possible wedge, escalation 1)


=== 5. AFTER: the hold is a bounded delay, never a lost alarm ===

FM_RUN_PROGRESS_HOLD_MAX=3, run reporting the same healthy verdict every window:
  window 1: held
  window 2: held
  window 3: held
  window 4: stale: test:fm-cap (idle 500s, possible wedge, escalation 1, validation run still progressing but this pane has been silent for 3 held windows: test running, last activity 14m23s ago (silent 863s, bound 1800s))
  window 5: held
Evidence: Away mode: same policy through the shared owner, captain-visible digest

crew parked on a progressing validation run -> away-mode digest: (nothing buffered - no digest would reach the captain) STILL ALARMS - a stranded validation step -> away-mode digest: stale persisted 500s (possible wedge, validation run stranded: test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)): sess:fm-stranded-w1 STILL ALARMS - an orphaned 'ci' step, run 'running' for days -> away-mode digest: stale persisted 500s (possible wedge, validation run stranded: ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound)): sess:fm-orphan-w1 STILL ALARMS - no active run at all -> away-mode digest: stale persisted 500s (possible wedge): sess:fm-norun-w1 the hold is capped in away mode too (FM_RUN_PROGRESS_HOLD_MAX=2): window 1: held window 2: held window 3: stale persisted 500s (possible wedge, validation run still progressing but this crew has been silent for 2 held windows: test running, last activity 14m23s ago (silent 863s, bound 1800s)): sess:fm-cap-w1


=== AWAY MODE: the same policy, through the same owner ===

crew parked on a progressing validation run -> away-mode digest:
  (nothing buffered - no digest would reach the captain)

STILL ALARMS - a stranded validation step -> away-mode digest:
  stale persisted 500s (possible wedge, validation run stranded: test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)): sess:fm-stranded-w1

STILL ALARMS - an orphaned 'ci' step, run 'running' for days -> away-mode digest:
  stale persisted 500s (possible wedge, validation run stranded: ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound)): sess:fm-orphan-w1

STILL ALARMS - no active run at all -> away-mode digest:
  stale persisted 500s (possible wedge): sess:fm-norun-w1

the hold is capped in away mode too (FM_RUN_PROGRESS_HOLD_MAX=2):
  window 1: held
  window 2: held
  window 3: stale persisted 500s (possible wedge, validation run still progressing but this crew has been silent for 2 held windows: test running, last activity 14m23s ago (silent 863s, bound 1800s)): sess:fm-cap-w1
Evidence: Read cost measured: escalation point, not poll path

watcher polled at FM_POLL=1 for 12s, FM_STALE_ESCALATE_SECS=240 wedge holds recorded : 1 fm-run-progress.sh calls : 1 no-mistakes axi status : 1 (~12 polls elapsed; one bounded read is the whole cost, because the hold restarts the timer and puts the next look a full 240s window away.)


=== read cost: escalation point, not poll path ===

watcher polled at FM_POLL=1 for 12s, FM_STALE_ESCALATE_SECS=240
  wedge holds recorded      : 1
  fm-run-progress.sh calls  : 1
  no-mistakes axi status    : 1

(~12 polls elapsed; one bounded read is the whole cost, because the hold
 restarts the timer and puts the next look a full 240s window away.)
Evidence: Reader contract suite (18 cases)
ok - an executing step with recent activity reads as progressing
ok - commas inside a quoted last_activity do not shift the parsed fields
ok - a step that has just started and not logged yet reads as progressing
ok - an executing step silent past the bound reads as stranded, naming the step
ok - the pipeline's quiet marker is read for its age, not taken as a stranded verdict
ok - a step that orphaned and never logged reads as stranded from its own age
ok - when nothing moved, the stranded verdict names the longest-silent step
ok - one executing step that moved recently proves the run as a whole is moving
ok - the stranded bound defaults to 1800s, states itself, and separates 29m50s from 30m10s
ok - the stranded bound is configurable and falls back to its default when malformed
ok - compact durations in d/h/m/s and ms all parse
ok - a run parked at a gate carries no progress evidence
ok - a non-executing active step (a gate) is not progress evidence
ok - a status answered for another branch carries no progress evidence
ok - unknown task, scout, torn-down worktree, and a failed read all carry no evidence
ok - an active step whose timings cannot be parsed carries no evidence
ok - the bounded read goes through the shared timeout owner, so its portable fallback arm both reads and bounds
ok - the reader writes nothing, exits 0 on any read, and exits non-zero only on usage
Evidence: Watcher triage suite (60 cases, 6 new wedge-gate cases)
ok - signal_reason_is_actionable: benign absorbed, captain verbs and coalesced batches surfaced
ok - stale_is_terminal: terminal status surfaces, non-terminal and no-status are benign
ok - scan_captain_relevant_statuses lists only captain-relevant statuses
ok - classifier primitives: keyed decisions and activity phases, captain relevance, window-to-task, and overrides
ok - crew_is_provably_working: only working+run-step/pane is provable; idle/finished/parked/failed/unknown surface
ok - status_is_paused: only the leading paused verb matches, and paused is not captain-relevant
ok - crew_absorb_class: working/paused/none from one read; crew_is_paused and crew_is_provably_working agree
ok - crew_wedge_progress: only a progressing run holds, a dead agent never does, everything else is none
ok - signal_crew_provably_working: benign only when every referenced crew is provably working
ok - a no-verb signal whose crew is provably working is absorbed (no exit, no queue, suppressor advanced, beacon present)
ok - a bare turn-end whose crew is provably working (busy pane) is absorbed
ok - a bare turn-end whose crew is not provably working is surfaced (the swallowed-finish fix)
ok - a no-verb working: note whose crew is idle with no running pipeline is surfaced
ok - captain-relevant signal is surfaced (queue + exit) and marked surfaced
ok - a stale pane sitting on a terminal status is surfaced (queue + exit)
ok - a stale terminal-looking status is overridden and absorbed while a run is actively working, then wedge-escalated
ok - a parked keyed decision surfaces once, not once per pane repaint
ok - a resolved keyed decision can reopen identically and surface again
ok - a new keyed decision on an already-parked pane still surfaces
ok - a parked crew that stops responding is still detected as a wedge suspect
ok - provably-working non-terminal stale is absorbed on first sight, then wedge-escalated past the threshold
ok - a crew parked on a demonstrably progressing validation run holds its wedge escalation
ok - a crew whose validation run has stranded still wedge-escalates, naming the step
ok - a crew wedged with no active run escalates exactly as it does today
ok - a confidently dead agent still escalates however well its validation run is moving
ok - consecutive run-progress holds are capped: past the cap the pane escalates anyway naming the still-moving run, and the count resets so the cadence repeats
ok - a busy pane past its turn-age bound is held while its run is moving and escalates anyway past the hold cap
ok - consecutive wedge escalations on the same pane accumulate and demand deep inspection at the threshold
ok - a pane becoming active again resets the consecutive wedge-escalation counter
ok - a busy worker below the turn-age bound remains working with no escalation
ok - a busy worker with a stable pane hash still escalates once its completed-turn age reaches the bound
ok - a busy worker whose pane hash changes every poll still escalates once its completed-turn age reaches the bound
ok - touching a busy worker's completed-turn marker resets the age and prevents an old-age escalation
ok - repeated busy turn-age escalations reuse the existing escalation counter and demand deep inspection at the threshold
ok - the production default busy-turn-age bound is 3600s (5min under does not wedge, 66min over does)
ok - a not-provably-working non-terminal stale is surfaced immediately (never left to wait out the timer)
ok - a declared pause is absorbed on first sight, then re-surfaced as a recheck past the threshold, never wedge-escalated
ok - pause_resurface_window doubles per unchanged recheck, caps, and clamps a misconfigured cap
ok - pause_streak_bump reconciles the wait on record before counting a recheck
ok - an unchanged declared wait rechecks on a widening cadence while a genuine wedge still escalates
ok - exited declared-pause and captain-held panes use bounded pause cadence while a live decision gate still surfaces once
ok - a declared paused secondmate re-surfaces on the bounded normal-mode cadence
ok - a non-paused secondmate retains normal stale suppression
ok - a resumed secondmate clears pause and stale tracking before stale exemption
ok - unchanged stale hashes reclassify when a crew enters or leaves pause
ok - a declared pause is periodically rechecked against authoritative active-run state
ok - a paused status overridden by authoritative working preserves its wedge timer and escalates
ok - matching non-terminal stale suppressors repair missing or corrupt stale-since timers
ok - triage log capping handles wc byte counts with leading spaces
ok - a captured process-event result wakes a healthy watcher proactively, with no manual drain
ok - a process-event wake is delivered once: no duplicate wake while queued, and none once handled
ok - complete process-event queue keys map to distinct seen markers
ok - queue revalidation, proactive output, and marker commit serialize with drain
/home/sungin/.no-mistakes/worktrees/bc8432f7c9f8/01KZ82ABJB3R0JZA1DHK506BP7/bin/fm-push-transition-lib.sh: line 92: echo: write error: Broken pipe
tests/wake-helpers.sh: line 286: 689367 Killed                  PATH="$dir/fakebin:$PATH" FM_HOME="$dir" FM_PROCEVENT_CLAIM_ROOT="$dir/claims" FM_CREW_STATE_BIN="$dir/fakebin/fm-crew-state.sh" FM_POLL=0.2 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out"
tests/wake-helpers.sh: line 286: 689927 Killed                  PATH="$dir/fakebin:$PATH" FM_HOME="$dir" FM_PROCEVENT_CLAIM_ROOT="$dir/claims" FM_CREW_STATE_BIN="$dir/fakebin/fm-crew-state.sh" FM_POLL=0.2 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 "$WATCH" > "$out"
ok - surfacing failures replay before marker commit and suppress only after delivered output
ok - marker failure exits through the shared wake owner, releases its lock, and replays later
ok - a heartbeat with no captain-relevant change is absorbed and backs off the cadence
ok - heartbeat backstop fail-safe surfaces a captain-relevant status the per-wake path missed
ok - the liveness beacon stays fresh while the watcher absorbs benign wakes (fm-guard never false-alarms)
ok - with .afk present the watcher reverts to one-shot so the daemon owns triage (no double-triage)
ok - AFK changed paused panes hand off plain stale identities for daemon-owned pause triage
- Evidence: Away-mode daemon suite (104 cases, 3 new hold/escalate/cap cases) (local file: /tmp/no-mistakes-evidence/01KZ82ABJB3R0JZA1DHK506BP7/fm-daemon.test.log) - Evidence: E2E harness sources (reproducible) (local file: /tmp/no-mistakes-evidence/01KZ82ABJB3R0JZA1DHK506BP7/e2e-issue-53.sh, /tmp/no-mistakes-evidence/01KZ82ABJB3R0JZA1DHK506BP7/e2e-issue-53-afk.sh, /tmp/no-mistakes-evidence/01KZ82ABJB3R0JZA1DHK506BP7/e2e-issue-53-readcost.sh)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • ⚠️ bin/fm-run-progress.sh:121 - The bounded-read ladder is two rungs (timeout, gtimeout) but the comment directly above it claims "Same bounding ladder as bin/fm-crew-state.sh's nm_run" - and that sibling has three rungs, adding a perl arm at bin/fm-crew-state.sh:281. bin/fm-timeout-lib.sh:29 states the reason explicitly: "macOS ships neither timeout nor gtimeout without coreutils, so the perl arm is the portability floor rather than a curiosity." README.md advertises macOS as a first-class platform. Concrete failure: on a stock macOS host, command -v timeout and command -v gtimeout both fail, so line 128 emits progress: none - no way to bound the status read on every call; crew_wedge_progress collapses that to none, and the wedge escalation is byte-identical to before this change. The entire fix is inert there and issue Supervision wedge-escalates a worker parked on a progressing validation run #53's five-in-a-row false escalations keep firing - while fm-crew-state.sh on that same host still reads the run correctly via its perl arm, which is exactly the asymmetry the intent complains about ("the authoritative current-state reader already knew; the escalation path simply never consulted it"). No test catches this because the suite runs where timeout exists. Fix at the earliest shared boundary rather than hand-rolling a third copy: source bin/fm-timeout-lib.sh (declared "the single owner of bounded external command execution", already carrying all three rungs) and use ( cd "$WT" && fm_run_timed "$NM_TIMEOUT" no-mistakes axi status ), treating its 125 return as the "could not bound" case. FM_TIMEOUT_FORCE_FALLBACK=1 then exercises the perl arm on hosts that do have timeout(1).
  • ⚠️ bin/fm-watch.sh:385 - The hold restarts the timer with no cap on consecutive holds, so "a delay, never a loss ... at most one extra window" holds only if the run itself stops progressing. Concrete reachable path: a crew launches its no-mistakes run, then its harness hangs mid-turn; tmux/herdr report the endpoint alive (not dead, so the crew_wedge_progress short-circuit at fm-classify-lib.sh:519 does not fire); the pipeline executes its own review/test/ci steps and keeps logging, so fm-run-progress.sh returns progressing on every look. Every wedge escalation is then held - watcher (this line) and daemon (fm-supervise-daemon.sh:1068) alike - for the entire remaining run, which can be an hour or more, with no captain-visible signal (triage_log and the daemon log are debug sinks, not wakes). Before this change that pane escalated every STALE_ESCALATE_SECS (240s default). The intent's design decision 4 treats the confidently dead agent as "the one shape 'the run is fine' would otherwise hide", but alive-but-hung is a second such shape and it is the one the suppressor trades away, which sits in tension with the stated hard constraint "WITHOUT weakening detection of one that is genuinely stuck". This is a policy call, not a code defect - please confirm it is intended. If a bound is wanted, the natural place is the same single owner: cap consecutive holds per pane (reusing the existing .wedge-escalations-<key> style counter) and escalate anyway past the cap, carrying the progress detail in the reason so the alarm stays informative.
  • ℹ️ bin/fm-watch.sh:1261 - wedge_timer_check is also the escalation path for a busy pane past FM_BUSY_TURN_MAX_SECS (this line and fm-watch.sh:1273), not only the stale/possible-wedge path the intent describes. The new gate therefore silently applies there too: a pane that is busy with no completed turn for over an hour is now held for as long as its validation run keeps progressing. That is arguably the right answer for a crew running no-mistakes in the foreground (the pane stays busy for the whole pipeline), but it is a behavior extension beyond the three shapes the intent enumerates, and docs/architecture.md still describes the busy-pane bound as routed through the wedge escalation "with the identical reason, escalation count, and demand-deep-inspection marker" without noting the new hold. Worth confirming it is intended, and adding a clause to that sentence if so.
  • ℹ️ docs/configuration.md:683 - FM_RUN_PROGRESS_BIN (bin/fm-classify-lib.sh:37) is a test-override seam introduced for exactly the same reason as FM_CREW_STATE_BIN, which is documented at docs/configuration.md:645. This change documents FM_RUN_STRANDED_SILENCE_SECS and FM_RUN_PROGRESS_NM_TIMEOUT but omits the seam itself, so tests/wake-helpers.sh's hermetic stub install has no documented contract to point at. Add a matching line beside the sibling: FM_RUN_PROGRESS_BIN=bin/fm-run-progress.sh # test override for the validation-run progress reader consulted at the wedge-escalation point.

🔧 Fix: bound the run-progress hold and use the shared timeout owner
1 info still open:

  • ℹ️ bin/fm-supervise-daemon.sh:216 - The cap's default and its numeric sanitization exist as two independent copies: bin/fm-watch.sh:364-365 (FM_RUN_PROGRESS_HOLD_MAX=${FM_RUN_PROGRESS_HOLD_MAX:-15} plus a case guard) and bin/fm-supervise-daemon.sh:216 (RUN_PROGRESS_HOLD_MAX_DEFAULT=15) resolved inline again at :1109-1110. There is an immediately competing semantic owner rather than a mere style preference: bin/fm-classify-lib.sh already owns crew_wedge_progress, which the intent names as the single owner of the wedge policy precisely "so the always-on watcher and the away-mode daemon cannot drift apart", and the cap is the other half of that same policy. The daemon's own comment concedes the risk by asking the reader to keep it "matching the always-on watcher's FM_RUN_PROGRESS_HOLD_MAX" - a comment is the only thing holding the two literals together today, so editing one default silently leaves the supervisors on different caps. Mechanical fix, no behavior change: resolve FM_RUN_PROGRESS_HOLD_MAX (default plus the non-numeric fallback) once in bin/fm-classify-lib.sh beside crew_wedge_progress, and have both supervisors read that. The value's reasoning comment stays where it is in bin/fm-watch.sh, matching how FM_RUN_STRANDED_SILENCE_SECS keeps its reasoning beside its own owner.
✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-run-progress.test.sh - 18 cases, the reader's parse/threshold/no-evidence contract
  • bash tests/fm-watch-triage.test.sh - 60 cases, including the 6 new watcher wedge-gate cases
  • bash tests/fm-daemon.test.sh - 104 cases, including the 3 new away-mode hold/escalate/cap cases
  • bash tests/fm-test-run.test.sh - test-selection routing for the newly registered bin/fm-run-progress.sh
  • bash tests/fm-documentation-audiences.test.sh - docs coverage for the new bin/ entry and configuration/architecture prose
  • E2E: real bin/fm-watch.sh + real bin/fm-run-progress.sh + fake no-mistakes axi status emitting the recorded TOON, run against base commit 5537370's bin/ tree (5 escalations reproduced) and this branch (0 wakes, 5 holds logged)
  • E2E: the four must-still-alarm shapes driven through the real watcher - stranded test step, orphaned ci step running for days, no active run, confidently dead agent while its run progresses
  • E2E: FM_RUN_PROGRESS_HOLD_MAX=3 over 5 windows - held x3, forced escalation naming the still-moving run, then the cadence repeats
  • E2E away mode: real housekeeping in bin/fm-supervise-daemon.sh + real reader, asserting the buffered digest (state/.subsuper-escalations) for progressing/stranded/orphaned/no-run plus the hold cap
  • Read-cost measurement: counting shim around the reader, watcher polling at FM_POLL=1 for 12s with FM_STALE_ESCALATE_SECS=240 - 1 reader call, 1 axi status call
  • Manual: bin/fm-run-progress.sh CLI transcript across progressing / lowered-bound / stranded / orphaned / no-run / torn-down-worktree
⚠️ **Document** - 2 infos
  • ℹ️ docs/fm-test-portable-shards.md:68 - docs/fm-test-portable-shards.md hand-copies derived facts from bin/fm-test-run.sh's partition (per-lane script counts 15/18/17/19 = 69, and estimated shard durations). Adding tests/fm-run-progress.test.sh puts 70 scripts in portable-serial and shifts one shard's estimate by its 333 ms hint, so those tables are now slightly off. I did not update them: git history over the last ~12 test-adding commits shows this doc is never touched when a test lands, the doc itself states membership is derived and that stale hints cost balance rather than coverage, and the numbers are presented as a dated 2026-08-02 measurement. Follow-up worth considering separately: generate those lane tables from bin/fm-test-run.sh (the authoritative partition) with a drift check, instead of hand-maintaining them.
  • ℹ️ docs/architecture.md:96 - docs/architecture.md:96 and docs/scripts.md:101 each enumerate bin/fm-classify-lib.sh's responsibilities ("captain-relevant status verbs, declared-external-wait vocabulary, and status-scan primitives" / "wake classification vocabulary") and neither now mentions crew_wedge_progress. I deliberately left both alone rather than extending them: docs/architecture.md:23 is the single authoritative statement that crew_wedge_progress is the shared wedge-policy owner, and adding it to two more summary enumerations would be synchronizing prose copies of the same fact. Flagging so the omission reads as intentional rather than missed.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Sungin Kim added 3 commits August 5, 2026 04:18
…rogressing

A worker that backgrounds a validation call and goes quiet was escalated as a
possible wedge every threshold - five times in a row on one pane - while the
authoritative current-state reader said `working · run-step · validating
(running)` the whole time. The escalation path never consulted it.

"There is an active run" cannot be the suppressor, because a run's own step can
strand: a test step sat nineteen minutes on its opening line with a live but
idle agent, and a `ci` step has orphaned and left a run reporting `running` for
days. So the gate reads PROGRESS, not status.

bin/fm-run-progress.sh reads the pipeline's own `active_steps[].active_for` and
`last_activity` and reports progressing, stranded, or none. Both supervisors
consult it through one policy owner (crew_wedge_progress) at the moment they
would raise the alarm - not on the poll path, so the bounded pipeline read costs
at most one call per window per pane. A progressing run holds the escalation and
restarts the timer; everything else escalates, and a stranded run now names the
step that stopped.

The hold only ever suppresses on positive evidence. No run, a run parked at a
gate, an unreadable status, and a confidently dead agent all escalate exactly as
before, so a crew wedged with no run at all keeps the alarm's original purpose
intact. A crew that wedges right after a hold waits at most one more window: a
delay, never a loss.

FM_RUN_STRANDED_SILENCE_SECS (1800) is the named bound, documented with its
reasoning beside it: three times the pipeline's own 10m `step_quiet_warning`,
which is deliberately a liveness clue rather than a verdict because review and
test steps routinely emit one opening line and then work silently for ten to
eighteen minutes.

Verified against a live mid-pipeline run on no-mistakes v1.41.2, recorded in
docs/verification/supervision.md. tests/fm-run-progress.test.sh pins every class
against that recorded shape, including the stranded and orphaned-step cases;
the watcher and daemon suites pin what still alarms.

Closes #53
@cursor

cursor Bot commented Aug 5, 2026

Copy link
Copy Markdown

Bugbot is not enabled for your account, so this pull request was not reviewed.

Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Supervision wedge-escalates a worker parked on a progressing validation run

1 participant