fix(bin): hold the wedge alarm while a crew's validation run is progressing - #55
Merged
Merged
Conversation
added 3 commits
August 5, 2026 04:18
…rogressing A worker that backgrounds a validation call and goes quiet was escalated as a possible wedge every threshold - five times in a row on one pane - while the authoritative current-state reader said `working · run-step · validating (running)` the whole time. The escalation path never consulted it. "There is an active run" cannot be the suppressor, because a run's own step can strand: a test step sat nineteen minutes on its opening line with a live but idle agent, and a `ci` step has orphaned and left a run reporting `running` for days. So the gate reads PROGRESS, not status. bin/fm-run-progress.sh reads the pipeline's own `active_steps[].active_for` and `last_activity` and reports progressing, stranded, or none. Both supervisors consult it through one policy owner (crew_wedge_progress) at the moment they would raise the alarm - not on the poll path, so the bounded pipeline read costs at most one call per window per pane. A progressing run holds the escalation and restarts the timer; everything else escalates, and a stranded run now names the step that stopped. The hold only ever suppresses on positive evidence. No run, a run parked at a gate, an unreadable status, and a confidently dead agent all escalate exactly as before, so a crew wedged with no run at all keeps the alarm's original purpose intact. A crew that wedges right after a hold waits at most one more window: a delay, never a loss. FM_RUN_STRANDED_SILENCE_SECS (1800) is the named bound, documented with its reasoning beside it: three times the pipeline's own 10m `step_quiet_warning`, which is deliberately a liveness clue rather than a verdict because review and test steps routinely emit one opening line and then work silently for ten to eighteen minutes. Verified against a live mid-pipeline run on no-mistakes v1.41.2, recorded in docs/verification/supervision.md. tests/fm-run-progress.test.sh pins every class against that recorded shape, including the stranded and orphaned-step cases; the watcher and daemon suites pin what still alarms. Closes #53
|
Bugbot is not enabled for your account, so this pull request was not reviewed. Enable Bugbot in the Cursor dashboard to get automatic reviews on future PRs. |
This was referenced Aug 7, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
GOAL (GitHub issue #53): stop firstmate's supervision raising a possible-wedge alarm for a worker that is parked on a validation run which is demonstrably progressing, WITHOUT weakening detection of one that is genuinely stuck. Observed 2026-08-05: one pane produced five consecutive 'stale: ... possible wedge, escalation N' escalations while bin/fm-crew-state.sh reported 'working - source: run-step - validating (running)' for the same task the whole time. The authoritative current-state reader already knew; the escalation path simply never consulted it.
THE HARD CONSTRAINT, AND THE BAR THIS CHANGE MUST BE JUDGED AGAINST. This NARROWS an alarm. The deliverable that earns it is the evidence of what STILL alarms, pinned by its own tests rather than by argument. All three of these must still escalate and each has a dedicated test: (1) a stranded step past the bound - a test step sat nineteen minutes on its opening line with a live but idle agent and was correctly aborted at thirty; (2) an orphaned step that never logged at all - firstmate has recorded a 'ci' step that orphaned and left a run reporting 'running' for days; (3) a worker wedged with no active run whatsoever, which is the alarm's original purpose. So 'status: running' is deliberately NOT accepted as evidence of health anywhere in this change - only observed PROGRESS is.
DELIBERATE DESIGN DECISIONS a reviewer reading only the diff would not know:
The gate is a SUPPRESSOR THAT REQUIRES POSITIVE EVIDENCE, not a filter. bin/fm-run-progress.sh reports exactly three classes: progressing, stranded, none. Only 'progressing' may ever quiet an alarm. Every no-evidence shape - no metadata, torn-down worktree, no-mistakes absent, a bounded read that could not complete, a status answered for another branch, a run parked at a gate, no active_steps rows, an unparseable duration - collapses to 'none', which leaves the escalation path byte-identical to what it was before this change. This is why the change cannot weaken the alarm by accident: a bug in the reader produces 'none', which alarms.
STRIPPING THE PIPELINE'S OWN 'quiet' MARKER RATHER THAN TRUSTING IT IS INTENTIONAL AND MUST SURVIVE REVIEW. no-mistakes prefixes last_activity with 'quiet' once its own step_quiet_warning (10m by default) elapses. That marker is a CLAIM about silence, not a MEASUREMENT of it, and no-mistakes' own guidance calls it 'a liveness clue, not permission to cancel'. Review and test steps routinely emit one opening line and then work silently for ten to eighteen minutes, so treating the marker as a wedge verdict would re-raise exactly the false alarms this change removes. The reader therefore strips the prefix and parses the AGE the field carries. tests/fm-run-progress.test.sh pins this directly (test_pipeline_quiet_marker_is_stripped_not_trusted): a step 14m quiet still reads 'progressing'. Do not 'simplify' this into trusting the marker.
THE THRESHOLD IS A NAMED, DOCUMENTED, CONFIGURABLE VALUE AND THAT IS HALF THE POINT OF THE CHANGE. FM_RUN_STRANDED_SILENCE_SECS defaults to 1800 and its reasoning sits beside it in bin/fm-run-progress.sh, in docs/configuration.md, and in docs/architecture.md: 3x the pipeline's own 10m quiet warning, comfortably past the longest healthy silence observed (18m), and the point at which an operator actually judged a stranded step stranded and aborted it. An unexplained magic number here would be a defect in its own right. Do not inline it as a literal, do not drop the reasoning comment, and do not fold it into another module's constant.
A CONFIDENTLY DEAD AGENT IS NEVER ABSORBED, however well its run is moving, and short-circuits before paying for the read. The pipeline executes its own steps, so a run can keep advancing with nobody left to answer its next gate - that is a wedge, and it is the one shape 'the run is fine' would otherwise hide. Pinned by test_dead_agent_escalates_even_while_its_run_progresses.
THE READ IS AT THE ESCALATION POINT, NOT ON THE POLL PATH. bin/fm-watch.sh documents in its own comments that the poll path must stay cheap and must NOT re-read crew state. bin/fm-run-progress.sh shells out to a bounded 'no-mistakes axi status', so both supervisors consult it only where an alarm would otherwise be raised - at most one bounded call per STALE_ESCALATE_SECS window per pane, rather than one per FM_POLL tick. A hold RESTARTS the timer rather than clearing it, so the next look is a full window away instead of one poll away. The bounded cost of that: a crew that wedges immediately after a hold waits at most one extra window. A delay, never a loss.
ONE POLICY OWNER FOR TWO SUPERVISORS. bin/fm-classify-lib.sh exists precisely so the always-on watcher and the away-mode daemon cannot drift apart on overlapping triage, so crew_wedge_progress lives there and both call it. I checked bin/fm-supervise-daemon.sh deliberately, as the task asked: away mode skips declared pauses in its stale loop, so it dodges the exact shape seen today, but it will still wedge-escalate a mid-run worker that never wrote a 'paused:' line - the same class of bug - so the same fix is applied there, through the same owner. Two new daemon tests pin the hold and what still escalates in away mode.
Firstmate's own bin/fm-crew-state.sh was deliberately NOT extended. It answers 'is there an active run', which is a different question from 'is that run moving', and its single-line output contract ('state: X - source: Y - detail') is parsed by many consumers; widening its detail field would be the fragile choice.
VERIFICATION. The reader was proven against the real thing before any test was written: run read-only against a live mid-pipeline task on no-mistakes v1.41.2, it returned 'progress: progressing - test running, last activity 14m23s ago (silent 863s, bound 1800s)', and 'stranded' with the bound lowered. That evidence, the exact TOON shape parsed, and the refresh trigger are recorded in docs/verification/supervision.md, which is the classified maintainer-verification owner. tests/fm-run-progress.test.sh (17 cases) pins every class portably against that recorded shape, including quoted commas inside last_activity that a naive comma split would mis-align.
TEST-SAFETY DECISION worth noting: tests/wake-helpers.sh installs a hermetic fm-run-progress stub next to the existing wedge-alarm recorder, following that same established pattern, so no suite can reach a real pipeline from a fixture. The stub defaults to 'none', which is why every pre-existing wedge assertion in the watcher and daemon suites still asserts today's alarm unchanged.
SCOPE EXCLUSIONS the user explicitly ruled out of this task, tracked separately, and NOT to be fixed here: the tests/lib.sh fm_test_tmproot EXIT-trap leak; tests/fm-teardown.test.sh emitting no completion marker; and the open upstream-merge PR #44. If PR #44 lands under this branch, reconcile by MERGE, never rebase.
Delivery: firstmate ships this project no-mistakes; the PR body must carry 'Closes #53' and reference #53.
What Changed
bin/fm-run-progress.sh, a read-only reader that answers whether a crew's validation run is moving (a different question fromfm-crew-state.sh's "is there an active run"). It parses the pipeline's ownactive_steps[].active_for/last_activityfrom a boundedno-mistakes axi statusand reports exactly three classes -progressing,stranded,none- stripping the pipeline'squietprefix and measuring the age it carries rather than trusting the marker, and separating slow from stranded at the named, documentedFM_RUN_STRANDED_SILENCE_SECS(1800s, 3x the pipeline's own 10m quiet warning). Every no-evidence shape (no metadata, torn-down worktree, unbounded read, run parked at a gate, unparseable output) collapses tonone.crew_wedge_progressinbin/fm-classify-lib.shthe single owner of the wedge policy, and wired both supervisors to it at the escalation point only - never on the poll path, so the bounded read costs at most one call per escalate window per pane.bin/fm-watch.shandbin/fm-supervise-daemon.shnow hold and restart the timer when the run isprogressing, escalate byte-identically on every other answer, short-circuit a confidently dead agent before paying for the read, name the stopped step when a run isstranded, and cap consecutive holds (FM_RUN_PROGRESS_HOLD_MAX/RUN_PROGRESS_HOLD_MAX_DEFAULT, 15) past which the pane escalates anyway carrying the progress detail - a bounded delay that repeats as a check-in cadence, never a lost alarm.tests/fm-run-progress.test.shcovers the reader's parse/threshold/no-evidence contract, plus new watcher wedge-gate cases and away-mode daemon hold/escalate/cap cases;tests/wake-helpers.shinstalls a hermeticfm-run-progressstub defaulting tononebeside the existing wedge-alarm recorder, so every pre-existing wedge assertion still asserts today's alarm and no fixture can reach a real pipeline. Registered the new script inbin/fm-test-run.shtest selection and documented the gate, its threshold, and the recorded live-pipeline verification shape indocs/architecture.md,docs/configuration.md,docs/scripts.md, anddocs/verification/supervision.md.Closes #53 - #53
Risk Assessment
✅ Low: Every round-1 finding is fully addressed with correct fail-closed semantics, complete state-file reset coverage across both supervisors, and tests pinning the hold, the cap, the forced escalation's progress detail, the count reset, the repeating cadence, and the busy path; the only item left is a mechanical consolidation of one duplicated constant.
Testing
Ran the three suites that own this change (18 reader cases, 60 watcher-triage cases, 104 daemon cases) plus the test-routing and documentation suites that it also touches - all pass. Unit tests alone would not have shown the intent working, so I built an end-to-end harness wiring the real fm-watch.sh to the real fm-run-progress.sh behind a fake
no-mistakesprinting the exactaxi statusTOON recorded in docs/verification/supervision.md, with nothing stubbed between the pipeline's status output and the wake queue. Against base commit 5537370's bin/ tree that fixture reproduces the reported bug exactly - five consecutivepossible wedge, escalation Nwakes escalating to demand-deep-inspection while fm-crew-state reportedworking - source: run-step - validating (running); on this branch the identical fixture produces zero wakes and fiveheld ... (progressing ...)triage lines. The bar the intent sets holds: a stranded test step, an orphanedcistep, a crew with no active run, and a dead agent whose run is progressing all still escalate through the real watcher, the first two now naming the step that stopped. I also verified the two decisions a reviewer could not confirm from the diff - the pipeline'squietmarker is stripped and its age parsed (14m quiet reads progressing; the same run reads stranded with the bound lowered), and the bounded read sits at the escalation point, measured at oneaxi statuscall across ~12 polls rather than one per tick. Away mode enforces the same policy through the shared owner, including the hold cap. No visual artifact: this is a shell supervision path with no rendered surface, so the end-user-visible output is the wake reason line and away-mode digest text, captured verbatim as CLI transcripts. One note on my own harness rather than the product: my first away-mode attempt showed the progressing crew escalating because tests/wake-helpers.sh exports FM_RUN_PROGRESS_BIN at source time and clobbered my override, so the daemon read the hermetic stub'snone- the safety seam is genuinely unforgeable from a fixture, which is what it was built for. Worktree left clean; all scratch trees removed.Evidence: Issue #53 reproduced on the base commit, then fixed on this branch (same fixture, real watcher + real reader)
=== 2. BEFORE (base commit 5537370): the reported bug === A crew parked on a validation run that is demonstrably progressing. fm-crew-state.sh, all five windows: state: working · source: run-step · validating (running) ------------------------------------------------------------------ window 1: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 1) window 2: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 2) window 3: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 3, demand-deep-inspection: same pane has wedge-escalated 3 times in a row - do not re-absorb on the run-step/pane state alone) window 4: stale: test:fm-issue53 (idle 500s, possible wedge, escalation 4, demand-deep-inspection: ...) window 5: stale: test:fm-issue53 (idle 501s, possible wedge, escalation 5, demand-deep-inspection: ...) === 3. AFTER (this change): the same fixture, byte for byte === window 1: (no wake - watcher still supervising) window 2: (no wake - watcher still supervising) window 3: (no wake - watcher still supervising) window 4: (no wake - watcher still supervising) window 5: (no wake - watcher still supervising) wake queue delivered to the captain: (empty above = nothing surfaced) what the watcher logged instead (state/.watch-triage.log): held non-terminal stale (provably working after a declared pause) wedge escalation (progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s), idle 500s, hold 1/15): test:fm-issue53 ... hold 2/15 ... hold 3/15 ... hold 4/15 ... hold 5/15Evidence: What still escalates: all four shapes the intent requires, through the real watcher
(1) a test step stranded past the bound, live but idle agent: reader says : progress: stranded · test running, last activity 31m0s ago (silent 1860s, past the 1800s bound) watcher wake: stale: test:fm-stranded (idle 500s, possible wedge, escalation 1, validation run stranded: test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)) (2) a ci step that orphaned and never logged, run still 'running' for days: reader says : progress: stranded · ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound) watcher wake: stale: test:fm-orphan (idle 500s, possible wedge, escalation 1, validation run stranded: ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound)) (3) a worker wedged with no active run whatsoever: reader says : progress: none · no actively-executing step (run parked, terminal, or between steps) watcher wake: stale: test:fm-norun (idle 500s, possible wedge, escalation 1) (4) a confidently dead agent, however well its run is moving: reader says : progress: progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s) watcher wake: stale: test:fm-dead (idle 500s, possible wedge, escalation 1) === 5. the hold is a bounded delay, never a lost alarm (FM_RUN_PROGRESS_HOLD_MAX=3) === window 1: held window 2: held window 3: held window 4: stale: test:fm-cap (idle 500s, possible wedge, escalation 1, validation run still progressing but this pane has been silent for 3 held windows: test running, last activity 14m23s ago (silent 863s, bound 1800s)) window 5: heldEvidence: The reader as an operator runs it: the quiet marker is measured, not trusted
$ bin/fm-run-progress.sh reader progress: progressing · test running, last activity 14m23s ago (silent 863s, bound 1800s) # the healthy pane from issue #53: 'quiet 14m23s ago', still moving $ FM_RUN_STRANDED_SILENCE_SECS=60 bin/fm-run-progress.sh reader progress: stranded · test running, last activity 14m23s ago (silent 863s, past the 60s bound) # same run, bound lowered - the class is a measurement, not a mood $ bin/fm-run-progress.sh reader # with the worktree torn down progress: none · worktree gone (torn down?) # every no-evidence shape collapses to none, which alarmsEvidence: Away mode: same policy through the shared owner, captain-visible digest
crew parked on a progressing validation run -> away-mode digest: (nothing buffered - no digest would reach the captain) STILL ALARMS - a stranded validation step -> away-mode digest: stale persisted 500s (possible wedge, validation run stranded: test running, last activity 31m0s ago (silent 1860s, past the 1800s bound)): sess:fm-stranded-w1 STILL ALARMS - an orphaned 'ci' step, run 'running' for days -> away-mode digest: stale persisted 500s (possible wedge, validation run stranded: ci ci, no activity in 3d4h12m (silent 274320s, past the 1800s bound)): sess:fm-orphan-w1 STILL ALARMS - no active run at all -> away-mode digest: stale persisted 500s (possible wedge): sess:fm-norun-w1 the hold is capped in away mode too (FM_RUN_PROGRESS_HOLD_MAX=2): window 1: held window 2: held window 3: stale persisted 500s (possible wedge, validation run still progressing but this crew has been silent for 2 held windows: test running, last activity 14m23s ago (silent 863s, bound 1800s)): sess:fm-cap-w1Evidence: Read cost measured: escalation point, not poll path
watcher polled at FM_POLL=1 for 12s, FM_STALE_ESCALATE_SECS=240 wedge holds recorded : 1 fm-run-progress.sh calls : 1 no-mistakes axi status : 1 (~12 polls elapsed; one bounded read is the whole cost, because the hold restarts the timer and puts the next look a full 240s window away.)Evidence: Reader contract suite (18 cases)
Evidence: Watcher triage suite (60 cases, 6 new wedge-gate cases)
/tmp/no-mistakes-evidence/01KZ82ABJB3R0JZA1DHK506BP7/fm-daemon.test.log) - Evidence: E2E harness sources (reproducible) (local file:/tmp/no-mistakes-evidence/01KZ82ABJB3R0JZA1DHK506BP7/e2e-issue-53.sh, /tmp/no-mistakes-evidence/01KZ82ABJB3R0JZA1DHK506BP7/e2e-issue-53-afk.sh, /tmp/no-mistakes-evidence/01KZ82ABJB3R0JZA1DHK506BP7/e2e-issue-53-readcost.sh)Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
bin/fm-run-progress.sh:121- The bounded-read ladder is two rungs (timeout, gtimeout) but the comment directly above it claims "Same bounding ladder as bin/fm-crew-state.sh's nm_run" - and that sibling has three rungs, adding a perl arm at bin/fm-crew-state.sh:281. bin/fm-timeout-lib.sh:29 states the reason explicitly: "macOS ships neither timeout nor gtimeout without coreutils, so the perl arm is the portability floor rather than a curiosity." README.md advertises macOS as a first-class platform. Concrete failure: on a stock macOS host,command -v timeoutandcommand -v gtimeoutboth fail, so line 128 emitsprogress: none - no way to bound the status readon every call; crew_wedge_progress collapses that tonone, and the wedge escalation is byte-identical to before this change. The entire fix is inert there and issue Supervision wedge-escalates a worker parked on a progressing validation run #53's five-in-a-row false escalations keep firing - while fm-crew-state.sh on that same host still reads the run correctly via its perl arm, which is exactly the asymmetry the intent complains about ("the authoritative current-state reader already knew; the escalation path simply never consulted it"). No test catches this because the suite runs wheretimeoutexists. Fix at the earliest shared boundary rather than hand-rolling a third copy: source bin/fm-timeout-lib.sh (declared "the single owner of bounded external command execution", already carrying all three rungs) and use( cd "$WT" && fm_run_timed "$NM_TIMEOUT" no-mistakes axi status ), treating its 125 return as the "could not bound" case. FM_TIMEOUT_FORCE_FALLBACK=1 then exercises the perl arm on hosts that do have timeout(1).bin/fm-watch.sh:385- The hold restarts the timer with no cap on consecutive holds, so "a delay, never a loss ... at most one extra window" holds only if the run itself stops progressing. Concrete reachable path: a crew launches its no-mistakes run, then its harness hangs mid-turn; tmux/herdr report the endpointalive(notdead, so the crew_wedge_progress short-circuit at fm-classify-lib.sh:519 does not fire); the pipeline executes its own review/test/ci steps and keeps logging, so fm-run-progress.sh returnsprogressingon every look. Every wedge escalation is then held - watcher (this line) and daemon (fm-supervise-daemon.sh:1068) alike - for the entire remaining run, which can be an hour or more, with no captain-visible signal (triage_log and the daemon log are debug sinks, not wakes). Before this change that pane escalated every STALE_ESCALATE_SECS (240s default). The intent's design decision 4 treats the confidently dead agent as "the one shape 'the run is fine' would otherwise hide", but alive-but-hung is a second such shape and it is the one the suppressor trades away, which sits in tension with the stated hard constraint "WITHOUT weakening detection of one that is genuinely stuck". This is a policy call, not a code defect - please confirm it is intended. If a bound is wanted, the natural place is the same single owner: cap consecutive holds per pane (reusing the existing .wedge-escalations-<key> style counter) and escalate anyway past the cap, carrying the progress detail in the reason so the alarm stays informative.bin/fm-watch.sh:1261- wedge_timer_check is also the escalation path for a busy pane past FM_BUSY_TURN_MAX_SECS (this line and fm-watch.sh:1273), not only the stale/possible-wedge path the intent describes. The new gate therefore silently applies there too: a pane that is busy with no completed turn for over an hour is now held for as long as its validation run keeps progressing. That is arguably the right answer for a crew runningno-mistakesin the foreground (the pane stays busy for the whole pipeline), but it is a behavior extension beyond the three shapes the intent enumerates, and docs/architecture.md still describes the busy-pane bound as routed through the wedge escalation "with the identical reason, escalation count, and demand-deep-inspection marker" without noting the new hold. Worth confirming it is intended, and adding a clause to that sentence if so.docs/configuration.md:683- FM_RUN_PROGRESS_BIN (bin/fm-classify-lib.sh:37) is a test-override seam introduced for exactly the same reason as FM_CREW_STATE_BIN, which is documented at docs/configuration.md:645. This change documents FM_RUN_STRANDED_SILENCE_SECS and FM_RUN_PROGRESS_NM_TIMEOUT but omits the seam itself, so tests/wake-helpers.sh's hermetic stub install has no documented contract to point at. Add a matching line beside the sibling:FM_RUN_PROGRESS_BIN=bin/fm-run-progress.sh # test override for the validation-run progress reader consulted at the wedge-escalation point.🔧 Fix: bound the run-progress hold and use the shared timeout owner
1 info still open:
bin/fm-supervise-daemon.sh:216- The cap's default and its numeric sanitization exist as two independent copies: bin/fm-watch.sh:364-365 (FM_RUN_PROGRESS_HOLD_MAX=${FM_RUN_PROGRESS_HOLD_MAX:-15}plus acaseguard) and bin/fm-supervise-daemon.sh:216 (RUN_PROGRESS_HOLD_MAX_DEFAULT=15) resolved inline again at :1109-1110. There is an immediately competing semantic owner rather than a mere style preference: bin/fm-classify-lib.sh already ownscrew_wedge_progress, which the intent names as the single owner of the wedge policy precisely "so the always-on watcher and the away-mode daemon cannot drift apart", and the cap is the other half of that same policy. The daemon's own comment concedes the risk by asking the reader to keep it "matching the always-on watcher's FM_RUN_PROGRESS_HOLD_MAX" - a comment is the only thing holding the two literals together today, so editing one default silently leaves the supervisors on different caps. Mechanical fix, no behavior change: resolveFM_RUN_PROGRESS_HOLD_MAX(default plus the non-numeric fallback) once in bin/fm-classify-lib.sh besidecrew_wedge_progress, and have both supervisors read that. The value's reasoning comment stays where it is in bin/fm-watch.sh, matching how FM_RUN_STRANDED_SILENCE_SECS keeps its reasoning beside its own owner.✅ **Test** - passed
✅ No issues found.
bash tests/fm-run-progress.test.sh- 18 cases, the reader's parse/threshold/no-evidence contractbash tests/fm-watch-triage.test.sh- 60 cases, including the 6 new watcher wedge-gate casesbash tests/fm-daemon.test.sh- 104 cases, including the 3 new away-mode hold/escalate/cap casesbash tests/fm-test-run.test.sh- test-selection routing for the newly registeredbin/fm-run-progress.shbash tests/fm-documentation-audiences.test.sh- docs coverage for the newbin/entry and configuration/architecture proseE2E: realbin/fm-watch.sh+ realbin/fm-run-progress.sh+ fakeno-mistakes axi statusemitting the recorded TOON, run against base commit 5537370'sbin/tree (5 escalations reproduced) and this branch (0 wakes, 5 holds logged)E2E: the four must-still-alarm shapes driven through the real watcher - stranded test step, orphanedcistep running for days, no active run, confidently dead agent while its run progressesE2E:FM_RUN_PROGRESS_HOLD_MAX=3over 5 windows - held x3, forced escalation naming the still-moving run, then the cadence repeatsE2E away mode: realhousekeepinginbin/fm-supervise-daemon.sh+ real reader, asserting the buffered digest (state/.subsuper-escalations) for progressing/stranded/orphaned/no-run plus the hold capRead-cost measurement: counting shim around the reader, watcher polling atFM_POLL=1for 12s withFM_STALE_ESCALATE_SECS=240- 1 reader call, 1axi statuscallManual:bin/fm-run-progress.shCLI transcript across progressing / lowered-bound / stranded / orphaned / no-run / torn-down-worktreedocs/fm-test-portable-shards.md:68- docs/fm-test-portable-shards.md hand-copies derived facts from bin/fm-test-run.sh's partition (per-lane script counts 15/18/17/19 = 69, and estimated shard durations). Adding tests/fm-run-progress.test.sh puts 70 scripts in portable-serial and shifts one shard's estimate by its 333 ms hint, so those tables are now slightly off. I did not update them: git history over the last ~12 test-adding commits shows this doc is never touched when a test lands, the doc itself states membership is derived and that stale hints cost balance rather than coverage, and the numbers are presented as a dated 2026-08-02 measurement. Follow-up worth considering separately: generate those lane tables from bin/fm-test-run.sh (the authoritative partition) with a drift check, instead of hand-maintaining them.docs/architecture.md:96- docs/architecture.md:96 and docs/scripts.md:101 each enumerate bin/fm-classify-lib.sh's responsibilities ("captain-relevant status verbs, declared-external-wait vocabulary, and status-scan primitives" / "wake classification vocabulary") and neither now mentions crew_wedge_progress. I deliberately left both alone rather than extending them: docs/architecture.md:23 is the single authoritative statement that crew_wedge_progress is the shared wedge-policy owner, and adding it to two more summary enumerations would be synchronizing prose copies of the same fact. Flagging so the omission reads as intentional rather than missed.✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.