Skip to content

fix(bin): stop wedge alarms on settled terminal tasks - #40

Merged
sbracewell64 merged 1 commit into
mainfrom
fm/stale-fires-on-finished-task
Aug 6, 2026
Merged

sbracewell64 merged 1 commit into
mainfrom
fm/stale-fires-on-finished-task

Conversation

@sbracewell64

@sbracewell64 sbracewell64 commented Aug 5, 2026 •

Copy link
Copy Markdown
Owner

Delivery disclosure

This shipped direct-PR without the no-mistakes pipeline and therefore carries NO attestation marker.
It has not been through automated code review, the pipeline's test/lint/docs gates, or any pipeline-driven fix cycle.
The evidence below is a local test run plus this PR's own CI, reported as-is, including every failure and where each one pre-exists on the base.

Problem

A task whose reconciled current state is terminal kept producing possible-wedge stale wakes, because the stale path decided from the status line's text rather than the reconciled state.

Seventeen such wakes were recorded on 2026-08-04 (joined ledger outcomes, defect stale-fires-on-finished-task), the second most expensive supervision defect measured that day. Every one fired on a task that was done, parked, or paused with the work complete; supervision read the state, found nothing to do, and recorded false-positive. An idle pane is the correct condition for those tasks, so the wake carried no information: a done task with its PR open and green awaiting merge approval, a task parked on a captain decision with work committed, a task blocked on a credential.

Fourteen of the seventeen are the reconciled-terminal shape this change addresses. The other three (seq 3144, 3238, 3240) describe an idle pane with a live test shell or long-running tool call underneath; their own ledger notes attribute them to wedge-detection-ignores-child-processes. Those still age and escalate here, correctly, and are out of scope.

Change

bin/fm-classify-lib.sh's crew_absorb_class already performs the one fm-crew-state.sh read on that path, so the new settled verdict is a fourth token from that same read rather than a new call:

Reconciled state Settled?
done yes, unconditionally: nothing is left for the crew to do on its own
parked, blocked only while a durable open decision (status_open_decisions) proves the next move belongs above the crew
everything else no

Both supervisors consume that one verdict, so they cannot drift:

  • Watcher (bin/fm-watch.sh): absorbs a settled stale with no wedge timer and records state/.settled-<key> so unchanged polls of the same pane hash absorb without re-reading the crew state. Any pane change drops the marker and reclassifies.
  • Away daemon (bin/fm-supervise-daemon.sh): classify_stale returns a new settled action that self-handles and drops rather than records wedge tracking, and housekeeping re-checks the reconciled state immediately before escalating an aged marker, which is where a task that settles after going stale would otherwise still alarm.

Why this shape

The alternative was to extend the seen-marker/terminal test on the status text. That is the cheaper-looking option but the wrong one: status text is exactly what is already wrong here, and it cannot see the reconciled state. It would have either missed the parked cases or re-broken the precedence fix that the existing crew-state read exists to protect. Consulting the reconciled state costs zero additional fm-crew-state.sh calls on the watcher path, because that read already happens there once per newly-classified stale hash.

Daemon cost is bounded and stated in-code: one read per stale wake in classify_stale (the watcher enqueues at most one stale wake per distinct pane hash, the same bound the always-on path pays), plus one read in housekeeping only at the moment of escalation.

Genuine wedges are preserved

This is the load-bearing property; breaking it would trade cheap false alarms for a silent stall.

  • failed is deliberately NOT settled. It also reconciles a cancelled run, the mid-supersession state in which a crew must recover custody and resume, so an idle pane there is a real stall.
  • unknown never settles, so a dead endpoint or torn-down worktree keeps aging.
  • A run parked at a gate the crew must answer itself opens no decision, so it keeps aging and still escalates.
  • Reading the reconciled state rather than the log text preserves the earlier precedence fix: a leftover pre-validation done: line under an active run reconciles as working and keeps its wedge timer.
  • The declared-paused: long-cadence recheck and the existing seen-marker dedupe are unchanged.

Nothing rots: an absorb deliberately does not mark the status surfaced, so the terminal status still reaches supervision through the captain-relevant signal path and, if that was missed, through the heartbeat catch-all scan, while a parked or blocked task keeps its open decision surfacing on every wake drain via scan_open_decisions.

Verification

Every new test was witnessed failing first against unmodified code on this exact base (fresh git archive of ad53e97 plus the new tests):

RED test_crew_absorb_class_settled_classifier       reconciled done not classed settled
RED test_settled_done_stale_absorbed                watcher exited -> stale: test:fm-shipped
RED test_settled_parked_stale_absorbed              watcher exited -> stale: test:fm-held-decision
RED test_settled_marker_cleared_when_pane_changes   watcher exited
RED test_stale_settled_terminal_self_handles        escalate|stale + terminal status
RED test_housekeeping_settled_stale_not_escalated   stale persisted 500s (possible wedge)

Results on this head:

Suite Result
fm-watch-triage 51/51
fm-daemon 104/104
fm-crew-state 49/49
fm-supervision-events 7/7
fm-wake-queue 12/12
fm-wake-drain-open-decisions 7/7
fm-watch-arm 3/3
fm-watch-checkpoint 4/4
fm-afk-return 5/5
fm-afk-launch 52/52
fm-guard-stale-banner 17/17
fm-decision-hold-lifecycle 9/9
fm-fleet-snapshot-view 15/15
fm-watcher-lock 19 ok, 1 FAILING
fm-session-start 5 ok, 1 FAILING

bin/fm-lint.sh clean (pinned ShellCheck 0.11.0), bin/fm-doc-audience-check.sh clean.

The two locally failing suites, reported as failures

Both are proven pre-existing on this base, not asserted. Each was run against a pristine git archive of ad53e97, with the relevant production and test files verified byte-identical to the base commit, and each fails there with the identical assertion and exit 1.

  1. tests/fm-watcher-lock.test.sh - restart did not attach to the verified healthy peer. Verified pristine: bin/fm-watch.sh, bin/fm-classify-lib.sh, tests/fm-watcher-lock.test.sh.
  2. tests/fm-session-start.test.sh - MISSING diagnostic did not appear at all. Verified pristine: bin/fm-bootstrap.sh, tests/fm-session-start.test.sh.

Both of these pass in CI (they run in the portable-serial lane, in shards that passed), so their local failures are specific to the machine this ran on. Either way they are not caused by this change and are not fixed by it.

CI on this PR

12 passed, 2 failed of 14. Both failures are accounted for:

  1. PR must be raised via no-mistakes - FAIL, and expected. This is the attestation gate. It fails precisely because this shipped direct-PR without the pipeline and carries no attestation marker, as disclosed at the top. It is not a code failure.
  2. Behavior portable serial 2 - FAIL, pre-existing on the base. The failing suite is tests/fm-remote-secondmate-lifecycle-e2e.test.sh (FM_TEST_END ... exit=1, FM_TEST_SUMMARY_FAMILY family=secondmate ... failed=1). The base commit's own push run on main (run 30999882189, same Behavior portable serial 2 shard) fails on the identical suite with the identical family summary. Not caused by this change.

The family covering every file this PR touches reports clean in the same shard: FM_TEST_SUMMARY_FAMILY family=watcher-wake-lock count=5 ... failed=0, and this PR's own new classifier test is visible passing in the CI log.

Behavioural coverage added:

  • a task reconciled done with an idle pane produces no stale wedge alarm, on first sight and past the wedge threshold
  • a task parked on an open captain decision produces none
  • a run parked at a gate with no open decision still surfaces
  • a settled absorb is dropped as soon as the pane changes and the state stops being settled
  • the away daemon reaches the identical verdict for each case, asserted per fixture
  • a declared paused: external wait still self-handles on its long cadence
  • an ordinary duplicate signal/stale pair still de-duplicates to one entry
  • an unreadable crew still escalates its persistent stale in away mode

Notes for review

This branch was rebased across the trunk reconciliation. AGENTS.md and docs/architecture.md each had a one-line collision with work already on the trunk; both were resolved keeping both behaviours (.pr-dirty-* alongside .settled-*; the conflicting-PR wake alongside the settled-terminal exclusion). Nothing on the trunk was reverted.

While the fork and upstream trunks were still diverged, this branch briefly targeted a base whose status_open_decisions had no symlink guard and no scan_open_decisions, and the in-code rationale was narrowed accordingly so it would not claim a mechanism that base lacked. The reconciled base carries both again, so the fuller rationale has been restored and the symlink concern no longer applies.

A task whose reconciled current state is terminal kept producing
possible-wedge stale wakes, because the stale path decided from the status
line's TEXT rather than the reconciled state. Seventeen such wakes were
recorded on 2026-08-04 (joined ledger outcomes, defect
stale-fires-on-finished-task); every one found nothing to do, because an idle
pane is the correct condition for those tasks: a done task with its PR open
and green, a task parked on a captain decision with its work committed, a
task blocked on a credential.

fm-classify-lib.sh's crew_absorb_class already performs the one
fm-crew-state.sh read on that path, so the settled verdict is a fourth token
from that same read rather than a new call:

  done            nothing is left for the crew to do on its own.
  parked, blocked the next move belongs above the crew, but ONLY while a
                  durable open decision (status_open_decisions) proves it.

Both supervisors consume that one verdict. The watcher absorbs a settled
stale with no wedge timer and records .settled-<key> so unchanged polls of
the same pane hash absorb without re-reading the crew state; any pane change
drops the marker and reclassifies. The away daemon self-handles the wake,
drops rather than records wedge tracking, and re-checks the reconciled state
immediately before escalating an aged marker, which is where a task that
settles after going stale would otherwise still alarm.

Genuine wedges are unaffected:
- failed is deliberately NOT settled, because it also reconciles a CANCELLED
  run, the mid-supersession state in which a crew must recover custody and
  resume, so an idle pane there is a real stall.
- unknown never settles, so a dead endpoint or torn-down worktree keeps aging.
- a run parked at a gate the crew must answer ITSELF opens no decision, so it
  keeps aging and still escalates.
- reading the reconciled state rather than the log text preserves the earlier
  precedence fix: a leftover pre-validation done: line under an active run
  reconciles as working and keeps its wedge timer.

Nothing rots as a result: an absorb does not mark the status surfaced, so the
terminal status still reaches firstmate through the captain-relevant signal
path and, if that was missed, through the heartbeat catch-all scan, while a
parked or blocked task keeps its open decision surfacing on every wake drain.

Every new test was witnessed failing on unmodified code first.
@sbracewell64
sbracewell64 merged commit 27c7a2d into main Aug 6, 2026
12 of 16 checks passed
sbracewell64 added a commit that referenced this pull request Aug 9, 2026
A task whose reconciled current state is terminal kept producing
possible-wedge stale wakes, because the stale path decided from the status
line's TEXT rather than the reconciled state. Seventeen such wakes were
recorded on 2026-08-04 (joined ledger outcomes, defect
stale-fires-on-finished-task); every one found nothing to do, because an idle
pane is the correct condition for those tasks: a done task with its PR open
and green, a task parked on a captain decision with its work committed, a
task blocked on a credential.

fm-classify-lib.sh's crew_absorb_class already performs the one
fm-crew-state.sh read on that path, so the settled verdict is a fourth token
from that same read rather than a new call:

  done            nothing is left for the crew to do on its own.
  parked, blocked the next move belongs above the crew, but ONLY while a
                  durable open decision (status_open_decisions) proves it.

Both supervisors consume that one verdict. The watcher absorbs a settled
stale with no wedge timer and records .settled-<key> so unchanged polls of
the same pane hash absorb without re-reading the crew state; any pane change
drops the marker and reclassifies. The away daemon self-handles the wake,
drops rather than records wedge tracking, and re-checks the reconciled state
immediately before escalating an aged marker, which is where a task that
settles after going stale would otherwise still alarm.

Genuine wedges are unaffected:
- failed is deliberately NOT settled, because it also reconciles a CANCELLED
  run, the mid-supersession state in which a crew must recover custody and
  resume, so an idle pane there is a real stall.
- unknown never settles, so a dead endpoint or torn-down worktree keeps aging.
- a run parked at a gate the crew must answer ITSELF opens no decision, so it
  keeps aging and still escalates.
- reading the reconciled state rather than the log text preserves the earlier
  precedence fix: a leftover pre-validation done: line under an active run
  reconciles as working and keeps its wedge timer.

Nothing rots as a result: an absorb does not mark the status surfaced, so the
terminal status still reaches firstmate through the captain-relevant signal
path and, if that was missed, through the heartbeat catch-all scan, while a
parked or blocked task keeps its open decision surfacing on every wake drain.

Every new test was witnessed failing on unmodified code first.
sbracewell64 added a commit that referenced this pull request Aug 9, 2026
A task whose reconciled current state is terminal kept producing
possible-wedge stale wakes, because the stale path decided from the status
line's TEXT rather than the reconciled state. Seventeen such wakes were
recorded on 2026-08-04 (joined ledger outcomes, defect
stale-fires-on-finished-task); every one found nothing to do, because an idle
pane is the correct condition for those tasks: a done task with its PR open
and green, a task parked on a captain decision with its work committed, a
task blocked on a credential.

fm-classify-lib.sh's crew_absorb_class already performs the one
fm-crew-state.sh read on that path, so the settled verdict is a fourth token
from that same read rather than a new call:

  done            nothing is left for the crew to do on its own.
  parked, blocked the next move belongs above the crew, but ONLY while a
                  durable open decision (status_open_decisions) proves it.

Both supervisors consume that one verdict. The watcher absorbs a settled
stale with no wedge timer and records .settled-<key> so unchanged polls of
the same pane hash absorb without re-reading the crew state; any pane change
drops the marker and reclassifies. The away daemon self-handles the wake,
drops rather than records wedge tracking, and re-checks the reconciled state
immediately before escalating an aged marker, which is where a task that
settles after going stale would otherwise still alarm.

Genuine wedges are unaffected:
- failed is deliberately NOT settled, because it also reconciles a CANCELLED
  run, the mid-supersession state in which a crew must recover custody and
  resume, so an idle pane there is a real stall.
- unknown never settles, so a dead endpoint or torn-down worktree keeps aging.
- a run parked at a gate the crew must answer ITSELF opens no decision, so it
  keeps aging and still escalates.
- reading the reconciled state rather than the log text preserves the earlier
  precedence fix: a leftover pre-validation done: line under an active run
  reconciles as working and keeps its wedge timer.

Nothing rots as a result: an absorb does not mark the status surfaced, so the
terminal status still reaches firstmate through the captain-relevant signal
path and, if that was missed, through the heartbeat catch-all scan, while a
parked or blocked task keeps its open decision surfacing on every wake drain.

Every new test was witnessed failing on unmodified code first.
sbracewell64 added a commit that referenced this pull request Aug 10, 2026
A task whose reconciled current state is terminal kept producing
possible-wedge stale wakes, because the stale path decided from the status
line's TEXT rather than the reconciled state. Seventeen such wakes were
recorded on 2026-08-04 (joined ledger outcomes, defect
stale-fires-on-finished-task); every one found nothing to do, because an idle
pane is the correct condition for those tasks: a done task with its PR open
and green, a task parked on a captain decision with its work committed, a
task blocked on a credential.

fm-classify-lib.sh's crew_absorb_class already performs the one
fm-crew-state.sh read on that path, so the settled verdict is a fourth token
from that same read rather than a new call:

  done            nothing is left for the crew to do on its own.
  parked, blocked the next move belongs above the crew, but ONLY while a
                  durable open decision (status_open_decisions) proves it.

Both supervisors consume that one verdict. The watcher absorbs a settled
stale with no wedge timer and records .settled-<key> so unchanged polls of
the same pane hash absorb without re-reading the crew state; any pane change
drops the marker and reclassifies. The away daemon self-handles the wake,
drops rather than records wedge tracking, and re-checks the reconciled state
immediately before escalating an aged marker, which is where a task that
settles after going stale would otherwise still alarm.

Genuine wedges are unaffected:
- failed is deliberately NOT settled, because it also reconciles a CANCELLED
  run, the mid-supersession state in which a crew must recover custody and
  resume, so an idle pane there is a real stall.
- unknown never settles, so a dead endpoint or torn-down worktree keeps aging.
- a run parked at a gate the crew must answer ITSELF opens no decision, so it
  keeps aging and still escalates.
- reading the reconciled state rather than the log text preserves the earlier
  precedence fix: a leftover pre-validation done: line under an active run
  reconciles as working and keeps its wedge timer.

Nothing rots as a result: an absorb does not mark the status surfaced, so the
terminal status still reaches firstmate through the captain-relevant signal
path and, if that was missed, through the heartbeat catch-all scan, while a
parked or blocked task keeps its open decision surfacing on every wake drain.

Every new test was witnessed failing on unmodified code first.
sbracewell64 added a commit that referenced this pull request Aug 11, 2026
A task whose reconciled current state is terminal kept producing
possible-wedge stale wakes, because the stale path decided from the status
line's TEXT rather than the reconciled state. Seventeen such wakes were
recorded on 2026-08-04 (joined ledger outcomes, defect
stale-fires-on-finished-task); every one found nothing to do, because an idle
pane is the correct condition for those tasks: a done task with its PR open
and green, a task parked on a captain decision with its work committed, a
task blocked on a credential.

fm-classify-lib.sh's crew_absorb_class already performs the one
fm-crew-state.sh read on that path, so the settled verdict is a fourth token
from that same read rather than a new call:

  done            nothing is left for the crew to do on its own.
  parked, blocked the next move belongs above the crew, but ONLY while a
                  durable open decision (status_open_decisions) proves it.

Both supervisors consume that one verdict. The watcher absorbs a settled
stale with no wedge timer and records .settled-<key> so unchanged polls of
the same pane hash absorb without re-reading the crew state; any pane change
drops the marker and reclassifies. The away daemon self-handles the wake,
drops rather than records wedge tracking, and re-checks the reconciled state
immediately before escalating an aged marker, which is where a task that
settles after going stale would otherwise still alarm.

Genuine wedges are unaffected:
- failed is deliberately NOT settled, because it also reconciles a CANCELLED
  run, the mid-supersession state in which a crew must recover custody and
  resume, so an idle pane there is a real stall.
- unknown never settles, so a dead endpoint or torn-down worktree keeps aging.
- a run parked at a gate the crew must answer ITSELF opens no decision, so it
  keeps aging and still escalates.
- reading the reconciled state rather than the log text preserves the earlier
  precedence fix: a leftover pre-validation done: line under an active run
  reconciles as working and keeps its wedge timer.

Nothing rots as a result: an absorb does not mark the status surfaced, so the
terminal status still reaches firstmate through the captain-relevant signal
path and, if that was missed, through the heartbeat catch-all scan, while a
parked or blocked task keeps its open decision surfacing on every wake drain.

Every new test was witnessed failing on unmodified code first.
sbracewell64 added a commit that referenced this pull request Aug 11, 2026
A task whose reconciled current state is terminal kept producing
possible-wedge stale wakes, because the stale path decided from the status
line's TEXT rather than the reconciled state. Seventeen such wakes were
recorded on 2026-08-04 (joined ledger outcomes, defect
stale-fires-on-finished-task); every one found nothing to do, because an idle
pane is the correct condition for those tasks: a done task with its PR open
and green, a task parked on a captain decision with its work committed, a
task blocked on a credential.

fm-classify-lib.sh's crew_absorb_class already performs the one
fm-crew-state.sh read on that path, so the settled verdict is a fourth token
from that same read rather than a new call:

  done            nothing is left for the crew to do on its own.
  parked, blocked the next move belongs above the crew, but ONLY while a
                  durable open decision (status_open_decisions) proves it.

Both supervisors consume that one verdict. The watcher absorbs a settled
stale with no wedge timer and records .settled-<key> so unchanged polls of
the same pane hash absorb without re-reading the crew state; any pane change
drops the marker and reclassifies. The away daemon self-handles the wake,
drops rather than records wedge tracking, and re-checks the reconciled state
immediately before escalating an aged marker, which is where a task that
settles after going stale would otherwise still alarm.

Genuine wedges are unaffected:
- failed is deliberately NOT settled, because it also reconciles a CANCELLED
  run, the mid-supersession state in which a crew must recover custody and
  resume, so an idle pane there is a real stall.
- unknown never settles, so a dead endpoint or torn-down worktree keeps aging.
- a run parked at a gate the crew must answer ITSELF opens no decision, so it
  keeps aging and still escalates.
- reading the reconciled state rather than the log text preserves the earlier
  precedence fix: a leftover pre-validation done: line under an active run
  reconciles as working and keeps its wedge timer.

Nothing rots as a result: an absorb does not mark the status surfaced, so the
terminal status still reaches firstmate through the captain-relevant signal
path and, if that was missed, through the heartbeat catch-all scan, while a
parked or blocked task keeps its open decision surfacing on every wake drain.

Every new test was witnessed failing on unmodified code first.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant