Skip to content

fix(supervision): defer stale wedge alarms for actively working crews - #18

Merged
knowttl merged 1 commit into
mainfrom
fm/fm-restore-wedge-defer-14
Aug 7, 2026
Merged

knowttl merged 1 commit into
mainfrom
fm/fm-restore-wedge-defer-14

Conversation

@knowttl

@knowttl knowttl commented Aug 7, 2026

Copy link
Copy Markdown
Owner

Intent

Restore firstmate fix #14 (commit fc69527, 'fix: defer stale wedge alarms for active crews'), which a later upstream sync silently dropped from the fork's default branch: its markers wedge_escalation_deferred and FM_WEDGE_WORKING_ESCALATE_SECS were absent from current main, so the watcher and the away-mode daemon again raise false 'possible wedge' / stale escalations against HEALTHY crews that are legitimately busy on a long gate (e.g. a long test suite or pipeline step) and therefore look idle - wasting supervision cycles and, in away mode, pestering the captain.

The task, as specified: re-integrate the defer-for-actively-working-crews behavior on top of CURRENT origin/main (which now includes the just-merged upstream sync, so the surrounding supervision code has changed) rather than blindly reverting; use git show fc69527 as the reference implementation; find where the wedge-alarm / stale-escalation logic now lives after the sync (bin/fm-watch.sh, bin/fm-classify-lib.sh, and the away-mode / sub-supervisor path) and re-apply it there, including the FM_WEDGE_WORKING_ESCALATE_SECS knob and the wedge_escalation_deferred path. PRESERVE genuine wedge detection: a truly stuck crew must STILL escalate after the bound. The fix only DEFERS alarms for crews provably actively working a known/declared long gate, on a bounded cadence - it must never mask a real wedge; that boundedness is the whole safety property. Restore or re-create the regression test proving BOTH: (a) a healthy actively-working crew on a long gate does NOT trigger a false wedge escalation within the defer window, and (b) a genuinely stuck crew STILL escalates after the bound, fail-first where practical.

Constraints stated by the captain: base on latest origin/main; repo style is one sentence per line in Markdown, plain dash (never an em dash), and no agent co-author on commits; bin/fm-lint.sh must pass; tests colocated in tests/ as .test.sh. The PR must be opened against knowttl/firstmate base main - NEVER the upstream kunchenguid repo and never another remote. GitHub Actions is disabled on this repo under a standing outage posture, so this pipeline's own suite run IS the verification and an empty forge check list is expected. Do NOT merge: firstmate merges separately on the captain's pre-authorization once the result is clean, and a red result is never merged.

Decisions and tradeoffs made while doing the work, which a reviewer reading only the diff would not know:

  1. This is deliberately a re-integration, not a revert. The fc69527 cherry-pick conflicted in AGENTS.md, bin/fm-watch.sh, docs/architecture.md, docs/configuration.md, and tests/fm-daemon.test.sh. Every conflict was resolved by KEEPING the newer upstream content and folding only the wedge-deferral change on top. In particular, fc69527's second commit (which deleted two fork-only Herdr native-idle supervisor-guard tests) was NOT re-applied: current main has since replaced that region with test_primary_busy_guard_is_harness_scoped, so HEAD was kept verbatim there. That old test deletion is out of scope for this task and was intentionally dropped.

  2. One genuine behavioral integration decision beyond a plain restore. The upstream sync added BUSY_TURN_MAX_SECS: a busy pane is normally unconditional proof of liveness, so that bound routes a busy pane with no completed turn through the SAME wedge_timer_check precisely because the pane's busy signal is no longer trusted past the bound. But a busy pane reports 'state: working / source: pane' to bin/fm-crew-state.sh, which is exactly what the deferral's re-verification (crew_is_provably_working) reads. Applying the deferral unchanged there would have re-affirmed the very signal BUSY_TURN_MAX_SECS exists to stop trusting, silently extending that bound from 3600s to 3600s plus the whole 3600s allowance and letting a hung foreground call stay hidden roughly twice as long. So wedge_timer_check gained an optional fifth 'nodefer' argument and the two busy-turn call sites pass it; every other call site keeps the deferral. The rationale is documented in the fm-watch.sh header and in docs/architecture.md. This was a deliberate choice to protect an existing guarantee, not an oversight.

  3. Ownership follows the original design: bin/fm-classify-lib.sh owns the wedge_escalation_deferred decision for BOTH the always-on watcher and the away-mode daemon, and each consumer owns only its own verification marker (state/.wedge-verified- for the watcher, state/.subsuper-wedge-verified- for the daemon), so the costly crew-state read happens at most once per threshold window per pane. FM_STALE_ESCALATE_SECS now selects re-examination rather than alarm; FM_WEDGE_WORKING_ESCALATE_SECS (default 3600s) is the ceiling and 0 disables the deferral entirely.

  4. Verification actually performed: tests/fm-watch-triage.test.sh, tests/fm-daemon.test.sh (101 passing), and tests/fm-wake-queue.test.sh all run green. Fail-first was confirmed twice - the new test_busy_turn_age_escalation_is_exempt_from_working_deferral fails with the nodefer argument removed, and the restored classifier assertions fail against a wedge_escalation_deferred stubbed to always escalate. bin/fm-doc-audience-check.sh passes. bin/fm-lint.sh could NOT be run in this environment because shellcheck 0.11.0 is not installed and fm-lint.sh correctly refuses to run under any other version; every touched script parses clean under bash -n, and this pipeline's lint step is the real check.

What Changed

  • Added wedge_escalation_deferred to bin/fm-classify-lib.sh as the single owner of the wedge-deferral policy: at the stale threshold it re-verifies the crew via crew_is_provably_working and defers the alarm for another threshold window, bounded by the new FM_WEDGE_WORKING_ESCALATE_SECS knob (default 3600s, 0 disables) so a genuinely stuck crew still escalates once the ceiling is reached.
  • Wired the deferral into both consumers, each owning its own verification marker so the crew read happens at most once per threshold window per pane: the always-on watcher (bin/fm-watch.sh, .wedge-verified-<key>) and the away-mode daemon (bin/fm-supervise-daemon.sh, .subsuper-wedge-verified-<key>), with FM_STALE_ESCALATE_SECS now selecting re-examination rather than immediate alarm and the markers cleared on resume/pause.
  • Gave wedge_timer_check an optional nodefer fifth argument used by the two busy-turn call sites, so a busy pane past BUSY_TURN_MAX_SECS escalates without re-affirming the busy signal that bound exists to stop trusting; extended tests/fm-watch-triage.test.sh and tests/fm-daemon.test.sh to prove a working crew is deferred within the window, a stuck crew still escalates after the bound, and the busy-turn path stays exempt from the deferral, with the rationale documented in the script headers, docs/architecture.md, and docs/configuration.md.

Risk Assessment

✅ Low: The change is a well-bounded re-integration whose critical safety invariant (a truly stuck crew still escalates at the ceiling) is correctly enforced by ordering the ceiling check before the cached-defer short-circuit, with thorough marker cleanup and behavioral regression tests covering both the defer and escalate paths.

Testing

Ran the two regression suites that own this change as targeted subsets (avoiding the full triage suite, which exceeds the timeout driving many real fm-watch.sh subprocesses): five watcher/classifier tests and two away-mode daemon tests all pass, proving a provably-working crew on a long gate is deferred within the allowance while a stopped or over-ceiling crew still escalates. I verified the boundedness safety property is genuinely guarded by sabotaging the deferral to always defer, which made the daemon test fail on the exact masked-wedge assertion, then restored the source (worktree clean). This is a CLI/shell supervision change with no rendered UI surface, so evidence is the captured test transcript rather than a screenshot. Transient test runners I created were removed and the worktree is clean.

Evidence: Targeted regression transcript (watcher/classifier + away-mode daemon deferral)

## A) Watcher + classifier ok - wedge_escalation_deferred: defers a working crew inside the allowance, escalates past it or once it stops ok - a live in-contract pipeline call is absorbed across repeated threshold windows instead of wedge-escalated ok - consecutive wedge escalations on the same pane accumulate and demand deep inspection at the threshold ok - a pane becoming active again resets the consecutive wedge-escalation counter ok - a busy pane past its completed-turn bound escalates even while the crew still reads as provably working ## B) Away-mode daemon ok - away mode defers a wedge alarm while an in-contract call verifies as working, and escalates once it stops ok - away mode escalates a frozen in-contract call once it passes the long-call allowance

# Wedge-deferral regression evidence (branch fm/fm-restore-wedge-defer-14, HEAD ca1d800)

## A) Watcher + classifier: healthy busy crew deferred; stuck crew still escalates; busy-turn bound exempt
ok - wedge_escalation_deferred: defers a working crew inside the allowance, escalates past it or once it stops
ok - a live in-contract pipeline call is absorbed across repeated threshold windows instead of wedge-escalated
ok - consecutive wedge escalations on the same pane accumulate and demand deep inspection at the threshold
ok - a pane becoming active again resets the consecutive wedge-escalation counter
ok - a busy pane past its completed-turn bound escalates even while the crew still reads as provably working

## B) Away-mode daemon: defers a working in-contract call, escalates once it stops / passes the ceiling
ok - away mode defers a wedge alarm while an in-contract call verifies as working, and escalates once it stops
ok - away mode escalates a frozen in-contract call once it passes the long-call allowance
Evidence: Fail-first proof (deferral sabotaged to always defer)
not ok - a stopped crew on a frozen pane was not escalated as a possible wedge # after forcing wedge_escalation_deferred to always return 0; source restored afterward

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-watch-triage.test.sh targeted subset: test_wedge_escalation_deferred_classifier, test_in_contract_pipeline_call_is_not_wedge_escalated, test_wedge_escalation_marks_demand_deep_inspection_after_threshold, test_wedge_escalation_resets_when_pane_becomes_active, test_busy_turn_age_escalation_is_exempt_from_working_deferral (all pass)
  • bash tests/fm-daemon.test.sh targeted subset: test_housekeeping_in_contract_stale_defers_then_escalates, test_housekeeping_in_contract_stale_escalates_past_allowance (both pass)
  • Fail-first safety check: patched bin/fm-classify-lib.sh wedge_escalation_deferred to always return 0 (always defer); the away-mode test correctly failed with not ok - a stopped crew on a frozen pane was not escalated as a possible wedge, then restored the file (git diff clean)
  • bash -n bin/fm-classify-lib.sh bin/fm-watch.sh bin/fm-supervise-daemon.sh (parse clean)
✅ **Document** - passed

✅ No issues found.

🔧 **Lint** - 1 issue found → auto-fixed ✅
  • ⚠️ linter found issues (exit code 127)

🔧 Fix: lint passes clean with ShellCheck 0.11.0; no code changes needed
✅ Re-checked - no issues remain.

✅ **Push** - passed

✅ No issues found.

Restores PR #14 (fc69527), which a later upstream sync dropped from the
fork's default branch: the markers wedge_escalation_deferred and
FM_WEDGE_WORKING_ESCALATE_SECS were absent from main, so the watcher and
the away-mode daemon again alarmed every long in-contract call as a
possible wedge.

A crewmate blocked on a long synchronous foreground call is idle by
design: a no-mistakes gate call blocks against its own multi-thousand-
second allowance, so its pane stays frozen far past the 240s threshold.
The threshold now selects re-examination, not alarm. On reaching it the
crew's current state is re-read once per threshold window; a crew still
provably working is deferred for another window, and only a crew that
stops verifying as working - or one that stays idle past
FM_WEDGE_WORKING_ESCALATE_SECS (default 3600s) - escalates. The deferral
trades false alarms for latency, never for silence, and
FM_WEDGE_WORKING_ESCALATE_SECS=0 disables it.

bin/fm-classify-lib.sh owns the decision for both supervisors; each
consumer owns only its own verification marker, so the crew state is read
at most once per threshold window per pane.

Re-integrated on top of the current code rather than reverted. The sync
added BUSY_TURN_MAX_SECS, which routes a busy pane with no completed turn
through the same wedge timer precisely because that pane's busy signal is
no longer trusted. Since a busy pane is exactly what the deferral's
re-verification reads, deferring there would re-affirm the discredited
signal and silently extend the bound by the whole allowance, so
wedge_timer_check takes a nodefer argument and the two busy-turn call
sites pass it.

Tests: the restored fc69527 coverage in tests/fm-watch-triage.test.sh and
tests/fm-daemon.test.sh proves both directions - repeated threshold
windows absorb a live in-contract call, while a stopped crew and a frozen
pipeline past the allowance still escalate. A new watcher test pins the
busy-turn exemption; it was watched failing with the nodefer argument
removed, and the classifier assertions were watched failing against a
wedge_escalation_deferred stubbed to never defer.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant