fix(supervision): defer stale wedge alarms for actively working crews - #18
Merged
Merged
Conversation
Restores PR #14 (fc69527), which a later upstream sync dropped from the fork's default branch: the markers wedge_escalation_deferred and FM_WEDGE_WORKING_ESCALATE_SECS were absent from main, so the watcher and the away-mode daemon again alarmed every long in-contract call as a possible wedge. A crewmate blocked on a long synchronous foreground call is idle by design: a no-mistakes gate call blocks against its own multi-thousand- second allowance, so its pane stays frozen far past the 240s threshold. The threshold now selects re-examination, not alarm. On reaching it the crew's current state is re-read once per threshold window; a crew still provably working is deferred for another window, and only a crew that stops verifying as working - or one that stays idle past FM_WEDGE_WORKING_ESCALATE_SECS (default 3600s) - escalates. The deferral trades false alarms for latency, never for silence, and FM_WEDGE_WORKING_ESCALATE_SECS=0 disables it. bin/fm-classify-lib.sh owns the decision for both supervisors; each consumer owns only its own verification marker, so the crew state is read at most once per threshold window per pane. Re-integrated on top of the current code rather than reverted. The sync added BUSY_TURN_MAX_SECS, which routes a busy pane with no completed turn through the same wedge timer precisely because that pane's busy signal is no longer trusted. Since a busy pane is exactly what the deferral's re-verification reads, deferring there would re-affirm the discredited signal and silently extend the bound by the whole allowance, so wedge_timer_check takes a nodefer argument and the two busy-turn call sites pass it. Tests: the restored fc69527 coverage in tests/fm-watch-triage.test.sh and tests/fm-daemon.test.sh proves both directions - repeated threshold windows absorb a live in-contract call, while a stopped crew and a frozen pipeline past the allowance still escalate. A new watcher test pins the busy-turn exemption; it was watched failing with the nodefer argument removed, and the classifier assertions were watched failing against a wedge_escalation_deferred stubbed to never defer.
This was referenced Aug 14, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Restore firstmate fix #14 (commit fc69527, 'fix: defer stale wedge alarms for active crews'), which a later upstream sync silently dropped from the fork's default branch: its markers wedge_escalation_deferred and FM_WEDGE_WORKING_ESCALATE_SECS were absent from current main, so the watcher and the away-mode daemon again raise false 'possible wedge' / stale escalations against HEALTHY crews that are legitimately busy on a long gate (e.g. a long test suite or pipeline step) and therefore look idle - wasting supervision cycles and, in away mode, pestering the captain.
The task, as specified: re-integrate the defer-for-actively-working-crews behavior on top of CURRENT origin/main (which now includes the just-merged upstream sync, so the surrounding supervision code has changed) rather than blindly reverting; use git show fc69527 as the reference implementation; find where the wedge-alarm / stale-escalation logic now lives after the sync (bin/fm-watch.sh, bin/fm-classify-lib.sh, and the away-mode / sub-supervisor path) and re-apply it there, including the FM_WEDGE_WORKING_ESCALATE_SECS knob and the wedge_escalation_deferred path. PRESERVE genuine wedge detection: a truly stuck crew must STILL escalate after the bound. The fix only DEFERS alarms for crews provably actively working a known/declared long gate, on a bounded cadence - it must never mask a real wedge; that boundedness is the whole safety property. Restore or re-create the regression test proving BOTH: (a) a healthy actively-working crew on a long gate does NOT trigger a false wedge escalation within the defer window, and (b) a genuinely stuck crew STILL escalates after the bound, fail-first where practical.
Constraints stated by the captain: base on latest origin/main; repo style is one sentence per line in Markdown, plain dash (never an em dash), and no agent co-author on commits; bin/fm-lint.sh must pass; tests colocated in tests/ as .test.sh. The PR must be opened against knowttl/firstmate base main - NEVER the upstream kunchenguid repo and never another remote. GitHub Actions is disabled on this repo under a standing outage posture, so this pipeline's own suite run IS the verification and an empty forge check list is expected. Do NOT merge: firstmate merges separately on the captain's pre-authorization once the result is clean, and a red result is never merged.
Decisions and tradeoffs made while doing the work, which a reviewer reading only the diff would not know:
This is deliberately a re-integration, not a revert. The fc69527 cherry-pick conflicted in AGENTS.md, bin/fm-watch.sh, docs/architecture.md, docs/configuration.md, and tests/fm-daemon.test.sh. Every conflict was resolved by KEEPING the newer upstream content and folding only the wedge-deferral change on top. In particular, fc69527's second commit (which deleted two fork-only Herdr native-idle supervisor-guard tests) was NOT re-applied: current main has since replaced that region with test_primary_busy_guard_is_harness_scoped, so HEAD was kept verbatim there. That old test deletion is out of scope for this task and was intentionally dropped.
One genuine behavioral integration decision beyond a plain restore. The upstream sync added BUSY_TURN_MAX_SECS: a busy pane is normally unconditional proof of liveness, so that bound routes a busy pane with no completed turn through the SAME wedge_timer_check precisely because the pane's busy signal is no longer trusted past the bound. But a busy pane reports 'state: working / source: pane' to bin/fm-crew-state.sh, which is exactly what the deferral's re-verification (crew_is_provably_working) reads. Applying the deferral unchanged there would have re-affirmed the very signal BUSY_TURN_MAX_SECS exists to stop trusting, silently extending that bound from 3600s to 3600s plus the whole 3600s allowance and letting a hung foreground call stay hidden roughly twice as long. So wedge_timer_check gained an optional fifth 'nodefer' argument and the two busy-turn call sites pass it; every other call site keeps the deferral. The rationale is documented in the fm-watch.sh header and in docs/architecture.md. This was a deliberate choice to protect an existing guarantee, not an oversight.
Ownership follows the original design: bin/fm-classify-lib.sh owns the wedge_escalation_deferred decision for BOTH the always-on watcher and the away-mode daemon, and each consumer owns only its own verification marker (state/.wedge-verified- for the watcher, state/.subsuper-wedge-verified- for the daemon), so the costly crew-state read happens at most once per threshold window per pane. FM_STALE_ESCALATE_SECS now selects re-examination rather than alarm; FM_WEDGE_WORKING_ESCALATE_SECS (default 3600s) is the ceiling and 0 disables the deferral entirely.
Verification actually performed: tests/fm-watch-triage.test.sh, tests/fm-daemon.test.sh (101 passing), and tests/fm-wake-queue.test.sh all run green. Fail-first was confirmed twice - the new test_busy_turn_age_escalation_is_exempt_from_working_deferral fails with the nodefer argument removed, and the restored classifier assertions fail against a wedge_escalation_deferred stubbed to always escalate. bin/fm-doc-audience-check.sh passes. bin/fm-lint.sh could NOT be run in this environment because shellcheck 0.11.0 is not installed and fm-lint.sh correctly refuses to run under any other version; every touched script parses clean under bash -n, and this pipeline's lint step is the real check.
What Changed
wedge_escalation_deferredtobin/fm-classify-lib.shas the single owner of the wedge-deferral policy: at the stale threshold it re-verifies the crew viacrew_is_provably_workingand defers the alarm for another threshold window, bounded by the newFM_WEDGE_WORKING_ESCALATE_SECSknob (default 3600s, 0 disables) so a genuinely stuck crew still escalates once the ceiling is reached.bin/fm-watch.sh,.wedge-verified-<key>) and the away-mode daemon (bin/fm-supervise-daemon.sh,.subsuper-wedge-verified-<key>), withFM_STALE_ESCALATE_SECSnow selecting re-examination rather than immediate alarm and the markers cleared on resume/pause.wedge_timer_checkan optionalnodeferfifth argument used by the two busy-turn call sites, so a busy pane pastBUSY_TURN_MAX_SECSescalates without re-affirming the busy signal that bound exists to stop trusting; extendedtests/fm-watch-triage.test.shandtests/fm-daemon.test.shto prove a working crew is deferred within the window, a stuck crew still escalates after the bound, and the busy-turn path stays exempt from the deferral, with the rationale documented in the script headers,docs/architecture.md, anddocs/configuration.md.Risk Assessment
✅ Low: The change is a well-bounded re-integration whose critical safety invariant (a truly stuck crew still escalates at the ceiling) is correctly enforced by ordering the ceiling check before the cached-defer short-circuit, with thorough marker cleanup and behavioral regression tests covering both the defer and escalate paths.
Testing
Ran the two regression suites that own this change as targeted subsets (avoiding the full triage suite, which exceeds the timeout driving many real fm-watch.sh subprocesses): five watcher/classifier tests and two away-mode daemon tests all pass, proving a provably-working crew on a long gate is deferred within the allowance while a stopped or over-ceiling crew still escalates. I verified the boundedness safety property is genuinely guarded by sabotaging the deferral to always defer, which made the daemon test fail on the exact masked-wedge assertion, then restored the source (worktree clean). This is a CLI/shell supervision change with no rendered UI surface, so evidence is the captured test transcript rather than a screenshot. Transient test runners I created were removed and the worktree is clean.
Evidence: Targeted regression transcript (watcher/classifier + away-mode daemon deferral)
## A) Watcher + classifier ok - wedge_escalation_deferred: defers a working crew inside the allowance, escalates past it or once it stops ok - a live in-contract pipeline call is absorbed across repeated threshold windows instead of wedge-escalated ok - consecutive wedge escalations on the same pane accumulate and demand deep inspection at the threshold ok - a pane becoming active again resets the consecutive wedge-escalation counter ok - a busy pane past its completed-turn bound escalates even while the crew still reads as provably working ## B) Away-mode daemon ok - away mode defers a wedge alarm while an in-contract call verifies as working, and escalates once it stops ok - away mode escalates a frozen in-contract call once it passes the long-call allowanceEvidence: Fail-first proof (deferral sabotaged to always defer)
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ **Review** - passed
✅ No issues found.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-watch-triage.test.shtargeted subset:test_wedge_escalation_deferred_classifier,test_in_contract_pipeline_call_is_not_wedge_escalated,test_wedge_escalation_marks_demand_deep_inspection_after_threshold,test_wedge_escalation_resets_when_pane_becomes_active,test_busy_turn_age_escalation_is_exempt_from_working_deferral(all pass)bash tests/fm-daemon.test.shtargeted subset:test_housekeeping_in_contract_stale_defers_then_escalates,test_housekeeping_in_contract_stale_escalates_past_allowance(both pass)Fail-first safety check: patchedbin/fm-classify-lib.shwedge_escalation_deferredto alwaysreturn 0(always defer); the away-mode test correctly failed withnot ok - a stopped crew on a frozen pane was not escalated as a possible wedge, then restored the file (git diff clean)bash -n bin/fm-classify-lib.sh bin/fm-watch.sh bin/fm-supervise-daemon.sh(parse clean)✅ **Document** - passed
✅ No issues found.
🔧 **Lint** - 1 issue found → auto-fixed ✅
🔧 Fix: lint passes clean with ShellCheck 0.11.0; no code changes needed
✅ Re-checked - no issues remain.
✅ **Push** - passed
✅ No issues found.