fix(watch): absorb declared-wait and running-validation wakes instead of forcing no-op turns - #37
Merged
Merged
Conversation
…ng validations A declared external wait (paused:/captain-held) and a still-running no-mistakes validation no longer force a supervision turn - the burn a Grok primary showed on 2026-08-15, where a healthy worker sitting on a long validation step or a declared captain-wait produced a stream of no-op "Captain, shipshape." replies. Two native watcher changes: - Signal path: a paused:/captain-held status append and the bare turn-end under it are absorbed as a declared external wait when the crew is not provably working, matching the stale path (which already absorbs declared waits) and the codified captain preference against alerts on declared captain-wait windows. The swallowed-finish guard is preserved: a crew that neither declared a wait nor is provably working still surfaces. - Stale path: a provably-working stale is re-verified against the authoritative current run state at each FM_STALE_ESCALATE_SECS boundary instead of blind-escalating on elapsed time. A run fm-crew-state.sh still reports as actively working legitimately sits on a static pane for a long CI/validation wait and keeps absorbing silently; only a run that reached a gate, finished, failed, or went unreadable escalates, so genuine stuck-worker detection survives. Because the watcher keeps blocking on these absorbs, the Grok monitor arm (bin/fm-watch-arm.sh, whose exit becomes a synthetic primary turn) never exits, proven end to end by a real arm + real watcher regression.
…y escalation gate
withally
added a commit
that referenced
this pull request
Aug 15, 2026
… survive compaction (#38) * fix(watch): hold the sitting wait for a live quiet crew and restore read-once after a compaction Two halves of one goal: the primary's sitting wait should end only when the captain would care, and a compaction should not force the primary to re-read the startup dump to remember it. Half A - the sitting wait ends only for a captain-relevant event. Every wake() exits the watcher, and on the Grok primary that exit becomes a synthetic turn, so each non-captain-relevant wake is a no-op turn. PR #37 absorbed declared waits, running validation, and empty rearms; the residual leak was the fail-safe immediate surface of a crew that merely went quiet - present identically on the signal path (a bare turn-end) and the stale path (the none class). It is one principle across both paths, not two more absorb types: a crew that is not provably working, declared no wait, has no captain-relevant status line, and whose agent is still live has only gone quiet between steps. The attended primary now absorbs it and hands it to the existing wedge timer, which surfaces it as a possible wedge only after STALE_ESCALATE_SECS, while a crew whose agent has confidently exited (fm_backend_agent_state dead/missing) still surfaces at once as a swallowed finish. Away mode is unchanged - its daemon keeps classifying every ambiguous signal. Captain-relevant lines, genuine wedges, genuine finishes, secondmate reports, and unknown/unmatched-decision status lines all still surface; the heartbeat backstop remains the safety net. Half B - read-once survives a compaction. A run-tier harness re-emits the digest on a clear/compact source, but Grok fires no session-open hook on an in-session compaction and discards SessionStart stdout, so nothing restores the dump and the primary re-reads captain.md/learnings.md file by file. The read-once contract in AGENTS.md section 3 (re-sent every turn, so it survives the compaction) now directs one fm-session-start.sh --reemit to restore the fleet state and durable dump, and the digest's own read-once block and docs/sessionstart-nudge.md carry the same restore path and the Grok limitation. Upstream (kunchenguid) had nothing importable that reaches this: the Grok turn-end guard is byte-identical, its wait/arm lineage diverged and is behind the fork's clean-close receipt, and its compaction refresh (kunchenguid#2163) is Pi-only. Regression tests drive a real fm-watch.sh and a real fm-session-start.sh --reemit, not stubs. * no-mistakes(review): gate quiet-crew absorb on a recorded window; fix stale secondmate comment * no-mistakes(document): align watcher-surface docs with quiet-crew absorb * no-mistakes: apply CI fixes
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Stop the no-op "Captain, shipshape." supervision turns a Grok primary showed on 2026-08-15, where a healthy worker sitting on a long validation step or in a declared captain-wait forced a stream of no-op primary turns. Goal: a healthy worker waiting on a declared external wait (paused:/captain-held) or a still-running validation step (a provably-working no-mistakes pipeline pane) must NOT force a Grok primary turn, while every real wake still lands unchanged - review-ready PRs, failures, blockers, needs-decision, credential needs, and genuinely wedged/unresponsive workers.
Grok-specific mechanism accounted for: on Grok the monitor path is a tracked background arm (bin/fm-watch-arm.sh) whose EXIT becomes a synthetic primary turn. The fix therefore makes the WATCHER keep blocking (absorb) on these benign wakes so the arm never exits - not merely relabel a wake. Proven end to end by a new real-arm + real-watcher regression (tests/fm-watch-arm.test.sh) asserting a declared-wait absorb leaves both processes alive while a genuine captain-relevant status still exits both.
This is a NATIVE fork fix, not an import. Upstream (kunchenguid) has no single fix for this class; the closest commits are c42cfe0 (recovery-ack livelock) and b0ad61e (duplicate-wake dedup), neither of which is the shipshape absorb burn.
Two changes, both in the always-on watcher, mirroring the away-mode daemon's proven semantics:
Deliberately EXCLUDED from this change (reported for a separate follow-up, not a gap to fix here): the third symptom in the evidence, an "empty monitor rearm" (check: rearm-resurface) from overlapping monitor cycles. An empty-suppress was prototyped and reverted because it swallowed genuine post-downtime recovery (a real wake), which violates the hard "every real wake still lands" requirement - proven by fm-watcher-lock and fm-wake-daemon-lifecycle-e2e, where a done: status that lands while the watcher is down must still surface via recovery. The correct fix is porting upstream c42cfe0's recovery-ack livelock fix (make the acknowledgement survive republication), a focused change to the guarded durable-wake recovery contract that is out of scope for this shipshape-absorb burn and borders the "do not redesign the supervision architecture" boundary. On a Grok primary specifically the overlapping-cycle rearm originates from a Claude stop-hook repair, so it is primarily a Claude-transition artifact rather than a steady-state Grok burn.
Existing regression tests in tests/fm-watch-triage.test.sh that encoded the old blind-240s wedge (escalate a still-provably-working stale purely on elapsed time) were updated to the corrected re-verify semantics and given an added phase proving the genuine wedge still fires once the crew is no longer provably working; this is a strengthening of the contract, not a weakening, and the captain authorized removing the false wedge. New coverage: the signal_all_declared_wait classifier, declared-paused and captain-held signal absorbs, and the real-arm no-synthetic-turn integration test. docs/architecture.md updated to keep the one-owner absorb-contract description accurate. bin/fm-lint.sh is clean.
Constraints: firstmate shared tracked material only (firstmate-coding-guidelines followed); no captain-facing copy changed beyond stopping the no-op turn itself; the no-blind-turn supervision contract, the single supervision-cycle model, and resurface-default-off are all preserved.
What Changed
bin/fm-watch.sh+ newsignal_all_declared_waitinbin/fm-classify-lib.sh): apaused:/captain-held status append and the bare turn-end under it are now absorbed as a declared external wait when the crew is not provably working, so the watcher keeps blocking and the tracked background arm never exits into a synthetic primary turn; the swallowed-finish guard still surfaces a crew that neither declared a wait nor is provably working.wedge_timer_checkgains an optional reverify-task): a provably-working stale is re-verified against the current run state at eachFM_STALE_ESCALATE_SECSboundary (bounded to once per threshold per window) instead of blind-escalating on elapsed time — an actively-running validation pane resets its timer and keeps absorbing, while a run that reached a gate, finished, failed, or went unreadable still escalates; the busy-pane turn-age wedge is intentionally left un-reverified.tests/fm-watch-arm.test.sh) asserting a declared-wait absorb leaves both processes alive while a captain-relevant status exits both, extendedtests/fm-watch-triage.test.shfor the new classifier and re-verify semantics, and updateddocs/architecture.md/docs/configuration.mdto match the absorb contract.Risk Assessment
✅ Low: A tightly scoped, well-tested watcher fix that adds two absorb guards behind the existing captain-relevant/provably-working checks, preserving the swallowed-finish and "every real wake lands" invariants and backed by real-arm and re-verify regression tests.
Testing
Ran the two targeted watcher suites the change adds/updates. tests/fm-watch-arm.test.sh (8/8) drives a real fm-watch.sh watcher plus a real attached fm-watch-arm.sh arm end to end and proves a declared-wait absorb keeps BOTH processes alive — so no synthetic Grok primary turn — while a genuine done: status still exits both and reports the signal. tests/fm-watch-triage.test.sh (62/62) covers the new signal_all_declared_wait classifier, declared-paused and captain-held signal absorbs when the crew is not provably working, the preserved swallowed-finish guard, and the stale-path re-verify that keeps absorbing a provably-working long validation wait but still fires the wedge escalation once the run is no longer provably working. All pass, zero failures; a CLI transcript evidence file was captured. No rendered UI surface exists for this supervision-daemon change, so evidence is the real-process CLI transcript rather than a screenshot.
Evidence: Real-arm+watcher absorb-burn test transcript
### fm-watch-arm.test.sh (real arm + real watcher) ### ok - watch-arm: a declared-wait absorb keeps the real attached arm alive (no synthetic turn); a genuine wake still exits it ### fm-watch-triage.test.sh - intent-critical assertions ### ok - signal_all_declared_wait: only all-paused/captain-held batches (status or turn-end) are declared waits ok - a bare turn-end whose crew is not provably working is surfaced (the swallowed-finish fix) ok - a declared paused: signal (status append plus turn-end) is absorbed even when the crew is not provably working ok - a verified captain-held signal is absorbed even when the crew is not provably working ok - a stale terminal-looking status overridden by an active run is re-verified past the threshold instead of burning a turn, and escalates only once no longer provably working ok - provably-working non-terminal stale is absorbed, re-verified working past the threshold instead of burning a turn, and escalates only once no longer provably workingPipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ **Review** - passed
✅ No issues found.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-watch-arm.test.sh— 8/8 pass, incl. real-arm+real-watchertest_attached_arm_stays_live_through_a_declared_wait_absorb(declared-wait absorb keeps both processes alive; genuine wake exits both)bash tests/fm-watch-triage.test.sh— 62/62 pass, incl.signal_all_declared_waitclassifier, declared-paused/captain-held signal absorbs while crew not provably working, swallowed-finish guard surfaces a not-working turn-end, and stale re-verify (...re-verified past the threshold instead of burning a turn, and escalates only once no longer provably working)✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.