Skip to content

fix(watch): absorb declared-wait and running-validation wakes instead of forcing no-op turns - #37

Merged
withally merged 2 commits into
mainfrom
fm/fm-shipshape-wake-burn-x1
Aug 15, 2026
Merged

withally merged 2 commits into
mainfrom
fm/fm-shipshape-wake-burn-x1

Conversation

@withally

Copy link
Copy Markdown
Owner

Intent

Stop the no-op "Captain, shipshape." supervision turns a Grok primary showed on 2026-08-15, where a healthy worker sitting on a long validation step or in a declared captain-wait forced a stream of no-op primary turns. Goal: a healthy worker waiting on a declared external wait (paused:/captain-held) or a still-running validation step (a provably-working no-mistakes pipeline pane) must NOT force a Grok primary turn, while every real wake still lands unchanged - review-ready PRs, failures, blockers, needs-decision, credential needs, and genuinely wedged/unresponsive workers.

Grok-specific mechanism accounted for: on Grok the monitor path is a tracked background arm (bin/fm-watch-arm.sh) whose EXIT becomes a synthetic primary turn. The fix therefore makes the WATCHER keep blocking (absorb) on these benign wakes so the arm never exits - not merely relabel a wake. Proven end to end by a new real-arm + real-watcher regression (tests/fm-watch-arm.test.sh) asserting a declared-wait absorb leaves both processes alive while a genuine captain-relevant status still exits both.

This is a NATIVE fork fix, not an import. Upstream (kunchenguid) has no single fix for this class; the closest commits are c42cfe0 (recovery-ack livelock) and b0ad61e (duplicate-wake dedup), neither of which is the shipshape absorb burn.

Two changes, both in the always-on watcher, mirroring the away-mode daemon's proven semantics:

  1. Signal path (bin/fm-watch.sh + new signal_all_declared_wait in bin/fm-classify-lib.sh): a paused:/captain-held status append and the bare turn-end under it are absorbed as a declared external wait when the crew is not provably working. This brings the signal path in line with the stale path, which already absorbs declared waits (pause_state_class), and honors the codified captain preference (data/captain.md, 2026-08-03) that a declared captain-wait window raises NO monitoring alert and resurface stays default-off. The swallowed-finish guard is deliberately preserved: a crew that neither declared a wait nor is provably working still surfaces immediately.
  2. Stale path (wedge_timer_check gains an optional reverify-task): a provably-working stale is re-verified against the authoritative current run state (fm-crew-state.sh) at each FM_STALE_ESCALATE_SECS boundary instead of blind-escalating on elapsed time. A run still reported as actively working via an active run-step legitimately sits on a static pane for a long CI/validation wait and keeps absorbing silently (timer reset), never forcing a no-op turn; only a run that reached a gate, finished, failed, or went unreadable escalates. Genuine stuck-worker detection survives: the escalation, its consecutive-count, and the demand-deep-inspection marker still fire once the crew is no longer provably working, and the busy-pane turn-age wedge path is intentionally NOT re-verified (a busy pane's only liveness evidence is its footer, the signal that bound exists to distrust). The re-read is bounded to once per threshold per window, not every poll.

Deliberately EXCLUDED from this change (reported for a separate follow-up, not a gap to fix here): the third symptom in the evidence, an "empty monitor rearm" (check: rearm-resurface) from overlapping monitor cycles. An empty-suppress was prototyped and reverted because it swallowed genuine post-downtime recovery (a real wake), which violates the hard "every real wake still lands" requirement - proven by fm-watcher-lock and fm-wake-daemon-lifecycle-e2e, where a done: status that lands while the watcher is down must still surface via recovery. The correct fix is porting upstream c42cfe0's recovery-ack livelock fix (make the acknowledgement survive republication), a focused change to the guarded durable-wake recovery contract that is out of scope for this shipshape-absorb burn and borders the "do not redesign the supervision architecture" boundary. On a Grok primary specifically the overlapping-cycle rearm originates from a Claude stop-hook repair, so it is primarily a Claude-transition artifact rather than a steady-state Grok burn.

Existing regression tests in tests/fm-watch-triage.test.sh that encoded the old blind-240s wedge (escalate a still-provably-working stale purely on elapsed time) were updated to the corrected re-verify semantics and given an added phase proving the genuine wedge still fires once the crew is no longer provably working; this is a strengthening of the contract, not a weakening, and the captain authorized removing the false wedge. New coverage: the signal_all_declared_wait classifier, declared-paused and captain-held signal absorbs, and the real-arm no-synthetic-turn integration test. docs/architecture.md updated to keep the one-owner absorb-contract description accurate. bin/fm-lint.sh is clean.

Constraints: firstmate shared tracked material only (firstmate-coding-guidelines followed); no captain-facing copy changed beyond stopping the no-op turn itself; the no-blind-turn supervision contract, the single supervision-cycle model, and resurface-default-off are all preserved.

What Changed

  • Signal path (bin/fm-watch.sh + new signal_all_declared_wait in bin/fm-classify-lib.sh): a paused:/captain-held status append and the bare turn-end under it are now absorbed as a declared external wait when the crew is not provably working, so the watcher keeps blocking and the tracked background arm never exits into a synthetic primary turn; the swallowed-finish guard still surfaces a crew that neither declared a wait nor is provably working.
  • Stale path (wedge_timer_check gains an optional reverify-task): a provably-working stale is re-verified against the current run state at each FM_STALE_ESCALATE_SECS boundary (bounded to once per threshold per window) instead of blind-escalating on elapsed time — an actively-running validation pane resets its timer and keeps absorbing, while a run that reached a gate, finished, failed, or went unreadable still escalates; the busy-pane turn-age wedge is intentionally left un-reverified.
  • Added a real-arm + real-watcher integration test (tests/fm-watch-arm.test.sh) asserting a declared-wait absorb leaves both processes alive while a captain-relevant status exits both, extended tests/fm-watch-triage.test.sh for the new classifier and re-verify semantics, and updated docs/architecture.md/docs/configuration.md to match the absorb contract.

Risk Assessment

✅ Low: A tightly scoped, well-tested watcher fix that adds two absorb guards behind the existing captain-relevant/provably-working checks, preserving the swallowed-finish and "every real wake lands" invariants and backed by real-arm and re-verify regression tests.

Testing

Ran the two targeted watcher suites the change adds/updates. tests/fm-watch-arm.test.sh (8/8) drives a real fm-watch.sh watcher plus a real attached fm-watch-arm.sh arm end to end and proves a declared-wait absorb keeps BOTH processes alive — so no synthetic Grok primary turn — while a genuine done: status still exits both and reports the signal. tests/fm-watch-triage.test.sh (62/62) covers the new signal_all_declared_wait classifier, declared-paused and captain-held signal absorbs when the crew is not provably working, the preserved swallowed-finish guard, and the stale-path re-verify that keeps absorbing a provably-working long validation wait but still fires the wedge escalation once the run is no longer provably working. All pass, zero failures; a CLI transcript evidence file was captured. No rendered UI surface exists for this supervision-daemon change, so evidence is the real-process CLI transcript rather than a screenshot.

Evidence: Real-arm+watcher absorb-burn test transcript

### fm-watch-arm.test.sh (real arm + real watcher) ### ok - watch-arm: a declared-wait absorb keeps the real attached arm alive (no synthetic turn); a genuine wake still exits it ### fm-watch-triage.test.sh - intent-critical assertions ### ok - signal_all_declared_wait: only all-paused/captain-held batches (status or turn-end) are declared waits ok - a bare turn-end whose crew is not provably working is surfaced (the swallowed-finish fix) ok - a declared paused: signal (status append plus turn-end) is absorbed even when the crew is not provably working ok - a verified captain-held signal is absorbed even when the crew is not provably working ok - a stale terminal-looking status overridden by an active run is re-verified past the threshold instead of burning a turn, and escalates only once no longer provably working ok - provably-working non-terminal stale is absorbed, re-verified working past the threshold instead of burning a turn, and escalates only once no longer provably working

### fm-watch-arm.test.sh (real arm + real watcher) ###
ok - watch-arm: an attached arm reports the wake its cycle delivered instead of a false failure
ok - watch-arm: a delivered wake consumed by the handling turn still closes the attached arm cleanly
ok - watch-arm: a cycle that delivered no wake of its own still fails loudly
ok - watch-arm: a lost ledger append is resolved from the cycle's own actionable close receipt
ok - watch-arm: a close receipt from another identity cannot resolve this cycle
ok - watch-arm: decision-only recovery resurfaces without a new queue row
ok - watch-arm: process-event delivery gets first refusal before recovery resurfacing
ok - watch-arm: a declared-wait absorb keeps the real attached arm alive (no synthetic turn); a genuine wake still exits it

### fm-watch-triage.test.sh - intent-critical assertions ###
ok - signal_all_declared_wait: only all-paused/captain-held batches (status or turn-end) are declared waits
ok - a bare turn-end whose crew is not provably working is surfaced (the swallowed-finish fix)
ok - a declared paused: signal (status append plus turn-end) is absorbed even when the crew is not provably working
ok - a verified captain-held signal is absorbed even when the crew is not provably working
ok - a stale terminal-looking status overridden by an active run is re-verified past the threshold instead of burning a turn, and escalates only once no longer provably working
ok - provably-working non-terminal stale is absorbed, re-verified working past the threshold instead of burning a turn, and escalates only once no longer provably working

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-watch-arm.test.sh — 8/8 pass, incl. real-arm+real-watcher test_attached_arm_stays_live_through_a_declared_wait_absorb (declared-wait absorb keeps both processes alive; genuine wake exits both)
  • bash tests/fm-watch-triage.test.sh — 62/62 pass, incl. signal_all_declared_wait classifier, declared-paused/captain-held signal absorbs while crew not provably working, swallowed-finish guard surfaces a not-working turn-end, and stale re-verify (...re-verified past the threshold instead of burning a turn, and escalates only once no longer provably working)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

…ng validations

A declared external wait (paused:/captain-held) and a still-running
no-mistakes validation no longer force a supervision turn - the burn a Grok
primary showed on 2026-08-15, where a healthy worker sitting on a long
validation step or a declared captain-wait produced a stream of no-op
"Captain, shipshape." replies.

Two native watcher changes:
- Signal path: a paused:/captain-held status append and the bare turn-end
  under it are absorbed as a declared external wait when the crew is not
  provably working, matching the stale path (which already absorbs declared
  waits) and the codified captain preference against alerts on declared
  captain-wait windows. The swallowed-finish guard is preserved: a crew that
  neither declared a wait nor is provably working still surfaces.
- Stale path: a provably-working stale is re-verified against the
  authoritative current run state at each FM_STALE_ESCALATE_SECS boundary
  instead of blind-escalating on elapsed time. A run fm-crew-state.sh still
  reports as actively working legitimately sits on a static pane for a long
  CI/validation wait and keeps absorbing silently; only a run that reached a
  gate, finished, failed, or went unreadable escalates, so genuine
  stuck-worker detection survives.

Because the watcher keeps blocking on these absorbs, the Grok monitor arm
(bin/fm-watch-arm.sh, whose exit becomes a synthetic primary turn) never
exits, proven end to end by a real arm + real watcher regression.
@withally
withally merged commit a34d151 into main Aug 15, 2026
13 checks passed
withally added a commit that referenced this pull request Aug 15, 2026
… survive compaction (#38)

* fix(watch): hold the sitting wait for a live quiet crew and restore read-once after a compaction

Two halves of one goal: the primary's sitting wait should end only when the
captain would care, and a compaction should not force the primary to re-read the
startup dump to remember it.

Half A - the sitting wait ends only for a captain-relevant event.
Every wake() exits the watcher, and on the Grok primary that exit becomes a
synthetic turn, so each non-captain-relevant wake is a no-op turn. PR #37
absorbed declared waits, running validation, and empty rearms; the residual leak
was the fail-safe immediate surface of a crew that merely went quiet - present
identically on the signal path (a bare turn-end) and the stale path (the none
class). It is one principle across both paths, not two more absorb types: a crew
that is not provably working, declared no wait, has no captain-relevant status
line, and whose agent is still live has only gone quiet between steps. The
attended primary now absorbs it and hands it to the existing wedge timer, which
surfaces it as a possible wedge only after STALE_ESCALATE_SECS, while a crew
whose agent has confidently exited (fm_backend_agent_state dead/missing) still
surfaces at once as a swallowed finish. Away mode is unchanged - its daemon keeps
classifying every ambiguous signal. Captain-relevant lines, genuine wedges,
genuine finishes, secondmate reports, and unknown/unmatched-decision status lines
all still surface; the heartbeat backstop remains the safety net.

Half B - read-once survives a compaction. A run-tier harness re-emits the digest
on a clear/compact source, but Grok fires no session-open hook on an in-session
compaction and discards SessionStart stdout, so nothing restores the dump and the
primary re-reads captain.md/learnings.md file by file. The read-once contract in
AGENTS.md section 3 (re-sent every turn, so it survives the compaction) now
directs one fm-session-start.sh --reemit to restore the fleet state and durable
dump, and the digest's own read-once block and docs/sessionstart-nudge.md carry
the same restore path and the Grok limitation.

Upstream (kunchenguid) had nothing importable that reaches this: the Grok
turn-end guard is byte-identical, its wait/arm lineage diverged and is behind the
fork's clean-close receipt, and its compaction refresh (kunchenguid#2163) is Pi-only.

Regression tests drive a real fm-watch.sh and a real fm-session-start.sh
--reemit, not stubs.

* no-mistakes(review): gate quiet-crew absorb on a recorded window; fix stale secondmate comment

* no-mistakes(document): align watcher-surface docs with quiet-crew absorb

* no-mistakes: apply CI fixes
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant