fix(watch): treat armed PR merge polls as external waits, not wedges - #24
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Stop the firstmate watcher (bin/fm-watch.sh) from escalating an idle ship worker as a possible wedge while its PR is open and its merge poll is armed.
Defect (observed four times in one night, 2026-09-22): a ship worker that reported
done: PR <url> checks greenand stopped has nothing left to do, yet its pane escalated as a possible wedge about every 30 minutes for as long as the captain took to merge; by the third escalation the watcher demanded a deep inspection that could only ever conclude "still waiting". Two remedies were tried and neither works: appendingpaused:does not help because pause_state_class decides run-step precedence BEFORE liveness (an actively-running pipeline outranks a declared wait), and the no-mistakes run stays active while it monitors for merge; exiting the agent does not help for the same reason. Raising config/stale-escalate-secs would weaken wedge detection for every real worker, so it is not acceptable.The signal that already exists: bin/fm-pr-check.sh records pr= in state/.meta and arms a validated merge poll (state/.pr-poll, state/.pr-poll-registration), so the watcher already knows this task waits on an external event it is itself polling for.
Required behaviour:
Tests: extend the existing watcher tests (tests/fm-watch-triage.test.sh, which covers pause_state_class / handle_paused_stale). Cover: armed poll + idle pane -> no wedge escalation within the window and a bounded recheck after it; same pane with no poll -> escalates as today; armed poll with dead endpoint -> still reported. Each assertion was proven able to fail by breaking the new branch (nine mutations, each turned its specific test red).
Constraints: do not touch the primary checkout; run bin/fm-lint.sh; update docs/turnend-guard.md or whichever doc owns stale-escalation semantics only where a statement becomes untrue (docs/architecture.md and docs/configuration.md were updated).
Implementation decisions made deliberately:
alive(unknown/unverified backends such as zellij/orca/cmux fall back to today's escalation rather than being silenced).pausedfor a holding PR wait with a cheap repeat path (poll registration re-validated every call; crew state and liveness re-read every STALE_ESCALATE_SECS); the declared-pause path lets a holding PR wait outrank the run-step precedence. No first surface is owed for a PR wait because the PR-ready report already reached firstmate.pausedclass to handle_paused_stale there, so unarmed terminal stale surfacing is byte-for-byte the same as before.What Changed
pr_merge_wait_holds/pause_state_class) and usehandle_paused_stale's boundedPAUSE_RESURFACE_SECScadence instead of wedge escalation, with a distinct PR-open recheck reason.pr=line) plus working/done crew state, non-blocked status, and a confidently alive agent; disarmed/invalid polls, dead endpoints, and tasks without an armed poll keep today's surface or escalation behavior.Risk Assessment
Testing
Drove the four PR-merge-wait watcher scenarios end-to-end against real bin/fm-watch.sh (armed via real bin/fm-pr-check.sh). All four suite tests passed in 91s, and a product transcript showed armed idle waits absorbed with a paused marker and zero wakes, then a single external-wait recheck; bare pr=, dead agent, and disarmed poll still wedge-escalated.
Evidence: PR merge wait suite results (all four ok)
ok - an idle crew whose PR has an armed merge poll is rechecked on the bounded cadence, never wedge-escalated ok - an idle provably-working crew without an armed merge poll still wedge-escalates as before ok - an armed PR merge wait never hides a dead agent ok - a PR merge wait re-surfaces as soon as its poll is disarmed or its run stops only waiting on the mergeEvidence: Watcher product transcript (absorb, recheck reason, wedge paths)
Evidence: Suite timing metadata
EXIT:0 SECS:91Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
✅ No issues found.
✅ **Test** - passed
✅ No issues found.
tests/fm-watch-triage.test.shfocused runner:test_pr_merge_wait_is_bounded_not_a_wedgetest_pr_merge_wait_without_armed_poll_still_escalatestest_pr_merge_wait_dead_endpoint_still_reportedtest_pr_merge_wait_resurfaces_when_it_stops_holdingManual transcript drive of realbin/fm-watch.sh+bin/fm-pr-check.shfor armed absorb/recheck, bare pr= wedge, dead-agent report, and disarmed-poll return-to-wedge✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.