Skip to content

fix(watch): let one TERM always stop the watcher on bash 5.2 - #31

Merged
Aviator-Coding merged 1 commit into
mainfrom
fm/fm-watch-triage-relaunch-window-flake
Sep 25, 2026
Merged

Aviator-Coding merged 1 commit into
mainfrom
fm/fm-watch-triage-relaunch-window-flake

Conversation

@Aviator-Coding

Copy link
Copy Markdown
Owner

Intent

Apply that same high standard to engineering excellence: lint, test failures, and test flakiness. If you see one, even if it is not caused by what you are working on right now, still get it fixed.

The failure this applies to: CI on main is red. Run https://github.com/Aviator-Coding/firstmate/actions/runs/36095026336, on commit 607a212, failed its "Behavior portable serial 1" job with not ok - the relaunch round escalated before its fresh window elapsed: stale: test:fm-wedge (idle 999s, possible wedge, escalation 1) in tests/fm-watch-triage.test.sh. The watcher-wake-lock family took 1257s in that run. Main should be green, and that test should pass reliably.

What Changed

  • Split signal handling in bin/fm-watch.sh into watcher_stop_signals, which leaves HUP/TERM on bash's native fatal-signal handling (so they run watcher_cleanup via the EXIT trap even mid-poll) while keeping an explicit exit 1 trap for INT, since a trap body for HUP/TERM is not reliably parsed/run on bash 5.2.
  • Narrowed the window in run_check_capture where stop signals are deferred via a trap to only the span before the check's process group is recorded, replacing the previous unconditional trap 'exit 1' HUP INT TERM.
  • Updated tests/fm-watch-triage.test.sh: reap() now waits with a bounded timeout (wait_for_exit ... 100) and fails loudly if the watcher doesn't exit within 10s of TERM instead of waiting unboundedly, and added test_term_stops_a_watcher_blocked_inside_a_poll, which blocks a watcher mid pane-capture on a FIFO and asserts TERM still stops it, releases its lock, and leaves an acknowledgeable stop record.
  • Documented the new signal-handling behavior and test coverage in docs/watcher-continuity.md.

Risk Assessment

✅ Low: The change is a narrowly-scoped, well-reasoned bash signal-handling fix (verified empirically: a custom bash trap body for TERM/HUP is deferred while blocked in a foreground read, so a blocked watcher never honors a stop request, while default-disposition fatal signals kill it immediately yet still run the EXIT trap), is consistently applied everywhere HUP/TERM are trapped in bin/fm-watch.sh, is covered by a new regression test that reproduces the exact blocked-poll scenario and asserts real observable cleanup (lock release, ack), and tightens a shared test helper (reap) to a bounded wait consistent with the file's existing 100-tick convention instead of an unbounded one.

Testing

Directly reproduced the reported regression (fails pre-fix, passes post-fix) on the isolated new test, then confirmed the fix holds under 10 idle repeats and 8 repeats under heavy CPU load with zero flakes; the full fm-watch-triage suite was 87/136 passing with zero failures (including the exact previously-failing assertion) when the run had to be reported, and the sibling fm-watcher-lock suite passed 32/32 end to end - no regressions or flakiness observed anywhere.

  • Live validation: ✅ go - 5 of 6 scenarios driven live against the product
Scenario Result Live Evidence
TERM stops a watcher blocked mid-poll and releases its lock (the reported flake's root cause) ✅ pass live Isolated run of the new test_term_stops_a_watcher_blocked_inside_a_poll against target bin/fm-watch.sh: ok - TERM stops a watcher blocked inside a poll and still runs its cleanup in ~2s; lock file a…
Adversarial: same scenario reproduces the failure on pre-fix code (regression proof) ✅ pass live Same isolated test against base commit's bin/fm-watch.sh: not ok - TERM did not stop a watcher blocked inside a poll after wait_for_exit's 10s budget plus KILL fallback (~23s wall), matching the rep…
Fix is not flaky under repetition or CPU contention ✅ pass live 10/10 consecutive passes idle, 8/8 passes while 8 yes processes loaded a 10-core host, all post-fix.
The originally-reported CI assertion now passes in a full suite run ✅ pass live Full tests/fm-watch-triage.test.sh run on target commit printed ok - a second death after a same-window relaunch reports in full without a live probe, and an unchanged dead pane stays silent (the…
No regression in the sibling watcher-lock signal/lifecycle suite sharing the same bin/fm-watch.sh code path ✅ pass live tests/fm-watcher-lock.test.sh on target commit: 32/32 pass, including ok - arm cleans child watcher and temp output on HUP and the SIGSTOP liveness-vs-stale-beacon case.
Full 136-test fm-watch-triage.test.sh suite completes green end-to-end in this environment ⏸️ untested no The background run reached 87/136 passing with zero failures (covering both the historically-flaky assertion and the new regression test) but had not reached test 136 by the time this report had to be…
Evidence: Regression reproduction and full-suite pass transcript

Source: Regression reproduction and full-suite pass transcript

fm/fm-watch-triage-relaunch-window-flake - live validation transcript
base:   5801243c2e3a8c8011a157b0ff14aadedc46f1fb
target: 424baf0e5b8006c861dbebbacd50c42ea1cb9bb6
bash:   GNU bash 5.3.9 (aarch64-apple-darwin25.4.0)

=== 1. Regression reproduction: isolated new test, base bin/fm-watch.sh (pre-fix) ===
$ bash tests/fm-watch-triage.test.sh   # only test_term_stops_a_watcher_blocked_inside_a_poll enabled
wait_for_exit: owned pid 56944 exceeded 100 polls; sending TERM
56944 56640 S    bash /tmp/fm-regress-check/repo/bin/fm-watch.sh
wait_for_exit: owned pid 56944 survived TERM; sending KILL
not ok - TERM did not stop a watcher blocked inside a poll
EXIT:1   (~23s wall)

=== 2. Same isolated test, target bin/fm-watch.sh (post-fix) ===
$ bash tests/fm-watch-triage.test.sh
ok - TERM stops a watcher blocked inside a poll and still runs its cleanup
EXIT:0   (~2s wall)

=== 3. Flake-resistance: 10 consecutive repeats, post-fix, idle machine ===
run 1..10: exit=0 ok - TERM stops a watcher blocked inside a poll and still runs its cleanup   (10/10)

=== 4. Flake-resistance under CPU load: 8x `yes` background load on a 10-core host, post-fix ===
load-run 1..8: exit=0 ok - TERM stops a watcher blocked inside a poll and still runs its cleanup   (8/8)

=== 5. Full tests/fm-watch-triage.test.sh suite, target commit (checked-out worktree) ===
$ bash tests/fm-watch-triage.test.sh
... 87/136 observed "ok", 0 "not ok", including:
ok - a second death after a same-window relaunch reports in full without a live probe, and an unchanged dead pane stays silent
     (this is the exact assertion that failed in the reported CI run:
      "the relaunch round escalated before its fresh window elapsed")
ok - TERM stops a watcher blocked inside a poll and still runs its cleanup
Run was still progressing (no failures) when this evidence was captured; see testing_summary.

=== 6. Sibling suite tests/fm-watcher-lock.test.sh, target commit (shares bin/fm-watch.sh signal handling) ===
$ bash tests/fm-watcher-lock.test.sh
32/32 ok, 0 not ok, including:
ok - arm cleans child watcher and temp output on HUP
ok - SIGSTOP distinguishes live PID from stale beacon and termination records the exit class
EXIT:0

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 5 of 6 scenarios driven live against the product
Scenario Result Live Evidence
TERM stops a watcher blocked mid-poll and releases its lock (the reported flake's root cause) ✅ pass live Isolated run of the new test_term_stops_a_watcher_blocked_inside_a_poll against target bin/fm-watch.sh: ok - TERM stops a watcher blocked inside a poll and still runs its cleanup in ~2s; lock file a…
Adversarial: same scenario reproduces the failure on pre-fix code (regression proof) ✅ pass live Same isolated test against base commit's bin/fm-watch.sh: not ok - TERM did not stop a watcher blocked inside a poll after wait_for_exit's 10s budget plus KILL fallback (~23s wall), matching the rep…
Fix is not flaky under repetition or CPU contention ✅ pass live 10/10 consecutive passes idle, 8/8 passes while 8 yes processes loaded a 10-core host, all post-fix.
The originally-reported CI assertion now passes in a full suite run ✅ pass live Full tests/fm-watch-triage.test.sh run on target commit printed ok - a second death after a same-window relaunch reports in full without a live probe, and an unchanged dead pane stays silent (the…
No regression in the sibling watcher-lock signal/lifecycle suite sharing the same bin/fm-watch.sh code path ✅ pass live tests/fm-watcher-lock.test.sh on target commit: 32/32 pass, including ok - arm cleans child watcher and temp output on HUP and the SIGSTOP liveness-vs-stale-beacon case.
Full 136-test fm-watch-triage.test.sh suite completes green end-to-end in this environment ⏸️ untested no The background run reached 87/136 passing with zero failures (covering both the historically-flaky assertion and the new regression test) but had not reached test 136 by the time this report had to be…
  • Isolated the new regression test and ran it against the pre-fix bin/fm-watch.sh (base commit 5801243): failed with not ok - TERM did not stop a watcher blocked inside a poll after a 10s TERM-then-KILL fallback (~23s wall), reproducing the class of hang/flake reported in CI.
  • Ran the same isolated test against the post-fix bin/fm-watch.sh (target commit 424baf0): passed cleanly in ~2s.
  • Repeated the post-fix isolated test 10 consecutive times on an idle machine: 10/10 pass, no flake.
  • Repeated the post-fix isolated test 8 times while 8 yes processes saturated CPU on a 10-core host (adversarial load condition matching the reported CI flake context): 8/8 pass, no flake.
  • bash tests/fm-watch-triage.test.sh (full 136-test suite) on the target commit: observed 87/136 tests pass with zero failures before report time, including the exact assertion that failed in the reported CI run (a second death after a same-window relaunch reports in full...) and the new regression test.
  • bash tests/fm-watcher-lock.test.sh (sibling suite exercising the same shared bin/fm-watch.sh signal-handling contract, including a HUP-cleanup test and a SIGSTOP liveness test) on the target commit: 32/32 pass.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Port of upstream kunchenguid/firstmate 52fca51 (kunchenguid#5362).

Bash 5.2 runs a pending trap from the parser entry of the next command
substitution it expands, where the trap body is parsed as the inside of
that substitution and fails ("trap: line 2: unexpected EOF while looking
for matching `)'") or is dropped, consuming the signal. The watcher's
`trap 'exit 1' HUP INT TERM` could therefore ignore a TERM and keep
polling while its stopper waited. In the triage suite that left reap
blocked until the relaunch round's watcher escalated on its own after
999s, which failed "the relaunch round escalated before its fresh window
elapsed" on main CI (Ubuntu 24.04, bash 5.2.21).

HUP and TERM now keep bash's native fatal-signal handling, which runs the
EXIT trap (watcher_cleanup) and exits. INT keeps its trap because bash
ignores a direct SIGINT while a child runs. The check-spawn deferral
window no longer contains a command substitution.

The triage suite's reap is now bounded and fails the case within 10s with
process evidence instead of hanging, and a new regression test proves
TERM stops a watcher blocked inside a poll's pane capture and still
releases its lock and records an acknowledgeable stop.
@Aviator-Coding
Aviator-Coding merged commit afe5791 into main Sep 25, 2026
19 checks passed
@Aviator-Coding
Aviator-Coding deleted the fm/fm-watch-triage-relaunch-window-flake branch September 25, 2026 23:38
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant