Skip to content

fix(bin): bound watcher shutdown and improve stop diagnostics - #4073

Closed
mremond wants to merge 6 commits into
kunchenguid:mainfrom
mremond:fm/fm-installed-timeout-watcher-will-not-stop
Closed

mremond wants to merge 6 commits into
kunchenguid:mainfrom
mremond:fm/fm-installed-timeout-watcher-will-not-stop

Conversation

@mremond

@mremond mremond commented Sep 9, 2026 •

Copy link
Copy Markdown
Contributor

Intent

FOUR OF OUR OPEN REQUESTS ARE STUCK BEHIND ONE RED, AND THE CAPTAIN HAS MADE UNBLOCKING THEM THE FLEET'S ABSOLUTE PRIORITY (routed 2026-09-09).

THE CHAIN, so you know what rests on this: pull request 4056 must go green before it can merge; 4006 rebases onto it and repairs the lane packing; the lanes that keep getting cut at twenty minutes then pass again, which unblocks 3844, 4019 and 4011. All of it waits on one assertion.

THE FAILURE, read from the job log of CI run 34319847102 rather than the summary page:

not ok - installed-timeout watcher did not stop after the direct check returned

in tests/fm-pr-check-security.test.sh, family pr-forge, script duration 125256ms, exit 1. The shard total was 945044ms - well under the twenty-minute ceiling - so this is a GENUINE FAILURE and not a lane that was cut. That corrects an earlier reading in which it was mistaken for the packing defect.

WHAT IS ALREADY ESTABLISHED. Do not re-derive it; do check it if you have reason to doubt it, and say so if you find it wrong.

  • Pull request 4056's own change cannot influence that assertion. Measured four ways. The two decisive grounds: the failing script runs at POSITION 1 in its batch, so no neighbour could have created or withheld anything for it, and it builds its own home and passes it explicitly.
  • The hypothesis that this is the process-stop defect repaired on pull request 4009 is KILLED ON MECHANISM: that repair touches only procevent code, and the failing watcher-drain path never reaches it.
  • We cannot re-run the job. Access to that repository is pull-only, verified at the permission level - admin, maintain, push and triage all report false - not inferred from an error message.

WHAT NOBODY HAS ESTABLISHED, and it is the entire deliverable: WHY the installed-timeout watcher did not stop after the direct check returned.

HOW WEAK OUR BOUND IS, stated so you do not inherit false confidence: no reproduction in three runs, and no prior sighting in five. Five is a small window. "No prior occurrence found" is not "first occurrence".

WHAT THIS DELIVERY ACTUALLY CONTAINS, resolved into substance so a reviewer reading only the diff has the same context: three items and nothing else.

  1. Bound the watcher's shutdown so it cannot wait forever.
  2. Make the stop assertion print what it saw instead of failing silently - and the failure diagnostic reports the DESCENDANT process tree rather than direct children, because this fixture deliberately reparents the very process the diagnostic exists to reveal.
  3. Fix the liveness helper that read an empty answer as ALIVE rather than UNKNOWN.
    Five mutations were proved before the first commit, establishing that the stop assertion can in fact go red for the reason it names.

DECISIONS ALREADY TAKEN ON THIS WORK, so they are not re-flagged as mistakes:

  • A public timeout override and an ambient one were BOTH REMOVED on purpose. The scope is the three items above and nothing else; configurable timing is deliberately not among them. Do not propose reintroducing a configurable or overridable timeout.
  • A fix round had added a blanket absent-lock guard broader than the problem needed, and it PERMANENTLY WEDGES: a stealer dying between removing the primary lock and replacing it leaves its steal lock behind, and that guard returns before reclaiming it, forever. For the watcher's own lock that means startup announcing "already running" while nothing owns it. The ruling was to DELETE that one line, keep the distinct resource code which already satisfies the earlier finding, and leave the steal path untouched. That is deliberate, not an omission. Worth stating plainly: a delivery whose subject is a watcher that cannot STOP nearly shipped a watcher that cannot START - we caught it, and a maintainer should see that we caught it.

WHAT THIS PARTICULAR RUN IS FOR, 2026-09-15: the branch is 7 ahead and 16 behind the trunk, and the captain ordered it REBASED onto current main and republished. The rebase is routed through the validation chain rather than done by hand, because the chain's own rebase step moves the base and its push step rebinds the attestation to the new head; a hand push would move the head without rebinding and leave the request asserting an attestation for a commit that no longer exists. The branch content must survive the base move unchanged - that is being proved by comparing the whole-branch diff patch-id against its own merge base before and after. No behaviour change, no cleanup, no new scope: if the rebase hits a conflict it is resolved as a conflict only. The request is NOT to be merged - it is classed new-default by the maintainer and goes to their own captain, so merge authority is theirs and green only clears the half they are waiting on.

PRIOR CONTEXT THAT A REVIEWER SHOULD NOT MISREAD: the forge Lint lane has been killed twice on the pre-rebase head (jobs 103900216344 and 103944331590, both exit 143, at 7m22s and 9m12s against a ~15-minute healthy runtime). Exit 143 is SIGTERM; bin/fm-lint.sh carries a TERM trap that converts an external kill into that exit status, which is why an external termination reaches the forge as a FAILURE on a check named Lint rather than as a cancellation. That is a termination, not a lint defect in this branch, and no lint configuration is to be touched to get past it.

What Changed

  • Bound watcher cleanup’s recovery-marker lock wait to five seconds, report persistence failures, and retain stale singleton-lock evidence. Return on lock-owner creation failures while preserving abandoned-stealer recovery.
  • Add elapsed time, stderr, and descendant process trees—including the recorded reparented child—to timeout watcher stop-failure diagnostics.
  • Share live/gone/unknown process checks across test helpers, retry empty ps responses, and update waits and assertions with regression coverage for liveness, held locks, and unwritable state.

Risk Assessment

✅ Low: The change remains within the approved scope, preserves branch behavior through the rebase, and has no substantiated material defects.

Testing

Targeted regressions, live shutdown and recovery faults, descendant diagnostics, and resource-mutation checks passed. Captured process transcripts and rebase evidence; corrected test-driver setup and removed temporary files. Real installed timeout was unavailable. No broad suite, lint, or other pipeline phase ran.

  • Live validation: ✅ go - 7 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Send TERM after a custom check returns: the watcher exits, drains its reparented descendant, and removes private files and its singleton lock. ✅ pass live Watcher validation and mutation evidence: fallback-returned-check
Stop a watcher after a returned custom check using a real installed timeout executable. ⏸️ untested no Neither timeout nor gtimeout is available. The existing regression passed with its timeout stand-in. Provide a real executable on PATH or rerun on a host with coreutils installed.
Hold the marker lock during shutdown or recovery-wake cleanup: both transitions stop, identify the holder, and preserve recovery evidence. ✅ pass live Watcher validation and mutation evidence: held-marker-new, held-marker-existing, release-lock-existing
Make state unwritable as a non-root user: acquisition returns nonzero and the signalled watcher reaches its deadline diagnostic. ✅ pass live Watcher validation and mutation evidence: unwritable-state and resource-fault RED to GREEN
Kill a stealer between primary-lock removal and replacement: the next acquirer recovers and a watcher starts; same-process marker reacquisition also succeeds. ✅ pass live Watcher validation and mutation evidence: abandoned-stealer and self-held
Supply ambient timeout overrides: ordinary publication still waits for its holder, while shutdown retains its internal five-second deadline. ✅ pass live Watcher validation and mutation evidence: ordinary-publish and release-lock-existing
Freeze the real watcher after direct-check completion: the failing stop assertion prints the watcher, reparented child, and grandchild. ✅ pass live Live descendant-tree diagnostic
Remove ps from the liveness command's PATH: the assertion reports UNKNOWN and unreadable liveness instead of accepting successful termination. ✅ pass live Unreadable liveness diagnostic
Evidence: Watcher validation and mutation evidence

Source: Watcher validation and mutation evidence

# Watcher validation on 31e4d35bec947717151b92359a2f04a542c05dd2

The real watcher and lock library ran in temporary homes inside the assigned worktree. No production file was changed.

## Fallback timeout: TERM after direct check returns

`` `json
{
  "scenario": "Fallback timeout: TERM after direct check returns",
  "result": "pass",
  "registration": "registered: state/custom.check.sh",
  "watcher_before": "65306 65247 65247 S    bash ~/.no-mistakes/worktrees/acf4a767348a/01M2JEJPE43XFXJCGYM3KFAX75/bin/fm-watch.sh",
  "watcher_exit": 1,
  "term_to_exit_seconds": 0.088,
  "recorded_descendant_pid": 65960,
  "descendant_after": "absent",
  "sentinel_exists": false,
  "private_files_left": [],
  "singleton_lock_exists": false
}
`` `

## held-marker-new

`` `json
{
  "scenario": "held-marker-new",
  "result": "pass",
  "watcher_pid": 67012,
  "watcher_exit": 1,
  "term_to_exit_seconds": 5.06,
  "diagnostic": "watcher: recovery state could not be persisted within 5s (marker lock held by pid 67569); stopping and retaining stale lock evidence",
  "holder_process": "67569 65247 65247 S    bash -c . \"$1\"; fm_lock_try_acquire \"$2\" || exit 11; trap 'fm_lock_release \"$2\"; exit' TERM; echo ready > \"$3\"; while :; do sleep .1; done _ ~/.no-mistakes/worktrees/acf4a767348a/01M2JEJPE43XFXJCGYM3KFAX75/bin/fm-wake-lib.sh ~/.no-mistakes/worktrees/acf4a767348a/01M2JEJPE43XFXJCGYM3KFAX75/.nm-test-phase/live/held-marker-new/state/.watcher-down.lock ~/.no-mistakes/worktrees/acf4a767348a/01M2JEJPE43XFXJCGYM3KFAX75/.nm-test-phase/live/held-marker-new/held-marker-new-holder.ready",
  "retained_watcher_lock_pid": "67012",
  "marker_before": "",
  "marker_after": ""
}
`` `

## held-marker-existing

`` `json
{
  "scenario": "held-marker-existing",
  "result": "pass",
  "watcher_pid": 70740,
  "watcher_exit": 1,
  "term_to_exit_seconds": 4.45,
  "diagnostic": "watcher: recovery state could not be persisted within 5s (marker lock held by pid 71323); stopping and retaining stale lock evidence",
  "holder_process": "71323 65247 65247 S    bash -c . \"$1\"; fm_lock_try_acquire \"$2\" || exit 11; trap 'fm_lock_release \"$2\"; exit' TERM; echo ready > \"$3\"; while :; do sleep .1; done _ ~/.no-mistakes/worktrees/acf4a767348a/01M2JEJPE43XFXJCGYM3KFAX75/bin/fm-wake-lib.sh ~/.no-mistakes/worktrees/acf4a767348a/01M2JEJPE43XFXJCGYM3KFAX75/.nm-test-phase/live/held-marker-existing/state/.watcher-down.lock ~/.no-mistakes/worktrees/acf4a767348a/01M2JEJPE43XFXJCGYM3KFAX75/.nm-test-phase/live/held-marker-existing/held-marker-existing-holder.ready",
  "retained_watcher_lock_pid": "70740",
  "marker_before": "pending:downtime:71276.1789474056.g9lNGq",
  "marker_after": "pending:downtime:71276.1789474056.g9lNGq"
}
`` `

## Resource fault: genuine unwritable state

`` `json
{
  "scenario": "Resource fault: genuine unwritable state",
  "result": "pass",
  "uid": 501,
  "state_mode": "0555",
  "primitive_output": "acquisition_rc=1",
  "primitive_elapsed_seconds": 0.046,
  "watcher_exit": 1,
  "term_to_exit_seconds": 4.402,
  "diagnostic": "watcher: recovery state could not be persisted within 5s (marker lock held by pid unknown); stopping and retaining stale lock evidence",
  "retained_watcher_lock_pid": "74210"
}
`` `

## Ordinary publication ignores ambient timeout overrides

`` `json
{
  "scenario": "Ordinary publication ignores ambient timeout overrides",
  "result": "pass",
  "ambient_values": {
    "FM_RECOVERY_MARKER_LOCK_TIMEOUT": "1",
    "FM_WATCHER_SHUTDOWN_LOCK_SECS": "1"
  },
  "blocked_after_seconds": 6,
  "publication_exit": 0,
  "marker": "pending:downtime:83505.1789474072.IbEHBa",
  "elapsed_seconds": 6.456
}
`` `

## Stealer dies after removing primary lock; next acquirer and watcher start

`` `json
{
  "scenario": "Stealer dies after removing primary lock; next acquirer and watcher start",
  "result": "pass",
  "stopped_stealer": "83708 65247 65247 T    bash -c . \"$1\"\\012set -T\\012trap 'if [ \"$BASH_COMMAND\" = \"rc=1\" ] && [ \"${lockdir:-}\" = \"$STATE/.watch.lock\" ] && [ ! -e \"$STATE/.watch.lock\" ] && [ -e \"$STATE/.watch.lock.steal\" ]; then printf \"%s\\n\" \"${BASHPID:-$$}\" > \"$STATE/steal-gap\"; kill -STOP \"${BASHPID:-$$}\"; fi' DEBUG\\012fm_lock_try_acquire \"$STATE/.watch.lock\"\\012 _ ~/.no-mistakes/worktrees/acf4a767348a/01M2JEJPE43XFXJCGYM3KFAX75/bin/fm-wake-lib.sh",
  "primary_absent_at_crash": true,
  "abandoned_steal_pid": "83708",
  "next_acquirer_output": "acquired_pid=83989\nrecovered_pid=",
  "next_acquirer_exit": 0,
  "steal_lock_reclaimed": true,
  "successor_watcher_pid": 84411,
  "successor_exit_after_term": 0,
  "successor_term_to_exit_seconds": 0.114
}
`` `

## Same process reclaims interrupted marker lock

`` `json
{
  "scenario": "Same process reclaims interrupted marker lock",
  "result": "pass",
  "exit": 0,
  "output": "marker=pending:downtime:85253.1789474074.f5o8ns\nreacquired_and_released"
}
`` `

## Recovery wake cleanup also times out and preserves the announced episode

`` `json
{
  "scenario": "Recovery wake cleanup also times out and preserves the announced episode",
  "result": "pass",
  "wake_output": "check: rearm-resurface",
  "transition": "release-lock-existing",
  "resume_to_exit_seconds": 4.786,
  "watcher_exit": 0,
  "diagnostic": "watcher: recovery state could not be persisted within 5s (marker lock held by pid 2838); stopping and retaining stale lock evidence",
  "ambient_values": {
    "FM_WATCHER_SHUTDOWN_LOCK_SECS": "300",
    "FM_RECOVERY_MARKER_LOCK_TIMEOUT": "300"
  },
  "marker_before": "announced:downtime:1636.1789474451.qAMl6i",
  "marker_after": "announced:downtime:1636.1789474451.qAMl6i",
  "holder_pid": 2838,
  "holder_survived": true,
  "stale_singleton_pid": "1665"
}
`` `

## Forced-stop diagnostic

The watcher used the host PATH; no fake timeout or backend tool was on its execution path.

`` `json
{
  "result": "pass",
  "injection": "SIGSTOP to real watcher from direct custom check; no production files modified",
  "ps_availability": "host ps",
  "assertion_exit": 1,
  "grandchild_pid": "75420",
  "recorded_reparented_child_pid": "75418",
  "reparented_child_ppid": "1"
}
`` `

`` `text
# fallback-timeout watcher still live 14s after TERM (150 polls); state below
# descendant tree (watcher pid 74563; recorded child pid 75418): pid ppid pgid stat wchan args
74563 74292 74288 T    -      bash ~/.no-mistakes/worktrees/acf4a767348a/01M2JEJPE43XFXJCGYM3KFAX75/bin/fm-watch.sh
75418     1 75413 S    -      perl -e $SIG{TERM}="IGNORE"; my $g=fork; die $! unless defined $g; if (!$g) {sleep 60; exit} open my $gf, ">", $ENV{NM_GRANDCHILD} or die $!; print {$gf} "$g\n"; close $gf; open my $ready, ">", $ENV{FM_TEST_DESCENDANT_READY} or die $!; print {$ready} "ready\n"; close $ready; sleep 60
  75420 75418 75413 S    -      perl -e $SIG{TERM}="IGNORE"; my $g=fork; die $! unless defined $g; if (!$g) {sleep 60; exit} open my $gf, ">", $ENV{NM_GRANDCHILD} or die $!; print {$gf} "$g\n"; close $gf; open my $ready, ">", $ENV{FM_TEST_DESCENDANT_READY} or die $!; print {$ready} "ready\n"; close $ready; sleep 60
# fallback-timeout watcher stderr tail:
not ok - fallback-timeout watcher did not stop after the direct check returned
`` `

## Unavailable ps fails closed

The watcher used the host PATH; no fake timeout or backend tool was on its execution path.

`` `json
{
  "result": "pass",
  "injection": "SIGSTOP to real watcher from direct custom check; no production files modified",
  "ps_availability": "removed from command-scoped PATH",
  "assertion_exit": 1,
  "grandchild_pid": "79423",
  "recorded_reparented_child_pid": "79421",
  "reparented_child_ppid": "1"
}
`` `

`` `text
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# fallback-timeout watcher still unreadable 14s after TERM (150 polls); state below
# descendant tree (watcher pid 78954; recorded child pid 79421): pid ppid pgid stat wchan args
78954 78854 74288 T    -      bash ~/.no-mistakes/worktrees/acf4a767348a/01M2JEJPE43XFXJCGYM3KFAX75/bin/fm-watch.sh
79421     1 79416 S    -      perl -e $SIG{TERM}="IGNORE"; my $g=fork; die $! unless defined $g; if (!$g) {sleep 60; exit} open my $gf, ">", $ENV{NM_GRANDCHILD} or die $!; print {$gf} "$g\n"; close $gf; open my $ready, ">", $ENV{FM_TEST_DESCENDANT_READY} or die $!; print {$ready} "ready\n"; close $ready; sleep 60
  79423 79421 79416 S    -      perl -e $SIG{TERM}="IGNORE"; my $g=fork; die $! unless defined $g; if (!$g) {sleep 60; exit} open my $gf, ">", $ENV{NM_GRANDCHILD} or die $!; print {$gf} "$g\n"; close $gf; open my $ready, ">", $ENV{FM_TEST_DESCENDANT_READY} or die $!; print {$ready} "ready\n"; close $ready; sleep 60
# fallback-timeout watcher stderr tail:
not ok - fallback-timeout watcher liveness was unreadable after the direct check returned
`` `

## Resource-fault RED to GREEN

`` `json
[
  {
    "mode": "mutant",
    "case": "test_lock_resource_failure_returns",
    "exit": 1,
    "seconds": 7.75,
    "output": "not ok - lock acquisition did not return on owner-directory creation failure"
  },
  {
    "mode": "mutant",
    "case": "test_shutdown_is_bounded_when_state_is_unwritable",
    "exit": 1,
    "seconds": 35.272,
    "output": "not ok - signaled watcher did not stop after owner-directory creation failed"
  },
  {
    "mode": "restored",
    "case": "test_lock_resource_failure_returns",
    "exit": 0,
    "seconds": 0.257,
    "output": "ok - lock acquisition returns nonzero when owner-directory creation fails"
  },
  {
    "mode": "restored",
    "case": "test_shutdown_is_bounded_when_state_is_unwritable",
    "exit": 0,
    "seconds": 16.381,
    "output": "ok - signaled watcher reaches its deadline on owner-directory creation failure (16s)"
  }
]
`` `

## Scope and limits

Targeted existing selectors: custom-check signal cleanup; returned descendants under both timeout fixture paths; held-marker shutdown; resource-fault acquisition and shutdown; live/gone/zombie/unreadable liveness.

The installed-timeout regression uses its existing timeout stand-in. Neither timeout nor gtimeout is installed on this host; the actual executable path is untested. Live custom-check validation used the real Perl fallback.

Resource mutations ran only against a disposable copy of bin/. Removing the resource distinction caused both expected hangs; restoring the target library passed both unchanged selectors. Existing polling budgets and ceilings were retained.

The alternate recovery-cleanup driver needed quoting and DEBUG-boundary corrections before its successful run. These were test setup issues; the final transcript records the successful live release-lock-existing case.

All nine runtime/test files are byte-identical to the pre-rebase merge. docs/watcher-continuity.md retains one upstream context change. Zero-context whole-branch patch IDs match; docs/configuration.md matches the new base. No new externally configurable timing remains.

The historical CI incident was not reproduced or assigned a new cause. This run establishes the current shutdown bounds, recovery behavior, and diagnostics.

No linter, formatter, static analyzer, full repository suite, remote CI, or pipeline-control command was run.
Evidence: Live descendant-tree diagnostic

Source: Live descendant-tree diagnostic

# fallback-timeout watcher still live 14s after TERM (150 polls); state below
# descendant tree (watcher pid 74563; recorded child pid 75418): pid ppid pgid stat wchan args
74563 74292 74288 T    -      bash ~/.no-mistakes/worktrees/acf4a767348a/01M2JEJPE43XFXJCGYM3KFAX75/bin/fm-watch.sh
75418     1 75413 S    -      perl -e $SIG{TERM}="IGNORE"; my $g=fork; die $! unless defined $g; if (!$g) {sleep 60; exit} open my $gf, ">", $ENV{NM_GRANDCHILD} or die $!; print {$gf} "$g\n"; close $gf; open my $ready, ">", $ENV{FM_TEST_DESCENDANT_READY} or die $!; print {$ready} "ready\n"; close $ready; sleep 60
  75420 75418 75413 S    -      perl -e $SIG{TERM}="IGNORE"; my $g=fork; die $! unless defined $g; if (!$g) {sleep 60; exit} open my $gf, ">", $ENV{NM_GRANDCHILD} or die $!; print {$gf} "$g\n"; close $gf; open my $ready, ">", $ENV{FM_TEST_DESCENDANT_READY} or die $!; print {$ready} "ready\n"; close $ready; sleep 60
# fallback-timeout watcher stderr tail:
not ok - fallback-timeout watcher did not stop after the direct check returned
Evidence: Unreadable liveness diagnostic

Source: Unreadable liveness diagnostic

# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# is_live_non_zombie: ps reported no state for present pid 78954; UNKNOWN, not live
# fallback-timeout watcher still unreadable 14s after TERM (150 polls); state below
# descendant tree (watcher pid 78954; recorded child pid 79421): pid ppid pgid stat wchan args
78954 78854 74288 T    -      bash ~/.no-mistakes/worktrees/acf4a767348a/01M2JEJPE43XFXJCGYM3KFAX75/bin/fm-watch.sh
79421     1 79416 S    -      perl -e $SIG{TERM}="IGNORE"; my $g=fork; die $! unless defined $g; if (!$g) {sleep 60; exit} open my $gf, ">", $ENV{NM_GRANDCHILD} or die $!; print {$gf} "$g\n"; close $gf; open my $ready, ">", $ENV{FM_TEST_DESCENDANT_READY} or die $!; print {$ready} "ready\n"; close $ready; sleep 60
  79423 79421 79416 S    -      perl -e $SIG{TERM}="IGNORE"; my $g=fork; die $! unless defined $g; if (!$g) {sleep 60; exit} open my $gf, ">", $ENV{NM_GRANDCHILD} or die $!; print {$gf} "$g\n"; close $gf; open my $ready, ">", $ENV{FM_TEST_DESCENDANT_READY} or die $!; print {$ready} "ready\n"; close $ready; sleep 60
# fallback-timeout watcher stderr tail:
not ok - fallback-timeout watcher liveness was unreadable after the direct check returned
Evidence: Rebase preservation

Source: Rebase preservation

{
  "target": "31e4d35bec947717151b92359a2f04a542c05dd2",
  "tree": "e650622bcce874254363c851c5e363be8ef96dd4",
  "base": "8b10b61e3feace8f275c6d0b3e490cdf7ab1f67d",
  "pre_rebase_head": "1129818ef720b6827ae75eb690957c2c7393d82b",
  "pre_rebase_base": "b182d0f908b78d08c7ccb8dce3775bdca8c5d657",
  "pre_rebase_patch_id": "4ad648a5cd1512046b5476ee9b69be0372f498fb",
  "post_rebase_patch_id": "3722f4b8850cf541d07c71c465e436bbbaa204bf",
  "changed_files_byte_identical": {
    "bin/fm-wake-lib.sh": true,
    "bin/fm-watch.sh": true,
    "docs/watcher-continuity.md": false,
    "tests/fm-daemon.test.sh": true,
    "tests/fm-pr-check-security.test.sh": true,
    "tests/fm-test-fixtures.test.sh": true,
    "tests/fm-watch-arm.test.sh": true,
    "tests/fm-watcher-lock.test.sh": true,
    "tests/lib.sh": true,
    "tests/wake-helpers.sh": true
  },
  "docs_configuration_matches_base": true,
  "pre_rebase_zero_context_patch_id": "1a6939634fb095c5fa6fbfeddce5da5aefffb1f9",
  "post_rebase_zero_context_patch_id": "1a6939634fb095c5fa6fbfeddce5da5aefffb1f9",
  "context_difference": "docs/watcher-continuity.md retains upstream host-timeout HUP/TERM/INT documentation; zero-context patch IDs compare branch additions/removals without that changed surrounding line."
}

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

🔧 **Rebase** - 3 issues found → auto-fixed ✅
  • ⚠️ bin/fm-watch.sh - merge conflict rebasing onto origin/main
  • ⚠️ docs/configuration.md - merge conflict rebasing onto origin/main
  • ⚠️ tests/fm-test-fixtures.test.sh - merge conflict rebasing onto origin/main

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 7 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Send TERM after a custom check returns: the watcher exits, drains its reparented descendant, and removes private files and its singleton lock. ✅ pass live Watcher validation and mutation evidence: fallback-returned-check
Stop a watcher after a returned custom check using a real installed timeout executable. ⏸️ untested no Neither timeout nor gtimeout is available. The existing regression passed with its timeout stand-in. Provide a real executable on PATH or rerun on a host with coreutils installed.
Hold the marker lock during shutdown or recovery-wake cleanup: both transitions stop, identify the holder, and preserve recovery evidence. ✅ pass live Watcher validation and mutation evidence: held-marker-new, held-marker-existing, release-lock-existing
Make state unwritable as a non-root user: acquisition returns nonzero and the signalled watcher reaches its deadline diagnostic. ✅ pass live Watcher validation and mutation evidence: unwritable-state and resource-fault RED to GREEN
Kill a stealer between primary-lock removal and replacement: the next acquirer recovers and a watcher starts; same-process marker reacquisition also succeeds. ✅ pass live Watcher validation and mutation evidence: abandoned-stealer and self-held
Supply ambient timeout overrides: ordinary publication still waits for its holder, while shutdown retains its internal five-second deadline. ✅ pass live Watcher validation and mutation evidence: ordinary-publish and release-lock-existing
Freeze the real watcher after direct-check completion: the failing stop assertion prints the watcher, reparented child, and grandchild. ✅ pass live Live descendant-tree diagnostic
Remove ps from the liveness command's PATH: the assertion reports UNKNOWN and unreadable liveness instead of accepting successful termination. ✅ pass live Unreadable liveness diagnostic
  • Compared pre/post-rebase zero-context patch IDs, changed-file contents, and docs/configuration.md against the new base.
  • bash tests/.nm-phase-fm-pr-check-security.sh: signal cleanup and returned descendants under both timeout fixture paths.
  • bash tests/.nm-phase-fm-watcher-lock.sh: held-marker shutdown, resource-fault acquisition, and unwritable-state shutdown.
  • bash tests/.nm-phase-fm-test-fixtures.sh: live, zombie, departed, and unreadable process states.
  • python3 ~/.no-mistakes/evidence/01M2JEJPE43XFXJCGYM3KFAX75/live-watcher-checks.py
  • python3 ~/.no-mistakes/evidence/01M2JEJPE43XFXJCGYM3KFAX75/diagnostic-check.py and the same command with unknown.
  • python3 ~/.no-mistakes/evidence/01M2JEJPE43XFXJCGYM3KFAX75/resource-mutation-check.py: both expected hangs went RED, then GREEN after restoration.
  • python3 ~/.no-mistakes/evidence/01M2JEJPE43XFXJCGYM3KFAX75/existing-recovery-check.py: recovery cleanup and ignored ambient overrides; corrected driver setup and re-drove successfully.
  • Removed temporary homes, selected harnesses, and disposable source copies; verified unchanged HEAD, clean git status, and no remaining test watchers.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@greptile-apps

greptile-apps Bot commented Sep 9, 2026

Copy link
Copy Markdown

RetriggerView in GreptileConfidence Score: 5/5

The PR appears safe to merge with no actionable correctness or security issues identified.

@mremond

mremond commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

Request: could you re-run Behavior portable serial 1?

That job is the only red on this PR, and no test failed in it. Reading the job log rather
than the checks summary (run 34351240268, job 102464763425):

FM_TEST_SUMMARY total=30 failed=0 skipped_gate=6 duration_ms=1186432

All 30 scripts passed. The job then wrote and uploaded its timing artifact and ran post-job
cleanup. Its conclusion is cancelled, not failure. The first log line is
2026-09-09T12:28:46.45Z and the last is 2026-09-09T12:48:46.85Z — 20 minutes and 0.4
seconds. The cut landed in the final second of cleanup, after everything that decides the
verdict had already passed. The other 12 jobs in the run all succeeded.

We can't re-run it ourselves — our access here is pull-only.

Before you decide: this PR is part of why that lane is tight

We'd rather volunteer this than have you find it. Our own tests pushed that lane closer to
the ceiling
, and by most of what was left:

merge base 40c50ea8 (run 34322484000) this PR fa8994aa (run 34351240268)
lane 1 result total=30 failed=0 — job succeeded total=30 failed=0 — job cancelled
lane 1 duration 1 146 927 ms 1 186 432 ms
headroom under the 1 200 000 ms ceiling 53 073 ms 13 568 ms

This change adds 39 505 ms — about 40 seconds — to that lane, consuming 39 of the 53
seconds of headroom it had and leaving about 14.

To be precise about what is not the explanation: we checked lane composition rather than
assuming it. --list --lane portable-serial-1of5 returns the identical 30 scripts at this
head and at the merge base, and the watcher-wake-lock family membership is identical too.
Nothing was repacked or reordered. The lane simply got about forty seconds heavier while
already sitting within a minute of the cap.

Two honest caveats, written before the re-run rather than after

A re-run may well not hold. At roughly 14 seconds of headroom on a 20-minute cap this is
close to a coin flip, so please read a re-run as worth attempting, not as a fix.

And if it does come back green, that is not proof this delivery fits the lane. It is one
sample at 14 seconds of headroom. We're saying that in advance so a green badge isn't later
read as evidence it wasn't.

The lane being cut at twenty minutes is also not specific to this PR: in the merge-base run
above, Behavior portable serial 4 was cut the same way on main. The ceiling catches
whichever lane is closest that day.

The actual fix, which is not this PR

The real remedy is the lane-packing repair, which is a separate delivery already in flight and
itself queued behind another change. We are deliberately not doing it here, and we have
also deliberately not trimmed our new tests to squeeze under the cap — that would buy a green
badge by deleting the coverage this change exists to add, and it would only hold until the next
test anyone adds.

@mremond

mremond commented Sep 9, 2026

Copy link
Copy Markdown
Contributor Author

Follow-up: we stopped waiting, and this is recorded as passed-with-override

No re-run arrived, so we set the request down rather than leave it open indefinitely. Recording
plainly what happened, because the difference matters more than the badge:

We asked for the verdict, we gave the facts including our own tightening in numbers, we
waited a stated period — two hours, from 12:56:56Z to 14:56:56Z — and we did not receive
it
. Our validation therefore records this as passed-with-override, with one lane carrying no
verdict
. Not "all checks green", and not "flaky".

The distinction we want left visible: this was not an answer that was unobtainable. It was
obtainable — one click — we asked for it, and we stopped waiting. That is a different thing from
a verdict nobody could have produced, and it should not be read as the first kind later.

Two notes on the numbers, since the checks summary misleads in both directions:

  • Nothing failed. Behavior portable serial 1 reported total=30 failed=0 and was then cut
    at 20 minutes 0.4 seconds, during post-job cleanup, after passing. The summary flattens that
    cancellation into "failed".
  • The report count is inflated by us. The summary counts check reports, not distinct
    checks: PR must be raised via no-mistakes appears three times because that workflow re-fires
    on every description edit, and our own disclosure edits caused two of them. There are 15
    distinct checks: 14 passed, 1 with no verdict.

The CI cost section in the description stays as it is. This change consumes 39 of the 53
seconds of that lane's headroom and leaves 14 — real whether or not the lane survives it, and
yours to weigh.

Re-running that lane is still welcome if you'd like a real verdict on it. The lane-packing
repair remains the actual fix and is a separate delivery, not this one.

@mremond mremond changed the title fix: bound watcher shutdown and improve failure diagnostics fix(bin): bound watcher shutdown on recovery lock failures Sep 14, 2026
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: first stamp on PR #4073 (mremond).

Head 1129818ef720b6827ae75eb690957c2c7393d82b (fm/fm-installed-timeout-watcher-will-not-stop, signed merge of main into tip) vs main b182d0f908b78d08c7ccb8dce3775bdca8c5d657. MERGEABLE / UNSTABLE. Author not blocked. Recent author activity 2026-09-14 (merge + attestation refresh) — not stale. Greptile Confidence 5/5 (earlier tip; no actionable security/correctness findings).

Attestation: MATCH — body <!-- no-mistakes-pipeline-attestation:v1 … "head_sha":"1129818ef720b6827ae75eb690957c2c7393d82b" …> equals tip. Tip Require no-mistakes SUCCESS (runs 34826728265 / 34826819671 after body refresh; earlier synchronize FAIL was pre-attestation).

Contract-class: new-default. Evidence: (1) WATCHER_SHUTDOWN_LOCK_SECS=5 is an always-on internal policy constant with no prior art in bin/fm-watch.sh — under a held recovery-marker lock the watcher now gives up after 5s, may leave the downtime marker unpublished, and retains stale singleton evidence (observable change vs unbounded wait that eventually published); author commit message itself states the case against restore (“5 is a new policy constant… If the maintainer classes this as new behavior we take the human gate”). (2) fm_lock_try_create now returns 2 for owner-creation/preparation failures vs 1 for contention — new lock-recovery branch semantics. Restorative intent (TERM must stop the watcher; existing “could not be persisted” branch made reachable; empty-ps → UNKNOWN not LIVE) is real but does not pull the PR below new-default while the always-on 5s bound and return-2 path land unconfigured. Never auto-merge.

VISION.md per-rule

  • One captain, one interface — aligns (stop/failure stays below-deck mechanics; diagnostics only on already-failing assertion paths).
  • Authority is explicit and never inferred — tension: always-on 5s bound changes downtime-marker publication under contention without captain opt-in — why new-default / captain card when otherwise ready.
  • Scripts own the mechanics, agents own the judgment — aligns (lock deadline + liveness states stay in scripts/test lib).
  • A restart is a non-event — aligns (bounded cleanup so a wedged marker lock cannot prevent watcher stop/restart; stale lock evidence retained for next arm).
  • Delegation with a spine — aligns (mutation-proven stop assertion; no scope creep into lane packing).
  • The fleet outlives any vendor — neutral/aligns (process/lock mechanics, not vendor UI).
  • Scope — aligns (watcher/wake-lib/docs/tests only). Author deliberately removed ambient/public timeout overrides — good refusal of scope growth.

CI/NM: Behavior portable (all serial/parallel) + Herdr + macOS + invariants + coverage + timing aggregate SUCCESS on CI run 34820326660. Lint FAILURE three times on the same run’s failed-job reruns — each time bin/fm-lint.sh starts ShellCheck 0.11.0 full extended analysis, then the runner sends shutdown (exit 143) at ~8–9 minutes with no ShellCheck findings emitted. This is infrastructure cancel, not an authored lint defect (by contrast #4427’s Lint on the same window completed SUCCESS in ~14m). Tip NM MATCH/SUCCESS. Merge-eligible N until Lint records a real completed verdict. Waiting-ci (maintainer will keep retrying Lint; author need not re-push solely for this cancel). Firstmate flag: no (not otherwise-ready while Lint is red; new-default captain card only once Lint is a true green).

Security: Diff review clean — lock/liveness hardening; no credential paths; no auth bypass; no .github/**. Greptile 5/5. No security captain gate. Security tip N/A.

Overlaps: Same watcher surface as open #4253 (bin/fm-watch.sh / bin/fm-wake-lib.sh) — different defect (shutdown bound vs turn-end token burn). Tip is ahead-of-main with merge already applied; #4253 is DIRTY behind main — rebase ordering will matter if both land. Related historical CI context in body cites #4056 / lane packing (#4006) as separate deliveries — not competing with this PR’s scope.

The watcher's EXIT trap took the recovery-marker lock with an unbounded
wait, so a live holder made SIGTERM a no-op: the watcher stayed up until
SIGKILL. Measured on the real script - with a live holder the signalled
watcher was still running after 20s, where an uncontended shutdown takes
0.15s. A supervisor that cannot stop its watcher by signalling it has no
supervision, so shutdown now refuses instead of blocking.

Three changes, one theme: a stop that cannot be observed is not a stop.

1. bin/fm-wake-lib.sh routes every recovery-marker lock acquisition
   through one helper, and bin/fm-watch.sh bounds only the shutdown
   transition (FM_WATCHER_SHUTDOWN_LOCK_SECS, default 5). Every other
   caller keeps the unbounded wait it relies on, because the default is
   empty. On the deadline watcher_cleanup's existing failure branch fires
   and names the holder, so the watcher stops and says what it could not
   persist, and it retains the stale lock evidence the next arm reclaims
   exactly as that branch already did for an unwritable marker.

2. tests/fm-pr-check-security.test.sh's watcher-stop assertion now prints
   what it saw - elapsed time, poll count, ps state and wchan, live
   descendants, and the watcher's stderr tail - on the failing path only.
   A CI sighting of that assertion was unexplainable after the fact
   because it reported a verdict with no evidence. The 150-poll budget is
   deliberately unchanged: the measured margin over a healthy shutdown is
   15-20x, so raising it would only hide a real failure.

3. tests/lib.sh now owns one liveness helper with three states: live,
   gone, and ps could not answer. The two former copies read an empty ps
   answer as LIVE, so a process reaped between the kill -0 and the ps
   read - the likeliest moment for a parent shell to reap it - was
   reported running, and a ps hiccup and a genuinely stuck process
   produced the same verdict and the same message. They are now different
   messages.

SUBSTITUTION, FLAGGED FOR REVIEW.
The obvious tool for (1) was fm_lock_acquire_wait_bounded, already in
this file. I built it that way first and it was WRONG, measured rather
than argued: tests/fm-pr-check-security.test.sh then failed
"signaled watcher left its singleton lock", deterministically on Linux
2/2 and macOS 2/2 in a whole-file run, where the unmodified tree passes
2/2 on both. Instrumented, the watcher reported status=1 with the marker
lock held by its own pid. That helper delegates acquisition to a CHILD
process, and from that child the caller's own abandoned hold is
indistinguishable from a live foreign holder, so it cannot acquire a lock
the caller itself already holds. The exit path is exactly that case: a
signal can land inside a recovery-marker critical section and the EXIT
trap then re-enters it, which is why fm_lock_try_acquire carries an
in-process self-held reclaim. The bound is therefore a deadline around
the ordinary in-process acquire. Same fm_lock_try_acquire, same stale
recovery and self-held reclaim, one added outcome (124 on the deadline).
Its dependency footprint is smaller, not larger: no helper process, no
external timeout binary, and nothing new to source from a dying shell.

ERREXIT, and why the new call sites are written the way they are.
tests/fm-pr-check-security.test.sh turns errexit on and off around
individual commands and leaves it ON for every case that follows, so in a
whole-file run a bare command returning non-zero ends the script with
status 1, no assertion and no message. My first version of the stop loop
called the liveness helper bare and read $?, which is exactly that shape:
the suite printed 25 oks and stopped, silently, in the case under change.
Every three-state call is therefore written `|| state=$?` - a condition
context, exempt from errexit - and the call sites say why. Worth a
maintainer's eye on its own: any bare command added to a later case in
that file truncates the suite the same way.

CONTRACT CLASS, with the case against it.
I read (1) as restoring intended behavior. watcher_cleanup already had a
"could not be persisted, retaining stale lock evidence" branch and a
non-zero cleanup status, so the design already contemplated this
transition not completing; the unbounded wait simply made that branch
unreachable for a held lock.
Against that reading: observable behavior does change. A watcher that
used to wait, and would have published the marker once the holder let go,
now gives up after 5s and leaves the downtime marker unpublished and the
singleton lock stale, and 5 is a new policy constant with no prior art in
this file. That is a fair reading. If the maintainer classes this as new
behavior we take the human gate rather than repackaging it.
I read (3) as new behavior in test infrastructure rather than a pure fix:
a third return code is genuinely new, and about 45 existing call sites
now read an unreadable ps differently than before. The case for calling
it a fix is that the old direction was simply wrong.
(2) changes no product behavior and prints only on an already-failing
path.

PROVEN BY MUTATION, each restored afterwards.
- Bound removed in watcher_cleanup: the new case goes red with "signaled
  watcher did not stop while the recovery marker lock was held".
- Deadline removed from the shared acquire helper: same red.  Both
  transitions route through that one helper, so there is no second site
  to miss; that is the structural answer to the trap-site hazard below
  rather than a second test.
- Both TERM traps in bin/fm-watch.sh disabled: the drain assertion goes
  red and now prints the state that makes it readable.
- Only the file-level TERM trap disabled: the drain case stays GREEN,
  because run_check_capture re-installs the disposition on every check.
  Both trap sites now carry a comment saying so, since a mutation applied
  to one alone proves nothing.
- Empty-ps direction reverted to LIVE: the liveness contract case goes
  red.

DELIBERATELY NOT CHANGED, and worth a maintainer's eye.
bin/fm-watch.sh takes the same marker lock unbounded at STARTUP
(fm_recovery_marker_reopen_announced and fm_recovery_marker_arm_check,
before the first beacon touch), and _fm_recovery_marker_arm_check holds
the wake-queue lock while it waits, so a held marker lock wedges a
starting watcher before it publishes any liveness at all. Blocking there
is defensible, because a watcher that cannot read its recovery state
arguably should not start, and the scope here is the exit trap. It is
left alone and flagged rather than widened.
@mremond
mremond force-pushed the fm/fm-installed-timeout-watcher-will-not-stop branch from 1129818 to 31e4d35 Compare September 15, 2026 12:20
@mremond mremond changed the title fix(bin): bound watcher shutdown on recovery lock failures fix(bin): bound watcher shutdown and improve stop diagnostics Sep 15, 2026
@mremond

mremond commented Sep 21, 2026

Copy link
Copy Markdown
Contributor Author

Checks are green on this PR now; the last triage stamp reads waiting-ci from before they finished. Flagging it for a re-check when the queue allows - nothing else is owed from our side.

@mremond

mremond commented Sep 25, 2026

Copy link
Copy Markdown
Contributor Author

Closing as superseded: the watcher shutdown flake this addressed is fixed on main by #5362 and #5381.

@mremond mremond closed this Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants