Skip to content

fix(bin): stop unbounded empty-wake resurface loop without dropping open decisions - #6

Merged
cfeddersen merged 6 commits into
mainfrom
fm/watcher-resurface-loop-r2
Sep 1, 2026
Merged

cfeddersen merged 6 commits into
mainfrom
fm/watcher-resurface-loop-r2

Conversation

@cfeddersen

@cfeddersen cfeddersen commented Sep 1, 2026 •

Copy link
Copy Markdown
Owner

Intent

Stop an unbounded empty-wake supervision loop (Claude Stop-hook auto-arm woken every 45-90s by 'check: rearm-resurface' draining empty) WITHOUT ever suppressing a real resurface. _fm_recovery_marker_publish preserves an already-acked marker on a routine close ONLY when there is no durable work a re-arm must surface. Atomicity: fm_wake_append publishes the marker BEFORE appending the row (base order; append-wake atomicity contract in tests/fm-wake-queue.test.sh, UNCHANGED, passes). The re-mint signal is an explicit caller-set literal 'work-to-surface' meaning 'there is durable work a re-arm must surface', not 'queue non-empty' - an OPEN DECISION is durable work that is not a queue row, so preserving over one would silently swallow it. Three callers signal it: fm_wake_append, watcher_cleanup's release-lock, and clear_stale_recorded_watcher_lock's clear-stale-lock (the last two check scan_open_decisions caller-side, outside the marker-lock critical section). The stale-lock steal is NOT susceptible (sets FM_LOCK_RECOVERED_PID). scan_open_decisions is the whole-file fold and never touches the drain's incremental cursor. Tests: tests/fm-wake-queue.test.sh and existing tests/fm-watch-arm.test.sh assertions UNCHANGED; a new test_restart_reused_pid_lock_surfaces_open_decision was ADDED (fails without the clear-stale-lock fix, passes with it); the pre-existing decision-only re-arm test pins the release-lock path. fm-watcher-lock and fm-watch-triage reconciliations from the loop fix remain. CI CONFIG CHANGE (deliberate, one job only): .github/workflows/ci.yml portable serial shard job timeout raised 20->30 minutes. Evidence: Behavior portable serial 4 on this head ran 20m17s and GitHub reported the timed-out job as conclusion 'cancelled' - it hit its own 20-minute ceiling, not an infrastructure cancellation. That ceiling had no real margin: main-branch baselines peaked at serial 3 = 19m48s and serial 4 = 17m18s against the 20-minute cap, and the regression tests added here (watcher-wake-lock family, serial lanes) consumed the remaining headroom and tipped serial 4 over. 30 restores a real tripwire margin. Only that job's timeout changed - not the matrix, shard count, other timeouts, or triggers.

What Changed

  • _fm_recovery_marker_publish (bin/fm-wake-lib.sh) now accepts an explicit work_to_surface argument; on a routine close over an already-acked marker (acked:handling:* / acked:downtime:*) it preserves the marker instead of re-minting a fresh recovery episode, but only when neither the wake queue is non-empty nor the caller explicitly signals work-to-surface — this stops the auto-arm Stop-hook from getting woken every 45-90s to drain an empty queue.
  • Three call sites are updated to signal work-to-surface when durable work exists that isn't visible as a queue row: fm_wake_append (enqueuing a wake row, still publishing the marker before the row is appended to keep the existing append-wake atomicity contract), watcher_cleanup's release-lock path in bin/fm-watch.sh, and clear_stale_recorded_watcher_lock in bin/fm-watch-arm.sh — the latter two check scan_open_decisions caller-side (outside the marker-lock critical section) so a still-open decision can't be silently swallowed by marker preservation.
  • fm_recovery_transition threads the new extra token through to release-lock and clear-stale-lock so callers can pass the work-to-surface signal.
  • CI: raised the "Behavior portable serial" shard job timeout in .github/workflows/ci.yml from 20 to 30 minutes to restore tripwire margin after the added regression tests pushed serial shard runtimes close to the prior cap.
  • Tests: added tests/fm-watch-recovery-loop.test.sh and a new test_restart_reused_pid_lock_surfaces_open_decision case in tests/fm-watch-arm.test.sh, plus reconciliation updates in tests/fm-watch-triage.test.sh and tests/fm-watcher-lock.test.sh; docs/watcher-continuity.md and docs/fm-test-portable-shards.md updated to describe the new behavior and CI timeout rationale.

Risk Assessment

✅ Low: The change is well-scoped to the recovery-marker publish/preserve logic and its three explicit call sites, preserves the pre-existing append-wake atomicity ordering, is backed by new tests that exercise the real fm-wake-lib.sh functions and observe actual marker/state transitions (not source-content matching), and the one-line CI timeout bump touches only the named job; tracing the acked-marker preserve/re-mint logic through fm_wake_append, watcher_cleanup's release-lock, and clear_stale_recorded_watcher_lock's clear-stale-lock shows each path correctly signals work-to-surface exactly when durable work (a queued wake or an open decision) exists, matching the stated intent.

Testing

Ran the full targeted regression matrix (fm-watch-recovery-loop, fm-watch-arm, fm-wake-queue, fm-watch-triage, fm-watcher-lock) live against the target commit; every test tied to this change's behavior passes, including three new library-driving tests (T3/T3b/T4) and one new end-to-end reused-pid-lock regression test that exercises the exact clear-stale-lock silent-loss scenario the fix addresses. Two unrelated test failures were found (a lock-reclaim mechanics test and an fm-procevent timing fixture); both were reproduced identically on the base commit in an isolated throwaway worktree, confirming they are pre-existing and untouched by this diff, not regressions. The CI timeout change and the untouched fm-wake-queue.test.sh file were also confirmed to match the stated intent exactly.

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-watch-recovery-loop.test.sh — all 5 pass, including the new T3/T3b/T4 tests that drive the real _fm_recovery_marker_publish/fm_recovery_marker_arm_check contract in a subprocess (acked+empty-queue close does not resurface; acked+queued-work still re-mints; unacked pending/announced still preserves and resurfaces)
  • bash tests/fm-watch-arm.test.sh — all pass, including the new test_restart_reused_pid_lock_surfaces_open_decision regression (clear-stale-lock path), verified this is a genuine before/after regression test by confirming the exact scenario is new in this diff
  • bash tests/fm-wake-queue.test.sh — one pre-existing unrelated failure (test_self_held_lock_reclaims_instead_of_deadlocking, a lock-reclaim mechanics test); reproduced the identical failure on base commit cb0e172 in an isolated worktree, confirming it predates and is untouched by this change; ran all remaining tests in the file (including test_wake_publish_requires_atomic_recovery_evidence, the append-wake atomicity contract) via a function-level runner skipping only that pre-existing failure — all pass
  • bash tests/fm-watch-triage.test.sh — one pre-existing unrelated failure (test_procevent_captured_result_surfaces_proactively, an fm-procevent.sh timing fixture unrelated to recovery-marker logic); reproduced the identical failure at the identical assertion on base commit cb0e172 in an isolated worktree, confirming it is pre-existing/environment-flaky and not caused by this diff; all other ~39 tests in the file pass
  • bash tests/fm-watcher-lock.test.sh — all tests pass
  • git diff cb0e172 5d6cb3f -- .github/workflows/ci.yml — confirmed the CI change is exactly the one-line portable-serial-shard timeout bump (20->30 minutes), no other timeout/matrix/trigger changes
  • git diff --stat cb0e172 5d6cb3f -- tests/fm-wake-queue.test.sh — confirmed empty diff, i.e. genuinely unchanged as the intent claims
⚠️ **Document** - 1 info
  • ℹ️ docs/verification/supervision.md:474 - Dated evidence block for tests/fm-watch-recovery-loop.test.sh predates the new settled-episode/durable-work/reused-pid-lock regression tests added in this change; refreshing it requires re-running the suite and capturing exact output, which is outside this document-only phase.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Christoph added 3 commits September 1, 2026 16:56
… close

A firstmate home on the Claude Stop-hook auto-arm model was woken every
45-90s by `check: rearm-resurface`, each draining to an empty queue - a
no-op supervision turn burning tokens continuously.

Root cause: on any watcher close that is not itself a rearm-resurface
delivery, watcher_cleanup runs the release-lock transition, whose
`_fm_recovery_marker_publish downtime` re-minted a fresh pending:downtime
episode over the `acked` marker the previous acknowledgement had just
written. The acknowledgement was therefore never durably consumed. Because
the Claude auto-arm re-arms a fresh NON-successor watcher every turn
(FM_WATCH_HANDLING_SUCCESSOR is never set on that path), the kunchenguid#2733
successor guard in resurface_after_downtime never engaged, so the next
watcher announced the re-minted episode and resurfaced with nothing to
handle. Persistent-adapter models (pi/opencode) were shielded only because
their successor guard suppressed the resurface; the same re-mint happened
there too.

Fix: a downtime publish preserves an already-acked marker when the durable
queue is empty. A clean close after acknowledgement is not a new
supervision gap, so a settled generation must not be reopened. Re-minting
stays justified only when durable work is actually queued: fm_wake_append
and the drain/arm-check adoption of an acked marker with a non-empty queue
re-mint through that same non-empty path, so real wakes are unaffected. A
mid-handshake death leaves pending/announced (never acked), so genuine
downtime recovery is untouched.

Scope is Defect A only; the separate teardown-marker residue is filed
independently and fm-teardown.sh is not touched here.

Regression coverage in tests/fm-watch-recovery-loop.test.sh, driving the
real recovery-marker library through the report's repro mechanism:
- a settled acked episode survives a clean close and a fresh non-successor
  arm does not resurface an empty queue (fails without this fix);
- an acked marker with durable queued work still re-mints an episode the
  drain can begin handling;
- the mirror safety property: a pending/announced mid-handshake death is
  preserved across a clean close and a pending gap still resurfaces.

docs/watcher-continuity.md, the recovery-episode contract owner, is updated
to state the preserved-on-empty-close behavior and the new coverage.
Follow-up to the acked-marker preservation fix. That change let a downtime
publish preserve an already-acknowledged episode when the durable queue is
empty. But fm_wake_append published the recovery marker BEFORE appending the
wake row, so its "is durable work queued" check ran against a queue that did
not yet contain the wake being enqueued. A genuine new wake arriving on an
acknowledged marker was therefore left unannounced until a later drain or arm
re-minted it - and CI's portable-serial suite caught it through two sibling
tests.

Fix: in fm_wake_append, enqueue the durable row (and its sequence) BEFORE
publishing the recovery marker, both under the queue lock. The publish now
sees the queued work and re-mints a fresh episode for it, exactly as before
the preservation fix. Ordering the append first also gives a stronger
invariant: the marker is never re-minted ahead of its row, so no observer can
see a published episode without the row that justified it. A clean close with
nothing queued still preserves the acknowledged marker (the empty-wake loop
fix), and a mid-handshake death still leaves pending/announced.

Test reconciliation for the genuinely-changed behavior: the old code re-minted
a phantom downtime episode on EVERY forced watcher stop, even one with nothing
outstanding. Two fixtures depended on draining that phantom away:

- fm-watcher-lock's cycle-exit ledger test drained a phantom after each forced
  HUP of a watcher holding a settled, empty-queue marker.
- fm-watch-triage's declared-pause test drained a phantom after a lifted-pause
  priming stop that surfaced nothing.

Both passed on the base and failed only under the fix, confirming the behavior
genuinely changed rather than a regression: a forced stop with nothing
outstanding now correctly leaves no recovery episode. The cleanly-classified
settled stops assert that positively via a new expect_settled_stopped_cycle
(no WAKE_ACK_REQUIRED, marker stays acked), which also guards against the
phantom ever returning. The many incidental cleanup acks in the triage suite -
fixture hygiene between phases, never the subject assertion, with each test
asserting its own wake surfacing separately via wait_for_exit and a grep on
the delivered output - are made tolerant of a settled stop while still failing
on a genuine drain error and still acknowledging a real queued wake.
…atomicity

Replaces the previous reorder. That reorder appended the wake row before
publishing the recovery marker, which inverted a deliberate atomicity
invariant: tests/fm-wake-queue.test.sh::test_wake_publish_requires_atomic_
recovery_evidence forces the marker's atomic write to fail mid-append and
asserts the wake row must NOT become durable, so recovery evidence is always
durable before the row. Appending first violated that.

Root defect was the inference, not the ordering: the preserve-acked check read
the queue to guess whether a wake was pending, and that read raced the very
append it was trying to observe. Fix removes the inference. fm_wake_append now
restores the base order - publish the recovery marker first, then append the
row, so a failed publish aborts before any row exists - and passes an explicit
wake-pending token to _fm_recovery_marker_publish. That token, set only by
fm_wake_append, re-mints a fresh episode over an acknowledged marker because
the caller states it is enqueuing real work; it is tested for the exact literal
so a stray or defaulted argument fails safe toward preservation. Every routine
close (release-lock, stale-lock steal) passes no token and keeps preserving an
acknowledged marker on an empty queue - the empty-wake loop fix - while the
drain and arm-check still re-mint an acknowledged marker that already has a
non-empty queue.

tests/fm-wake-queue.test.sh is unchanged and passes: the append-wake atomicity
contract is preserved exactly as base rather than rewritten to match the code.
The recovery-loop regression tests still hold both directions (settled leads to
no resurface; pending or announced still resurfaces), and the fm-watcher-lock
and fm-watch-triage reconciliations remain required - the clean-close
preservation they cover is unchanged by this follow-up. docs/watcher-
continuity.md is updated to describe the explicit-flag mechanism.
… open

Follow-up to the acked-marker preservation fix. Preserving a settled marker on
a routine close correctly stops the empty rearm-resurface loop, but the
condition used "the wake queue is non-empty" as its proxy for "there is durable
work a re-arm must surface". An OPEN DECISION the captain still owes an answer
to is durable supervision work that does NOT live as a wake-queue row, so a
decision-only re-arm saw an empty queue, preserved the acked marker, stayed
silent, and swallowed the decision - a silent-loss regression worse than the
noise the loop fix removed. CI's tests/fm-watch-arm.test.sh caught it via
test_rearm_resurfaces_durable_queue_and_remote_open_decision.

Fix: the flag now means "there is durable work a re-arm must surface", not "the
queue file is non-empty", and it is renamed from wake-pending to work-to-surface
so the name says what it means. The caller signals it, because callers know
things the queue does not. Two publish callers that could preserve an acked
marker over an open decision now check for one, caller-side and OUTSIDE the
marker-lock critical section (scan_open_decisions forks; that section must not):

- watcher_cleanup's release-lock (the clean-close path, the CI failure).
- clear_stale_recorded_watcher_lock's clear-stale-lock (a --restart reused-pid
  recovery path): it REMOVES the lock, so the fresh watcher gets no
  FM_LOCK_RECOVERED_PID and relies solely on the marker; the same swallow was
  reproduced there and is now closed. The stale-lock steal in fm_lock_try_acquire
  was investigated and is NOT susceptible - the steal sets FM_LOCK_RECOVERED_PID,
  so the stealing watcher resurfaces regardless of the marker.

fm_wake_append still passes the flag (it is enqueuing a wake row), the queued-row
case still re-mints, and a mid-handshake death still leaves pending/announced.
scan_open_decisions is the whole-file fold and provably never touches the drain's
incremental cursor, so this cannot itself become a silent-loss path.

Regression coverage, both directions proven (each fails without its fix, passes
with it): the pre-existing decision-only re-arm test pins the release-lock path,
and a new test_restart_reused_pid_lock_surfaces_open_decision pins the
clear-stale-lock path. tests/fm-wake-queue.test.sh and the existing assertions in
tests/fm-watch-arm.test.sh are unchanged - only a new test case was added.
docs/watcher-continuity.md is updated for the generalized flag.
@cfeddersen cfeddersen changed the title fix(bin): preserve acked recovery marker to break the empty-wake resurface loop fix(bin): preserve acked recovery marker only when no durable work needs surfacing Sep 1, 2026
Christoph added 2 commits September 1, 2026 21:14
Behavior portable serial 4 on this branch ran 20m17s and GitHub reported the
timed-out job as conclusion "cancelled" - the job hit its own 20-minute ceiling,
not an infrastructure cancellation. That ceiling already had no real margin:
main-branch baselines for these shards peaked at serial 3 = 19m48s and serial 4 =
17m18s against the 20-minute cap. The regression tests added on this branch land
in the watcher-wake-lock family on the serial lanes and consumed what little
headroom was left, tipping serial 4 over. Raise the cap to 30 to restore a real
tripwire margin. Only this job's timeout changes.
@cfeddersen cfeddersen changed the title fix(bin): preserve acked recovery marker only when no durable work needs surfacing fix(bin): stop unbounded empty-wake resurface loop without dropping open decisions Sep 1, 2026
@cfeddersen
cfeddersen merged commit 86db779 into main Sep 1, 2026
13 checks passed
cfeddersen added a commit that referenced this pull request Sep 4, 2026
…pen decisions (#6)

* fix(bin): preserve an acknowledged recovery marker on a clean watcher close

A firstmate home on the Claude Stop-hook auto-arm model was woken every
45-90s by `check: rearm-resurface`, each draining to an empty queue - a
no-op supervision turn burning tokens continuously.

Root cause: on any watcher close that is not itself a rearm-resurface
delivery, watcher_cleanup runs the release-lock transition, whose
`_fm_recovery_marker_publish downtime` re-minted a fresh pending:downtime
episode over the `acked` marker the previous acknowledgement had just
written. The acknowledgement was therefore never durably consumed. Because
the Claude auto-arm re-arms a fresh NON-successor watcher every turn
(FM_WATCH_HANDLING_SUCCESSOR is never set on that path), the kunchenguid#2733
successor guard in resurface_after_downtime never engaged, so the next
watcher announced the re-minted episode and resurfaced with nothing to
handle. Persistent-adapter models (pi/opencode) were shielded only because
their successor guard suppressed the resurface; the same re-mint happened
there too.

Fix: a downtime publish preserves an already-acked marker when the durable
queue is empty. A clean close after acknowledgement is not a new
supervision gap, so a settled generation must not be reopened. Re-minting
stays justified only when durable work is actually queued: fm_wake_append
and the drain/arm-check adoption of an acked marker with a non-empty queue
re-mint through that same non-empty path, so real wakes are unaffected. A
mid-handshake death leaves pending/announced (never acked), so genuine
downtime recovery is untouched.

Scope is Defect A only; the separate teardown-marker residue is filed
independently and fm-teardown.sh is not touched here.

Regression coverage in tests/fm-watch-recovery-loop.test.sh, driving the
real recovery-marker library through the report's repro mechanism:
- a settled acked episode survives a clean close and a fresh non-successor
  arm does not resurface an empty queue (fails without this fix);
- an acked marker with durable queued work still re-mints an episode the
  drain can begin handling;
- the mirror safety property: a pending/announced mid-handshake death is
  preserved across a clean close and a pending gap still resurfaces.

docs/watcher-continuity.md, the recovery-episode contract owner, is updated
to state the preserved-on-empty-close behavior and the new coverage.

* fix(bin): re-mint the recovery marker for a genuinely queued wake

Follow-up to the acked-marker preservation fix. That change let a downtime
publish preserve an already-acknowledged episode when the durable queue is
empty. But fm_wake_append published the recovery marker BEFORE appending the
wake row, so its "is durable work queued" check ran against a queue that did
not yet contain the wake being enqueued. A genuine new wake arriving on an
acknowledged marker was therefore left unannounced until a later drain or arm
re-minted it - and CI's portable-serial suite caught it through two sibling
tests.

Fix: in fm_wake_append, enqueue the durable row (and its sequence) BEFORE
publishing the recovery marker, both under the queue lock. The publish now
sees the queued work and re-mints a fresh episode for it, exactly as before
the preservation fix. Ordering the append first also gives a stronger
invariant: the marker is never re-minted ahead of its row, so no observer can
see a published episode without the row that justified it. A clean close with
nothing queued still preserves the acknowledged marker (the empty-wake loop
fix), and a mid-handshake death still leaves pending/announced.

Test reconciliation for the genuinely-changed behavior: the old code re-minted
a phantom downtime episode on EVERY forced watcher stop, even one with nothing
outstanding. Two fixtures depended on draining that phantom away:

- fm-watcher-lock's cycle-exit ledger test drained a phantom after each forced
  HUP of a watcher holding a settled, empty-queue marker.
- fm-watch-triage's declared-pause test drained a phantom after a lifted-pause
  priming stop that surfaced nothing.

Both passed on the base and failed only under the fix, confirming the behavior
genuinely changed rather than a regression: a forced stop with nothing
outstanding now correctly leaves no recovery episode. The cleanly-classified
settled stops assert that positively via a new expect_settled_stopped_cycle
(no WAKE_ACK_REQUIRED, marker stays acked), which also guards against the
phantom ever returning. The many incidental cleanup acks in the triage suite -
fixture hygiene between phases, never the subject assertion, with each test
asserting its own wake surfacing separately via wait_for_exit and a grep on
the delivered output - are made tolerant of a settled stop while still failing
on a genuine drain error and still acknowledging a real queued wake.

* fix(bin): re-mint via an explicit caller flag, restoring append-wake atomicity

Replaces the previous reorder. That reorder appended the wake row before
publishing the recovery marker, which inverted a deliberate atomicity
invariant: tests/fm-wake-queue.test.sh::test_wake_publish_requires_atomic_
recovery_evidence forces the marker's atomic write to fail mid-append and
asserts the wake row must NOT become durable, so recovery evidence is always
durable before the row. Appending first violated that.

Root defect was the inference, not the ordering: the preserve-acked check read
the queue to guess whether a wake was pending, and that read raced the very
append it was trying to observe. Fix removes the inference. fm_wake_append now
restores the base order - publish the recovery marker first, then append the
row, so a failed publish aborts before any row exists - and passes an explicit
wake-pending token to _fm_recovery_marker_publish. That token, set only by
fm_wake_append, re-mints a fresh episode over an acknowledged marker because
the caller states it is enqueuing real work; it is tested for the exact literal
so a stray or defaulted argument fails safe toward preservation. Every routine
close (release-lock, stale-lock steal) passes no token and keeps preserving an
acknowledged marker on an empty queue - the empty-wake loop fix - while the
drain and arm-check still re-mint an acknowledged marker that already has a
non-empty queue.

tests/fm-wake-queue.test.sh is unchanged and passes: the append-wake atomicity
contract is preserved exactly as base rather than rewritten to match the code.
The recovery-loop regression tests still hold both directions (settled leads to
no resurface; pending or announced still resurfaces), and the fm-watcher-lock
and fm-watch-triage reconciliations remain required - the clean-close
preservation they cover is unchanged by this follow-up. docs/watcher-
continuity.md is updated to describe the explicit-flag mechanism.

* fix(bin): never preserve an acked recovery marker while a decision is open

Follow-up to the acked-marker preservation fix. Preserving a settled marker on
a routine close correctly stops the empty rearm-resurface loop, but the
condition used "the wake queue is non-empty" as its proxy for "there is durable
work a re-arm must surface". An OPEN DECISION the captain still owes an answer
to is durable supervision work that does NOT live as a wake-queue row, so a
decision-only re-arm saw an empty queue, preserved the acked marker, stayed
silent, and swallowed the decision - a silent-loss regression worse than the
noise the loop fix removed. CI's tests/fm-watch-arm.test.sh caught it via
test_rearm_resurfaces_durable_queue_and_remote_open_decision.

Fix: the flag now means "there is durable work a re-arm must surface", not "the
queue file is non-empty", and it is renamed from wake-pending to work-to-surface
so the name says what it means. The caller signals it, because callers know
things the queue does not. Two publish callers that could preserve an acked
marker over an open decision now check for one, caller-side and OUTSIDE the
marker-lock critical section (scan_open_decisions forks; that section must not):

- watcher_cleanup's release-lock (the clean-close path, the CI failure).
- clear_stale_recorded_watcher_lock's clear-stale-lock (a --restart reused-pid
  recovery path): it REMOVES the lock, so the fresh watcher gets no
  FM_LOCK_RECOVERED_PID and relies solely on the marker; the same swallow was
  reproduced there and is now closed. The stale-lock steal in fm_lock_try_acquire
  was investigated and is NOT susceptible - the steal sets FM_LOCK_RECOVERED_PID,
  so the stealing watcher resurfaces regardless of the marker.

fm_wake_append still passes the flag (it is enqueuing a wake row), the queued-row
case still re-mints, and a mid-handshake death still leaves pending/announced.
scan_open_decisions is the whole-file fold and provably never touches the drain's
incremental cursor, so this cannot itself become a silent-loss path.

Regression coverage, both directions proven (each fails without its fix, passes
with it): the pre-existing decision-only re-arm test pins the release-lock path,
and a new test_restart_reused_pid_lock_surfaces_open_decision pins the
clear-stale-lock path. tests/fm-wake-queue.test.sh and the existing assertions in
tests/fm-watch-arm.test.sh are unchanged - only a new test case was added.
docs/watcher-continuity.md is updated for the generalized flag.

* ci: raise portable serial shard timeout from 20 to 30 minutes

Behavior portable serial 4 on this branch ran 20m17s and GitHub reported the
timed-out job as conclusion "cancelled" - the job hit its own 20-minute ceiling,
not an infrastructure cancellation. That ceiling already had no real margin:
main-branch baselines for these shards peaked at serial 3 = 19m48s and serial 4 =
17m18s against the 20-minute cap. The regression tests added on this branch land
in the watcher-wake-lock family on the serial lanes and consumed what little
headroom was left, tipping serial 4 over. Raise the cap to 30 to restore a real
tripwire margin. Only this job's timeout changes.

* no-mistakes(document): Fix stale CI timeout facts and add missing third recovery-marker caller to docs

---------

Co-authored-by: Christoph <astaran@herr-der-ringe-film.de>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant