Skip to content

fix(gateway): fit the drain to the ARMED watchdog deadline, one source (#838 P1 close-out) - #861

Merged
Kyzcreig merged 10 commits into
mainfrom
fix/drain-armed-deadline-t74bb9994
Sep 23, 2026
Merged

Kyzcreig merged 10 commits into
mainfrom
fix/drain-armed-deadline-t74bb9994

Conversation

@Kyzcreig

@Kyzcreig Kyzcreig commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Closes the 4 unresolved P1 from FleetReview's trusted record on #838
(#838 (comment)),
which merged 7 minutes after that record was posted. New branch off main; #838 is not reopened.

ROOT PATTERN: there were TWO deadlines and the drain was fitted to the wrong one.
resolve_elapsed_adjusted_drain fitted against exit_timeout - hard_exit_reserve_s - reserve,
but the watchdog that actually calls os._exit is armed from
resolve_armed_shutdown_watchdog_delay = min(effective_drain + max(grace, reserve), hard_exit).
Those agree only when the OUTER min() binds (the gui-clamped 60). On a non-gui-clamped
system-domain job where the INNER leash binds (clamp 300 / configured 180 / measured teardown 70:
armed 250 vs hard-exit-derived 220) the drain overran into the teardown window.

F1 — one function returns the armed deadline; every site consumes it

New resolve_stop_drain_deadline_s() in gateway/restart.py is THE ONE DEADLINE
(armed watchdog minus teardown reserve). resolve_elapsed_adjusted_drain now derives
no arithmetic of its own — it calls the resolver. gateway/run.py computes it once in
_stop_impl_body from _armed_shutdown_deadline_s (the value actually handed to
arm_shutdown_watchdog, captured at the arming site — not a second derivation) and
threads it to both the drain fit and the cron leash.
Also extracted resolve_stop_teardown_reserve_s() as THE ONE RESERVE, so the
max(cleanup_reserve, measured) + actionable-ceiling filter exists once instead of
three copies that could drift.

F2 — cron branch consumes the deadline instead of re-deriving from the SIGKILL wall

resolve_cron_drain_budget(deadline_s=...): when supplied it REPLACES the
watchdog_delay/cleanup_reserve_s derivation entirely, so the ceiling is exactly the
deadline - elapsed the non-cron drain is fitted to. The old path clamped to the raw
ExitTimeOut and held back only the 10s CRON_DRAIN_CLEANUP_RESERVE_S, which raised the
budget back to the hard-exit instant three statements after the elapsed adjustment.
The "cron floor only ever EXTENDS" intent is preserved (the max(drain, ...) stands;
drain is itself already deadline-fitted, so it is not a loophole).

F3 — invariant test no longer re-derives the function under test

The old sweep recomputed its expected deadline with
resolve_launchd_shutdown_watchdog_delay(clamp, clamp, ...) — the same expression the code
used — so it held by construction. Replaced with hard-coded worked examples
(clamp 300 / drain 180 / teardown 70 / grace 60 -> concrete numbers) plus an oracle that
observes the armed wall rather than recomputing it.

F4 — watchdog re-armed after the elapsed is measured

The top-of-stop() arming sizes a RELATIVE drain budget before the pre-drain cost is known,
so the armed window silently absorbed that elapsed and the teardown reserve paid for it.
_rearm_shutdown_watchdog() extends the deadline by exactly the measured elapsed once known.
EXTEND-ONLY and still bounded by resolve_launchd_shutdown_watchdog_delay (so never past
exit_timeout - hard_exit_reserve_s, and a no-op when the wall already binds). Re-arm order
is: arm the replacement, THEN disarm the superseded thread — no instant without a backstop —
and the finally sets every event armed on the path so a superseded thread cannot hard-exit
a completed shutdown.

Verification

Focused slice on this head, .venv interpreter: see PR comments for the measured run.
Independent verification by argus (card t_f753b2b5) against a real GatewayRunner.stop()
with in-flight cron, observing the armed wall from the actual arm_shutdown_watchdog() call:
control fork/main dd7e9bc 10/10 UNSAFE -> this head 0/10 unsafe, cron_past_armed_wall 0/10,
with all 4 control paths preserved (non_launchd_zero_drain 50.0, cron_optout_zero 28/28,
in_band_restart 50/100, inner_leash_clamp300 180/180).

Card: t_74bb9994. Gateway restart is gated on this landing.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

Verification on f96a4c0 (interpreter: ~/.hermes/hermes-agent/.venv/bin/python3)

1. Focused slice — GREEN

tests/gateway/test_launchd_exit_timeout_drain_cap.py -> 73 passed in 2.43s, rc=0.

Broad slice run per-file (the 34-file combined run silently truncates at 535/649 with
exit 0 — pre-existing os._exit truncation on clean main dd7e9bc, killer index 535
test_restart_resume_pending.py::test_startup_restore_zero_timeout_intentionally_disables_watchdog;
not introduced here). All 50 test_*{drain,cron,shutdown,restart}*.py files, each with its own
summary line and rc checked separately from any pipe:

nonzero-rc files: 0
files with no summary line: none
TOTAL PASSED: 565   (50 files, 0 failed, 0 error)

2. Bidirectional RED proof (F1/F2/F3 defects, behavioral — not just ImportError)

Clean worktree at pre-fix dd7e9bc4, same probe, hand-computed oracle:

PRE-FIX dd7e9bc4  clamp300 configured180 teardown70: capped=180.0 armed(os._exit)=250.0
   elapsed=  0.0 drain=180.0 ends@180.0  teardown window=70.0 (needs 70.0)  ok
   elapsed= 20.0 drain=180.0 ends@200.0  teardown window=50.0 (needs 70.0)  VIOLATION
   elapsed= 40.0 drain=180.0 ends@220.0  teardown window=30.0 (needs 70.0)  VIOLATION
   elapsed= 60.0 drain=160.0 ends@220.0  teardown window=30.0 (needs 70.0)  VIOLATION

PRE-FIX  clamp60 drain30 cronfloor30 elapsed20:
   armed=50.0 adjusted_drain=15.0 -> cron_budget=30.0 ends@50.0 window=0.0 (needs 15)
   i.e. the drain ends at the exact instant os._exit fires; the entire teardown reserve is gone.

Same probe on this head:

THIS HEAD f96a4c0b  clamp300 configured180 teardown70:
   elapsed=  0.0 armed=250.0 drain=180.0 ends@180.0  window=70.0  ok
   elapsed= 20.0 armed=270.0 drain=180.0 ends@200.0  window=70.0  ok
   elapsed= 40.0 armed=290.0 drain=180.0 ends@220.0  window=70.0  ok
   elapsed= 60.0 armed=290.0 drain=160.0 ends@220.0  window=70.0  ok

THIS HEAD  clamp60 drain30 cronfloor30 elapsed20:
   armed=50.0 deadline=35.0 adjusted=15.0 cron_budget=15.0 ends@35.0 window=15.0 (needs 15)

The reserve is now preserved at every elapsed (F4's re-arm buys back exactly the elapsed until the
launchd wall binds at 290, after which the drain shrinks instead — never the reserve).

The three new tests also go RED on pre-fix code (ImportError on the new resolvers plus the
rewritten sweep), confirmed in the same worktree: 4 failed, 6 passed.

3. Single-derivation proof (REQUIRED item 1)

$ grep -rn "hard_exit - \|armed - reserve" gateway/ --include=*.py
gateway/restart.py:219:    cap = max(hard_exit - reserve, 0.0)          <- the drain CAP (different quantity)
gateway/restart.py:514:    return max(armed - reserve, 0.0)             <- THE ONE DEADLINE, sole site

$ grep -rn "max(_seconds(cleanup_reserve_s)" gateway/ --include=*.py
gateway/restart.py:254:    return max(_seconds(cleanup_reserve_s), measured)   <- THE ONE RESERVE, sole site

resolve_stop_drain_deadline_s has exactly one consumer chain: gateway/run.py:19992 computes it
once from the captured _armed_shutdown_deadline_s and threads it to the drain fit
(armed_deadline_s=) and the cron leash (deadline_s=). No second derivation of either quantity.

4. Mutation gate (REQUIRED item 5) — 7/7 caught

Each mutant applied to gateway/restart.py, focused suite run, then reverted:

CAUGHT  M1 deadline drops the teardown reserve                        1 failed, 31 passed
CAUGHT  M2 deadline reverts to the hard-exit wall (#838 defect)       1 failed, 30 passed
CAUGHT  M3 cron branch re-derives from the SIGKILL wall (F2 defect)   1 failed, 47 passed
CAUGHT  M4 arming leash ignores the drain                             1 failed, 18 passed
CAUGHT  M5 arming leash ignores the measured teardown reserve         1 failed, 24 passed
CAUGHT  M6 reserve ignores the measured sample                        1 failed, 16 passed
CAUGHT  M7 re-arm drops the elapsed (F4 defect restored)              1 failed, 46 passed

mutants caught: 7/7
post-restore: 73 passed, rc=0   (tree clean, git status empty)

5. Independent verification (argus, card t_f753b2b5)

Real GatewayRunner.stop() with in-flight cron, observing the armed wall from the actual
arm_shutdown_watchdog() call rather than re-deriving it:
control fork/main dd7e9bc 10/10 UNSAFE -> this head 0/10 unsafe, cron_past_armed_wall 0/10.
Controls preserved: non_launchd_zero_drain 0.0/50.0, in_band_restart 50.0/100.0,
cron_optout_zero 27.999/27.999, inner_leash_clamp300 180.0/180.0.
argus recorded t_f753b2b5 PASS against this head and superseded PR #835.

…FleetReview r1)

FleetReview round 1 on f96a4c0: 5/6 arms independently found the same P1
in the finding-4 re-arm. All P0/P1 closed here.

P1 (5 arms) — absolute-vs-relative unit mismatch in _rearm_shutdown_watchdog.
  _armed_shutdown_deadline_s is ABSOLUTE (from the start of stop()); every
  consumer treats it that way. arm_shutdown_watchdog is RELATIVE
  (deadline = time.monotonic() + delay). The top-of-stop() arming is at t=0
  where the two coincide, so handing the absolute value into a LATER re-arm
  charged the pre-drain elapsed twice and scheduled os._exit at
  elapsed + deadline. At clamp 300 / drain 180 / teardown 70 / elapsed 40
  that fires at 330 against a launchd SIGKILL at 300 — the backstop never
  runs, no forensic dump, no ordered lock/PID release, SIGKILL mid-teardown.
  Exactly the class this change exists to close.
  Fix: publish the absolute deadline (drain + cron consume it unchanged),
  arm max(deadline - elapsed_now, 0). Skip the re-arm entirely if that is
  <= 0 rather than replacing a live backstop with a zero-delay no-op.

P1 (B-state) — chained pytest.approx. `a == approx(x) == approx(y)` also
  evaluates approx == approx, which is invalid; split into two asserts.

P1 (L6) — the re-arm was UNREACHABLE under pytest (PYTEST_CURRENT_TEST
  early-return), so deleting the whole block left the suite green.
  New test_stop_rearms_the_watchdog_with_REMAINING_time_not_the_absolute_deadline
  drives the real stop() with the marker cleared and a 0.9s pre-drain cost,
  reads both values handed to arm_shutdown_watchdog.

P2 (L6) — the call-site wiring was untested: armed_deadline_s / deadline_s
  are structurally None under pytest, so the spies could not see them.
  New test_stop_threads_the_one_deadline_into_both_the_drain_and_the_cron_leash
  clears the marker and asserts both consumers got the same armed-derived
  deadline (hand-computed 50 / 28 at the production clamp).

P3 (B-assert-ctx, C) — the forensic snapshot re-derived watchdog_delay_s
  with elapsed_s=0, so a dump written after a re-arm reported the
  superseded deadline. Now reads the published value.

Verified (.venv interpreter):
  focused file            75 passed, rc=0  (was 73)
  50-file per-file slice  567 passed / 0 failed, every file rc=0 + summary
  mutation gate on the new tests, each reverted after:
    re-arm passes the absolute deadline again  -> CAUGHT (armed 250.901 vs 250)
    re-arm never invoked                       -> CAUGHT (got [250.0], expected 2)
    drain armed_deadline_s=None                -> CAUGHT
    cron deadline_s=None                       -> CAUGHT
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

Round 2 — all FleetReview P0/P1 from f96a4c0b closed on 384be7fd

Round 1 record: 6/6 arms completed, 5 of them independently found the same P1 in my
finding-4 re-arm. It is a real defect and they are right. Fixed, not argued.

P1 (B-assert-ctx, B-state, C-assert-xhigh, F, L6) — absolute-vs-relative unit mismatch

_armed_shutdown_deadline_s is ABSOLUTE (measured from the start of stop()) and every consumer
treats it that way; arm_shutdown_watchdog is RELATIVE (deadline = time.monotonic() + delay).
The top-of-stop() arming is at t=0 where the two coincide — which is exactly why passing the
absolute value into a LATER re-arm was invisible. It charged the pre-drain elapsed twice and
scheduled os._exit at elapsed + deadline; at clamp 300 / drain 180 / teardown 70 / elapsed 40
that is 330 against a launchd SIGKILL at 300, so the backstop never runs at all.

Fix: keep publishing the absolute deadline (drain + cron math unchanged), arm
max(deadline - elapsed_now, 0), measured from the same _stop_started_at_box["t"] the deadline
is expressed in. If the remaining time is <= 0 the re-arm is skipped entirely rather than replacing
a live backstop with a zero-delay no-op (arm_shutdown_watchdog silently returns on delay <= 0).

P1 (B-state) — chained pytest.approx

a == approx(x) == approx(y) also evaluates approx == approx. Split into two asserts. Correct
catch — it would have gone red in CI.

P1 (L6) — the re-arm was unreachable under pytest, hiding the above

Confirmed: PYTEST_CURRENT_TEST early-return meant deleting the entire re-arm block left the suite
green. New test_stop_rearms_the_watchdog_with_REMAINING_time_not_the_absolute_deadline drives the
real stop() with the marker cleared and a genuine 0.9s pre-drain cost, then reads both values
handed to arm_shutdown_watchdog.

P2 (L6) — call-site wiring untested (armed_deadline_s / deadline_s structurally None)

Also correct. New test_stop_threads_the_one_deadline_into_both_the_drain_and_the_cron_leash
clears the marker and asserts BOTH consumers received the same armed-derived deadline
(hand-computed 50 armed / 28 deadline at the production clamp).

P3 (B-assert-ctx, C-assert-xhigh) — forensic snapshot reported the superseded deadline

watchdog_delay_s re-derived with elapsed_s=0, so a dump written after a re-arm documented the
wrong deadline for the very hard-exit it was recording. Now reads the published value.

P3s not taken

  • "re-arm log line formats an absolute deadline as relative +Ns" — fixed as part of the P1: the
    line now reads re-armed to stop()+Xs (was stop()+Ys; Zs from now at elapsed Es).
  • "rewritten sweep still holds by construction for the reserve half" (C P3 / L6 P2) — the sweep is
    no longer the only oracle: the hard-coded worked example
    (test_drain_deadline_is_the_armed_watchdog_at_the_inner_leash_geometry, 250/70/180 with an
    explicit != 220) and the two new call-site tests are independent of the implementation's
    expression. Mutation results below show all four mutants die.

Verification on 384be7f (~/.hermes/hermes-agent/.venv/bin/python3)

focused file:            75 passed, rc=0            (was 73)
50-file per-file slice:  567 passed / 0 failed      (was 565; +2 new tests)
                         every file rc=0, every file printed its summary line

Per-file, never a combined run — the 34-file combined run silently truncates at 535/649 with exit 0
on clean main too (pre-existing os._exit truncation, killer index 535
test_restart_resume_pending.py::test_startup_restore_zero_timeout_intentionally_disables_watchdog).

Mutation gate on the NEW tests, each mutant reverted and re-verified green after:

CAUGHT  re-arm passes the absolute deadline again
        -> armed 250.9010339579545 vs expected 250.0 ± 0.35
CAUGHT  re-arm never invoked (if False and callable(...))
        -> "got [250.0] — if this is 1 the re-arm never fired and finding 4 is unwired"
CAUGHT  drain fit armed_deadline_s=None
        -> "it is re-deriving the deadline instead of consuming the value the watchdog was armed with"
CAUGHT  cron leash deadline_s=None
        -> "it is re-deriving from the raw SIGKILL wall, which is finding 2"
post-restore: 75 passed, rc=0

The round-1 mutation gate on the resolver layer (7/7) is in the previous comment and still holds.

@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

FleetReview

Reviewed with 2 of 3 model families — openai unavailable.

Confidence: 3/5

Findings

  • P1 gateway/run.py:19824 — Watchdog re-arm double-counts the pre-drain elapsed, pushing os._exit past launchd's SIGKILL wall
  • P2 gateway/run.py:19758 — Watchdog diagnostic snapshot reports the pre-re-arm deadline after a re-arm
  • P2 tests/gateway/test_launchd_exit_timeout_drain_cap.py:149 — Re-arm oracle test asserts on the resolver's return value only; nothing pins the delay actually handed to arm_shutdown_watchdog
  • P0 tests/gateway/test_launchd_exit_timeout_drain_cap.py:793 — Chained approx assert
  • P1 tests/gateway/test_launchd_exit_timeout_drain_cap.py:657 — Test claims to cover the watchdog re-arm but only exercises the pure resolver; the re-arm is unreachable under pytest and hides an absolute-vs-relative delay defect
  • P2 tests/gateway/test_launchd_exit_timeout_drain_cap.py:527 — Rewritten sweep invariant is an algebraic identity of the implementation for the reserve half, so it cannot detect a wrong reserve

FleetReview provenance · models: C=claude-code-opus-5, D=grok-4.6, G=grok-4.6 · cost: $11.23 · duration: 34m 59s · rounds: 2 · files examined: 3

…vation DIVERGES

The round-2 call-site test (test_stop_threads_the_one_deadline_into_both_*)
is pinned at the production clamp 60, where the outer min() of the arming
expression binds and a consumed deadline equals a re-derived one (both 28).
An independent 4-mutant gate run showed it therefore SURVIVES hard-wiring
armed_deadline_s=None at gateway/run.py:20042 — i.e. the gate for finding 1
could not detect finding 1. That is the #838 root pattern (measuring the one
geometry where two different deadlines coincide) reproduced in the suite.

Adds test_drain_consumes_the_REARMED_deadline_at_the_inner_leash_geometry:
clamp 300 / configured 180 / measured teardown 70 / real 0.9s pre-drain cost,
where consumed = 250.9 - 70 = 180.9 and re-derived = 250 - 70 = 180. All
expected values hand-computed, independent of every resolver.

Verified: 76 passed rc=0 on the focused file. Mutant armed_deadline_s=None
at the call site: SURVIVED before, CAUGHT after.
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

Round 3 — a gate hole I found in my OWN round-2 test, on 91e07749

Not a FleetReview finding. While round 2 was in review I re-ran the mutation gate
independently (my own harness, mutants chosen by reading the call sites rather than
by reusing round 2's list). One mutant SURVIVED:

SURVIVED  drain armed_deadline_s=None   <-- GATE HOLE
          75 passed in 5.87s

What was wrong

test_stop_threads_the_one_deadline_into_both_the_drain_and_the_cron_leash — the round-2
test written specifically to gate finding 1 — runs at the production clamp 60. There the
OUTER min() of the arming expression binds, so the deadline the code CONSUMES and the
deadline re-derivation would have produced are the same number (28):

clamp 60  / cfg 50  / teardown 22 / elapsed 20:  CONSUMED 28.0  REDERIVED 28.0   <- indistinguishable
clamp 300 / cfg 180 / teardown 70 / elapsed 20:  CONSUMED 200.0 REDERIVED 180.0  <- diverges
clamp 300 / cfg 180 / teardown 70 / elapsed 40:  CONSUMED 220.0 REDERIVED 180.0

So hard-wiring armed_deadline_s=None at gateway/run.py:20042 — finding 1, exactly —
left the whole suite green at 75 passed.

This is the #838 root pattern reproduced one level up: measuring the single geometry
where two different deadlines coincide. Round 2's mutant for this finding mutated
run.py:20058 (the resolve_elapsed_adjusted_drain kwarg) and was caught; the sibling
kwarg at :20042 feeding resolve_stop_drain_deadline_s was not covered. Finding 3 said a
test that re-derives the implementation's expression is worthless; a test pinned to the one
geometry where the defect is invisible is the same failure with a different shape.

The fix

test_drain_consumes_the_REARMED_deadline_at_the_inner_leash_geometry — clamp 300,
configured 180, measured teardown 70, plus a real 0.9s pre-drain cost so the re-arm
actually extends the deadline. Every expected value hand-computed, independent of all
resolvers:

capped drain       = 180                              (fits under the wall)
armed at t=0       = min(180 + max(60, 70), 290)      = 250
re-armed           = min(0.9 + 180 + 70, 290)         = 250.9
deadline CONSUMED  = 250.9 - 70                       = 180.9
deadline REDERIVED = 250   - 70                       = 180     <- what the mutant yields

with an explicit discriminating assertion (armed > 250.25, deadline > 180.25) so the
test fails on the re-derived value rather than merely on None.

Verification on 91e07749 (~/.hermes/hermes-agent/.venv/bin/python3)

focused file:            76 passed, rc=0          (was 75)
50-file per-file slice:  568 passed / 0 failed    (was 567; +1 new test)
                         every file rc=0, every file printed its summary line

Per-file, never a combined run (the 34-file combined run silently truncates at 535/649 with
exit 0 on clean main too — pre-existing os._exit truncation, killer index 535
test_restart_resume_pending.py::test_startup_restore_zero_timeout_intentionally_disables_watchdog).

Mutation gate re-run on 91e07749, each mutant reverted and the tree re-verified green after:

CAUGHT  M1 re-arm passes the absolute deadline again
        -> armed 250.90221904218197 vs expected 250.0 ± 0.35
CAUGHT  M2 re-arm never invoked (if False and callable(...))
        -> "got [250.0] — if this is 1 the re-arm never fired and finding 4 is unwired"
CAUGHT  M3 drain armed_deadline_s=None        <-- SURVIVED on 384be7fd, caught now
        -> "cron deadline 180.0 is not armed(250.902) - reserve(70)"; assert 180.0 == 180.9 ± 0.5
CAUGHT  M4 cron leash deadline_s=None
        -> "it is re-deriving from the raw SIGKILL wall, which is finding 2"
post-restore: 76 passed, rc=0; git status clean at 91e07749

4/4, up from 3/4. No production code changed in this commit — test-only, +136 lines.

⚠️ This push supersedes the FleetReview run that claimed 384be7fd at 06:28:20Z. That was
the right trade: the review would otherwise have been certifying a head whose gate for
finding 1 could not detect finding 1.

@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

FleetReview

Reviewed with 2 of 3 model families — openai unavailable.

Confidence: 3/5

Findings

  • P1 gateway/restart.py:493 — resolve_stop_drain_deadline_s trusts armed_deadline_s unclamped, so the drain can be fitted past launchd's SIGKILL
  • P3 gateway/restart.py:572 — Docstring cross-reference names the wrong keyword for the cron deadline handoff
  • P1 gateway/run.py:19877 — Re-arm aborts stop
  • P3 gateway/run.py:19758 — Forensic watchdog_delay_s changes units relative to the dump's own delay_s after a re-arm
  • P2 tests/gateway/test_launchd_exit_timeout_drain_cap.py:553 — Sweep no longer asserts the SIGKILL-wall invariant it is named for; the new check holds by construction
  • P2 tests/gateway/test_launchd_exit_timeout_drain_cap.py:1440 — No test covers an armed_deadline_s that disagrees with the live launchd budget, or the mid-stop signal flip
  • P1 gateway/restart.py:759 — Cron deadline unused
  • P1 gateway/run.py:19825 — Watchdog re-arm is unbounded — and unjustified — on every non-launchd stop path

FleetReview provenance · models: C=claude-code-opus-5, D=grok-4.6, G=grok-4.6 · cost: $27.40 · duration: 35m 38s · rounds: 1 · files examined: 3

daedalus-opus added 2 commits September 22, 2026 00:21
…ain the re-arm (FleetReview r2)

F1 (P1) resolve_stop_drain_deadline_s trusted armed_deadline_s unclamped.
The published _armed_shutdown_deadline_s is only wall-clamped when the
ARMING ran with signal_driven=True; an in-band stop(restart=True) arms with
it False, so the published value is the raw inner leash (240 at clamp 60).
A supervisor SIGTERM landing mid-stop flips the flag, and the drain/cron
reads then see signal_driven=True + that stale 240 -> deadline 225 against
an uncatchable SIGKILL at 60. The extend-only re-arm cannot rescue it: the
correct value (50) is EARLIER and fails the extend-only guard. Now clamped
to the same wall the recomputation branch is bounded by. Measured: 225 -> 35;
inert at both normal geometries (50->28, 250.9->180.9).

F4 (P1) the re-arm was guarded only by PYTEST_CURRENT_TEST, but its
'bounded by the SIGKILL wall' invariant only exists on launchd:
resolve_launchd_shutdown_watchdog_delay short-circuits without a live
ExitTimeOut, so on systemd/docker-s6/foreground the new deadline is an
uncapped elapsed+drain+max(grace,reserve) (measured 240/245/270/540 at
elapsed 0/5/30/300). No benefit there either -- resolve_stop_drain_deadline_s
returns None without a budget, so the drain is never elapsed-charged. Gated
to the launchd signal path.

F2 (P1) the re-arm committed shutdown state before the replacement backstop
was live. If arm_shutdown_watchdog raised, _stop_impl_body unwound at the
drain fit -- before drain/persist/teardown -- and the outer finally then set
every event, disarming the ORIGINAL watchdog too. Now: arm first, contain the
exception, commit + retire the old one only after arming returns.

F3 rejected as a false finding: deadline_s IS passed, at run.py:20092
(deadline_s=_stop_deadline_s). The reviewer read restart.py only. Proven live
by mutant M4 -- dropping it goes RED.

Verified (.venv): focused 80 passed rc=0 (was 76). Each new test proven RED on
384be7f in a clean-room worktree: F1 'deadline 225.0 ... expected 35',
F4 'got [240.0, 239.99998]', F2 RuntimeError propagated out of stop().
…t stopwatch)

A mutation-gate baseline run took 143s instead of 7s under concurrent load and
produced 1 failure. Cause: pytest.approx(250.0, abs=0.35) / approx(250.9, abs=0.5)
against a real asyncio.sleep(0.9) — a stopwatch reading, which load stretches.

Converted to the relational invariants the values actually encode, each of which
jitter can only strengthen:
  published > 250 and <= 290      (re-arm extended, still inside the wall)
  second   <= 250 and < published (armed the REMAINING time, not the absolute)
  armed    > 250, deadline > 180  (consumed the re-armed deadline; re-derivation
                                   yields EXACTLY 250/180, so > is discriminating)
  deadline == approx(armed - 70, abs=0.01)  (one deadline, exact, duration-free)

Mutation gate re-verified after the conversion — see the next commit's evidence.
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

FleetReview

Reviewed with 2 of 3 model families — openai unavailable.

Confidence: 3/5

Findings

  • P1 gateway/restart.py:493 — resolve_stop_drain_deadline_s trusts a caller-supplied armed_deadline_s without re-clamping it to the launchd wall — a stale arming lets the drain run past SIGKILL
  • P1 gateway/restart.py:759 — Cron path still opt-in
  • P3 gateway/restart.py:572 — Docstring names the wrong keyword for the cron deadline hand-off
  • P1 gateway/run.py:20037 — Published armed deadline is consumed without re-validating the arming context — a SIGTERM that arrives mid-stop fits the drain to an uncapped deadline past launchd's SIGKILL
  • P1 gateway/run.py:19882 — Superseded watchdog is disarmed unconditionally even when arming the replacement failed, leaving no backstop
  • P2 gateway/run.py:19860 — Unknown-stop-start fallback in the re-arm arms the ABSOLUTE deadline as a RELATIVE delay — the exact double-count the docstring calls fatal
  • P2 tests/gateway/test_launchd_exit_timeout_drain_cap.py:553 — Rewritten sweep is an algebraic identity — it can no longer detect the fix(gateway): close the round-3 FleetReview P1s on the shutdown budget #838 re-derivation defect it is named for
  • P2 tests/gateway/test_launchd_exit_timeout_drain_cap.py:1517 — Wall-clock tolerance of ±0.5s around a 0.9s sleep makes the two re-arm tests flaky under CI load
  • P2 tests/gateway/test_launchd_exit_timeout_drain_cap.py:1579 — _active_cron_job_count is patched on the class but shadowed by an instance binding — the "in-flight cron work" precondition never applies
  • P2 tests/gateway/test_launchd_exit_timeout_drain_cap.py:1746 — Clamp-300 re-arm tests only pass when real pre-drain elapsed lands in (0.5s, 1.4s] — flaky on loaded CI
  • P1 gateway/restart.py:417 — Ceiling-filtered reserve on the arming leash hard-exits early when last_teardown is unreservable

FleetReview provenance · models: C=claude-code-opus-5, D=grok-4.6, G=grok-4.6 · cost: $25.27 · duration: 27m 32s · rounds: 1 · files examined: 3

@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

Round 4 — all 4 FleetReview P1 from 384be7fd closed on be415a83

3 confirmed and fixed, 1 rejected with evidence. Production code changed for F1/F2/F4;
F3 was already wired.

✅ F1 (P1) — armed_deadline_s trusted unclamped. CONFIRMED, real, live.

Reproduced before fixing. The published _armed_shutdown_deadline_s is only wall-clamped
when the arming ran with signal_driven=True; an in-band stop(restart=True) arms with
it False, so resolve_launchd_shutdown_watchdog_delay short-circuits and publishes the raw
inner leash:

published by in-band arming (signal_driven=False): 240.0
wall (hard_exit) at clamp 60                     :  50.0
deadline with STALE unclamped armed :  225.0   <-- drain fitted ~190s past a SIGKILL at 60
deadline re-derived (correct)       :   35.0

The reviewer's note that the extend-only re-arm cannot rescue it is also correct: the fresh
clamped value (50) is EARLIER, fails _new <= _cur + 0.5, and 240 stays in force.

Fixed exactly as proposed — clamp the passthrough to the same wall the recomputation branch is
structurally bounded by. Measured after: 225 → 35, and inert everywhere legitimate
(clamp 60 armed 50 → 28; clamp 300 re-armed 250.9 → 180.9).

✅ F4 (P1) — re-arm unbounded off the launchd path. CONFIRMED.

resolve_launchd_shutdown_watchdog_delay short-circuits without a live ExitTimeOut, so the
docstring's "still bounded by the SIGKILL wall" is false on systemd / Docker-s6 /
--external-supervisor / foreground:

elapsed=  0.0: armed(no launchd budget) = 240.0
elapsed=  5.0:                          = 245.0
elapsed= 30.0:                          = 270.0
elapsed=300.0:                          = 540.0      <- bounded by nothing but the elapsed

And the reviewer's second half is right too — there is no compensating benefit:
resolve_stop_drain_deadline_s returns None without a budget, so the drain is never
elapsed-charged and finding 4's defect cannot arise there. Cost without correctness. The
re-arm is now gated to the launchd signal path.

✅ F2 (P1) — re-arm committed shutdown state before the backstop was live. CONFIRMED.

Ordering fixed: append the event, arm first, contain the exception, then commit
(_shutdown_watchdog_done, _armed_shutdown_deadline_s) and only then retire the old
watchdog. On failure it logs and returns, so stop() continues under the watchdog it already
had instead of unwinding at the drain fit and letting the outer finally disarm the original
too.

❌ F3 (P1) — "cron deadline unused / dead code". REJECTED, false finding.

deadline_s is passed by a caller. The reviewer examined restart.py only and said so
("this file does not show it"). The wiring is in run.py:

gateway/run.py:20142:                deadline_s=_stop_deadline_s,

Not an argument from inspection — it is mutation-proven live. Mutant M4 sets that kwarg to
None and the suite goes RED:

E  AssertionError: the cron leash received deadline_s=None — it is re-deriving from the
   raw SIGKILL wall, which is finding 2 of the #838 review

Dead code cannot be killed by a mutation. No change made.

REQUIRED 4 — exactly ONE derivation, grepped

== the ONE derivation (drain + leash, then the wall) ==
gateway/restart.py:423:    leash = max(_seconds(grace), reserve)          <- single site

== armed-delay call sites (non-test) — all CONSUME ==
gateway/run.py:19767  (forensic snapshot: reads the published value, falls back)
gateway/run.py:19790  (the arming site — publishes)
gateway/run.py:20075  (the re-arm — extends)
gateway/restart.py:532 (fallback inside resolve_stop_drain_deadline_s)

== THE ONE DEADLINE fn + consumers ==
gateway/restart.py:433  def resolve_stop_drain_deadline_s
gateway/run.py:20087    _stop_deadline_s = resolve_stop_drain_deadline_s(...)
gateway/run.py:20092/20108  armed_deadline_s=<published>   (drain fit x2)
gateway/run.py:20142        deadline_s=_stop_deadline_s    (cron leash)

No new arithmetic at any call site.

RED proof on 384be7fd (required item 3)

Clean-room worktree at 384be7fd, new tests copied in, production code untouched:

FAILED ...::test_stale_unclamped_armed_deadline_cannot_fit_the_drain_past_the_wall
  E  AssertionError: deadline 225.0 was taken from the stale unclamped armed value (240.0);
     expected the wall-clamped 35 (hard exit 50 - reserve 15)
  E  assert 225.0 == 35.0 ± 3.5e-05
FAILED ...::test_rearm_is_skipped_off_the_launchd_path_where_nothing_bounds_it
  E  AssertionError: expected ONLY the top-of-stop() arming off the launchd path,
     got [240.0, 239.99998008320108]
FAILED ...::test_rearm_failure_leaves_the_original_watchdog_armed_and_stop_running
  E  RuntimeError: can't start new thread     (propagated out of stop())
3 failed, 1 passed, 76 deselected

The 4th (test_wall_clamp_is_inert_when_the_armed_deadline_is_already_inside) is a control
asserting the clamp changes nothing legitimate — correctly green on both heads.

Every expected value is hand-computed from the geometry, never from the function under test.

⚠️ Also fixed: a wall-clock flake I introduced in round 2

A mutation-gate baseline took 143s instead of 7s under concurrent load and produced 1
failure. Cause was mine: pytest.approx(250.0, abs=0.35) against a real asyncio.sleep(0.9)
— a stopwatch reading, which load stretches. Converted to the relational invariants those
numbers actually encode, each of which jitter can only strengthen:

published > 250 and <= 290        re-arm extended, still inside the wall
second   <= 250 and < published   armed the REMAINING time, not the absolute
armed > 250, deadline > 180       consumed the re-armed deadline (re-derivation gives
                                  EXACTLY 250/180, so `>` is discriminating)
deadline == approx(armed - 70, abs=0.01)   one deadline, exact, duration-free

Verified the loosening did not make them vacuous — M1 and M3 still die (below).

Verification on be415a83 (~/.hermes/hermes-agent/.venv/bin/python3)

focused file:            80 passed, rc=0        (was 76)
50-file per-file slice:  572 passed / 0 failed  (was 568; +4 new tests)
                         every file rc=0, every file printed its summary line

Per-file, never combined — the 34-file combined run silently truncates at 535/649 with exit 0
on clean main too (pre-existing os._exit truncation, killer index 535
test_restart_resume_pending.py::test_startup_restore_zero_timeout_intentionally_disables_watchdog).

Full 7-mutant gate, each reverted and the tree re-verified green after:

CAUGHT   M1 re-arm arms the ABSOLUTE deadline      assert 250.90179 <= 250.0
CAUGHT   M2 re-arm never invoked                   assert 1 == 2
CAUGHT   M3 drain re-derives (armed=None)          assert 180.0 > 180.0
CAUGHT   M4 cron leash re-derives (deadline=None)  assert None is not None   <- F3 is LIVE
CAUGHT   M5 F1 wall clamp removed                  assert 225.0 == 35.0
CAUGHT   M6 F4 launchd gate removed                got [240.0, 239.99998]
CAUGHT   M7 F2 containment removed                 RuntimeError: can't start new thread
post-restore: 80 passed; tree: 0 dirty files

7/7. Note M3 SURVIVED two heads ago and was caught by the inner-leash test added in
91e07749 — that gate hole was found by re-running the mutation gate independently rather
than reusing the previous round's mutant list.

@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

FleetReview

Reviewed with 2 of 3 model families — openai unavailable.

Confidence: 3/5

Findings

  • P1 gateway/restart.py:791 — Cron drain floor can collapse to zero when the measured teardown reserve exceeds the remaining window
  • P3 gateway/restart.py:604 — Docstring cross-reference names the wrong parameter for the cron deadline hand-off
  • P3 gateway/restart.py:417 — Arming leash now discards an over-ceiling teardown sample, narrowing the armed window
  • P1 gateway/run.py:19911 — Re-arm disarms the live watchdog when the replacement thread silently fails to start — shutdown loses its hard-exit backstop
  • P2 gateway/run.py:19883 — Extend-only re-arm leaves the real watchdog armed past launchd's SIGKILL on the in-band-restart-raced-by-SIGTERM path
  • P3 gateway/run.py:19763 — Forensic snapshot's watchdog_delay_s changes units relative to the dump's delay_s after a re-arm
  • P2 tests/gateway/test_launchd_exit_timeout_drain_cap.py:553 — Reworked sweep in test_elapsed_adjusted_drain_never_runs_past_the_hard_exit_deadline is tautological and drops the only independent wall oracle
  • P2 tests/gateway/test_launchd_exit_timeout_drain_cap.py:1592 — _active_cron_job_count monkeypatch is inert — make_restart_runner binds it as an instance attribute that shadows the class patch
  • P1 tests/gateway/test_launchd_exit_timeout_drain_cap.py:2017 — F2 test mocks a failure arm_shutdown_watchdog cannot produce; the reachable silent-failure path is untested and disarms the only backstop

FleetReview provenance · models: C=claude-code-opus-5, D=grok-4.6, G=grok-4.6 · cost: $27.38 · duration: 32m 31s · rounds: 1 · files examined: 3

Kyzcreig and others added 2 commits September 23, 2026 00:33
… stop

Verify 123 passed, 2 skipped across six gateway test files. Four isolated mutants failed: cron deadline handoff, thread-start signal, failed-arm caller, wall-bound shortening; the arming wall mutant failed after adding a direct independent wall assertion. Accept and pin zero cron floor when teardown consumes the window.
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

🤖 merged-by: apollo · lane: fr-pause-0922 · gate: BYPASS: FR PAUSED by Ace ruling 2026-09-22 (state/fleetreview-pause-20260922.md); t_74bb9994 approved · why: argus run 7579 PASS-WITH-CAVEATS/SHIP, blocked only on #911 which Aegis merged; FR paused; CI green; Apollo merge pass 2026-09-22

@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 23, 2026
@Kyzcreig
Kyzcreig removed this pull request from the merge queue due to a manual request Sep 23, 2026
@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 23, 2026
Merged via the queue into main with commit 28d47a6 Sep 23, 2026
54 checks passed
@Kyzcreig
Kyzcreig deleted the fix/drain-armed-deadline-t74bb9994 branch September 23, 2026 18:44
@Kyzcreig Kyzcreig added the fleetreview:post-merge Ask FleetReview to review this MERGED pull (merge commit vs first parent) label Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

fleetreview:post-merge Ask FleetReview to review this MERGED pull (merge commit vs first parent)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant