Gateway replies, cron deliveries and the boot outbox are proven exactly-once through faults, crashes and DST (delivery E2E suites + CI) - #120344
Conversation
a704342 to
61267ce
Compare
૮ >ﻌ< ა ci reviewran on 3863b96 — test(e2e): cron soak fails on a broken fire-claim heartbeat; ℹ️ InfoCI-sensitive file review · View jobPR touches sensitive files, but the Sensitive files changed: debug infoCI timingsCI timings · View report · View jobWall time 5m48s vs 5m59s (-3.1%). 10 job(s) slower, 7 faster, 2 unchanged.
|
2cf408e to
f601e22
Compare
61267ce to
dede940
Compare
|
Review findings addressed — head Merge-order safety. New Each probe checked against every fix PR's head. Every probe reports open on this branch, and on each fix head only that PR's probe flips: Coin-flip xfails removed. zzz_unclean_restart compares Two-process fire-claim contention. In the soak, the tick lock admits one ticker per instant, so the second process never contended for a fire (a neutralised claim stayed green). The new Full files (
The 3 xfails left with every fix merged are strict live gaps that have no fix PR ( |
e658a83 to
5535d03
Compare
dede940 to
90c7c03
Compare
5535d03 to
276f56e
Compare
90c7c03 to
55833c8
Compare
55833c8 to
9af8d41
Compare
…safe probes A strict xfail(raises=AssertionError) also swallowed boot failures and timeouts, and flips main red (XPASS) the moment its fix merges. Each gap now has a probe that reproduces the defect's mechanism on the tree under test; the xfail applies only while the probe reports the defect open, and only for the bug-specific exception class the cell raises when it observes that leak (TenantLeak, LaunchProfileBleed, RoutingLeak, PromptLeftPending). Same pattern as #120344's delivery suite.
Runs the real InProcessCronScheduler loop (due scan, pending slots, fire claims + heartbeat, executions ledger, run_one_job, delivery routing, mark_job_run, manual run / trigger_job, a second process on the same HERMES_HOME, SIGKILL + boot recovery) under a file-backed virtual clock for ~30 virtual days per scenario across DST transitions and TZ configs (unset, foreign process TZ, Asia/Shanghai +08:00, America/New_York), outages, restarts mid-run, manual runs between ticks, runs longer than the fire-claim TTL and two contending processes. Only run_job and the platform send are faked. Oracle, independent of cron/jobs.py: croniter over wall time in the job's zone. After every tick the fired-occurrence set equals the model (no early, late, duplicate or skipped fire beyond the catch-up policy), every stored next_run_at equals the model and is in the future; at the end every execution is terminal and truthful (delivered <=> the fake platform got exactly that execution's message once). Red-proven against re-injected #105690 manual-run re-stamp, the fire-claim heartbeat deadlock shape, UTC next_run, skipped boot recovery, missing cross-process exclusion and the DST fold due-compare bug (#120314; newyork_on_shanghai_fall carries a strict xfail 'fixed by #120314' until it lands). #119969 is pinned as a strict xfail.
…-once Replaces the wave-2 skip stub. Drives the real delivery ledger, the real adapter record/finalize methods and the real GatewayRunner boot claim + redeliver halves in child processes on a temp HERMES_HOME, SIGKILLing at recorded-unsent, attempting-unsent, sent-unacked and inside a boot's own redelivery, plus 3 concurrent rebooters over 12 rows. Only the transport is fake (fsynced journal = ground truth). Invariants per obligation: <= 1 unmarked copy, >= 1 copy, terminal ledger row after the first clean boot, later reboots claim/send nothing, integrity_check ok. FIRE (strict xfail, not fixed): sweep_recoverable claims a 'pending' row without moving it to 'attempting' and the boot redelivery never calls mark_attempting, so a boot killed after the platform accepted a PLAIN redelivery leaves the row pending and the next boot resends it unmarked - two unmarked copies. Red-proven against dropped recovery marker, dropped claim CAS, skipped mark_delivered and re-claiming delivered rows.
A child-process GatewayRunner (own HERMES_HOME, SIGKILL + restart) runs the real agent against the scripted fake LLM provider and three instances of a FakePlatformAdapter that implements the gateway/platforms/base.py contract (4096/2000 limits, edits vs no edits, threads, streaming on/off). Only the transport is fake; its fsynced op journal is ground truth. Faults per op: timeout, ack lost, 429 retry_after, connection reset, too long, and park points for SIGKILL before/after the platform applies an op or between provider completion and the delivery ledger. Oracle per inbound: exactly one complete unmarked visible reply (extra copies only with the ledger's duplicate marker), visible text == persisted assistant text, user row persisted once, model turn run once, no stray or stale partial reply; plus a whole-run audit across chats. Scenarios: 13 faults x 3 platforms, /queue and interrupt follow-ups on a slow turn, parallel threads, re-delivered inbound ids (during/after a turn, after a runner reconnect), five crash points with restart catch-up, and an unclean restart with nothing in flight. Red-proven against the reverted stream fixes (#120315), a hand-revert of c961e5b (#91653), retrying timed-out sends, boot redelivery without the marker, double-sent finals, a dropped queued follow-up and disabled inbound dedup. Five strict xfails pin four live gaps (streaming ack-lost duplicate, dedup lost on reconnect, streaming crash after accept, and the unclean-restart recency fallback that re-answers every session active in the last 120 s). Eight cells stay red until #120315 lands (every streamed reply over the limit, and the first-send timeout) and carry xfail 'fixed by #120315': strict where the bug fires every run, non-strict where it depends on the stream/edit tick. The Director counts a resume note merged into the original tagged user message as a resume, not a second run of that inbound.
The e2e job already discovers tests/e2e/core/delivery/ (C12 messaging exactly-once, C13 cron virtual-clock soak) through `run_tests.sh --include-integration tests/e2e`; both need the 900 s per-file budget under load (C12 150-590 s on a loaded 20-core box). Cell 5 lives under tests/conformance/ and runs in the unit job.
…y cells A strict xfail on a gap whose fix is an open PR turns main red the moment that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces the defect's mechanism on the tree under test in a throwaway interpreter; expect_gap() applies the strict xfail only while the probe still reproduces it, so the cell becomes a plain test once the fix is in the tree, whatever the merge order. Covered: #120314, #119970 (soak), #120315, #120377, #120444 (C12) and #120450 (cell 5). Each probe was checked against every fix head: it flips on its own PR and on no other. C12 cells made deterministic (identical outcome on every run): - long_split streams the whole reply as one chunk; long_streamed and stream_timeout_first_send pace chunks so each lands in its own consumer tick. The five former coin-flip xfails are now two plain cells and three strict #120315 gap cells. - sent_ack_lost waits until the answer is persisted before the kill, so it pins the #120377 recovery; the streamed-before-persisted order is its own cell (stream_accepted_unpersisted, a strict live gap with a stalled provider stream). - zzz_unclean_restart compares director.resumes against a snapshot taken before its kill instead of requiring it empty: crash cells on their own homes may legitimately resume. - the whole-run audit skips the reconnect replay only while #120444's gap is open.
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.
test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.
The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
… gaps with no fix PR #120314, #120377, #120444 and #120450 are on main: their PROBES entries, every expect_gap naming them and the gap_open(120444) audit branch go, so those cells are plain tests again (two of those probes read source text, which the suite must not do). Cell 5's README row no longer claims a strict xfail. The two static strict xfails with no probe and no fix PR (STREAM_ACK_LOST_GAP, STREAM_CRASH_AFTER_ACCEPT_GAP) become run-time xfails via _pending_fixes.known_failure: only the final assertions run under it, after every wait (restart, catch-up, settle) has succeeded, and only an AssertionError matching the gap's own signature XFAILs; any other failure stays red and a fixed tree simply passes. The whole-run audit now skips only tokens whose cell actually XFAILed this run.
…xfail scoped to its deviation ev_long_run used to flag a missing claim refresh as stalled_heartbeat and carry on, so neutralising heartbeat_fire_claim in _heartbeat_loop left all five scenarios green. Every virtual 30 s of the 15-minute hold now (a) waits for the run's heartbeat to restamp the claim and fails if it never does, and (b) has a contender call the real claim_job_for_fire and asserts it loses, also once the first stamp is older than the TTL. With the heartbeat replaced by True all 5 scenarios are red on the refresh assertion; on main they are green. The #119970 cell no longer carries a whole-scenario strict xfail. While its behavioural probe reproduces, only the oracle's fired-set and next_run_at checks may deviate: each deviation is recorded, the model follows the stored slot, and the soak runs every virtual day with all other checks live, XFAILing at the end only if a deviation was seen (and failing if the probe says open but none was). The #120314 entry is gone (merged). A scenario that fails mid-hold now releases its held run before teardown so it cannot bleed into the next scenario.
9af8d41 to
3863b96
Compare
…safe probes A strict xfail(raises=AssertionError) also swallowed boot failures and timeouts, and flips main red (XPASS) the moment its fix merges. Each gap now has a probe that reproduces the defect's mechanism on the tree under test; the xfail applies only while the probe reports the defect open, and only for the bug-specific exception class the cell raises when it observes that leak (TenantLeak, LaunchProfileBleed, RoutingLeak, PromptLeftPending). Same pattern as #120344's delivery suite.
…safe probes A strict xfail(raises=AssertionError) also swallowed boot failures and timeouts, and flips main red (XPASS) the moment its fix merges. Each gap now has a probe that reproduces the defect's mechanism on the tree under test; the xfail applies only while the probe reports the defect open, and only for the bug-specific exception class the cell raises when it observes that leak (TenantLeak, LaunchProfileBleed, RoutingLeak, PromptLeftPending). Same pattern as #120344's delivery suite.
…safe probes A strict xfail(raises=AssertionError) also swallowed boot failures and timeouts, and flips main red (XPASS) the moment its fix merges. Each gap now has a probe that reproduces the defect's mechanism on the tree under test; the xfail applies only while the probe reports the defect open, and only for the bug-specific exception class the cell raises when it observes that leak (TenantLeak, LaunchProfileBleed, RoutingLeak, PromptLeftPending). Same pattern as #120344's delivery suite.
Gateway replies, cron deliveries and the boot outbox are now checked end to end for exactly-once delivery through faults, SIGKILLs, restarts and DST transitions; the only expected failures left are the gaps that are still open, and each one is merge-order safe.
Tests + CI only, with no production code. Fixes these suites found that are already on main: #120314, #120377, #120444, #120450 (their cells are plain passing tests now). Still open: #120315 (streaming overflow / uneditable preview) and #119970 (no-tz DST).
What lands
tests/e2e/core/delivery/test_messaging_exactly_once.py(58 tests) plus_fake_platform.py.GatewayRunnerruns in a child OS process: ownHERMES_HOME, SIGKILL-able, restartable.tests/e2e/core/delivery/test_cron_virtual_clock_soak.pyplus_cron_clock.py.InProcessCronSchedulerloop runs under a file-backed virtual clock for about 30 virtual days × 4 TZ configs, across DST, outages, a restart mid-run, manual/trigger runs, a run longer than the fire-claim TTL, two contending processes and a SIGKILL.tests/conformance/persistence/test_cell5_delivery_outbox_exactly_once.py.e2ejob already runsscripts/run_tests.sh --include-integration tests/e2ewithHERMES_TEST_FILE_TIMEOUT=900, sotests/e2e/core/delivery/is picked up with the 900 s per-file budget it needs. The only change is naming these suites in that budget's comment; adding a second step would run them twice. Cell 5 runs in the unit job (21–76 s).Expected failures: open gaps only, each merge-order safe
fault_matrix[burst_overflow-fk_tg/dc],[long_streamed-fk_dc],[stream_timeout_first_send-fk_tg/dc]_pending_fixes.PROBES, runsGatewayStreamConsumerin a throwaway interpreter); strict xfail applied after every wait, right before the final assertions, only while the probe reproducesfixed→ plain testsoak[unset_on_newyork_spring]compute_next_run); while it reproduces, ONLY the schedule oracle's fired-set /next_run_atchecks may deviate: each deviation is recorded, the model adopts the stored slot and the soak runs all 12 virtual days with every other check live; XFAIL at the end only if a deviation was seen (and a FAIL if the probe says open but none was)fault_matrix[ack_lost-fk_tg/dc]_pending_fixes.known_failure) around the final assertions only, keyed on the gap's own message^\d+ unmarked copies of A-fk_(tg|dc)\.ack_lost \(silent duplicate\); any other failure stays redcrash[stream_accepted_unpersisted-fk_tg]2 unmarked complete replies after catch-up/both the held answer and an auto-resumed re-answer are visiblewith the heldA-…AND a resumedR-…copy in the dumpGone (fixed on main, probes + every
expect_gapremoved; two of them read source text): #120314 (C13newyork_on_shanghai_fall), #120377 (C12crash[sent_ack_lost-fk_tg],test_zzz_unclean_restart_reruns_nothing), #120444 (C12redelivered[after_reconnect]), #120450 (cell 5 plain-redelivery kill). The whole-run audit now skips only tokens whose cell actually XFAILed in that run.Re-review fixes (head
3863b96049f1, rebased on main)expect_gapnaming them,GAPS['newyork_on_shanghai_fall']and thegap_open(120444)audit branch; cell 5 README row updated. Remaining probes (fix(cron): keep server-local jobs on their wall clock across DST changes #119970, Streamed replies that overflow the platform limit are no longer duplicated or left as a stale partial (#25349, salvage #114151) #120315) are behavioural only (fix(cron): keep server-local jobs on their wall clock across DST changes #119970 and Streamed replies that overflow the platform limit are no longer duplicated or left as a stale partial (#25349, salvage #114151) #120315 both still OPEN).crash[stream_accepted_unpersisted-fk_tg]: no marker up front; the run-time xfail wraps only the final assertions, after fault, inject, kill -9, restart, catch-up and settle waits all succeeded. fix(cron): keep server-local jobs on their wall clock across DST changes #119970 soak: see table above, XFAIL only on its deviation, rest of the soak live.ev_long_run: every virtual 30 s of the 15-min hold asserts the claim was re-stamped by the run's heartbeat AND that a contender's realclaim_job_for_fireloses (also past the 300 s TTL). A failure mid-hold now releases the held run so it can't bleed into the next scenario.STREAM_ACK_LOST_GAPandSTREAM_CRASH_AFTER_ACCEPT_GAPhave no fix PR, so they are run-time xfails on the known failure message (_pending_fixes.known_failure), not probes: other failures stay red, a fixed tree passes.Sabotage proof for (b) —
heartbeat_fire_claim(...)incron/scheduler.py::_heartbeat_loopreplaced byTrue(both calls):timed out … brief-0900: fire-claim heartbeat refresh), replicas 1 passedunset_on_newyork_spring, #119970)Suites at this head (
HERMES_TEST_FILE_TIMEOUT=900 scripts/run_tests.sh --include-integration …, retries off): 3 files, 61 passed, 0 failed.test_messaging_exactly_once.pyburst_overflow-fk_tg/dc,long_streamed-fk_dc,stream_timeout_first_send-fk_tg/dc; run-time:ack_lost-fk_tg/dc,crash[stream_accepted_unpersisted-fk_tg]test_cron_virtual_clock_soak.pyunset_on_newyork_spring(1 deviation:next_run_at 2026-03-08T09:00:00-05:00 != expected 13:00Z)test_cell5_delivery_outbox_exactly_once.pyRed proofs: 19 total. Each bug was re-injected at its current location, and each suite goes red.
(n/n)(#120315 fix 1 reverted)burst_overflow-fk_tg/dcvisible reply != persisted assistant text … first diff at 4002burst_overflow-fk_dc,long_split-fk_tgvisible reply != persisted … first diff at 7182send_timeout-fk_netimed out … waiting for fk_ne.send_timeout ledger to settlequeued_followup[fk_tg],[fk_ne]answer A-fk_tg.q2 persisted 0xack_lost-fk_ne,stream_ack_lost_final_edit-fk_ne2 unmarked copies … (silent duplicate)needs_markercrash[sent_ack_lost-fk_ne]2 unmarked complete replies after catch-upclean-fk_neunmarked=2 marked=0redelivered[during_turn],[after_turn]answer … persisted 0xstream_timeout_first_send-fk_tg/dca truncated/stale partial rendition … stayed visibleunset_utc_springfired set mismatch … expected=[('brief-0900', None, …)]newyork_on_shanghai_springfired set mismatch: actual=[] expected=[('gap-0230', …)]shanghai_on_la_fallnext_run_at 2026-10-23T09:00:00+00:00 != expected 2026-10-23T01:00:00+00:00newyork_on_shanghai_fallevery-2h fired at 01:14 EDT for the 01:11 EST slot (57 min early)recover_interrupted_executionsunset_utc_springrestart left the SIGKILLed run non-terminalnewyork_on_shanghai_fallbrief-0900 fired twice2 UNMARKED copies reached the platform36 claims for 12 obligationsmark_deliveredskipped after redeliveryledger row not terminal after a clean boot sweepfirst clean boot claimed [4 ids], owed [3 ids]Stability, earlier head
61267ce3f4a1(10 consecutive runs,scripts/run_tests.sh --include-integration --file-timeout 900 --file-retries 0, on a shared 20-core box at load 50–240)C12: 10/10 at head
61267ce3f4a1, retries off.The xpassed cells are the non-strict
fixed by #120315marks on runs where the timing-dependent bug did not trigger. Every strict xfail failed as expected on every run.5 files, 15 tests passed, 0 failed.Python tests / e2eat this head: C1244✓ 9xf 4xp, 138 sand C133✓ 2xf, 48 s.Harness fixes in this revision
send_timeout-fk_tgtimed out, 2 times in 3 runs at load 110–240. A stub that is not followed by its completion is still reported as stale, and that path is what R9 proves.Not covered
Compaction mid-drain; real platform adapter code (only the base contract is faked);
busy_input_modevariants beyond interrupt and/queue; runtime flood redelivery timing; multi-profileadapter_profilerouting; cron's live-adapter lane; the cron catch-up outbox; live-owner protection in the boot sweep. Details are in the suites' docstrings.Infographic