Skip to content

Gateway replies, cron deliveries and the boot outbox are proven exactly-once through faults, crashes and DST (delivery E2E suites + CI) - #120344

Merged
teknium1 merged 8 commits into
mainfrom
tests/core-delivery
Sep 23, 2026
Merged

teknium1 merged 8 commits into
mainfrom
tests/core-delivery

Conversation

@teknium1

@teknium1 teknium1 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Gateway replies, cron deliveries and the boot outbox are now checked end to end for exactly-once delivery through faults, SIGKILLs, restarts and DST transitions; the only expected failures left are the gaps that are still open, and each one is merge-order safe.

Tests + CI only, with no production code. Fixes these suites found that are already on main: #120314, #120377, #120444, #120450 (their cells are plain passing tests now). Still open: #120315 (streaming overflow / uneditable preview) and #119970 (no-tz DST).

What lands

  • C12, tests/e2e/core/delivery/test_messaging_exactly_once.py (58 tests) plus _fake_platform.py.
    • A real GatewayRunner runs in a child OS process: own HERMES_HOME, SIGKILL-able, restartable.
    • It uses the real agent against the scripted fake LLM provider, and the real adapter, ledger, busy/queue/interrupt, stream consumer, boot redelivery, auto-resume and reconnect machinery.
    • Only the transport is fake. Three platform profiles: 4096 with edits, 2000 with edits and threads, 4096 without edits. The fsynced op journal is the ground truth.
    • Scenarios: 13 faults × 3 platforms, /queue and interrupt follow-ups, parallel threads, re-delivered inbound ids, 5 crash points with restart catch-up, and an unclean restart with nothing in flight.
    • Oracle: one complete unmarked visible reply per inbound, and visible text == persisted text.
  • C13, tests/e2e/core/delivery/test_cron_virtual_clock_soak.py plus _cron_clock.py.
    • The real InProcessCronScheduler loop runs under a file-backed virtual clock for about 30 virtual days × 4 TZ configs, across DST, outages, a restart mid-run, manual/trigger runs, a run longer than the fire-claim TTL, two contending processes and a SIGKILL.
    • Oracle: an independent croniter model checked after every tick, plus a truthful terminal status.
  • Conformance cell 5, tests/conformance/persistence/test_cell5_delivery_outbox_exactly_once.py.
    • Real delivery ledger, adapter record/finalize, and runner boot claim/redeliver in child processes.
    • SIGKILL at 4 points, plus 3 concurrent rebooters.
    • The README row is updated. The wave-2 stub file was already removed on the base.
  • CI: the e2e job already runs scripts/run_tests.sh --include-integration tests/e2e with HERMES_TEST_FILE_TIMEOUT=900, so tests/e2e/core/delivery/ is picked up with the 900 s per-file budget it needs. The only change is naming these suites in that budget's comment; adding a second step would run them twice. Cell 5 runs in the unit job (21–76 s).

Expected failures: open gaps only, each merge-order safe

cell(s) gap mechanism when the fix lands
C12 fault_matrix[burst_overflow-fk_tg/dc], [long_streamed-fk_dc], [stream_timeout_first_send-fk_tg/dc] streamed overflow / uneditable preview (#25349) behavioural probe for #120315 (_pending_fixes.PROBES, runs GatewayStreamConsumer in a throwaway interpreter); strict xfail applied after every wait, right before the final assertions, only while the probe reproduces probe says fixed → plain test
C13 soak[unset_on_newyork_spring] #119969 no-tz DST fixed offset behavioural probe for #119970 (calls compute_next_run); while it reproduces, ONLY the schedule oracle's fired-set / next_run_at checks may deviate: each deviation is recorded, the model adopts the stored slot and the soak runs all 12 virtual days with every other check live; XFAIL at the end only if a deviation was seen (and a FAIL if the probe says open but none was) plain test
C12 fault_matrix[ack_lost-fk_tg/dc] streaming first send applied, ack lost → the gateway's final send re-sends it unmarked (#53449/#25349 family) no fix PR, no probe: run-time xfail (_pending_fixes.known_failure) around the final assertions only, keyed on the gap's own message ^\d+ unmarked copies of A-fk_(tg|dc)\.ack_lost \(silent duplicate\); any other failure stays red passes on its own
C12 crash[stream_accepted_unpersisted-fk_tg] streaming turn visible, SIGKILL before persist → auto-resume answers again, unmarked no fix PR, no probe: run-time xfail after fault, kill -9, restart, catch-up and settle waits have all succeeded; matches only 2 unmarked complete replies after catch-up / both the held answer and an auto-resumed re-answer are visible with the held A-… AND a resumed R-… copy in the dump passes on its own

Gone (fixed on main, probes + every expect_gap removed; two of them read source text): #120314 (C13 newyork_on_shanghai_fall), #120377 (C12 crash[sent_ack_lost-fk_tg], test_zzz_unclean_restart_reruns_nothing), #120444 (C12 redelivered[after_reconnect]), #120450 (cell 5 plain-redelivery kill). The whole-run audit now skips only tokens whose cell actually XFAILed in that run.

Re-review fixes (head 3863b96049f1, rebased on main)

Sabotage proof for (b) — heartbeat_fire_claim(...) in cron/scheduler.py::_heartbeat_loop replaced by True (both calls):

tree soak result
main + heartbeat neutralised 5 failed (every scenario: timed out … brief-0900: fire-claim heartbeat refresh), replicas 1 passed
main 5 passed, 1 xfailed (unset_on_newyork_spring, #119970)

Suites at this head (HERMES_TEST_FILE_TIMEOUT=900 scripts/run_tests.sh --include-integration …, retries off): 3 files, 61 passed, 0 failed.

file passed xfailed xfails (all open gaps)
C12 test_messaging_exactly_once.py 50 8 #120315 probe: burst_overflow-fk_tg/dc, long_streamed-fk_dc, stream_timeout_first_send-fk_tg/dc; run-time: ack_lost-fk_tg/dc, crash[stream_accepted_unpersisted-fk_tg]
C13 test_cron_virtual_clock_soak.py 5 1 #119970 probe: unset_on_newyork_spring (1 deviation: next_run_at 2026-03-08T09:00:00-05:00 != expected 13:00Z)
Cell 5 test_cell5_delivery_outbox_exactly_once.py 6 0 —

Red proofs: 19 total. Each bug was re-injected at its current location, and each suite goes red.

suite injected bug red tests first failing assertion
C12 R1 overflow tail keeps (n/n) (#120315 fix 1 reverted) burst_overflow-fk_tg/dc visible reply != persisted assistant text … first diff at 4002
C12 R2 no re-split after seal (#120315 fix 2 reverted) burst_overflow-fk_dc, long_split-fk_tg visible reply != persisted … first diff at 7182
C12 R3 hand-revert of c961e5b (#91653): only flood refusals re-arm timed redelivery send_timeout-fk_ne timed out … waiting for fk_ne.send_timeout ledger to settle
C12 R4 queued follow-up dropped queued_followup[fk_tg], [fk_ne] answer A-fk_tg.q2 persisted 0x
C12 R5 timed-out sends retried inline ack_lost-fk_ne, stream_ack_lost_final_edit-fk_ne 2 unmarked copies … (silent duplicate)
C12 R6 boot redelivery ignores needs_marker crash[sent_ack_lost-fk_ne] 2 unmarked complete replies after catch-up
C12 R7 every successful final double-sent clean-fk_ne unmarked=2 marked=0
C12 R8 inbound dedup disabled redelivered[during_turn], [after_turn] answer … persisted 0x
C12 R9 uneditable preview (#120315 fix 3 reverted) stream_timeout_first_send-fk_tg/dc a truncated/stale partial rendition … stayed visible
C13 R1 off-tick manual run stamps the next occurrence (#105690) unset_utc_spring fired set mismatch … expected=[('brief-0900', None, …)]
C13 R2 fire-claim heartbeat renewal reports lost newyork_on_shanghai_spring fired set mismatch: actual=[] expected=[('gap-0230', …)]
C13 R3 next_run computed in UTC, not the configured zone shanghai_on_la_fall next_run_at 2026-10-23T09:00:00+00:00 != expected 2026-10-23T01:00:00+00:00
C13 R4 same-zone wall-clock due compare (#120314 reverted) newyork_on_shanghai_fall every-2h fired at 01:14 EDT for the 01:11 EST slot (57 min early)
C13 R5 gateway boot skips recover_interrupted_executions unset_utc_spring restart left the SIGKILLed run non-terminal
C13 R6 no cross-process exclusion newyork_on_shanghai_fall brief-0900 fired twice
Cell 5 1 no recovery marker on boot redelivery 4 of 5 cases 2 UNMARKED copies reached the platform
Cell 5 2 claim CAS dropped concurrent 36 claims for 12 obligations
Cell 5 3 mark_delivered skipped after redelivery all 5 ledger row not terminal after a clean boot sweep
Cell 5 4 sweep re-claims delivered rows all 5 first clean boot claimed [4 ids], owed [3 ids]

Stability, earlier head 61267ce3f4a1 (10 consecutive runs, scripts/run_tests.sh --include-integration --file-timeout 900 --file-retries 0, on a shared 20-core box at load 50–240)

C12: 10/10 at head 61267ce3f4a1, retries off.

run rc wall load avg at start per-file result
1 0 390 s 157.6 44 passed, 11 xfailed, 2 xpassed
2 0 224 s 132.8 44 passed, 11 xfailed, 2 xpassed
3 0 302 s 72.6 44 passed, 9 xfailed, 4 xpassed
4 0 222 s 103.5 44 passed, 10 xfailed, 3 xpassed
5 0 166 s 85.7 44 passed, 10 xfailed, 3 xpassed
6 0 154 s 31.1 44 passed, 9 xfailed, 4 xpassed
7 0 187 s 25.7 44 passed, 9 xfailed, 4 xpassed
8 0 175 s 50.0 44 passed, 9 xfailed, 4 xpassed
9 0 219 s 36.0 44 passed, 10 xfailed, 3 xpassed
10 0 174 s 73.9 44 passed, 10 xfailed, 3 xpassed

The xpassed cells are the non-strict fixed by #120315 marks on runs where the timing-dependent bug did not trigger. Every strict xfail failed as expected on every run.

  • C13 + cell 5 + conformance cells 1–3 at this head: 5 files, 15 tests passed, 0 failed.
  • CI Python tests / e2e at this head: C12 44✓ 9xf 4xp, 138 s and C13 3✓ 2xf, 48 s.
  • Earlier loops: C13 10/10 and cell 5 10/10, both from the lane (see REPORT).

Harness fixes in this revision

  • The Director counts a request whose last user message carries "previous turn was interrupted" as a resume, before looking for inbound tokens. Under load, the resumed turn's request merged the original tagged user message into the resume note, the Director scored it as a second run of that inbound, and the crash scenario waited forever for the resume reply.
  • The oracle joins a preview cut inside the reply header with the fallback continuation that completes it. When a final edit times out after such a preview, the gateway's documented continuation resends from the original cut (a prefix with no word boundary is not backed up). The reply was complete, but the oracle read it as a stale stub and send_timeout-fk_tg timed out, 2 times in 3 runs at load 110–240. A stub that is not followed by its completion is still reported as stale, and that path is what R9 proves.

Not covered

Compaction mid-drain; real platform adapter code (only the base contract is faked); busy_input_mode variants beyond interrupt and /queue; runtime flood redelivery timing; multi-profile adapter_profile routing; cron's live-adapter lane; the cron catch-up outbox; live-owner protection in the boot sweep. Details are in the suites' docstrings.

Infographic

delivery E2E suites

@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 3863b96 — test(e2e): cron soak fails on a broken fire-claim heartbeat;

ℹ️ Info

CI-sensitive file review · View job

PR touches sensitive files, but the ci-reviewed label has been added, approving them.

Sensitive files changed:


debug info

CI timings

CI timings · View report · View job

Wall time 5m48s vs 5m59s (-3.1%). 10 job(s) slower, 7 faster, 2 unchanged.

  • Python tests / e2e: +134.0s
  • JS & TS checks / JS & TS checks: +55.0s
  • Python tests / Run tests: +31.0s
  • OS-specific tests / Windows-only tests: -14.0s
  • Check contributors / check-attribution: -13.0s

@alt-glitch alt-glitch added type/test Test coverage or test infrastructure P3 Low — cosmetic, nice to have comp/gateway Gateway runner, session dispatch, delivery comp/cron Cron scheduler and job management sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages sweeper:risk-automation Sweeper risk: may affect CI, automerge, label sync, or maintainer automation labels Sep 23, 2026
@teknium1 teknium1 added the ci-reviewed applied to manually approve dangerous changes label Sep 23, 2026
@teknium1

Copy link
Copy Markdown
Collaborator Author

Review findings addressed — head dede9401047, rebased onto tests/core-e2e e658a831c98 (source tree == f601e220acc).

Merge-order safety. New tests/e2e/core/delivery/_pending_fixes.py: one probe per pending fix, each reproducing the defect's mechanism on the tree under test in a throwaway interpreter (behavioural for #120314 / #119970 / #120315 / #120450; for #120377 / #120444 it checks whether the buggy code path is still wired). expect_gap(request, pr, reason) applies the strict xfail only while the probe still reproduces the defect. Once the fix is in the tree the cell is a plain test. Wired: #120314 + #119970 (soak DST scenarios), #120315 (5 stream cells), #120377 (sent_ack_lost-fk_tg, whose reason names #120377, and zzz_unclean_restart), #120444 (after_reconnect, plus the whole-run audit's fk_tg.redeliver-after_reconnect exclusion), #120450 (cell 5 plain redelivery).

Each probe checked against every fix PR's head. Every probe reports open on this branch, and on each fix head only that PR's probe flips:

tree      120314 119970 120315 120377 120444 120450
this PR   open   open   open   open   open   open
#120314   FIXED  open   open   open   open   open
#119970   open   FIXED  open   open   open   open
#120315   open   open   FIXED  open   open   open
#120377   open   open   open   FIXED  open   open
#120444   open   open   open   open   FIXED  open
#120450   open   open   open   open   open   FIXED

Coin-flip xfails removed. long_split now streams one chunk. long_streamed (2.5k chunks, 1 s apart) and stream_timeout_first_send (42-char chunks, 0.5 s apart) pace their chunks so each lands in its own consumer tick (edit interval 0.05 s). That leaves 2 plain cells and 3 strict #120315 cells. sent_ack_lost waits for the answer to be persisted before the kill, so it deterministically pins #120377. The streamed-before-persisted order is its own strict live-gap cell, stream_accepted_unpersisted-fk_tg, which uses a stalled provider stream. --runxfail, 2 runs each at load 100–230: base failed the identical 7 cells both times; with #120315 all 12 stream cells passed both times; with #120377 only stream_accepted_unpersisted failed (it has no fix PR).

zzz_unclean_restart compares director.resumes against a snapshot taken before its own kill instead of requiring it empty, so resumes left by crash cells on their own homes no longer count. With #120377 applied it passed in -k "crash or zzz" both times.

Two-process fire-claim contention. In the soak, the tick lock admits one ticker per instant, so the second process never contended for a fire (a neutralised claim stayed green). The new test_two_replicas_contend_for_every_fire drives CronScheduler.fire_due ("exactly one of N replicas") from this process and a second OS process for 12 due occurrences of two jobs. A file barrier parks both inside the real claim_job_for_fire (it only delays), and the winner's run is held open. Each round asserts: 2 claim attempts, exactly 1 win, 1 run for that slot delivered once, the loser recorded as "not acquired", and the store re-armed past now. Sabotage: skipping _claim_is_live → red in round 1 ([True, True]); making the cross-process jobs flock a no-op → red in round 1, 3/3 runs.

Full files (scripts/run_tests.sh --include-integration, no retries):

tree cell 5 C12 messaging C13 soak
this PR 5✓ 1xf 47✓ 11xf 4✓ 2xf
+ #120315 only 52✓ 6xf
+ #120377 only 49✓ 9xf
+ all six fix PRs 6✓ 55✓ 3xf 6✓

The 3 xfails left with every fix merged are strict live gaps that have no fix PR (ack_lost on fk_tg/fk_dc, stream_accepted_unpersisted-fk_tg). There is 0 XPASS in any tree.

Base automatically changed from tests/core-e2e to main September 23, 2026 21:54
@teknium1
teknium1 requested a review from a team September 23, 2026 21:54
teknium1 added a commit that referenced this pull request Sep 23, 2026
…safe probes

A strict xfail(raises=AssertionError) also swallowed boot failures and timeouts, and
flips main red (XPASS) the moment its fix merges. Each gap now has a probe that
reproduces the defect's mechanism on the tree under test; the xfail applies only
while the probe reports the defect open, and only for the bug-specific exception
class the cell raises when it observes that leak (TenantLeak, LaunchProfileBleed,
RoutingLeak, PromptLeftPending). Same pattern as #120344's delivery suite.
Runs the real InProcessCronScheduler loop (due scan, pending slots, fire
claims + heartbeat, executions ledger, run_one_job, delivery routing,
mark_job_run, manual run / trigger_job, a second process on the same
HERMES_HOME, SIGKILL + boot recovery) under a file-backed virtual clock for
~30 virtual days per scenario across DST transitions and TZ configs (unset,
foreign process TZ, Asia/Shanghai +08:00, America/New_York), outages,
restarts mid-run, manual runs between ticks, runs longer than the fire-claim
TTL and two contending processes. Only run_job and the platform send are
faked.

Oracle, independent of cron/jobs.py: croniter over wall time in the job's
zone. After every tick the fired-occurrence set equals the model (no early,
late, duplicate or skipped fire beyond the catch-up policy), every stored
next_run_at equals the model and is in the future; at the end every
execution is terminal and truthful (delivered <=> the fake platform got
exactly that execution's message once). Red-proven against re-injected
#105690 manual-run re-stamp, the fire-claim heartbeat deadlock shape, UTC
next_run, skipped boot recovery, missing cross-process exclusion and the
DST fold due-compare bug (#120314; newyork_on_shanghai_fall carries a strict
xfail 'fixed by #120314' until it lands). #119969 is pinned as a strict xfail.
…-once

Replaces the wave-2 skip stub. Drives the real delivery ledger, the real
adapter record/finalize methods and the real GatewayRunner boot claim +
redeliver halves in child processes on a temp HERMES_HOME, SIGKILLing at
recorded-unsent, attempting-unsent, sent-unacked and inside a boot's own
redelivery, plus 3 concurrent rebooters over 12 rows. Only the transport is
fake (fsynced journal = ground truth). Invariants per obligation: <= 1
unmarked copy, >= 1 copy, terminal ledger row after the first clean boot,
later reboots claim/send nothing, integrity_check ok.

FIRE (strict xfail, not fixed): sweep_recoverable claims a 'pending' row
without moving it to 'attempting' and the boot redelivery never calls
mark_attempting, so a boot killed after the platform accepted a PLAIN
redelivery leaves the row pending and the next boot resends it unmarked -
two unmarked copies. Red-proven against dropped recovery marker, dropped
claim CAS, skipped mark_delivered and re-claiming delivered rows.
A child-process GatewayRunner (own HERMES_HOME, SIGKILL + restart) runs the
real agent against the scripted fake LLM provider and three instances of a
FakePlatformAdapter that implements the gateway/platforms/base.py contract
(4096/2000 limits, edits vs no edits, threads, streaming on/off). Only the
transport is fake; its fsynced op journal is ground truth. Faults per op:
timeout, ack lost, 429 retry_after, connection reset, too long, and park
points for SIGKILL before/after the platform applies an op or between
provider completion and the delivery ledger.

Oracle per inbound: exactly one complete unmarked visible reply (extra
copies only with the ledger's duplicate marker), visible text == persisted
assistant text, user row persisted once, model turn run once, no stray or
stale partial reply; plus a whole-run audit across chats. Scenarios: 13
faults x 3 platforms, /queue and interrupt follow-ups on a slow turn,
parallel threads, re-delivered inbound ids (during/after a turn, after a
runner reconnect), five crash points with restart catch-up, and an unclean
restart with nothing in flight.

Red-proven against the reverted stream fixes (#120315), a hand-revert
of c961e5b (#91653), retrying timed-out sends, boot redelivery without
the marker, double-sent finals, a dropped queued follow-up and disabled
inbound dedup. Five strict xfails pin four live gaps (streaming
ack-lost duplicate, dedup lost on reconnect, streaming crash after accept,
and the unclean-restart recency fallback that re-answers every session
active in the last 120 s). Eight cells stay red until #120315 lands (every
streamed reply over the limit, and the first-send timeout) and carry xfail 'fixed by #120315':
strict where the bug fires every run, non-strict where it depends on the
stream/edit tick. The Director counts a resume note merged into the original
tagged user message as a resume, not a second run of that inbound.
The e2e job already discovers tests/e2e/core/delivery/ (C12 messaging
exactly-once, C13 cron virtual-clock soak) through
`run_tests.sh --include-integration tests/e2e`; both need the 900 s
per-file budget under load (C12 150-590 s on a loaded 20-core box).
Cell 5 lives under tests/conformance/ and runs in the unit job.
…y cells

A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
The soak's second ticker process never contends for a fire: the tick lock
admits one ticker per instant and the winner advances next_run_at before it
releases, so neutralising claim_job_for_fire left the soak green.

test_two_replicas_contend_for_every_fire drives the path the claim actually
guards, CronScheduler.fire_due ("exactly one of N replicas runs a job"): this
process and a second OS process receive the same fire for every due
occurrence of two jobs. A file barrier parks each replica inside
claim_job_for_fire until both are there (it only delays), and the winner's
run is held open, so the loser always meets a live claim. Per round exactly
one replica claims, one run starts for that slot and is delivered once, the
loser's attempt is recorded as not acquired, and the store re-arms past now.
Skipping the live-claim check or the cross-process jobs lock turns it red
on the first round.

The soak's DST scenarios gate their xfails on the #120314 / #119970 probes.
… gaps with no fix PR

#120314, #120377, #120444 and #120450 are on main: their PROBES entries,
every expect_gap naming them and the gap_open(120444) audit branch go, so
those cells are plain tests again (two of those probes read source text,
which the suite must not do). Cell 5's README row no longer claims a
strict xfail.

The two static strict xfails with no probe and no fix PR
(STREAM_ACK_LOST_GAP, STREAM_CRASH_AFTER_ACCEPT_GAP) become run-time
xfails via _pending_fixes.known_failure: only the final assertions run
under it, after every wait (restart, catch-up, settle) has succeeded,
and only an AssertionError matching the gap's own signature XFAILs; any
other failure stays red and a fixed tree simply passes. The whole-run
audit now skips only tokens whose cell actually XFAILed this run.
…xfail scoped to its deviation

ev_long_run used to flag a missing claim refresh as stalled_heartbeat
and carry on, so neutralising heartbeat_fire_claim in _heartbeat_loop
left all five scenarios green. Every virtual 30 s of the 15-minute hold
now (a) waits for the run's heartbeat to restamp the claim and fails if
it never does, and (b) has a contender call the real claim_job_for_fire
and asserts it loses, also once the first stamp is older than the TTL.
With the heartbeat replaced by True all 5 scenarios are red on the
refresh assertion; on main they are green.

The #119970 cell no longer carries a whole-scenario strict xfail. While
its behavioural probe reproduces, only the oracle's fired-set and
next_run_at checks may deviate: each deviation is recorded, the model
follows the stored slot, and the soak runs every virtual day with all
other checks live, XFAILing at the end only if a deviation was seen
(and failing if the probe says open but none was). The #120314 entry
is gone (merged). A scenario that fails mid-hold now releases its held
run before teardown so it cannot bleed into the next scenario.
@teknium1
teknium1 merged commit 21999f0 into main Sep 23, 2026
35 checks passed
@teknium1
teknium1 deleted the tests/core-delivery branch September 23, 2026 23:21
teknium1 added a commit that referenced this pull request Sep 24, 2026
…safe probes

A strict xfail(raises=AssertionError) also swallowed boot failures and timeouts, and
flips main red (XPASS) the moment its fix merges. Each gap now has a probe that
reproduces the defect's mechanism on the tree under test; the xfail applies only
while the probe reports the defect open, and only for the bug-specific exception
class the cell raises when it observes that leak (TenantLeak, LaunchProfileBleed,
RoutingLeak, PromptLeftPending). Same pattern as #120344's delivery suite.
teknium1 added a commit that referenced this pull request Sep 24, 2026
…safe probes

A strict xfail(raises=AssertionError) also swallowed boot failures and timeouts, and
flips main red (XPASS) the moment its fix merges. Each gap now has a probe that
reproduces the defect's mechanism on the tree under test; the xfail applies only
while the probe reports the defect open, and only for the bug-specific exception
class the cell raises when it observes that leak (TenantLeak, LaunchProfileBleed,
RoutingLeak, PromptLeftPending). Same pattern as #120344's delivery suite.
teknium1 added a commit that referenced this pull request Sep 24, 2026
…safe probes

A strict xfail(raises=AssertionError) also swallowed boot failures and timeouts, and
flips main red (XPASS) the moment its fix merges. Each gap now has a probe that
reproduces the defect's mechanism on the tree under test; the xfail applies only
while the probe reports the defect open, and only for the bug-specific exception
class the cell raises when it observes that leak (TenantLeak, LaunchProfileBleed,
RoutingLeak, PromptLeftPending). Same pattern as #120344's delivery suite.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-reviewed applied to manually approve dangerous changes comp/cron Cron scheduler and job management comp/gateway Gateway runner, session dispatch, delivery P3 Low — cosmetic, nice to have sweeper:risk-automation Sweeper risk: may affect CI, automerge, label sync, or maintainer automation sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/test Test coverage or test infrastructure

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants