Skip to content

fix(gateway): a crash mid boot-redelivery no longer sends a second unmarked copy of the reply (salvage #70700) - #120450

Merged
teknium1 merged 4 commits into
mainfrom
fix/delivery-outbox-redelivery-claim
Sep 23, 2026
Merged

teknium1 merged 4 commits into
mainfrom
fix/delivery-outbox-redelivery-claim

Conversation

@teknium1

@teknium1 teknium1 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

A gateway killed while its boot sweep is redelivering a stored reply no longer makes the next boot send a second, unmarked copy of that reply: any copy after the first is labeled with the "♻️ Recovered reply" marker.

Salvages #70700 by @fangliquanflq (authorship kept on the fix commit), rebased onto the current ledger (gateway/run.py → gateway/run_startup.py split) and trimmed.

Changes

  • gateway/delivery_ledger.py::sweep_recoverable: the claim UPDATE moves every claimed row to 'attempting' in the same owner-stamp CAS (it used to do that only for flood rows). needs_marker still reads the pre-claim state, so a never-claimed pending row goes out plain exactly once. A pending row with attempts > 0, left by an older build that already claimed and maybe sent it, is also marked.
  • Sibling path in the same mechanism: a boot-claimed failed row used to stay failed and owned by the live process while its boot send was in flight. The runtime reconnect sweep (sweep_failed_for_runtime, which takes failed rows) could then claim it and send it again. The row is now attempting for the whole send, so the boot claim stays exclusive.
  • Dropped from fix(gateway): label redelivery after mid-recovery crash #70700: the extra mark_attempting call right before adapter.send in the redelivery loop. Both callers of _redeliver_claimed_obligations (the boot sweep and the runtime sweep) now hand over rows that are already attempting, so the call was a redundant write.
  • Updated the comment in _obligation_adapter and the messaging docs bullet on mid-send recovery (website/docs/user-guide/messaging/index.md).

Root cause

sweep_recoverable spent an attempt and re-stamped the owner of a pending row without moving it to attempting. A crash after the platform accepted the resend therefore still looked like "never sent" to the next boot.

Validation

Live repro: before: on origin/main @ 03544a7, the cell-5 real-process harness from #120344 (tests/conformance/persistence/test_cell5_delivery_outbox_exactly_once.py) was run with --runxfail. It uses a real state.db, the real GatewayRunner claim and redeliver halves, a journaling BasePlatformAdapter that fsyncs each accepted send, and a real SIGKILL inside boot 1's plain resend. It fails with 'cell5 reply 0' (pending): 2 UNMARKED copies reached the platform — a silent duplicate (roles=['boot-1', 'boot-2']). After: the same harness on the fix passes all 6 kill-point cases (6 passed in 216s), including concurrent reboots, where claims still sum to exactly one per obligation.

Scenario Before After
Boot killed inside the plain redelivery of a pending row, then reboot 2 plain copies (silent duplicate) 1 plain + 1 marked copy
pending row left by an older build with attempts=1 redelivered plain redelivered marked
Boot-claimed failed row, then adapter reconnect during the boot send runtime sweep re-claims it (second send) not claimable until the boot send settles
Never-claimed pending row plain plain (unchanged)

Tests (2 invariants in tests/gateway/test_delivery_ledger.py; each was red with the origin/main ledger and is green with the fix):

  • test_redelivery_after_an_earlier_boot_claim_is_marked[killed_inside_send|older_build_left_pending]: drives the real GatewayRunner._redeliver_pending_obligations. Boot 1's send is interrupted after acceptance (CancelledError, so no mark_delivered), and the next boot's copy must be RECOVERED_MARKER + reply.
  • test_boot_claimed_row_is_not_reclaimed_by_runtime_sweep_mid_send: after a boot claim, sweep_failed_for_runtime returns nothing.

No existing assertions changed. scripts/run_tests.sh tests/gateway/: 8665 passed and 12 failed at host load 100–220. None of the 12 are delivery-ledger tests. 6 of them fail the same way on a pristine origin/main worktree, and the other 6 pass when their files are re-run alone. The delivery, restart-resume and flood-invariant files pass 93/93. ruff, check_no_tmp_literals, check-windows-footguns --all, check_compat_pointers, git diff --check and audit_pr_attribution are all clean.

Sibling surfaces checked: the cron outbox (cron/delivery_queue.claim_next) already moves pending to delivering in its claim CAS. The turn-side producer (BasePlatformAdapter._record_delivery_obligation) already marks rows attempting before the send. The runtime sweep already claims into attempting.

Heads-up for #120344: its strict xfail on test_boot_killed_inside_plain_redelivery_never_double_delivers will XPASS once this lands, so that marker should be removed.

Gap: if a startup claim's adapter disappears between the claim and the send, the row now stays attempting until the next boot, which redelivers it marked. Before, a claimed failed row in that state could also be picked up by the runtime timer. The adapter set is fixed at claim time, so this window is very small.

Infographic

Exactly-once outbox

fangliquanflq and others added 2 commits September 23, 2026 09:06
sweep_recoverable claimed a dead-owner 'pending' row by re-stamping the
owner and spending an attempt but left state='pending', and the boot
redelivery path never marked it attempting before adapter.send. A boot
killed after the platform accepted that plain resend (before
mark_delivered) left the row 'pending', so the next boot resent it
UNMARKED: a silent duplicate reply.

The claim UPDATE now moves every claimed row to 'attempting' in the same
owner-stamp CAS; needs_marker still reads the pre-claim state, so a
never-claimed pending row is redelivered plainly once and any later copy
carries RECOVERED_MARKER. A pending row that an older build already
claimed (attempts > 0) is also marked, since it may have been sent.
Two invariants on the boot outbox sweep, both red on the previous ledger:
- a pending row whose boot redelivery was interrupted after the platform
  accepted it (or that an older build claimed) is redelivered with the
  recovered marker, never a second plain copy;
- a boot-claimed failed row is not re-claimable by the runtime reconnect
  sweep while the boot send is in flight (it used to stay 'failed' and
  owned by this process, so a reconnect could send it a second time).

Also corrects the startup-claim comment in _obligation_adapter and the
messaging docs bullet on mid-send recovery.
@github-actions

github-actions Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 209911e — fix(gateway): release a boot claim whose adapter vanished be

⚠️ Warnings

CI timings · View report · View job

Wall time 6m41s vs 4m59s (+34.1%). 6 job(s) slower, 7 faster,

  • Python tests / e2e: -61.0s
  • OS-specific tests / Windows-only tests: +50.0s
  • OS-specific tests / macOS-only tests: -38.0s
  • Detect affected areas: +36.0s
  • Python tests / Run tests: -33.0s

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/gateway Gateway runner, session dispatch, delivery sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Sep 23, 2026
@teknium1 teknium1 added the ci-reviewed applied to manually approve dangerous changes label Sep 23, 2026
…atch

The boot sweep now moves every claimed row to attempting. If the platform went
fatal between the claim and the send (restart notification, flood sleep),
_obligation_adapter skipped the row without releasing it, so it sat in
attempting owned by this live process: the reconnect sweep only takes failed
rows and the boot sweep skips live owners, so the reply waited for the next
restart (and then carried a false duplicate marker). Release any undispatched
claim, not only runtime ones.

Found by independent review of #120450.
teknium1 added a commit that referenced this pull request Sep 23, 2026
…y cells

A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
@teknium1
teknium1 merged commit 4e96650 into main Sep 23, 2026
34 checks passed
@teknium1
teknium1 deleted the fix/delivery-outbox-redelivery-claim branch September 23, 2026 17:46
teknium1 added a commit that referenced this pull request Sep 23, 2026
…y cells

A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
teknium1 added a commit that referenced this pull request Sep 23, 2026
…y cells

A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
teknium1 added a commit that referenced this pull request Sep 23, 2026
…y cells

A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
teknium1 added a commit that referenced this pull request Sep 23, 2026
…y cells

A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
teknium1 added a commit that referenced this pull request Sep 23, 2026
… gaps with no fix PR

#120314, #120377, #120444 and #120450 are on main: their PROBES entries,
every expect_gap naming them and the gap_open(120444) audit branch go, so
those cells are plain tests again (two of those probes read source text,
which the suite must not do). Cell 5's README row no longer claims a
strict xfail.

The two static strict xfails with no probe and no fix PR
(STREAM_ACK_LOST_GAP, STREAM_CRASH_AFTER_ACCEPT_GAP) become run-time
xfails via _pending_fixes.known_failure: only the final assertions run
under it, after every wait (restart, catch-up, settle) has succeeded,
and only an AssertionError matching the gap's own signature XFAILs; any
other failure stays red and a fixed tree simply passes. The whole-run
audit now skips only tokens whose cell actually XFAILed this run.
teknium1 added a commit that referenced this pull request Sep 23, 2026
…y cells

A strict xfail on a gap whose fix is an open PR turns main red the moment
that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap
now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces
the defect's mechanism on the tree under test in a throwaway interpreter;
expect_gap() applies the strict xfail only while the probe still reproduces
it, so the cell becomes a plain test once the fix is in the tree, whatever
the merge order. Covered: #120314, #119970 (soak), #120315, #120377,
#120444 (C12) and #120450 (cell 5). Each probe was checked against every
fix head: it flips on its own PR and on no other.

C12 cells made deterministic (identical outcome on every run):
- long_split streams the whole reply as one chunk; long_streamed and
  stream_timeout_first_send pace chunks so each lands in its own consumer
  tick. The five former coin-flip xfails are now two plain cells and three
  strict #120315 gap cells.
- sent_ack_lost waits until the answer is persisted before the kill, so it
  pins the #120377 recovery; the streamed-before-persisted order is its own
  cell (stream_accepted_unpersisted, a strict live gap with a stalled
  provider stream).
- zzz_unclean_restart compares director.resumes against a snapshot taken
  before its kill instead of requiring it empty: crash cells on their own
  homes may legitimately resume.
- the whole-run audit skips the reconnect replay only while #120444's gap
  is open.
teknium1 added a commit that referenced this pull request Sep 23, 2026
… gaps with no fix PR

#120314, #120377, #120444 and #120450 are on main: their PROBES entries,
every expect_gap naming them and the gap_open(120444) audit branch go, so
those cells are plain tests again (two of those probes read source text,
which the suite must not do). Cell 5's README row no longer claims a
strict xfail.

The two static strict xfails with no probe and no fix PR
(STREAM_ACK_LOST_GAP, STREAM_CRASH_AFTER_ACCEPT_GAP) become run-time
xfails via _pending_fixes.known_failure: only the final assertions run
under it, after every wait (restart, catch-up, settle) has succeeded,
and only an AssertionError matching the gap's own signature XFAILs; any
other failure stays red and a fixed tree simply passes. The whole-run
audit now skips only tokens whose cell actually XFAILed this run.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci-reviewed applied to manually approve dangerous changes comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants