Repository navigation
fix(gateway): a platform replay after a reconnect is no longer answered twice - #120444
Merged
Merged
Conversation
…ed twice The reconnect watcher replaces a failed adapter with a NEW instance, and every adapter keeps its inbound MessageDeduplicator on the instance. The rebuilt adapter started with an empty cache, so a platform re-delivering a recent inbound ID right after the reconnect (websocket resume replay, webhook retry, unacked poll batch) got it processed and answered again. The reconnect queue entry now holds the retired adapter's MessageDeduplicator attributes by reference, and the rebuilt adapter absorbs their live IDs before it connects. The multiplex secondary-profile reconnect path gets the same handover. Any adapter using the shared helper is covered without per-adapter code.
૮ >ﻌ< ა ci reviewran on 04efcac — fix(gateway): a platform replay after a reconnect is no long debug infoCI timingsCI timings · View report · View jobWall time 7m42s vs 6m10s (+24.9%). 10 job(s) slower, 2 faster, 1 unchanged.
|
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…y cells A strict xfail on a gap whose fix is an open PR turns main red the moment that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces the defect's mechanism on the tree under test in a throwaway interpreter; expect_gap() applies the strict xfail only while the probe still reproduces it, so the cell becomes a plain test once the fix is in the tree, whatever the merge order. Covered: #120314, #119970 (soak), #120315, #120377, #120444 (C12) and #120450 (cell 5). Each probe was checked against every fix head: it flips on its own PR and on no other. C12 cells made deterministic (identical outcome on every run): - long_split streams the whole reply as one chunk; long_streamed and stream_timeout_first_send pace chunks so each lands in its own consumer tick. The five former coin-flip xfails are now two plain cells and three strict #120315 gap cells. - sent_ack_lost waits until the answer is persisted before the kill, so it pins the #120377 recovery; the streamed-before-persisted order is its own cell (stream_accepted_unpersisted, a strict live gap with a stalled provider stream). - zzz_unclean_restart compares director.resumes against a snapshot taken before its kill instead of requiring it empty: crash cells on their own homes may legitimately resume. - the whole-run audit skips the reconnect replay only while #120444's gap is open.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…y cells A strict xfail on a gap whose fix is an open PR turns main red the moment that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces the defect's mechanism on the tree under test in a throwaway interpreter; expect_gap() applies the strict xfail only while the probe still reproduces it, so the cell becomes a plain test once the fix is in the tree, whatever the merge order. Covered: #120314, #119970 (soak), #120315, #120377, #120444 (C12) and #120450 (cell 5). Each probe was checked against every fix head: it flips on its own PR and on no other. C12 cells made deterministic (identical outcome on every run): - long_split streams the whole reply as one chunk; long_streamed and stream_timeout_first_send pace chunks so each lands in its own consumer tick. The five former coin-flip xfails are now two plain cells and three strict #120315 gap cells. - sent_ack_lost waits until the answer is persisted before the kill, so it pins the #120377 recovery; the streamed-before-persisted order is its own cell (stream_accepted_unpersisted, a strict live gap with a stalled provider stream). - zzz_unclean_restart compares director.resumes against a snapshot taken before its kill instead of requiring it empty: crash cells on their own homes may legitimately resume. - the whole-run audit skips the reconnect replay only while #120444's gap is open.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…y cells A strict xfail on a gap whose fix is an open PR turns main red the moment that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces the defect's mechanism on the tree under test in a throwaway interpreter; expect_gap() applies the strict xfail only while the probe still reproduces it, so the cell becomes a plain test once the fix is in the tree, whatever the merge order. Covered: #120314, #119970 (soak), #120315, #120377, #120444 (C12) and #120450 (cell 5). Each probe was checked against every fix head: it flips on its own PR and on no other. C12 cells made deterministic (identical outcome on every run): - long_split streams the whole reply as one chunk; long_streamed and stream_timeout_first_send pace chunks so each lands in its own consumer tick. The five former coin-flip xfails are now two plain cells and three strict #120315 gap cells. - sent_ack_lost waits until the answer is persisted before the kill, so it pins the #120377 recovery; the streamed-before-persisted order is its own cell (stream_accepted_unpersisted, a strict live gap with a stalled provider stream). - zzz_unclean_restart compares director.resumes against a snapshot taken before its kill instead of requiring it empty: crash cells on their own homes may legitimately resume. - the whole-run audit skips the reconnect replay only while #120444's gap is open.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…y cells A strict xfail on a gap whose fix is an open PR turns main red the moment that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces the defect's mechanism on the tree under test in a throwaway interpreter; expect_gap() applies the strict xfail only while the probe still reproduces it, so the cell becomes a plain test once the fix is in the tree, whatever the merge order. Covered: #120314, #119970 (soak), #120315, #120377, #120444 (C12) and #120450 (cell 5). Each probe was checked against every fix head: it flips on its own PR and on no other. C12 cells made deterministic (identical outcome on every run): - long_split streams the whole reply as one chunk; long_streamed and stream_timeout_first_send pace chunks so each lands in its own consumer tick. The five former coin-flip xfails are now two plain cells and three strict #120315 gap cells. - sent_ack_lost waits until the answer is persisted before the kill, so it pins the #120377 recovery; the streamed-before-persisted order is its own cell (stream_accepted_unpersisted, a strict live gap with a stalled provider stream). - zzz_unclean_restart compares director.resumes against a snapshot taken before its kill instead of requiring it empty: crash cells on their own homes may legitimately resume. - the whole-run audit skips the reconnect replay only while #120444's gap is open.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…y cells A strict xfail on a gap whose fix is an open PR turns main red the moment that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces the defect's mechanism on the tree under test in a throwaway interpreter; expect_gap() applies the strict xfail only while the probe still reproduces it, so the cell becomes a plain test once the fix is in the tree, whatever the merge order. Covered: #120314, #119970 (soak), #120315, #120377, #120444 (C12) and #120450 (cell 5). Each probe was checked against every fix head: it flips on its own PR and on no other. C12 cells made deterministic (identical outcome on every run): - long_split streams the whole reply as one chunk; long_streamed and stream_timeout_first_send pace chunks so each lands in its own consumer tick. The five former coin-flip xfails are now two plain cells and three strict #120315 gap cells. - sent_ack_lost waits until the answer is persisted before the kill, so it pins the #120377 recovery; the streamed-before-persisted order is its own cell (stream_accepted_unpersisted, a strict live gap with a stalled provider stream). - zzz_unclean_restart compares director.resumes against a snapshot taken before its kill instead of requiring it empty: crash cells on their own homes may legitimately resume. - the whole-run audit skips the reconnect replay only while #120444's gap is open.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
… gaps with no fix PR #120314, #120377, #120444 and #120450 are on main: their PROBES entries, every expect_gap naming them and the gap_open(120444) audit branch go, so those cells are plain tests again (two of those probes read source text, which the suite must not do). Cell 5's README row no longer claims a strict xfail. The two static strict xfails with no probe and no fix PR (STREAM_ACK_LOST_GAP, STREAM_CRASH_AFTER_ACCEPT_GAP) become run-time xfails via _pending_fixes.known_failure: only the final assertions run under it, after every wait (restart, catch-up, settle) has succeeded, and only an AssertionError matching the gap's own signature XFAILs; any other failure stays red and a fixed tree simply passes. The whole-run audit now skips only tokens whose cell actually XFAILed this run.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
…y cells A strict xfail on a gap whose fix is an open PR turns main red the moment that fix merges (XPASS), and the non-strict ones guarded nothing. Each gap now has a probe (tests/e2e/core/delivery/_pending_fixes.py) that reproduces the defect's mechanism on the tree under test in a throwaway interpreter; expect_gap() applies the strict xfail only while the probe still reproduces it, so the cell becomes a plain test once the fix is in the tree, whatever the merge order. Covered: #120314, #119970 (soak), #120315, #120377, #120444 (C12) and #120450 (cell 5). Each probe was checked against every fix head: it flips on its own PR and on no other. C12 cells made deterministic (identical outcome on every run): - long_split streams the whole reply as one chunk; long_streamed and stream_timeout_first_send pace chunks so each lands in its own consumer tick. The five former coin-flip xfails are now two plain cells and three strict #120315 gap cells. - sent_ack_lost waits until the answer is persisted before the kill, so it pins the #120377 recovery; the streamed-before-persisted order is its own cell (stream_accepted_unpersisted, a strict live gap with a stalled provider stream). - zzz_unclean_restart compares director.resumes against a snapshot taken before its kill instead of requiring it empty: crash cells on their own homes may legitimately resume. - the whole-run audit skips the reconnect replay only while #120444's gap is open.
teknium1
added a commit
that referenced
this pull request
Sep 23, 2026
… gaps with no fix PR #120314, #120377, #120444 and #120450 are on main: their PROBES entries, every expect_gap naming them and the gap_open(120444) audit branch go, so those cells are plain tests again (two of those probes read source text, which the suite must not do). Cell 5's README row no longer claims a strict xfail. The two static strict xfails with no probe and no fix PR (STREAM_ACK_LOST_GAP, STREAM_CRASH_AFTER_ACCEPT_GAP) become run-time xfails via _pending_fixes.known_failure: only the final assertions run under it, after every wait (restart, catch-up, settle) has succeeded, and only an AssertionError matching the gap's own signature XFAILs; any other failure stays red and a fixed tree simply passes. The whole-run audit now skips only tokens whose cell actually XFAILed this run.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
When a platform drops and the gateway reconnects it, a message the platform sends again right after the reconnect is now dropped instead of being processed and answered a second time.
Root cause: the reconnect watcher replaces a failed adapter with a new instance, and inbound dedup (
MessageDeduplicator) lives on the instance, so the new adapter started with an empty cache.Changes
gateway/platforms/helpers.py:MessageDeduplicator.absorb()copies another cache's live IDs, keeping their original seen times.inbound_dedup_caches(adapter)collects an adapter'sMessageDeduplicatorattributes by reference.carry_inbound_dedup()seeds a rebuilt adapter from them.gateway/run_adapters.py: the primary reconnect queue entry stores the retired adapter's caches, and_reconnect_failed_platformseeds the new adapter before it connects. The multiplex secondary-profile reconnect (_schedule_secondary_profile_reconnect→_run_secondary_profile_reconnect→_secondary_reconnect_attempt) gets the same handover.website/docs/developer-guide/adding-platform-adapters.mdgets a new "Inbound Deduplication" pattern, andgateway/platforms/ADDING_A_PLATFORM.mdgets a checklist bullet saying the reconnect carries this state and a hand-rolled cache does not.Validation
Live repro: before / after on the real gateway, using the C12 exactly-once harness from #120344. A child-process
GatewayRunnerruns the real agent against the scripted fake LLM provider and a fake transport. The test injects inboundin-X, the turn gets answered, the adapter hits a retryable fatal, the real reconnect watcher installs a new adapter, and the samein-Xis injected again.03544a73test_redelivered_inbound_id_processed_once[after_reconnect](--runxfail)inboundop admitted →an extra model turn … replied in this chat: ['<<…after_reconnect-extra2>> unscripted extra turn …']test_platform_reconnect.py::…test_replayed_inbound_id_after_runner_reconnect_is_droppedassert False is True)test_multiplex_adapter_registry.py::…test_secondary_reconnect_keeps_inbound_dedupassert False is True)_handle_adapter_fatal_error/_handle_profile_adapter_fatal_error→ reconnect path. Each checks that the replacement adapter drops the replayed ID, and that a fresh ID still gets through as a control. No existing assertions changed; the multiplex fixture adapter gained a_dedupattribute.check_no_tmp_literals.py,check-windows-footguns.py --allandgit diff --checkall clean.scripts/run_tests.sh tests/gateway(host load ~380): 8533 passed, 12 failed. Re-running the failing files alone:test_compression_failure_session_sync,test_session_db_recovery,test_approve_deny_commands,test_api_server_active_work_drain,test_session_hygiene_turnhold_adoption.test_buzz_websocketpasses 20/20 on both trees with plain pytest, but occasionally times out underrun_tests.shon either tree.after_reconnectcell is a strict xfail (RECONNECT_DEDUP_GAP) and will XPASS once this lands. Drop that marker andfk_tg.redeliver-after_reconnectfromKNOWN_GAP_TOKENS.Not covered
whatsapp_cloud._seen_wamids, Slack's_processed_message_tsedit guard) still starts empty on a rebuilt adapter. They are documented as such.Related
Infographic