Repository navigation
fix(gateway): keep in-flight final delivery single-writer - #88166
michaelversluis wants to merge 2 commits into
Conversation
…send decision
Incident (2026-08-05 15:15 ET, Telegram group, 4070-char final response
delivered twice):
15:15:10 MarkdownV2 edit failed, falling back to plain text: Flood
control exceeded. Retry in 37 seconds
15:15:10 Telegram flood control, waiting 37.0s
15:15:10 Telegram flood control on send (attempt 1/3), retrying in
37.0s: Flood control exceeded
15:15:15 gateway.platforms.base: Sending response (4070 chars) to
-1003725014629
15:15:15 Telegram flood control on send (attempt 1/3), retrying in
32.0s
gateway/run.py waited a flat 5 seconds (asyncio.wait_for(stream_task,
timeout=5.0)) for the stream consumer's finalize/fallback send before the
duplicate-send decision, then cancelled and moved on. Telegram's own
retry_after can mandate waits far longer than 5s, so the gateway routinely
gave up while the send was still in flight: final_response_sent /
final_content_delivered were still False (the attempt hadn't resolved
yet), suppression didn't trigger, and the gateway sent its own "final"
response. If the abandoned send still reached the platform after being
locally cancelled, the user got the answer twice.
Fix: GatewayStreamConsumer now tracks final_delivery_in_progress for the
duration of the got_done finalize/fallback sequence. gateway/run.py's new
_await_stream_task_before_final_decision waits a short base timeout first
(unchanged latency for the common case), and only when the consumer
reports an attempt still in flight extends the wait up to 90s before
giving up. Uses asyncio.wait (not wait_for) so checking the flag doesn't
race the task's own cancellation-triggered unwind. Applied at all three
call sites that raced this decision, including the proxy-relay path.
Regression test (tests/gateway/test_final_send_flood_race.py) drives the
real GatewayStreamConsumer and gateway helper directly, coordinating the
mocked flood-controlled send via asyncio.Event rather than sleep-timing
comparisons, so the race is deterministic in both directions. Verified
failing on unpatched main (2 duplicate messages, 3/3 runs) and passing
with the fix (1 message, 5/5 runs).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Reviewed the diff. Correct fix for a duplicate-send race with a genuinely subtle async detail handled properly: the old flat 5s Nit: the no-timeout branch means an adapter with an unbounded retry loop would hold the turn forever; the comment leans on "adapters own bounded retries" - worth one assert/comment at the adapter layer stating retries are bounded by policy. |
Summary
Prevent duplicate gateway replies when a streamed final delivery is still in flight under a platform
RetryAfterlonger than the gateway's own wait ceiling.This preserves and builds on the implementation from #79592 (original commit/authorship retained), then closes the remaining long-
RetryAftergap found in a live Telegram incident.Root cause
The stream consumer and the normal gateway final-send path can both become writers for the same final response:
RetryAfterand the adapter keeps that attempt alive;#79592 correctly introduced
final_delivery_in_progressand waits past the original five seconds, but capped the total at 90 seconds. A production case on v0.20.2 returned 202–212 second waits, reproducing the duplicate beyond that cap:Changes
final_delivery_in_progressstate and all three gateway call-site conversions.TDD evidence
The beyond-ceiling regression was run first against the #79592 base and failed with two visible messages:
After the single-writer change:
Additional checks:
A local full
tests/gateway/run was attempted but the minimal dev environment lacks optional platform extras (aiohttp, Matrix, etc.).uv sync --all-extrasis blocked on macOS by the existingpython-olm==3.2.16native build (cmake/gmakeplus a const-qualified C++ compile error). The affected production and regression surfaces above are green; upstream CI should exercise the complete platform matrix.Scope and attribution
RetryAfterproduction reproduction, RED test, single-writer completion, failure fallback, and non-final cancellation coverage: this PR.Related: #74248, #76494, #79592.
Nederlandse samenvatting: bij Telegram-wachttijden boven 90 seconden kon de bestaande fix alsnog twee eindberichten versturen. Deze versie houdt één actieve eindlevering als enige schrijver, maar laat na een echte mislukking nog steeds precies één normale fallback toe.