Repository navigation
fix(gateway): await in-flight streamed final send before duplicate-resend decision - #79592
0xjarvisagent wants to merge 1 commit into
Conversation
…send decision
Incident (2026-08-05 15:15 ET, Telegram group, 4070-char final response
delivered twice):
15:15:10 MarkdownV2 edit failed, falling back to plain text: Flood
control exceeded. Retry in 37 seconds
15:15:10 Telegram flood control, waiting 37.0s
15:15:10 Telegram flood control on send (attempt 1/3), retrying in
37.0s: Flood control exceeded
15:15:15 gateway.platforms.base: Sending response (4070 chars) to
-1003725014629
15:15:15 Telegram flood control on send (attempt 1/3), retrying in
32.0s
gateway/run.py waited a flat 5 seconds (asyncio.wait_for(stream_task,
timeout=5.0)) for the stream consumer's finalize/fallback send before the
duplicate-send decision, then cancelled and moved on. Telegram's own
retry_after can mandate waits far longer than 5s, so the gateway routinely
gave up while the send was still in flight: final_response_sent /
final_content_delivered were still False (the attempt hadn't resolved
yet), suppression didn't trigger, and the gateway sent its own "final"
response. If the abandoned send still reached the platform after being
locally cancelled, the user got the answer twice.
Fix: GatewayStreamConsumer now tracks final_delivery_in_progress for the
duration of the got_done finalize/fallback sequence. gateway/run.py's new
_await_stream_task_before_final_decision waits a short base timeout first
(unchanged latency for the common case), and only when the consumer
reports an attempt still in flight extends the wait up to 90s before
giving up. Uses asyncio.wait (not wait_for) so checking the flag doesn't
race the task's own cancellation-triggered unwind. Applied at all three
call sites that raced this decision, including the proxy-relay path.
Regression test (tests/gateway/test_final_send_flood_race.py) drives the
real GatewayStreamConsumer and gateway helper directly, coordinating the
mocked flood-controlled send via asyncio.Event rather than sleep-timing
comparisons, so the race is deterministic in both directions. Verified
failing on unpatched main (2 duplicate messages, 3/3 runs) and passing
with the fix (1 message, 5/5 runs).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
|
I reproduced a production case on Hermes Agent v0.20.2 that outlives this PR's fixed 90-second maximum, so the current patch still duplicates the final response for the reported class. Sanitized Telegram gateway timeline (2026-08-17, one inbound, one Codex call): Telegram showed two separate messages: the first was the partial/live delivery, the second the full normal final response. The screenshot and ledger confirm separate sends rather than one edited message. I also tested this PR's own regression harness at head adapter, already_confirmed = await _run_flood_race(
monkeypatch,
base_timeout=0.05,
max_timeout=0.10,
release_delay=0.15,
)
assert len(adapter.visible_full_response_messages()) == 1
assert already_confirmed is TrueCanonical run: This is the same deterministic mechanism as the PR's existing test; the only changed variable is that delivery resolves after the total wait ceiling. In production the observed RetryAfter was 202–212 seconds, while the patch stops waiting after 90 seconds, cancels the consumer, and lets the outer path send again. Requested adjustment before merge:
The PR's in-flight state tracking and |
|
I opened #88166 as a strict superset after reproducing a production |
Incident
2026-08-05 15:15 ET, Telegram group — a 4070-char final response was delivered twice.
Mechanism
gateway/run.pywaited a flat 5 seconds (asyncio.wait_for(stream_task, timeout=5.0)) for the stream consumer's finalize/fallback send before deciding whether to send its own "final" response, then cancelled the consumer's task on timeout and moved on. Telegram's ownretry_aftercan mandate waits far longer than 5s (37s here). So the gateway routinely gave up while the send was still in flight:final_response_sent/final_content_deliveredwere stillFalse(the attempt hadn't resolved yet), suppression never triggered, and the gateway sent its own "final" response. If the abandoned send still reached the platform after being locally cancelled, the user got the answer twice.Fix
GatewayStreamConsumer(gateway/stream_consumer.py) now tracksfinal_delivery_in_progressfor the duration of thegot_donefinalize/fallback sequence (cleared in the outerfinally, so every exit path — normal return, cancellation, exception — resets it).gateway/run.py's new_await_stream_task_before_final_decisionwaits a short base timeout first (unchanged latency for the common, non-flood case), and only when the consumer reports an attempt still in flight extends the wait up to 90s before giving up. It usesasyncio.waitrather thanasyncio.wait_for—wait_forcancels the awaited task on timeout and waits for it to fully unwind before raisingTimeoutError, so by the time a caller's except clause runs, the consumer's own cancellation handling has already reset the in-flight flag; checking it there would always seeFalse.asyncio.waitjust reports done/pending on timeout without touching the task, so the flag reflects genuinely live state._run_agent_via_proxy).Sibling flood paths checked
gateway/stream_consumer.py's_fallback_flood_retry_delayusers (_send_fallback_final,_send_empty_fallback_final) are unaffected — they already bound their own retry sleep to_max_fallback_flood_retry_seconds(5s) and bail out early on long waits, deliberately leaving final delivery to the gateway. That's the correct pattern; this fix makes the gateway side honor it properly.gateway/run.py's progress-edit flood path (~line 4085-4106, insidesend_progress_messages) is unaffected — it's cancelled harmlessly in thefinallyblock alongsideprogress_task; it only ever sends non-critical progress/status messages, never the final answer, so there's no duplicate-content risk there.Tests
New:
tests/gateway/test_final_send_flood_race.py::test_flood_controlled_finalize_delivers_final_response_onceDrives the real
GatewayStreamConsumerand the real gateway helper directly. The mocked flood-controlledsend()blocks on anasyncio.Event(send_release) that only the test controls, and signalssend_startedthe instant it begins blocking — this replaced an earlier sleep-duration-vs-scaled-timeout version that was flaky under scheduling jitter. The test awaitssend_started(no timing guess), lets the gateway's base timeout genuinely elapse via a one-sided sleep (nothing races it, since the send cannot complete on its own), then releases the send — resolution after that is driven by the event loop waking the extended wait on task completion, not more sleeping.Before fix (
git stashthe fix, keep the test) — 3/3 runs, always fails with the duplicate:After fix — 5/5 runs, passes:
Full suite run (
/home/deploy/.hermes/hermes-agent/venv/bin/python -m pytest):Also ran
tests/gateway/test_stream_consumer.py(66 passed) to confirm no regression in the consumer's broader finalize/fallback logic.🤖 Generated with Claude Code