Skip to content

fix(gateway): flush the undelivered tail when the first send fails, not just an edit - #80990

Open
briandevans wants to merge 1 commit into
NousResearch:mainfrom
briandevans:fix/gateway-stream-first-send-tail-8124
Open

briandevans wants to merge 1 commit into
NousResearch:mainfrom
briandevans:fix/gateway-stream-first-send-tail-8124

Conversation

@briandevans

Copy link
Copy Markdown

This is a sibling follow-up to commit 1d1e1277e (issue #8124)

  • What 1d1e1277e covered: a failed edit mid-segment. When an edit was rejected (flood control) and a tool boundary arrived before the retry, the guard at the got_segment_break reset flushed _accumulated as a fresh message so the pre-boundary tail was not lost.
  • What it did NOT touch: a failed first send, where _message_id is never assigned; and the commentary reset, which has no guard at all.
  • What this adds: both sites go through one shared helper that drops the _message_id truthiness term while keeping the __no_edit__ exclusion, plus a regression that pins the sentinel so the fix cannot start double-sending.

What does this PR do?

_reset_segment_state sets self._message_id = None at every tool boundary, so the next text segment enters the first-send branch of _send_or_edit. When that send fails, the branch does:

else:
    # Initial send failed — disable streaming for this session
    self._edit_supported = False
    return False

_message_id is never assigned, so it stays None. _last_sent_text is never assigned either (that happens above the failure return), so _reset_segment_state's _delivered_segment_texts bookkeeping records nothing.

The #8124 recovery flush was gated on:

if (
    self._accumulated
    and not current_update_visible
    and self._message_id                      # <-- never true after a failed first send
    and self._message_id != "__no_edit__"
):
    await self._flush_segment_tail_on_edit_failure()
self._reset_segment_state(preserve_no_edit=True)

The self._message_id term is exactly the condition a failed first send can never satisfy, so the flush was skipped and the reset immediately wiped _accumulated — the only copy of the text. The commentary reset a few lines above (if commentary_text is not None: self._reset_segment_state()) had no guard at all and destroyed the same buffer.

Symptom: on a multi-tool turn where one send is rejected, the paragraph written between two tool calls ("Here's what I found, let me check X") never appears and is never re-sent. The final answer still lands, so the reply visibly jumps from tool bubble to tool bubble with the explanation missing. Every chat user on the default transport: "edit" path is exposed; no unusual config or exotic state is required.

The existing helper needs no change and is already safe with _message_id = None: _try_strip_cursor() early-returns on a falsy id, and _visible_prefix() returns "" when _last_sent_text is "", so the prefix trim is skipped and the full buffer is sent as a new message. Only the caller's guard was excluding it.

Related Issue

Relates to #8124 (closed; 1d1e1277e was its fix — this completes the same guard for the first-send path).

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)
  • ✨ New feature (non-breaking change that adds functionality)
  • 🔒 Security fix
  • 📝 Documentation update
  • ✅ Tests (adding or improving test coverage)
  • ♻️ Refactor (no behavior change)
  • 🎯 New skill (bundled or hub)

Changes Made

  • gateway/stream_consumer.py — new _flush_undelivered_tail_before_reset() next to _flush_segment_tail_on_edit_failure(). It keeps the _accumulated / current_update_visible / __no_edit__ conditions and drops only the _message_id truthiness term.
  • gateway/stream_consumer.py — the got_segment_break reset now calls the helper instead of carrying the inline guard; the surrounding #8124 comment is updated to say the flush covers a failed first send as well as a failed edit.
  • gateway/stream_consumer.py — the commentary reset calls the helper before _reset_segment_state().
  • tests/gateway/test_stream_consumer_first_send_tail.py — new, 3 tests (below).

Sibling-site sweep

grep -n "_reset_segment_state" gateway/stream_consumer.py → 5 hits. Every call site is accounted for:

site verdict
:285 a comment, not a call site — excluded
:552 the definition — excluded
:1122 (commentary, pre-send) COVERED
:1125 (commentary, post-send) excluded — runs after _send_commentary delivered; _accumulated is already empty
:1157 (segment break) COVERED

Plus the manual reset inside _suppress_silence_marker (~:2047), which is excluded on purpose: it is a deliberate retraction of the preview when the agent emits a bare [SILENT]/NO_REPLY marker, and it even deletes the previously-sent preview messages. Flushing there would resurrect text the agent chose not to send.

How to Test

pytest tests/gateway/test_stream_consumer_first_send_tail.py -v

Three tests, all asserting on delivered payloads rather than internal attributes. The fake adapter fails only the editable-preview send path (the one the consumer marks with expect_edits in _metadata_for_send) and accepts plain sends — mirroring a real platform split, since Telegram keeps editable previews on the legacy send path.

  1. test_failed_first_send_tail_survives_tool_boundary — deltas → tool boundary → deltas → done. Every preview send fails, so _message_id is still None at the boundary.
  2. test_failed_first_send_tail_survives_commentary_reset — same, driven through the commentary path so the unguarded reset is the one under test. Also asserts the prose is delivered before the commentary that interrupted it.
  3. test_no_edit_sentinel_does_not_double_send — adapter accepts the send but returns no message_id, driving _message_id = "__no_edit__". Asserts the segment is delivered exactly once, pinning the sentinel exclusion so this change cannot regress into double-sending.

Regression guard, both directions verified. Reverting only the production hunks (tests untouched) turns 1 and 2 red:

E   AssertionError: text buffered before the commentary reset was dropped: delivered=['Using the shell tool...']
E   assert 0 == 1

The prose is entirely absent from what reached the user. Restoring the production change makes all three pass. Test 3 passes in both states by design — it pins existing behavior rather than demonstrating the bug.

Adjacent suites, all green with the change applied:

tests/gateway/test_stream_consumer_first_send_tail.py
tests/gateway/test_stream_consumer.py
tests/gateway/test_stream_consumer_draft.py
tests/gateway/test_stream_consumer_fresh_final.py
tests/gateway/test_stream_consumer_silence.py
tests/gateway/test_stream_consumer_thread_routing.py
tests/gateway/test_stale_finalize_suppression.py
tests/gateway/test_telegram_final_delivery.py
tests/gateway/test_telegram_overflow_partial.py
tests/gateway/test_duplicate_reply_suppression.py
tests/gateway/test_code_fence_tracking.py
tests/gateway/test_escape_reasoning_fences.py
tests/gateway/test_fence_chunker.py
tests/gateway/relay/test_relay_adapter.py
tests/gateway/relay/test_relay_slack_dm_streaming.py
tests/gateway/test_relay_capability_surface.py
tests/gateway/test_slack.py

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits (fix(scope):, feat(scope):, etc.)
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix/feature (no unrelated commits)
  • I've run pytest tests/ -q and all tests pass — ran the focused + adjacent gateway suites listed above instead of the full tree
  • I've added tests for my changes (required for bug fixes, strongly encouraged for features)
  • I've tested on my platform: macOS 15 (Darwin 25.4), Python 3.11

Documentation & Housekeeping

  • I've updated relevant documentation (README, docs/, docstrings) — docstrings on the new helper and the updated #8124 comment block
  • N/A — no config keys added or changed
  • N/A — no architecture or workflow change
  • N/A — pure Python, no platform-specific code paths touched
  • N/A — no tool descriptions or schemas changed

Related / Positioning

#8116 (@chinadbo, "preserve accumulated text on chunk send failure") is the closest open PR by concern. Its hunk is @@ -343,13 in the _send_new_chunk chunk-splitting path, i.e. a different failure site; this PR covers the segment-reset path and both of its call sites, so the two are disjoint and this is the broader of the two. Its test file (tests/gateway/test_stream_consumer_text_loss.py) is deliberately not the path used here, so there is no filename collision on merge.

Other open PRs touching gateway/stream_consumer.py were hunk-checked and are disjoint from 1112-1160: #80823 (@@1078, @@2239), #79592 (@@261, @@362, @@812, @@1164), #80173 (@@971, @@1433), #10194 (leading-newline stripping), #71725 (gateway/run.py only).

…ot just an edit

`_reset_segment_state` clears `_message_id` at every tool boundary, so the
next text segment re-enters the first-send branch of `_send_or_edit`. That
branch sets `_edit_supported = False` and returns without ever assigning
`_message_id`, leaving it `None`.

The NousResearch#8124 recovery flush was gated on `self._message_id` being truthy, which
is exactly the condition a failed first send can never satisfy. The flush was
therefore skipped and `_reset_segment_state` wiped `_accumulated` -- the only
copy of prose the user never saw. The commentary reset had no guard at all
and dropped the same buffer.

Route both resets through one helper that keeps the `__no_edit__` exclusion
(where the reset deliberately preserves state for `_send_fallback_final`) but
drops the `_message_id` truthiness term. Symptom: on a multi-tool turn where
one send is rejected, the paragraph written between two tool calls never
appears and is never re-sent, so the reply jumps from tool bubble to tool
bubble with the explanation missing.
Copilot AI lite review requested due to automatic review settings August 7, 2026 11:44

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR closes a remaining content-loss edge case in the gateway streaming consumer by ensuring buffered assistant text is flushed before segment state is reset when the initial send fails (i.e., no _message_id is ever assigned). It complements the earlier #8124 fix that handled failed edits mid-segment by extending the same “flush-before-reset” protection to the failed first send path and to the commentary reset.

Changes:

  • Add a shared _flush_undelivered_tail_before_reset() helper that flushes _accumulated when it was never made visible, without requiring a truthy _message_id (while still respecting the __no_edit__ sentinel).
  • Route both the segment-break reset and the commentary reset through the helper to avoid wiping the only copy of undelivered text.
  • Add a focused regression test suite covering tool-boundary loss, commentary-reset loss, and guarding against __no_edit__ double-sends.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated no comments.

File Description
gateway/stream_consumer.py Introduces and applies a shared flush-before-reset helper to prevent silent loss of buffered text when the first send fails.
tests/gateway/test_stream_consumer_first_send_tail.py Adds regression coverage for failed-first-send tail preservation across tool boundaries and commentary resets, plus a __no_edit__ non-duplication guard.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/gateway Gateway runner, session dispatch, delivery area/streaming Streaming responses: gateway delivery, provider wire sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Aug 7, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/streaming Streaming responses: gateway delivery, provider wire comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants