Skip to content

fix(grpc): stop the open Messages content block before another starts - #2715

Merged
slin1237 merged 5 commits into
smg-project:mainfrom
yechank-nvidia:fix/messages-content-block-overlap
Oct 1, 2026
Merged

slin1237 merged 5 commits into
smg-project:mainfrom
yechank-nvidia:fix/messages-content-block-overlap

Conversation

@yechank-nvidia

@yechank-nvidia yechank-nvidia commented Sep 30, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Problem

In short. In Messages streaming a response is a sequence of content blocks. Each block is sent as content_block_start, its deltas, then content_block_stop, and the next block starts only after the previous one stops; clients build message.content from the block indexes. The gRPC path could start the next block without stopping the open one, sometimes at the same index (the specific-tool case below):

expected: start 0 thinking, delta 0, stop 0, start 1 tool_use, delta 1, stop 1
main:     start 0 thinking, delta 0, start 0 tool_use, stop 0, delta 1, stop 1

Most streams still came out right: a response with a single block, or reasoning followed by text in a later chunk, never overlaps, and a client that only appends deltas by index rebuilds the content anyway. Where blocks overlap or an index is reused, deltas can reach the wrong block or a block that was never started: depending on the client the stream fails or a tool call arrives without its arguments, and a client that tracks a single open block sees blocks that never stop.

gRPC Messages streaming (process_messages_streaming_chunks) keeps one open flag per block kind (thinking_block_open, text_block_open, tool_block_open) and one current_block_index. Before this change:

  • only the tool_use start sites stopped an open block first, and only a text or tool_use block;
  • before the end of the stream, a thinking block was stopped only when a later chunk carried normal text outside reasoning;
  • a new thinking or text block stopped nothing.

A block could then start at the index of a block that was still open. The deltas and stops that followed went to an index that was never started. Three cases reach this with the parsers on main.

Text after a tool call. qwen_xml returns what follows </tool_call>, even a lone newline, as normal text. With one character per chunk, <tool_call>\n<function=lookup>\n<parameter=q>\n1\n</parameter>\n</function>\n</tool_call>\nDone. gives:

start 0 tool_use, delta 0 input_json_delta, delta 0 input_json_delta, start 0 text, delta 0 text_delta, ..., stop 0, stop 1

A specific tool right after reasoning. Parsers deepseek_r1 + deepseek (or qwen3 + qwen when the output opens with <think>), tool_choice: {"type": "tool", "name": "lookup"}, chunks ["plan", "</think>", "{}"]:

start 0 thinking, delta 0 thinking_delta, start 0 tool_use, stop 0, delta 1 input_json_delta, stop 1

The </think> chunk carries no text, so the thinking block is still open when the specific-tool path starts the tool_use block. This needs a backend that lets the model reason before the tool's JSON schema applies.

Reasoning again after text or a tool call. deepseek_v41 enters reasoning again on <think> in content. inkling can send a thinking message after a text or tool message. With deepseek_v41 and chunks ["<think>plan", "</think>", "answer", "<think>", "more", "</think>", "done"]:

start 0 thinking, ..., stop 0, start 1 text, delta 1 text_delta, start 1 thinking, delta 1 thinking_delta, stop 1, delta 2 text_delta, stop 2

Solution

Before every content_block_start, stop whichever block is open and move to the next index. A new helper, stop_open_block, does this and replaces the existing text/tool_use stops at the three tool_use start sites.

Two changes keep tool arguments in their block now that a new block stops the open one:

  • With multi-token chunks, a tool parser can return the arguments that finish the open call together with the text after the call. qwen_xml does this when one chunk holds the end of a parameter, </tool_call> and the text after it. The router emitted the text first. With the stops above alone, the text block stopped the tool_use block, the arguments went to the text block, and clients built the tool input without them (on main, the text block started on top of the tool_use block instead). The calls before the first one with a name continue the open call, and now go to the open tool_use block before the text. The text in such a result follows the end of the call, so this keeps the model's order.
  • Tool arguments are sent only while a tool_use block is open, and are dropped otherwise. A thinking block can stop the tool_use block in the middle of a call, and the arguments after it have no block to go to. See Out of scope for when this happens.

A stream that never started a block while another was open emits the same events as before.

Changes

  • model_gateway/src/routers/grpc/regular/streaming.rs
    • Add StreamingProcessor::stop_open_block.
    • Call it before all seven content_block_start sites in process_messages_streaming_chunks: thinking, text, text from the tool parser, the leftover text at end of stream, and the tool_use starts of the specific-tool, incremental and end-of-stream paths.
    • Incremental tool path: send the calls that continue the open call before the chunk's text.
    • Add StreamingProcessor::send_tool_arguments, which sends a call's arguments only while a tool_use block is open, and drops them with a debug log otherwise. The specific-tool, incremental and end-of-stream paths send arguments through it.
    • The prefill-decode path goes through the same function.
  • model_gateway/src/routers/grpc/regular/streaming/eof_tests.rs
    • messages_blocks_and_inputs: runs a scripted gRPC stream through process_messages_streaming_chunks. It returns the block starts and stops, and the input of each tool_use block joined from its input_json_deltas as a client SDK builds it (the joined text if it is not complete JSON). It asserts that a block starts only when none is open, and that every delta goes to the open block and matches its type (text_delta to text, input_json_delta to tool_use, thinking_delta to thinking). messages_blocks returns only the blocks.
    • messages_blocks_do_not_overlap_when_reasoning_calls_and_text_alternate: stub parsers for three orders:
      • reasoning, then a call, text and reasoning again;
      • text returned while the parser stays in reasoning;
      • reasoning right after a call.
    • messages_blocks_do_not_overlap_with_deepseek_parsers: the registered parsers from the cases above (deepseek_r1 + deepseek with a specific tool, and deepseek_v41), and json + deepseek_v41, where the tool parser releases held text at the end of the stream after reasoning started again.
    • messages_tool_arguments_precede_text_in_the_same_chunk: qwen_xml with multi-token chunks: two parallel calls, arguments followed by Done., and a second parameter followed by the close and a newline. It checks the tool inputs.
    • messages_tool_arguments_need_an_open_block: qwen + qwen3 with a specific tool and <think> split across chunks, and qwen_xml + qwen3 with <think> inside an argument, both in the middle of the stream and at its end (see Out of scope).

The message of the first commit says that a thinking block was stopped only "when reasoning ended in a chunk that also carried text". The accurate statement is the second bullet under Problem: before the end of the stream, a thinking block was stopped only when a later chunk carried normal text outside reasoning. The commit is already pushed, so it is left as is.

Out of scope

A specific tool before reasoning. With a specific tool_choice, the tool_use block starts in the first chunk that is not in reasoning, even when that chunk has no text. If reasoning starts only after that, as when the first chunk holds a partial <think>, the thinking start now stops the tool_use block, and the arguments after reasoning have no open block. They are dropped. For example, qwen3 + qwen with chunks ["<thi", "nk>plan", "</think>", "{}"] gives:

start 0 tool_use, stop 0, start 1 thinking, delta 1 thinking_delta, stop 1

The client gets a tool_use block with empty input, as on main. There, the thinking block starts on top of the tool_use block (start 0 tool_use, start 0 thinking, ...), the arguments go to index 1, and the Python and TypeScript SDKs drop them. With the block stops alone, the arguments went to index 2, which was never started, and the Python SDK's stream accumulator raised IndexError on that delta. Fixing it means deciding when the specific-tool block should start, so it is left for a follow-up.

<think> inside tool arguments. Until it strips a <think>, a reasoning parser such as qwen3 takes the first <think> anywhere in the output as the start of reasoning. When the output does not open with <think>, as when thinking is enabled for a model that answers without it, a <think> inside a tool argument starts reasoning in the middle of the call. The thinking block now stops the tool_use block, and the rest of the call's arguments are dropped. qwen3 + qwen_xml with chunks ["<tool_call>\n<function=lookup>\n<parameter=q>\nA ", "<think>x</think> B\n</parameter>\n", "</function>\n</tool_call>"] gives:

start 0 tool_use, stop 0, start 1 thinking, delta 1 thinking_delta, stop 1

The client gets a tool_use block with empty input, as on main. There, the thinking block starts on top of the tool_use block (start 0 tool_use, start 0 thinking, ...) and the arguments go to index 1. With the block stops alone, they went to index 2, which was never started, and the Python SDK raised IndexError. When reasoning lasts to the end of the stream, what is dropped is the closing brace that qwen_xml releases there, which with the block stops alone went to the open thinking block. The Python SDK parses the rest as partial JSON and builds the same input as on main. Parsers that hold back the end of the last argument until the end of the stream, such as qwen, json, mistral and llama, lose that part too. On main it reached the client, because the thinking block had started at the index of the tool_use block. Keeping these arguments would mean changing how the reasoning parser reads <think>, which this PR does not touch.

Reasoning and text switching inside one chunk. The reasoning parser returns a chunk's reasoning and normal text as two strings, and the router emits the reasoning first. So when one chunk holds </think>answer<think>more, the blocks no longer overlap, but they come out in the wrong order. deepseek_v41 with chunks ["<think>plan", "</think>answer<think>more", "rest</think>done"] gives:

start 0 thinking, delta 0 "plan", delta 0 "more", stop 0, start 1 text, delta 1 "answer", stop 1, start 2 thinking, delta 2 "rest", stop 2, start 3 text, delta 3 "done", stop 3

more joins the first thinking block, and one reasoning span is split over two blocks. Normal text in a chunk that ends in reasoning also skips the tool parser and follows the chunk's reasoning. This is existing behavior, and the Chat Completions path does the same.

Merging with #2538. #2538 adds ReasoningParser::prompt_reasoning without a default, and the field MessagesResponseSpec::starts_in_reasoning. Whichever of the two lands second needs fn prompt_reasoning(&self, _: &str) -> PromptReasoning { PromptReasoning::Absent } in the ReasoningButText test stub and starts_in_reasoning: false in the MessagesResponseSpec literal of messages_blocks_and_inputs.

Test Plan

Negative control. In scratch copies, I ran the new tests with the streaming.rs of main (61b250f7) and of the first commit (46448999). I also ran each case as its own test.

Test main 46448999 This PR
messages_blocks_do_not_overlap_when_reasoning_calls_and_text_alternate fail pass pass
messages_blocks_do_not_overlap_with_deepseek_parsers fail pass pass
messages_tool_arguments_precede_text_in_the_same_chunk fail fail pass
messages_tool_arguments_need_an_open_block fail fail pass

The cases added in this revision:

Case Parsers Chunks main 46448999
parallel calls qwen_xml <tool_call>\n<function=lookup>\n<parameter=q>\n1, \n</parameter>\n</function>\n</tool_call>\n<tool_call>\n<function=lookup>\n, <parameter=q>\n2\n</parameter>\n</function>\n</tool_call> start 0 tool_use, start 0 text the first call's arguments go to text block 1, and its input is {} instead of {"q": 1}
arguments, then text qwen_xml <tool_call>\n<function=lookup>\n<parameter=q>\nPar, is\n</parameter>\n</function>\n</tool_call>\nDone. start 0 tool_use, start 0 text input {} instead of {"q": "Paris"}
second parameter, close and newline qwen_xml <tool_call>\n<function=lookup>\n<parameter=a>\nx\n</parameter>\n, <parameter=b>\ny\n</parameter>\n</function>\n</tool_call>\n start 0 tool_use, start 0 text input {"a": "x"} instead of {"a": "x", "b": "y"}
specific tool, split <think> qwen3 + qwen <thi, nk>plan, </think>, {} start 0 tool_use, start 0 thinking input_json_delta at index 2, which was never started
<think> inside an argument qwen3 + qwen_xml <tool_call>\n<function=lookup>\n<parameter=q>\nA , <think>x</think> B\n</parameter>\n, </function>\n</tool_call> start 0 tool_use, start 0 thinking input_json_delta at index 2, which was never started
<think> inside an argument, then the end of the stream qwen3 + qwen_xml <tool_call>\n<function=lookup>\n<parameter=q>\nA\n</parameter>\n, <parameter=r>\nB , <think>x start 0 tool_use, start 0 thinking input_json_delta to the open thinking block
reasoning right after a call stubs a, b start 0 thinking, start 0 tool_use pass
text released at end of stream json + deepseek_v41 <think>plan, </think>, {, <think>, more start 0 thinking, stop 0, start 1 thinking, start 1 text pass
text after a call, one character per chunk (scratch only) qwen_xml the first example under Problem start 0 tool_use, start 0 text pass

The other cases of the first revision (the stub specific-tool case was replaced) still fail on main and pass with this PR.

Mutation check. In a scratch copy, I disabled one change at a time and ran eof_tests (test names shortened):

Disabled Failing tests
stop_open_block before the thinking start ..._alternate, ..._deepseek_parsers, ..._need_an_open_block
the same call, stopping only a thinking or text block ..._alternate, ..._need_an_open_block
before the specific-tool tool_use start ..._deepseek_parsers
before the text start in the tool path ..._alternate, ..._precede_text_in_the_same_chunk
before the incremental tool_use start ..._alternate, ..._precede_text_in_the_same_chunk
before the text start without tools ..._alternate
before the leftover text at end of stream ..._deepseek_parsers
before the end-of-stream tool_use start none
arguments of the open call before the text ..._precede_text_in_the_same_chunk
the open-block check in send_tool_arguments ..._need_an_open_block
the same check for the specific-tool, incremental or end-of-stream arguments alone ..._need_an_open_block (each)

Only a parser whose get_unstreamed_tool_args returns a named call reaches the end-of-stream tool_use start while another block is open. Every registered parser returns unnamed items there, so no test covers that call.

Commands. All ran offline on CPU.

Check Result
cargo +nightly fmt --all -- --check pass
cargo test -p smg --lib -- eof_tests 11 passed (7 existing, 4 for this PR)
cargo test -p smg --lib 2,034 passed
cargo test -p smg --test messages_streaming_test / messages_test / grpc_responses_stream_contract_test / api_tests / spec_test / grpc_context_length_test / grpc_pd_fanout_test 21 / 11 / 29 / 109 / 98 / 17 / 13 passed
cargo clippy --workspace --all-targets -- -D warnings; -p smg --all-targets with default features and with grpc-server,jemalloc-profiling,test-util; CI's two --no-default-features variants pass
pre-commit run --all-files with CI's SKIP list pass

The offline crate cache did not have the three dependency bumps on main: lru 0.18.5, cc 1.5.1 and the opentelemetry-proto 0.33 dev-dependency. So the build copy used the Cargo.lock from before those bumps and the dev-dependency at 0.32. No source file differed.

Checklist
  • cargo +nightly fmt passes
  • cargo clippy --all-targets -- -D warnings passes for the workspace, -p smg and CI's two --no-default-features variants. --all-features needs OpenCV and was not run offline.
  • (Optional) Documentation updated: not needed, no user-facing option or API changed
  • (Optional) Please join us on Slack #sig-smg to discuss, review, and merge PRs

Messages streaming stopped an open text or tool_use block before a new
tool_use block, and stopped an open thinking block only when reasoning
ended in a chunk that also carried text. A new thinking or text block
stopped nothing. The new block then started at the index of the block
that was still open, and the deltas and stops that followed went to an
index that was never started.

With the existing parsers this happens when a specific tool_choice
follows reasoning whose last chunk has no text after `</think>`
(deepseek_r1 with the deepseek tool parser, or qwen3 with qwen when the
output opens with `<think>`), and when a reasoning parser enters
reasoning again after text or a tool call (deepseek_v41, inkling), also
within one chunk.

Stop whichever block is open, and move to the next index, before every
content_block_start.

Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>
@coderabbitai

coderabbitai Bot commented Sep 30, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Advanced

Run ID: 0a1c1bcd-4382-4577-a1ef-22e505018142

📥 Commits

Reviewing files that changed from the base of the PR and between 4644899 and 00e8603.

📒 Files selected for processing (2)
  • model_gateway/src/routers/grpc/regular/streaming.rs
  • model_gateway/src/routers/grpc/regular/streaming/eof_tests.rs

Included review availability: This review used your included allowance. Your plan provides up to 4 included reviews per hour; 3 remain after this review.


📝 Summary

Summary by CodeRabbit

  • Bug Fixes
    • Improved streaming transitions between reasoning, text, and tool-call content so blocks close correctly and do not overlap, including at the end of a response.
    • Fixed ordering when tool arguments and text arrive together, ensuring arguments are sent before subsequent text.
    • Tool arguments received without an open tool-call block are no longer emitted.

Walkthrough

Messages streaming now closes an open reasoning, text, or tool block before starting another. Tool arguments are emitted only when a tool block is open. Tests cover block transitions, argument ordering, and parser-flush output.

Changes

Messages content block boundaries

Layer / File(s) Summary
Manage block transitions
model_gateway/src/routers/grpc/regular/streaming.rs
Shared helpers close open content blocks and guard tool argument emission. Streaming transitions and end-of-stream recovery use these helpers.
Test block transitions
model_gateway/src/routers/grpc/regular/streaming/eof_tests.rs
Parser fixtures and stream tests cover non-overlapping content blocks, tool argument ordering, and arguments received without an open tool block.

Priority: ➖ Normal

Estimated code review effort: 3 (Moderate) | ~20 minutes

Change: Bug fix

Merge Risk: ⚪ Minimal · up to 00e86

The change closes open Messages blocks before transitions and keeps tool arguments within active tool blocks. No merge-blocking issue is established; merge after normal checks pass.

Security Architecture Review

Security architecture risk: 🔵 Low · up to 00e86

The change improves stream ordering without adding tool-execution authority or shared request state. Interrupted tool calls can still finish with empty or incomplete arguments; whether downstream applications reject those calls remains unverified.

Retained concerns
No architecture-level concerns identified.

Security review details

Security Blast Radius

  • inferred — The directly affected boundary is the client-facing Messages response produced by this path, including prefill/decode responses. The inspected code establishes no additional execution authority or cross-request state exposure; downstream tool privileges cannot be determined from this producer.

Trust Boundaries and Controls

  • observed — Parser-produced arguments pass through an active-tool-block check before becoming client-visible input deltas. This control prevents emission into a non-tool block, but it does not validate complete JSON, required parameters, or authorization for downstream execution.

Resilience and Maintainability Implications

  • observed — The response stage retains its existing load-guard and reservation attachment for disconnect/error cleanup. Normal reservation settlement remains after successful terminal event emission.

Hardening Proposals

  • proposed — Define interrupted tool calls as non-executable outcomes, or require downstream consumers to validate complete, schema-conforming arguments before execution. Block closure and a ToolUse stop reason should not alone establish executable completeness.
🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 56.52% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 23 functions across 2 files. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely describes the main change: stopping an open Messages content block before starting another.
Description check ✅ Passed The description is directly related to the changeset. It explains the streaming bug, the implementation, out-of-scope cases, and test results.
  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
🛠️ Fix failing CI checks 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Create a new PR

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added grpc gRPC client and router changes model-gateway Model gateway crate changes labels Sep 30, 2026
@yechank-nvidia yechank-nvidia self-assigned this Sep 30, 2026
…reaming

With multi-token chunks, a tool parser can return the arguments that
finish the open call together with the text after the call. qwen_xml
does this for a chunk like `1\n</parameter>\n</function>\n</tool_call>\n`.
Messages streaming emitted the text first. Since the previous commit,
starting that text block stops the open tool_use block, so the
arguments went to the text block and clients built the tool input
without them.

Send the calls before the first one with a name, which continue the
open call, to the tool_use block before the text. The test helper now
also checks that each delta matches the type of its block, and returns
each tool input as a client SDK builds it.

Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>
With a specific tool_choice, the tool_use block starts at the first
chunk outside reasoning. If reasoning starts only after that, as when
`<think>` is split across chunks, the thinking block stops the tool_use
block, and the arguments that follow went to an index that was never
started. The official Python SDK raises IndexError on such a delta.

Send the arguments only while the tool_use block is open, and log and
drop them otherwise. The client gets the tool_use block with empty
input, as it did before the block stops were added.

Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>
Cover a thinking block that follows a tool_use block, and text that the
tool parser releases at the end of the stream after reasoning started
again. Replace the stub specific-tool case, which the deepseek_r1 case
already covers, and correct two test comments.

Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>
Until it strips a `<think>`, a reasoning parser such as qwen3 takes the
first `<think>` anywhere in the output as the start of reasoning. When
the output does not open with `<think>`, one inside a tool argument
starts a thinking block in the middle of the call, and that block stops
the tool_use block. The rest of the call's arguments, which the tool
parser returns without a name, went to an index that was never started,
or at the end of the stream to the open thinking block. The official
Python SDK raises IndexError on the former.

Send tool arguments through one helper that drops them when no tool_use
block is open, as the specific tool path already did. The client gets
the tool_use block without those arguments. On main the thinking block
started on top of the tool_use block, so the part of the last argument
that JSON-style parsers release at the end of the stream still reached
the client there; it is dropped now. Log the drop at debug level, since
it repeats for every chunk of the call. The open-block test now covers
both cases.

Signed-off-by: yechank <161688079+yechank-nvidia@users.noreply.github.com>
@slin1237
slin1237 merged commit 5e5e207 into smg-project:main Oct 1, 2026
100 of 110 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

grpc gRPC client and router changes model-gateway Model gateway crate changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants