Skip to content

fix(server): emit terminal finish_reason chunk when parsers swallow the finished output - #673

Merged
waybarrios merged 3 commits into
waybarrios:mainfrom
funkymonkeymonk:fix/stream-terminal-finish-reason
Aug 15, 2026
Merged

waybarrios merged 3 commits into
waybarrios:mainfrom
funkymonkeymonk:fix/stream-terminal-finish-reason

Conversation

@funkymonkeymonk

@funkymonkeymonk funkymonkeymonk commented Aug 3, 2026 •

Copy link
Copy Markdown
Contributor

Refs #672

Problem

A streamed tool call can end with the tool_calls chunk (finish_reason: null) followed directly by data: [DONE] — no chunk ever carries finish_reason, and strict OpenAI clients (e.g. the pi coding agent) abort with Stream ended without finish_reason and retry the turn.

Root cause: in stream_chat_completion, the only code attaching finish_reason to a chunk lives in the per-output loop. When the engine's finished=True output is consumed by a parser continue — e.g. gemma4 emits the complete canonical call (<|tool_call>call:fn{...}<tool_call|>) in one delta, then a bare <turn|> end-of-turn token as the terminal delta, which the tool parser suppresses with continue (nothing new to emit) — the terminal chunk is never emitted. There is no post-loop recovery.

This is a bug class, not gemma4-specific: any parser continue on a finished=True output (tool parser, reasoning parser buffering) loses the stream's finish_reason. Non-deterministic in production: bites when the model keeps its thought channel open through the tool call, so the terminal delta travels the reasoning path into the tool parser.

Fix

Post-loop terminal-chunk guard: track whether any emitted chunk carried finish_reason; if the engine's finished output was suppressed, emit the terminal chunk (finish_reason="tool_calls" when tool calls were detected, else the engine's reason or "stop").

Tests

Two regression tests in TestStreamChatCompletion, covering both swallow branches:

  • test_stream_terminal_finish_reason_when_tool_parser_suppresses_eot — plain tool-parser branch, real Gemma4ToolParser, exact production delta sequence
  • test_stream_terminal_finish_reason_when_reasoning_path_suppresses_eot — reasoning-parser branch

Both fail on unpatched main (assert None == 'tool_calls') and pass with the fix. Full tests/test_server.py suite: 121 passed.

Live verification

Reproduced against a real server (gemma-4-31b-it-4bit, production agent payload with 21 tools): unpatched stream ended tool_calls(finish_reason=null) → [DONE]; patched stream ends tool_calls → finish_reason: "tool_calls" → [DONE].

Related

…he finished output

A finished=True engine output consumed by a parser `continue` in
stream_chat_completion (e.g. gemma4 emits the complete tool call in one
delta, then a bare <turn|> end-of-turn token as the terminal delta,
which the tool parser suppresses) ends the stream with the tool_calls
chunk (finish_reason=null) followed by [DONE] and no finish_reason
chunk. Strict OpenAI clients (pi coding agent) abort with
"Stream ended without finish_reason" and retry the turn.

Track whether any emitted chunk carried a *non-null* finish_reason;
after the streaming loop, if the engines finished output was
suppressed, emit the terminal chunk (finish_reason="tool_calls" when
tool calls were detected, else the engine's reason or "stop").

Also cherry-pick the engine-side fix from waybarrios#681 (Thump604) so that
SimpleEngine stamps finish_reason="stop" on natural-stop streaming
epilogue instead of leaving it null.

Refs waybarrios#672
@funkymonkeymonk
funkymonkeymonk force-pushed the fix/stream-terminal-finish-reason branch from 85f444d to 7aaea58 Compare August 9, 2026 06:18
@funkymonkeymonk

Copy link
Copy Markdown
Contributor Author

@Thump604 thank you for the thorough review and for catching the tracker bug. I have applied all the requested changes:

  1. Tracker now reflects actual non-null serialized finish_reason — switched from to so a chunk does not falsely claim emission.

  2. Added regression test () — verifies the guard still stamps when the engine leaves the reason null.

  3. Cherry-picked the engine-side fix from fix(simple-engine): preserve streaming finish reasons #681 — now stamps on natural-stop streaming epilogue and on token-limit hits, so the combined path is green.

  4. **Changed → ** in the commit message until the full combined path is merged.

Please let me know if anything else needs adjustment.

@funkymonkeymonk
funkymonkeymonk marked this pull request as ready for review August 9, 2026 06:19
@waybarrios

Copy link
Copy Markdown
Owner

just updated this branch from main. Ill review it soon, meanwhile ci is running.

Comment thread vllm_mlx/server.py
),
)
],
usage=get_usage(last_output),

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I found one edge case here: when the engine finishes with finish_reason=None, the normal path already sends usage at the preceding final chunk. Since no finish reason was emitted, this fallback then sends another terminal chunk with the same usage.

That can make clients double-count prompt and completion tokens. Could we ensure usage is emitted only once, ideally on the fallback chunk when it supplies "stop"? A regression assertion that exactly one payload contains usage would cover this.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in Thump604/vllm-mlx@7e454a0. Finished outputs with no terminal reason no longer attach usage to the preceding nonterminal chunk; the fallback stop chunk is now the sole usage-bearing payload. The regression asserts exactly one usage payload with the expected prompt/completion totals. Focused terminal matrix: 4 passed; full tests/test_server.py: 126 passed, 3 deselected. The author can cherry-pick 7e454a0.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Thump604 — cherry-picked 7e454a0 onto the PR head (commit f525d2d), attribution preserved. Verified:

  • RED: without the usage guard, the new assertion in test_stream_terminal_finish_reason_when_engine_emits_none fails assert 2 == 1 — exactly the double-usage payload you diagnosed; the engine-emits-none path attaches usage to the preceding nonterminal chunk and the fallback stop chunk.
  • GREEN: with the guard, both affected tests pass; full tests/test_server.py = 126 passed, 3 deselected, matching your run. ruff/black clean, no new mypy errors at the changed lines.

Comment thread vllm_mlx/server.py
# after a completed tool call), no chunk carried finish_reason and
# OpenAI clients abort with "stream ended without finish_reason".
# Emit the terminal chunk now.
if (

@waybarrios waybarrios Aug 14, 2026 •

Copy link
Copy Markdown
Owner

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Could we add coverage for the reasoning parser swallowing the final delta directly? The relevant early continue is at vllm_mlx/server.py#L6172-L6174.

The current tests cover tool-parser suppression, but not the delta_msg is None path. A small regression test should verify that the stream still ends with finish_reason: "stop" and includes the final usage values.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Added in Thump604/vllm-mlx@7e454a0. The new regression drives the direct reasoning-parser delta_msg is None branch on the finished output and verifies the fallback emits finish_reason="stop" with exactly one final usage payload. Focused terminal matrix: 4 passed; full tests/test_server.py: 126 passed, 3 deselected. The author can cherry-pick 7e454a0.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @Thump604 — same cherry-pick as above (f525d2d). The new test_stream_terminal_when_reasoning_parser_swallows_finished_delta drives the reasoning_parser branch where extract_reasoning_streaming returns None on the finished delta (continue), so no normal-path chunk is emitted and the post-loop guard stamps finish_reason="stop" with exactly one usage payload (prompt=5, completion=2, total=7). Passes with the fix; full suite 126 passed, 3 deselected.

@funkymonkeymonk

Copy link
Copy Markdown
Contributor Author

Applied both review suggestions by cherry-picking Thump604's verified commit 7e454a0 (f525d2d, attribution preserved):

  1. EEC-001 (waybarrios) — duplicate terminal stream usage: the normal-path usage guard now only attaches usage when the chunk actually carries a finish_reason (tool_calls_detected or output.finish_reason), so a swallowed finished output defers usage to the sole fallback stop chunk. Reduces double-counting of prompt/completion tokens.
  2. TC-001 (waybarrios) — added test_stream_terminal_when_reasoning_parser_swallows_finished_delta covering the reasoning-parser delta_msg is None continue branch on the finished delta.

Verification (local, arm64 py3.12 via mise, pip install -e ".[dev]"):

  • RED proven: pre-fix, test_stream_terminal_finish_reason_when_engine_emits_none fails assert 2 == 1 (double usage).
  • GREEN: 4 affected tests pass; full tests/test_server.py = 126 passed, 3 deselected (matches @Thump604's run).
  • ruff/black clean; no new mypy errors at the changed lines.

@waybarrios

Copy link
Copy Markdown
Owner

All set. ready to go

@waybarrios
waybarrios merged commit 61c78dc into waybarrios:main Aug 15, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants