Skip to content

[Bugfix][Frontend] Preserve abort finish_reason for scale-out token streams - #47933

Merged
sfeng33 merged 4 commits into
vllm-project:mainfrom
Sunt-ing:34
Oct 4, 2026
Merged

sfeng33 merged 4 commits into
vllm-project:mainfrom
Sunt-ing:34

Conversation

@Sunt-ing

@Sunt-ing Sunt-ing commented Jul 8, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

In token mode, /inference/v1/generate streams skip engine outputs that have no new token IDs. An abort always ends with such an output (token_ids=[], finish_reason="abort"), so an aborted stream ends with [DONE] and no finish_reason. The only exception is when the abort happens to be merged into a pending token delta. Text mode (#58588) and the Rust frontend already emit every terminal output.

  • Every output that has a finish_reason is now emitted, in both output modes.
  • The per-choice token counter is now sized by sampling_params.n. It was sized from the first output, which with n > 1 may hold only choice 0. The next choice then raised IndexError, and the stream ended in a 500 error chunk.

Test Plan

Live server on H200 with Qwen/Qwen3-0.6B. The same requests went to main (bc21cba) and to this PR:

vllm serve Qwen/Qwen3-0.6B --max-model-len 2048 --no-enable-prefix-caching --tokens-only
vllm serve Qwen/Qwen3-0.6B --max-model-len 2048 --no-enable-prefix-caching --enable-scale-out --max-num-seqs 1

Each request was a streamed /inference/v1/generate call with max_tokens: 900, ignore_eos: true and include_usage. It was aborted through /abort_requests (or /inference/v1/abort_requests) once the first token chunk arrived. With --max-num-seqs 1, a second request could also be aborted while still queued.

Test Result

Case main this PR
tokens mode, abort mid-stream (also with logprobs / flat_logprobs) no finish_reason unless the abort merged into a token chunk final chunk token_ids: [], finish_reason: "abort" (logprobs: {"content": []})
abort while queued (0 tokens) no choice chunk one abort chunk
n=2, abort mid-stream list index out of range 500 chunk in 4/7 runs both choices get abort, 12/12 runs
text-mode abort, non-stream abort, length / EOS stop — unchanged

Usage completion_tokens matched the streamed token count in every case.

…treams

Signed-off-by: Ting Sun <suntcrick@gmail.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added frontend bug Something isn't working labels Jul 8, 2026
@Sunt-ing

Sunt-ing commented Jul 8, 2026

Copy link
Copy Markdown
Contributor Author

@DarkLight1337 A simple PR. PTAL. Thanks~

@mergify

mergify Bot commented Sep 17, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @Sunt-ing.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 17, 2026
Signed-off-by: Ting Sun <suntcrick@gmail.com>
@Sunt-ing

Copy link
Copy Markdown
Contributor Author

Hi @aoshen02, could you please take a look at this simple bugfix? Thanks~

@mergify

mergify Bot commented Oct 3, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @Sunt-ing.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Oct 3, 2026
…e streams

Token-mode /inference/v1/generate streams dropped engine outputs without
new token IDs, so an aborted request ended without finish_reason="abort".
Emit every output that has a finish reason, matching text mode and the
Rust frontend.

Also size the per-choice token counter by sampling_params.n: with n > 1
the first output may not include every choice, which raised IndexError.

Signed-off-by: Flora Feng <4florafeng@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
@sfeng33

sfeng33 commented Oct 4, 2026

Copy link
Copy Markdown
Contributor

/ci run --allow-stale

@sfeng33
sfeng33 enabled auto-merge (squash) October 4, 2026 01:43
@github-actions github-actions Bot added the ready ONLY add when PR is ready to merge/full CI is needed label Oct 4, 2026
@github-actions

github-actions Bot commented Oct 4, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #92775 for commit 1dcd0c41d357.

⚠️ This PR is 6 commits behind upstream main. Running CI at your own risk because --allow-stale was requested; outdated CI configuration may cause failures. Before merging, merge or rebase onto the latest main, then rerun /ci run on the latest PR commit.

@vllm-agent

Copy link
Copy Markdown
Contributor

CI selector (shadow): 2 test steps (5 jobs) instead of 35 (48 jobs)

Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works.

Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.

steps (jobs) Today's rules Selector Would skip Would add
NVIDIA, CPU and others 35 (48) 2 (5) 33 (43) 0 (0)
AMD mirrors 34 (42) 2 (5) 32 (37) 0 (0)
Selector would run (2)
  • entrypoints-integration-api-server ×4
  • rust-frontend-serve-admin-coverage
Would skip (today's rules run them) (33)
  • basic-correctness ×2
  • basic-correctness-cpu-offload
  • basic-correctness-cumem
  • basic-correctness-prefetch-offload
  • basic-correctness-sleep-mode
  • basic-models-test-other-cpu
  • basic-models-tests-initialization
  • basic-models-tests-other
  • benchmarks-cli-test
  • cpu-language-generation-and-pooling-model-tests ×3
  • entrypoints-integration-api-server-generate
  • entrypoints-integration-api-server-openai-chat_completion
  • entrypoints-integration-api-server-openai-completion
  • entrypoints-integration-llm
  • entrypoints-integration-multimodal
  • entrypoints-integration-pooling
  • entrypoints-integration-responses-api
  • entrypoints-integration-speech_to_text
  • entrypoints-unit-tests
  • examples
  • kernels-fla-ops-test-b200
  • kernels-mhc-test-b200
  • kernels-root-misc-test-b200
  • language-models-tests-granite-l4-compatibility
  • language-models-tests-hybrid ×2
  • language-models-tests-standard
  • multi-modal-models-standard-1-qwen2
  • multi-modal-models-standard-2-qwen3-gemma
  • multi-modal-models-standard-3-llava-qwen2-vl
  • multi-modal-models-standard-4-other-whisper
  • multi-modal-processor ×4
  • multi-modal-processor-cpu ×4
  • scale-out-ec-e2e-2-gpus
Would add (today's rules do not run them) (0)

none

AMD mirrors: would skip (32)
  • basic-correctness ×2
  • basic-correctness-cpu-offload
  • basic-correctness-cumem
  • basic-correctness-prefetch-offload
  • basic-correctness-sleep-mode
  • basic-models-tests-initialization
  • basic-models-tests-other
  • benchmarks-cli-test
  • entrypoints-integration-api-server-generate
  • entrypoints-integration-api-server-openai-chat_completion
  • entrypoints-integration-api-server-openai-completion
  • entrypoints-integration-llm
  • entrypoints-integration-multimodal
  • entrypoints-integration-pooling
  • entrypoints-integration-responses-api
  • entrypoints-integration-speech_to_text
  • entrypoints-unit-tests
  • examples
  • kernels-fla-ops-test-b200
  • kernels-mhc-test-b200
  • kernels-root-misc-test-b200
  • language-models-tests-granite-l4-compatibility
  • language-models-tests-hybrid ×2
  • language-models-tests-standard
  • multi-modal-models-standard-1-qwen2
  • multi-modal-models-standard-2-qwen3-gemma
  • multi-modal-models-standard-3-llava-qwen2-vl
  • multi-modal-models-standard-4-other-whisper
  • multi-modal-processor ×4
  • platform-tests
  • pytorch-compilation-passes-unit-tests
  • scale-out-ec-e2e-2-gpus
AMD mirrors: would add (0)

none

3 changed files · base bc21cba967 · head 1dcd0c41d3 · Python record: build 92769 at 84bcbc6264 · kernel record: table 84bcbc6 (build 92769), map 84bcbc6 · not counted: 10 build steps, 5 A100 steps the generator no longer emits, 0 optional steps the selector would also run

@sfeng33
sfeng33 merged commit 50b404e into vllm-project:main Oct 4, 2026
96 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working frontend ready ONLY add when PR is ready to merge/full CI is needed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants