Repository navigation
[Bugfix][Frontend] Preserve abort finish_reason for scale-out token streams - #47933
Conversation
…treams Signed-off-by: Ting Sun <suntcrick@gmail.com>
|
@DarkLight1337 A simple PR. PTAL. Thanks~ |
|
This pull request has merge conflicts that must be resolved before it can be |
Signed-off-by: Ting Sun <suntcrick@gmail.com>
|
Hi @aoshen02, could you please take a look at this simple bugfix? Thanks~ |
|
This pull request has merge conflicts that must be resolved before it can be |
…e streams Token-mode /inference/v1/generate streams dropped engine outputs without new token IDs, so an aborted request ended without finish_reason="abort". Emit every output that has a finish reason, matching text mode and the Rust frontend. Also size the per-choice token counter by sampling_params.n: with n > 1 the first output may not include every choice, which raised IndexError. Signed-off-by: Flora Feng <4florafeng@gmail.com>
Signed-off-by: Flora Feng <4florafeng@gmail.com>
|
/ci run --allow-stale |
|
✅ Triggered Buildkite CI #92775 for commit
|
CI selector (shadow): 2 test steps (5 jobs) instead of 35 (48 jobs)Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works. Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.
Selector would run (2)
Would skip (today's rules run them) (33)
Would add (today's rules do not run them) (0)none AMD mirrors: would skip (32)
AMD mirrors: would add (0)none 3 changed files · base |
Purpose
In token mode,
/inference/v1/generatestreams skip engine outputs that have no new token IDs. An abort always ends with such an output (token_ids=[],finish_reason="abort"), so an aborted stream ends with[DONE]and nofinish_reason. The only exception is when the abort happens to be merged into a pending token delta. Text mode (#58588) and the Rust frontend already emit every terminal output.finish_reasonis now emitted, in both output modes.sampling_params.n. It was sized from the first output, which withn > 1may hold only choice 0. The next choice then raisedIndexError, and the stream ended in a 500 error chunk.Test Plan
Live server on H200 with
Qwen/Qwen3-0.6B. The same requests went tomain(bc21cba) and to this PR:Each request was a streamed
/inference/v1/generatecall withmax_tokens: 900,ignore_eos: trueandinclude_usage. It was aborted through/abort_requests(or/inference/v1/abort_requests) once the first token chunk arrived. With--max-num-seqs 1, a second request could also be aborted while still queued.Test Result
logprobs/flat_logprobs)finish_reasonunless the abort merged into a token chunktoken_ids: [], finish_reason: "abort"(logprobs: {"content": []})abortchunkn=2, abort mid-streamlist index out of range500 chunk in 4/7 runsabort, 12/12 runslength/ EOSstopUsage
completion_tokensmatched the streamed token count in every case.