[Frontend] Add output_mode to /inference/v1/generate (RFC #56851 Phase 1) - #58588
Conversation
|
Documentation preview: https://vllm--58588.org.readthedocs.build/en/58588/ |
Section 12: TTFT/TPOT benchmarksSetup. 1×A100-40G (GKE), Qwen/Qwen3-8B, hermes tool parser + qwen3
Images: Results (TTFT p50/p95 | TPOT p50/p95, ms):
Shapes: chat = 128–512 output tok, tools = tool-calling 64–256 tok, ¹ 6 rps exceeds this rig's sustainable chat rate (~5.4 rps measured Findings:
Note on the ~+15% E2E figure in #56851's motivation. That figure was Caveats. One model on one GPU; reasoning shapes above 1 rps not attempted |
|
Thanks @shimib for the benchmarks. |
|
This pull request has merge conflicts that must be resolved before it can be |
|
Moving to draft for now as have some updates to do. |
40471fd to
835d037
Compare
|
❌ This PR is 1 commit behind upstream |
Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
|
Hi @hickeyma, the pre-commit checks have failed. Please run: uv pip install pre-commit>=4.5.1
pre-commit install
pre-commit run --all-filesThen, commit the changes and push to your branch. For future commits, |
Head branch was pushed to by a user without write access
|
/ci run |
|
✅ Triggered Buildkite CI #92616 for commit |
CI selector (shadow): 61 test steps (70 jobs) instead of 39 (52 jobs)Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works. Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.
Selector would run (61)
Would skip (today's rules run them) (24)
Would add (today's rules do not run them) (46)
AMD mirrors: would skip (20)
AMD mirrors: would add (43)
18 changed files · base |
|
/ci retry |
|
✅ Queued 8 failed job(s) for retry in Buildkite CI #92616. |
test_derender_stream.py still used GenerateResponseChoice and built GenerateResponse directly. Switch those to GenerateTokensChoice and GenerateTokensResponse to match the tokens/text split in the protocol. Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
b9111d3 to
81e60d2
Compare
|
/ci run |
|
❌ This PR is 32 commits behind upstream |
|
/ci run |
|
✅ Triggered Buildkite CI #92670 for commit |
The test compared streamed token IDs against a separate non-streaming generate call which assumes two greedy runs pick identical tokens. This led to flaky test failures from time to time. Check the streamed text against the non-streaming derender of the same token IDs instead. Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
3a3fd1c to
7cd5f1d
Compare
|
/ci run |
|
✅ Triggered Buildkite CI #92693 for commit |
|
/ci retry |
|
✅ Queued 1 failed job(s) for retry in Buildkite CI #92693. |
Purpose
Phase 1 of #56851. A generate request can now set
output_mode: "text"to get detokenized text next to the token IDs, so it doesn't need a separate/derendercall. The default stays"tokens", so existing clients see the same response with one new field,output_mode.GenerateResponseandGenerateStreamResponseare now unions of a tokens class and a text class, picked byoutput_mode. A body without the field still parses as tokens, so/derenderkeeps working with older responses./derendernow returns a 400 for text responses because a text response already has its stop string cut and decoding itstoken_idsagain would put it back.text, logprobs carry decoded tokens andbytes.--return-tokens-as-token-idskeeps thetoken_id:Nplaceholders.output_mode="text"returns 400 on a server without a tokenizer or whensampling_params.detokenizeis false./inference/v1/abort_requestsis registered wherever generate is served and requires the API key when--api-keyis set. The unauthenticated/abort_requestsstays--tokens-onlyonly, as before.docs/serving/online_serving/token_in_token_out.md.Test Plan
The new tests cover:
/derenderrejecting text responses/inference/v1/abort_requests--tokens-only/v1/completions/derender, batch and streaming, including logprobs and stop stringsTest Result
TTFT/TPOT benchmark results in #58588 (comment) (from @shimib)