Skip to content

[Frontend] Add output_mode to /inference/v1/generate (RFC #56851 Phase 1) - #58588

Merged
DarkLight1337 merged 9 commits into
vllm-project:mainfrom
hickeyma:add-outputmode-generate-api
Oct 3, 2026
Merged

DarkLight1337 merged 9 commits into
vllm-project:mainfrom
hickeyma:add-outputmode-generate-api

Conversation

@hickeyma

@hickeyma hickeyma commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Phase 1 of #56851. A generate request can now set output_mode: "text" to get detokenized text next to the token IDs, so it doesn't need a separate /derender call. The default stays "tokens", so existing clients see the same response with one new field, output_mode.

  • GenerateResponse and GenerateStreamResponse are now unions of a tokens class and a text class, picked by output_mode. A body without the field still parses as tokens, so /derender keeps working with older responses. /derender now returns a 400 for text responses because a text response already has its stop string cut and decoding its token_ids again would put it back.
  • At text, logprobs carry decoded tokens and bytes. --return-tokens-as-token-ids keeps the token_id:N placeholders.
  • Text streams also send a chunk when there's new text or a finish reason but no new token IDs. That covers text held back for stop strings and the final chunk after an abort.
  • output_mode="text" returns 400 on a server without a tokenizer or when sampling_params.detokenize is false.
  • /inference/v1/abort_requests is registered wherever generate is served and requires the API key when --api-key is set. The unauthenticated /abort_requests stays --tokens-only only, as before.
  • New docs page: docs/serving/online_serving/token_in_token_out.md.

Test Plan

pytest tests/entrypoints/scale_out/token_in_token_out/ -v
pytest tests/entrypoints/scale_out/derender/test_derender_parity.py -v
pytest tests/entrypoints/scale_out/derender/test_derender_stream.py -v
pytest tests/entrypoints/scale_out/test_factories.py -v

The new tests cover:

  • protocol parsing and validation, including /derender rejecting text responses
  • text responses and stream chunks, including chunks that only carry a finish reason or only carry text
  • a text stream getting its abort chunk after /inference/v1/abort_requests
  • abort paths: the API key protected one wherever generate runs, the bare one only with --tokens-only
  • parity with /v1/completions/derender, batch and streaming, including logprobs and stop strings

Test Result

TTFT/TPOT benchmark results in #58588 (comment) (from @shimib)

@mergify

mergify Bot commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

Documentation preview: https://vllm--58588.org.readthedocs.build/en/58588/

@shimib

shimib commented Sep 25, 2026

Copy link
Copy Markdown
Contributor

Section 12: TTFT/TPOT benchmarks

Setup. 1×A100-40G (GKE), Qwen/Qwen3-8B, hermes tool parser + qwen3
reasoning parser where applicable. Seeded greedy traffic, open-loop
(rate-driven) load, 300 s window per cell, one run per cell, all cells
measured on the same rig within one day. TTFT/TPOT measured client-side;
TPOT = (last delta − first delta) / (output_tokens − 1), which is robust to
chunk coalescing on the derender path. Four configurations over the same
render → generate flow:

  • G — generate only: client consumes the token SSE stream directly
    (--tokens-only worker). Baseline.
  • T — this PR: output_mode="text" inline on the generate server
    (worker with tokenizer loaded, --enable-scale-out).
  • B — generate + per-chunk streaming /derender, parsers configured on
    the derender tier (--renderer-num-workers=8).
  • BP — same as B with no parsers on the tier (plain detok path).

Images: vllm/vllm-openai:nightly (digest ed3c505d…) for G/B/BP and the
derender tier; this PR (head acde4917ff) as a pure-Python overlay on that
same nightly for the T worker. The render step is included in all four
configs, so columns differ only in where token→text happens.

Results (TTFT p50/p95 | TPOT p50/p95, ms):

shape @ rps G (gen only) T (inline text, this PR) B (derender, parsed) BP (derender, plain)
chat @0.5 116/163 | 13.8/14.5 113/156 | 13.8/14.5 222/293 | 13.5/14.2 214/294 | 13.4/14.1
chat @2 122/186 | 16.6/19.2 121/187 | 16.6/19.3 263/3746 | 16.6/31.4 243/354 | 16.3/18.8
chat @6 5185/16994 | 53.7/56.9 ¹ 4907/15893 | 53.7/56.8 ¹ did not drain ² did not drain ²
tools @0.5 49/57 | 13.1/13.2 47/56 | 13.1/13.2 150/161 | 12.4/12.7 n/a ³
tools @2 52/59 | 13.1/13.2 52/59 | 13.1/13.2 152/162 | 12.5/12.8 n/a ³
tools @6 57/69 | 13.4/13.8 54/66 | 13.4/13.8 163/5921 | 12.9/65.9 n/a ³
reasoning @0.5 73/89 | 17.2/18.0 74/91 | 17.3/18.4 29183/59263 | 69.3/101.7 194/219 | 16.7/17.7
reasoning @1 109/138 | 23.0/24.4 119/158 | 26.5/31.0 29499/614902 | 67.9/372.9 272/333 | 22.7/24.6

Shapes: chat = 128–512 output tok, tools = tool-calling 64–256 tok,
reasoning = 2,048–8,192 tok.

¹ 6 rps exceeds this rig's sustainable chat rate (~5.4 rps measured
previously); the G/T figures are queue-dominated but the streams survive.
² Three attempts including a 2,400 s budget; the request queue never drains
once the derender path's added cost is stacked on an already-oversubscribed
workload.
³ Structural: a parser-less render server rejects tool requests with 400 at
/render, so a plain tier cannot serve tool traffic end-to-end.

Findings:

  1. Inline text is free. T matches G within noise on every cell — TTFT
    within ±4 ms, TPOT identical — from light load through overload. On
    latency there is no observable cost to detokenizing inline on the decode
    pod for this traffic, consistent with the CPU headroom numbers already in
    the RFC.
  2. The derender hop itself costs a flat ~+100 ms TTFT (render + first
    per-chunk call) and nothing on TPOT — the BP column. Benign at these
    rates.
  3. Parser replay is the dominant cost, isolated from transport. On long
    reasoning, parsed derender reaches 29 s TTFT p50 / 69 ms TPOT while plain
    detok on identical traffic sits at ~200 ms / 16.7 ms. The tier was sized
    at 8 renderer workers; this is replay compute plus the queueing it causes,
    not the HTTP hop.
  4. Tier pressure is visible well before collapse: B's TTFT p95 at 2–6 rps
    (3.7 s chat, 5.9 s tools) against clean p50s.

Note on the ~+15% E2E figure in #56851's motivation. That figure was
measured on the pre-merge #50550 revision, whose replay fed the echoed
history through one parse_delta call per chunk. The merged implementation
replays one call per original chunk boundary (needed for parity and stable
tool-call IDs). A same-day A/B on this rig — only the image differing — puts
the merged replay at ~2.4× the pre-merge E2E on long reasoning (72 s → 176 s
p50 at 0.5 rps). The +15% should be read as a lower bound for what is now on
main; today's table is internally consistent and should not be compared
numerically against the earlier run. This also sharpens the case for #57571
(parser cache / bidi) independently of this PR.

Caveats. One model on one GPU; reasoning shapes above 1 rps not attempted
(oversubscribed by construction); tool cells reuse the request shapes from
the earlier CPU benchmark. Raw per-request logs and cell summaries retained —
happy to share or re-run specific cells.

@hickeyma

Copy link
Copy Markdown
Contributor Author

Thanks @shimib for the benchmarks.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

Comment thread docs/serving/online_serving/token_in_token_out.md
Comment thread docs/serving/online_serving/token_in_token_out.md Outdated
Comment thread docs/serving/online_serving/token_in_token_out.md Outdated
Comment thread docs/serving/online_serving/token_in_token_out.md Outdated
Comment thread vllm/entrypoints/scale_out/token_in_token_out/protocol.py Outdated
Comment thread vllm/entrypoints/scale_out/token_in_token_out/protocol.py Outdated
Comment thread vllm/entrypoints/scale_out/token_in_token_out/api_router.py Outdated
@mergify

mergify Bot commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @hickeyma.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 29, 2026
@hickeyma
hickeyma marked this pull request as draft September 29, 2026 20:24
@hickeyma

Copy link
Copy Markdown
Contributor Author

Moving to draft for now as have some updates to do.

@hickeyma
hickeyma force-pushed the add-outputmode-generate-api branch from 40471fd to 835d037 Compare September 30, 2026 08:53
@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

❌ This PR is 1 commit behind upstream main. Your branch must contain every commit currently on upstream main. No new CI build was started. Merge or rebase onto the latest main, then rerun /ci run. To test this branch at your own risk, use /ci run --allow-stale.

Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
@mergify

mergify Bot commented Oct 2, 2026

Copy link
Copy Markdown
Contributor

Hi @hickeyma, the pre-commit checks have failed. Please run:

uv pip install pre-commit>=4.5.1
pre-commit install
pre-commit run --all-files

Then, commit the changes and push to your branch.

For future commits, pre-commit will run automatically on changed files before each commit.

auto-merge was automatically disabled October 2, 2026 08:25

Head branch was pushed to by a user without write access

@mergify mergify Bot removed the needs-rebase label Oct 2, 2026
@hickeyma hickeyma changed the title [Frontend] Add output_mode to /inference/v1/generate (RFC #56851 Phase 1) [Frontend] Add output_mode to /inference/v1/generate (RFC #56851 Phase 1) Oct 2, 2026
@hickeyma

hickeyma commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #92616 for commit b9111d35fdc9.

@vllm-agent

vllm-agent commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

CI selector (shadow): 61 test steps (70 jobs) instead of 39 (52 jobs)

Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works.

Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.

steps (jobs) Today's rules Selector Would skip Would add
NVIDIA, CPU and others 39 (52) 61 (70) 24 (34) 46 (52)
AMD mirrors 34 (42) 57 (66) 20 (25) 43 (49)
Selector would run (61)
  • basic-correctness-cpu-offload
  • basic-correctness-prefetch-offload
  • batch-invariance-b200
  • benchmarks-cli-test
  • cpu-reasoning-renderers
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus
  • distributed-dp-tests-2-gpus
  • distributed-dp-tests-4-gpus
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus
  • distributed-mooncakeconnector-pd-accuracy-4-gpus
  • distributed-nixlconnector-pd-accuracy-4-gpus
  • distributed-torchrun-examples-4-gpus
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus
  • elastic-ep-scaling-test
  • engine
  • entrypoints-integration-api-server ×4
  • entrypoints-integration-api-server-generate
  • entrypoints-integration-api-server-openai-chat_completion
  • entrypoints-integration-api-server-openai-completion
  • entrypoints-integration-multimodal
  • entrypoints-integration-pooling
  • entrypoints-integration-responses-api
  • entrypoints-integration-speech_to_text
  • entrypoints-unit-tests
  • fault-tolerance-e2e-2xh100
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus
  • kv-offload-large
  • kv-offload-medium
  • kv-offload-small
  • lm-eval-dspark-watermark-2xh100
  • lm-eval-small-models
  • lm-eval-turboquant-k3v4nc
  • lm-eval-turboquant-k8v4
  • lm-eval-turboquant-t3nc
  • lm-eval-turboquant-t4nc
  • lm-eval-watermarking
  • model-executor
  • mooncake-ec-tcp-e2e-2-gpus
  • mrcr-eval-small-models
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus
  • nixlconnector-pd-edge-cases-2-gpus
  • openai-api-correctness
  • pipeline-context-parallelism-4-gpus
  • plugin-tests-2-gpus
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus
  • pytorch-compilation-unit-tests
  • pytorch-fullgraph-test
  • quantization ×4
  • quantized-moe-test-b200
  • rayexecutorv2-4-gpus
  • rust-frontend-openai-coverage
  • rust-frontend-serve-admin-coverage
  • scale-out-ec-e2e-2-gpus
  • v1-core
  • v1-kv-connectors ×4
  • v1-logits-oracle
  • v1-metrics-lmeval
  • v1-others-cpu
  • v1-sample
Would skip (today's rules run them) (24)
  • basic-correctness ×2
  • basic-correctness-cumem
  • basic-correctness-sleep-mode
  • basic-models-test-other-cpu
  • basic-models-tests-initialization
  • basic-models-tests-other
  • cpu-language-generation-and-pooling-model-tests ×3
  • cpu-multimodal-config
  • cpu-params-env-tokenizers-parser
  • cpu-tool-parsers
  • entrypoints-integration-llm
  • examples
  • kernels-fla-ops-test-b200
  • kernels-mhc-test-b200
  • kernels-root-misc-test-b200
  • language-models-tests-granite-l4-compatibility
  • language-models-tests-hybrid ×2
  • language-models-tests-standard
  • multi-modal-models-standard-1-qwen2
  • multi-modal-models-standard-2-qwen3-gemma
  • multi-modal-models-standard-3-llava-qwen2-vl
  • multi-modal-models-standard-4-other-whisper
  • multi-modal-processor ×4
  • multi-modal-processor-cpu ×4
Would add (today's rules do not run them) (46)
  • batch-invariance-b200 (Python record)
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • distributed-dp-tests-2-gpus (code map)
  • distributed-dp-tests-4-gpus (Python record)
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-mooncakeconnector-pd-accuracy-4-gpus (Python record)
  • distributed-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-torchrun-examples-4-gpus (Python record)
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • elastic-ep-scaling-test (Python record)
  • engine (code map)
  • fault-tolerance-e2e-2xh100 (Python record)
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus (Python record)
  • kv-offload-large (Python record)
  • kv-offload-medium (Python record)
  • kv-offload-small (Python record)
  • lm-eval-dspark-watermark-2xh100 (Python record)
  • lm-eval-small-models (Python record)
  • lm-eval-turboquant-k3v4nc (Python record)
  • lm-eval-turboquant-k8v4 (Python record)
  • lm-eval-turboquant-t3nc (Python record)
  • lm-eval-turboquant-t4nc (Python record)
  • lm-eval-watermarking (Python record)
  • model-executor (Python record)
  • mooncake-ec-tcp-e2e-2-gpus (Python record)
  • mrcr-eval-small-models (Python record)
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus (Python record)
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus (Python record)
  • nixlconnector-pd-edge-cases-2-gpus (Python record)
  • openai-api-correctness (Python record)
  • pipeline-context-parallelism-4-gpus (Python record)
  • plugin-tests-2-gpus (code map)
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus (Python record)
  • pytorch-compilation-unit-tests (Python record)
  • pytorch-fullgraph-test (Python record)
  • quantization ×4 (Python record)
  • quantized-moe-test-b200 (Python record)
  • rayexecutorv2-4-gpus (Python record)
  • rust-frontend-openai-coverage (code map)
  • v1-core (Python record)
  • v1-kv-connectors ×4 (code map)
  • v1-logits-oracle (Python record)
  • v1-metrics-lmeval (Python record)
  • v1-others-cpu (code map)
  • v1-sample (Python record)
AMD mirrors: would skip (20)
  • basic-correctness ×2
  • basic-correctness-cumem
  • basic-correctness-sleep-mode
  • basic-models-tests-initialization
  • basic-models-tests-other
  • entrypoints-integration-llm
  • examples
  • kernels-fla-ops-test-b200
  • kernels-mhc-test-b200
  • kernels-root-misc-test-b200
  • language-models-tests-granite-l4-compatibility
  • language-models-tests-hybrid ×2
  • language-models-tests-standard
  • multi-modal-models-standard-1-qwen2
  • multi-modal-models-standard-2-qwen3-gemma
  • multi-modal-models-standard-3-llava-qwen2-vl
  • multi-modal-models-standard-4-other-whisper
  • multi-modal-processor ×4
  • platform-tests
  • pytorch-compilation-passes-unit-tests
AMD mirrors: would add (43)
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • distributed-dp-tests-2-gpus (code map)
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-mooncakeconnector-pd-accuracy-4-gpus (Python record)
  • distributed-nixlconnector-pd-accuracy-4-gpus (Python record)
  • distributed-torchrun-examples-4-gpus (Python record)
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • engine (code map)
  • fusion-and-compile-unit-tests-2xb200 (Python record)
  • fusion-e2e-tp2-asynctp-config-sweep-h100 (Python record)
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus (Python record)
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus (Python record)
  • kv-offload-large (Python record)
  • kv-offload-medium (Python record)
  • kv-offload-small (Python record)
  • lm-eval-dspark-watermark-2xh100 (Python record)
  • lm-eval-small-models (Python record)
  • lm-eval-turboquant-k3v4nc (Python record)
  • lm-eval-turboquant-k8v4 (Python record)
  • lm-eval-turboquant-t3nc (Python record)
  • lm-eval-turboquant-t4nc (Python record)
  • lm-eval-watermarking (Python record)
  • model-executor (Python record)
  • mooncake-ec-tcp-e2e-2-gpus (Python record)
  • mrcr-eval-small-models (Python record)
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus (Python record)
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus (Python record)
  • nixlconnector-pd-edge-cases-2-gpus (Python record)
  • nixlconnector-pd-spec-decode-acceptance-2-gpus (Python record)
  • openai-api-correctness (Python record)
  • pipeline-context-parallelism-4-gpus (Python record)
  • plugin-tests-2-gpus (code map)
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus (Python record)
  • pytorch-compilation-unit-tests (Python record)
  • pytorch-fullgraph-test (Python record)
  • quantization ×4 (Python record)
  • quantized-moe-test-b200 (Python record)
  • rust-frontend-openai-coverage (code map)
  • v1-core (Python record)
  • v1-kv-connectors ×4 (code map)
  • v1-logits-oracle (Python record)
  • v1-metrics-lmeval (Python record)
  • v1-sample (Python record)

18 changed files · base 111f71a6db · head d8f421b5af · Python record: build 92561 at 4056c8ac1f · kernel record: table 4056c8a (build 92561), map 4056c8a · not counted: 10 build steps, 5 A100 steps the generator no longer emits, 63 optional steps the selector would also run

@hickeyma

hickeyma commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

/ci retry

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

✅ Queued 8 failed job(s) for retry in Buildkite CI #92616.

test_derender_stream.py still used GenerateResponseChoice and built
GenerateResponse directly. Switch those to GenerateTokensChoice and
GenerateTokensResponse to match the tokens/text split in the protocol.

Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
@hickeyma
hickeyma force-pushed the add-outputmode-generate-api branch from b9111d3 to 81e60d2 Compare October 2, 2026 16:38
@hickeyma

hickeyma commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

❌ This PR is 32 commits behind upstream main. Your branch must contain every commit currently on upstream main. No new CI build was started. Merge or rebase onto the latest main, then rerun /ci run. To test this branch at your own risk, use /ci run --allow-stale.

@hickeyma

hickeyma commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #92670 for commit 3a3fd1cb0887.

The test compared streamed token IDs against a separate non-streaming
generate call which assumes two greedy runs pick identical tokens.
This led to flaky test failures from time to time.
Check the streamed text against the non-streaming derender of the same
token IDs instead.

Signed-off-by: Martin Hickey <martin.hickey@ie.ibm.com>
@hickeyma
hickeyma force-pushed the add-outputmode-generate-api branch from 3a3fd1c to 7cd5f1d Compare October 2, 2026 19:52
@hickeyma

hickeyma commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

github-actions Bot commented Oct 2, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #92693 for commit d8f421b5afaf.

@hickeyma

hickeyma commented Oct 3, 2026

Copy link
Copy Markdown
Contributor Author

/ci retry

@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

✅ Queued 1 failed job(s) for retry in Buildkite CI #92693.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation frontend ready ONLY add when PR is ready to merge/full CI is needed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants