Skip to content

[Frontend] Integer token IDs for generate output logprobs (GenerateLogProbs) - #58181

Merged
DarkLight1337 merged 7 commits into
vllm-project:mainfrom
yinli-systems:generate-logprobs-shape
Oct 6, 2026
Merged

DarkLight1337 merged 7 commits into
vllm-project:mainfrom
yinli-systems:generate-logprobs-shape

Conversation

@yinli-systems

@yinli-systems yinli-systems commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Fixes #57574
Fixes #59241

Implements the GenerateLogProbs shape agreed in #57574. /inference/v1/generate is a vLLM-defined token-in/token-out API, but its output logprobs used the OpenAI shape: the generate server has no tokenizer, so it wrote every token as a "token_id:N" placeholder string that derender parsed back out, and the two frontends disagreed on bytes (Python left it unset, the Rust frontend set it to the UTF-8 bytes of the placeholder itself).

Scope is exactly the split @hickeyma described in #57574 (comment): this PR is content only. #57442 stacks GenerateLogProbs.sampled on top afterwards.

Shape

class GenerateLogProb(BaseModel):
    token_id: int
    logprob: float
    rank: int | None = None

class GenerateLogProbsContent(GenerateLogProb):
    top_logprobs: list[GenerateLogProb] = []

class GenerateLogProbs(BaseModel):
    content: list[GenerateLogProbsContent] | None = None
  • top_logprobs is a list, not a dict: JSON turns dict keys into strings and the order would be implicit. It keeps the engine's order: the sampled token first, then the remaining candidates in rank order. With non-greedy sampling the sampled token can be outside the top k (e.g. ranks [5, 1, 2]); it then takes one of the k slots and rank k is left out, as on the OpenAI endpoints and the main engine path.
  • rank is kept: the engine already computes it and the Rust frontend's gRPC CandidateTokenInfo carries it. It was being dropped on the floor before.
  • No bytes, no string token.
  • logprob is required, as in [Feature]: Integer token IDs for logprobs in /inference/v1/generate responses (GenerateLogProbs) #57574 and the Rust frontend. When the sampled token is absent from the engine's top-k map the server passes -9999.0 explicitly, the value the OpenAI shapes already put on the wire for that case, so it is unchanged.

What changed

  • Python generate server (scale_out/token_in_token_out): _create_tokens_logprobs builds the new type; GenerateResponseChoice.logprobs and GenerateResponseStreamChoice.logprobs are GenerateLogProbs | None, streaming and non-streaming.
  • Rust frontend (routes/inference/generate): same field swap, position_to_generate_logprobs_content emits ids + rank instead of format_token_id.
  • Derender: _resolve_logprobs converts GenerateLogProbs → ChatCompletionLogProbs, filling token/bytes from the tokenizer, and _convert_chat_logprobs_to_completion_logprobs then produces CompletionLogProbs as before. The U+FFFD byte-fallback correction (_correct_decoded_token) is untouched — it needs the preceding sampled token ids as context, and the integer shape now carries those directly instead of reconstructing them from placeholder strings.
  • Placeholder layer deleted from the generate path: _parse_token_id_placeholder and resolve_token_id_placeholder are gone; the id-taking core is now decode_token_id in entrypoints/generate/base/serving.py.
  • Out of scope, unchanged: prompt_logprobs (already integer-keyed) and return_tokens_as_token_ids on /v1/chat/completions / /v1/completions, which is a user-facing OpenAI option that uses token_id:N on purpose.

Release note

Breaking: output logprobs on /inference/v1/generate are now GenerateLogProbs with integer token_id instead of ChatCompletionLogProbs with "token_id:N" placeholder strings, in both the Python and Rust frontends, streaming and non-streaming. rank is now reported and bytes is gone. Clients reading only content[i].logprob are unaffected; clients that parsed the placeholder should read content[i].token_id. Derender output (/v1/chat/completions/derender, /v1/completions/derender) is unchanged. return_tokens_as_token_ids on the OpenAI endpoints is unchanged.

Documented in docs/serving/online_serving/renderer.md (new "Generate Output Logprobs" section, with the breaking-change note) and derenderer.md.

Test Plan

  • tests/entrypoints/scale_out/derender/test_generate_logprobs_conversion.py (new): unit tests over _resolve_logprobs with a stub tokenizer — decoding + bytes + rank-order preservation, the U+FFFD byte-fallback path where the second half of a two-byte character is only decodable with the preceding sampled id as context, and content=None.
  • tests/entrypoints/scale_out/token_in_token_out/test_tokens_logprobs.py: sampled entry carries token_id/rank, per-candidate ids, logprobs=0 still emits the sampled token, sampled token missing from the top-k map keeps the -9999.0 sentinel.
  • tests/entrypoints/scale_out/derender/test_derender.py, test_derender_parity.py: derender chat/completion logprob tests and the coupled-vs-derender parity test now feed the generate-shaped payload; parity assertions on token/bytes are unchanged.
  • tests/entrypoints/scale_out/token_in_token_out/test_generate_stream.py, test_serving_tokens.py: streaming chunks and the sampling-mask test read token_id.
  • Rust: routes::tests::non_stream_raw_generate_returns_token_output_envelope and stream_raw_generate_returns_sse_chunks_and_usage assert integer token_id + rank and a null token/bytes on the wire.

Test Result

Rust, locally:

cargo check -p vllm-server --tests          # clean
cargo test -p vllm-server raw_generate      # 9 passed; 0 failed

The two updated assertions run through the mock engine and confirm the wire shape end to end ("token_id": 33, "rank": 1, token/bytes absent), streaming and non-streaming.

Python, locally (CPU-only box, no GPU):

pytest tests/entrypoints/scale_out/derender/test_generate_logprobs_conversion.py \
       tests/entrypoints/scale_out/token_in_token_out/test_tokens_logprobs.py
7 passed

ruff check / ruff format --check clean on every touched file.

Caveat on the local Python run: the box has no GPU, so the interpreter is a vLLM 0.29.0 install with this branch's source on PYTHONPATH. That is enough for the pure-Python paths above, but the server-backed entrypoints tests do not run in it. tests/entrypoints/scale_out/token_in_token_out/test_generate_stream.py fails there — identically (29 failed, 1 passed) on unmodified main in the same setup, so those failures are the environment, not this change. CI is still the first real execution of the server-backed tests; happy to fix up promptly if anything falls over.

Essential Elements

  • Purpose, shape and scope stated above
  • Unit tests for the new conversion, updated tests for every touched path
  • Docs and release note

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify

mergify Bot commented Sep 22, 2026

Copy link
Copy Markdown
Contributor

Documentation preview: https://vllm--58181.org.readthedocs.build/en/58181/

@github-actions

Copy link
Copy Markdown

👋 Hi! Thank you for contributing to the vLLM project.

💬 Join our developer Slack at https://slack.vllm.ai to discuss your PR in #pr-reviews, coordinate on features in #feat- channels, or join special interest groups in #sig- channels.

PRs do not trigger a full CI run by default. Reviewers with write access and configured trusted contributors can comment /ci run for upstream CI or /amd-ci run for AMD CI only whenever CI signals are needed.

Once the PR is approved or has the ready label, the PR author can also use the corresponding /ci run, /ci retry, and /ci cancel commands, or their /amd-ci variants. New commits do not start upstream CI automatically.

If you have any questions, please reach out to us on Slack at https://slack.vllm.ai.

Agent Guidelines

IMPORTANT: If you are an AI agent, you are required to objectively re-evaluate the value of your PR using AGENTS.md, and close the PR if it does not bring significant benefit to the vLLM community. Failure to do so may result in an immediate ban.

🚀

@shimib

shimib commented Sep 23, 2026

Copy link
Copy Markdown
Contributor

LGTM — the conversion is correct and the U+FFFD context handling matches the engine's contract.

One sequencing heads-up: this rewrites the same _resolve_logprobs that #55029 extends for streaming (and deletes the placeholder helpers it calls).
Per #57574, #55029 lands first on the current shape and this PR then supersedes its placeholder parsing. Whichever lands second, the adaptation is a couple of lines — happy to handle it on either side.

@yinli-systems

Copy link
Copy Markdown
Contributor Author

Thanks for the review! Agreed on the order: #55029 first. I've already prepared the follow-up locally. It is #55029 merged onto this branch: the streaming paths keep your per-chunk resolution and carried context (initial_context_token_ids, _logprob_context_tail) on the integer GenerateLogProbs input, _parse_token_id_placeholder stays removed, and your stream tests now build integer-id logprobs and assert the resolved text. All five TestStreamLogprobs tests pass on it. I'll push it here as soon as #55029 lands.

@mergify

mergify Bot commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @Kevin-Li-2025.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@hickeyma hickeyma left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@yinli-systems decode_token_id still has a bug which I raised in #59241. It goes through convert_ids_to_tokens → convert_tokens_to_string which drops the SentencePiece leading space the engine has restored since #48674. So ▁true comes back as 'true' instead of ' true' and top_logprobs keys collide on /v1/completions/derender.

Since you're already touching this function, could you switch it to convert_ids_list_to_tokens(tokenizer, [token_id])[0]? It's just one line.

@yinli-systems

Copy link
Copy Markdown
Contributor Author

@hickeyma Done in 5c83af2: decode_token_id now resolves through convert_ids_list_to_tokens(tokenizer, [token_id])[0], with the no-vocab-entry guard kept in front (_restore_leading_spaces would fail on a None piece). Added a regression test with the tiny Llama tokenizer: ▁true resolves to " true" and stays a distinct top_logprobs key from true on the completions path. Your repro now matches the engine for every token (including a bare ▁, which used to come back empty). Marked it Fixes #59241.

@mergify

mergify Bot commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @yinli-systems.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 29, 2026
@mergify mergify Bot added the needs-rebase label Oct 3, 2026
@hickeyma

hickeyma commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

@yinli-systems Do you mind resolving the merge conflict?

Yin Li and others added 4 commits October 5, 2026 15:17
…gProbs)

`/inference/v1/generate` is a vLLM-defined token-in/token-out API, but its
output logprobs used the OpenAI shape: the generate server has no tokenizer,
so it wrote every token as a `"token_id:N"` placeholder string that derender
parsed back out. The two frontends also disagreed on `bytes` (Python left it
unset, Rust set it to the bytes of the placeholder itself).

Replace `GenerateResponseChoice.logprobs` and
`GenerateResponseStreamChoice.logprobs` with `GenerateLogProbs`:

    GenerateLogProb:        token_id, logprob, rank
    GenerateLogProbsContent(GenerateLogProb): top_logprobs: list[GenerateLogProb]
    GenerateLogProbs:       content: list[GenerateLogProbsContent] | None

`top_logprobs` is a list in rank order, not a dict, so the ordering is explicit
and the keys stay integers on the wire. `rank` is kept: the engine already
computes it and the Rust frontend's gRPC `CandidateTokenInfo` carries it.
`content=None` is the normal "no per-token candidates requested" state.

Both frontends emit the new shape, streaming and non-streaming. Derender
converts it to `ChatCompletionLogProbs` / `CompletionLogProbs`, filling `token`
and `bytes` from the tokenizer with the existing U+FFFD byte-fallback
correction; that correction needs the preceding sampled token ids as context,
which the integer shape now carries directly. The placeholder parsing
(`_parse_token_id_placeholder`, `resolve_token_id_placeholder`) drops out of
the generate path; `decode_token_id` replaces it.

`prompt_logprobs` on the same response is unchanged, and
`return_tokens_as_token_ids` on `/v1/chat/completions` and `/v1/completions` is
untouched: it is a user-facing OpenAI option that uses `token_id:N` on purpose.

Breaking change for direct generate consumers, replaced in one release per the
discussion in vllm-project#57574. Code that reads only `content[i].logprob` is unaffected;
code that parsed the placeholder reads `content[i].token_id` instead.

Closes vllm-project#57574

Co-authored-by: Claude <noreply@anthropic.com>
Signed-off-by: Yin Li <kevinli.ai.work@gmail.com>
decode_token_id went through convert_ids_to_tokens ->
convert_tokens_to_string, which drops the Metaspace leading space the
engine restores since vllm-project#48674 (`▁true` came back as "true"). On
/v1/completions/derender that made `▁true` and `true` collide as
top_logprobs keys and shifted text_offset by one per spaced token; chat
derender token/bytes did not match /v1/chat/completions.

Resolve through convert_ids_list_to_tokens, the engine's per-token
detokenization. The no-vocab-entry guard stays in front of it.

Fixes vllm-project#59241.

Signed-off-by: Yin Li <kevinli.ai.work@gmail.com>
- GenerateLogProb.logprob is required again, as in vllm-project#57574 and the Rust
  frontend; _create_tokens_logprobs passes -9999.0 explicitly in its one
  fallback branch, so a derender payload missing logprob is rejected
  instead of silently becoming -9999.
- Describe top_logprobs order as the engine's (sampled token first, then
  rank order) in the Python and Rust docs and renderer.md, and drop the
  `content: null` note, which is not a state this PR produces.
- Decode a position's sampled and top-k ids in one batch
  (decode_token_ids): 2 + n tokenizer calls per position instead of 3n.
- Switch the two oversized-logprobs derender tests to token_id payloads
  so they reach the bound checks instead of failing schema validation.
- Tests: sampled token outside the top k comes first, logprob required,
  an unknown top id does not shift its neighbours.

Signed-off-by: Yin Li <kevinli.ai.work@gmail.com>
The Rust generate route mapped every engine entry at a position, so it
returned the sampled token twice when it was also in the top k and never
cut to the requested count (k+1 entries where Python returns k). It now
drops repeated token ids (keeping the sampled entry first) and cuts to
max(logprobs, 1), or all for -1, matching the Python generate server.
ResponseOptions carries the requested count instead of a bool.

Document that rank is the vocabulary rank on every entry and that the
list is not sorted by it.

Signed-off-by: Yin Li <kevinli.ai.work@gmail.com>
@yinli-systems
yinli-systems force-pushed the generate-logprobs-shape branch from de1c0cd to 4a6748f Compare October 5, 2026 11:25
@yinli-systems

Copy link
Copy Markdown
Contributor Author

@hickeyma Rebased onto main after #58588. Tokens mode now carries GenerateLogProbs (integer ids) on GenerateTokensChoice / GenerateTokensStreamChoice; text mode keeps the decoded ChatCompletionLogProbs from #58588 unchanged (that builder is now _create_text_logprobs). I updated token_in_token_out.md, which described tokens-mode logprobs as token_id:N placeholders.

@DarkLight1337

Copy link
Copy Markdown
Member

@claude review

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings marked 🟡 are optional suggestions and need no follow-up push.

Comment thread rust/src/server/src/routes/inference/generate.rs
Comment thread rust/src/server/src/routes/inference/generate.rs Outdated
@mergify mergify Bot removed the needs-rebase label Oct 5, 2026

@hickeyma hickeyma left a comment •

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@yinli-systems Getting there, thanks for your patience. Do you mind addressing the claude bot review comments?

- Rust generate: drop the repeated sampled entry with one HashSet pass
  instead of rescanning the earlier entries for every entry. With
  logprobs=-1 the position holds the whole vocabulary, where the rescan was
  quadratic (~2e10 comparisons per token for a 200k vocab); the dedup of a
  200k-entry position now takes ~0.1 s in a debug test build.
- Both frontends: the engine reports rank 0 for a sampled token whose
  logprob is NaN; send it as rank null, the way the NaN logprob itself is
  clamped, and document it.

Signed-off-by: Yin Li <kevinli.ai.work@gmail.com>
@yinli-systems

Copy link
Copy Markdown
Contributor Author

@hickeyma Both bot findings are addressed in 871a7cd (linear dedup with a full-vocab regression test; rank 0 from a NaN logprob sent as null in both frontends). Thanks!

@hickeyma hickeyma left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, thanks @yinli-systems

@hickeyma

hickeyma commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

/ready

@DarkLight1337

Copy link
Copy Markdown
Member

@claude review

@DarkLight1337 DarkLight1337 added the verified Run pre-commit for new contributors without triggering other tests label Oct 6, 2026

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline NaN-logprob finding, I also checked whether the now-strict GenerateLogProbs/GenerateTokensChoice schema (protocol.py:404) would break /inference/v1/generate parsing during a mixed-version disaggregated rollout (old generate server emitting the old ChatCompletionLogProbs shape against a new derender, or vice versa) — that is exactly the breaking change the PR's release note already documents, not a new defect introduced here, so I'm not flagging it separately.

Extended reasoning...

Verified the confirmed NaN-logprob finding independently: _create_tokens_logprobs (serving.py) computes max(step_token.logprob, -9999.0) with the possibly-NaN value as the first argument, and Python's max() keeps the first operand when a comparison against NaN is involved, so the sentinel never applies and a literal NaN reaches the JSON response; the PR's own new test only asserts rank is None, not the logprob value, so this slipped through. Checked git history: the two findings from my prior review (Rust O(n^2) dedup, rank=0 on NaN) were addressed in commit 871a7cd, whose message claims the logprob was "already clamped" — that belief is what left this new bug unaddressed. The mixed-version 422 candidate was examined and ruled out as the documented breaking-change behavior itself, not a separate bug.

Comment thread vllm/entrypoints/scale_out/token_in_token_out/serving.py Outdated
max(nan, -9999.0) is nan in Python, so the clamps in _create_tokens_logprobs
and the text-mode builder let a NaN logprob through: JSONResponse rejected it
(500, non-streaming) and streaming wrote "logprob": null, which derender then
refused as logprob is required. The Rust frontend already gets -9999.0 from
f32::max. Add _clamp_logprob (NaN -> -9999.0, else max(x, -9999.0)) and use
it at every call site, tokens and text mode. Tests assert -9999.0 for the
content and top_logprobs entries and that the result serializes as strict
JSON.

Signed-off-by: Yin Li <kevinli.ai.work@gmail.com>
@DarkLight1337

Copy link
Copy Markdown
Member

@claude review

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

@DarkLight1337
DarkLight1337 enabled auto-merge (squash) October 6, 2026 15:48
@DarkLight1337

Copy link
Copy Markdown
Member

/ci run

@github-actions github-actions Bot added the ready ONLY add when PR is ready to merge/full CI is needed label Oct 6, 2026
@github-actions

github-actions Bot commented Oct 6, 2026

Copy link
Copy Markdown

✅ Triggered Buildkite CI #93120 for commit a62aab00c78b.

@vllm-agent

Copy link
Copy Markdown
Contributor

CI selector (shadow): no narrower answer, every step runs: 193 test steps (269 jobs) instead of 45 (58 jobs)

Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works.

Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.

steps (jobs) Today's rules Selector Would skip Would add
NVIDIA, CPU and others 45 (58) 193 (269) 0 (0) 148 (211)
AMD mirrors 40 (48) 163 (227) 0 (0) 123 (179)

Why: rust: rust/src/server/src/routes/inference/generate.rs: root crate bucket; 18 gate-env steps + 12 hardware-image steps; + the PyO3 bridge claim (vllm/tool_parsers/rust_tool_parser.py)

Selector would run (193)
  • amd-fp8-moe-kernels-mi355
  • amd-kernels-mi355
  • amd-lm-eval-small-models-harness
  • amd-native-quantization-kernels-mi355
  • amd-qwen3-next-mtp-async-eplb-accuracy
  • arm-cpu-test ×3
  • ascend-npu-test
  • async-engine-inputs-utils-worker
  • basic-correctness ×2
  • basic-correctness-cpu-offload
  • basic-correctness-cumem
  • basic-correctness-prefetch-offload
  • basic-correctness-sleep-mode
  • basic-models-test-other-cpu
  • basic-models-tests-extra-initialization ×14
  • basic-models-tests-initialization
  • basic-models-tests-other
  • batch-invariance-b200
  • batch-invariance-h100
  • benchmarks-cli-test
  • cpu-distributed-tests-dp-tp
  • cpu-distributed-tests-pp-tp
  • cpu-kernel-tests ×2
  • cpu-language-generation-and-pooling-model-tests ×3
  • cpu-multi-modal-model-tests-n ×4
  • cpu-multimodal-config
  • cpu-params-env-tokenizers-parser
  • cpu-quantization-model-tests
  • cpu-qwen2-5-vl-multimodal-tests
  • cpu-reasoning-renderers
  • cpu-spec-decode-tests
  • cpu-tool-parsers
  • crcr-report
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus
  • cudagraph
  • deepseek-v4-kernel-test-b200
  • deepseek-v4-kernel-test-h100
  • distributed-comm-ops
  • distributed-compile-comm-4-gpus
  • distributed-compile-rpc-tests-2-gpus
  • distributed-compile-unit-tests-2xh100
  • distributed-dp-tests-2-gpus
  • distributed-dp-tests-4-gpus
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus
  • distributed-model-tests-2-gpus ×3
  • distributed-mooncakeconnector-pd-accuracy-4-gpus
  • distributed-nixlconnector-pd-accuracy-4-gpus
  • distributed-tests-8xh100
  • distributed-torchrun-examples-4-gpus
  • distributed-torchrun-shutdown-tests-2-gpus
  • docker-build-metadata
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus
  • e2e-core-1-gpu
  • e2e-core-large-memory
  • e2e-scheduling-1-gpu
  • e2e-scheduling-accuracy-1-gpu
  • elastic-ep-scaling-test
  • engine
  • engine-1-gpu
  • entrypoints-integration-api-server ×4
  • entrypoints-integration-api-server-generate
  • entrypoints-integration-api-server-openai-chat_completion
  • entrypoints-integration-api-server-openai-completion
  • entrypoints-integration-llm
  • entrypoints-integration-multimodal
  • entrypoints-integration-pooling
  • entrypoints-integration-responses-api
  • entrypoints-integration-speech_to_text
  • entrypoints-unit-tests
  • eplb-algorithm
  • eplb-execution
  • examples
  • extract-hidden-states-integration
  • extract-hidden-states-integration-2-gpus
  • fault-tolerance-e2e-2xh100
  • fusion-and-compile-unit-tests-2xb200
  • fusion-e2e-config-sweep-h100
  • fusion-e2e-quick-h100
  • fusion-e2e-tp2-ar-rms-config-sweep-h100
  • fusion-e2e-tp2-ar-rms-dynamo-partition-amd
  • fusion-e2e-tp2-ar-rms-inductor-partition-amd
  • fusion-e2e-tp2-asynctp-config-sweep-h100
  • fusion-e2e-tp2-b200
  • fusion-e2e-tp2-quick-h100
  • gemm-rs-ar-2xb200
  • glm5next-unit-tests
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus
  • inkling-unit-tests-b200
  • intel-hpu-test
  • jit-monitor-no-runtime-jit
  • kernels-attention-diffkv-test-h100
  • kernels-attention-test ×7
  • kernels-b200 ×3
  • kernels-core-operation-test ×3
  • kernels-deepgemm-test-h100
  • kernels-fla-ops-test-b200
  • kernels-flashmla-test-h100
  • kernels-fusedmoe-layer-test-2-b200s
  • kernels-fusedmoe-layer-test-2-h100s
  • kernels-helion-test ×5
  • kernels-mamba-test
  • kernels-mhc-test-b200
  • kernels-minimax-reduce-rms-test-2-gpus
  • kernels-moe-test ×5
  • kernels-quantization-test ×6
  • kernels-root-misc-test-b200
  • kimi-k3-prefix-cache-4xb200 ×2
  • kimi-k3-unit-tests-b200
  • kv-offload-large
  • kv-offload-medium
  • kv-offload-small
  • language-models-tests-extra-standard ×2
  • language-models-tests-granite-l4-compatibility
  • language-models-tests-hybrid ×2
  • language-models-tests-standard
  • lm-eval-dspark-watermark-2xh100
  • lm-eval-small-models
  • lm-eval-turboquant-k3v4nc
  • lm-eval-turboquant-k8v4
  • lm-eval-turboquant-t3nc
  • lm-eval-turboquant-t4nc
  • lm-eval-watermarking
  • lora ×4
  • lora-tp-distributed ×4
  • metrics-tracing-2-gpus
  • model-executor
  • mooncake-ec-tcp-e2e-2-gpus
  • mrcr-eval-small-models
  • multi-modal-accuracy-eval-small-models
  • multi-modal-models-standard-1-qwen2
  • multi-modal-models-standard-2-qwen3-gemma
  • multi-modal-models-standard-3-llava-qwen2-vl
  • multi-modal-models-standard-4-other-whisper
  • multi-modal-processor ×4
  • multi-modal-processor-cpu ×4
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus
  • nixlconnector-pd-edge-cases-2-gpus
  • openai-api-correctness
  • pipeline-context-parallelism-4-gpus
  • platform-tests
  • plugin-tests-2-gpus
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus
  • pytorch-compilation-dynamic-shapes
  • pytorch-compilation-passes-unit-tests
  • pytorch-compilation-unit-tests
  • pytorch-compilation-unit-tests-h100
  • pytorch-fullgraph-cudagraph-l4-compatibility
  • pytorch-fullgraph-test
  • pytorch-nightly-dependency-override-check
  • quantization ×4
  • quantized-fusions
  • quantized-models-test
  • quantized-moe-test-b200
  • qwen4-exp-unit-tests
  • qwen4-exp-unit-tests-cpu
  • ray-dependency-compatibility-check
  • rayexecutorv2-4-gpus
  • regression
  • replayssm-e2e
  • rust-frontend-cargo-style-clippy
  • rust-frontend-cargo-tests
  • rust-frontend-core-correctness
  • rust-frontend-distributed
  • rust-frontend-openai-coverage
  • rust-frontend-serve-admin-coverage
  • rust-frontend-tool-use
  • samplers-multimodal-beam-search
  • samplers-test
  • scale-out-ec-e2e-2-gpus
  • sharded-rdt-weight-transfer
  • spec-decode-draft-model ×4
  • spec-decode-eagle-1-deepseek-qwen
  • spec-decode-eagle-2-llama3-qwen-vl-other
  • spec-decode-mtp-deepseek-mimo
  • spec-decode-mtp-gemma4
  • spec-decode-mtp-qwen3-5
  • spec-decode-ngram-suffix
  • spec-decode-speculators
  • torch-stable-abi-audit
  • v1-attention-b200 ×2
  • v1-attention-h100-mi300 ×2
  • v1-core
  • v1-executor-worker
  • v1-kv-connectors ×4
  • v1-kv-offload
  • v1-logits-oracle
  • v1-metrics-lmeval
  • v1-others-cpu
  • v1-sample
  • v1-spec-decode
  • vllm-ir-tests
Would skip (today's rules run them) (0)

none

Would add (today's rules do not run them) (148)
  • amd-fp8-moe-kernels-mi355 (code map)
  • amd-kernels-mi355 (code map)
  • amd-lm-eval-small-models-harness (code map)
  • amd-native-quantization-kernels-mi355 (code map)
  • amd-qwen3-next-mtp-async-eplb-accuracy (code map)
  • arm-cpu-test ×3 (code map)
  • ascend-npu-test (code map)
  • async-engine-inputs-utils-worker (code map)
  • basic-models-tests-extra-initialization ×14 (code map)
  • batch-invariance-b200 (code map)
  • batch-invariance-h100 (code map)
  • cpu-distributed-tests-dp-tp (code map)
  • cpu-distributed-tests-pp-tp (code map)
  • cpu-kernel-tests ×2 (code map)
  • cpu-multi-modal-model-tests-n ×4 (code map)
  • cpu-quantization-model-tests (code map)
  • cpu-qwen2-5-vl-multimodal-tests (code map)
  • cpu-spec-decode-tests (code map)
  • crcr-report (code map)
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • cudagraph (code map)
  • deepseek-v4-kernel-test-b200 (code map)
  • deepseek-v4-kernel-test-h100 (code map)
  • distributed-comm-ops (code map)
  • distributed-compile-comm-4-gpus (code map)
  • distributed-compile-rpc-tests-2-gpus (code map)
  • distributed-compile-unit-tests-2xh100 (code map)
  • distributed-dp-tests-2-gpus (code map)
  • distributed-dp-tests-4-gpus (code map)
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus (code map)
  • distributed-model-tests-2-gpus ×3 (code map)
  • distributed-mooncakeconnector-pd-accuracy-4-gpus (code map)
  • distributed-nixlconnector-pd-accuracy-4-gpus (code map)
  • distributed-tests-8xh100 (code map)
  • distributed-torchrun-examples-4-gpus (code map)
  • distributed-torchrun-shutdown-tests-2-gpus (code map)
  • docker-build-metadata (code map)
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • e2e-core-1-gpu (code map)
  • e2e-core-large-memory (code map)
  • e2e-scheduling-1-gpu (code map)
  • e2e-scheduling-accuracy-1-gpu (code map)
  • elastic-ep-scaling-test (code map)
  • engine (code map)
  • engine-1-gpu (code map)
  • eplb-algorithm (code map)
  • eplb-execution (code map)
  • extract-hidden-states-integration (code map)
  • extract-hidden-states-integration-2-gpus (code map)
  • fault-tolerance-e2e-2xh100 (code map)
  • fusion-and-compile-unit-tests-2xb200 (code map)
  • fusion-e2e-config-sweep-h100 (code map)
  • fusion-e2e-quick-h100 (code map)
  • fusion-e2e-tp2-ar-rms-config-sweep-h100 (code map)
  • fusion-e2e-tp2-ar-rms-dynamo-partition-amd (code map)
  • fusion-e2e-tp2-ar-rms-inductor-partition-amd (code map)
  • fusion-e2e-tp2-asynctp-config-sweep-h100 (code map)
  • fusion-e2e-tp2-b200 (code map)
  • fusion-e2e-tp2-quick-h100 (code map)
  • gemm-rs-ar-2xb200 (code map)
  • glm5next-unit-tests (code map)
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus (code map)
  • inkling-unit-tests-b200 (code map)
  • intel-hpu-test (code map)
  • jit-monitor-no-runtime-jit (code map)
  • kernels-attention-diffkv-test-h100 (code map)
  • kernels-attention-test ×7 (code map)
  • kernels-b200 ×3 (code map)
  • kernels-core-operation-test ×3 (code map)
  • kernels-deepgemm-test-h100 (code map)
  • kernels-flashmla-test-h100 (code map)
  • kernels-fusedmoe-layer-test-2-b200s (code map)
  • kernels-fusedmoe-layer-test-2-h100s (code map)
  • kernels-helion-test ×5 (code map)
  • kernels-mamba-test (code map)
  • kernels-minimax-reduce-rms-test-2-gpus (code map)
  • kernels-moe-test ×5 (code map)
  • kernels-quantization-test ×6 (code map)
  • kimi-k3-prefix-cache-4xb200 ×2 (code map)
  • kimi-k3-unit-tests-b200 (code map)
  • kv-offload-large (code map)
  • kv-offload-medium (code map)
  • kv-offload-small (code map)
  • language-models-tests-extra-standard ×2 (code map)
  • lm-eval-dspark-watermark-2xh100 (code map)
  • lm-eval-small-models (code map)
  • lm-eval-turboquant-k3v4nc (code map)
  • lm-eval-turboquant-k8v4 (code map)
  • lm-eval-turboquant-t3nc (code map)
  • lm-eval-turboquant-t4nc (code map)
  • lm-eval-watermarking (code map)
  • lora ×4 (code map)
  • lora-tp-distributed ×4 (code map)
  • metrics-tracing-2-gpus (code map)
  • model-executor (code map)
  • mooncake-ec-tcp-e2e-2-gpus (code map)
  • mrcr-eval-small-models (code map)
  • multi-modal-accuracy-eval-small-models (code map)
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus (code map)
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus (code map)
  • nixlconnector-pd-edge-cases-2-gpus (code map)
  • openai-api-correctness (code map)
  • pipeline-context-parallelism-4-gpus (code map)
  • platform-tests (code map)
  • plugin-tests-2-gpus (code map)
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus (code map)
  • pytorch-compilation-dynamic-shapes (code map)
  • pytorch-compilation-passes-unit-tests (code map)
  • pytorch-compilation-unit-tests (code map)
  • pytorch-compilation-unit-tests-h100 (code map)
  • pytorch-fullgraph-cudagraph-l4-compatibility (code map)
  • pytorch-fullgraph-test (code map)
  • pytorch-nightly-dependency-override-check (code map)
  • quantization ×4 (code map)
  • quantized-fusions (code map)
  • quantized-models-test (code map)
  • quantized-moe-test-b200 (code map)
  • qwen4-exp-unit-tests (code map)
  • qwen4-exp-unit-tests-cpu (code map)
  • ray-dependency-compatibility-check (code map)
  • rayexecutorv2-4-gpus (code map)
  • regression (code map)
  • replayssm-e2e (code map)
  • samplers-multimodal-beam-search (code map)
  • samplers-test (code map)
  • sharded-rdt-weight-transfer (code map)
  • spec-decode-draft-model ×4 (code map)
  • spec-decode-eagle-1-deepseek-qwen (code map)
  • spec-decode-eagle-2-llama3-qwen-vl-other (code map)
  • spec-decode-mtp-deepseek-mimo (code map)
  • spec-decode-mtp-gemma4 (code map)
  • spec-decode-mtp-qwen3-5 (code map)
  • spec-decode-ngram-suffix (code map)
  • spec-decode-speculators (code map)
  • torch-stable-abi-audit (code map)
  • v1-attention-b200 ×2 (code map)
  • v1-attention-h100-mi300 ×2 (code map)
  • v1-core (code map)
  • v1-executor-worker (code map)
  • v1-kv-connectors ×4 (code map)
  • v1-kv-offload (code map)
  • v1-logits-oracle (code map)
  • v1-metrics-lmeval (code map)
  • v1-others-cpu (code map)
  • v1-sample (code map)
  • v1-spec-decode (code map)
  • vllm-ir-tests (code map)
AMD mirrors: would skip (0)

none

AMD mirrors: would add (123)
  • async-engine-inputs-utils-worker (code map)
  • basic-models-tests-extra-initialization ×14 (code map)
  • batch-invariance-h100 (code map)
  • crosslayer-kv-layout-distributed-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • cudagraph (code map)
  • deepseek-v4-kernel-test-b200 (code map)
  • deepseek-v4-kernel-test-h100 (code map)
  • distributed-comm-ops (code map)
  • distributed-compile-comm-4-gpus (code map)
  • distributed-compile-rpc-tests-2-gpus (code map)
  • distributed-compile-unit-tests-2xh100 (code map)
  • distributed-dp-tests-2-gpus (code map)
  • distributed-dp-tests-4-gpus (code map)
  • distributed-flashinfer-nixlconnector-pd-accuracy-4-gpus (code map)
  • distributed-model-tests-2-gpus ×3 (code map)
  • distributed-mooncakeconnector-pd-accuracy-4-gpus (code map)
  • distributed-nixlconnector-pd-accuracy-4-gpus (code map)
  • distributed-tests-8xh100 (code map)
  • distributed-torchrun-examples-4-gpus (code map)
  • distributed-torchrun-shutdown-tests-2-gpus (code map)
  • docker-build-metadata (code map)
  • dp-ep-distributed-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • e2e-core-1-gpu (code map)
  • e2e-core-large-memory (code map)
  • e2e-scheduling-1-gpu (code map)
  • e2e-scheduling-accuracy-1-gpu (code map)
  • engine (code map)
  • engine-1-gpu (code map)
  • eplb-algorithm (code map)
  • eplb-execution (code map)
  • extract-hidden-states-integration (code map)
  • extract-hidden-states-integration-2-gpus (code map)
  • fault-tolerance-e2e-2xh100 (code map)
  • fusion-and-compile-unit-tests-2xb200 (code map)
  • fusion-e2e-config-sweep-h100 (code map)
  • fusion-e2e-quick-h100 (code map)
  • fusion-e2e-tp2-ar-rms-config-sweep-h100 (code map)
  • fusion-e2e-tp2-asynctp-config-sweep-h100 (code map)
  • fusion-e2e-tp2-b200 (code map)
  • fusion-e2e-tp2-quick-h100 (code map)
  • gemm-rs-ar-2xb200 (code map)
  • glm5next-unit-tests (code map)
  • hybrid-ssm-nixlconnector-pd-accuracy-tests-4-gpus (code map)
  • hybrid-ssm-nixlconnector-pd-prefix-cache-2-gpus (code map)
  • inkling-unit-tests-b200 (code map)
  • jit-monitor-no-runtime-jit (code map)
  • kernels-attention-diffkv-test-h100 (code map)
  • kernels-attention-test ×7 (code map)
  • kernels-b200 ×3 (code map)
  • kernels-core-operation-test ×3 (code map)
  • kernels-flashmla-test-h100 (code map)
  • kernels-fp4-moe-test-b200 (code map)
  • kernels-fusedmoe-layer-test-2-b200s (code map)
  • kernels-fusedmoe-layer-test-2-h100s (code map)
  • kernels-helion-test ×5 (code map)
  • kernels-mamba-test (code map)
  • kernels-minimax-reduce-rms-test-2-gpus (code map)
  • kernels-moe-test ×5 (code map)
  • kernels-quantization-test ×6 (code map)
  • kimi-k3-unit-tests-b200 (code map)
  • kv-offload-large (code map)
  • kv-offload-medium (code map)
  • kv-offload-small (code map)
  • language-models-tests-extra-standard ×2 (code map)
  • lm-eval-dspark-watermark-2xh100 (code map)
  • lm-eval-small-models (code map)
  • lm-eval-turboquant-k3v4nc (code map)
  • lm-eval-turboquant-k8v4 (code map)
  • lm-eval-turboquant-t3nc (code map)
  • lm-eval-turboquant-t4nc (code map)
  • lm-eval-watermarking (code map)
  • lora ×4 (code map)
  • lora-tp-distributed ×4 (code map)
  • metrics-tracing-2-gpus (code map)
  • model-executor (code map)
  • mooncake-ec-tcp-e2e-2-gpus (code map)
  • mrcr-eval-small-models (code map)
  • multi-modal-accuracy-eval-small-models (code map)
  • multiconnector-nixl-offloading-pd-accuracy-2-gpus (code map)
  • multiconnector-nixl-offloading-pd-edge-cases-2-gpus (code map)
  • nixlconnector-pd-edge-cases-2-gpus (code map)
  • nixlconnector-pd-spec-decode-acceptance-2-gpus (code map)
  • openai-api-correctness (code map)
  • pipeline-context-parallelism-4-gpus (code map)
  • plugin-tests-2-gpus (code map)
  • push-nixlconnector-pp-prefill-pd-accuracy-4-gpus (code map)
  • pytorch-compilation-dynamic-shapes (code map)
  • pytorch-compilation-unit-tests (code map)
  • pytorch-compilation-unit-tests-h100 (code map)
  • pytorch-fullgraph-cudagraph-l4-compatibility (code map)
  • pytorch-fullgraph-test (code map)
  • pytorch-nightly-dependency-override-check (code map)
  • quantization ×4 (code map)
  • quantized-fusions (code map)
  • quantized-models-test (code map)
  • quantized-moe-test-b200 (code map)
  • qwen4-exp-unit-tests (code map)
  • ray-dependency-compatibility-check (code map)
  • rayexecutorv2-4-gpus (code map)
  • regression (code map)
  • samplers-multimodal-beam-search (code map)
  • samplers-test (code map)
  • sharded-rdt-weight-transfer (code map)
  • spec-decode-draft-model ×4 (code map)
  • spec-decode-eagle-1-deepseek-qwen (code map)
  • spec-decode-eagle-2-llama3-qwen-vl-other (code map)
  • spec-decode-mtp-deepseek-mimo (code map)
  • spec-decode-mtp-gemma4 (code map)
  • spec-decode-mtp-qwen3-5 (code map)
  • spec-decode-ngram-suffix (code map)
  • spec-decode-speculators (code map)
  • torch-stable-abi-audit (code map)
  • v1-attention-b200 ×2 (code map)
  • v1-attention-h100-mi300 ×2 (code map)
  • v1-core (code map)
  • v1-executor-worker (code map)
  • v1-kv-connectors ×4 (code map)
  • v1-kv-offload (code map)
  • v1-logits-oracle (code map)
  • v1-metrics-lmeval (code map)
  • v1-sample (code map)
  • v1-spec-decode (code map)
  • vllm-ir-tests (code map)

18 changed files · base aad4579bbe · head a62aab00c7 · Python record: build 93039 at 1e5d0ea888 · kernel record: table 1e5d0ea (build 93039), map 1e5d0ea · not counted: 10 build steps, 5 A100 steps the generator no longer emits, 138 optional steps the selector would also run

@DarkLight1337
DarkLight1337 merged commit 7b665f7 into vllm-project:main Oct 6, 2026
107 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation frontend ready ONLY add when PR is ready to merge/full CI is needed rust verified Run pre-commit for new contributors without triggering other tests

Projects

None yet

6 participants