[Model] Support top_k and top_p sampling for DiffusionGemma - #45429
Merged
Isotr0py merged 14 commits intoJul 26, 2026
Conversation
guan404ming
force-pushed
the
feat/diffusion-top-k-top-p
branch
2 times, most recently
from
June 15, 2026 04:32
e0b6a58 to
e53d6f3
Compare
hsjlyj
pushed a commit
to hsjlyj/vllm
that referenced
this pull request
Jun 15, 2026
…allback Why not duplicate: open PRs vllm-project#45429 and vllm-project#45417 address sampling and generation config, not the startup crash 'Argument input_ids not found in the forward method of DiffusionGemmaDecoderModel'. Duplicate checks run with 'gh pr list --repo vllm-project/vllm --state open --search "DiffusionGemma input_ids"' and 'gh pr list --repo vllm-project/vllm --state open --search "DiffusionGemma transformers backend"', both empty. Tests run: python3 -m py_compile vllm/model_executor/models/transformers/base.py. Also reproduced the user-facing failure on vLLM 0.23.0 before the patch via vllm serve against RedHatAI/diffusiongemma-26B-A4B-it-FP8-dynamic; root cause was the torch-compile decoration path assuming a decoder forward(input_ids=...) signature for the Transformers fallback. AI assistance used. Co-authored-by: Claude <noreply@anthropic.com> Signed-off-by: hsjlyj <173873397+hsjlyj@users.noreply.github.com>
Contributor
Author
|
Hi @LucasWilkinson, could you help take a look, thanks! |
guan404ming
force-pushed
the
feat/diffusion-top-k-top-p
branch
2 times, most recently
from
June 23, 2026 15:57
ee488c0 to
047369e
Compare
guan404ming
force-pushed
the
feat/diffusion-top-k-top-p
branch
4 times, most recently
from
July 2, 2026 13:28
81836b9 to
e703dff
Compare
guan404ming
force-pushed
the
feat/diffusion-top-k-top-p
branch
2 times, most recently
from
July 7, 2026 11:31
ab257e5 to
bfb750a
Compare
Contributor
Author
|
Hi @Isotr0py could you help take a look at this one, thanks! |
guan404ming
force-pushed
the
feat/diffusion-top-k-top-p
branch
2 times, most recently
from
July 16, 2026 13:07
823b7e0 to
7c1a121
Compare
Contributor
Author
|
Hello @Isotr0py just gentle ping could you help take look at this, thanks! |
guan404ming
force-pushed
the
feat/diffusion-top-k-top-p
branch
2 times, most recently
from
July 20, 2026 10:39
035dd70 to
459de27
Compare
Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com>
guan404ming
force-pushed
the
feat/diffusion-top-k-top-p
branch
from
July 21, 2026 06:00
459de27 to
ccad439
Compare
Isotr0py
enabled auto-merge (squash)
July 22, 2026 12:46
Contributor
Author
|
Never mind, thanks! |
edwinlim0919
pushed a commit
to chaeminlim-mb/vllm
that referenced
this pull request
Jul 29, 2026
…ject#45429) Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com>
aoshen02
pushed a commit
to zllion/vllm
that referenced
this pull request
Aug 1, 2026
…ject#45429) Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com> Signed-off-by: aoshen02 <aoshen@inferact.ai>
pranavthakur0-0
pushed a commit
to pranavthakur0-0/vllm
that referenced
this pull request
Aug 4, 2026
…ject#45429) Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com>
itej89
pushed a commit
to itej89/vllm
that referenced
this pull request
Aug 4, 2026
…ject#45429) Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com> Signed-off-by: Tej Kiran <kiran.tej@amd.com>
aditi-amd
pushed a commit
to aditi-amd/vllm
that referenced
this pull request
Aug 4, 2026
…ject#45429) Signed-off-by: Guan-Ming (Wesley) Chiu <105915352+guan404ming@users.noreply.github.com> Signed-off-by: root <root@smci355-ccs-aus-m02-09.cs-aus.dcgpu>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Support per-request
top_k/top_pfor DiffusionGemma by filtering logits before the compiled denoise step, mirroring the AR sampler. Masked tokens become-inf; committed argmax (top-1) unaffected. Padding switched tomasked_fill_to avoid-inf * 0 = NaNon truncated canvases.Related: #45163. Validation relaxation lives in the companion PR.
Test Plan
GPU behavioral check: per-request filtering, default fast path, truncated-canvas padding with
-inf.Test Result
All passed: k finite logits per filtered row, top-1 preserved, no NaN.
Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.