Skip to content

[Bugfix][Qwen-Image] Restore RotaryEmbedding CUDA RoPE for Diffusers e2e - #7513

Merged
yenuo26 merged 6 commits into
vllm-project:mainfrom
NumberWan:fix/qwen-image-cuda-rope-7494
Sep 16, 2026
Merged

yenuo26 merged 6 commits into
vllm-project:mainfrom
NumberWan:fix/qwen-image-cuda-rope-7494

Conversation

@NumberWan

@NumberWan NumberWan commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Local L20X, same nightly recipe (FA3 hub, --fa-deterministic, 512² / 20-step / seed 42). Absolute numbers are not H100; this box is stable:

tree SSIM PSNR result
#7185 038c9b948f 0.987344 34.317 dB PASS
#7185 + only the #7230 CUDA RoPE hunk 0.948505 25.981 dB FAIL (PSNR < 27)
revert CUDA branch to RotaryEmbedding 0.987344 34.317 dB PASS

25.98 is the same band as H100 26.13. Not lowering PSNR_THRESHOLD.

Test plan

  • Local L20X: tests/e2e/accuracy/test_qwen_image.py::test_qwen_image_matches_diffusers after revert → SSIM 0.987344 / PSNR 34.317
  • Nightly H100 (Buildkite #15339, nightly-test): Diffusion X2I(&A&T) · Accuracy Test
    • test_qwen_image_matches_diffusers → SSIM 0.978375 / PSNR 31.83 (then raised SSIM_THRESHOLD to 0.97; PSNR_THRESHOLD stays 27)
    • test_diffusers_backend_t2i_matches_diffusers passed
  • Ready (path-triggered, needs ready): Diffusion · Qwen Image Test
    pytest -s -v tests/e2e/online_serving/test_qwen_image.py -m 'core_model' --run-level 'core_model'
  • Merge (path-triggered, needs merge-test): Diffusion · Qwen Image Test
    pytest -s -v tests/e2e/online_serving/test_qwen_image.py -m 'advanced_model and cuda' --run-level 'advanced_model'
    Ready/merge serving jobs did not run on #15339 (nightly pipeline only). They should run when those labels are applied; this change only restores eager CUDA RoPE.

Test Result

  • Local: SSIM 0.987344 / PSNR 34.317 vs the [Rebase] Rebase to vLLM 0.29.0 #7230 helper regression (SSIM 0.948505 / PSNR 25.981).
  • H100 #15339 accuracy: SSIM 0.978375 / PSNR 31.83.
  • Ready/merge Qwen Image serving: not executed on #15339; pending ready / merge-test CI.

@NumberWan
NumberWan marked this pull request as ready for review September 14, 2026 09:22
@NumberWan
NumberWan requested a review from wtomin as a code owner September 14, 2026 09:22
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/offloader.md, docs/design/module/diffusion/diffusion_model_integration.md, docs/design/module/diffusion/index.md.

Module owners: @Bounty-hunter @fhfuih @wtomin

Routing: @Bounty-hunter via module of the changed files, CODEOWNERS; @fhfuih via module of the changed files, CODEOWNERS; @wtomin via module of the changed files, CODEOWNERS

@NumberWan, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 14, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 1ec6cc77c6ad produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@hsliuustc0106 hsliuustc0106 added bug Something isn't working diffusion codes related to diffusion models labels Sep 15, 2026
@NumberWan

Copy link
Copy Markdown
Contributor Author

@yenuo26 PTAL

@fhfuih fhfuih left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I do remember having tested this code when reviewing #6110. The CUDAs-specific branch is introduced when Rebasing onto 0.29.0, and this PR actually reverts it. When testing it back then, I remember CUDA does pass the accuracy test. Together with the latest test results in the PR description, this PR looks good to me.

QwenImageTransformerBlock. That matches the Diffusers helper in unit
tests but drops Omni vs Diffusers pipeline PSNR below 27. Use the
pre-vllm-project#7230 RotaryEmbedding path in forward again; keep the helper for CPU tests.

Fixes vllm-project#7494

Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Forward no longer calls _apply_qwen_image_rotary_emb. Remove the
function and the CPU test that only covered it.

Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Keep vllm-project#5931 fused QK-RoPE. Eager CUDA fallback uses RotaryEmbedding
like other devices. Tests keep a local FP32 reference for the fused kernel.

Signed-off-by: NumberWan <wantszkin2003@gmail.com>
@yenuo26

yenuo26 commented Sep 16, 2026

Copy link
Copy Markdown
Collaborator

please fix pre-commit

Signed-off-by: NumberWan <wantszkin2003@gmail.com>
@NumberWan

Copy link
Copy Markdown
Contributor Author

please fix pre-commit

@yenuo26 Thanks — pre-commit is fixed on this branch (4ba9d93).
The GitHub pre-commit check has not gone green yet because the workflow is still waiting for a runner.

@yenuo26 yenuo26 added nightly-test label to trigger buildkite nightly test CI cuda-test Used to trigger vllm-omni cuda CI separately. labels Sep 16, 2026
Nightly test_qwen_image_matches_diffusers on Buildkite #15339 scored
SSIM 0.978375. Raise SSIM_THRESHOLD so a drop is visible.

Signed-off-by: NumberWan <wantszkin2003@gmail.com>
@yenuo26 yenuo26 removed nightly-test label to trigger buildkite nightly test CI cuda-test Used to trigger vllm-omni cuda CI separately. labels Sep 16, 2026
@yenuo26

yenuo26 commented Sep 16, 2026

Copy link
Copy Markdown
Collaborator

@vllm-omni-review-bot

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Direct under experiment vllm-omni-strict-5050-20260829.

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

PR description

This bugfix restores Qwen-Image’s non-fused CUDA RoPE path to shared RotaryEmbedding (BF16/activation-dtype cos/sin) after #7230’s FP32 complex multiply helper regressed Omni↔Diffusers pipeline PSNR below the nightly gate. The dead _apply_qwen_image_rotary_emb helper and its unit test are removed; fused Q/K path and e2e thresholds are left intact. User-visible effect is recovering Diffusers-matching image quality on the CUDA eager RoPE path without loosening PSNR_THRESHOLD.

Change flow

flowchart TD
  A["[EXISTING] QwenImageCrossAttention.forward<br/>vid/txt freqs + qk_norm"]:::existing
  B["[CHANGED] _qwen_image_qk_norm_rope<br/>eager: RotaryEmbedding on all devices"]:::changed
  C["[REMOVED] _apply_qwen_image_rotary_emb<br/>CUDA FP32 complex multiply"]:::removed
  D["[CHANGED] benchmark + fused fallback tests<br/>native/eager pin RotaryEmbedding"]:::changed
  E["[EXISTING] test_qwen_image_matches_diffusers<br/>SSIM/PSNR vs Diffusers pipeline"]:::existing
  A --> B
  B -.-> C
  B --> D
  B --> E
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

See inline comments below.

Comment thread vllm_omni/diffusion/models/qwen_image/qwen_image_transformer.py
@yenuo26 yenuo26 added the ready label to trigger buildkite CI label Sep 16, 2026

@yenuo26 yenuo26 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@yenuo26 yenuo26 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@yenuo26
yenuo26 enabled auto-merge (squash) September 16, 2026 14:02
@yenuo26
yenuo26 merged commit 81b5be2 into vllm-project:main Sep 16, 2026
7 of 9 checks passed
yuweih205 added a commit to yuweih205/vllm-omni that referenced this pull request Sep 17, 2026
…into one

main (vllm-project#5931/vllm-project#7513) already fuses Q/K RMSNorm + RoPE per stream and then
concatenates text and image into the joint sequence. Replace the two launches
and the three cats with a single fused_joint_qkv_norm_rope call that writes the
joint [B, S_txt+S_img, H, D] Q/K/V attention consumes, with no intermediate
copies. The table carries the same fp32 coefficients the per-stream path
builds, so the joint launch is bitwise equal to it. Any sequence parallelism
keeps the per-stream chain: vid_freqs is sharded while txt_freqs is not.

Signed-off-by: HuangYuwei <yuweih205@gmail.com>
yuweih205 added a commit to yuweih205/vllm-omni that referenced this pull request Sep 17, 2026
…into one

main (vllm-project#5931/vllm-project#7513) already fuses Q/K RMSNorm + RoPE per stream and then
concatenates text and image into the joint sequence. Replace the two launches
and the three cats with a single fused_joint_qkv_norm_rope call that writes the
joint [B, S_txt+S_img, H, D] Q/K/V attention consumes, with no intermediate
copies. The table carries the same fp32 coefficients the per-stream path
builds, so the joint launch is bitwise equal to it. Any sequence parallelism
keeps the per-stream chain: vid_freqs is sharded while txt_freqs is not.

Signed-off-by: HuangYuwei <yuweih205@gmail.com>
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
…e2e (vllm-project#7513)

Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…e2e (vllm-project#7513)

Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Co-authored-by: wangyu <53896905+yenuo26@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working diffusion codes related to diffusion models ready label to trigger buildkite CI

Projects

None yet

5 participants