Skip to content

[Bugfix] Gate Qwen-Image fused QK-norm+RoPE on short sequences (#7780) - #7965

Merged
Gaohan123 merged 4 commits into
vllm-project:mainfrom
NumberWan:fix/qwen-image-fused-qk-rope-min-tokens
Sep 23, 2026
Merged

Gaohan123 merged 4 commits into
vllm-project:mainfrom
NumberWan:fix/qwen-image-fused-qk-rope-min-tokens

Conversation

@NumberWan

@NumberWan NumberWan commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Fixes the Qwen-Image nightly latency regression in #7780 after always-on fused QK-norm+RoPE (#5931 / 952022c8e).
  • On short sequences the fused path is host-bound (custom-op dispatch + Triton launch), so end-to-end cost can exceed the kernel savings. Gate fusion on B*S >= 2048 (same helper as Boogu: VLLM_OMNI_FUSED_QK_NORM_ROPE_MIN_TOKENS, 0 = always fuse) instead of reverting [Diffusion][Perf] Add Qwen-Image QK RoPE Triton path #5931.

CI cases from #7780

All failing rows are still bench=512x512_steps20* (short seq). From the issue body and congw729's 2026-09-21 update (still failing 5 days; adds high_concurrency):

Test Config Metrics failing on H100 (vs baseline)
test_qwen_image_single_device 512x512_steps20, c=1, n=10 e2e_latency_ms ~−26%, throughput_qps ~−20%
test_qwen_image_single_device_step_execution 512x512_steps20, c=1, n=10 e2e_latency_ms ~−21%, throughput_qps ~−18%
test_qwen_image_single_device_step_execution 512x512_steps20_high_concurrency, c=1, n=20 e2e_latency_ms ~−16%, throughput_qps ~−14%

Same root cause for all three: 512² stays under the fuse crossover, so the gate should put them back on the eager path. Local verification below covers only the first config; the other two share the same DiT call shape.

Local reproduction (bisect)

NVIDIA L20X, FLASH_ATTN, test_qwen_image_single_device / 512x512_steps20, c=1, n=10.

Bisect pinned the regression to #5931 (952022c8e); its parent a3dee6fdf is clean:

Build Commit latency_mean (s) latency_median (s) latency_p99 (s) QPS diffuse_mean (s)
Parent (pre-#5931) a3dee6fdf 2.137 2.122 2.220 0.468 1.940
Culprit (#5931) 952022c8e 2.370 2.368 2.430 0.422 2.183

→ +10.9% end-to-end latency (median also moves; slowdown sits in QwenImagePipeline.diffuse).

Gap vs CI (~11% local vs ~20–26% H100): not all of the golden-baseline gap is #5931.

Nightly tips around the cliff (~02:00 HKT):

Nightly tip Has #5931? Notes
9/14 02:00 e284d907b no near golden
9/15 02:00 5f25d9865 no already ~10% off golden
9/16 02:00 78934753f yes (merged 9/15 14:28 UTC) much worse

9/14→9/15 is before #5931. In that window the only Qwen-Image tree change is #7461 (drop dead _get_qwen_prompt_embeds in the Edit pipeline — not on this T2I hot path). So that first ~10% looks like the same class of H100 host / node variance already discussed on #7309 (same test / same 2406.5 ms golden; recovered without a code fix). #5931 then adds a reproducible short-seq fuse tax on top (this PR).

Expect this gate to walk back the #5931 chunk; do not expect it alone to always land inside 10% of golden 2406 on H100 if the pre-#5931 host gap is still present.

Fix result (same harness)

Build latency_mean (s) latency_median (s) QPS diffuse_mean (s) vs parent
Parent 2.137 2.122 0.468 1.940 —
#5931 2.370 2.368 0.422 2.183 +10.9%
This PR (gate @ 2048) 2.187 2.179 0.457 1.991 +2.3%

Gate recovers most of the local regression; residual ~2% vs parent is within normal noise on this box.

Unit tests: short seq stays on eager; B*S >= 2048 takes fused; fused-kernel correctness tests force min_tokens=0.

Test plan

  • Local bisect parent vs [Diffusion][Perf] Add Qwen-Image QK RoPE Triton path #5931 on test_qwen_image_single_device / 512x512_steps20 c=1 n=10 (+10.9%)
  • Same harness after gate (+2.3% vs parent)
  • pytest tests/diffusion/models/qwen_image/test_qwen_image_fused_qk_norm_rope.py (13 passed)
  • H100 / nightly: test_qwen_image_single_device + test_qwen_image_single_device_step_execution for 512x512_steps20 and 512x512_steps20_high_concurrency (expect CI ~20% gap may not fully close if something beyond the fuse gate remains)

Fixes #7780

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/offloader.md, docs/design/module/diffusion/diffusion_model_integration.md, docs/design/module/diffusion/index.md.

Module owners: @wtomin @david6666666 @Isotr0py

Routing: @wtomin via module of the changed files, module named in the PR description, CODEOWNERS; @david6666666 via module of the changed files, module named in the PR description; @Isotr0py via module of the changed files, module named in the PR description

@NumberWan, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 22, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 4d72f6802947 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@NumberWan

Copy link
Copy Markdown
Contributor Author

This PR appears to belong to: docs/design/module/diffusion/offloader.md, docs/design/module/diffusion/diffusion_model_integration.md, docs/design/module/diffusion/index.md.

Module owners: @wtomin @david6666666 @Isotr0py

Routing: @wtomin via module of the changed files, module named in the PR description, CODEOWNERS; @david6666666 via module of the changed files, module named in the PR description; @Isotr0py via module of the changed files, module named in the PR description

@NumberWan, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

This PR is add a min gate for the Triton fused kernel, since before this PR, any token length will used Triton fused kernel, but the Triton fused kernel only gain when the token len is enough.
Boogu model used the same way to handle this issue

@congw729 congw729 added ready label to trigger buildkite CI cuda-test Used to trigger vllm-omni cuda CI separately. labels Sep 22, 2026
@hsliuustc0106 hsliuustc0106 added the bug Something isn't working label Sep 23, 2026

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123 Gaohan123 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI cuda-test Used to trigger vllm-omni cuda CI separately. labels Sep 23, 2026
@Gaohan123
Gaohan123 enabled auto-merge (squash) September 23, 2026 03:35
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: CI is red on this head

@NumberWan required checks failed on 01d8073cb818:

Please fix the failure and push again; this note is updated in place when the head goes green or moves.

@NumberWan

Copy link
Copy Markdown
Contributor Author

Omni ReviewBot: CI is red on this head

@NumberWan required checks failed on 01d8073cb818:

Please fix the failure and push again; this note is updated in place when the head goes green or moves.

The Simple · Diffusion Test failure on build 15934 is unrelated to this PR.

The job is 6 failed / 6543 passed. All six failures are in tests/diffusion/distributed/test_wan_vae_fastpath_install.py (KeyError: 'WanRMS_norm', ImportError: cannot import name 'RMSNormVAE', and installer installed=True assertions). This PR only gates fused QK-norm+RoPE on short sequences; it does not touch Wan VAE fastpath or vllm_omni.diffusion.layers.norm.

@NumberWan

Copy link
Copy Markdown
Contributor Author

AMD build 12594 is the same class of failure as CUDA Simple Diffusion on 15934, not this PR.

AMD L2 runs pytest tests/diffusion -m 'core_model and cpu' in four Simple Diffusion shards. CUDA already failed 6 tests in test_wan_vae_fastpath_install.py (WanRMS_norm / RMSNormVAE). Those are CPU tests, so the ROCm shards hit the same installer/norm mismatch.

This PR only gates fused QK-norm+RoPE on short sequences; it does not touch Wan VAE fastpath. Intel CI on this head is green.

@Gaohan123 Gaohan123 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 23, 2026
@Gaohan123
Gaohan123 merged commit 1934d12 into vllm-project:main Sep 23, 2026
6 of 9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Nightly CI, Qwen-image, performance metrics regressed by more than 10% compared to the baseline in some scenarios

5 participants