Repository navigation
Conversation
…project#7780) Signed-off-by: NumberWan <wantszkin2003@gmail.com>
|
This PR appears to belong to: docs/design/module/diffusion/offloader.md, docs/design/module/diffusion/diffusion_model_integration.md, docs/design/module/diffusion/index.md. Module owners: @wtomin @david6666666 @Isotr0py Routing: @wtomin via module of the changed files, module named in the PR description, CODEOWNERS; @david6666666 via module of the changed files, module named in the PR description; @Isotr0py via module of the changed files, module named in the PR description @NumberWan, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
This PR is add a min gate for the Triton fused kernel, since before this PR, any token length will used Triton fused kernel, but the Triton fused kernel only gain when the token len is enough. |
Omni ReviewBot: CI is red on this head@NumberWan required checks failed on Please fix the failure and push again; this note is updated in place when the head goes green or moves. |
The Simple · Diffusion Test failure on build 15934 is unrelated to this PR. The job is 6 failed / 6543 passed. All six failures are in tests/diffusion/distributed/test_wan_vae_fastpath_install.py (KeyError: 'WanRMS_norm', ImportError: cannot import name 'RMSNormVAE', and installer installed=True assertions). This PR only gates fused QK-norm+RoPE on short sequences; it does not touch Wan VAE fastpath or vllm_omni.diffusion.layers.norm. |
|
AMD build 12594 is the same class of failure as CUDA Simple Diffusion on 15934, not this PR. AMD L2 runs pytest tests/diffusion -m 'core_model and cpu' in four Simple Diffusion shards. CUDA already failed 6 tests in test_wan_vae_fastpath_install.py (WanRMS_norm / RMSNormVAE). Those are CPU tests, so the ROCm shards hit the same installer/norm mismatch. This PR only gates fused QK-norm+RoPE on short sequences; it does not touch Wan VAE fastpath. Intel CI on this head is green. |
Summary
952022c8e).B*S >= 2048(same helper as Boogu:VLLM_OMNI_FUSED_QK_NORM_ROPE_MIN_TOKENS,0= always fuse) instead of reverting [Diffusion][Perf] Add Qwen-Image QK RoPE Triton path #5931.CI cases from #7780
All failing rows are still
bench=512x512_steps20*(short seq). From the issue body and congw729's 2026-09-21 update (still failing 5 days; adds high_concurrency):test_qwen_image_single_device512x512_steps20, c=1, n=10e2e_latency_ms~−26%,throughput_qps~−20%test_qwen_image_single_device_step_execution512x512_steps20, c=1, n=10e2e_latency_ms~−21%,throughput_qps~−18%test_qwen_image_single_device_step_execution512x512_steps20_high_concurrency, c=1, n=20e2e_latency_ms~−16%,throughput_qps~−14%Same root cause for all three: 512² stays under the fuse crossover, so the gate should put them back on the eager path. Local verification below covers only the first config; the other two share the same DiT call shape.
Local reproduction (bisect)
NVIDIA L20X,
FLASH_ATTN,test_qwen_image_single_device/512x512_steps20, c=1, n=10.Bisect pinned the regression to #5931 (
952022c8e); its parenta3dee6fdfis clean:a3dee6fdf952022c8e→ +10.9% end-to-end latency (median also moves; slowdown sits in
QwenImagePipeline.diffuse).Gap vs CI (~11% local vs ~20–26% H100): not all of the golden-baseline gap is #5931.
Nightly tips around the cliff (~02:00 HKT):
e284d907b5f25d986578934753f9/14→9/15 is before #5931. In that window the only Qwen-Image tree change is #7461 (drop dead
_get_qwen_prompt_embedsin the Edit pipeline — not on this T2I hot path). So that first ~10% looks like the same class of H100 host / node variance already discussed on #7309 (same test / same 2406.5 ms golden; recovered without a code fix). #5931 then adds a reproducible short-seq fuse tax on top (this PR).Expect this gate to walk back the #5931 chunk; do not expect it alone to always land inside 10% of golden 2406 on H100 if the pre-#5931 host gap is still present.
Fix result (same harness)
Gate recovers most of the local regression; residual ~2% vs parent is within normal noise on this box.
Unit tests: short seq stays on eager;
B*S >= 2048takes fused; fused-kernel correctness tests forcemin_tokens=0.Test plan
test_qwen_image_single_device/512x512_steps20c=1 n=10 (+10.9%)pytest tests/diffusion/models/qwen_image/test_qwen_image_fused_qk_norm_rope.py(13 passed)test_qwen_image_single_device+test_qwen_image_single_device_step_executionfor512x512_steps20and512x512_steps20_high_concurrency(expect CI ~20% gap may not fully close if something beyond the fuse gate remains)Fixes #7780