Skip to content

[Kernel][Qwen-Image] Fuse text/image QK RMSNorm + RoPE + cat into one Triton launch - #7594

Open
yuweih205 wants to merge 5 commits into
vllm-project:mainfrom
yuweih205:perf/qwen-image-fused-qk-norm-rope
Open

yuweih205 wants to merge 5 commits into
vllm-project:mainfrom
yuweih205:perf/qwen-image-fused-qk-norm-rope

Conversation

@yuweih205

@yuweih205 yuweih205 commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Independent merge update (2026-10-08)

Current head: cd2a996a on main c548a110; shared operator commits: bf9acacd and 24f9e25e (64-bit indexing). The six PRs #7560, #7594, #7595, #7596, #7597, and #7600 now each contain this identical shared commit plus only their own model integration. Any one of the six can merge first. The other five do not import code from that model.

The shared operator, test, and documentation files are byte-identical across all six heads. Local pre-commit checks passed (including ruff, mypy, SPDX, and markdownlint). Git merge simulations found no conflicts for all 15 pairs or for all 30 ordered pairs with the first PR squash-merged onto main. This Mac host has no CUDA/PyTorch test environment, so the GPU results below are historical results on earlier revisions; the new heads await CI and any GPU rerun.


Historical synchronization update (2026-09-21)

Current head: 2a061959c. Merged upstream main
1b87115c9ede2d29702c847b6da98f9e854fa593 while preserving commit history.
At that revision #7560 still landed first. The independent-merge update above supersedes this merge order.

  • Preserved main's large-storage-offset fix and regression test; extended the same
    64-bit row addressing to the joint QKV kernel. Added four CUDA regression cases
    covering either large-offset stream and both RoPE pairings. Floating-point
    arithmetic, model wiring, sequence-parallel behavior, and default gates are unchanged.
  • All applicable changed-file pre-commit hooks passed locally, including mypy,
    ruff, SPDX, test-mark checks, and markdownlint. A Python AST comparison also
    verified that the existing model integrations and shared floating-point operations
    were preserved (apart from widening address indices).

Operator-level GPU regression update (2026-09-21)

Operator-level correctness and regression tests for the current revision have now passed on real H200 hardware.

  • Environment: Python 3.12.12, vLLM 0.29.0, Torch 2.13.0+cu130, CUDA 13.0, and Triton 3.7.1. The repository's native pytest suites ran in a CUDA base image with an isolated venv. This is independent GPU validation, not a Buildkite CI Docker-image run or the earlier CPU-shim check.
  • Shared operator suite: the complete tests/diffusion/layers/test_fused_qk_norm_rope.py passed on [Kernel][Flux.2] Fuse text/image QK RMSNorm + cat + RoPE into one Triton launch #7560 at 08dc20e68: 30 passed. The 6 large-storage-offset cases also passed separately, including the 4 new dual-stream cases. Those 6 are included in the 30 and are not counted twice. The shared source and test files are byte-identical across the six current PR heads.
  • Coverage includes 64-bit address-index boundaries, joint/per-stream bitwise parity, eager-reference tolerances, the Flux.2 eager chain, torch.compile / CUDA graph, and shape, gate, and fallback behavior. Bitwise guarantees apply only to the corresponding exact-parity cases, not to every fused/eager comparison or the full model.
  • This PR's Qwen-Image operator-integration suite at 2a061959c: 14 passed, covering FP32/FP16 fallback, BF16 and packed-view cases, torch.compile fullgraph, and joint/per-stream parity. These are component/attention-level regressions, not full-model end-to-end tests.
  • Every suite above has 0 failures / 0 errors / 0 skipped. The original JUnit reports, pinned commits, file hashes, and actual import paths were cross-checked. The model-specific suites use randomly initialized parameters/inputs with offline mode enabled; no pretrained weights were downloaded for this run.

Scope: operator-level and model operator-integration regressions are now complete. Pretrained-model end-to-end accuracy and performance ABBA were not rerun. All historical accuracy and performance results below, including the two earlier measurements, are retained with their original revisions and environments; they are not new performance measurements for this revision.

To reproduce the shared operator run, execute from the tests/ directory of #7560 at 08dc20e68, with the compatible CUDA/Triton environment and repository test dependencies above. The large-storage-offset cases require slightly more than 4 GiB of free GPU memory:

python -m pytest -s -v diffusion/layers/test_fused_qk_norm_rope.py \
  -m "core_model and cuda" --run-level=core_model -p no:cacheprovider

Run the full operator-integration suite from the tests/ directory of this PR's pinned head:

python -m pytest -s -v diffusion/models/qwen_image/test_qwen_image_fused_qk_norm_rope.py \
  -m core_model --run-level=core_model -p no:cacheprovider

Use -m core_model for the Qwen file without adding and cuda: two joint GPU cases do not carry the CUDA marker and would otherwise be omitted.

AI assistance for this synchronization: Codex resolved the merge, updated the indexing
regression coverage, ran the local checks and GPU regressions above, and verified the JUnit results.


Merge order update (2026-10-08): Any of the six model PRs can merge first. Each branch includes the same shared operator commit and only its own model integration.

Update 2026-09-17 — rebased, and the baseline moved. #5931 and #7513 landed after this PR
was opened. main now fuses Q/K RMSNorm + RoPE per stream (_qwen_image_qk_norm_rope) and then
concatenates, and its eager fallback rotates with BF16 RotaryEmbedding instead of the FP32
complex helper this description refers to below. The PR is rebased on that main and its increment
is now two per-stream launches plus three cats → one joint launch, no cats. Everything under
"Update: against the new main" is the current state; the original sections below are kept for
history and describe the pre-#5931 baseline.

Update: against the new main

The joint launch produces the same Q/K/V as main's two per-stream launches followed by the three
cats, so nothing about the images changes — only the launch count and the copies.

Numerics (H200, full 60-block transformer, random init, txt 512). Eager: joint vs main's
per-stream path is bitwise identical (torch.equal, relative L2 exactly 0) at B = 1 and 2,
512² and 1024². The packed table carries the same FP32 coefficients _qwen_image_qk_norm_rope
builds, so in eager this is equality by construction rather than a tolerance.

Under regional torch.compile the two paths are not bitwise equal: relative L2 2.17–2.18% on
the same random-init model, consistently across all three shapes. Only the surroundings differ —
the per-stream arm hands Inductor three cats to fuse and lay out, the joint arm hands it none —
so the graphs either side of the custom op are not the same and the ≤1-ulp consequences amplify
over 60 layers (the same scale as the 1.5–1.6% measured at full depth for FLUX.1 in #7595). We have
not measured the e2e SSIM against Diffusers on the compiled path, so the gate #7513 raised to
0.97 is not something this PR can claim to leave untouched; the eager equality above is what is
established.

Runtime path (asserted in both benchmarks): with the gate off, 120 per-stream launches per
forward (60 blocks × 2 streams) and 0 joint launches; with it on, 60 joint launches and 0
per-stream launches.

Transformer forward, CUDA-event p50, A B B A (3 × 10 after warmup):

B image tokens (side²) main (per stream + cats) this PR (joint) saving
1 4,096 (1024²) 164.28 ms 158.82 ms −3.33%
2 4,096 (1024²) 324.81 ms 314.62 ms −3.14%
2 1,024 (512²) 105.25 ms 102.39 ms −2.72%

Same measurement with the blocks under vLLM-Omni's regionally_compile (60/60 blocks), which is
what a default deployment runs:

B image tokens (side²) main (per stream + cats) this PR (joint) saving
1 4,096 (1024²) 164.49 ms 158.99 ms −3.34%
2 4,096 (1024²) 325.13 ms 314.55 ms −3.25%
2 1,024 (512²) 105.73 ms 103.32 ms −2.28%

Peak allocation unchanged (76.2 GiB at B = 2). This is the per-denoising-step cost; the −5.3% per
image reported below was measured against the fully eager pre-#5931 chain and no longer describes
the increment over main. We have not re-run the end-to-end pipeline on the new baseline — with
byte-identical outputs the only change there is this forward saving spread over the denoising loop.

Unit tests: 14 passed on H200 — main's 12 plus two new ones: the joint launch is bitwise equal
to per-stream-plus-cat at (512 + 4096) and (77 + 1024), and the table packer returns None without
allocating where the CUDA kernel cannot run.

Sequence parallelism: any SP keeps the per-stream chain. _sp_plan shards vid_freqs while
txt_freqs stays replicated, so a joint table built from the two would not line up with the local
tokens.


Purpose

Follow-up to #7560 (Flux.2), same optimization for Qwen-Image (QwenImageCrossAttention, all 60
blocks are dual-stream): the per-block chain of four nn.RMSNorms + four fp32 complex RoPE
applications + three torch.cats
becomes one Triton launch whose outputs are the joint
[B, S_txt+S_img, H, D] Q/K/V attention consumes
, with no intermediate copies.

  • The attention takes an optional qk_norm_rope_table; the block forwards it from
    joint_attention_kwargs, and the model packs it once per forward from the complex
    (txt_freqs, vid_freqs) as an fp32 [B*S, D] = [cos | sin] table (text rows first, as the
    eager cat orders them). Kept in fp32 because the eager path also rotates with unrounded
    fp32 frequencies (_apply_qwen_image_rotary_emb).
  • The eager chain stays for sequence parallelism (RoPE per shard, text through joint_*
    metadata) and for any dtype/geometry the CUDA kernel cannot take; default gate
    _FUSED_MIN_TOKENS = 0, VLLM_OMNI_FUSED_QK_NORM_ROPE_MIN_TOKENS still overrides.
  • Qwen-Image's eager chain already rounds once (torch RMSNorm, fp32 RoPE), so fused and eager
    differ only through fp32 reduction order in sum(x²) (see numerics).

Independent branch: shared operator commit plus this model integration, based directly on main.

Test Plan

vLLM Version: 0.28.0 (torch 2.13.0+cu130, Triton 3.7.1), on the last pre-vLLM-0.29 tree with
qwen_image/ and fused_qk_norm_rope.py from main — see #7560 for why.

vLLM-Omni Commit: branch on 1b6cd28 + #7560.

Hardware: 1× NVIDIA H200.

  • tests/diffusion/models/qwen_image/test_qwen_image_fused_qk_norm_rope.py: table geometry and
    gate; fused op vs the attention's actual chain (nn.RMSNorm + _apply_qwen_image_rotary_emb,
    text-first cat) at (B=1, 77+1024) and (B=2, 512+4096).
  • End-to-end Qwen/Qwen-Image text-to-image through Omni(model=..., mode="text-to-image")
    (QwenImagePipeline), true CFG 4.0 with a negative prompt, 1024², 50 and 20 steps, 3 prompts ×
    2 seeds × 2 repeats per arm, gate off vs on, each arm its own process, A B B A. Wall-clock per
    generate() including text encoder (Qwen2.5-VL-7B) and VAE.

Test Result

Unit tests: 3 passed on H200 (with the op-level suite from #7560: 19 passed).

End-to-end, 1024², CFG (B=2 through the transformer):

steps eager p50 fused p50 per image per denoising step
50 13.05 s 12.37 s −0.69 s (−5.3%) 258.3 → 244.7 ms (−5.3%)
20 5.30 s 5.03 s −0.28 s (−5.3%)

Fixed cost outside the denoising loop ≈ 0.14 s; peak allocation unchanged (56.7 GiB).

Numerics:

  • Each arm is bitwise deterministic across runs.

  • Op level (unit test): fused vs eager Q/K differ on < 2% of elements, by 1 bf16 ulp — the
    fp32 reduction-order residue only, since both sides round once.

  • Images, eager vs fused on identical (prompt, seed): PSNR median 37.1 dB (24 pairs), minimum
    19.1 dB. The low-PSNR pairs are the same scene with small pose/detail differences (e.g. a
    horse's leg positions), the usual outcome of a ulp-level perturbation amplified by 20–50 CFG
    sampling steps — not a systematic shift.

    Control arm. To separate "kernel error" from "any ulp-level change gets amplified", a third
    arm runs the eager path with every RMSNorm replaced by a plain torch implementation of the
    same single-rounding order (a sitecustomize patch in the worker; nothing else changes). That
    arm vs eager: PSNR median 39.9 dB, min 19.1 dB (12 pairs) — the same spread as fused vs eager
    (37.1 / 19.1). Fused vs the control arm: 36.1 / 26.9 dB. So a reduction-order-only change
    already moves the samples exactly as much as the fused kernel does; there is no additional
    error attributable to the kernel.

🤖 Generated with Claude Code

https://claude.ai/code/session_01G8uaASBAbnS8TsFvC5gCUD

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/diffusion_model_integration.md, docs/design/module/diffusion/index.md.

Module owners: @david6666666 @wtomin @Isotr0py

Routing: @david6666666 via module of the changed files, module named in the PR description, CODEOWNERS; @wtomin via module of the changed files, module named in the PR description, CODEOWNERS; @Isotr0py via module of the changed files, module named in the PR description

@yuweih205, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 16, 2026
@yuweih205
yuweih205 force-pushed the perf/qwen-image-fused-qk-norm-rope branch from ee0501e to 333add5 Compare September 16, 2026 06:13
@yuweih205

Copy link
Copy Markdown
Contributor Author

Self-review: QwenImageCrossAttention takes an optional qk_norm_rope_table and replaces its 4 RMSNorm + 4 complex RoPE + 3 cat with one fused_joint_qkv_norm_rope launch (text first, as the eager cat); the model packs the fp32 table once per forward, only when fused_qk_norm_rope_available() passes. SP and unsupported dtype/geometry keep the untouched eager chain. Tests: table geometry/order on CUDA, CPU skip, fused op vs the attention's real chain. End-to-end on Qwen/Qwen-Image in the description (−5.3% per image; control arm shows the difference is RMSNorm rounding order only).

@Dong1017 Dong1017 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Scope. Follow-up to #7560 making Qwen-Image the next consumer of the fused QK-norm-RoPE family: the per-block four-RMSNorm + four-RoPE + three-cat chain becomes one packed launch, with the fp32 [cos | sin] table built once per forward and the eager chain kept as fallback (_FUSED_MIN_TOKENS = 0, env override).

Measured on H200 (nightly diffusers-parity recipe: 512×512, 20 steps, CFG 4.0, seed 42 — complements the PR's fused-vs-eager numerics with vs-diffusers gate numbers):

path SSIM PSNR
main@21d86ec9 default (fused, fp32 table) 0.964473 28.772
this PR default (packed table) 0.964473 28.772 — trajectory-identical to main
this PR, gate off (eager) 0.964340 28.861 — same regime, consistent
bf16-coefficient path (reference band, #7513's forced-eager) 0.985766 33.768

So the packed refactor preserves diffusers-parity behavior exactly — consistent with the op-level result and the control-arm reasoning in the PR description.

One question this data raises, probably for #7382 rather than here: the "kept in fp32" rationale (identical coefficients across fused/eager) is about internal consistency; vs diffusers, fp32 coefficients diverge on 28.09% of rotated elements (max 0.03125, op-level 1 ulp either way), and the fp32 band (0.9645) sits below #7513's proposed SSIM 0.97 gate while the bf16 band passes with margin (and #7494 shows the same main tree scoring 26.13 on the nightly infra). Since #7513 also rewrites _qwen_image_qk_norm_rope, the regime choice and that threshold seem coupled — might make sense to settle them together as a #7382 contract decision; happy to help with either direction.

Comment thread vllm_omni/diffusion/models/qwen_image/qwen_image_transformer.py Outdated
Comment thread vllm_omni/diffusion/models/qwen_image/qwen_image_transformer.py Outdated
@yuweih205
yuweih205 force-pushed the perf/qwen-image-fused-qk-norm-rope branch from 333add5 to 0867925 Compare September 17, 2026 06:17
@RuixiangMa

RuixiangMa commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

@yuweih205 Please resolve the merge conflicts with the latest main branch

@yuweih205
yuweih205 force-pushed the perf/qwen-image-fused-qk-norm-rope branch 2 times, most recently from 9f11491 to f3afda8 Compare September 17, 2026 11:04
@yuweih205

Copy link
Copy Markdown
Contributor Author

@RuixiangMa Done — rebased on latest main and reworked on top of it. main now fuses Q/K RMSNorm + RoPE per stream (#5931) with a BF16 RotaryEmbedding eager fallback (#7513), so this PR's increment is now two per-stream launches plus three cats → one joint launch with no cats.

On H200, full 60-block forward: −3.3% / −3.1% / −2.7% eager and −3.3% / −3.3% / −2.3% compiled, at (B=1, 1024²), (B=2, 1024²), (B=2, 512²). In eager the joint output is bitwise identical to the per-stream path (same FP32 coefficients); under regionally_compile they differ by relative L2 2.2% on a random-init model, since Inductor sees three cats on one side and none on the other. 14 unit tests pass. Details and the superseded pre-#5931 numbers are in the description.

@yuweih205

yuweih205 commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

@RuixiangMa One thing worth flagging while you are looking at these: #7560 needs to land first as things stand. It carries the shared op — fused_joint_qkv_norm_rope plus the table helpers in vllm_omni/diffusion/layers/fused_qk_norm_rope.py and the op-level tests — and this PR contains those same commits plus only the Qwen-Image wiring. Merging this one first would pull the op and Flux.2's own integration in with it, which is not what this PR describes.

After #7560 lands, this PR and the other four consumers (#7595 FLUX.1, #7596 HunyuanVideo-1.5, #7597 Z-Image, #7600 Ovis-Image) are independent of each other and can go in any order. I have added the same note to each description.

That said, the order is not fixed — tell me which you prefer and I will restructure:

  1. Keep [Kernel][Flux.2] Fuse text/image QK RMSNorm + cat + RoPE into one Triton launch #7560 first (current state): it merges the op together with the Flux.2 / klein consumer.
  2. Split the op into its own PR: a model-independent change on the critical path, with Flux.2 becoming a sixth consumer alongside the rest.
  3. Let this PR go first: I drop the Flux.2 / klein wiring from it so it is just the shared op plus Qwen-Image, and [Kernel][Flux.2] Fuse text/image QK RMSNorm + cat + RoPE into one Triton launch #7560 then becomes the Flux.2-only consumer. Same for any other consumer you would rather take first.

Happy with whichever fits your review flow best.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 21, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit dfc7f368a121 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@hsliuustc0106 hsliuustc0106 added the Kernel optimization Codes related to optimize kernel execution to improve hardware utilization label Sep 22, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 9 days

@yuweih205 this pull request has had no human commit, comment or review since 2026-09-19. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on zcode (GLM-5.3-Flash) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

Signed-off-by: HuangYuwei <yuweih205@gmail.com>
@yuweih205
yuweih205 force-pushed the perf/qwen-image-fused-qk-norm-rope branch from 2a06195 to 4277824 Compare October 8, 2026 04:15
Signed-off-by: HuangYuwei <yuweih205@gmail.com>
Signed-off-by: HuangYuwei <yuweih205@gmail.com>
@yuweih205
yuweih205 force-pushed the perf/qwen-image-fused-qk-norm-rope branch from 4277824 to cd2a996 Compare October 8, 2026 06:13
@yuweih205

Copy link
Copy Markdown
Contributor Author

@RuixiangMa The merge conflicts you flagged have been resolved; GitHub currently reports the latest head (ec39cdd0) as mergeable. Could you please review the updated PR when you have a chance?

The six-PR series (#7560, #7594, #7595, #7596, #7597, #7600) has also been restructured: each PR contains the same shared operator, tests and documentation changes, plus only its own model integration. There is no longer a prerequisite PR within this series; any one of the six can merge first, once its review and CI requirements are satisfied. This supersedes the earlier comments saying #7560 had to land first. The shared files are identical across the six heads, and local merge simulations passed for all 15 pairs and all 30 ordered pairs with the first PR squash-merged onto main.

Signed-off-by: HuangYuwei <yuweih205@gmail.com>
Signed-off-by: HuangYuwei <yuweih205@gmail.com>
@yuweih205
yuweih205 force-pushed the perf/qwen-image-fused-qk-norm-rope branch from ebfa21a to dfc7f36 Compare October 8, 2026 07:42
@yuweih205

Copy link
Copy Markdown
Contributor Author

Reproducible eager bitwise case: retained full-transformer results

The already measured eager configuration preserves the full returned tensor against the historical main per-stream fused baseline, while retaining a 2.72–3.33% transformer-forward saving. The original harness, source overlay and three raw records have now been recovered and made reproducible. These are the September 17 measurements already summarized in this PR; no new GPU measurement is being reported, and the existing compiled and end-to-end performance results are retained.

Configuration change: run without the harness's --compile flag, and set VLLM_OMNI_FUSED_QK_NORM_ROPE_MIN_TOKENS=0 for the joint arm. The historical comparison arm uses 1000000000 for that joint gate and still executes main's per-stream fused kernels. Runtime counts confirm 120 per-stream + 0 joint launches → 0 per-stream + 60 joint launches per forward. On this pinned historical source, the large joint gate does not disable the per-stream fusion. Current main changed that gate behavior; the reproduction below therefore pins the original source rather than applying the historical switch to a newer tree.

Exact settings:

  • One H200 141 GB; Python 3.12.3, vLLM 0.28.0, Torch 2.13.0+cu130, CUDA 13.0, Triton 3.7.1, Diffusers 0.40.0. The retained environment metadata was read without executing the model.
  • QwenImageTransformer2DModel(OmniDiffusionConfig(model="test", dtype=torch.bfloat16), num_layers=60), .eval(), no compilation, one process, all other constructor and attention defaults. This gives 24 attention heads of dimension 128, 64 image-input channels and 3,584 text-input channels.
  • Constructor seed 0. One CUDA generator with seed 1 then sequentially initializes parameters: 1D norm weights use 1 + 0.05 * randn; other 1D parameters use 0.01 * randn; matrices use 0.02 * randn in FP32, cast to BF16. The same generator continues into input generation.
  • hidden_states=[B, image_tokens, 64], encoder_hidden_states=[B, 512, 3584], an all-true [B,512] mask, BF16 timestep [B] filled with 0.5; img_shapes=[[(1, sqrt(image_tokens), sqrt(image_tokens))]] * B, txt_seq_lens=[512] * B.
Batch / image tokens Main per-stream + cats Joint Forward saving Returned tensor
1 / 4096 164.278221 ms 158.815163 ms 3.325492% torch.equal=True, relative L2 = 0
2 / 4096 324.814789 ms 314.622177 ms 3.137976% torch.equal=True, relative L2 = 0
2 / 1024 105.247505 ms 102.388176 ms 2.716767% torch.equal=True, relative L2 = 0

Timing is CUDA-event p50, A B B A, ten forward calls per segment, with three warmups before each timed segment. Comparison covers the entire returned transformer tensor, converted from BF16 to FP32 before torch.equal; the old harness did not record an integer-view signed-zero audit. This is a random-weight full transformer, not pretrained end-to-end generation. The compiled arms were not bitwise equal and are not included in this equality claim.

Reproduction uses base ff6e906a25de1ca3122c0d0d5250fde5a0ddbc7b plus the 21-file overlay from f3afda8aa8b841e2d72a842dea534d63fc65555a. Every retained overlay file matches that commit exactly. See the staging instructions, harness, provenance and file manifest. The retained harness includes the optional compile flag added after the eager measurements; its saved hash identifies the recovered copy rather than a hash recorded at the original run.

PYTHONPATH="$RUN_DIR" "$PYTHON" "$RUN_DIR/qwen_joint_bench.py" --layers 60 --batch 1 --txt 512 --img 4096 --out "$RUN_DIR/b1_4096.json"
PYTHONPATH="$RUN_DIR" "$PYTHON" "$RUN_DIR/qwen_joint_bench.py" --layers 60 --batch 2 --txt 512 --img 4096 --out "$RUN_DIR/b2_4096.json"
PYTHONPATH="$RUN_DIR" "$PYTHON" "$RUN_DIR/qwen_joint_bench.py" --layers 60 --batch 2 --txt 512 --img 1024 --out "$RUN_DIR/b2_1024.json"

Raw records: B1 / 4096, B2 / 4096, B2 / 1024. They retain the zero relative-L2 values, exact launch counters and original timings. These establish the stated historical eager case; they are not measurements of the October 8 heads.

AI assistance: Codex recovered and checked the existing harness/source hashes/raw records and prepared this additive reproduction note; no GPU case was rerun for this note.

@yuweih205

Copy link
Copy Markdown
Contributor Author

Fresh eager full-transformer measurement: bitwise joint fusion

Reran the Qwen-Image eager per-stream-versus-joint case on the current prepared source, with strict storage-byte checks for the complete returned BF16 tensor.

Tested source 8b57eb5986c93f6562fd730e3552f2c04e418b47 over public base dfc7f368a12170f9b5e41c33245d512ccd8605d7. The shared numerical-preset patch is outside the public PR head; both Qwen arms use the existing fast mode, so this measurement does not depend on switching to the CUDA-order preset.

Configuration: full 60-layer random-weight transformer, 24 heads × 128 dimensions, BF16, 512 text + 4096 image tokens, B1/B2, H200 141 GB, world/TP/SP size 1. Constructor seed 0; CUDA generator seed 1 initializes parameters and then inputs. vLLM 0.29.0, Torch 2.13.0+cu130, CUDA 13.0, Triton 3.7.1.

The reference uses the existing per-stream fused chain plus cats, with only the joint table packer suppressed. Both arms set VLLM_OMNI_FUSED_QK_NORM_ROPE_MIN_TOKENS=0. Runtime counts verify 120 per-stream + 0 joint launches versus 0 per-stream + 60 joint launches. Model RoPE generation and per-forward packing are included.

Batch Per-stream (ms) Joint (ms) Forward latency reduction Full BF16 storage bits
1 167.449280 161.765442 3.394% 0 byte mismatches
2 329.690540 319.401697 3.121% 0 byte mismatches

Both arms repeat exactly; strict byte parity, including signed zero, is checked before and after timing. Counter wrappers are removed for the measured forwards. CUDA-event medians: three warmups per arm, three ABBA rounds, ten forward calls/sample, six samples/arm. This is a random-weight full-transformer result.

B1 raw record, B2 raw record, harness, and the reproduction procedure and all raw results.

AI assistance: Codex reran and checked the GPU measurements and prepared this additive evidence note.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 1ef71a57-d4cf-4f07-b058-7cdf689e4956) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 120s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: aef4a307-2449-4045-aec8-49662a3fbf3a) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 600s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 9394ce5d-38cc-4733-bc5f-dc8fd2621749) — check zcode login and the CLI version; falling back to direct/cursor/auto).

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

Changes since the previous review

  • 1 new inline finding(s); 1 finding(s) below.

CI at dfc7f368a121 (2026-10-10T10:29:18.793635+00:00): required check(s) blocking: buildkite/vllm-omni (missing).

Note: The assigned review arm strict/zcode/GLM-5.3-Flash could not complete this review, so it was produced by the fallback arm direct/cursor/auto. It is excluded from the routing experiment.

Full review analysis

PR description

This change adds a shared two-stream Triton op that RMSNorms and RoPE-rotates text and image Q/K and gathers V into the joint sequence attention already consumes. Qwen-Image packs one float32 [cos | sin] table per transformer forward and, when that table exists, runs the joint op from every cross-attention block instead of the existing per-stream norm/RoPE path plus three concatenations. Sequence-parallel, non-BF16, and non-CUDA forwards keep the per-stream path.

Change flow

flowchart LR
  freqs["[EXISTING] text and image complex freqs"]:::existing
  pack["[NEW] packed joint FP32 RoPE table"]:::new
  attn["[CHANGED] QwenImageCrossAttention"]:::changed
  joint["[NEW] fused_joint_qkv_norm_rope"]:::new
  perstream["[EXISTING] per-stream RMSNorm and RoPE"]:::existing
  qkv["[CHANGED] joint Q, K, V"]:::changed
  freqs --> pack --> attn
  attn --> joint --> qkv
  attn --> perstream --> qkv
classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

Findings

  • [P1] Joint gate skips the short-sequence BF16 RoPE path — vllm_omni/diffusion/models/qwen_image/qwen_image_transformer.py:755
    Existing thread: #7594 (comment)
Evidence for Joint gate skips the short-sequence BF16 RoPE path

_JOINT_FUSED_MIN_TOKENS is 0, and use_fused_joint launches fused_joint_qkv_norm_rope whenever the packed FP32 table exists. The per-stream helper it replaces still fuses only when B*S >= fused_qk_norm_rope_min_tokens(2048); below that, _qwen_image_qk_norm_rope stays on activation-dtype RotaryEmbedding (torch.real/torch.imag cast to BF16), which is the Diffusers-aligned CUDA path from #7513. Typical text lengths fail that gate (BATCH*512 = 1024 is the existing short-seq test), and a 512² image stream does too, so the default joint path rotates those tokens with the FP32 Triton coefficients instead. test_joint_launch_matches_per_stream_then_cat hides this: force_always_fuse plus use_fused=True compares only with the forced fused kernel. The published bitwise forwards also set VLLM_OMNI_FUSED_QK_NORM_ROPE_MIN_TOKENS=0 on both arms. Enable the joint op only when each stream would already take the fused per-stream kernel, or re-measure the compiled Omni-vs-Diffusers gate (SSIM ≥ 0.97, PSNR ≥ 30) on this default path before treating the FP32 joint launch as the serving default.


🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!


@pytest.mark.skipif(not torch.cuda.is_available(), reason="CUDA required")
@pytest.mark.parametrize("txt_len,img_len", [(512, 4096), (77, 1024)])
def test_joint_launch_matches_per_stream_then_cat(txt_len, img_len, force_always_fuse):

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] New joint parity tests are invisible to the CUDA model step

Evidence and suggested fix

test_joint_launch_matches_per_stream_then_cat is marked core_model via the module pytestmark and has only a CUDA skipif. It does not carry @pytest.mark.cuda. The ready pipeline step Diffusion · Model Test in .buildkite/cuda/test-ready.yml runs pytest -sv tests/diffusion/models/ -m 'core_model and cuda' --run-level "core_model", so those two cases are dropped. The PR's own command is python -m pytest .../test_qwen_image_fused_qk_norm_rope.py -m core_model --run-level=core_model, and the body says adding and cuda would omit them. The new shared-kernel cases live in tests/diffusion/layers/test_fused_qk_norm_rope.py, which has the CUDA marker, but Diffusion · Other Test only selects tests/diffusion/*.py, tests/diffusion/ar_diffusion, and tests/diffusion/layers/test_adalayernorm_*.py, so that file is outside the ready CUDA commands as well. Add @pytest.mark.cuda to the joint Qwen cases so the model step actually runs them.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: finding feedback

[p1] Joint gate skips the short-sequence BF16 RoPE path — vllm_omni/diffusion/models/qwen_image/qwen_image_transformer.py:755

See the review for details. If you are the PR author and disagree, react 👎 here; the maintainer will see your disagreement.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: finding feedback

[p1] New joint parity tests are invisible to the CUDA model step — tests/diffusion/models/qwen_image/test_qwen_image_fused_qk_norm_rope.py:268

See the review for details. If you are the PR author and disagree, react 👎 here; the maintainer will see your disagreement.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

high priority high priority issue, needs to be done asap Kernel optimization Codes related to optimize kernel execution to improve hardware utilization

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants