Skip to content

[Kernel][Z-Image] Fuse QK RMSNorm + RoPE into one Triton launch - #7597

Open
yuweih205 wants to merge 4 commits into
vllm-project:mainfrom
yuweih205:perf/z-image-fused-qk-norm-rope
Open

yuweih205 wants to merge 4 commits into
vllm-project:mainfrom
yuweih205:perf/z-image-fused-qk-norm-rope

Conversation

@yuweih205

@yuweih205 yuweih205 commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Independent merge update (2026-10-08)

Current head: 02984be8 on main c548a110; shared operator commits: bf9acacd and 24f9e25e (64-bit indexing). The six PRs #7560, #7594, #7595, #7596, #7597, and #7600 now each contain this identical shared commit plus only their own model integration. Any one of the six can merge first. The other five do not import code from that model.

The shared operator, test, and documentation files are byte-identical across all six heads. Local pre-commit checks passed (including ruff, mypy, SPDX, and markdownlint). Git merge simulations found no conflicts for all 15 pairs or for all 30 ordered pairs with the first PR squash-merged onto main. This Mac host has no CUDA/PyTorch test environment, so the GPU results below are historical results on earlier revisions; the new heads await CI and any GPU rerun.


Historical synchronization update (2026-09-21)

Current head: ee9f92405. Merged upstream main
1b87115c9ede2d29702c847b6da98f9e854fa593 while preserving commit history.
At that revision #7560 still landed first. The independent-merge update above supersedes this merge order.

  • Preserved main's large-storage-offset fix and regression test; extended the same
    64-bit row addressing to the joint QKV kernel. Added four CUDA regression cases
    covering either large-offset stream and both RoPE pairings. Floating-point
    arithmetic, model wiring, sequence-parallel behavior, and default gates are unchanged.
  • All applicable changed-file pre-commit hooks passed locally, including mypy,
    ruff, SPDX, test-mark checks, and markdownlint. A Python AST comparison also
    verified that the existing model integrations and shared floating-point operations
    were preserved (apart from widening address indices).

Operator-level GPU regression update (2026-09-21)

Operator-level correctness and regression tests for the current revision have now passed on real H200 hardware.

  • Environment: Python 3.12.12, vLLM 0.29.0, Torch 2.13.0+cu130, CUDA 13.0, and Triton 3.7.1. The repository's native pytest suites ran in a CUDA base image with an isolated venv. This is independent GPU validation, not a Buildkite CI Docker-image run or the earlier CPU-shim check.
  • Shared operator suite: the complete tests/diffusion/layers/test_fused_qk_norm_rope.py passed on [Kernel][Flux.2] Fuse text/image QK RMSNorm + cat + RoPE into one Triton launch #7560 at 08dc20e68: 30 passed. The 6 large-storage-offset cases also passed separately, including the 4 new dual-stream cases. Those 6 are included in the 30 and are not counted twice. The shared source and test files are byte-identical across the six current PR heads.
  • Coverage includes 64-bit address-index boundaries, joint/per-stream bitwise parity, eager-reference tolerances, the Flux.2 eager chain, torch.compile / CUDA graph, and shape, gate, and fallback behavior. Bitwise guarantees apply only to the corresponding exact-parity cases, not to every fused/eager comparison or the full model.
  • This PR's Z-Image operator-integration suite at ee9f92405: 4 passed, covering row-zero semantics, sequence-parallel table construction, and fused/eager attention comparisons. These are component/attention-level regressions, not full-model end-to-end tests.
  • Every suite above has 0 failures / 0 errors / 0 skipped. The original JUnit reports, pinned commits, file hashes, and actual import paths were cross-checked. The model-specific suites use randomly initialized parameters/inputs with offline mode enabled; no pretrained weights were downloaded for this run.

Scope: operator-level and model operator-integration regressions are now complete. Pretrained-model end-to-end accuracy and performance ABBA were not rerun. All historical accuracy and performance results below, including the two earlier measurements, are retained with their original revisions and environments; they are not new performance measurements for this revision.

To reproduce the shared operator run, execute from the tests/ directory of #7560 at 08dc20e68, with the compatible CUDA/Triton environment and repository test dependencies above. The large-storage-offset cases require slightly more than 4 GiB of free GPU memory:

python -m pytest -s -v diffusion/layers/test_fused_qk_norm_rope.py \
  -m "core_model and cuda" --run-level=core_model -p no:cacheprovider

Run the full operator-integration suite from the tests/ directory of this PR's pinned head:

python -m pytest -s -v diffusion/models/z_image/test_z_image_fused_qk_norm_rope.py \
  -m core_model --run-level=core_model -p no:cacheprovider

AI assistance for this synchronization: Codex resolved the merge, updated the indexing
regression coverage, ran the local checks and GPU regressions above, and verified the JUnit results.


Merge order update (2026-10-08): Any of the six model PRs can merge first. Each branch includes the same shared operator commit and only its own model integration.

Purpose

Follow-up to #7560, for the single-stream Z-Image transformer (ZImageAttention in the 30 main
blocks and the noise/context refiners): the per-site two RMSNorms + two RoPE passes become one
launch of the shared fused_qk_norm_rope(..., interleaved=True) op (#6982), writing the rotated
Q/K attention consumes directly.

  • One packed [cos | sin] table per attention site per forward (x, cap, unified), built
    from row 0 of the padded [B, S, D/2] cos/sin exactly as RotaryEmbedding applies them
    (_prepare_half_head_dim_cos_sin uses cos[0] for every batch element), in the activation dtype.
  • Sequence parallelism needs no special case. The refiner sites are not parallelized at all
    (_sp_plan shards only unified_prepare's outputs), and at the unified site the table is packed
    after that sharding, so cos/sin are already this rank's shard — the same coefficients the
    eager chain rotates this rank's tokens with. All three sites fuse under SP; verified on 2 GPUs
    below. Any unsupported dtype/geometry still keeps the eager chain. Default gate
    _FUSED_MIN_TOKENS = 0; env override unchanged.

Independent branch: shared operator commit plus this model integration, based directly on main.

Test Plan

vLLM Version: 0.28.0 (torch 2.13.0+cu130, Triton 3.7.1), on the last pre-vLLM-0.29 tree with
z_image/ and fused_qk_norm_rope.py from main — see #7560.

vLLM-Omni Commit: branch on 1b6cd28 + #7560. Hardware: 1× NVIDIA H200.

  • tests/diffusion/models/z_image/test_z_image_fused_qk_norm_rope.py: table follows cos[0],
    gate, CPU skip, and a table is still packed under a sequence-parallel forward context;
    ZImageAttention fused vs eager output.
  • Sequence-parallel check on 2× H200: the transformer (tiny random-init config, 2 main layers +
    both refiners) with real NCCL groups and the model's own _sp_plan hooks, world 1 vs world 2
    (ulysses_degree=2), gate off vs on in the same process, plus a runtime counter asserting the
    fused op really runs (and never runs in the gate-off arm).
  • End-to-end Tongyi-MAI/Z-Image-Turbo through Omni(model=..., mode="text-to-image"), 1024²,
    8 steps (Turbo default) and 28 steps, 3 prompts × 2 seeds × 2 repeats per arm, gate off vs on,
    A B B A, wall-clock per generate() including text encoder and VAE.

Test Result

Unit-test inventory: the module now contains 4 tests: CPU fallback,
row-zero/gate/dtype behavior, table packing under a sequence-parallel forward context,
and fused-vs-eager attention. Three require CUDA and Triton.

Review follow-up validation (2026-09-20): the current module returned
1 passed, 3 skipped on CPU (Python 3.12.3, vLLM 0.29.0, torch 2.13.0+cu130,
diffusers 0.40.0, transformers 5.14.1). All three skips explicitly require CUDA.

CUDA_VISIBLE_DEVICES= VLLM_TARGET_DEVICE=cpu python -m pytest -q -rs \
  tests/diffusion/models/z_image/test_z_image_fused_qk_norm_rope.py

The environment-variable reference now includes Z-Image's default gate of 0.
All applicable changed-file pre-commit hooks passed, including markdownlint and typos.
The CUDA unit suite was not rerun for this documentation-only follow-up; the H200
end-to-end and sequence-parallel measurements below are the original implementation
validation, not new measurements on the documentation follow-up.

End-to-end, 1024²:

steps eager p50 fused p50 per image
8 1.643 s 1.615 s −28 ms (−1.7%)
28 5.362 s 5.259 s −103 ms (−1.9%)

Peak allocation unchanged (21.6 GiB). The saving is modest because Z-Image only has the
single-stream chain (two norms + two RoPE per block, no cats) and 30-head × 128 blocks where
attention/MLP dominate; it is nevertheless free.

Sequence parallelism (H200, tiny random-init transformer: 2 main layers + both refiners, so 4
attention sites per forward).
Real NCCL groups and the model's own _sp_plan hooks; gate off vs
on in the same process; a runtime counter asserts the fused op actually runs (and never runs in the
gate-off arm):

parallelism ranks fused-op calls / forward (off → on) fused vs eager vs world-1 output eager fused saving
none 1 0 → 4 bitwise identical — 6.648 ms 6.167 ms −7.2%
ulysses 2 2 0 → 4 bitwise identical 0.465% 7.153 ms 6.609 ms −7.6%
ulysses 4 4 0 → 4 bitwise identical 0.470% 7.170 ms 6.604 ms −7.9%
ulysses 8 8 0 → 4 bitwise identical 0.534% 7.220 ms 6.670 ms −7.6%
ring 2 2 0 → 4 bitwise identical 0.085% 9.267 ms 8.781 ms −5.2%

All four sites (both refiners and the two main layers) take the fused path on every rank of every
configuration. The "vs world-1" column is the same number to every digit in the gate-off and
gate-on arms (e.g. 0.004646282005149976 for ulysses 2), i.e. it is sequence parallelism's own
reduction order, not the fused path. _sp_plan declares no auto_pad, so a sequence that does not
divide evenly is rejected before any of this; the shards are exact and cos/sin are split by the
same rule as the hidden states.

"Bitwise identical" here is this small configuration: the ≤1-ulp rounding difference described
under Numerics does not survive to the output at this depth. The comparison across parallelism
settings is the point — sequence parallelism changes neither the relationship between the two arms
nor the size of the saving. Absolute times are from that small model and are not an end-to-end
figure.

Numerics: arms bitwise deterministic across runs; eager vs fused images PSNR median 33.5 dB
(24 pairs), minimum 21.6 dB on one prompt/seed where the fox's pose shifts slightly — the same
≤1-ulp rounding-order difference of vLLM's rms_norm vs the kernel analysed in #7560, amplified
by sampling.

Control arm. A third arm runs the eager path with vLLM's RMSNorm swapped for a torch
implementation with the kernel's single-rounding order (sitecustomize patch in the worker,
nothing else changed). That arm vs eager: PSNR median 33.2 dB, min 20.8 dB (24 pairs) — the
same spread as fused vs eager (33.5 / 21.6). Fused vs the control arm: 33.2 / 23.6 dB. The
rounding order alone accounts for the whole difference.

🤖 Generated with Claude Code

https://claude.ai/code/session_01G8uaASBAbnS8TsFvC5gCUD

AI assistance for this review follow-up: Codex updated the gate documentation, checked the four-test inventory, ran the CPU and pre-commit checks above, and refreshed this validation record.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/diffusion_model_integration.md, docs/design/module/diffusion/index.md.

Module owners: @wtomin @RuixiangMa @david6666666

Routing: @wtomin via module of the changed files, module named in the PR description, semantic router, CODEOWNERS; @RuixiangMa via module of the changed files, module named in the PR description, semantic router; @david6666666 via module of the changed files, module named in the PR description, CODEOWNERS

@yuweih205, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 16, 2026
@yuweih205
yuweih205 force-pushed the perf/z-image-fused-qk-norm-rope branch from d7cdd26 to 73a881a Compare September 16, 2026 06:13
@yuweih205

Copy link
Copy Markdown
Contributor Author

Self-review: ZImageAttention takes an explicit qk_norm_rope_table and uses fused_qk_norm_rope(interleaved=True); one table per site (x, cap, unified) per forward from cos[0]/sin[0] as RotaryEmbedding applies them, only when the CUDA path would run. SP (via forward context) keeps the eager chain. Tests: table follows cos[0], CPU skip, attention fused vs eager. End-to-end on Tongyi-MAI/Z-Image-Turbo in the description (−1.7% / −1.9% per image; control arm shows RMSNorm rounding order only).

@yuweih205
yuweih205 force-pushed the perf/z-image-fused-qk-norm-rope branch from 73a881a to 4033cd9 Compare September 17, 2026 06:18
@Dong1017

Copy link
Copy Markdown
Contributor

Independent verification (H200, vllm 0.29.0 / torch 2.13.0+cu130 / Triton 3.7.1; e2e on 73a881af, units re-run on the current head after the force-push):

  • Unit: z_image tests 3/3 (description says "2 passed"; head has 3 test functions), op suite 20/20. No rope.py/kernel drift between the e2e base 1b6cd28 and main; the force-push's kernel/flux2 commits leave the Z-Image call path unchanged (the new packer params default to the previous behavior at this call site).
  • E2E Z-Image-Turbo 1024², 8 steps, 12 img/arm, gate off/on, A B B A: arms bitwise-deterministic, fused engaged 12/12, PSNR min 19.2 / median 28.7 / max 33.0 dB, p50 1.636 → 1.607 s (−1.75%) — matches the reported −28 ms (−1.7%).
  • SP=2 (Ulysses, 2×H200): both arms complete; an in-worker probe confirms sequence_parallel_size=2 is visible at every attention site, so the eager fallback engages as described. (SP runs are not bitwise-deterministic run-to-run — pre-existing, 18–37 dB across runs.)
  • NPU not run: designed no-op there (packer requires CUDA+bf16; NPU-eligible half-split path untouched).

Findings — nothing blocking (details inline): the SP check in the z_image helper duplicates the packer's new sequence_parallel_size gate / ForwardContext.sp_active and forfeits valid fusion at the refiner sites; the SP branch has no in-repo test; fuse-by-default (_FUSED_MIN_TOKENS = 0) now spans Boogu, Flux.2/klein and Z-Image — worth one explicit maintainer ratification.

RFC #7382 — fits as a consumer-adoption PR (Goal 3, no kernel surface added; the Phase-1 list names only H3/Boogu, covered for later consumers by the RFC's "subsequent work" clause).

Coordination — nothing blocking: linking RFC #7382/#7417/#7422 in the description would help public-entry reviewers, and this PR can follow #7560's rebase to the canonical diffusion.layers.ops import in the same pass (#7417 → #7560 → #7422).

Comment thread vllm_omni/diffusion/models/z_image/z_image_transformer.py Outdated
Comment thread tests/diffusion/models/z_image/test_z_image_fused_qk_norm_rope.py
Comment thread docs/configuration/environment_variables.md Outdated
Comment thread tests/diffusion/models/z_image/test_z_image_fused_qk_norm_rope.py
@hsliuustc0106 hsliuustc0106 added diffusion codes related to diffusion models Kernel optimization Codes related to optimize kernel execution to improve hardware utilization labels Sep 19, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 7 days

@yuweih205 this pull request has had no human commit, comment or review since 2026-09-21. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 7 days

@yuweih205 this pull request has had no human commit, comment or review since 2026-09-28. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on cursor (cursor-grok-4.6-high) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

Signed-off-by: HuangYuwei <yuweih205@gmail.com>
@yuweih205
yuweih205 force-pushed the perf/z-image-fused-qk-norm-rope branch from ee9f924 to 5522ee9 Compare October 8, 2026 04:15
@vllm-omni-review-bot

vllm-omni-review-bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit ffdb71041bae produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

Signed-off-by: HuangYuwei <yuweih205@gmail.com>
Signed-off-by: HuangYuwei <yuweih205@gmail.com>
@yuweih205
yuweih205 force-pushed the perf/z-image-fused-qk-norm-rope branch from 5522ee9 to 02984be Compare October 8, 2026 06:14
@yuweih205

Copy link
Copy Markdown
Contributor Author

@hsliuustc0106 Both of your review comments are addressed on the current head (c89e0b34): the environment-variable documentation explicitly lists Z-Image's default as 0, and the PR description records the four-test inventory and the reported results (4 passed on H200; 1 passed and 3 skipped on CPU). Both review threads are resolved. Could you please re-review when you have a chance?

The six-PR series (#7560, #7594, #7595, #7596, #7597, #7600) has also been restructured: each PR contains the same shared operator, tests and documentation changes, plus only its own model integration. There is no longer a prerequisite PR within this series; any one of the six can merge first, once its review and CI requirements are satisfied. This supersedes the earlier comments saying #7560 had to land first. The shared files are identical across the six heads, and local merge simulations passed for all 15 pairs and all 30 ordered pairs with the first PR squash-merged onto main.

@yuweih205

Copy link
Copy Markdown
Contributor Author

@hsliuustc0106 One additional numerical point for your re-review:

If the concern is that fusion changes floating-point reduction/rounding behavior, operator-level bitwise alignment with a fixed uncompiled eager reference is achievable in principle: we can adjust the kernel's reduction configuration, intermediate precision and rounding points to reproduce that reference.

The shared-operator experiment already demonstrated this for the tested vLLM 0.29 CUDA RMSNorm + RoPE contract on H200 (BF16, head dimension 128, epsilon 1e-6): an experimental single/joint candidate passed 42/42 strict Q/K/V bitwise checks. Its preparation latency was 9.0–10.8% higher for joint and 32.1–33.4% higher for single than the current fusion. The candidate is not included in the current PRs; these are operator-level results on the documented configurations, so each model/provider would need its own validation.

If strict eager operator parity is the preferred acceptance criterion, please let me know; I can adapt and validate that variant for this PR. Equality against an Inductor-compiled full model would require a separate check.

Signed-off-by: HuangYuwei <yuweih205@gmail.com>
@yuweih205
yuweih205 force-pushed the perf/z-image-fused-qk-norm-rope branch from c89e0b3 to ffdb710 Compare October 8, 2026 07:42
@yuweih205

Copy link
Copy Markdown
Contributor Author

Reproducible eager bitwise case: retained tiny-model results

The already measured tiny-model eager configuration has equal gate-off/gate-on output on the tensor checked by the original harness, with fusion active at all four attention sites. Its retained single-GPU transformer-forward timing is 6.647760 → 6.167152 ms, a 7.229623% saving. The original script, source overlay and 17 SP/rank records have now been recovered. These are September 17 measurements already summarized in this PR; no new GPU measurement is being reported, and the existing full pretrained-model performance results are retained.

Configuration change for this observed case: use the tiny random-weight constructor below, eager execution without torch.compile, TORCH_SDPA, BF16, and change VLLM_OMNI_FUSED_QK_NORM_ROPE_MIN_TOKENS from 1000000000 to 0. Runtime counters show 0 → 4 fused calls per forward on every rank, covering both refiners and two main layers.

ZImageTransformer2DModel(
    all_patch_size=(2,), all_f_patch_size=(1,), in_channels=16,
    dim=512, n_layers=2, n_refiner_layers=1,
    n_heads=4, n_kv_heads=4, cap_feat_dim=64,
).eval()

Exact settings:

  • H200 141 GB; Python 3.12.3, vLLM 0.28.0, Torch 2.13.0+cu130, CUDA 13.0, Triton 3.7.1, Diffusers 0.40.0. The retained environment metadata was read without executing the model.
  • Constructor seed 42. For each parameter in sorted named_parameters(), initialize normal_(mean=0, std=0.02) using a new CUDA generator seeded with 7 for that parameter, then cast the model to BF16. This generator-reset rule is part of the observed case.
  • Batch 2. One CPU generator with seed 123 creates both x tensors first, each [16,1,16,16], followed by both caption-feature tensors, each [32,64]; cast them to CUDA BF16. BF16 timestep [2] is filled with 0.5.
  • Explicit TORCH_SDPA; TP, PP, DP and CFG degrees are 1. The SP sweep uses the model's real _sp_plan hooks and NCCL groups: none, Ulysses 2/4/8, and ring 2. Compare the same model and inputs within each SP configuration.

Equality scope: the old harness checks the first returned image tensor only, after the complete tiny-model forward. It converts the BF16 tensor to FP32 and applies torch.equal; no integer-view signed-zero audit was recorded. All 17 retained SP/rank records report torch.equal=True, relative L2 = 0, maximum absolute difference = 0, and 4 fused calls. This does not establish equality for the second image, a pretrained full model, or an arbitrary input/model configuration. Outputs from different SP sizes differ numerically; the equality is gate-off versus gate-on within the same SP configuration.

The original rounded SP timings remain:

Parallelism Eager forward Fused forward Forward saving
None 6.648 ms 6.167 ms 7.2%
Ulysses 2 7.153 ms 6.609 ms 7.6%
Ulysses 4 7.170 ms 6.604 ms 7.9%
Ulysses 8 7.220 ms 6.670 ms 7.6%
Ring 2 9.267 ms 8.781 ms 5.2%

Timing uses CUDA-event medians, three warmups per arm, then A B B A with ten complete forwards per segment. These small-model timings are not end-to-end image-generation timings. The existing pretrained-model 1.7% / 1.9% performance figures remain separate; those pretrained outputs are not bitwise equal to the eager baseline.

Reproduction uses base ff6e906a25de1ca3122c0d0d5250fde5a0ddbc7b plus the retained eight-file overlay. Seven files match 6dc3f91353a2f879c5a0981788f148b29a252ced exactly; its environment-variable inventory removes VLLM_OMNI_ABORT_TIMEOUT and VLLM_OMNI_BREEZE_GOLDEN_DIR for old-base compatibility. See the staging instructions, exact harness, provenance, file manifest and all raw records.

PYTHONPATH="$RUN_DIR" ZSP_REPORT="$RUN_DIR/report.json" "$PYTHON" "$RUN_DIR/zimage_sp_check.py" '2:ulysses,4:ulysses,8:ulysses,2:ring'

The retained harness SHA256 is 4c258e9faeb5e7d1ff68507f7d5a3f801f2ab810b9e77d2e8fc8625abcb8deee; the raw report SHA256 is f36cec1c8cb5a69e27d25a3e25bfd45ef14c3a407a7cdd6977c8107d9007578a. These establish the stated historical eager case; they are not measurements of the October 8 heads.

AI assistance: Codex recovered and checked the existing harness/source hashes/raw records and prepared this additive reproduction note; no GPU case was rerun for this note.

@yuweih205

Copy link
Copy Markdown
Contributor Author

Pretrained eager generation: bitwise parity and measured acceleration

@hsliuustc0106 The numerical attribution and performance follow-up are complete. Matching the CUDA RMSNorm reduction order preserves bitwise output while retaining a full-request speedup in the tested Z-Image generation cases.

Tested implementation: prepared source cbe0e7f03f39f54b13e46ba3084fc085cf89425e over public base ffdb71041baeefe3242edcc4d36efbcab949fda5. This exact numerical preset is still outside the public PR head; the prepared shared patch and reproducible GPU evidence are available for review.

Configuration: pretrained Tongyi-MAI/Z-Image-Turbo, revision f332072aa78be7aecdf3ee76d5c247082da564a6, 1024×1024, guidance 0, seed 7, one fixed prompt (recorded in the raw JSON), 8/28 steps. Eager execution, H200 141 GB, BF16, head dimension 128, world/TP/SP size 1; vLLM 0.29.0, Torch 2.13.0+cu130, CUDA 13.0, Triton 3.7.1.

Original: MIN_TOKENS=1000000000000, NUMERICS=fast; exact: MIN_TOKENS=0, NUMERICS=vllm_cuda_128, under the VLLM_OMNI_FUSED_QK_NORM_ROPE_ prefix. The latter selects CUDA-order reduction and enable_fp_fusion=False. The model invokes the prepared shared operator directly.

Steps Original request (s) Bitwise preset (s) Latency reduction VAE output / RGB image
8 1.050633 1.002427 4.588% Both 0 byte mismatches
28 3.338334 3.186369 4.552% Both 0 byte mismatches

Wall request latency includes text encoding, all denoising steps, VAE and PIL image output; loading the model is excluded. One complete warmup request per arm precedes three original/exact/exact/original rounds, six timed requests per arm. Debug traces, tensor saves, provider hooks and counter wrappers are absent from timed calls.

The complete VAE float output and RGB uint8 image are compared by storage bytes before and after timing. Every arm repeats exactly. Actual fused calls are 0 for original and 272/952 for fusion at 8/28 steps. Every timed exact request is faster than every paired original sample in these two cases. This establishes parity and performance for the pinned checkpoint, prompt, seed and eager configurations above.

The preceding same-input investigation located the first difference at QK RMSNorm FP32 reduction: raw QKV and RoPE coefficients matched; CUDA and default Triton summation orders differed. Reproducing CUDA’s order closed the eager full-generation gap. The current fast arm’s separate parity and timing records are retained in the raw JSON.

Raw full-generation parity and ABBA samples, harness, and the reproduction procedure and all raw results.

AI assistance: Codex ran and checked the GPU investigation and performance measurements and prepared this additive evidence note.

Please review the prepared numerical preset and the measured eager cases when convenient.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (failed; retrying strict/cursor/cursor-grok-4.6-high in 120s (try 2 of 3)).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (failed; retrying strict/cursor/cursor-grok-4.6-high in 600s (try 3 of 3)).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (failed; falling back to direct/cursor/auto).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (failed; retrying strict/cursor/cursor-grok-4.6-high in 120s (try 2 of 3)).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (failed; retrying strict/cursor/cursor-grok-4.6-high in 600s (try 3 of 3)).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (failed; falling back to direct/cursor/auto).

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

Changes since the previous review

  • 1 new inline finding(s); 0 finding(s) below.

CI at ffdb71041bae (2026-10-10T18:07:13.168845+00:00): required check(s) blocking: buildkite/vllm-omni (missing).

Note: The assigned review arm strict/cursor/cursor-grok-4.6-high could not complete this review, so it was produced by the fallback arm direct/cursor/auto. It is excluded from the routing experiment.

Full review analysis

PR description

This change turns Z-Image's per-site Q/K RMSNorm and interleaved RoPE into one CUDA Triton launch of the shared fused_qk_norm_rope operator. Each forward packs one [cos | sin] table from row 0 of the padded cos/sin pair, matching the existing half-head RoPE broadcast, and passes it through the noise refiner, context refiner, and main blocks. The token gate defaults to 0, so CUDA bf16 sites fuse unless VLLM_OMNI_FUSED_QK_NORM_ROPE_MIN_TOKENS is raised; other devices and dtypes keep the eager chain. The same head also carries the shared joint QKV kernel, packer, and operator tests so this model PR can merge without waiting on the other consumers.

Change flow

flowchart LR
  A["[EXISTING] Z-Image x, cap, and unified sites"]:::existing
  B["[NEW] pack row-0 cos/sin table"]:::new
  C["[CHANGED] ZImageAttention fused gate"]:::changed
  D["[CHANGED] fused_qk_norm_rope Triton launch"]:::changed
  E["[EXISTING] vLLM RMSNorm then RoPE"]:::existing
  F["[NEW] Z-Image fusion tests"]:::new
  A --> B --> C
  C --> D
  C --> E
  D --> F
  E --> F
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

See inline comments below.


🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!

# RMSNorm -> RoPE chain; fuse by default (the fused path won at every size
# measured on H200 for this chain, see Flux.2 / Boogu-Image) and keep the
# gate for VLLM_OMNI_FUSED_QK_NORM_ROPE_MIN_TOKENS overrides.
_FUSED_MIN_TOKENS = 0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Default Z-Image fusion changes sampled images

Evidence and suggested fix

_FUSED_MIN_TOKENS = 0 packs a table for every CUDA bf16 forward, so ZImageAttention replaces vLLM RMSNorm plus apply_rope_to_qk with the Triton reduction. The in-repo attention check only allows atol/rtol 0.05, which still passes when Q/K differ by far more than one ulp. This PR's own Z-Image-Turbo record reports eager-versus-fused PSNR median 33.5 dB and minimum 21.6 dB, including a visible pose change, and the later CUDA-order bitwise preset is explicitly not in public head ffdb710. Set the Z-Image default above real sequence lengths, or land that measured reduction preset, so the fused launch does not become the default image contract.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion codes related to diffusion models high priority high priority issue, needs to be done asap Kernel optimization Codes related to optimize kernel execution to improve hardware utilization

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants