Skip to content

[Kernel][Ovis-Image] Fuse text/image QK RMSNorm + cat + RoPE into one Triton launch - #7600

Open
yuweih205 wants to merge 4 commits into
vllm-project:mainfrom
yuweih205:perf/ovis-image-fused-qk-norm-rope
Open

yuweih205 wants to merge 4 commits into
vllm-project:mainfrom
yuweih205:perf/ovis-image-fused-qk-norm-rope

Conversation

@yuweih205

@yuweih205 yuweih205 commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Independent merge update (2026-10-08)

Current head: 26a3eac0 on main c548a110; shared operator commits: bf9acacd and 24f9e25e (64-bit indexing). The six PRs #7560, #7594, #7595, #7596, #7597, and #7600 now each contain this identical shared commit plus only their own model integration. Any one of the six can merge first. The other five do not import code from that model.

The shared operator, test, and documentation files are byte-identical across all six heads. Local pre-commit checks passed (including ruff, mypy, SPDX, and markdownlint). Git merge simulations found no conflicts for all 15 pairs or for all 30 ordered pairs with the first PR squash-merged onto main. This Mac host has no CUDA/PyTorch test environment, so the GPU results below are historical results on earlier revisions; the new heads await CI and any GPU rerun.


Historical synchronization update (2026-09-21)

Current head: bbf831efa. Merged upstream main
1b87115c9ede2d29702c847b6da98f9e854fa593 while preserving commit history.
At that revision #7560 still landed first. The independent-merge update above supersedes this merge order.

  • Preserved main's large-storage-offset fix and regression test; extended the same
    64-bit row addressing to the joint QKV kernel. Added four CUDA regression cases
    covering either large-offset stream and both RoPE pairings. Floating-point
    arithmetic, model wiring, sequence-parallel behavior, and default gates are unchanged.
  • All applicable changed-file pre-commit hooks passed locally, including mypy,
    ruff, SPDX, test-mark checks, and markdownlint. A Python AST comparison also
    verified that the existing model integrations and shared floating-point operations
    were preserved (apart from widening address indices).

Operator-level GPU regression update (2026-09-21)

Operator-level correctness and regression tests for the current revision have now passed on real H200 hardware.

  • Environment: Python 3.12.12, vLLM 0.29.0, Torch 2.13.0+cu130, CUDA 13.0, and Triton 3.7.1. The repository's native pytest suites ran in a CUDA base image with an isolated venv. This is independent GPU validation, not a Buildkite CI Docker-image run or the earlier CPU-shim check.
  • Shared operator suite: the complete tests/diffusion/layers/test_fused_qk_norm_rope.py passed on [Kernel][Flux.2] Fuse text/image QK RMSNorm + cat + RoPE into one Triton launch #7560 at 08dc20e68: 30 passed. The 6 large-storage-offset cases also passed separately, including the 4 new dual-stream cases. Those 6 are included in the 30 and are not counted twice. The shared source and test files are byte-identical across the six current PR heads.
  • Coverage includes 64-bit address-index boundaries, joint/per-stream bitwise parity, eager-reference tolerances, the Flux.2 eager chain, torch.compile / CUDA graph, and shape, gate, and fallback behavior. Bitwise guarantees apply only to the corresponding exact-parity cases, not to every fused/eager comparison or the full model.
  • This PR's Ovis-Image operator-integration suite at bbf831efa: 2 passed, covering fused/eager attention comparisons for both double- and single-stream blocks. These are component/attention-level regressions, not full-model end-to-end tests.
  • Every suite above has 0 failures / 0 errors / 0 skipped. The original JUnit reports, pinned commits, file hashes, and actual import paths were cross-checked. The model-specific suites use randomly initialized parameters/inputs with offline mode enabled; no pretrained weights were downloaded for this run.

Scope: operator-level and model operator-integration regressions are now complete. Pretrained-model end-to-end accuracy and performance ABBA were not rerun. All historical accuracy and performance results below, including the two earlier measurements, are retained with their original revisions and environments; they are not new performance measurements for this revision.

To reproduce the shared operator run, execute from the tests/ directory of #7560 at 08dc20e68, with the compatible CUDA/Triton environment and repository test dependencies above. The large-storage-offset cases require slightly more than 4 GiB of free GPU memory:

python -m pytest -s -v diffusion/layers/test_fused_qk_norm_rope.py \
  -m "core_model and cuda" --run-level=core_model -p no:cacheprovider

Run the full operator-integration suite from the tests/ directory of this PR's pinned head:

python -m pytest -s -v diffusion/models/ovis_image/test_ovis_image_fused_qk_norm_rope.py \
  -m core_model --run-level=core_model -p no:cacheprovider

AI assistance for this synchronization: Codex resolved the merge, updated the indexing
regression coverage, ran the local checks and GPU regressions above, and verified the JUnit results.


Merge order update (2026-10-08): Any of the six model PRs can merge first. Each branch includes the same shared operator commit and only its own model integration.

Purpose

Follow-up to #7560 (Flux.2), same optimization for Ovis-Image (OvisImageAttention, 6 double +
27 single blocks): the double blocks' four RMSNorms + three torch.cats + two RoPE passes and
the single blocks' two RMSNorms + two RoPE passes each become one Triton launch whose outputs
are the joint [B, S_txt+S_img, H, D] Q/K/V attention consumes
, with no intermediate copies.

  • Double blocks call fused_joint_qkv_norm_rope (text first, as the eager cat); single blocks
    call fused_qk_norm_rope(..., interleaved=True) over the already-joint sequence.
  • One packed [cos | sin] table per forward from OvisImagePosEmbed's [S, D/2] output via the
    shared pack_qk_norm_rope_table helper, passed to the blocks through joint_attention_kwargs
    (the model forward now passes that argument; the blocks already accepted it).
  • Eager chain kept for unsupported dtype/geometry (bitwise identical to main otherwise); default
    gate _FUSED_MIN_TOKENS = 0, VLLM_OMNI_FUSED_QK_NORM_ROPE_MIN_TOKENS still overrides.

Independent branch: shared operator commit plus this model integration, based directly on main.

Test Plan

vLLM Version: 0.28.0 (torch 2.13.0+cu130, Triton 3.7.1), on the last pre-vLLM-0.29 tree with
ovis_image/ (byte-identical to main) and fused_qk_norm_rope.py from main — see #7560.

vLLM-Omni Commit: branch on 1b6cd28 + #7560. Hardware: 1× NVIDIA H200.

  • tests/diffusion/models/ovis_image/test_ovis_image_fused_qk_norm_rope.py: OvisImageAttention
    fused vs eager output, double-stream and single-stream configuration.
  • End-to-end AIDC-AI/Ovis-Image-7B through Omni(model=..., mode="text-to-image"), true CFG 5.0
    with a negative prompt, 1024², 50 and 20 steps, 3 prompts × 2 seeds × 2 repeats per arm, gate
    off vs on, A B B A, plus a control arm (eager path with vLLM's RMSNorm swapped for a torch
    single-rounding implementation). Wall-clock per generate() including text encoder and VAE.

Test Result

Unit tests: 2 passed on H200 (plus #7560's op suite).

End-to-end, 1024², true CFG 5.0 (B=2 through the transformer):

steps eager p50 fused p50 per image per denoising step
50 4.68 s 4.51 s −0.17 s (−3.5%) 91.4 → 88.0 ms (−3.7%)
20 1.94 s 1.87 s −0.06 s (−3.3%)

Fixed cost outside the denoising loop ≈ 0.11 s; peak allocation unchanged (19.5 GiB).

Numerics:

  • Images, eager vs fused on identical (prompt, seed): PSNR median 48.8 dB (24 pairs), minimum
    40.0 dB — visually identical.
  • Control arm (eager path with vLLM's RMSNorm swapped for a torch single-rounding
    implementation, nothing else changed) vs eager: 47.6 / 41.9 dB; fused vs control: 47.7 / 40.6 dB.
    The rounding order alone accounts for the whole difference, as in [Kernel][Flux.2] Fuse text/image QK RMSNorm + cat + RoPE into one Triton launch #7560.
  • Run-to-run reproducibility: the eager arm was bitwise identical across processes. Of five
    fused-arm processes (two jobs), four were bitwise identical to each other and one (the first
    of the first job) differed from them at 41–52 dB — the same ulp scale as torch.compile's own
    compiled-vs-enforce_eager variation measured on the eager path (38.6 / 45.8 dB), so it
    looks like an Inductor autotuning choice rather than the kernel (our Triton kernel has a fixed
    launch configuration and is deterministic). With enforce_eager=True the fused path was
    bitwise reproducible across processes.

🤖 Generated with Claude Code

https://claude.ai/code/session_01G8uaASBAbnS8TsFvC5gCUD

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/diffusion_model_integration.md, docs/design/module/diffusion/index.md.

Module owners: @wtomin @RuixiangMa @david6666666

Routing: @wtomin via module of the changed files, module named in the PR description, semantic router, CODEOWNERS; @RuixiangMa via module of the changed files, module named in the PR description, semantic router; @david6666666 via module of the changed files, module named in the PR description, CODEOWNERS

@yuweih205, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

Comment thread vllm_omni/diffusion/models/ovis_image/ovis_image_transformer.py
@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 16, 2026
yuweih205 added a commit to yuweih205/vllm-omni that referenced this pull request Sep 16, 2026
…l cannot run

Same gate as the shared helper (review on vllm-project#7600): no [B*S, D] allocation
or copy on CPU/NPU/ROCm or non-bf16 activations.

Signed-off-by: HuangYuwei <yuweih205@gmail.com>
@yuweih205
yuweih205 force-pushed the perf/ovis-image-fused-qk-norm-rope branch from 35d7b9a to a2b44d0 Compare September 16, 2026 06:13
@yuweih205

yuweih205 commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor Author

@RuixiangMa Thanks, done in #7560: pack_qk_norm_rope_table now returns None (no allocation) unless the CUDA kernel would run, so Flux.2/klein and this PR are covered through the rebase.

@yuweih205

Copy link
Copy Markdown
Contributor Author

Self-review: OvisImageAttention uses the joint op in double blocks and fused_qk_norm_rope(interleaved=True) in single blocks; the model forward builds joint_attention_kwargs only when pack_qk_norm_rope_table returns a table (CUDA + bf16, per @RuixiangMa's comment). Eager chain untouched otherwise. Tests: attention fused vs eager, double- and single-stream. End-to-end on AIDC-AI/Ovis-Image-7B in the description (−3.5% / −3.3% per image; control arm and cross-process reproducibility note included).

@Dong1017

Copy link
Copy Markdown
Contributor

Reviewed at a2b44d0 (delta over #7560: ovis_image_transformer.py + the new test only). Kernel-level: sound. The interleaved tile keeps one documented rounding contract (bf16-rounded normalized value, fp32 rotation, one final rounding), supports any even rotary_dim ≤ head_dim ≤ 256 with padded lanes provably out-of-domain, takes explicit strides for the chunked-QKV views (zero-copy merge), allocates fresh outputs (mutates_args=[]), and is deterministic (fixed launch config, no atomics, no autotune). Double-block order (text-first in, joint out) matches every eager cat/split site; the per-forward table gate and per-block _fused_cuda_supported share one predicate, so CPU/NPU/ROCm/non-bf16 stay bitwise-main eager. Verified on H200: op suite 20 passed, new test 2 passed, model-level probe shows the table packed once per forward and both ops live (strict shape validation active). A/B evidence is comparable (same-head toggle, A B B A, control arm); the one divergent fused process is plausibly Inductor autotune, bounded by the enforce_eager control.

Non-blocking: pack_qk_norm_rope_table is called without head_dim, so its gate assumes head_dim == rotary_dim (fine for every Ovis config; passing it would make the gate exact if a variant ever exceeds 256). Coordination for #7417/#7422: the legacy-module imports here plus the second torch.ops.vllm_omni registration join the shim surface, and fused_qk_norm_rope_available vs #7422's fused_qk_norm_rope_supported should converge to one API.

Comment thread vllm_omni/diffusion/models/ovis_image/ovis_image_transformer.py
yuweih205 added a commit to yuweih205/vllm-omni that referenced this pull request Sep 17, 2026
…l cannot run

Same gate as the shared helper (review on vllm-project#7600): no [B*S, D] allocation
or copy on CPU/NPU/ROCm or non-bf16 activations.

Signed-off-by: HuangYuwei <yuweih205@gmail.com>
@yuweih205
yuweih205 force-pushed the perf/ovis-image-fused-qk-norm-rope branch 2 times, most recently from b4590fe to e1dc249 Compare September 17, 2026 11:05
@hsliuustc0106 hsliuustc0106 added the Kernel optimization Codes related to optimize kernel execution to improve hardware utilization label Sep 22, 2026

@0z5a 0z5a left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cross-architecture validation on an RTX 5090 against bbf831e (torch 2.13.0+cu130, BF16): the two PR attention tests and an additional one-double/one-single-block OvisImageTransformer2DModel.forward test passed (3/3). The model-level test exercised one RoPE-table pack, one fused joint attention call, and one fused single-stream call; fused and eager outputs matched at atol=rtol=0.05.

Median synchronized wall time for the same synthetic transformer, 12 runs per path:

Image tokens Total tokens Eager (ms) Fused (ms) Speedup
16 32 4.933 3.874 1.27×
64 80 5.193 3.908 1.33×
256 272 4.992 3.905 1.28×
1024 1040 5.034 3.933 1.28×

These are RTX 5090 synthetic transformer results; I did not run a pretrained Ovis checkpoint or measure H200 performance. I found no issue in the tested path. LGTM.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 7 days

@yuweih205 this pull request has had no human commit, comment or review since 2026-09-28. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on zcode (GLM-5.3-Flash) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

Signed-off-by: HuangYuwei <yuweih205@gmail.com>
@yuweih205
yuweih205 force-pushed the perf/ovis-image-fused-qk-norm-rope branch from bbf831e to 237a36a Compare October 8, 2026 04:15
@vllm-omni-review-bot

vllm-omni-review-bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit aba7d60a3074 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

Signed-off-by: HuangYuwei <yuweih205@gmail.com>
Signed-off-by: HuangYuwei <yuweih205@gmail.com>
@yuweih205
yuweih205 force-pushed the perf/ovis-image-fused-qk-norm-rope branch from 237a36a to 26a3eac Compare October 8, 2026 06:14
@0z5a

0z5a commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

I also do some optimization on these models, will check yours soon.

If your work looks better and more brief, my optimization will be superseded

@yuweih205

Copy link
Copy Markdown
Contributor Author

@RuixiangMa Your table-allocation feedback is addressed on the current head (96ec51b2): the shared pack_qk_norm_rope_table helper returns None before allocating when the fused CUDA path cannot run, so Ovis-Image and Flux.2/klein retain their eager fallback without constructing that table. The review thread is resolved. This fix is now included directly in this PR. Could you please re-review when you have a chance?

The six-PR series (#7560, #7594, #7595, #7596, #7597, #7600) has also been restructured: each PR contains the same shared operator, tests and documentation changes, plus only its own model integration. There is no longer a prerequisite PR within this series; any one of the six can merge first, once its review and CI requirements are satisfied. This supersedes the earlier comments saying #7560 had to land first. The shared files are identical across the six heads, and local merge simulations passed for all 15 pairs and all 30 ordered pairs with the first PR squash-merged onto main.

@yuweih205

Copy link
Copy Markdown
Contributor Author

@RuixiangMa One additional numerical point for your re-review:

If the concern is that fusion changes floating-point reduction/rounding behavior, operator-level bitwise alignment with a fixed uncompiled eager reference is achievable in principle: we can adjust the kernel's reduction configuration, intermediate precision and rounding points to reproduce that reference.

The shared-operator experiment already demonstrated this for the tested vLLM 0.29 CUDA RMSNorm + RoPE contract on H200 (BF16, head dimension 128, epsilon 1e-6): an experimental single/joint candidate passed 42/42 strict Q/K/V bitwise checks. Its preparation latency was 9.0–10.8% higher for joint and 32.1–33.4% higher for single than the current fusion. The candidate is not included in the current PRs; these are operator-level results on the documented configurations, so each model/provider would need its own validation.

If strict eager operator parity is the preferred acceptance criterion, please let me know; I can adapt and validate that variant for this PR. Equality against an Inductor-compiled full model would require a separate check.

Signed-off-by: HuangYuwei <yuweih205@gmail.com>
@yuweih205
yuweih205 force-pushed the perf/ovis-image-fused-qk-norm-rope branch from 96ec51b to aba7d60 Compare October 8, 2026 07:42
@yuweih205

yuweih205 commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor Author

Supplemental eager bitwise measurements and queued production check

Existing PR performance results remain as recorded. The retained operator scripts and raw results stay pinned to their original artifact revision.

Measured operator case: H200; B1, 256 text + 4096 image tokens, 24 heads, head dimension 128; BF16 activations/weights/RoPE table, full interleaved RoPE, epsilon 1e-6, seed 20261008, world/TP/SP size 1. The candidate has zero integer-view Q/K bit mismatches against the original eager CUDA RMSNorm/RoPE chain; routed V also matches exactly. Original and candidate repeat checks pass.

Path Original chain (ms) Current fusion (ms) Reference candidate (ms) Candidate saving vs original Candidate extra latency vs current
Joint 0.306889 0.106426 0.106892 65.17% 0.44%
Single 0.185626 0.066534 0.067409 63.69% 1.32%

Eager CUDA-event medians: six balanced-order samples, 20 warmup calls, 100 calls per sample. These durations cover Q/K(/V) preparation with the packed RoPE table reused; table packing is excluded. They establish operator parity and latency for this case.

Prepared production setting: change VLLM_OMNI_FUSED_QK_NORM_ROPE_NUMERICS=fast to vllm_cuda_128; this resolves rms_norm_reduction="triton" → "vllm_cuda_128" and enable_fp_fusion=True → False. The common patch and validation records are prepared locally and have not been pushed to the six PRs. The retained table measured the isolated candidate.

Actual r3 target: prepared source 53b1d8da4cf07fc40f55613a381f7b401873bbea, using model_production_bitwise.py, SHA256 1d4eaff2c5de50b0c09a098adfab736bd936dba1c01653f870bd7dbb4b3efed9. The model's imported production functions are wrapped only for counting and receive unchanged arguments; no experimental operator is substituted.

The three eager arms use the same initialized model/state/inputs:

Arm MIN_TOKENS NUMERICS
Original chain 1000000000000 fast
Current fusion 0 fast
CUDA reference preset 0 vllm_cuda_128

H200, vLLM 0.29.0, Torch 2.13.0/CUDA 13.0, Triton 3.7.1; BF16, head dimension 128, B1/B2, seed 20261008, world/TP/SP size 1; two dual and two single blocks, 64 configured text tokens and a 16×16 image/video grid. Original RMSNorm is pinned to vllm_c; the harness records source/operator hashes and actual RMSNorm/RoPE forwards/providers. Complete output storage bits must match between the CUDA reference and original; every arm must repeat exactly and meet its fused-call counts before eager ABBA timing (20 warmup calls, 20 calls/sample, three rounds), including model RoPE generation and per-forward table packing.

Source manifests and reproduction instructions and the eight-worker runner pin the settings. The eight-GPU r3 job is queued in MOVA2.0纯交付分区; full-model parity/performance is not yet established.

Copy link
Copy Markdown
Contributor Author

Thanks for checking! This optimization was also validated on our internal diffusion workloads before upstreaming to vLLM-Omni. Feel free to take a look at the implementation — happy to discuss any suggestions or potential improvements.

@yuweih205

Copy link
Copy Markdown
Contributor Author

Fresh eager measurements: bitwise reference parity with retained speedup

@RuixiangMa The H200 performance follow-up is now complete. The measured CUDA-reference numerical preset has zero output storage-byte mismatches versus the original eager chain, with a measurable transformer-forward saving.

Tested implementation: prepared source 53b1d8da4cf07fc40f55613a381f7b401873bbea. The numerical preset patch is still outside the public PR head; the exact shared patch is retained for review.

Settings: MIN_TOKENS=1000000000000, NUMERICS=fast for original; MIN_TOKENS=0, NUMERICS=vllm_cuda_128 for exact fusion, using the VLLM_OMNI_FUSED_QK_NORM_ROPE_ prefix. The preset selects rms_norm_reduction="vllm_cuda_128" and enable_fp_fusion=False. Model calls use the prepared operator directly.

Scope: eager, random-weight, full-width shallow Ovis-Image transformer, 24 heads × 128 dimensions, BF16, 4096 image/video tokens and 512 configured text tokens, B1/B2, seed 20261008, world/TP/SP size 1. Two joint/dual blocks and two single blocks.
These timings include the complete shallow transformer forward, model RoPE generation and per-forward table packing. Text encoding, sampling and VAE are outside this measurement.

Batch Original chain (ms) Bitwise preset (ms) Forward latency reduction Output storage bits
1 15.091319 14.552750 3.569% 0 byte mismatches
2 29.961728 29.195374 2.558% 0 byte mismatches

The original RMSNorm provider is pinned and checked as vllm_c. Each arm repeats exactly; exact versus original is checked before and after timing. Source/operator hashes, state/input hashes, real norm/RoPE dispatch and fused-call counts are recorded. Provider hooks and counter wrappers are removed during timing.

Timing: H200 141 GB, vLLM 0.29.0, Torch 2.13.0+cu130, CUDA 13.0, Triton 3.7.1; CUDA-event medians, 10 warmup forwards per arm, three ABBA rounds, 10 calls/sample, six samples/arm. These are the stated random-weight eager cases; they do not establish pretrained full image/video generation parity.

Raw records: B1, B2. See the reproduction procedure and all raw results. The fast arm and its separate timing/parity records are retained in those JSON files.

AI assistance: Codex ran and checked the GPU measurements and prepared this additive evidence note.

Please review the prepared numerical preset and the measured eager cases when convenient.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: d05ee350-2109-49a0-82ba-b1de2e29212a) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 120s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 161fec01-cd8b-4c51-9818-3de0a54afb59) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 600s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 7752e74d-c475-4109-900a-9a5208eb50c3) — check zcode login and the CLI version; falling back to direct/cursor/auto).

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

Changes since the previous review

  • 0 new inline finding(s); 0 finding(s) below.

CI at aba7d60a3074 (2026-10-10T11:36:40.024473+00:00): required check(s) blocking: buildkite/vllm-omni (missing).

Note: The assigned review arm strict/zcode/GLM-5.3-Flash could not complete this review, so it was produced by the fallback arm direct/cursor/auto. It is excluded from the routing experiment.

Full review analysis

PR description

Ovis-Image attention now builds one packed [cos | sin] table per forward and uses it in every block. Double-stream blocks launch the new joint Triton op so text and image Q/K RMSNorm, the text-then-image concatenation, and RoPE are written straight into the joint Q, K, and V that attention reads. Single-stream blocks run the existing interleaved fused op on the already-concatenated sequence. CUDA bf16 forwards take that path by default (_FUSED_MIN_TOKENS = 0); other devices, dtypes, and geometries keep the eager RMSNorm, cat, and RoPE chain, and VLLM_OMNI_FUSED_QK_NORM_ROPE_MIN_TOKENS can raise the token gate.

Change flow

flowchart LR
  pos["[EXISTING] OvisImagePosEmbed<br/>txt then img cos/sin"]:::existing
  pack["[NEW] pack_qk_norm_rope_table"]:::new
  attn["[CHANGED] OvisImageAttention"]:::changed
  joint["[NEW] fused_joint_qkv_norm_rope"]:::new
  single["[CHANGED] fused_qk_norm_rope<br/>interleaved joint sequence"]:::changed
  eager["[EXISTING] Eager RMSNorm cat RoPE"]:::existing
  out["[EXISTING] Attention consumes Q/K/V"]:::existing
  pos --> pack
  pack --> attn
  attn --> joint
  attn --> single
  attn --> eager
  joint --> out
  single --> out
  eager --> out
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

high priority high priority issue, needs to be done asap Kernel optimization Codes related to optimize kernel execution to improve hardware utilization

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants