Skip to content

[Kernel] Fuse LongCat paired Q/K RoPE - #7500

Merged
lishunyang12 merged 2 commits into
vllm-project:mainfrom
dongbo910220:perf/longcat-qk-rope-exact
Sep 17, 2026
Merged

lishunyang12 merged 2 commits into
vllm-project:mainfrom
dongbo910220:perf/longcat-qk-rope-exact

Conversation

@dongbo910220

@dongbo910220 dongbo910220 commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

What does this PR do?

This PR adds a bit-exact Triton kernel that applies LongCat-Image's adjacent-interleaved RoPE to normalized query and key tensors together, replacing two eager RoPE launches with one fused launch on eligible single-GPU CUDA inference.

The optimization is intentionally limited to eager inference. Regional torch.compile\ retain the existing native path.

Test

python -m pytest -q tests/diffusion/layers/test_fused_qk_rope.py
python -m pytest -q tests/diffusion/models/longcat_image/test_longcat_image_transformer.py
python -m pytest -q tests/diffusion/layers/test_fused_qk_norm_rope.py
python -m pytest -q tests/diffusion/models/boogu_image/test_boogu_image_transformer.py
python -m pytest -q tests/diffusion/cache/test_cache_dit.py

all passed

Measured on 1× NVIDIA RTX PRO 6000 Blackwell Server Edition with meituan-longcat/LongCat-Image, batch size 1, 1024×1024 output, 50 inference steps, guidance scale 4.5, and seed 42:

Benchmark Reference This PR Result
Eager E2E median 14.5265 s 12.9382 s 10.93% lower latency (1.1228×)
Paired Q/K RoPE median (B=1, S=4608, H=24, D=128) 634.9 μs 66.6 μs 89.51% lower latency (9.53×)

Design & Code Changes

  • Add a registered Triton custom op for paired full-width interleaved Q/K RoPE while preserving the native FP32 operation and BF16 rounding order.
  • Keep LongCat's existing RMSNorm operations unchanged and fuse only the following query and key RoPE work for eligible dual-stream and single-stream attention blocks.
  • Verify each eligible runtime signature against the complete native output once, then permanently fall back for any mismatch or execution error.
  • Preserve the native implementation for compile, capture, autograd, sequence-parallel, small-token, non-CUDA, and unsupported tensor paths.

Signed-off-by: dongbo910220 <1275604947@qq.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/diffusion_model_integration.md, docs/design/module/diffusion/index.md.

Module owners: @wtomin @Bounty-hunter @fhfuih

Routing: @wtomin via module of the changed files, semantic router, CODEOWNERS; @Bounty-hunter via module of the changed files, semantic router; @fhfuih via module of the changed files, semantic router

@dongbo910220, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: three questions on the performance claim

@dongbo910220 this PR reads as a performance or value claim:

  • claim: | Eager E2E median | 14.5265 s | 12.9382 s | 10.93% lower latency (1.1228×) |
  • claim: | Paired Q/K RoPE median (B=1, S=4608, H=24, D=128) | 634.9 μs | 66.6 μs | 89.51% lower latency (9.53×) |

Before the full evidence checklist, three short questions:

  1. Bottleneck — what is the current bottleneck, and which profile, trace or per-stage measurement shows it?
  2. Value — what does the change buy the user or the system (latency, throughput, memory, cost), and at which workload?
  3. A/B or ablation — is there a same-workload, same-head/config comparison that isolates each main claim on its own? For stacked optimizations, one number per item rather than a blended delta.

When you answer, the evidence that settles it is: base and head SHA, hardware, model, workload, warm-up and repeat count, mean or percentiles with their spread, and a correctness/quality-equivalence signal; an end-to-end claim also needs stage attribution.

@dongbo910220

Copy link
Copy Markdown
Contributor Author

Self-review completed. I reviewed the full diff for scope and correctness, checked the eager-only eligibility and native fallbacks for compile, autograd, sequence parallelism, CUDA graph capture, small inputs, and unsupported layouts, verified the one-time bit-exactness guard and permanent failure fallback, and confirmed the targeted tests, local pre-commit checks, DCO, and recorded E2E and operator benchmark numbers. I also confirmed that the PR contains only the intended implementation and test files.

@hsliuustc0106 hsliuustc0106 added Kernel optimization Codes related to optimize kernel execution to improve hardware utilization diffusion codes related to diffusion models labels Sep 15, 2026
return output

reference = _apply_qk_rope_reference(query, key, rotary_pair)
if torch.equal(output[0], reference[0]) and torch.equal(output[1], reference[1]):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Keep full-output parity checks in tests; remove these inference-time GPU synchronizations.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed the inference-time parity check, so eligible inputs now use the fused operator without native recomputation or torch.equal synchronization. Focused CUDA tests still verify bit-exact Q/K outputs and confirm that the fused path is exercised.

# hardware-specific crossover without adding another environment variable.
_FUSED_MIN_TOKENS = 512
_FUSED_QK_ROPE = HAS_TRITON and current_platform.is_cuda()
_VERIFIED_QK_ROPE_SIGNATURES: set[tuple] = set()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Add eviction limits to both signature caches before storing request-specific shapes.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed the verified-signature cache together with the runtime parity check. The remaining launch-failure cache is capped at 128 entries with LRU eviction, and test_failed_qk_rope_signature_cache_is_bounded covers the limit and eviction.

Signed-off-by: dongbo910220 <1275604947@qq.com>
@lishunyang12

Copy link
Copy Markdown
Collaborator

show visual comparision to make sure there is no quality degradation.

@dongbo910220

Copy link
Copy Markdown
Contributor Author

@lishunyang12, here is a side-by-side visual comparison for the same fixed-seed LongCat-Image workload. The fusion-off and fusion-on 1024×1024 RGB PNGs are byte-identical, so no visual or pixel-level quality difference was observed.

Side-by-side visual comparison of fusion disabled and fusion enabled LongCat-Image outputs

@lishunyang12 lishunyang12 added the ready label to trigger buildkite CI label Sep 17, 2026
@dongbo910220

Copy link
Copy Markdown
Contributor Author

@lishunyang12, Intel CI #8555 failed during checkout because the runner ran out of disk space. Could you retrigger it? I also saw the same failure in #5921 (Intel CI #8301) and #7544 (Intel CI #8342 and #8463), and it is tracked in #7587.

@lishunyang12
lishunyang12 merged commit 97912de into vllm-project:main Sep 17, 2026
8 of 9 checks passed
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
Signed-off-by: dongbo910220 <1275604947@qq.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Signed-off-by: dongbo910220 <1275604947@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion codes related to diffusion models Kernel optimization Codes related to optimize kernel execution to improve hardware utilization ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants