Repository navigation
[Kernel] Fuse LongCat paired Q/K RoPE - #7500
Conversation
Signed-off-by: dongbo910220 <1275604947@qq.com>
|
This PR appears to belong to: docs/design/module/diffusion/diffusion_model_integration.md, docs/design/module/diffusion/index.md. Module owners: @wtomin @Bounty-hunter @fhfuih Routing: @wtomin via module of the changed files, semantic router, CODEOWNERS; @Bounty-hunter via module of the changed files, semantic router; @fhfuih via module of the changed files, semantic router @dongbo910220, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot: three questions on the performance claim@dongbo910220 this PR reads as a performance or value claim:
Before the full evidence checklist, three short questions:
When you answer, the evidence that settles it is: base and head SHA, hardware, model, workload, warm-up and repeat count, mean or percentiles with their spread, and a correctness/quality-equivalence signal; an end-to-end claim also needs stage attribution. |
|
Self-review completed. I reviewed the full diff for scope and correctness, checked the eager-only eligibility and native fallbacks for compile, autograd, sequence parallelism, CUDA graph capture, small inputs, and unsupported layouts, verified the one-time bit-exactness guard and permanent failure fallback, and confirmed the targeted tests, local pre-commit checks, DCO, and recorded E2E and operator benchmark numbers. I also confirmed that the PR contains only the intended implementation and test files. |
| return output | ||
|
|
||
| reference = _apply_qk_rope_reference(query, key, rotary_pair) | ||
| if torch.equal(output[0], reference[0]) and torch.equal(output[1], reference[1]): |
There was a problem hiding this comment.
Keep full-output parity checks in tests; remove these inference-time GPU synchronizations.
There was a problem hiding this comment.
Removed the inference-time parity check, so eligible inputs now use the fused operator without native recomputation or torch.equal synchronization. Focused CUDA tests still verify bit-exact Q/K outputs and confirm that the fused path is exercised.
| # hardware-specific crossover without adding another environment variable. | ||
| _FUSED_MIN_TOKENS = 512 | ||
| _FUSED_QK_ROPE = HAS_TRITON and current_platform.is_cuda() | ||
| _VERIFIED_QK_ROPE_SIGNATURES: set[tuple] = set() |
There was a problem hiding this comment.
Add eviction limits to both signature caches before storing request-specific shapes.
There was a problem hiding this comment.
Removed the verified-signature cache together with the runtime parity check. The remaining launch-failure cache is capped at 128 entries with LRU eviction, and test_failed_qk_rope_signature_cache_is_bounded covers the limit and eviction.
Signed-off-by: dongbo910220 <1275604947@qq.com>
|
show visual comparision to make sure there is no quality degradation. |
|
@lishunyang12, here is a side-by-side visual comparison for the same fixed-seed LongCat-Image workload. The fusion-off and fusion-on 1024×1024 RGB PNGs are byte-identical, so no visual or pixel-level quality difference was observed. |
Signed-off-by: dongbo910220 <1275604947@qq.com> Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
Signed-off-by: dongbo910220 <1275604947@qq.com>

What does this PR do?
This PR adds a bit-exact Triton kernel that applies LongCat-Image's adjacent-interleaved RoPE to normalized query and key tensors together, replacing two eager RoPE launches with one fused launch on eligible single-GPU CUDA inference.
The optimization is intentionally limited to eager inference. Regional
torch.compile\ retain the existing native path.Test
all passed
Measured on 1× NVIDIA RTX PRO 6000 Blackwell Server Edition with
meituan-longcat/LongCat-Image, batch size 1, 1024×1024 output, 50 inference steps, guidance scale 4.5, and seed 42:B=1, S=4608, H=24, D=128)Design & Code Changes