Skip to content

[rotary] Fix the fused Qwen3.5 RoPE kernel discarding mrope height and width - #34446

Merged
BBuf merged 6 commits into
sgl-project:mainfrom
modal-projects:jwiem/fix-mrope-fused-rope
Aug 30, 2026
Merged

BBuf merged 6 commits into
sgl-project:mainfrom
modal-projects:jwiem/fix-mrope-fused-rope

Conversation

@jason136

@jason136 jason136 commented Aug 11, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

fused_qk_gemma_rmsnorm_rope_gate loads one position per token,
tl.load(positions_ptr + token). Multimodal Qwen3.5 passes mrope positions shaped
[3, T], one row per axis, so that offset only ever lands in row 0: height and width
were discarded, and every image token rotated as if it sat at its temporal position on
all three axes.

Text tokens hold the same position on all three rows, so they were already correct and no
text benchmark could catch this. The path is not behind a flag —
Qwen3_5AttentionDecoderLayer.self_attention takes it whenever
_is_cuda and self.attn_output_gate (default True) — so
Qwen3_5ForConditionalGeneration and Qwen3_5MoeForConditionalGeneration on CUDA were
affected for every request carrying an image. forward_native, the HIP/XPU/CPU/NPU
branches, and text-only users of the same layer (qwen3_5_mtp, minicpmv,
interns2preview) are unaffected.

Modifications

  • fused_qk_rmsnorm_rope_gate.py: the kernel takes mrope_axis_map and an
    MROPE: tl.constexpr; each rotary lane loads the axis that owns it and indexes
    positions by that row. The row stride comes from the tensor rather than T, since
    CUDA-graph decode replays mrope_positions[:, :num_tokens] out of a
    [3, max_num_token] buffer.
  • MRotaryEmbedding._build_axis_map: generalized from GLM-interleaved only to all three
    layouts, which differ only in which axis owns which lane.
  • _legacy_axis_map: triton_mrope_fused and the out-of-tree
    sgl_kernel.multimodal_rotary_embedding keep receiving None outside GLM. This PR
    changes no input to either.
  • Ernie4_5_VLRotaryEmbedding._build_axis_map returns None: it reads mrope_section
    as h, w, t while the shared builder assumes t, h, w.
  • qwen3_5.py: passes the map when positions are 2-D.

Accuracy Tests

Both kernels against MRotaryEmbedding.forward_native, at Qwen3.6-35B-A3B's shape
(16 q heads, 2 kv heads, head_dim 256, rotary_dim 64) and its shipped interleaved
mrope_section [11, 11, 10], bf16, with error grouped by the axis owning each lane:

positions kernel max abs error t-lanes h-lanes w-lanes
distinct t, h, w stock 19.875 0.0625 5.891 19.875
distinct t, h, w this PR 0.0625 0.0625 0.031 0.0625
t == h == w stock 0.0625 0.0625 0.031 0.0625

0.0625 is the bf16 floor here. Stock sits on it for temporal lanes and up to 300x above
it for height and width, which is one row applied and two ignored; this PR brings every
lane to the floor. The last row is why it shipped.

New test_fused_qk_rmsnorm_rope_gate.py compares the kernel with forward_native over
interleaved and contiguous [11, 11, 10] and [24, 20, 20] (no pass-through tail), with
positions sliced out of a wider buffer to reproduce the CUDA-graph stride, plus 1-D parity
and rejection of both illegal position/map pairings. New test_mrope_axis_map.py checks
each layout against the function that consumes it, and pins GLM's order (its consumer is
out of tree) and Ernie's opt-out.

End to end, which is what shows the model is wired to the map: Qwen/Qwen3.6-35B-A3B at
995ad96, 1xB200, greedy, three words asked for separately on a generated page. The
reference column disables the fused branch, dropping CUDA to the plain path that ropes
through MRotaryEmbedding — same weights and hardware, only the rope kernel differs.

word reference (unfused) this PR before
AURORA 680, 509, 781, 522 680, 509, 781, 523 800, 480, 900, 492
BASALT 640, 508, 717, 524 640, 509, 719, 528 800, 480, 900, 494
CINDER 760, 508, 810, 520 752, 508, 810, 520 800, 487, 882, 501

Before, all three share x1 = 800: the word being asked about does not move the box.
After, each tracks the unfused path within 8 px. This is parity between the two paths, not
a grounding score.

Speed Tests and Profiling

MROPE is a tl.constexpr, so the 1-D path keeps its own specialization. Same shape,
triton.testing.do_bench, median per launch on 1xB200:

tokens 1-D before 1-D after mrope 1-D delta
3072 (one page of prefill) 65.4 us 65.4 us 65.5 us -0.08%
32 (decode graph batch) 8.2 us 8.2 us 8.3 us +0.11%

The mrope specialization costs one int64 load per lane, with no extra launches,
synchronization or allocations.


CI States

Latest PR Test (Base): ❌ Run #33127023374
Latest PR Test (Extra): ✅ Run #33127023235
Latest PR Test (AMD ROCm 7.2): ❌ Run #33127023285

@gongy

gongy commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator

/rerun-test test/registered/kernels/ops/attention/test_fused_qk_rmsnorm_rope_gate.py test/registered/rotary/test_mrope_axis_map.py test/registered/cpu/test_rope.py test/registered/kernels/ops/attention/test_rope.py test/registered/unit/multimodal/test_mrope_encoder_utils.py test/registered/rotary/test_rope_cache_invalidation.py test/registered/unit/multimodal/rust/qwen/test_token_layout_mrope.py test/registered/vlm/test_vision_openai_server_a.py

@github-actions

github-actions Bot commented Aug 14, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/kernels/ops/attention/test_fused_qk_rmsnorm_rope_gate.py test/registered/rotary/test_mrope_axis_map.py test/registered/cpu/test_rope.py test/registered/kernels/ops/attention/test_rope.py test/registered/unit/multimodal/test_mrope_encoder_utils.py test/registered/rotary/test_rope_cache_invalidation.py test/registered/unit/multimodal/rust/qwen/test_token_layout_mrope.py test/registered/vlm/test_vision_openai_server_a.py:

⛔ test/registered/kernels/ops/attention/test_fused_qk_rmsnorm_rope_gate.py, test/registered/kernels/ops/attention/test_rope.py, test/registered/vlm/test_vision_openai_server_a.py: Dispatch failed: 422

⛔ test/registered/rotary/test_mrope_axis_map.py, test/registered/cpu/test_rope.py, test/registered/unit/multimodal/test_mrope_encoder_utils.py, test/registered/rotary/test_rope_cache_invalidation.py, test/registered/unit/multimodal/rust/qwen/test_token_layout_mrope.py: Dispatch failed: 422

@jason136
jason136 force-pushed the jwiem/fix-mrope-fused-rope branch from c155342 to db41e22 Compare August 14, 2026 19:52
@jason136
jason136 force-pushed the jwiem/fix-mrope-fused-rope branch from db41e22 to ac694cd Compare August 17, 2026 18:50
@gongy

gongy commented Aug 17, 2026

Copy link
Copy Markdown
Collaborator

/rerun-test test/registered/kernels/ops/attention/test_fused_qk_rmsnorm_rope_gate.py test/registered/rotary/test_mrope_axis_map.py test/registered/cpu/test_rope.py test/registered/kernels/ops/attention/test_rope.py test/registered/unit/multimodal/test_mrope_encoder_utils.py test/registered/rotary/test_rope_cache_invalidation.py test/registered/unit/multimodal/rust/qwen/test_token_layout_mrope.py test/registered/vlm/test_vision_openai_server_a.py

@github-actions

github-actions Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/kernels/ops/attention/test_fused_qk_rmsnorm_rope_gate.py test/registered/rotary/test_mrope_axis_map.py test/registered/cpu/test_rope.py test/registered/kernels/ops/attention/test_rope.py test/registered/unit/multimodal/test_mrope_encoder_utils.py test/registered/rotary/test_rope_cache_invalidation.py test/registered/unit/multimodal/rust/qwen/test_token_layout_mrope.py test/registered/vlm/test_vision_openai_server_a.py:

🚀 1-gpu-h100 (3 tests): ✅ View workflow run

cd test/ && python3 registered/kernels/ops/attention/test_fused_qk_rmsnorm_rope_gate.py
cd test/ && python3 registered/kernels/ops/attention/test_rope.py
cd test/ && python3 registered/vlm/test_vision_openai_server_a.py

🚀 ubuntu-latest (5 tests): ❌ View workflow run

cd test/ && python3 registered/rotary/test_mrope_axis_map.py
cd test/ && python3 registered/cpu/test_rope.py
cd test/ && python3 registered/unit/multimodal/test_mrope_encoder_utils.py
cd test/ && python3 registered/rotary/test_rope_cache_invalidation.py
cd test/ && python3 registered/unit/multimodal/rust/qwen/test_token_layout_mrope.py

@jason136
jason136 force-pushed the jwiem/fix-mrope-fused-rope branch from 71c5338 to 0929d4c Compare August 17, 2026 21:51
@gongy

gongy commented Aug 17, 2026

Copy link
Copy Markdown
Collaborator

/rerun-test test/registered/rotary/test_mrope_axis_map.py test/registered/kernels/ops/attention/test_fused_qk_rmsnorm_rope_gate.py test/registered/vlm/test_vision_openai_server_a.py

@github-actions

github-actions Bot commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/rotary/test_mrope_axis_map.py test/registered/kernels/ops/attention/test_fused_qk_rmsnorm_rope_gate.py test/registered/vlm/test_vision_openai_server_a.py:

🚀 ubuntu-latest (1 test): ✅ View workflow run

cd test/ && python3 registered/rotary/test_mrope_axis_map.py

🚀 1-gpu-h100 (2 tests): ✅ View workflow run

cd test/ && python3 registered/kernels/ops/attention/test_fused_qk_rmsnorm_rope_gate.py
cd test/ && python3 registered/vlm/test_vision_openai_server_a.py

@BBuf BBuf left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Great fix, thanks.

@BBuf BBuf added run-ci CI: run the baseline test suite on this PR bypass-fastfail labels Aug 26, 2026
@BBuf BBuf added the run-ci-extra CI: also run the extra suite (requires run-ci) label Aug 26, 2026
@BBuf
BBuf merged commit e635577 into sgl-project:main Aug 30, 2026
181 of 205 checks passed
@jason136
jason136 deleted the jwiem/fix-mrope-fused-rope branch August 31, 2026 18:10
TobyMint added a commit to TobyMint/sglang that referenced this pull request Sep 1, 2026
…gl-project#34446)

The fused kernel loaded a single position per token, so with image
inputs ([3, T] temporal/height/width positions) every rotary lane
silently read the temporal row — wrong RoPE on image tokens in all
full-attention layers of Qwen3.5/3.8 hybrids. Text was unaffected
(the three rows coincide), so this never shows in text-only smoke.

Port: the kernel takes an mrope_axis_map ([rotary_dim//2] lane->axis)
and reads positions[axis[lane], t] when positions is 2-D;
MRotaryEmbedding now builds the axis map for every mrope_section style
(contiguous, interleaved, GLM round-robin) instead of GLM only, while
the legacy sgl_kernel call sites keep the GLM-only map via
_legacy_axis_map. Unit test mirrors the kernel math bitwise for 1-D
and mrope positions and checks both axis-map styles.
StevenChenSE pushed a commit to StevenChenSE/sglang that referenced this pull request Sep 6, 2026
efschu pushed a commit to efschu/htsglang that referenced this pull request Sep 26, 2026
…cks, chronologisch)

Grundlage: Präsenz-Scan aller 230 27B-Commits seit 76f8deb gegen diesen Baum
(Stichprobe der hinzugefügten Zeilen je Commit); die 94 fehlenden minus die bewusst
anders gewählten Formen (76e87ac/4ae11ababd -> S3 form.calibration_identity;
479f6ec/d7f588e017/d0fba8955f/34892e3017 -> S2 NF-Formen; 7f81f09/3c14481318 ->
S4/S7a; 3dbb790 line_gate_27b (Werkzeug, Schritt 9); 8604d13 W100-by-name
(Nutzer: bleibt aus); 6545e2c flashinfer-Pin in pyproject (Image-Frage, nicht Baum)).
Liste: 92bbccb eb5d044 829ebd0 431fcbc ef4d11f 6816062 a233e50 2cc593c f0c8451 87cc4fb c529777 31f2dbe c50085a 3301a96 036b368 e1d1fe9 03c68af 6dddc2e 06932b5 87389c4 58a7490 f85ac55 fc64aa5 5aa24dd 97c0e9a 159333c d9f1532 f3c685b 8550655 e50fb59 db2c2ef f09dc0d c255e10 51b810e 28a55a2 34965fc ff3d9cc 340a018 bee5e10 67b6352 fdade85 ed6630d f1c9a43 b434831 517f26d 0b6b60b a40837f 644de86 aff2b7c 197b701 856024b 238512a 9738626 b857a22 1f8c24d d294b3e 810239d b429dfd e714c95 9efd974 3d63e0a d3cfcf3 fee6134 7985b56 49a14e9 fc45706 19c720e 5306bee 6f1235a c98eaa3 93bc802 328349e f9fb3a2 572af73 94fa8b4 d342caa 2fd7d7e 3ebbb96 871d55f 78c2f16 3babf51 196f6a8 e70af54 22eccfc

Inhalt: Upstream-Ports (sgl-project#33758 sgl-project#37818 sgl-project#36738 sgl-project#33459/sgl-project#30096 sgl-project#34446 sgl-project#36267 sgl-project#33778 sgl-project#34859
sgl-project#36415 sgl-project#35255/sgl-project#36638 sgl-project#39858/sgl-project#40259 sgl-project#31417 sgl-project#34892 sgl-project#32225 sgl-project#30832/sgl-project#36626 sgl-project#39574 sgl-project#29579
sgl-project#31468 sgl-project#32575 sgl-project#31648); xsn409/410-Wake-Verdikte; Vision-Linie V1-V3b + xsn438
(SGLANG_WEG2_VISION_FLIP_URGENT); D-Planer L6 (159333c); DFLASH-Window-Pool
sync-frei, PLAN_SYNC_FREE, D-Kollektive (vocab-argmax, a2a-Merge, deferred rebuild),
#DGAP/D_DEFER_SEQ_LENS_CPU; Mamba-Anker Raster 4096 + Per-Path-Cap + Inner-Release;
P-TRIM (--p-trim-end-anchor); FP8 uniform Marlin; ModelOpt/NVFP4 RadixArk; GGUF G1-G6
+ F1/F2; native-mixed sgl-project#38 (sm_8x W4A8, sm_12x CUTLASS/W4A16); RC1-Capture-Set; sgl-project#49
Agent-Turns; dynchunk (--p-chunk-policy, --p-chunk-dynamic-min-tokens).

Auflösungen (Gabel -> Form, Grund):
- L6 d_operating_point_rows: 27B (d) "Token-Vektor auf jeder Position aus der Kapazität"
  nur bei TP-symmetrischem D (Profil d_layout paged_dcp); sonst NF-sgl-project#1293-Pin + NF-Anker-
  Klausel. mamba_ssm_dtype aus EARLY_READ_FACTS nur bei Profil early_read_flags.
  Overhead-Kalibrierung liest mit form.CalibrationIdentity statt LineIdentity.
- RC1 Capture-Set: neuer RecordKey-Term d_capture_set (qwen27b), Leser
  CalibrationIdentity.d_max_running_requests; nextflash unverändert.
- URC Carrier-Hold: 27B _weg2_carrier_hold entfällt (S2 NF-Rotation), Inner-Release und
  Per-Path-Cap bleiben (Env, Default aus; 27b.env setzt sie).
- scheduler_pp_mixin/overlap_utils/batch_result_processor: NF H49/H58 und 27B #PGAP/#DGAP
  komponiert (beide Instrumente getrennt schaltbar).
- schedule_policy: P-TRIM-Kurzschluss vor NF H63-Fold/QSA-Korn; sgl-project#36415 Hoist + NF
  computed_input_len.
- gdn_backend sgl-project#33778: 27B-strided-Verify; NF-Ring flacht beide Layouts ab.
- flashinfer_backend: RC9-Datei + NF-Form-A-Waiver (27B-intern mehrfach gegabelt).
- checkpoint_census: GGUF-Leser + NF exclude_segments (PLE) in einer Aggregation.
- xchg_manifest: FLAT_SEGMENTS in beiden (dst/src) NF-Breitenbedingungen ausgenommen.
- vram_peak_window: NF-Kumulativ-Peak liest über den 27B-Fast-Read.
- FP8: 8c86eb8-Rest nachgezogen (private Workspace-Registry entfernt, wie RC9).
- argv_d: vision= an allen drei Aufrufstellen (inkl. NF --d-only).
- census_checkpoint_decision (W161) jetzt für beide Profile aktiv.

Gates: py_compile aller geänderten Dateien; ruff F821/F811 ohne neue Funde gegenüber
dem Vorgänger (PendingSeqLensCpu ist String-Annotation wie in RC9); dup_defs_gate 0 neu.
Tests: 73 portierte Testdateien, Lauf nach dem Ruhefenster (Boot aktiv).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bypass-fastfail jit-kernel run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants