[rotary] Fix the fused Qwen3.5 RoPE kernel discarding mrope height and width - #34446
Conversation
|
/rerun-test test/registered/kernels/ops/attention/test_fused_qk_rmsnorm_rope_gate.py test/registered/rotary/test_mrope_axis_map.py test/registered/cpu/test_rope.py test/registered/kernels/ops/attention/test_rope.py test/registered/unit/multimodal/test_mrope_encoder_utils.py test/registered/rotary/test_rope_cache_invalidation.py test/registered/unit/multimodal/rust/qwen/test_token_layout_mrope.py test/registered/vlm/test_vision_openai_server_a.py |
|
Results for ⛔ ⛔ |
c155342 to
db41e22
Compare
db41e22 to
ac694cd
Compare
|
/rerun-test test/registered/kernels/ops/attention/test_fused_qk_rmsnorm_rope_gate.py test/registered/rotary/test_mrope_axis_map.py test/registered/cpu/test_rope.py test/registered/kernels/ops/attention/test_rope.py test/registered/unit/multimodal/test_mrope_encoder_utils.py test/registered/rotary/test_rope_cache_invalidation.py test/registered/unit/multimodal/rust/qwen/test_token_layout_mrope.py test/registered/vlm/test_vision_openai_server_a.py |
|
Results for 🚀 🚀 |
71c5338 to
0929d4c
Compare
|
/rerun-test test/registered/rotary/test_mrope_axis_map.py test/registered/kernels/ops/attention/test_fused_qk_rmsnorm_rope_gate.py test/registered/vlm/test_vision_openai_server_a.py |
|
Results for 🚀 🚀 |
…gl-project#34446) The fused kernel loaded a single position per token, so with image inputs ([3, T] temporal/height/width positions) every rotary lane silently read the temporal row — wrong RoPE on image tokens in all full-attention layers of Qwen3.5/3.8 hybrids. Text was unaffected (the three rows coincide), so this never shows in text-only smoke. Port: the kernel takes an mrope_axis_map ([rotary_dim//2] lane->axis) and reads positions[axis[lane], t] when positions is 2-D; MRotaryEmbedding now builds the axis map for every mrope_section style (contiguous, interleaved, GLM round-robin) instead of GLM only, while the legacy sgl_kernel call sites keep the GLM-only map via _legacy_axis_map. Unit test mirrors the kernel math bitwise for 1-D and mrope positions and checks both axis-map styles.
…cks, chronologisch) Grundlage: Präsenz-Scan aller 230 27B-Commits seit 76f8deb gegen diesen Baum (Stichprobe der hinzugefügten Zeilen je Commit); die 94 fehlenden minus die bewusst anders gewählten Formen (76e87ac/4ae11ababd -> S3 form.calibration_identity; 479f6ec/d7f588e017/d0fba8955f/34892e3017 -> S2 NF-Formen; 7f81f09/3c14481318 -> S4/S7a; 3dbb790 line_gate_27b (Werkzeug, Schritt 9); 8604d13 W100-by-name (Nutzer: bleibt aus); 6545e2c flashinfer-Pin in pyproject (Image-Frage, nicht Baum)). Liste: 92bbccb eb5d044 829ebd0 431fcbc ef4d11f 6816062 a233e50 2cc593c f0c8451 87cc4fb c529777 31f2dbe c50085a 3301a96 036b368 e1d1fe9 03c68af 6dddc2e 06932b5 87389c4 58a7490 f85ac55 fc64aa5 5aa24dd 97c0e9a 159333c d9f1532 f3c685b 8550655 e50fb59 db2c2ef f09dc0d c255e10 51b810e 28a55a2 34965fc ff3d9cc 340a018 bee5e10 67b6352 fdade85 ed6630d f1c9a43 b434831 517f26d 0b6b60b a40837f 644de86 aff2b7c 197b701 856024b 238512a 9738626 b857a22 1f8c24d d294b3e 810239d b429dfd e714c95 9efd974 3d63e0a d3cfcf3 fee6134 7985b56 49a14e9 fc45706 19c720e 5306bee 6f1235a c98eaa3 93bc802 328349e f9fb3a2 572af73 94fa8b4 d342caa 2fd7d7e 3ebbb96 871d55f 78c2f16 3babf51 196f6a8 e70af54 22eccfc Inhalt: Upstream-Ports (sgl-project#33758 sgl-project#37818 sgl-project#36738 sgl-project#33459/sgl-project#30096 sgl-project#34446 sgl-project#36267 sgl-project#33778 sgl-project#34859 sgl-project#36415 sgl-project#35255/sgl-project#36638 sgl-project#39858/sgl-project#40259 sgl-project#31417 sgl-project#34892 sgl-project#32225 sgl-project#30832/sgl-project#36626 sgl-project#39574 sgl-project#29579 sgl-project#31468 sgl-project#32575 sgl-project#31648); xsn409/410-Wake-Verdikte; Vision-Linie V1-V3b + xsn438 (SGLANG_WEG2_VISION_FLIP_URGENT); D-Planer L6 (159333c); DFLASH-Window-Pool sync-frei, PLAN_SYNC_FREE, D-Kollektive (vocab-argmax, a2a-Merge, deferred rebuild), #DGAP/D_DEFER_SEQ_LENS_CPU; Mamba-Anker Raster 4096 + Per-Path-Cap + Inner-Release; P-TRIM (--p-trim-end-anchor); FP8 uniform Marlin; ModelOpt/NVFP4 RadixArk; GGUF G1-G6 + F1/F2; native-mixed sgl-project#38 (sm_8x W4A8, sm_12x CUTLASS/W4A16); RC1-Capture-Set; sgl-project#49 Agent-Turns; dynchunk (--p-chunk-policy, --p-chunk-dynamic-min-tokens). Auflösungen (Gabel -> Form, Grund): - L6 d_operating_point_rows: 27B (d) "Token-Vektor auf jeder Position aus der Kapazität" nur bei TP-symmetrischem D (Profil d_layout paged_dcp); sonst NF-sgl-project#1293-Pin + NF-Anker- Klausel. mamba_ssm_dtype aus EARLY_READ_FACTS nur bei Profil early_read_flags. Overhead-Kalibrierung liest mit form.CalibrationIdentity statt LineIdentity. - RC1 Capture-Set: neuer RecordKey-Term d_capture_set (qwen27b), Leser CalibrationIdentity.d_max_running_requests; nextflash unverändert. - URC Carrier-Hold: 27B _weg2_carrier_hold entfällt (S2 NF-Rotation), Inner-Release und Per-Path-Cap bleiben (Env, Default aus; 27b.env setzt sie). - scheduler_pp_mixin/overlap_utils/batch_result_processor: NF H49/H58 und 27B #PGAP/#DGAP komponiert (beide Instrumente getrennt schaltbar). - schedule_policy: P-TRIM-Kurzschluss vor NF H63-Fold/QSA-Korn; sgl-project#36415 Hoist + NF computed_input_len. - gdn_backend sgl-project#33778: 27B-strided-Verify; NF-Ring flacht beide Layouts ab. - flashinfer_backend: RC9-Datei + NF-Form-A-Waiver (27B-intern mehrfach gegabelt). - checkpoint_census: GGUF-Leser + NF exclude_segments (PLE) in einer Aggregation. - xchg_manifest: FLAT_SEGMENTS in beiden (dst/src) NF-Breitenbedingungen ausgenommen. - vram_peak_window: NF-Kumulativ-Peak liest über den 27B-Fast-Read. - FP8: 8c86eb8-Rest nachgezogen (private Workspace-Registry entfernt, wie RC9). - argv_d: vision= an allen drei Aufrufstellen (inkl. NF --d-only). - census_checkpoint_decision (W161) jetzt für beide Profile aktiv. Gates: py_compile aller geänderten Dateien; ruff F821/F811 ohne neue Funde gegenüber dem Vorgänger (PendingSeqLensCpu ist String-Annotation wie in RC9); dup_defs_gate 0 neu. Tests: 73 portierte Testdateien, Lauf nach dem Ruhefenster (Boot aktiv).
Motivation
fused_qk_gemma_rmsnorm_rope_gateloads one position per token,tl.load(positions_ptr + token). Multimodal Qwen3.5 passes mrope positions shaped[3, T], one row per axis, so that offset only ever lands in row 0: height and widthwere discarded, and every image token rotated as if it sat at its temporal position on
all three axes.
Text tokens hold the same position on all three rows, so they were already correct and no
text benchmark could catch this. The path is not behind a flag —
Qwen3_5AttentionDecoderLayer.self_attentiontakes it whenever_is_cuda and self.attn_output_gate(defaultTrue) — soQwen3_5ForConditionalGenerationandQwen3_5MoeForConditionalGenerationon CUDA wereaffected for every request carrying an image.
forward_native, the HIP/XPU/CPU/NPUbranches, and text-only users of the same layer (
qwen3_5_mtp,minicpmv,interns2preview) are unaffected.Modifications
fused_qk_rmsnorm_rope_gate.py: the kernel takesmrope_axis_mapand anMROPE: tl.constexpr; each rotary lane loads the axis that owns it and indexespositionsby that row. The row stride comes from the tensor rather thanT, sinceCUDA-graph decode replays
mrope_positions[:, :num_tokens]out of a[3, max_num_token]buffer.MRotaryEmbedding._build_axis_map: generalized from GLM-interleaved only to all threelayouts, which differ only in which axis owns which lane.
_legacy_axis_map:triton_mrope_fusedand the out-of-treesgl_kernel.multimodal_rotary_embeddingkeep receivingNoneoutside GLM. This PRchanges no input to either.
Ernie4_5_VLRotaryEmbedding._build_axis_mapreturnsNone: it readsmrope_sectionas h, w, t while the shared builder assumes t, h, w.
qwen3_5.py: passes the map when positions are 2-D.Accuracy Tests
Both kernels against
MRotaryEmbedding.forward_native, at Qwen3.6-35B-A3B's shape(16 q heads, 2 kv heads,
head_dim256,rotary_dim64) and its shipped interleavedmrope_section [11, 11, 10], bf16, with error grouped by the axis owning each lane:0.0625 is the bf16 floor here. Stock sits on it for temporal lanes and up to 300x above
it for height and width, which is one row applied and two ignored; this PR brings every
lane to the floor. The last row is why it shipped.
New
test_fused_qk_rmsnorm_rope_gate.pycompares the kernel withforward_nativeoverinterleaved and contiguous
[11, 11, 10]and[24, 20, 20](no pass-through tail), withpositions sliced out of a wider buffer to reproduce the CUDA-graph stride, plus 1-D parity
and rejection of both illegal position/map pairings. New
test_mrope_axis_map.pycheckseach layout against the function that consumes it, and pins GLM's order (its consumer is
out of tree) and Ernie's opt-out.
End to end, which is what shows the model is wired to the map:
Qwen/Qwen3.6-35B-A3Bat995ad96, 1xB200, greedy, three words asked for separately on a generated page. Thereference column disables the fused branch, dropping CUDA to the plain path that ropes
through
MRotaryEmbedding— same weights and hardware, only the rope kernel differs.Before, all three share
x1 = 800: the word being asked about does not move the box.After, each tracks the unfused path within 8 px. This is parity between the two paths, not
a grounding score.
Speed Tests and Profiling
MROPEis atl.constexpr, so the 1-D path keeps its own specialization. Same shape,triton.testing.do_bench, median per launch on 1xB200:The mrope specialization costs one
int64load per lane, with no extra launches,synchronization or allocations.
CI States
Latest PR Test (Base): ❌ Run #33127023374
Latest PR Test (Extra): ✅ Run #33127023235
Latest PR Test (AMD ROCm 7.2): ❌ Run #33127023285