Repository navigation
dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels - #41019
Merged
HaiShaw merged 3 commits intoSep 27, 2026
Merged
Conversation
kevin-mii
requested review from
Alisehen,
AniZpZ,
BBuf,
DarkSharpness,
Edwardf0t1,
FlamingoPg,
Fridge003,
HaiShaw,
HydraQYH,
JustinTong0323,
OrangeRedeng,
Qiaolin-Yu,
Ying1123,
alphabetc1,
b8zhong,
celve,
ch-wan,
hanming-lu,
hebiao064,
hnyls2002,
huangtingwei9988,
hzh0425,
ispobock,
merrymercy,
mmangkad,
wisclmy0611,
xiezhq-hermann,
yizhang2077,
yuan-luo and
zijiexia
as code owners
September 24, 2026 03:03
kevin-mii
removed request for
Alisehen,
Fridge003,
JustinTong0323,
OrangeRedeng,
Qiaolin-Yu,
alphabetc1,
b8zhong,
ch-wan,
hanming-lu,
huangtingwei9988,
hzh0425,
mmangkad and
sogalin
September 25, 2026 19:29
Collaborator
Author
|
/tag-and-rerun-ci |
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
main moved kernels/ops/attention/dsv4/gemm.py to kernels/ops/gemm/bf16_fp32.py (sgl-project#41243); this PR's linear_bf16_fp32 change is carried to the new path. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The split-K bf16 GEMV and its partial reduce move from moe/rocm_router_gate.py to gemm/router_gemv_hip.py, beside main's router_gemv.py; the gate stays in moe. - The V4.1 KV quant oracle moves to sglang.test.kernels.deepseek_v4, main's shared DSV4 kernel-test helpers. - sort_selection_rows is removed: the AOT top-k sorts in its epilogue, so nothing calls it. - test_router_gemv_hip.py holds the GEMV accuracy / batch-invariance case and a linear_bf16_fp32 check that fails on the bf16-rounding aiter route this PR removed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Collaborator
Author
|
/rerun-failed-ci |
HaiShaw
approved these changes
Sep 27, 2026
This was referenced Sep 27, 2026
5 tasks
1 task done
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
include/sgl_kernel/deepseek_v4/kv_layout.cuh,fp4_utils.cuh):v41::store_rowquantizesV41(fp8, ue8m0 per 32) andV41_FP4(e2m1, e4m3 per 16) rows on gfx950 with round-to-nearest-even as on CUDA, sofused_store_cache,compress_norm_rope_store,fused_k_norm_rope_flashmlaanddequantize_k_cache_paged_v41serve both layouts there.fp4_indexer_rope.cuh);fp4_indexer_rope_hip.cuhwrites the split FlyDSL index-K layout;fp4_indexer_hip.pyadds the split writer and reader, the FlyDSL query packer, a one-launch index-Q pack + head weights and a page-table bucket width.fp4_indexer_schedule_hip.py, next to the existing in-tree prefill schedule):build_decode_schedulewrites aitercompute_varctx_schedule'scta_infoin two dispatches whose registers do not grow with the row count. aiter's single kernel spills past ~1K rows and could not reserve its scratch at DSpark batch 256 (1,536 rows), which crashed decode graph capture.ops/gemm/router_gemv_hip.py, beside main'srouter_gemv.py) and a sqrtsoftplus top-k gate over its partials bitwise equal to aiter'stopk_gating_kernel_opt(ops/moe/rocm_router_gate.py). It is here becauseindexer_head_weightsruns the GEMV.c1.cuh,c2.cuh,c2_decode_pool.py) run on ROCm, taking the fp8 pool as bytes there (low_ratio_compress.py);fused_k_norm_rope_flashmla(q=...)ropes the query heads in the K launch on ROCm;sglang/test/kernels/deepseek_v4/dsv41_kv_quant_reference.pyis an independent torch oracle forV41_FP4pages.Changes to existing kernels
main_norm_rope.cuhgains a defaultedkRopeQtemplate flag;fp4_indexer.pyandfp4_rope_fake_quant.pymove their row math into@triton.jithelpers (positions=Nonekeeps the old call);c1.cuhandfp4_indexer_rope.cuhmatchkDLGPU, which iskDLCUDAon CUDA;c2_decode_pool.pydivides withtl.div_rninstead oflibdevice.div_rn(both IEEE round-to-nearest; Triton's HIP libdevice has nodiv_rn).SGLANG_USE_AITER,linear_bf16_fp32now takes the default fp32torch.mmpath instead of aiter's bf16-roundedtgemm(V4 compressor and router logits), and the aiter branch is removed. The V4.1 paths are asserted gfx950-only, and gfx942 code is unchanged.Verification
test_v41_kv_store.py(byte-exact stores and dequant for both layouts, query rope in the K launch),test_rocm_router_gate.py(bitwise to aiter on ties and non-finite logits, the fused split-K gate),test_router_gemv_hip.py(the batch-invariant GEMV, andlinear_bf16_fp32staying in fp32, which fails on the removed aiter route);test_fp4_indexer_hip.pyis extended (split index-K writer bitwise to the Triton RMSNorm + RoPE + FP4 chain, decode schedule bitwise to aiter).test_c4_v2.py,test_c128_v2.py,test_deepseek_v4_compress_state_runtime_shapes.py,test_dsv4_unified_fp8_compress_store.py,test_dsv4_fp8_cast.py) pass on this branch, as do the AMD-registered DSV4 attention, MXFP8, mHC and HiCache kernel tests.test_v41_kv_store.py(SM100),test_fp4_indexer.py,test_deepseek_v4.py.Stack
main. Independent of dsv4.1-amd: gfx950 MXFP8 matmul kernels and fp8-grid producers #41018 (gfx950 MXFP8 matmul kernels).🤖 Generated with Claude Code
CI States
Latest PR Test (Base): ✅ Run #36227764374
Latest PR Test (Extra): ✅ Run #36227764031
Latest PR Test (AMD ROCm 10): ❌ Run #36227764383