Repository navigation
dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels - #41021
Merged
Merged
Conversation
kevin-mii
requested review from
Alisehen,
AniZpZ,
BBuf,
DarkSharpness,
Edwardf0t1,
FlamingoPg,
Fridge003,
HaiShaw,
HydraQYH,
OrangeRedeng,
Qiaolin-Yu,
Ying1123,
alphabetc1,
b8zhong,
celve,
ch-wan,
fzyzcjy,
hanming-lu,
hebiao064,
hnyls2002,
huangtingwei9988,
hzh0425,
ispobock,
merrymercy,
mmangkad,
wisclmy0611,
xiezhq-hermann,
yizhang2077,
yuan-luo and
zijiexia
as code owners
September 24, 2026 03:04
kevin-mii
removed request for
Alisehen,
FlamingoPg,
Fridge003,
JustinTong0323,
OrangeRedeng,
Qiaolin-Yu,
alphabetc1,
ch-wan,
hanming-lu,
hzh0425,
ispobock,
sogalin,
wisclmy0611,
yizhang2077 and
zijiexia
September 26, 2026 16:17
HaiShaw
reviewed
Sep 27, 2026
| // Adapted from AITER custom_all_reduce.cuh: split-H, native HIP post rounding. | ||
| #include "aiter_enum.h" | ||
| #include "aiter_stream.h" | ||
| #include "rocm_ops.hpp" |
Collaborator
There was a problem hiding this comment.
where above 3 header files come from?
Collaborator
Author
There was a problem hiding this comment.
They were AITER's (csrc/include): the kernel was compiled by AITER's JIT. In 38783a1 it builds with SGLang's load_jit like the other JIT kernels, so aiter_enum.h, aiter_stream.h and rocm_ops.hpp are gone. The one AITER header left is custom_all_reduce.cuh, for the CustomAllreduce whose peer buffers and signals the kernel reduces over (SGLang's ROCm custom all-reduce).
The kernel was compiled by AITER's JIT (compile_ops, pybind, aiter_tensor_t) from an in-tree .cu. It is now a load_jit/tvm-ffi module like the other JIT kernels: TensorMatcher checks, LaunchKernel, cache_once. The one AITER header left is custom_all_reduce.cuh, for the CustomAllreduce whose peer buffers and signals SGLang's ROCm custom all-reduce owns. Output and latency are unchanged. Adds a TP4 test: bitwise equal to the unfused all-reduce + hc_post, eager and under graph capture. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
1 task done
…est on MI35x The wrapper rejects operands off the communicator's device or not bf16/fp32 before the launch, as the JIT kernel guide asks of Python wrappers. The test registers on the MI35x stage-c suite like test_deepseek_v4_amd_tp4.py and skips off gfx950. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
HaiShaw
approved these changes
Sep 28, 2026
Collaborator
Author
|
https://github.com/sgl-project/sglang/actions/runs/36335839486 all PR base passed except for JIT test failure from main |
…all-reduce + hc_post kernel test_mhc_kernels.py runs only on CUDA CI, and this PR leaves the CUDA hc_mix_stats_sinkhorn untouched (every mhc.py change is behind _is_hip), so no change here could fail it. The HIP path keeps its coverage in test_hc_boundary_hip.py (fp64 agreement, batch invariance). all_reduce_mhc_hip.cuh copies no AITER or vLLM code: it calls AITER's start_sync / packed_reduce / end_sync through custom_all_reduce.cuh and adds the hc_post epilogue. Replace the copied AMD / vLLM license block with a one-line note, as SGLang does for kernels built on another project. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tf32x3 The old comment said CDNA has no TF32; gfx942 (CDNA3) has xf32 and Triton enables tf32 there, but no AMD target accepts the CUDA default tf32x3. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
hnyls2002
approved these changes
Sep 28, 2026
Collaborator
Author
|
/rerun-failed-ci |
This was referenced Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
ops/layernorm/mhc_boundary_hip.py: the fused mHC sublayer boundary for ROCm. One launch applies the pendinghc_postonto the residual, collapses it with the previouspreand takes the split-K mixing statistics. The reduce + Sinkhorn stays pending inHcCoefficients, andrmsnorm_with_sinkhornruns it in the same launch as the next RMSNorm (optionally with the gfx950 fp8-grid fake quant). A row's result does not depend on M, so it is batch invariant like main's mHC.jit/csrc/deepseek_v4/mhc_boundary_gfx95.cuh: the prefill regime (M >= 1024) of that boundary on gfx950, with LDS DMA andv_permlane*_swap. It follows the Triton kernel's operation order, so each row is bitwise equal in both regimes.ops/communication/all_reduce_mhc_hip.py+jit/csrc/distributed/all_reduce_mhc_hip.cuh: a TP4 all-reduce fused withhc_postfor 1-8 rows of hidden size 5120, on aiter'sCustomAllreducepeer buffers. aiter'sfused_allreduce_mhc_post_one_stagehas the same math, but it rejects hidden sizes above 4096 in bf16 (512 packs), so it cannot serve V4.1's 5120. It is aload_jitmodule like the other JIT kernels; the one aiter header it includes iscustom_all_reduce.cuh, for theCustomAllreducewhose peer buffers and signals SGLang's ROCm custom all-reduce owns. It is compiled with-ffp-contract=off, so it is bitwise equal to the unfused all-reduce followed byhc_post.Changes to existing kernels
ops/layernorm/mhc.py, HIP only:hc_mix_stats_sinkhornusesieeedot precision (CDNA has no TF32) and one row tile at every M, and reduces through the boundary's reduce + Sinkhorn row. CUDA is unchanged.Verification
test_hc_boundary_hip.pycovers the boundary forms against torch and fp64, batch invariance of the boundary and of HIPhc_mix_stats_sinkhorn, the hosted norm against the standalone launches (fp8-grid, MXFP8 and no fake quant), and the gfx950 prefill kernel bitwise to the Triton kernel on full and partial row blocks.main+ this PR.test_all_reduce_mhc_hip.py(four GPUs) checks the fused all-reduce bitwise against the unfused all-reduce +hc_postfor 1, 3 and 8 rows, eager and under graph capture. dsv4.1-amd: serve DeepSeek-V4.1 on gfx950 #41308's four-GPUtest_deepseek_v4_amd_tp4.pyalso covers it: it compares eager (registered-buffer copy-in) and graph-replay outputs of the fused and unfused handoff; all 10 cases pass.Stack
main. Independent of dsv4.1-amd: KV cache layouts, FP4 indexer, compressor and router kernels #41019 (KV cache layouts, FP4 indexer, compressor and router kernels) and dsv4.1-amd: gfx950 sparse decode attention and sorted top-k #41020 (gfx950 sparse decode attention and sorted top-k).🤖 Generated with Claude Code
CI States
Latest PR Test (Base): ⏳ Run #36488300396
Latest PR Test (Extra): ❌ Run #36488300042
Latest PR Test (AMD ROCm 10): ⏳ Run #36488300408