Skip to content

dsv4.1-amd: fused mHC boundary and all-reduce + mHC post kernels - #41021

Merged
hnyls2002 merged 6 commits into
sgl-project:mainfrom
kevin-mii:dsv41-amd-4-model
Sep 28, 2026
Merged

hnyls2002 merged 6 commits into
sgl-project:mainfrom
kevin-mii:dsv41-amd-4-model

Conversation

@kevin-mii

@kevin-mii kevin-mii commented Sep 24, 2026 •

Copy link
Copy Markdown
Collaborator

This PR was "stack 4/4" of the DeepSeek-V4.1 AMD series (#41018 to #41021). It now follows the CUDA dsv4.1 layout (#39646, #39652, #39653, #39656, #39664, then #38798): kernel PRs by domain, each on main with no callers, then one integration PR. This is the fourth kernel PR; #41018 has merged, and it does not depend on #41019 or #41020. The model wiring moved to the integration PR, #41308.

Summary

  • ops/layernorm/mhc_boundary_hip.py: the fused mHC sublayer boundary for ROCm. One launch applies the pending hc_post onto the residual, collapses it with the previous pre and takes the split-K mixing statistics. The reduce + Sinkhorn stays pending in HcCoefficients, and rmsnorm_with_sinkhorn runs it in the same launch as the next RMSNorm (optionally with the gfx950 fp8-grid fake quant). A row's result does not depend on M, so it is batch invariant like main's mHC.
  • jit/csrc/deepseek_v4/mhc_boundary_gfx95.cuh: the prefill regime (M >= 1024) of that boundary on gfx950, with LDS DMA and v_permlane*_swap. It follows the Triton kernel's operation order, so each row is bitwise equal in both regimes.
  • ops/communication/all_reduce_mhc_hip.py + jit/csrc/distributed/all_reduce_mhc_hip.cuh: a TP4 all-reduce fused with hc_post for 1-8 rows of hidden size 5120, on aiter's CustomAllreduce peer buffers. aiter's fused_allreduce_mhc_post_one_stage has the same math, but it rejects hidden sizes above 4096 in bf16 (512 packs), so it cannot serve V4.1's 5120. It is a load_jit module like the other JIT kernels; the one aiter header it includes is custom_all_reduce.cuh, for the CustomAllreduce whose peer buffers and signals SGLang's ROCm custom all-reduce owns. It is compiled with -ffp-contract=off, so it is bitwise equal to the unfused all-reduce followed by hc_post.
  • Model call sites are in the integration PR of this stack.

Changes to existing kernels

  • ops/layernorm/mhc.py, HIP only: hc_mix_stats_sinkhorn uses ieee dot precision (CDNA has no TF32) and one row tile at every M, and reduces through the boundary's reduce + Sinkhorn row. CUDA is unchanged.

Verification

  • New tests: test_hc_boundary_hip.py covers the boundary forms against torch and fp64, batch invariance of the boundary and of HIP hc_mix_stats_sinkhorn, the hosted norm against the standalone launches (fp8-grid, MXFP8 and no fake quant), and the gfx950 prefill kernel bitwise to the Triton kernel on full and partial row blocks.
  • MI350X (ROCm 10.0.0, torch 2.11, triton 3.8, aiter at the Dockerfile pin): these tests pass on main + this PR.
  • test_all_reduce_mhc_hip.py (four GPUs) checks the fused all-reduce bitwise against the unfused all-reduce + hc_post for 1, 3 and 8 rows, eager and under graph capture. dsv4.1-amd: serve DeepSeek-V4.1 on gfx950 #41308's four-GPU test_deepseek_v4_amd_tp4.py also covers it: it compares eager (registered-buffer copy-in) and graph-replay outputs of the fused and unfused handoff; all 10 cases pass.
  • With dsv4.1-amd: serve DeepSeek-V4.1 on gfx950 #41308's callers, DeepSeek-V4.1-Flash at TP4 / EP4 with DSpark scores 0.905 on GSM8K (200 questions) and serves random 4K/1K at 521 / 1,840 / 3,307 tok/s at concurrency 1 / 8 / 32; see dsv4.1-amd: serve DeepSeek-V4.1 on gfx950 #41308.

Stack

🤖 Generated with Claude Code


CI States

Latest PR Test (Base): ⏳ Run #36488300396
Latest PR Test (Extra): ❌ Run #36488300042
Latest PR Test (AMD ROCm 10): ⏳ Run #36488300408

// Adapted from AITER custom_all_reduce.cuh: split-H, native HIP post rounding.
#include "aiter_enum.h"
#include "aiter_stream.h"
#include "rocm_ops.hpp"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

where above 3 header files come from?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They were AITER's (csrc/include): the kernel was compiled by AITER's JIT. In 38783a1 it builds with SGLang's load_jit like the other JIT kernels, so aiter_enum.h, aiter_stream.h and rocm_ops.hpp are gone. The one AITER header left is custom_all_reduce.cuh, for the CustomAllreduce whose peer buffers and signals the kernel reduces over (SGLang's ROCm custom all-reduce).

The kernel was compiled by AITER's JIT (compile_ops, pybind, aiter_tensor_t) from an in-tree .cu.
It is now a load_jit/tvm-ffi module like the other JIT kernels: TensorMatcher checks, LaunchKernel,
cache_once. The one AITER header left is custom_all_reduce.cuh, for the CustomAllreduce whose peer
buffers and signals SGLang's ROCm custom all-reduce owns. Output and latency are unchanged.

Adds a TP4 test: bitwise equal to the unfused all-reduce + hc_post, eager and under graph capture.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…est on MI35x

The wrapper rejects operands off the communicator's device or not bf16/fp32 before the
launch, as the JIT kernel guide asks of Python wrappers. The test registers on the MI35x
stage-c suite like test_deepseek_v4_amd_tp4.py and skips off gfx950.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kevin-mii

Copy link
Copy Markdown
Collaborator Author

https://github.com/sgl-project/sglang/actions/runs/36335839486 all PR base passed except for JIT test failure from main

kevin-mii and others added 3 commits September 28, 2026 13:09
…all-reduce + hc_post kernel

test_mhc_kernels.py runs only on CUDA CI, and this PR leaves the CUDA
hc_mix_stats_sinkhorn untouched (every mhc.py change is behind _is_hip), so
no change here could fail it. The HIP path keeps its coverage in
test_hc_boundary_hip.py (fp64 agreement, batch invariance).

all_reduce_mhc_hip.cuh copies no AITER or vLLM code: it calls AITER's
start_sync / packed_reduce / end_sync through custom_all_reduce.cuh and adds
the hc_post epilogue. Replace the copied AMD / vLLM license block with a
one-line note, as SGLang does for kernels built on another project.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tf32x3

The old comment said CDNA has no TF32; gfx942 (CDNA3) has xf32 and Triton
enables tf32 there, but no AMD target accepts the CUDA default tf32x3.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@kevin-mii

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

amd bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) bypass-fastfail deepseek documentation Improvements or additions to documentation jit-kernel memory-pool quant LLM Quantization run-ci CI: run the baseline test suite on this PR sgl-kernel

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants