Skip to content

[AMD][V4.1][*/N] Switch the fp8 dense GEMMs on gfx950 to aiter's MXFP8 GEMM - #41970

Merged
HaiShaw merged 3 commits into
sgl-project:mainfrom
RolaoDenthu:dsv41/aiter-mxfp8-gemm
Oct 1, 2026
Merged

HaiShaw merged 3 commits into
sgl-project:mainfrom
RolaoDenthu:dsv41/aiter-mxfp8-gemm

Conversation

@RolaoDenthu

@RolaoDenthu RolaoDenthu commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

On gfx950, the DeepSeek-V4.1 fp8 dense linears (32x32-block weights with ue8m0 scales) run on sglang's native MXFP8 route.

ROCm/aiter#5750 adds a native MXFP8 GEMM to aiter (e4m3 with e8m0 scales per 32 along K, 32x32-block weight scales), reached through aiter.gemm_a8w8_blockscale. This PR adds it as an opt-in route: with --fp8-gemm-backend aiter, the V4.1 fp8 dense linears on gfx950 run on aiter's kernel.

Modifications

Add aiter's MXFP8 GEMM as a gfx950 backend for the DeepSeek-V4.1 fp8 dense linears, plus tests.

  • Opt-in with --fp8-gemm-backend aiter
  • Test: TestMxfp8AiterRouteGfx95 in test_mxfp8_amd_gfx95.py.

Accuracy Tests

gsm8k 1319: 0.909

Speed Tests and Profiling

kernel level test

With DSpark (block size 5), decode M is 6 x concurrency.

M concurrency native aiter aiter / native
6 1 35.4 37.2 1.05x
12 2 39.2 39.8 1.02x
24 4 58.1 37.6 0.65x
48 8 106.2 45.3 0.43x
96 16 109.3 52.8 0.48x
192 32 116.2 64.1 0.55x
512 prefill 148.0 105.7 0.71x

e2e test

8k1k

Concurrency Metric Baseline (native) aiter Change
1 TTFT (ms) 400.29 396.40 −1.0%
1 TPOT (ms) 1.66 1.66 0
1 ITL (ms) 1.66 1.66 0
1 Output throughput (tok/s) 411.46 430.21 +4.6%
4 TTFT (ms) 425.99 420.38 −1.3%
4 TPOT (ms) 3.76 3.59 −4.5%
4 ITL (ms) 2.68 2.51 −6.3%
4 Output throughput (tok/s) 839.72 875.57 +4.3%
16 TTFT (ms) 1096.42 943.27 −14.0%
16 TPOT (ms) 9.39 9.00 −4.2%
16 ITL (ms) 4.67 4.31 −7.7%
16 Output throughput (tok/s) 1349.40 1414.23 +4.8%

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #36810199110
Latest PR Test (Extra): ❌ Run #36810198773
Latest PR Test (AMD ROCm 10): ❌ Run #36810198918

Depends on ROCm/aiter#5750, which adds a native MXFP8 GEMM (e4m3 with e8m0
scales per 32 along K, 32x32-block weight scales) behind
aiter.gemm_a8w8_blockscale. With --fp8-gemm-backend aiter, the DSv4.1 fp8
dense linears on gfx950 run on that kernel instead of the sglang native MXFP8
route (mxfp8_gemv / dot_scaled / hipBLASLt bf16).
On the aiter route the stacked wkv weights fell through to the Triton block
fp8 GEMM, whose ue8m0 activation quant is CUDA-only. Stack the ue8m0 block
scales too and run the stacked GEMM through aiter.
@RolaoDenthu RolaoDenthu changed the title [AMD] Switch the DSv4.1 fp8 dense GEMMs on gfx950 to aiter's MXFP8 GEMM [AMD][V4.1][*/N] Switch the fp8 dense GEMMs on gfx950 to aiter's MXFP8 GEMM Oct 1, 2026
@HaiShaw
HaiShaw merged commit 73ba651 into sgl-project:main Oct 1, 2026
94 of 110 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants