Skip to content

[Config] Add tuned a8w4 fused-MoE configs for DeepSeek V4.1-Flash (TP4 + EP, gfx950) - #5904

Merged
yifehuan merged 2 commits into
ROCm:mainfrom
cpersson-amd:dsv41_fmoe_tunings
Sep 30, 2026
Merged

yifehuan merged 2 commits into
ROCm:mainfrom
cpersson-amd:dsv41_fmoe_tunings

Conversation

@cpersson-amd

Copy link
Copy Markdown

Add tuned a8w4 fused-MoE configs for DeepSeek V4.1-Flash (TP4 + EP, gfx950)

Adds tuned fused-MoE kernel selections for the DeepSeek V4.1-Flash a8w4 MoE shape
(model_dim=5120, inter_dim=2304, 96 local experts, topk=6, FP8 activations x MXFP4
weights, QuantType.per_1x32) on gfx950 / 256 CU.

What changed

  • aiter/configs/model_configs/dsv41_fp8_tuned_fmoe.csv — 13 tuned rows covering token
    buckets 1 through 16384.
  • aiter/configs/model_configs/dsv41_fp8_untuned_fmoe.csv — the corresponding shape list.

Results

End-to-end serving benchmark on MI355X, TP4 with expert parallelism, same node and same
vLLM build as the baseline, with the tuned CSV as the only difference (metrics are given as the mean):

Metric Δ vs default heuristics Improved at
TPOT -2.5% 12/12 points
Output throughput +2.7% 12/12 points
TTFT -4.7% 12/12 points

The gain is consistent across the whole sweep and grows modestly with concurrency (-2.1% at 4 up to -3.4% at 256).

input/output conc. TTFT base (ms) TTFT tuned (ms) TTFT Δ TPOT base (ms) TPOT tuned (ms) TPOT Δ Thr. base (tok/s) Thr. tuned (tok/s) Thr. Δ
8192/1024 4 275.7 225.7 -18.1% 10.17 9.83 -3.3% 374.9 389.2 +3.8%
8192/1024 8 289.3 273.8 -5.3% 11.37 11.29 -0.7% 669.9 675.3 +0.8%
8192/1024 16 422.1 405.7 -3.9% 13.53 13.11 -3.2% 1112.5 1148.4 +3.2%
8192/1024 32 930.8 828.7 -11.0% 17.18 16.94 -1.4% 1668.4 1702.5 +2.0%
8192/1024 64 1461.7 1432.2 -2.0% 24.75 23.78 -3.9% 2364.7 2456.2 +3.9%
8192/1024 128 3528.9 3444.0 -2.4% 36.15 34.81 -3.7% 3105.1 3219.7 +3.7%
8192/1024 256 6796.4 6649.6 -2.2% 58.28 56.57 -2.9% 3801.9 3913.4 +2.9%
60000/600 4 1791.4 1722.3 -3.9% 15.19 14.84 -2.3% 215.0 220.9 +2.7%
60000/600 16 3336.7 3314.2 -0.7% 38.71 37.70 -2.6% 352.5 360.9 +2.4%
60000/600 64 12538.5 11853.9 -5.5% 125.87 123.77 -1.7% 425.8 435.5 +2.3%
128000/1024 4 3670.2 3607.2 -1.7% 17.55 17.20 -2.0% 183.7 187.3 +2.0%
128000/1024 16 7415.6 7408.2 -0.1% 49.96 48.72 -2.5% 271.9 277.9 +2.2%

Fused-MoE kernel, expert-parallel benchmark

Kernels were picked with a fused_moe benchmark that reproduces how vLLM calls it
under TP4 + EP. The benchmark times the whole fused_moe call: sorting, both stages and
the intermediate quantization. Each token count is timed with the tuned
lookup and with AITER_BYPASS_TUNE_CONFIG=1 (default heuristics) in the same process,
over 3 repeats with random routing.

tokens local routes tuned (us) default (us) speedup
1 1 38.7 38.6 0.999
4 6 44.9 47.2 1.050
16 26 94.8 101.1 1.067
32 40 118.2 123.9 1.048
64 92 178.1 194.8 1.094
128 211 269.8 289.1 1.072
256 378 295.4 315.3 1.067
512 809 306.8 333.3 1.086
1024 1586 327.5 346.7 1.059
2048 3080 373.7 412.3 1.103
4096 6187 505.2 588.3 1.165
8192 12506 727.8 870.3 1.196
16384 24985 1180.6 1361.7 1.153

Geometric-mean speedup over the 13 tuned rows is 1.088×. The largest gains are at
the prefill-sized buckets (1.10–1.20× from 2048 tokens up). Outputs match the
default path to within 1.5e-4 max abs difference at every token count. Token counts
without a row (2, 8, 32768) run the same kernels in both modes and measure 0.993×,
which is the benchmark's noise floor.

Validation

Accuracy remains unchanged on GSM8K after tuning. GSM8K, 5-shot, via lm_eval:

Filter Default Tuned Δ
flexible-extract 0.9227 ± 0.0074 0.9249 ± 0.0073 +0.0022
strict-match 0.9219 ± 0.0074 0.9257 ± 0.0072 +0.0038

@cpersson-amd
cpersson-amd requested a review from a team September 28, 2026 07:53
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5904 --add-label <label>

One backend per PR:
A PR changes one kernel backend: [Triton/Gluon] (Triton and Gluon count as one), [HIP], [ASM], [CK], [OPUS] or [FlyDSL]. If the title ends up with two backend tags, split the PR -- as stacked pull requests when one part cannot merge without the other.

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to stop the title rewrites; labels stay in sync either way.

@github-actions github-actions Bot changed the title Add tuned a8w4 fused-MoE configs for DeepSeek V4.1-Flash (TP4 + EP, gfx950) [Config] Add tuned a8w4 fused-MoE configs for DeepSeek V4.1-Flash (TP4 + EP, gfx950) Sep 28, 2026
@zufayu
zufayu requested a review from yifehuan September 29, 2026 00:32

@yifehuan yifehuan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@yifehuan
yifehuan merged commit 1e04738 into ROCm:main Sep 30, 2026
50 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants