You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
aiter/configs/model_configs/dsv41_fp8_untuned_fmoe.csv — the corresponding shape list.
Results
End-to-end serving benchmark on MI355X, TP4 with expert parallelism, same node and same
vLLM build as the baseline, with the tuned CSV as the only difference (metrics are given as the mean):
Metric
Δ vs default heuristics
Improved at
TPOT
-2.5%
12/12 points
Output throughput
+2.7%
12/12 points
TTFT
-4.7%
12/12 points
The gain is consistent across the whole sweep and grows modestly with concurrency (-2.1% at 4 up to -3.4% at 256).
input/output
conc.
TTFT base (ms)
TTFT tuned (ms)
TTFT Δ
TPOT base (ms)
TPOT tuned (ms)
TPOT Δ
Thr. base (tok/s)
Thr. tuned (tok/s)
Thr. Δ
8192/1024
4
275.7
225.7
-18.1%
10.17
9.83
-3.3%
374.9
389.2
+3.8%
8192/1024
8
289.3
273.8
-5.3%
11.37
11.29
-0.7%
669.9
675.3
+0.8%
8192/1024
16
422.1
405.7
-3.9%
13.53
13.11
-3.2%
1112.5
1148.4
+3.2%
8192/1024
32
930.8
828.7
-11.0%
17.18
16.94
-1.4%
1668.4
1702.5
+2.0%
8192/1024
64
1461.7
1432.2
-2.0%
24.75
23.78
-3.9%
2364.7
2456.2
+3.9%
8192/1024
128
3528.9
3444.0
-2.4%
36.15
34.81
-3.7%
3105.1
3219.7
+3.7%
8192/1024
256
6796.4
6649.6
-2.2%
58.28
56.57
-2.9%
3801.9
3913.4
+2.9%
60000/600
4
1791.4
1722.3
-3.9%
15.19
14.84
-2.3%
215.0
220.9
+2.7%
60000/600
16
3336.7
3314.2
-0.7%
38.71
37.70
-2.6%
352.5
360.9
+2.4%
60000/600
64
12538.5
11853.9
-5.5%
125.87
123.77
-1.7%
425.8
435.5
+2.3%
128000/1024
4
3670.2
3607.2
-1.7%
17.55
17.20
-2.0%
183.7
187.3
+2.0%
128000/1024
16
7415.6
7408.2
-0.1%
49.96
48.72
-2.5%
271.9
277.9
+2.2%
Fused-MoE kernel, expert-parallel benchmark
Kernels were picked with a fused_moe benchmark that reproduces how vLLM calls it
under TP4 + EP. The benchmark times the whole fused_moe call: sorting, both stages and
the intermediate quantization. Each token count is timed with the tuned
lookup and with AITER_BYPASS_TUNE_CONFIG=1 (default heuristics) in the same process,
over 3 repeats with random routing.
tokens
local routes
tuned (us)
default (us)
speedup
1
1
38.7
38.6
0.999
4
6
44.9
47.2
1.050
16
26
94.8
101.1
1.067
32
40
118.2
123.9
1.048
64
92
178.1
194.8
1.094
128
211
269.8
289.1
1.072
256
378
295.4
315.3
1.067
512
809
306.8
333.3
1.086
1024
1586
327.5
346.7
1.059
2048
3080
373.7
412.3
1.103
4096
6187
505.2
588.3
1.165
8192
12506
727.8
870.3
1.196
16384
24985
1180.6
1361.7
1.153
Geometric-mean speedup over the 13 tuned rows is 1.088×. The largest gains are at
the prefill-sized buckets (1.10–1.20× from 2048 tokens up). Outputs match the
default path to within 1.5e-4 max abs difference at every token count. Token counts
without a row (2, 8, 32768) run the same kernels in both modes and measure 0.993×,
which is the benchmark's noise floor.
Validation
Accuracy remains unchanged on GSM8K after tuning. GSM8K, 5-shot, via lm_eval:
All standard extended tests (excludes ci:atom_full)
Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5904 --add-label <label>
One backend per PR:
A PR changes one kernel backend: [Triton/Gluon] (Triton and Gluon count as one), [HIP], [ASM], [CK], [OPUS] or [FlyDSL]. If the title ends up with two backend tags, split the PR -- as stacked pull requests when one part cannot merge without the other.
PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to stop the title rewrites; labels stay in sync either way.
github-actionsBot
changed the title
Add tuned a8w4 fused-MoE configs for DeepSeek V4.1-Flash (TP4 + EP, gfx950)
[Config] Add tuned a8w4 fused-MoE configs for DeepSeek V4.1-Flash (TP4 + EP, gfx950)
Sep 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Add tuned a8w4 fused-MoE configs for DeepSeek V4.1-Flash (TP4 + EP, gfx950)
Adds tuned fused-MoE kernel selections for the DeepSeek V4.1-Flash a8w4 MoE shape
(
model_dim=5120, inter_dim=2304, 96 local experts, topk=6, FP8 activations x MXFP4weights,
QuantType.per_1x32) on gfx950 / 256 CU.What changed
aiter/configs/model_configs/dsv41_fp8_tuned_fmoe.csv— 13 tuned rows covering tokenbuckets 1 through 16384.
aiter/configs/model_configs/dsv41_fp8_untuned_fmoe.csv— the corresponding shape list.Results
End-to-end serving benchmark on MI355X, TP4 with expert parallelism, same node and same
vLLM build as the baseline, with the tuned CSV as the only difference (metrics are given as the mean):
The gain is consistent across the whole sweep and grows modestly with concurrency (-2.1% at 4 up to -3.4% at 256).
Fused-MoE kernel, expert-parallel benchmark
Kernels were picked with a
fused_moebenchmark that reproduces how vLLM calls itunder TP4 + EP. The benchmark times the whole
fused_moecall: sorting, both stages andthe intermediate quantization. Each token count is timed with the tuned
lookup and with
AITER_BYPASS_TUNE_CONFIG=1(default heuristics) in the same process,over 3 repeats with random routing.
Geometric-mean speedup over the 13 tuned rows is 1.088×. The largest gains are at
the prefill-sized buckets (1.10–1.20× from 2048 tokens up). Outputs match the
default path to within 1.5e-4 max abs difference at every token count. Token counts
without a row (2, 8, 32768) run the same kernels in both modes and measure 0.993×,
which is the benchmark's noise floor.
Validation
Accuracy remains unchanged on GSM8K after tuning. GSM8K, 5-shot, via
lm_eval: