Skip to content

[Config] Extend the DeepSeek-V4 a8w8 blockscale GEMM tunings for gfx950 - #5485

Merged
yifehuan merged 2 commits into
ROCm:mainfrom
LiuYinfeng01:dsv4-a8w8-blockscale-tuning
Sep 15, 2026
Merged

yifehuan merged 2 commits into
ROCm:mainfrom
LiuYinfeng01:dsv4-a8w8-blockscale-tuning

Conversation

@LiuYinfeng01

@LiuYinfeng01 LiuYinfeng01 commented Sep 13, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Two commits against aiter/configs/model_configs/dsv4_a8w8_blockscale_tuned_gemm.csv:

  1. Add 734 rows for shapes the file does not cover.
  2. Retune 4 rows that it does, all N=65536 K=1536.

Addressing @valarLip's review on #5473 ("move new configs to model_configs"): my
first attempt put these in the global aiter/configs/a8w8_blockscale_tuned_gemm.csv.
That was wrong twice over. The per-model file already exists, and 81 of the rows
I proposed collided with keys already in it. get_config_file merges
model_configs/ with configs/ at load time, so per-model rows belong here.

This continues #5473, which I closed by accident with a force-push of a commit
that had lost its parent. GitHub refuses to reopen it: the reopen check runs
against that PR's recorded head sha, which is the parentless commit, so no state
of the branch can satisfy it. The review there still applies and is answered
here.

Commit 1: the missing shapes

The file already covers four of the six (N, K) families this model drives, but
only at a coarse M ladder, and carries no entry at all for the two vocabulary
projections:

N K before added after
7168 7168 61 594 655
6144 7168 15 39 54
7168 3072 15 39 54
65536 1536 15 39 54
129280 7168 0 16 16
16160 7168 0 7 7

Pure addition: the 310 existing rows are byte-identical, no key is duplicated,
every added row has errRatio = 0.0. Tuned with aiter's own tuner at the default
ERR_RATIO on gfx950 with cu_num=256, matching the rest of the file. M values
follow the ladder the file already uses, restricted to the range each family
reaches at runtime.

Without these rows the entry point reports not found tuned config for every
distinct (M, N, K) DeepSeek-V4 issues; in one 3600-second AgentX run on MI355X
that fires 14066 times.

Commit 2: the four retuned rows

Of the 81 colliding keys, tuning produced a different kernel pick for 19. Rather
than trust us values that come from two separate tuning runs, I re-timed all 19
on one MI355X: both kernels back to back in a single process on the same tensors,
arms interleaved, medians over 7 repeats of 50 iterations after 20 warmup calls.

Only 4 are real, all N=65536 K=1536, where the pick moves from the ck kernels
to the ck_tile aqrm variants:

M N K current proposed speedup
512 65536 1536 78.01us 61.75us 1.26x
128 65536 1536 25.17us 21.11us 1.19x
4 65536 1536 17.15us 16.22us 1.06x
256 65536 1536 42.07us 40.50us 1.04x

The other 15 landed within 1% of the row already in the file and are left
untouched. 13 of them resolve to the identical ck_tile tile config
192x256x128_4x2x1_16x16x128_intrawave_0x1x0 and differ only in a trailing
variant index, so the spread in the CSV was tuner run-to-run noise, not a kernel
difference. No shape regressed.

These four are worth roughly 22us per occurrence set, far below what an
end-to-end run resolves, so nothing below is attributed to them.

Per-shape effect

Same entry point the model calls, aiter.gemm_a8w8_blockscale, run twice in the
same image: once against the current table, where all 40 sampled shapes report
not found tuned config, and once with these rows present. 40 shapes sampled
across the M range of all six families.

M N K fallback tuned speedup
4 6144 7168 25.31 us 21.72 us 1.17x
112 6144 7168 43.64 30.72 1.42x
448 6144 7168 114.00 49.88 2.29x
16384 6144 7168 3458.92 647.58 5.34x
64 7168 3072 12.42 13.15 0.94x
448 7168 3072 58.43 26.25 2.23x
16384 7168 3072 1706.08 362.31 4.71x
7332 7168 7168 1822.48 362.20 5.03x
22832 7168 7168 5620.46 958.78 5.86x
32768 7168 7168 8046.69 1509.39 5.33x
272 16160 7168 225.69 62.01 3.64x
368 16160 7168 295.69 59.81 4.94x
112 65536 1536 59.67 22.06 2.71x
16384 65536 1536 7715.77 2649.00 2.91x
8 129280 7168 160.42 157.02 1.02x
128 129280 7168 609.75 196.51 3.10x

Over the 40 sampled shapes: median 2.79x, mean 2.94x, max 5.86x, and
50,860 us to 11,700 us in aggregate. The gain grows with M, which is the
expected shape of the problem.

One shape regresses, M=64 N=7168 K=3072 at 0.94x. It is the only one of the
40 and the absolute difference is 0.7 us, within run-to-run spread, but it is
listed rather than omitted.

In situ

The same effect on a live run, summing every blockscale-family GEMM kernel per
prefill step per rank from a captured trace:

GEMM per step largest contributor
current table 500.1 ms ck::kernel_gemm_xdl_cshuffle_v3 352.2 ms, n=171, 2059 us/call
+ these rows 244.8 ms same kernel: 25.6 ms, n=46, 555 us/call

GEMM time per prefill step halves, and it is the untuned
kernel_gemm_xdl_cshuffle_v3 fallback that disappears as the work moves onto
tuned cktile kernels.

End to end

DeepSeek-V4-Pro, MI355X x8, TP1/DP8, FP8 KV, MTP-3, concurrency 64, full
3600-second window, everything else held byte-identical:

Configuration tokens/$ P90 interactivity requests a8w8 fallbacks
current table 27,003,552 7.055 2836 14066
+ these rows 30,319,229 8.654 3140 7180

+12.3% tokens/$ and +22.7% P90 interactivity from a data file alone.

The remaining 7180 fallbacks are M values between ladder points. The lookup
still resolves those, they are simply not exact hits.

Relation to #4664

#4664 added DeepSeek-V4 coverage but left these six families out. The rows here
are diffed against current main, not against that PR.

Scope

Tuned on gfx950 with cu_num=256 only. The shipped table is gfx950/256
throughout so this matches it, but these rows say nothing about other
architectures.

@LiuYinfeng01
LiuYinfeng01 requested a review from a team September 13, 2026 16:42
@github-actions github-actions Bot changed the title Extend the DeepSeek-V4 a8w8 blockscale GEMM tunings for gfx950 [Config] Extend the DeepSeek-V4 a8w8 blockscale GEMM tunings for gfx950 Sep 13, 2026
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5485 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

@LiuYinfeng01

Copy link
Copy Markdown
Contributor Author

Closing: reopening #5473 instead, which already carries @valarLip's review. Same commit lands there.

@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5485 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

@LiuYinfeng01

Copy link
Copy Markdown
Contributor Author

Reopened. Disregard my close note above: #5473 turned out to be permanently un-reopenable, so this PR is the one to review.

GitHub rejects reopening #5473 under two mutually exclusive rules. Restoring its recorded head gives the branch has no history in common with ROCm:main (that commit had lost its parent, which is what auto-closed it). Pointing the branch at the corrected commit gives the branch was force-pushed or recreated. There is no branch state that satisfies both.

Content here is unchanged from what I described: 734 rows added and 4 updated in aiter/configs/model_configs/dsv4_a8w8_blockscale_tuned_gemm.csv, addressing @valarLip's review on #5473.

@LiuYinfeng01

Copy link
Copy Markdown
Contributor Author

Closing in favour of a cleaner branch: the shape additions and the four retuned rows are now separate commits, so the new rows can be reviewed apart from the changes to existing ones. Same content, same measurements.

@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5485 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

@LiuYinfeng01

Copy link
Copy Markdown
Contributor Author

@valarLip this carries on from your review on #5473 (move new configs to model_configs). That PR is stuck closed through my own mistake, so the work is here.

Your point is applied, and chasing it down showed I had a second problem: aiter/configs/model_configs/dsv4_a8w8_blockscale_tuned_gemm.csv already exists and already held 81 of the keys I had proposed for the global table. So the change is now 734 added and 4 updated rows in that per-model file, not 815 added to aiter/configs/a8w8_blockscale_tuned_gemm.csv.

Split into two commits so the new shapes can be read apart from the four rows that touch existing entries. Those four were re-timed on hardware rather than compared across tuning runs; 15 other differing picks turned out to be tuner noise and are left alone.

@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5485 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

@zufayu
zufayu requested a review from yifehuan September 14, 2026 01:43
The file already covers four of the six (N, K) families this model drives, but
only at a coarse M ladder, and carries no entry at all for the two vocabulary
projections:

  N       K     before  added  after
  7168    7168      61    594    655
  6144    7168      15     39     54
  7168    3072      15     39     54
  65536   1536      15     39     54
  129280  7168       0     16     16
  16160   7168       0      7      7

Pure addition: the 310 existing rows are byte-identical, no key is duplicated,
and every added row has errRatio 0.0. Tuned with aiter's own tuner at the
default ERR_RATIO on gfx950 with cu_num=256, matching the rest of the file. M
values follow the ladder the file already uses, restricted to the range each
family reaches at runtime.

Without these rows the entry point reports "not found tuned config" for every
distinct (M, N, K) DeepSeek-V4 issues; in one 3600-second AgentX run on MI355X
that fires 14066 times, each falling back to a kernel chosen without knowledge
of the tile.

Per-shape, 40 shapes sampled across the M range of all six families, measured
against the current table where all 40 fall back: median 2.79x, mean 2.94x,
max 5.86x, 50860us -> 11700us in aggregate; the gain grows with M. One shape,
M=64 N=7168 K=3072, is 0.94x, within run-to-run spread at 12us.

In a captured prefill trace the blockscale-family GEMM time per step per rank
falls from 500.1ms to 244.8ms, with the untuned ck::kernel_gemm_xdl_cshuffle_v3
fallback going from 171 calls at 2059us to 46 at 555us.

End to end on DeepSeek-V4-Pro, MI355X x8, TP1/DP8, FP8 KV, MTP-3, concurrency
64, full 3600-second window with everything else byte-identical: 27,003,552 ->
30,319,229 tokens/$ (+12.3%) and 7.055 -> 8.654 P90 interactivity (+22.7%).
At M = 4, 128, 256 and 512 the best pick moves off the ck kernels. Timed here
against the rows they replace, both kernels back to back in one process on the
same tensors with the arms interleaved, medians over 7 repeats of 50 iterations
after 20 warmup calls, on MI355X gfx950 cu_num=256:

  M    current   proposed  speedup
  512   78.01us   61.75us    1.26x
  128   25.17us   21.11us    1.19x
    4   17.15us   16.22us    1.06x
  256   42.07us   40.50us    1.04x

Tuning also produced a different pick for 15 further shapes already in the file.
Re-timing them the same way put every one within 1% of the row already there, so
they are left untouched: 13 resolve to the identical ck_tile tile config
192x256x128_4x2x1_16x16x128_intrawave_0x1x0 and differ only in a trailing
variant index, meaning the spread between the two tuning runs was noise rather
than a kernel difference. No shape regressed.
@LiuYinfeng01
LiuYinfeng01 force-pushed the dsv4-a8w8-blockscale-tuning branch from 8c3921d to 10aada0 Compare September 15, 2026 00:29

@yifehuan yifehuan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@yifehuan
yifehuan merged commit 9252f46 into ROCm:main Sep 15, 2026
135 of 143 checks passed
vgokhale added a commit that referenced this pull request Sep 17, 2026
…ing fixes, fp8 MQA logits split-k and DSv4 tunings (#5573, #5295, #5558, #5603, #5627, #5485) (#5638)

Cherry-picks six already-merged `main` PRs onto `release/v0.1.22` for the `v0.1.22.post1` post release.

| PR | `main` commit | Backport commit | What |
|---|---|---|---|
| #5485 | `9252f4672` | `155534984` | Extend the DeepSeek-V4 a8w8 blockscale GEMM tunings for gfx950 (tuning CSV only) |
| #5573 | `22d2c7c91` | `6b23ba866` | Pad the MXFP4 A4W4 MoE sort extent to a block_size multiple (fixes a HIP illegal memory access) |
| #5295 | `972c8e1fd` | `7d68b0edb` | Skip invalid expert IDs in MoE sorting |
| #5558 | `3fdfca11e` | `dd83a9d17` | Fix MoE routing kernel compile failure |
| #5603 | `a84bd368c` | `a96461997` | Add split-k support for fp8 MQA logits on gfx950 |
| #5627 | `f5ed7dc54` | `a41214712` | Follow-up to #5603: drop chunking when summation folding is unavailable (fixes Triton 3.6 compile) |

Original PRs:
- #5485: #5485
- #5573: #5573
- #5295: #5295
- #5558: #5558
- #5603: #5603
- #5627: #5627

To be published as `v0.1.22.post1` once merged (tag on the merge commit, release automation builds the wheel set).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants