Repository navigation
[Config] [AMD][DSV4] [tune][gfx950] wo_b/wq_b a8w8 blockscale bpreshuffle configs - #5279
Merged
Merged
Conversation
Contributor
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
PR title tags & labels: |
karverma-amd
force-pushed
the
karverma/dsv4-gfx950-gemm-tune
branch
from
September 4, 2026 21:02
6398c51 to
7496e7b
Compare
wo_b (o_proj up, N=7168 K=1024) and wq_b (q up-proj, N=8192 K=1536) had no gfx950 tuned rows and fell back to aiter's default config on MI355X. Add tuned rows across the M ladder spanning concurrency 4-256 (decode M=1 .. prefill M=16384); the tuner selects tuned cktile at large M (wo_b 224->148us @m=16384, ~34% vs default). Validated on DeepSeek-V4-Pro (MI355X, TP8): +5% TTFT at 128k/1k long-context; GSM8K accuracy unchanged (0.948). Co-authored-by: Cursor <cursoragent@cursor.com>
karverma-amd
force-pushed
the
karverma/dsv4-gfx950-gemm-tune
branch
from
September 4, 2026 21:10
7496e7b to
feb952f
Compare
wq_b previously had only M=4,8 in the small/mid range before jumping to 1024, so decode/low-concurrency shapes fell back to aiter's default config. Add 20 autotuned rows (M=1,2,16,32,48,64,80,96,112,128,144,160,176,192,208,224,240, 256,384,512), bringing wq_b to parity with wo_b's ladder. ck wins at small M (~6us), asm BpreShuffle tiles at mid M; all errRatio=0.0. Tuned with: gemm_a8w8_blockscale_tune.py --preshuffle --libtype all (gfx950, MI355X). Co-authored-by: Cursor <cursoragent@cursor.com>
…le config The wq_b shape is comprehensively tuned for gfx950 in a8w8_blockscale_bpreshuffle_tuned_gemm_dsv3.csv (64 rows), which merges into the same runtime lookup (dedup key gfx,M,N,K). The dsv4 wq_b rows were therefore cross-file duplicates and tripped the prebuild merge/dedup check (RuntimeError: duplicate shape entries). dsv3's wq_b timings are within noise of this tune, so removing them loses nothing. Keep only wo_b (7168x1024), which is genuinely uncovered by any blockscale-family config. Net PR is now +15 wo_b rows; merged set has 0 duplicate shapes. Co-authored-by: Cursor <cursoragent@cursor.com>
HaiShaw
previously approved these changes
Sep 6, 2026
coderfeli
reviewed
Sep 6, 2026
Collaborator
|
add 32k config or <32k will go into default |
Reviewers (coderfeli, HaiShaw) noted the wo_b ladder stopped at M=16384, so prefill shapes above 16384 fell back to aiter's default config. Add the autotuned M=32768 row: cktile a8w8_blockscale_cktile_192x256x128_4x2x1_ 16x16x128_intrawave_0x1x0_1, 284.5us, 1690 TFLOP/s, errRatio=0.0 (gfx950, MI355X, --preshuffle --libtype all). Same kernel family as the M=16384 winner, ~linear scaling. wo_b ladder now 1..32768; no duplicate shapes in the blockscale_bpreshuffle merge set. Co-authored-by: Cursor <cursoragent@cursor.com>
coderfeli
approved these changes
Sep 7, 2026
2 of 4 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add MI355X (gfx950) tuned
a8w8block-scale bpreshuffle GEMM configs for twoDeepSeek-V4-Pro attention projections that had no gfx950 rows and were falling back
to aiter's default config:
wo_b(o_proj up)wq_b(q up-proj)Tuned across the M ladder spanning concurrency 4–256 (decode
M=1… prefillM=16384).The tuner selects tuned
cktileat large M.Kernel result (MI355X, gfx950)
wo_b(7168×1024)wq_b(8192×1536)End-to-end result (DeepSeek-V4-Pro, TP8, 128k ISL / 1k OSL)
Long-context prefill is where these projections run at large M most often (a 128k prompt
chunks into 16 prefill forwards at
M=8192). TTFT, 3 reps, conc 8:Tuned is faster in every rep (TTFT and total throughput). Decode is unaffected
(decode
Mis tiny), and GSM8K accuracy is unchanged (0.948).How the config was generated
# untuned shapes: wo_b (7168,1024) + wq_b (8192,1536) across M = 1,2,4,...,16384 python3 csrc/ck_gemm_a8w8_blockscale/gemm_a8w8_blockscale_tune.py \ --preshuffle --libtype all \ -i aiter/configs/a8w8_blockscale_bpreshuffle_untuned_gemm.csv \ -o aiter/configs/model_configs/dsv4_a8w8_blockscale_bpreshuffle_tuned_gemm.csvHow it was validated
Server (DeepSeek-V4-Pro, MI355X, TP8):
Accuracy — GSM8K (1319 questions, 1319 parallel):
python3 benchmark/gsm8k/bench_sglang.py --num-questions 1319 --parallel 1319 --port 8888 # -> Accuracy: 0.948 (unchanged vs default config)Perf — 128k ISL / 1k OSL, streaming TTFT probe (conc 8, 3 reps):
Notes / scope
dsv4_a8w8_blockscale_bpreshuffle_tuned_gemm.csvonly; no existingrows modified.
neutral for small-M decode where these projections are a negligible fraction of the step.