Skip to content

[Config] [Perf] Tune DeepSeek-V4.1 Flash BF16 GEMMs for gfx950 - #5618

Closed
Fangzhou-Ai wants to merge 2 commits into
ROCm:mainfrom
Fangzhou-Ai:tune/dsv41-flash-gfx950-bf16-gemm-upstream
Closed

Fangzhou-Ai wants to merge 2 commits into
ROCm:mainfrom
Fangzhou-Ai:tune/dsv41-flash-gfx950-bf16-gemm-upstream

Conversation

@Fangzhou-Ai

Copy link
Copy Markdown
Contributor

Summary

Add 14 gfx950/CU256 BF16 GEMM configurations for missing DeepSeek-V4.1-Flash TP4 shapes observed in c1–c128 serving logs. These select existing FlyDSL, ASM, Triton and hipBLASLt kernels instead of the native untuned fallback. The change adds one model CSV and preserves all 3,701 existing BF16 rows.

Kernel results

Measured on MI355X, one GPU at a time, with native AITER dispatch and the full merged configuration table. All 14 exact targets were faster in 31/31 paired measurements; paired median latency reductions range from 6.06% to 29.34%.

M N K Backend Baseline µs Selected µs Paired reduction Faster pairs
6 32 5120 flydsl 5.750 4.300 24.57% 31/31
16380 32 5120 flydsl 28.658 25.920 9.58% 31/31
16384 32 5120 flydsl 28.745 25.793 10.17% 31/31
12 128 512 triton 4.010 2.873 29.34% 31/31
24 128 512 flydsl 4.072 3.200 21.40% 31/31
48 128 512 asm 4.160 3.765 9.54% 31/31
96 128 512 flydsl 4.217 3.605 14.31% 31/31
192 128 512 flydsl 4.200 3.543 15.52% 31/31
384 128 512 flydsl 4.298 3.853 10.70% 31/31
768 128 512 flydsl 4.393 3.540 19.47% 31/31
8192 128 512 asm 6.053 4.912 18.53% 31/31
16380 128 512 hipblaslt 7.925 7.043 10.37% 31/31
192 32320 5120 hipblaslt 85.488 80.323 6.06% 31/31
8 129280 256 flydsl 14.825 11.653 21.12% 31/31

Latencies are separate medians; reduction is the median of the 31 per-pair percentages, so it need not equal the percentage calculated from the two displayed medians. Each HIP graph contains 16 calls, with alternating baseline/candidate order and warm caches. Both before/after identical-kernel control median paired changes remained within ±2%.

AITER measured at b0ced008d89f3fe8ad9fa9388d78b615fd53cb0b; publication base a84bd368cfe014718c6e8c59b51c841aa2e2ab72 has identical measured BF16 sources and configuration tables (its intervening change is FP8 attention). Runtime: Torch 2.12.0+rocm10.0.0, HIP 7.15.26333, Triton 3.8.0+git4cff872c.rocm10.0.0, FlyDSL 0.3.2, hipBLASLt 1.4.1/8d1ae90e. The installed hipBLASLt library SHA256 is 726921b73776933606610d0d812fcbe4d448423b16057352f7c9902b152cda38; solution IDs are specific to that library build.

Neighboring input sizes

Native padded lookup can route smaller M values to a tuning row. The final combined-table run validated 69 exact/neighbor cases across 16 candidates. Two rows were excluded despite improving at their exact target: (8192,32,5120) slowed sampled neighbors by 9–34%, and (128,129280,256) slowed sampled neighbors by 10–13%.

The 14 retained rows cover 57 measured cases (14 exact targets plus 43 neighbors). None of those cases regressed by more than 2%; the worst was −0.90%. Removing the two rejected rows does not expand any retained row's dispatch domain. M is a GEMM dimension, not request concurrency. Unmeasured sizes remain unvalidated.

Sampled M values and remaining unmeasured ranges
CSV M,N,K Measured actual M Reduction range Unmeasured actual M
6,32,5120 6 24.57% to 24.57% none
16380,32,5120 16380 9.58% to 9.58% none
16384,32,5120 8193, 12288, 16368, 16372, 16376, 16384 -0.90% to 10.59% 8194–12287, 12289–16367, 16369–16371, 16373–16375, 16377–16379, 16381–16383
12,128,512 12 29.34% to 29.34% none
24,128,512 24 21.40% to 21.40% none
48,128,512 33, 34, 36, 40, 42, 48 9.28% to 10.11% 35, 37–39, 41, 43–47
96,128,512 81, 84, 88, 90, 92, 96 13.30% to 14.84% 82–83, 85–87, 89, 91, 93–95
192,128,512 177, 181, 182, 184, 190, 192 14.68% to 15.79% 178–180, 183, 185–189, 191
384,128,512 353, 368, 370, 372, 377, 384 6.99% to 10.70% 354–367, 369, 371, 373–376, 378–383
768,128,512 737, 741, 752, 756, 761, 768 19.03% to 20.30% 738–740, 742–751, 753–755, 757–760, 762–767
8192,128,512 4097, 4217, 4264, 4580, 6144, 8192 18.53% to 25.59% 4098–4216, 4218–4263, 4265–4579, 4581–6143, 6145–8191
16380,128,512 16380 10.37% to 10.37% none
192,32320,5120 177, 181, 184, 186, 187, 192 5.90% to 6.83% 178–180, 182–183, 185, 188–191
8,129280,256 5, 6, 7, 8 20.57% to 21.12% none

Validation

  • All 69 final cases completed with native baseline/candidate selection checked against the full CSVs. Eager calls, captured graph outputs, changed/stressed/restored inputs, finite outputs and unchanged inputs/weights were checked.
  • Every checked element passed atol=rtol=0.02 against the rounded FP32 reference and native baseline; no mismatch fraction was allowed. err_ratio=0 records the observed mismatch fraction under those checks, not bit-exact output. The native BF16 tuner uses the looser .05/.05 tolerance.
  • CPU publication audit passed: all old rows and measured source files preserved, exact kernel selections and timing metadata reproduced, native config discovery checked, no duplicate keys or expanded retained lookup domains.
  • Existing TestConfigShapeCollision.test_bf16, test_selfcheck_detects_planted_duplicate, and test_selfcheck_passes_on_clean: 3 passed, 0 skipped. Executed with .venv/bin/python through a local unittest runner that verifies imported source paths, under a private /tmp mount.
  • bash .githooks/pre-commit and git diff --cached --check: passed for the CSV-only change.

This is a bounded synthetic kernel benchmark. Full-model accuracy and end-to-end serving gains have not been measured. CSV bw is the conventional effective BF16 bytes/time calculation, not measured DRAM traffic.

Related work

Open PRs were checked by title/body and relevant CSV diffs on 2026-09-17. No overlapping keys were found, including #4203, #4222, #5353 and #5602. #4663 uses N129280/K7168; these rows use K256. #3836 targets FP32 outputs. #5159 removes different GPT-OSS bias-enabled rows. The MoE configuration work in #5609 is a separate operator.

AI assistance was used for tuning automation, validation analysis and this description.

Add 14 missing configurations validated against native fallback and sampled
neighboring input sizes. Exclude two rows with alias regressions.

Co-authored-by: Codex
Signed-off-by: fai <fangzhouai@gmail.com>
@Fangzhou-Ai
Fangzhou-Ai requested a review from a team September 17, 2026 04:12
@github-actions github-actions Bot changed the title [Perf] Tune DeepSeek-V4.1 Flash BF16 GEMMs for gfx950 [Config] [Perf] Tune DeepSeek-V4.1 Flash BF16 GEMMs for gfx950 Sep 17, 2026
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5618 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

@Fangzhou-Ai
Fangzhou-Ai marked this pull request as draft September 17, 2026 14:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant