Repository navigation
[Config] [Perf] Tune DeepSeek-V4.1 Flash BF16 GEMMs for gfx950 - #5618
Closed
Fangzhou-Ai wants to merge 2 commits into
Closed
Fangzhou-Ai wants to merge 2 commits into
Fangzhou-Ai wants to merge 2 commits into
Conversation
Add 14 missing configurations validated against native fallback and sampled neighboring input sizes. Exclude two rows with alias regressions. Co-authored-by: Codex Signed-off-by: fai <fangzhouai@gmail.com>
Contributor
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
PR title tags & labels: |
Fangzhou-Ai
marked this pull request as draft
September 17, 2026 14:27
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Add 14 gfx950/CU256 BF16 GEMM configurations for missing DeepSeek-V4.1-Flash TP4 shapes observed in c1–c128 serving logs. These select existing FlyDSL, ASM, Triton and hipBLASLt kernels instead of the native untuned fallback. The change adds one model CSV and preserves all 3,701 existing BF16 rows.
Kernel results
Measured on MI355X, one GPU at a time, with native AITER dispatch and the full merged configuration table. All 14 exact targets were faster in 31/31 paired measurements; paired median latency reductions range from 6.06% to 29.34%.
Latencies are separate medians; reduction is the median of the 31 per-pair percentages, so it need not equal the percentage calculated from the two displayed medians. Each HIP graph contains 16 calls, with alternating baseline/candidate order and warm caches. Both before/after identical-kernel control median paired changes remained within ±2%.
AITER measured at
b0ced008d89f3fe8ad9fa9388d78b615fd53cb0b; publication basea84bd368cfe014718c6e8c59b51c841aa2e2ab72has identical measured BF16 sources and configuration tables (its intervening change is FP8 attention). Runtime: Torch2.12.0+rocm10.0.0, HIP7.15.26333, Triton3.8.0+git4cff872c.rocm10.0.0, FlyDSL0.3.2, hipBLASLt1.4.1/8d1ae90e. The installed hipBLASLt library SHA256 is726921b73776933606610d0d812fcbe4d448423b16057352f7c9902b152cda38; solution IDs are specific to that library build.Neighboring input sizes
Native padded lookup can route smaller M values to a tuning row. The final combined-table run validated 69 exact/neighbor cases across 16 candidates. Two rows were excluded despite improving at their exact target:
(8192,32,5120)slowed sampled neighbors by 9–34%, and(128,129280,256)slowed sampled neighbors by 10–13%.The 14 retained rows cover 57 measured cases (14 exact targets plus 43 neighbors). None of those cases regressed by more than 2%; the worst was −0.90%. Removing the two rejected rows does not expand any retained row's dispatch domain. M is a GEMM dimension, not request concurrency. Unmeasured sizes remain unvalidated.
Sampled M values and remaining unmeasured ranges
Validation
atol=rtol=0.02against the rounded FP32 reference and native baseline; no mismatch fraction was allowed.err_ratio=0records the observed mismatch fraction under those checks, not bit-exact output. The native BF16 tuner uses the looser.05/.05tolerance.TestConfigShapeCollision.test_bf16,test_selfcheck_detects_planted_duplicate, andtest_selfcheck_passes_on_clean: 3 passed, 0 skipped. Executed with.venv/bin/pythonthrough a local unittest runner that verifies imported source paths, under a private/tmpmount.bash .githooks/pre-commitandgit diff --cached --check: passed for the CSV-only change.This is a bounded synthetic kernel benchmark. Full-model accuracy and end-to-end serving gains have not been measured. CSV
bwis the conventional effective BF16 bytes/time calculation, not measured DRAM traffic.Related work
Open PRs were checked by title/body and relevant CSV diffs on 2026-09-17. No overlapping keys were found, including #4203, #4222, #5353 and #5602. #4663 uses N129280/K7168; these rows use K256. #3836 targets FP32 outputs. #5159 removes different GPT-OSS bias-enabled rows. The MoE configuration work in #5609 is a separate operator.
AI assistance was used for tuning automation, validation analysis and this description.