Repository navigation
[tune] DSv4 bf16: add gfx950 LM-head GEMM configs (N=129280, K=7168) - #4663
Draft
jiacao-amd wants to merge 2 commits into
Draft
jiacao-amd wants to merge 2 commits into
jiacao-amd wants to merge 2 commits into
Conversation
Adds 16 tuned rows for the DeepSeek-V4 LM head projection (vocab 129280 x hidden 7168) on gfx950, covering an M sweep of 1..256. This shape had no entry in dsv4_bf16_tuned_gemm.csv, so the lookup fell back to an untuned default. The sweep picks the torch/native path for most M, with an opus mono-tile kernel winning at M=128. Append-only: no existing rows are modified or removed. Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
Contributor
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
|
The opus row added for M=128 also captures runtime M in [65,80] through
the padded-M fallback: getPaddedM(gl=0) rounds 65..80 up to 80, which has
no row, so the lookup falls through to gl=1 (nextPow2) and lands on 128.
Measured on MI355X (gfx950, cu_num=256), median of 5 rounds of 100 iters,
N=129280, K=7168, bf16:
M torch/native opus 6401 opus vs torch
66 303.6 us 321.5 us -5.9%
72 305.3 us 324.0 us -6.1%
80 304.7 us 326.3 us -7.1%
98 332.0 us 332.2 us -0.1%
104 333.4 us 335.2 us -0.5%
112 333.4 us 337.8 us -1.3%
120 406.4 us 339.3 us +16.5%
124 406.1 us 342.4 us +15.7%
128 404.5 us 343.7 us +15.1%
So the opus tile is a clear win from M=120 up (torch steps from ~333 us to
~405 us there and the 128-wide tile absorbs the step), roughly neutral in
97..112, and a 6-7% loss in 65..80. Adding an explicit torch row at M=80
keeps 65..80 on torch while leaving 97..128 on opus.
The DSv4 logs we sampled only ever hit this shape at M=32, so nothing
regresses today; this just bounds the new row so a future batch size in
that window cannot land on the slower path.
Contributor
|
No activity for 15 days, so this is now labelled |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds 16 tuned rows to
dsv4_bf16_tuned_gemm.csvfor the DeepSeek-V4 LM-head projection on gfx950:N=129280, K=7168(vocab 129280 x hidden 7168), M sweep 1..256.This shape had no entry in the file, so lookups fell back to an untuned default.
Details
torch/nativefor most M, with anopusmono-tile kernel winning at M=128.(gfx, cu_num, M, N, K)keys and no malformed rows after the append.Test Plan
Tuning data only — no code paths change. The new rows are picked up by
AITER_CONFIG.get_config_file(), which globsmodel_configs/*dsv4_bf16_tuned_gemm*.csv.Collected on MI355X (gfx950).