Skip to content

[tune] DSv4 bf16: add gfx950 LM-head GEMM configs (N=129280, K=7168) - #4663

Draft
jiacao-amd wants to merge 2 commits into
ROCm:mainfrom
jiacao-amd:dsv4-bf16-lmhead-tuning
Draft

jiacao-amd wants to merge 2 commits into
ROCm:mainfrom
jiacao-amd:dsv4-bf16-lmhead-tuning

Conversation

@jiacao-amd

Copy link
Copy Markdown
Contributor

Summary

Adds 16 tuned rows to dsv4_bf16_tuned_gemm.csv for the DeepSeek-V4 LM-head projection on gfx950: N=129280, K=7168 (vocab 129280 x hidden 7168), M sweep 1..256.

This shape had no entry in the file, so lookups fell back to an untuned default.

Details

  • The sweep selects torch/native for most M, with an opus mono-tile kernel winning at M=128.
  • Append-only: 16 insertions, 0 deletions. No existing row is modified or removed, and per-gfx counts are otherwise unchanged.
  • Verified no duplicate (gfx, cu_num, M, N, K) keys and no malformed rows after the append.

Test Plan

Tuning data only — no code paths change. The new rows are picked up by AITER_CONFIG.get_config_file(), which globs model_configs/*dsv4_bf16_tuned_gemm*.csv.

Collected on MI355X (gfx950).

Adds 16 tuned rows for the DeepSeek-V4 LM head projection
(vocab 129280 x hidden 7168) on gfx950, covering an M sweep of
1..256. This shape had no entry in dsv4_bf16_tuned_gemm.csv, so
the lookup fell back to an untuned default.

The sweep picks the torch/native path for most M, with an opus
mono-tile kernel winning at M=128.

Append-only: no existing rows are modified or removed.
Signed-off-by: jiacao-amd <jiahui.cao@amd.com>
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 4663 --add-label <label>

The opus row added for M=128 also captures runtime M in [65,80] through
the padded-M fallback: getPaddedM(gl=0) rounds 65..80 up to 80, which has
no row, so the lookup falls through to gl=1 (nextPow2) and lands on 128.

Measured on MI355X (gfx950, cu_num=256), median of 5 rounds of 100 iters,
N=129280, K=7168, bf16:

    M     torch/native   opus 6401   opus vs torch
    66      303.6 us      321.5 us      -5.9%
    72      305.3 us      324.0 us      -6.1%
    80      304.7 us      326.3 us      -7.1%
    98      332.0 us      332.2 us      -0.1%
   104      333.4 us      335.2 us      -0.5%
   112      333.4 us      337.8 us      -1.3%
   120      406.4 us      339.3 us     +16.5%
   124      406.1 us      342.4 us     +15.7%
   128      404.5 us      343.7 us     +15.1%

So the opus tile is a clear win from M=120 up (torch steps from ~333 us to
~405 us there and the 128-wide tile absorbs the step), roughly neutral in
97..112, and a 6-7% loss in 65..80. Adding an explicit torch row at M=80
keeps 65..80 on torch while leaving 97..128 on opus.

The DSv4 logs we sampled only ever hit this shape at M=32, so nothing
regresses today; this just bounds the new row so a future batch size in
that window cannot land on the slower path.
@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

No activity for 15 days, so this is now labelled stale. A push or a comment clears it; keep-open exempts it.

@github-actions github-actions Bot added the stale The PR hasn't been updated in 2 weeks label Oct 7, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

stale The PR hasn't been updated in 2 weeks

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant