Skip to content

[Config] Modify kimik3_bf16_tuned_gemm.csv - #5353

Merged
yifehuan merged 2 commits into
ROCm:mainfrom
hmahmad26:kimik3-bf16-tuned-gemm
Sep 19, 2026
Merged

yifehuan merged 2 commits into
ROCm:mainfrom
hmahmad26:kimik3-bf16-tuned-gemm

Conversation

@hmahmad26

@hmahmad26 hmahmad26 commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Modify aiter/configs/model_configs/kimik3_bf16_tuned_gemm.csv

Adds BF16-tuned GEMMs for shapes collected using Conc 1 for the Kimi-K3 model.

How the shapes were collected

The shapes come from a Kimi-K3 concurrency-1 agentic run with AITER_TUNE_GEMM=1, which records every gemm_a16w16 shape the model executes. They were then tuned on gfx950 (cu_num=256), with the aiter build from the pinned vLLM ROCm image.

What this adds

1,464 tuned rows, all gfx950 / cu_num=256, by M band:

Band Count
decode M<=64 572
M 65-512 77
M 513-4096 557
M>4096 258

Backends: flydsl 902, asm 379, torch 146, triton 26, opus 11.

Rows already in the file are kept, including all gfx1250 entries. 23 gfx950 shapes existed in both; those take the newly tuned kernel.

Validation

We timed each BF16 GEMM by itself. The tuner builds random A and W tensors and calls gemm_a16w16 (gemm_a16w16_tune.py --run_config). This is kernel time only. It is not a full e2e model run.

Default is the kernel aiter already picks from the shipped model_configs tables. For small token counts it often pads M (for example 9 → 16). Tuned is the kernel from this PR's CSV. The last column is how much faster tuned is: (default − tuned) / default.

The microbenchmark results are shown below. The MLA / KDA / MoE Shared / DSpark breakdown is a granular analysis for decode GEMMs only.

  • MLA / KDA / MoE shared: the BF16 GEMMs in one target decoder layer of that type (projections and the shared expert). Same ops at every repeating layer.
  • DSpark: the draft model's BF16 GEMMs, counted once (input concat, in-proj, q_b, o_proj, MLP, lm_head, Markov head). The draft has 5 layers; this row is not 5× that cost.
  • Total: sum of the four decode rows above.
  • Prefill GEMMs: timed at prefill M. Not included in Total.
Dtype: BF16
Default µsTuned µsvs default
MLA GEMMs28.8927.40+5.2%
KDA GEMMs26.1925.12+4.1%
MoE Shared GEMMs14.5714.18+2.7%
DSpark GEMMs223.84213.97+4.4%
Total293.49280.67+4.4%
Prefill GEMMs206.13201.77+2.1%

Co-authored-by: Cursor <cursoragent@cursor.com>
@hmahmad26
hmahmad26 requested a review from a team September 8, 2026 18:06
@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5353 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

@github-actions github-actions Bot changed the title Modify kimik3_bf16_tuned_gemm.csv [Config] Modify kimik3_bf16_tuned_gemm.csv Sep 8, 2026
@github-actions github-actions Bot added the Config label Sep 8, 2026
@zufayu
zufayu requested a review from amd-ruitang3 September 9, 2026 01:24
@zufayu
zufayu requested review from yifehuan and removed request for amd-ruitang3 September 17, 2026 05:22
@yifehuan
yifehuan requested review from demonsan and removed request for demonsan September 17, 2026 08:14
These 9 entries (N=2112, K=7168, gfx950) duplicate shapes already present
in kimi/kimik2/dsv3_bf16_tuned_gemm.csv. Since the bf16 tables are merged
into one table keyed on (M,N,K,bias,dtype,outdtype,scaleAB,bpreshuffle,
cu_num,gfx), they tripped the duplicate-shape check in jit/core.py at
build time.

The pre-existing entries are faster on 7 of the 9 shapes, and all 9 remain
covered by those tables, so dropping them here resolves the merge without
losing any tuned coverage.

@yifehuan yifehuan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@yifehuan
yifehuan merged commit 729fe95 into ROCm:main Sep 19, 2026
53 of 56 checks passed
XiaobingSuper added a commit that referenced this pull request Sep 21, 2026
@XiaobingSuper

Copy link
Copy Markdown
Contributor

@hmahmad26 I have reverted this PR because the tuned config has the older flydsl kernel name(flydsl_gemm*) which can't be passed(will fallback the native path)

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants