Repository navigation
[Config] Modify kimik3_bf16_tuned_gemm.csv - #5353
Merged
Merged
Conversation
Co-authored-by: Cursor <cursoragent@cursor.com>
Contributor
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
PR title tags & labels: |
These 9 entries (N=2112, K=7168, gfx950) duplicate shapes already present in kimi/kimik2/dsv3_bf16_tuned_gemm.csv. Since the bf16 tables are merged into one table keyed on (M,N,K,bias,dtype,outdtype,scaleAB,bpreshuffle, cu_num,gfx), they tripped the duplicate-shape check in jit/core.py at build time. The pre-existing entries are faster on 7 of the 9 shapes, and all 9 remain covered by those tables, so dropping them here resolves the merge without losing any tuned coverage.
Contributor
|
@hmahmad26 I have reverted this PR because the tuned config has the older flydsl kernel name(flydsl_gemm*) which can't be passed(will fallback the native path) |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
aiter/configs/model_configs/kimik3_bf16_tuned_gemm.csvAdds BF16-tuned GEMMs for shapes collected using Conc 1 for the Kimi-K3 model.
How the shapes were collected
The shapes come from a Kimi-K3 concurrency-1 agentic run with
AITER_TUNE_GEMM=1, which records everygemm_a16w16shape the model executes. They were then tuned on gfx950 (cu_num=256), with the aiter build from the pinned vLLM ROCm image.What this adds
1,464 tuned rows, all gfx950 /
cu_num=256, by M band:Backends: flydsl 902, asm 379, torch 146, triton 26, opus 11.
Rows already in the file are kept, including all gfx1250 entries. 23 gfx950 shapes existed in both; those take the newly tuned kernel.
Validation
We timed each BF16 GEMM by itself. The tuner builds random A and W tensors and calls
gemm_a16w16(gemm_a16w16_tune.py --run_config). This is kernel time only. It is not a full e2e model run.Default is the kernel aiter already picks from the shipped
model_configstables. For small token counts it often pads M (for example 9 → 16). Tuned is the kernel from this PR's CSV. The last column is how much faster tuned is:(default − tuned) / default.The microbenchmark results are shown below. The MLA / KDA / MoE Shared / DSpark breakdown is a granular analysis for decode GEMMs only.