Repository navigation
[CK] blockscale split-K: pick the pipeline from the per-split loop count - #5847
Merged
yifehuan merged 1 commit intoSep 28, 2026
Merged
Conversation
…op count gemm_a8w8_blockscale_cktile chose has_hot_loop/tail_num from the full K, but each split only runs K / (k_batch * K_Tile) loops. With two loops per split the pipeline did not match and split-K outputs were wrong (69%-97% of elements with the default gfx942 kernel). Compute the loop count per split, as the other CK-Tile launchers in aiter do, and make test_splitk_correctness assert and cover short splits. Co-authored-by: Cursor <cursoragent@cursor.com>
Contributor
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
PR title tags & labels: |
3 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
gemm_a8w8_blockscale_cktilepicks the pipeline variant (has_hot_loop,tail_num) from the loop count of the full K. With split-K, each split only runsK / (k_batch * K_Tile)loops. When a split is short (two K loops), the selected variant does not match the loops it actually runs, and the output is silently wrong.This PR computes the loop count per split, the same way the other CK-Tile launchers in aiter already do:
ck_deepgemm,ck_tile_gemm_moe_2stagesandcktile_gemm_a8w8_bpreshuffleall usek_grain = k_batch * K_Tile.Results (MI325X, gfx942)
Public
gemm_a8w8_blockscale_cktilewith the default kernel (K_Tile 128), split-K output compared against split-K 0, fraction of mismatched elements:Every CK-Tile tuning candidate on gfx942 (80 instances × splitK 0–3 × 8 shapes, M from 1 to 16384, (N, K) = (128, 6144) and (256, 4096); 2560 runs), compared against a torch reference:
The tuner already rejected these candidates through errRatio, so no tuned config was affected. Direct split-K calls with short splits were.
Test plan
op_tests/test_gemm_a8w8_blockscale.py:test_splitk_correctnessnow asserts on the error ratio and covers three shapes with two K loops per split. Before this change the new cases fail (cktile_err=0.8716); after it, all cases pass. The existing cases pass either way.Found by Hyperloom; reviewed and measured by hand.
It showed up while tuning the gfx942 table in #5839, which started from Hyperloom optimizing
Qwen/Qwen3.5-397B-A17B-FP8.