Repository navigation
[Config] gptoss bf16 tuned gemm: drop the losing large-M QKV rows on gfx950 (tuned rows lose 16-27% to the hipBLASLt fallback) - #5159
Open
alexnails wants to merge 1 commit into
Conversation
The gfx950 rows for the gpt-oss QKV projection shape (N=5120, K=2880, bias) at M=4096/8192/16384 pick kernels that lose to the plain hipBLASLt fallback on MI355X: at M=61366 (padded to the 8192 row) the flydsl pick runs 1.744 ms vs 1.271 ms for torch.addmm (-27%), the 4096 row loses 16.5% at M=4065, and the 16384 'triton auto' row loses 22%. The o-projection shape, which has no tuned row, already runs at 93% of the measured GEMM ceiling through the same fallback. Delete the three rows so large-M lookups fall through to hipBLASLt; the small-M rows (M<=64, decode) keep flydsl, which wins there (+11.5% at M=64). Verified on MI355X: the shapes now take the hipBLASLt path with outputs identical to torch.addmm (max_abs 0.0), and serving TTFT for a 61k-token gpt-oss prefill improves ~16 ms/request (36 launches). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Contributor
🏷️ CI GuideRuns automatically on every PR:
Extended tests (opt-in via labels):
PR title tags & labels: |
Contributor
|
No activity for 15 days, so this is now labelled |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
The gfx950 rows in
gptoss_bf16_tuned_gemm.csvfor the gpt-oss QKV projection shape (N=5120, K=2880, bias) at M=4096/8192/16384 pick kernels that lose to the plain hipBLASLt fallback on MI355X:torch.addmm)The o-projection shape — which has no tuned row — already runs at 93% of the measured 1424 TF/s GEMM ceiling through the same fallback, which is the tell. Delete the three rows so large-M lookups fall through to hipBLASLt; the small-M decode rows (M ≤ 64) keep flydsl, which wins there (+11.5% at M=64).
Verified on MI355X: the shapes take the hipBLASLt path with outputs identical to
torch.addmm(max_abs 0.0), worth ~16 ms/request on a 61k-token gpt-oss prefill (36 launches). Suggests a broader guard for the tuning flow: reject a candidate row at table-generation time if it does not beat the untuned fallback for that shape.🤖 Generated with Claude Code