Skip to content

[Config] gptoss bf16 tuned gemm: drop the losing large-M QKV rows on gfx950 (tuned rows lose 16-27% to the hipBLASLt fallback) - #5159

Open
alexnails wants to merge 1 commit into
ROCm:mainfrom
alexnails:gptoss-gemm-csv-large-m
Open

alexnails wants to merge 1 commit into
ROCm:mainfrom
alexnails:gptoss-gemm-csv-large-m

Conversation

@alexnails

Copy link
Copy Markdown

What

The gfx950 rows in gptoss_bf16_tuned_gemm.csv for the gpt-oss QKV projection shape (N=5120, K=2880, bias) at M=4096/8192/16384 pick kernels that lose to the plain hipBLASLt fallback on MI355X:

M (lookup) tuned pick hipBLASLt (torch.addmm) Δ
61366 → 8192 row flydsl, 1.744 ms 1.271 ms −27%
4065 → 4096 row flydsl −16.5%
16384 row "triton auto" −22%

The o-projection shape — which has no tuned row — already runs at 93% of the measured 1424 TF/s GEMM ceiling through the same fallback, which is the tell. Delete the three rows so large-M lookups fall through to hipBLASLt; the small-M decode rows (M ≤ 64) keep flydsl, which wins there (+11.5% at M=64).

Verified on MI355X: the shapes take the hipBLASLt path with outputs identical to torch.addmm (max_abs 0.0), worth ~16 ms/request on a 61k-token gpt-oss prefill (36 launches). Suggests a broader guard for the tuning flow: reject a candidate row at table-generation time if it does not beat the untuned fallback for that shape.

🤖 Generated with Claude Code

The gfx950 rows for the gpt-oss QKV projection shape (N=5120, K=2880,
bias) at M=4096/8192/16384 pick kernels that lose to the plain
hipBLASLt fallback on MI355X: at M=61366 (padded to the 8192 row) the
flydsl pick runs 1.744 ms vs 1.271 ms for torch.addmm (-27%), the 4096
row loses 16.5% at M=4065, and the 16384 'triton auto' row loses 22%.
The o-projection shape, which has no tuned row, already runs at 93% of
the measured GEMM ceiling through the same fallback. Delete the three
rows so large-M lookups fall through to hipBLASLt; the small-M rows
(M<=64, decode) keep flydsl, which wins there (+11.5% at M=64).

Verified on MI355X: the shapes now take the hipBLASLt path with
outputs identical to torch.addmm (max_abs 0.0), and serving TTFT for a
61k-token gpt-oss prefill improves ~16 ms/request (36 launches).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@alexnails
alexnails requested a review from a team September 1, 2026 00:58
@github-actions

github-actions Bot commented Sep 1, 2026

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every PR:

  • ✅ Pre-checks (submodule verification, code formatting)
  • ✅ Aiter op tests (gfx942 + gfx950)
  • ✅ Triton tests on MI35X (only when aiter/ops/triton/** or related paths are changed)

Extended tests (opt-in via labels):

Label Tests
ci:gfx1250-ffm-triton Run the five-shard gfx1250 FFM Triton test suite
ci:triton-300x Run an additional Triton test job on MI300X in PRs; main branch always runs both MI35X and MI300X
multigpu Aiter multi-GPU tests on the 8-GPU runner
ci:sglang SGLang integration tests: DeepSeek-R1-MXFP4 accuracy, Qwen 3.5 accuracy
ci:atom ATOM benchmark: DeepSeek-R1-0528, GPT-OSS-120B
ci:atom_full ATOM accuracy suite for PR and main models from ATOM models_accuracy.json
ci:vllm vLLM benchmark: GPT-OSS-120B, DeepSeek-R1-0528, Kimi-K2.5
ci:all All standard extended tests (excludes ci:atom_full)

Only add ci:atom_full for FlyDSL or Triton upgrades.
Add labels via the sidebar or gh pr edit 5159 --add-label <label>

PR title tags & labels:
Component tags ([Triton/Gluon], [HIP], [CK], [ASM], ...) are added to the PR title and as PR labels automatically from the changed files and re-synced on every push — change-type tags like [fix]/[Perf], op tags like [MLA], and human labels (ci:*) are left untouched. Add the no-auto-title label to opt this PR out.

@github-actions github-actions Bot changed the title gptoss bf16 tuned gemm: drop the losing large-M QKV rows on gfx950 (tuned rows lose 16-27% to the hipBLASLt fallback) [Config] gptoss bf16 tuned gemm: drop the losing large-M QKV rows on gfx950 (tuned rows lose 16-27% to the hipBLASLt fallback) Sep 1, 2026
@github-actions github-actions Bot added the Config label Sep 1, 2026
@zufayu
zufayu requested a review from amd-ruitang3 September 2, 2026 01:22
@zufayu
zufayu requested review from yifehuan and removed request for amd-ruitang3 September 17, 2026 05:22

@yifehuan yifehuan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@github-actions

github-actions Bot commented Oct 7, 2026

Copy link
Copy Markdown
Contributor

No activity for 15 days, so this is now labelled stale. A push or a comment clears it; keep-open exempts it. Nothing is closed automatically.

@github-actions github-actions Bot added the stale The PR hasn't been updated in 2 weeks label Oct 7, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Config stale The PR hasn't been updated in 2 weeks

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants