Repository navigation
[ROCm][Perf][GLM-5.3-Flash] BF16 splitk kernel for sparse MLA decode - #58584
dllehr-amd merged 20 commits into
Conversation
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
f813949 to
50a3ade
Compare
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
|
@claude review |
| sparse_len, | ||
| 32, | ||
| ) | ||
| if num_splits == 1: |
There was a problem hiding this comment.
Can you see if we can check num_splits in the original decode check to avoid having this function call appear twice in the same file?
|
Can you also double check with expert parallelism/dcp etc. Just to be sure we aren't exposed anywhere |
…rse-decode Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
| def test_rocm_sparse_triton_decode_routes_on_num_splits( | ||
| monkeypatch, num_splits, num_decode_tokens, expected | ||
| ): | ||
| """Split-K needs both a pure-decode batch and a heuristic asking for splits.""" |
There was a problem hiding this comment.
Please let me know if you think these pure "is the kernel we expect to be called really called" tests are not useful, happy to remove to trim
There was a problem hiding this comment.
Hi @simondanielsson, I just took another look at this, this test does seem close to trivial. If there's no good way to bolster its coverage, we can probably remove it.
There was a problem hiding this comment.
Thanks for having a look, I'll remove it
|
/ci run |
dllehr-amd
left a comment
There was a problem hiding this comment.
Looks good. Thanks @simondanielsson
|
✅ Triggered Buildkite CI #93362 for commit |
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
|
/ci run |
|
✅ Triggered Buildkite CI #93367 for commit |
CI selector (shadow): 15 test steps (26 jobs) instead of 75 (99 jobs)Shadow mode: this changes nothing about what CI runs. It shows what the evidence-based selector would pick for this PR, next to today's rules. How it works. Feedback welcome: reply here if it would skip a step this change needs, or runs something unrelated.
Selector would run (15)
Would skip (today's rules run them) (66)
Would add (today's rules do not run them) (6)
AMD mirrors: would skip (0)none AMD mirrors: would add (88)
4 changed files · base |
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
|
/ci run |
|
✅ Triggered Buildkite CI #93432 for commit |
Purpose
Decode rows in the bf16 sparse MLA path go through the ragged prefill kernel, whose grid is
(q_len, heads_blocks). GLM-5.3-Flash at TP8 has one head block, so during decode with concurrency 4 launches only 4 workgroups on 256 CUs. Clearly underutilized.This new kernel instead uses splitK+reduce over the context dim, i.e. with a grid
(q_len, num_kv_splits, heads_blocks). DSv4 has a similar one.Enabled if topk>=512, as microbenchmarks suggest (GLM uses topk=2048).
Perf: e2e -5-15% ITL. Microbenchmarks: 6x speedup on topk=2048
Test Result
Tested on MI350 unless mentioned otherwise.
gsm8k
MI350
This branch:
Main:
MI300
This branch:
Main:
gpqa
Note no thinking mode enabled. I think low score on main is related to #59413.
This branch
Main
e2e
MI350
MI300, limited bench just for completeness. TPOT improves
8k/1k, conc 32
This branch
Main
Microbenchmark, MI350
topk 128. Here spitK loses, hence we don't use it in this domain
topk 512, splitk wins. Hence _SPARSE_DECODE_BF16_MIN_SPLIT_LEN = 512
topk 2048 (GLM-5.3-Flash)
Comparing this vs
VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLAadded in #53492Testing this branch vs Oct 6 nightly
vllm/vllm-openai-rocm:nightly-bb87d227d4b964abb2a966cdf9194f3d376d9bbe.64K/1K, conc 4:
1k/1k. conc 32:
=> very similar performance, whereas the added splitk kernel is enabled by default and the aiter kernel is opt-in at the moment.
Traces
Before: 191us

After: 17us

Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.