Repository navigation
[Perf][ROCm] Use AITER gluon sparse MLA for rope-free BF16 (GLM-5.3-Flash) - #57590
simondanielsson wants to merge 1 commit into
Conversation
Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
|
@simondanielsson Hello, I am the original author of the MLA kernel on the aiter side. I have a draft ready here: #53492 I was waiting for my PR to be merged to mark the PR ready. I am planning to extend this API for dsv4, dsv4-flash, GLM-5.x and GLM 5.3-flash. So, could we prioritize my PR over this? |
@cagrikymk Thanks for letting me know! Indeed let's prio yours and I'll close this one. I think the tests I added in this PR we can probably port to yours as well, wdyt? |
Purpose
Awaiting; AITER bump
GLM-5.3-Flash doesn't use RoPE and falls back to the default triton impl for BF16 KV MLA.
ROCm/aiter#4919 added a gfx950 gluon kernel tailored for BF16 KV non-rope models. We add it here.
Perf gain: TODO
Fixes #57588.
Test Plan
Run GLM-5.3-Flash on MI350 (gfx950), TP8, BF16 KV cache.
1. GSM8k vs main
Expand for details
VLLM_ROCM_USE_AITER=1 vllm serve zai-org/GLM-5.3-Flash \ --tensor-parallel-size 8 \ --max-num-seqs 512 \ --gpu-memory-utilization 0.8 \ --attention-backend ROCM_AITER_MLA_SPARSE \ --max-num-batched-tokens 16384lm_eval --model local-completions --tasks gsm8k \ --model_args model=zai-org/GLM-5.3-Flash,base_url=http://127.0.0.1:8000/v1/completions \ --num_fewshot 52. bench serve 1k/1k and 8k/1k vs main
Both prefill and decode move to the new kernel, so TTFT and TPOT are both
reported.
Expand for details
vllm bench serve \ --backend vllm \ --model zai-org/GLM-5.3-Flash \ --dataset-name random \ --random-input-len 1000 \ --random-output-len 1000 \ --ignore-eos \ --seed 5678 \ --max-concurrency $CONC \ --num-prompts $((4 * CONC)) --num-warmups $CONC3. Non-regression on the separated-rope path
DeepSeek-V4-Flash on the same backend, to confirm the ASM and Triton routes are
unchanged.
Expand for details
VLLM_ROCM_USE_AITER=1 vllm serve deepseek-ai/DeepSeek-V4-Flash \ --tensor-parallel-size 8 --kv-cache-dtype fp8_ds_mla \ --attention-backend ROCM_AITER_MLA_SPARSETest Result
1. GSM8k
TODO
2. 1k/1k
TODO (P50s)
3. Non-regression on the separated-rope path
TODO
Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.