Skip to content

[Perf][ROCm] Use AITER gluon sparse MLA for rope-free BF16 (GLM-5.3-Flash) - #57590

Closed
simondanielsson wants to merge 1 commit into
vllm-project:mainfrom
simondanielsson:feat/sparse-mla-gfx950
Closed

simondanielsson wants to merge 1 commit into
vllm-project:mainfrom
simondanielsson:feat/sparse-mla-gfx950

Conversation

@simondanielsson

@simondanielsson simondanielsson commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Awaiting; AITER bump

GLM-5.3-Flash doesn't use RoPE and falls back to the default triton impl for BF16 KV MLA.

ROCm/aiter#4919 added a gfx950 gluon kernel tailored for BF16 KV non-rope models. We add it here.

Perf gain: TODO

Fixes #57588.

Test Plan

Run GLM-5.3-Flash on MI350 (gfx950), TP8, BF16 KV cache.

1. GSM8k vs main

Expand for details
VLLM_ROCM_USE_AITER=1 vllm serve zai-org/GLM-5.3-Flash \
    --tensor-parallel-size 8 \
    --max-num-seqs 512 \
    --gpu-memory-utilization 0.8 \
    --attention-backend ROCM_AITER_MLA_SPARSE \
    --max-num-batched-tokens 16384
lm_eval --model local-completions --tasks gsm8k \
    --model_args model=zai-org/GLM-5.3-Flash,base_url=http://127.0.0.1:8000/v1/completions \
    --num_fewshot 5

2. bench serve 1k/1k and 8k/1k vs main

Both prefill and decode move to the new kernel, so TTFT and TPOT are both
reported.

Expand for details
vllm bench serve \
    --backend vllm \
    --model zai-org/GLM-5.3-Flash \
    --dataset-name random \
    --random-input-len 1000 \
    --random-output-len 1000 \
    --ignore-eos \
    --seed 5678 \
    --max-concurrency $CONC \
    --num-prompts $((4 * CONC)) --num-warmups $CONC

3. Non-regression on the separated-rope path

DeepSeek-V4-Flash on the same backend, to confirm the ASM and Triton routes are
unchanged.

Expand for details
VLLM_ROCM_USE_AITER=1 vllm serve deepseek-ai/DeepSeek-V4-Flash \
    --tensor-parallel-size 8 --kv-cache-dtype fp8_ds_mla \
    --attention-backend ROCM_AITER_MLA_SPARSE

Test Result

1. GSM8k

TODO

2. 1k/1k

TODO (P50s)

Concurrency Variant QPS TTFT (ms) TPOT (ms) % TPOT improved

3. Non-regression on the separated-rope path

TODO


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

Signed-off-by: simondanielsson <simon.danielsson99@hotmail.com>
@cagrikymk

Copy link
Copy Markdown
Contributor

@simondanielsson Hello, I am the original author of the MLA kernel on the aiter side.

I have a draft ready here: #53492 I was waiting for my PR to be merged to mark the PR ready.

I am planning to extend this API for dsv4, dsv4-flash, GLM-5.x and GLM 5.3-flash. So, could we prioritize my PR over this?

@simondanielsson

Copy link
Copy Markdown
Contributor Author

@simondanielsson Hello, I am the original author of the MLA kernel on the aiter side.

I have a draft ready here: #53492 I was waiting for my PR to be merged to mark the PR ready.

I am planning to extend this API for dsv4, dsv4-flash, GLM-5.x and GLM 5.3-flash. So, could we prioritize my PR over this?

@cagrikymk Thanks for letting me know! Indeed let's prio yours and I'll close this one.

I think the tests I added in this PR we can probably port to yours as well, wdyt?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

glm rocm Related to AMD ROCm

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

[Feature][ROCm][Perf]: Add AITER gluon sparse MLA for rope-free BF16 (GLM-5.3-Flash, gfx950)

2 participants