Skip to content

[ROCm][Perf] Add AITER paged MXFP4 sparse indexing on gfx950 - #57517

Draft
LiuYinfeng01 wants to merge 4 commits into
vllm-project:mainfrom
LiuYinfeng01:rocm-dsv41-mxfp4-indexer
Draft

LiuYinfeng01 wants to merge 4 commits into
vllm-project:mainfrom
LiuYinfeng01:rocm-dsv41-mxfp4-indexer

Conversation

@LiuYinfeng01

@LiuYinfeng01 LiuYinfeng01 commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Add opt-in AITER paged MXFP4 sparse indexing for DeepSeek V4/V4.1 on ROCm gfx950.

Select with:

--attention-config '{"indexer_kv_dtype":"mxfp4"}'

FP8 remains the default. This PR changes only the sparse indexer's Q/K cache precision and scorer path. It does not change the main attention KV cache and does not overlap the NVFP4 compressed main-KV implementation in #57463.

The AITER scorer implementation, shared-pool page strides, and 64-bit physical-page rebasing are provided by ROCm/aiter#5518. This vLLM PR contains only the integration:

  • packed MXFP4 indexer K-cache production on ROCm;
  • packed Q production and scale layout;
  • zero-copy cache views and AITER prefill/decode dispatch;
  • graph-stable decode schedule reuse across indexer layers;
  • gfx950/geometry/dependency gates and FP8 fallback;
  • focused dispatch and layout contract tests.

This clean branch supersedes #56834. It deliberately excludes the main-attention compressed-KV work from #57489, which overlaps #57463.

Performance and accuracy

DeepSeek-V4.1-Flash, 8x MI355X, TP8+EP8, 8K input / 1K output,
--block-size 128, prefix caching disabled. Each measured concurrency is preceded
by an unmeasured full request batch with a different random seed so runtime JIT
compilation is outside the timed interval. One measured run per arm (n=1):

Concurrency FP8 tok/s MXFP4 tok/s Throughput change FP8 TPOT MXFP4 TPOT
1 112.25 112.10 -0.13% 8.69 ms 8.71 ms
4 405.66 405.19 -0.11% 9.20 ms 9.21 ms
16 1240.42 1244.42 +0.32% 10.88 ms 10.87 ms
64 2561.00 2563.72 +0.11% 17.81 ms 17.87 ms

All performance requests completed with zero failures. The MXFP4 indexer is
throughput-neutral at model level for this V4.1 configuration; its isolated
scorer is faster, but the indexer is a small fraction of total model time.
Earlier negative results were invalid: the timed request compiled uncached
MXFP4 helper kernels and the FP8 run reused prompt prefixes.

Isolated V4.1 scorer shapes on MI355X:

Path FP8 gather+score MXFP4 page 128 Speedup
Decode B64/H32/D128/8K 9.33 us 6.30 us 1.48x
Prefill 8192 queries/H32/8K 610.05 us 309.00 us 1.97x

Both MXFP4 kernels match the exact FP4-dequant reference at cosine 1.000000.

GSM8K-1319, 5-shot greedy, temperature 0, one run per arm:

Arm Accuracy Invalid Output tok/s
FP8 indexer 90.9780% 0% 912.77
MXFP4 indexer 90.6748% 0.0758% 910.73

The accuracy delta is -0.30 percentage points, within single-run noise. The one
invalid MXFP4 response is retained in the reported denominator.

Test plan

python -m pytest -q \
  tests/v1/attention/test_indexer_fp4_dispatch.py \
  tests/kernels/test_fused_indexer_q_dispatch.py

pre-commit run --files \
  docs/design/attention_backends.md \
  tests/kernels/test_fused_indexer_q_dispatch.py \
  tests/v1/attention/test_indexer_fp4_dispatch.py \
  vllm/model_executor/layers/sparse_attn_indexer.py \
  vllm/models/deepseek_v4/common/ops/fused_compress_quant_cache.py \
  vllm/models/deepseek_v4/common/ops/fused_indexer_q.py \
  vllm/v1/attention/backends/mla/indexer.py

Result: 39 passed; applicable pre-commit hooks passed. End-to-end MXFP4 completed all four service loads and GSM8K-1319.

AITER #5518 separately covers exact FP4-dequant agreement, eager and graph replay, shuffled pages, nonzero offsets, 2/4/4.608 GB address boundaries, and page-map/scale sensitivity.

Dependency and non-duplication

AI assistance

AI assistance was used at the author's request. Human review remains required before marking ready.

Signed-off-by: LiuYinfeng01 LiuYinfeng01@users.noreply.github.com

LiuYinfeng01 and others added 3 commits September 18, 2026 07:08
Select AITER paged MXFP4 indexing through
`--attention-config '{"indexer_kv_dtype":"mxfp4"}'` on gfx950 C4 when the
installed AITER build exposes stride-capable FlyDSL FP4 kernels; otherwise
fall back to the default fp8 indexer. Match Q/K scale layouts, preserve
request-local top-k semantics, and flatten speculative causal bounds before
C4 length division.

Assisted-by: OpenAI API coding assistant

Signed-off-by: LiuYinfeng01 <LiuYinfeng01@users.noreply.github.com>
Signed-off-by: Liuyinfeng01 <yinfeliu@amd.com>
Signed-off-by: Liuyinfeng01 <yinfeliu@amd.com>
@mergify

mergify Bot commented Sep 18, 2026

Copy link
Copy Markdown
Contributor

Documentation preview: https://vllm--57517.org.readthedocs.build/en/57517/

@mergify mergify Bot added documentation Improvements or additions to documentation deepseek Related to DeepSeek models DSv4 rocm Related to AMD ROCm labels Sep 18, 2026
@github-project-automation github-project-automation Bot moved this to Todo in AMD Sep 18, 2026
@LiuYinfeng01
LiuYinfeng01 force-pushed the rocm-dsv41-mxfp4-indexer branch 2 times, most recently from 3c9cedc to e34b1f4 Compare September 18, 2026 08:28
@mergify mergify Bot added the DSv4.1 Related to DeepSeek-V4.1 models label Sep 18, 2026
@LiuYinfeng01
LiuYinfeng01 force-pushed the rocm-dsv41-mxfp4-indexer branch 11 times, most recently from 66b61c5 to e91ec67 Compare September 18, 2026 11:21
@LiuYinfeng01
LiuYinfeng01 force-pushed the rocm-dsv41-mxfp4-indexer branch 3 times, most recently from a485029 to 64ce46f Compare September 18, 2026 13:38
@LiuYinfeng01
LiuYinfeng01 force-pushed the rocm-dsv41-mxfp4-indexer branch 4 times, most recently from a6cbb99 to 6f47d5b Compare September 18, 2026 14:19
Signed-off-by: Liuyinfeng01 <yinfeliu@amd.com>
@mergify

mergify Bot commented Sep 24, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @LiuYinfeng01.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deepseek Related to DeepSeek models documentation Improvements or additions to documentation DSv4 DSv4.1 Related to DeepSeek-V4.1 models needs-rebase rocm Related to AMD ROCm

Projects

Status: Todo

Development

Successfully merging this pull request may close these issues.

1 participant