Repository navigation
[ROCm][Perf] Add AITER paged MXFP4 sparse indexing on gfx950 - #57517
Draft
LiuYinfeng01 wants to merge 4 commits into
Draft
LiuYinfeng01 wants to merge 4 commits into
LiuYinfeng01 wants to merge 4 commits into
Conversation
Select AITER paged MXFP4 indexing through
`--attention-config '{"indexer_kv_dtype":"mxfp4"}'` on gfx950 C4 when the
installed AITER build exposes stride-capable FlyDSL FP4 kernels; otherwise
fall back to the default fp8 indexer. Match Q/K scale layouts, preserve
request-local top-k semantics, and flatten speculative causal bounds before
C4 length division.
Assisted-by: OpenAI API coding assistant
Signed-off-by: LiuYinfeng01 <LiuYinfeng01@users.noreply.github.com>
Signed-off-by: Liuyinfeng01 <yinfeliu@amd.com>
Signed-off-by: Liuyinfeng01 <yinfeliu@amd.com>
This was referenced Sep 18, 2026
Contributor
|
Documentation preview: https://vllm--57517.org.readthedocs.build/en/57517/ |
LiuYinfeng01
force-pushed
the
rocm-dsv41-mxfp4-indexer
branch
2 times, most recently
from
September 18, 2026 08:28
3c9cedc to
e34b1f4
Compare
LiuYinfeng01
force-pushed
the
rocm-dsv41-mxfp4-indexer
branch
11 times, most recently
from
September 18, 2026 11:21
66b61c5 to
e91ec67
Compare
LiuYinfeng01
force-pushed
the
rocm-dsv41-mxfp4-indexer
branch
3 times, most recently
from
September 18, 2026 13:38
a485029 to
64ce46f
Compare
LiuYinfeng01
force-pushed
the
rocm-dsv41-mxfp4-indexer
branch
4 times, most recently
from
September 18, 2026 14:19
a6cbb99 to
6f47d5b
Compare
Signed-off-by: Liuyinfeng01 <yinfeliu@amd.com>
LiuYinfeng01
force-pushed
the
rocm-dsv41-mxfp4-indexer
branch
from
September 18, 2026 14:28
6f47d5b to
317c7fc
Compare
Contributor
|
This pull request has merge conflicts that must be resolved before it can be |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Add opt-in AITER paged MXFP4 sparse indexing for DeepSeek V4/V4.1 on ROCm gfx950.
Select with:
--attention-config '{"indexer_kv_dtype":"mxfp4"}'FP8 remains the default. This PR changes only the sparse indexer's Q/K cache precision and scorer path. It does not change the main attention KV cache and does not overlap the NVFP4 compressed main-KV implementation in #57463.
The AITER scorer implementation, shared-pool page strides, and 64-bit physical-page rebasing are provided by ROCm/aiter#5518. This vLLM PR contains only the integration:
This clean branch supersedes #56834. It deliberately excludes the main-attention compressed-KV work from #57489, which overlaps #57463.
Performance and accuracy
DeepSeek-V4.1-Flash, 8x MI355X, TP8+EP8, 8K input / 1K output,
--block-size 128, prefix caching disabled. Each measured concurrency is precededby an unmeasured full request batch with a different random seed so runtime JIT
compilation is outside the timed interval. One measured run per arm (
n=1):All performance requests completed with zero failures. The MXFP4 indexer is
throughput-neutral at model level for this V4.1 configuration; its isolated
scorer is faster, but the indexer is a small fraction of total model time.
Earlier negative results were invalid: the timed request compiled uncached
MXFP4 helper kernels and the FP8 run reused prompt prefixes.
Isolated V4.1 scorer shapes on MI355X:
Both MXFP4 kernels match the exact FP4-dequant reference at cosine 1.000000.
GSM8K-1319, 5-shot greedy, temperature 0, one run per arm:
The accuracy delta is -0.30 percentage points, within single-run noise. The one
invalid MXFP4 response is retained in the reported denominator.
Test plan
Result: 39 passed; applicable pre-commit hooks passed. End-to-end MXFP4 completed all four service loads and GSM8K-1319.
AITER #5518 separately covers exact FP4-dequant agreement, eager and graph replay, shuffled pages, nonzero offsets, 2/4/4.608 GB address boundaries, and page-map/scale sensitivity.
Dependency and non-duplication
AI assistance
AI assistance was used at the author's request. Human review remains required before marking ready.
Signed-off-by: LiuYinfeng01 LiuYinfeng01@users.noreply.github.com