Skip to content

[ROCm][DSv4.1] Paged MXFP4 sparse indexer on aiter's MQA-logits kernel - #58671

Merged
shen-shanshan merged 36 commits into
vllm-project:mainfrom
cagrikymk:cagri/rocm-mxfp4-indexer-aiter-cache-ops
Sep 29, 2026
Merged

shen-shanshan merged 36 commits into
vllm-project:mainfrom
cagrikymk:cagri/rocm-mxfp4-indexer-aiter-cache-ops

Conversation

@cagrikymk

@cagrikymk cagrikymk commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds a ROCm (gfx950) path for DeepSeek V4.1's indexer with an MXFP4 K cache, built on aiter's paged MXFP4 MQA-logits Gluon kernel. The CUDA path is unchanged. There are two modes:

  • Regular MXFP4 (indexer_kv_dtype="mxfp4"): dense logits over the whole context, read straight from the paged K cache.
  • Sparse MXFP4 (opt-in, also set indexer_sparse_logits=True): the candidate source takes block maxima from the same pass that writes its logits, and the consumer layers score only the candidate pool instead of the whole context. Steps whose contexts are too short for the pool to pay off keep the regular pass, so it helps mainly at long contexts.

Both modes:

  • The logits kernel reads the preshuffled paged K cache in place, so no K gather is needed, prefill included.
  • K is written to the cache by aiter's k_norm_rope_mxfp4_cache, and Q is quantized by q_rope_mxfp4_quant.
  • Each prefill chunk is packed into one launch through query_start_loc. Uniform spec-decode steps run as next_n-row sequences, with the decode schedule built once per step.
  • 128-token KV pages are preferred.

In sparse mode, the candidate pool is resolved once per step and shared by every consumer layer.

Requires aiter with the paged MXFP4 MQA-logits kernel and the cache ops (merged to aiter main) (ROCm/aiter#5834)

Tests

  • tests/kernels/attention/test_rocm_paged_mxfp4_indexer.py
  • tests/v1/attention/test_rocm_paged_mxfp4_indexer_plan.py

Performance

Benchmark: C=32, 128 prompts of 350K tokens, 350 out, spec. decoding = 5+1

metric tot_baseline mxfp4_dense Δ mxfp4_sparse Δ
total tok/s 19,225 22,796 +18.6% 24,863 +29.3%
output tok/s 19.21 22.77 +18.6% 24.84 +29.3%
requests/s 0.0549 0.0651 +18.6% 0.0710 +29.3%
duration s 2,333 1,967 -15.7% 1,804 -22.7%
TTFT mean s 378.2 319.4 -15.5% 292.3 -22.7%
TTFT median s 400.8 339.3 -15.4% 309.0 -22.9%
TTFT p99 s 556.9 473.2 -15.0% 436.7 -21.6%
TPOT mean ms 490.8 412.1 -16.0% 380.5 -22.5%
TPOT median ms 513.8 429.9 -16.3% 399.1 -22.3%
TPOT p99 ms 638.6 525.8 -17.7% 467.4 -26.8%
ITL median ms 816.2 695.0 -14.8% 638.7 -21.7%
E2E mean s 549.5 463.3 -15.7% 425.1 -22.6%
spec acceptance length 1.64 1.65 +0.9% 1.64 +0.3%

Note, since this is done on artificial data, AL is lower (1.65) compared to more realistic 3.5.

8k/1k, NP=10*C, TP4, using mxfp4 dense path (too short kv for sparse), no spec decoding:

C FP8 output tok/s MXFP4 output tok/s Δ FP8 mean TTFT MXFP4 mean TTFT Δ FP8 mean TPOT MXFP4 mean TPOT Δ
1 102.0 109.3 +7.2% 291 ms 290 ms -0.3% 9.53 ms 8.88 ms -6.8%
16 1027.9 1061.0 +3.2% 1611 ms 1580 ms -1.9% 14.00 ms 13.55 ms -3.2%
64 2005.7 2065.9 +3.0% 2048 ms 2025 ms -1.1% 29.92 ms 29.01 ms -3.0%
128 2485.5 2564.0 +3.2% 2835 ms 2798 ms -1.3% 48.73 ms 47.18 ms -3.2%

Accuracy

check tot_baseline mxfp4_sparse mxfp4_dense
GSM8K flexible-extract 0.9651 ± 0.0051 0.9613 ± 0.0053 0.9598 ± 0.0054

Long context accuracy check

  • The text: the first 131,072 tokens of War and Peace (Project Gutenberg), tokenized with the model's own tokenizer. The token ids are saved, so every run scores exactly the same sequence.
  • One request to the running server: /v1/completions with the ids as the prompt, max_tokens=1, temperature=0, prompt_logprobs=0. For every position the server returns the log-probability the model gave to the token that actually comes next, given everything before it. Nothing is sampled; it is one teacher-forced pass over the whole book.
  • The score: NLL per token is minus that log-probability, averaged in four position buckets: 0-16K, 16K-32K, 32K-64K and 64K-128K.
  • Two runs per configuration on the same server.
run 0K-16K 16K-32K 32K-64K 64K-128K
tot_baseline run1 0.0710 0.1302 0.1323 0.1460
tot_baseline run2 0.0717 0.1293 0.1332 0.1484
mxfp4_sparse run1 0.0709 0.1274 0.1334 0.1465
mxfp4_sparse run2 0.0711 0.1269 0.1300 0.1449
mxfp4_dense run1 0.0715 0.1313 0.1353 0.1489
mxfp4_dense run2 0.0721 0.1302 0.1330 0.1494

Not a duplicate

#57517 and #56834 also add a ROCm MXFP4 indexer K cache, with dense scoring. The main
addition here is sparse MXFP4 (indexer_sparse_logits=True): the candidate source takes
block maxima from the same pass that writes its logits, and the consumer layers score only
the candidate pool, read in place from the paged MXFP4 cache. Neither PR supports this
(#56834 rejects candidate blocks outright). For the dense path, this PR uses aiter's paged
MXFP4 MQA-logits kernel (ROCm/aiter#5834, in aiter v0.1.23) and leaves vLLM's shared Q/K
kernels untouched.

Follows the main idea from #56254 while going direct to the paged K for prefill.

AI assistance

Claude is used during development, and the PR author reviewed the AI driven changes.

… kernel

With indexer_kv_dtype="mxfp4" on gfx950 the DeepSeek V4.1 indexer stores
its K cache in the dot-operand order of aiter's paged MXFP4 MQA-logits
kernel (ROCm/aiter cagri/gluon_paged_mxfp4_mqa_logits), and every indexer
layer reads it in place, so prefill no longer gathers K into a workspace.

- The K writer and the fused Q kernel round to e2m1 in plain Triton on
  ROCm; the existing conversion is PTX.
- The candidate source takes its block maxima from its own logits walk.
- With indexer_sparse_logits the four consumers score only the candidate
  pool through the kernel's gather, resolved once per step and shared.
  They fall back to the dense walk masked to the pool below a context
  length gate, and whenever the pool is too large for the gather's int32
  offsets.
- A prefill chunk takes one launch per run of requests with equal query
  rows; decode is flattened to a row per query token.

aiter has to take the page stride from kv_cache.stride(0): the indexer
cache is a view into vLLM's block-major KV pool. Until it does, the first
forward raises instead of reading the wrong pages.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
aiter's paged MXFP4 MQA-logits kernel tiles 64 entries, so at 64-token
pages a ratio-2 indexer page (32 entries) halves its tile. With
indexer_kv_dtype="mxfp4" on ROCm the V4.1 sparse-MLA backend now prefers
128-token blocks, and both it and the MXFP4 indexer backend accept 64 and
128 as kernel block sizes, so neither group splits a block. The FP8
default is unchanged at 64, and an explicit --block-size still wins.

The kernel tests run at both page sizes.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
The dense decode launches read the same rows through the same page geometry
in every layer of a KV cache group, so the metadata builder builds their
work schedule once per step into a preallocated buffer: one for layers 2, 8
and 14, one for layer 20, which the consumers' masked fallback reuses. The
gathering consumers take none. Below four rows, or when the rows already
fill the machine, aiter returns none and the launch keeps the static grid.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
The MXFP4 indexer flattened every decode step to one row per query token,
so each dense layer walked a request's context once per draft token. On a
uniform step (every decode request has next_n rows, CUDA-graph padding
requests only after them) the dense launches (layers 2/8/14, the candidate
source and the masked consumers) now take the same rows as [B, next_n]
sequences: per-request compressed context, each request's first flattened
block-table row as a strided view (so the base builder's block remap
carries over), and the flattened row bounds as cu_ends. At 32 heads and
next_n = 6, aiter plans BLOCK_M = 6, so a workgroup reads each KV tile once
for all six rows. The step's decode schedule is built for that shape.

The choice depends only on the step's shape, so a FULL graph captured on a
uniform batch replays the native launch on a smaller one padded up to it.
Ragged steps, the consumers' gather, and adaptive verification (which
replays FULL graphs on reallocated drafts) keep the flattened rows.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
aiter's build_candidate_gather now widens its resolved offsets to int64 when
the pool outgrows int32 (offset_dtype, ROCm/aiter
cagri/gluon_paged_mxfp4_mqa_logits 7a514b055; 07d4cccdc keeps the buffer
loads under 2 GiB). With it, the reach guard stops pinning the candidate
consumers to the masked dense walk at deployment pool sizes, about 103 GiB
per GPU at 350K TP4. Older aiter keeps the int32 bound.

aiter allocates the candidate lists itself: at int64, 24 B per candidate
block against the compact logits' 4 B per pool column, up to 0.75x the
prefill logits budget at 8-token blocks. The profiling run now reserves that
for the consumer layers.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
ROCm/aiter cagri/gluon_paged_mxfp4_mqa_logits renamed the per-row key bound
cu_ends to row_ends (3f11d618f) and made build_schedule size its slices from
max_model_len, capping a dynamic slice the way the static grid caps one
(5ca52d336). The decode schedule now gets the layer group's logits width, and
its persistent buffer is sized for up to SCHED_SLOT_CAP descriptors, since
the slice cap can take a schedule past target_wgs.

Startup rejects an aiter without the row_ends / query_start_loc API instead of
failing at the first launch. Every aiter that has it takes the page stride
from the cache tensor and widens gather offsets to int64, so the page-stride
check and the consumers' int32 reach guard go.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
aiter's paged MXFP4 kernel takes packed rows through query_start_loc and,
since b2a2fa441, runs them on its static grid without a schedule: one unit
per BLOCK_M rows plus a spare per sequence, each finding its sequence by a
binary search over query_start_loc. A prefill chunk whose requests do not
all share one query length now goes to each dense layer as one such launch,
instead of one launch per run of equal-length requests.

The chunk-local query_start_loc is the step's device copy clamped to the
chunk and rebased: two small device ops per ragged chunk in the metadata
builder. Chunk boundaries, the logits tile and the host-side planning are
unchanged. A chunk with a single run keeps its [B, next_n] launch, and the
consumers' gather keeps launching per request, since varlen does not take a
gather.

VLLM_ROCM_MXFP4_INDEXER_VARLEN=0 goes back to one launch per run.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
…l loop

The loop that marks each block's page bytes reused `block`, the page size,
as its counter, so the cache views after it were built with num_blocks - 1
entries per page. The writer's slot lookup then indexed past the block
table, a device-side assert on the first GPU run.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
The ops, backend and layer modules wrap aiter's paged MXFP4 MQA-logits
kernel, which reads the indexer K cache in place where the FP8 prefill
path first gathers K into a contiguous buffer. Name them after it:

  ops/rocm_mxfp4_indexer.py          -> ops/rocm_paged_mxfp4_indexer.py
  backends/mla/rocm_mxfp4_indexer.py -> backends/mla/rocm_paged_mxfp4_indexer.py
  layers/rocm_sparse_mqa_indexer.py  -> layers/rocm_paged_mxfp4_sparse_mqa_indexer.py

and their tests likewise. The DSv4.1 attention module now imports the
ROCm consumer layer inside its ROCm branch, so importing the model on
other platforms no longer loads it or the ops module.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
The K writer took only the MFMA N from aiter, through mfma_nonk_dim, and
hardcoded the rest of the order the paged MXFP4 kernel reads: the 16-byte
K chunk, borrowed from the FP8 tile, and the wave64 scale split. Read the
whole layout from aiter's cache_format() in one place instead: a frozen
RocmPagedMxfp4CacheLayout (tokens per MFMA N tile, K chunk bytes, scale
lanes) in the indexer ops module, built once per head count, head size
and page size. The backend's page and candidate-block checks, the
attention layer and the tests read it; the shared writer takes it and
uses it only in its ROCm MXFP4 branch. The FP8 tiled order keeps its own
constants, and other platforms pass no layout.

cache_format() is a new aiter API, so the adapter fails closed: a missing
cache_format or key, or a layout the writer does not implement (another
scale order, a chunk or lane split that does not tile a key), stops
startup instead of misordering the cache. It does not report the scale
order or the lane split yet; until it does they default to what aiter's
preshuffle_scales assumes (order 1, 64 // n_per_tile lanes).

For DeepSeek V4.1 (32 heads of 128, 64- and 128-entry pages) the layout
is the constants the writer had, so its kernel is unchanged there.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
…ound

The indexer layer built its K page layout at construction, from
cache_config.block_size // compress_ratio. The block size is not final
there: the first boot on it asked aiter's cache_format() for 8-entry
pages, which it rejects, and the engine failed to start. The GPU tests
pass the layout for their own caches, so they never reached this path.

Resolve it in the indexer cache's bind_kv_cache instead, from the bound
cache's page size, and hand the writer that. Only the ROCm MXFP4 cache
does this; other platforms and the FP8 cache are unchanged.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
…olds

The paged MXFP4 indexer reads the K cache in place, but its prefill chunker
still split like the FP8 path, which gathers K: a chunk's rows were charged
the sum of its requests' contexts, and the candidate consumers walked the
dense layers' split although their logits are only as wide as the pool. At
344K that is 42 consumer launches per layer for one 16K chunk.

- Dense layers: a chunk's logits are [rows, widest compressed context] fp32
  within VLLM_SPARSE_INDEXER_MAX_LOGITS_MB, so the requests packed into a
  chunk no longer add their contexts up. A logits tensor stays under 2 GiB.
- Consumers: their own launches. Each request's rows are rejoined across
  the dense chunks and cut at the rows whose [rows, pool] logits fit the
  logits budget, 8,192 at 512 MiB. Each consumer resolves its launch's pool,
  so no pool outlives its launch and the profiling run's reservation (the
  budget plus one launch's candidate lists) covers the peak.
- The pool-width clamp on the dense split goes: the consumers no longer walk
  that split.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
A prefill chunk launched once, packed through query_start_loc, only when its
requests' row counts differed and VLLM_ROCM_MXFP4_INDEXER_VARLEN was on;
otherwise it launched once per run of requests with equal rows. The packed
launch covers both, so it is now the only one: every chunk launches once, the
kernel finding each row's request from the chunk's query_start_loc. Drops
the environment variable and the per-run launches.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
…prep ops

The ROCm MXFP4 indexer wrote its key cache and quantized its query through
MXFP4 paths added to the shared indexer_k_norm_rope_store and
fused_indexer_q_rope_quant kernels. Those ops now live in aiter, next to the
paged MXFP4 MQA-logits kernel that reads them
(aiter.ops.triton.attention.pa_mqa_logits_mxfp4_cache), and both shared
kernels are back to upstream, so the CUDA path does not change.

The rest of the ROCm wiring moves out of shared files too:
- DeepseekV4Indexer picks its K store, Q quant and indexer layers once. On
  ROCm MXFP4 they come from model_executor/layers/rocm_paged_mxfp4_indexer.py,
  which holds RocmSparseAttnIndexer (the MXFP4 forward_hip that was patched
  into SparseAttnIndexer) and RocmSparseMQAIndexer, built as
  SparseMQAIndexer is.
- sparse_attn_indexer.py and deepseek_v4/attention.py are back to upstream.
  DeepSeek V4.0 refuses the MXFP4 indexer on ROCm from its AMD model.
- vLLM still reads the K page layout from aiter's cache_format() and hands it
  to aiter's K cache op as the shuffle pattern.
- The capability check also requires aiter's two cache-prep ops.

aiter's ops write the same bytes as the kernels they replace: the K pools and
the Q values, scales and weights match exactly at DSv4.1 shapes.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
Every consumer layer of a step reads the same rows through the same page
geometry, so the first one's candidate lists serve the rest, as decode
already does. The reservation now covers a whole step's lists.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
…nce layout

aiter's cache-op tests pin the page order and its round trip through the
logits kernel, and the layer tests here read vLLM's writes back through that
kernel, so only the check that the writer stays inside its pages is left.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
aiter now hands back one int32 slot and one int32 position per candidate
block, whatever the pool's span.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
Plain imports inside cached getters instead of importlib on fixed module
paths and a SimpleNamespace, a signature probe that fails closed, and a
plainer description of the K page's value order.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
…nature

mypy flags the override for dropping kv_cache_spec.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
aiter's kernel is faster on the candidate top-k and from 128 rows on every
indexer shape measured; below that a wide decode step is faster on vLLM's.
Shape alone decides, so FULL graphs no longer need the longest-context
workaround.

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
@mergify mergify Bot added deepseek Related to DeepSeek models DSv4 DSv4.1 Related to DeepSeek-V4.1 models rocm Related to AMD ROCm labels Sep 25, 2026
@github-project-automation github-project-automation Bot moved this to Todo in AMD Sep 25, 2026
@mergify

mergify Bot commented Sep 27, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @cagrikymk.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

cagrikymk and others added 2 commits September 29, 2026 01:46
…ling

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
… off ROCm

The check stays == 4: aiter's paged MXFP4 MQA-logits kernel is gfx950
code, and gfx1250 reports CDNA 5.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
@Fangzhou-Ai
Fangzhou-Ai self-requested a review September 29, 2026 02:15

@Fangzhou-Ai Fangzhou-Ai left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Dead code removed

@cagrikymk

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

Copy link
Copy Markdown

❌ This PR is 1 commit behind upstream main. Your branch must contain every commit currently on upstream main. No new CI build was started. Merge or rebase onto the latest main, then rerun /ci run. To test this branch at your own risk, use /ci run --allow-stale.

@shen-shanshan

Copy link
Copy Markdown
Contributor

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #91719 for commit cd39d3d3672a.

The K store and Q quant held on the instance broke the tests that call
DeepseekV4Indexer's methods on a fake and patch the module functions. The
base indexer calls those functions again, and the ROCm attention layer
builds DeepseekV41RocmMxfp4Indexer when the index cache is MXFP4.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
Fangzhou-Ai added a commit to SemiAnalysisAI/InferenceX that referenced this pull request Sep 29, 2026
#58671 isn't merging in time. Keep #53492's VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA
env var and the image unpin for #58208/#58655, which are unrelated. Drop the
--attention-config flag and its --block-size 128 workaround, both of which
existed only to exercise #58671's ROCm paged MXFP4 sparse-logits indexer.

Signed-off-by: Fangzhou Ai <31551580+Fangzhou-Ai@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Fangzhou-Ai added a commit to SemiAnalysisAI/InferenceX that referenced this pull request Sep 29, 2026
#58671 isn't merging in time. Keep #53492's VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA
env var and the image unpin for #58208/#58655, which are unrelated. Drop the
--attention-config flag and its --block-size 128 workaround, both of which
existed only to exercise #58671's ROCm paged MXFP4 sparse-logits indexer.

Signed-off-by: Fangzhou Ai <31551580+Fangzhou-Ai@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Fangzhou-Ai added a commit to SemiAnalysisAI/InferenceX that referenced this pull request Sep 29, 2026
Stacked on #3555. Re-adds the --attention-config flag enabling #58671's
ROCm paged MXFP4 sparse-logits indexer and the --block-size 128 workaround
its active fp4 indexer needs (vLLM's block-size auto-resolution otherwise
picks 64 across this model's 4+ attention backends). Draft until #58671
merges upstream and lands in a nightly image, so #3555 can run now and this
PR can re-sweep immediately once it's available.

Signed-off-by: Fangzhou Ai <31551580+Fangzhou-Ai@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@cagrikymk

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #91744 for commit 8fc68e0c64bb.

@shen-shanshan shen-shanshan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed by @Fangzhou-Ai

@shen-shanshan
shen-shanshan merged commit 63b9931 into vllm-project:main Sep 29, 2026
235 checks passed
linxiLIFE pushed a commit to linxiLIFE/vllm that referenced this pull request Sep 29, 2026
vllm-project#58671)

Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Shanshan Shen <467638484@qq.com>
Fangzhou-Ai added a commit to Fangzhou-Ai/recipes that referenced this pull request Sep 30, 2026
…parse indexer on MI355X

Add to the MI355X TP override:
- VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=1 for vllm-project/vllm#53492's
  gfx950-only Gluon sparse-MLA kernel.
- --attention-config '{"indexer_kv_dtype":"mxfp4","indexer_sparse_logits":true}'
  for vllm-project/vllm#58671's ROCm paged MXFP4 sparse indexer, replacing
  the dense fp8 indexer path.
- --block-size 128, which the fp4 indexer and SparseMLA backend prefer once
  active (auto-resolution otherwise falls back to 64).

Both vLLM PRs are merged and gfx950-scoped. Mirrors
SemiAnalysisAI/InferenceX#3555 and
SemiAnalysisAI/InferenceX#3571.

Signed-off-by: fai <fangzhouai@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Fangzhou-Ai added a commit to SemiAnalysisAI/InferenceX that referenced this pull request Sep 30, 2026
Stacked on #3555. Re-adds the --attention-config flag enabling #58671's
ROCm paged MXFP4 sparse-logits indexer and the --block-size 128 workaround
its active fp4 indexer needs (vLLM's block-size auto-resolution otherwise
picks 64 across this model's 4+ attention backends). Draft until #58671
merges upstream and lands in a nightly image, so #3555 can run now and this
PR can re-sweep immediately once it's available.

Signed-off-by: Fangzhou Ai <31551580+Fangzhou-Ai@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Fangzhou-Ai added a commit to SemiAnalysisAI/InferenceX that referenced this pull request Sep 30, 2026
Stacked on #3555. Re-adds the --attention-config flag enabling #58671's
ROCm paged MXFP4 sparse-logits indexer and the --block-size 128 workaround
its active fp4 indexer needs (vLLM's block-size auto-resolution otherwise
picks 64 across this model's 4+ attention backends). Draft until #58671
merges upstream and lands in a nightly image, so #3555 can run now and this
PR can re-sweep immediately once it's available.

Signed-off-by: Fangzhou Ai <31551580+Fangzhou-Ai@users.noreply.github.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Fangzhou-Ai added a commit to SemiAnalysisAI/InferenceX that referenced this pull request Oct 1, 2026
…cludes #58671)

Enables vllm-project/vllm#58671's ROCm paged MXFP4 sparse-logits indexer via
--attention-config and --block-size 128, replacing the dense fp8 indexer
path. A live A/B test (TP2 c16, matched 900s window) measured #58671 giving
+14.7/+14.9% p50/p90 interactivity and -9.5/-10.8% p50/p90 e2e latency over
the dense path, with throughput/GPU unchanged. Also drops c128 from both
TP2 and TP4 conc-lists for this recipe.

Co-authored-by: Cursor <cursoragent@cursor.com>
Fangzhou-Ai added a commit to SemiAnalysisAI/InferenceX that referenced this pull request Oct 1, 2026
…cludes #58671)

Enables vllm-project/vllm#58671's ROCm paged MXFP4 sparse-logits indexer via
--attention-config and --block-size 128, replacing the dense fp8 indexer
path. A live A/B test (TP2 c16, matched 900s window) measured #58671 giving
+14.7/+14.9% p50/p90 interactivity and -9.5/-10.8% p50/p90 e2e latency over
the dense path, with throughput/GPU unchanged. Also drops c128 from both
TP2 and TP4 conc-lists for this recipe.

Co-authored-by: Cursor <cursoragent@cursor.com>
chunfangamd added a commit to SemiAnalysisAI/InferenceX that referenced this pull request Oct 2, 2026
…er / DSv4.1-Flash vLLM:启用 ROCm 分页 MXFP4 稀疏 indexer (#3571)

* Repin DSv4.1-Flash MI355X vLLM AgentX to nightly-rocm100-ac9126e5 (includes #58671)

Enables vllm-project/vllm#58671's ROCm paged MXFP4 sparse-logits indexer via
--attention-config and --block-size 128, replacing the dense fp8 indexer
path. A live A/B test (TP2 c16, matched 900s window) measured #58671 giving
+14.7/+14.9% p50/p90 interactivity and -9.5/-10.8% p50/p90 e2e latency over
the dense path, with throughput/GPU unchanged. Also drops c128 from both
TP2 and TP4 conc-lists for this recipe.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Repin DSv4.1-Flash MI355X vLLM AgentX to nightly-rocm100-ac9126e5 (includes #58671)

Enables vllm-project/vllm#58671's ROCm paged MXFP4 sparse-logits indexer via
--attention-config and --block-size 128, replacing the dense fp8 indexer
path. A live A/B test (TP2 c16, matched 900s window) measured #58671 giving
+14.7/+14.9% p50/p90 interactivity and -9.5/-10.8% p50/p90 e2e latency over
the dense path, with throughput/GPU unchanged. Also drops c128 from both
TP2 and TP4 conc-lists for this recipe.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(dsv41flash): align MI355X vLLM changelog and docs with the recipe

- perf-changelog: make the new entry English-only, as AGENTS.md now
  requires; record VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True, which this
  PR also sets; state why c128 is dropped; spell out the vLLM PR
  references.
- amd-master.yaml and configuration-procedures (EN and ZH): replace the
  remaining c128 statements with the current c1-c64 range.

No recipe or master-config value changes.

中文:
- perf-changelog:按 AGENTS.md 的最新要求将新条目改为仅英文;记录本 PR
  同时设置的 VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True;说明移除 c128
  的原因;写全 vLLM PR 引用。
- amd-master.yaml 与 configuration-procedures(中英文):将残留的 c128
  描述更新为当前的 c1-c64 范围。

配方与主配置的取值均无变化。

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(dsv41flash): update MI355X GPU validation for the ac9126e5 pin

Address review feedback: the GPU-validation paragraph still named the
superseded nightly-rocm100-29468dde pin. Point it at
nightly-rocm100-ac9126e5 and run 36824408313 (TP4 and TP2, concurrency
1-64; eval-only TP4 c64 GSM8K 0.9712 strict / 0.9704 flexible), list the
earlier sweeps as superseded, and cite the merged upstream recipe
vllm-project/recipes#1049. The English and Chinese pages change together.

中文:根据评审意见,GPU 验证段落仍写着已被取代的 nightly-rocm100-29468dde。
现改为 nightly-rocm100-ac9126e5 与运行 36824408313(TP4 与 TP2,并发 1-64;
仅评测的 TP4 c64 GSM8K 为 strict 0.9712 / flexible 0.9704),将此前的 sweep
标注为已被取代,并引用已合并的上游配方 vllm-project/recipes#1049。中英文页面
同步更新。

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Chun Fang <chun.fang@amd.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deepseek Related to DeepSeek models DSv4 DSv4.1 Related to DeepSeek-V4.1 models ready ONLY add when PR is ready to merge/full CI is needed rocm Related to AMD ROCm

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

4 participants