Repository navigation
[ROCm][DSv4.1] Paged MXFP4 sparse indexer on aiter's MQA-logits kernel - #58671
Merged
shen-shanshan merged 36 commits intoSep 29, 2026
Merged
shen-shanshan merged 36 commits into
shen-shanshan merged 36 commits into
Conversation
… kernel With indexer_kv_dtype="mxfp4" on gfx950 the DeepSeek V4.1 indexer stores its K cache in the dot-operand order of aiter's paged MXFP4 MQA-logits kernel (ROCm/aiter cagri/gluon_paged_mxfp4_mqa_logits), and every indexer layer reads it in place, so prefill no longer gathers K into a workspace. - The K writer and the fused Q kernel round to e2m1 in plain Triton on ROCm; the existing conversion is PTX. - The candidate source takes its block maxima from its own logits walk. - With indexer_sparse_logits the four consumers score only the candidate pool through the kernel's gather, resolved once per step and shared. They fall back to the dense walk masked to the pool below a context length gate, and whenever the pool is too large for the gather's int32 offsets. - A prefill chunk takes one launch per run of requests with equal query rows; decode is flattened to a row per query token. aiter has to take the page stride from kv_cache.stride(0): the indexer cache is a view into vLLM's block-major KV pool. Until it does, the first forward raises instead of reading the wrong pages. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
aiter's paged MXFP4 MQA-logits kernel tiles 64 entries, so at 64-token pages a ratio-2 indexer page (32 entries) halves its tile. With indexer_kv_dtype="mxfp4" on ROCm the V4.1 sparse-MLA backend now prefers 128-token blocks, and both it and the MXFP4 indexer backend accept 64 and 128 as kernel block sizes, so neither group splits a block. The FP8 default is unchanged at 64, and an explicit --block-size still wins. The kernel tests run at both page sizes. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
The dense decode launches read the same rows through the same page geometry in every layer of a KV cache group, so the metadata builder builds their work schedule once per step into a preallocated buffer: one for layers 2, 8 and 14, one for layer 20, which the consumers' masked fallback reuses. The gathering consumers take none. Below four rows, or when the rows already fill the machine, aiter returns none and the launch keeps the static grid. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
The MXFP4 indexer flattened every decode step to one row per query token, so each dense layer walked a request's context once per draft token. On a uniform step (every decode request has next_n rows, CUDA-graph padding requests only after them) the dense launches (layers 2/8/14, the candidate source and the masked consumers) now take the same rows as [B, next_n] sequences: per-request compressed context, each request's first flattened block-table row as a strided view (so the base builder's block remap carries over), and the flattened row bounds as cu_ends. At 32 heads and next_n = 6, aiter plans BLOCK_M = 6, so a workgroup reads each KV tile once for all six rows. The step's decode schedule is built for that shape. The choice depends only on the step's shape, so a FULL graph captured on a uniform batch replays the native launch on a smaller one padded up to it. Ragged steps, the consumers' gather, and adaptive verification (which replays FULL graphs on reallocated drafts) keep the flattened rows. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
aiter's build_candidate_gather now widens its resolved offsets to int64 when the pool outgrows int32 (offset_dtype, ROCm/aiter cagri/gluon_paged_mxfp4_mqa_logits 7a514b055; 07d4cccdc keeps the buffer loads under 2 GiB). With it, the reach guard stops pinning the candidate consumers to the masked dense walk at deployment pool sizes, about 103 GiB per GPU at 350K TP4. Older aiter keeps the int32 bound. aiter allocates the candidate lists itself: at int64, 24 B per candidate block against the compact logits' 4 B per pool column, up to 0.75x the prefill logits budget at 8-token blocks. The profiling run now reserves that for the consumer layers. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
ROCm/aiter cagri/gluon_paged_mxfp4_mqa_logits renamed the per-row key bound cu_ends to row_ends (3f11d618f) and made build_schedule size its slices from max_model_len, capping a dynamic slice the way the static grid caps one (5ca52d336). The decode schedule now gets the layer group's logits width, and its persistent buffer is sized for up to SCHED_SLOT_CAP descriptors, since the slice cap can take a schedule past target_wgs. Startup rejects an aiter without the row_ends / query_start_loc API instead of failing at the first launch. Every aiter that has it takes the page stride from the cache tensor and widens gather offsets to int64, so the page-stride check and the consumers' int32 reach guard go. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
aiter's paged MXFP4 kernel takes packed rows through query_start_loc and, since b2a2fa441, runs them on its static grid without a schedule: one unit per BLOCK_M rows plus a spare per sequence, each finding its sequence by a binary search over query_start_loc. A prefill chunk whose requests do not all share one query length now goes to each dense layer as one such launch, instead of one launch per run of equal-length requests. The chunk-local query_start_loc is the step's device copy clamped to the chunk and rebased: two small device ops per ragged chunk in the metadata builder. Chunk boundaries, the logits tile and the host-side planning are unchanged. A chunk with a single run keeps its [B, next_n] launch, and the consumers' gather keeps launching per request, since varlen does not take a gather. VLLM_ROCM_MXFP4_INDEXER_VARLEN=0 goes back to one launch per run. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
…l loop The loop that marks each block's page bytes reused `block`, the page size, as its counter, so the cache views after it were built with num_blocks - 1 entries per page. The writer's slot lookup then indexed past the block table, a device-side assert on the first GPU run. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
The ops, backend and layer modules wrap aiter's paged MXFP4 MQA-logits kernel, which reads the indexer K cache in place where the FP8 prefill path first gathers K into a contiguous buffer. Name them after it: ops/rocm_mxfp4_indexer.py -> ops/rocm_paged_mxfp4_indexer.py backends/mla/rocm_mxfp4_indexer.py -> backends/mla/rocm_paged_mxfp4_indexer.py layers/rocm_sparse_mqa_indexer.py -> layers/rocm_paged_mxfp4_sparse_mqa_indexer.py and their tests likewise. The DSv4.1 attention module now imports the ROCm consumer layer inside its ROCm branch, so importing the model on other platforms no longer loads it or the ops module. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
The K writer took only the MFMA N from aiter, through mfma_nonk_dim, and hardcoded the rest of the order the paged MXFP4 kernel reads: the 16-byte K chunk, borrowed from the FP8 tile, and the wave64 scale split. Read the whole layout from aiter's cache_format() in one place instead: a frozen RocmPagedMxfp4CacheLayout (tokens per MFMA N tile, K chunk bytes, scale lanes) in the indexer ops module, built once per head count, head size and page size. The backend's page and candidate-block checks, the attention layer and the tests read it; the shared writer takes it and uses it only in its ROCm MXFP4 branch. The FP8 tiled order keeps its own constants, and other platforms pass no layout. cache_format() is a new aiter API, so the adapter fails closed: a missing cache_format or key, or a layout the writer does not implement (another scale order, a chunk or lane split that does not tile a key), stops startup instead of misordering the cache. It does not report the scale order or the lane split yet; until it does they default to what aiter's preshuffle_scales assumes (order 1, 64 // n_per_tile lanes). For DeepSeek V4.1 (32 heads of 128, 64- and 128-entry pages) the layout is the constants the writer had, so its kernel is unchanged there. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
…ound The indexer layer built its K page layout at construction, from cache_config.block_size // compress_ratio. The block size is not final there: the first boot on it asked aiter's cache_format() for 8-entry pages, which it rejects, and the engine failed to start. The GPU tests pass the layout for their own caches, so they never reached this path. Resolve it in the indexer cache's bind_kv_cache instead, from the bound cache's page size, and hand the writer that. Only the ROCm MXFP4 cache does this; other platforms and the FP8 cache are unchanged. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
…olds The paged MXFP4 indexer reads the K cache in place, but its prefill chunker still split like the FP8 path, which gathers K: a chunk's rows were charged the sum of its requests' contexts, and the candidate consumers walked the dense layers' split although their logits are only as wide as the pool. At 344K that is 42 consumer launches per layer for one 16K chunk. - Dense layers: a chunk's logits are [rows, widest compressed context] fp32 within VLLM_SPARSE_INDEXER_MAX_LOGITS_MB, so the requests packed into a chunk no longer add their contexts up. A logits tensor stays under 2 GiB. - Consumers: their own launches. Each request's rows are rejoined across the dense chunks and cut at the rows whose [rows, pool] logits fit the logits budget, 8,192 at 512 MiB. Each consumer resolves its launch's pool, so no pool outlives its launch and the profiling run's reservation (the budget plus one launch's candidate lists) covers the peak. - The pool-width clamp on the dense split goes: the consumers no longer walk that split. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
A prefill chunk launched once, packed through query_start_loc, only when its requests' row counts differed and VLLM_ROCM_MXFP4_INDEXER_VARLEN was on; otherwise it launched once per run of requests with equal rows. The packed launch covers both, so it is now the only one: every chunk launches once, the kernel finding each row's request from the chunk's query_start_loc. Drops the environment variable and the per-run launches. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
…prep ops The ROCm MXFP4 indexer wrote its key cache and quantized its query through MXFP4 paths added to the shared indexer_k_norm_rope_store and fused_indexer_q_rope_quant kernels. Those ops now live in aiter, next to the paged MXFP4 MQA-logits kernel that reads them (aiter.ops.triton.attention.pa_mqa_logits_mxfp4_cache), and both shared kernels are back to upstream, so the CUDA path does not change. The rest of the ROCm wiring moves out of shared files too: - DeepseekV4Indexer picks its K store, Q quant and indexer layers once. On ROCm MXFP4 they come from model_executor/layers/rocm_paged_mxfp4_indexer.py, which holds RocmSparseAttnIndexer (the MXFP4 forward_hip that was patched into SparseAttnIndexer) and RocmSparseMQAIndexer, built as SparseMQAIndexer is. - sparse_attn_indexer.py and deepseek_v4/attention.py are back to upstream. DeepSeek V4.0 refuses the MXFP4 indexer on ROCm from its AMD model. - vLLM still reads the K page layout from aiter's cache_format() and hands it to aiter's K cache op as the shuffle pattern. - The capability check also requires aiter's two cache-prep ops. aiter's ops write the same bytes as the kernels they replace: the K pools and the Q values, scales and weights match exactly at DSv4.1 shapes. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
Every consumer layer of a step reads the same rows through the same page geometry, so the first one's candidate lists serve the rest, as decode already does. The reservation now covers a whole step's lists. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
…nce layout aiter's cache-op tests pin the page order and its round trip through the logits kernel, and the layer tests here read vLLM's writes back through that kernel, so only the check that the writer stays inside its pages is left. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
aiter now hands back one int32 slot and one int32 position per candidate block, whatever the pool's span. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
Plain imports inside cached getters instead of importlib on fixed module paths and a SimpleNamespace, a signature probe that fails closed, and a plainer description of the K page's value order. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
…nature mypy flags the override for dropping kv_cache_spec. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
aiter's kernel is faster on the candidate top-k and from 128 rows on every indexer shape measured; below that a wide decode step is faster on vLLM's. Shape alone decides, so FULL graphs no longer need the longest-context workaround. Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
Contributor
|
This pull request has merge conflicts that must be resolved before it can be |
…ling Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
… off ROCm The check stays == 4: aiter's paged MXFP4 MQA-logits kernel is gfx950 code, and gfx1250 reports CDNA 5. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
Fangzhou-Ai
self-requested a review
September 29, 2026 02:15
Contributor
Author
|
/ci run |
|
❌ This PR is 1 commit behind upstream |
Contributor
|
/ci run |
|
✅ Triggered Buildkite CI #91719 for commit |
The K store and Q quant held on the instance broke the tests that call DeepseekV4Indexer's methods on a fake and patch the module functions. The base indexer calls those functions again, and the ROCm attention layer builds DeepseekV41RocmMxfp4Indexer when the index cache is MXFP4. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com>
Fangzhou-Ai
added a commit
to SemiAnalysisAI/InferenceX
that referenced
this pull request
Sep 29, 2026
#58671 isn't merging in time. Keep #53492's VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA env var and the image unpin for #58208/#58655, which are unrelated. Drop the --attention-config flag and its --block-size 128 workaround, both of which existed only to exercise #58671's ROCm paged MXFP4 sparse-logits indexer. Signed-off-by: Fangzhou Ai <31551580+Fangzhou-Ai@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Fangzhou-Ai
added a commit
to SemiAnalysisAI/InferenceX
that referenced
this pull request
Sep 29, 2026
#58671 isn't merging in time. Keep #53492's VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA env var and the image unpin for #58208/#58655, which are unrelated. Drop the --attention-config flag and its --block-size 128 workaround, both of which existed only to exercise #58671's ROCm paged MXFP4 sparse-logits indexer. Signed-off-by: Fangzhou Ai <31551580+Fangzhou-Ai@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
4 of 10 tasks
Fangzhou-Ai
added a commit
to SemiAnalysisAI/InferenceX
that referenced
this pull request
Sep 29, 2026
Stacked on #3555. Re-adds the --attention-config flag enabling #58671's ROCm paged MXFP4 sparse-logits indexer and the --block-size 128 workaround its active fp4 indexer needs (vLLM's block-size auto-resolution otherwise picks 64 across this model's 4+ attention backends). Draft until #58671 merges upstream and lands in a nightly image, so #3555 can run now and this PR can re-sweep immediately once it's available. Signed-off-by: Fangzhou Ai <31551580+Fangzhou-Ai@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Contributor
Author
|
/ci run |
|
✅ Triggered Buildkite CI #91744 for commit |
shen-shanshan
approved these changes
Sep 29, 2026
shen-shanshan
left a comment
Contributor
There was a problem hiding this comment.
Reviewed by @Fangzhou-Ai
linxiLIFE
pushed a commit
to linxiLIFE/vllm
that referenced
this pull request
Sep 29, 2026
vllm-project#58671) Signed-off-by: Mehmet Cagri Kaymak <mehmet.kaymak@amd.com> Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com> Co-authored-by: Shanshan Shen <467638484@qq.com>
Fangzhou-Ai
added a commit
to Fangzhou-Ai/recipes
that referenced
this pull request
Sep 30, 2026
…parse indexer on MI355X Add to the MI355X TP override: - VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=1 for vllm-project/vllm#53492's gfx950-only Gluon sparse-MLA kernel. - --attention-config '{"indexer_kv_dtype":"mxfp4","indexer_sparse_logits":true}' for vllm-project/vllm#58671's ROCm paged MXFP4 sparse indexer, replacing the dense fp8 indexer path. - --block-size 128, which the fp4 indexer and SparseMLA backend prefer once active (auto-resolution otherwise falls back to 64). Both vLLM PRs are merged and gfx950-scoped. Mirrors SemiAnalysisAI/InferenceX#3555 and SemiAnalysisAI/InferenceX#3571. Signed-off-by: fai <fangzhouai@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Fangzhou-Ai
added a commit
to SemiAnalysisAI/InferenceX
that referenced
this pull request
Sep 30, 2026
Stacked on #3555. Re-adds the --attention-config flag enabling #58671's ROCm paged MXFP4 sparse-logits indexer and the --block-size 128 workaround its active fp4 indexer needs (vLLM's block-size auto-resolution otherwise picks 64 across this model's 4+ attention backends). Draft until #58671 merges upstream and lands in a nightly image, so #3555 can run now and this PR can re-sweep immediately once it's available. Signed-off-by: Fangzhou Ai <31551580+Fangzhou-Ai@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Fangzhou-Ai
added a commit
to SemiAnalysisAI/InferenceX
that referenced
this pull request
Sep 30, 2026
Stacked on #3555. Re-adds the --attention-config flag enabling #58671's ROCm paged MXFP4 sparse-logits indexer and the --block-size 128 workaround its active fp4 indexer needs (vLLM's block-size auto-resolution otherwise picks 64 across this model's 4+ attention backends). Draft until #58671 merges upstream and lands in a nightly image, so #3555 can run now and this PR can re-sweep immediately once it's available. Signed-off-by: Fangzhou Ai <31551580+Fangzhou-Ai@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com>
1 task
Fangzhou-Ai
added a commit
to SemiAnalysisAI/InferenceX
that referenced
this pull request
Oct 1, 2026
…cludes #58671) Enables vllm-project/vllm#58671's ROCm paged MXFP4 sparse-logits indexer via --attention-config and --block-size 128, replacing the dense fp8 indexer path. A live A/B test (TP2 c16, matched 900s window) measured #58671 giving +14.7/+14.9% p50/p90 interactivity and -9.5/-10.8% p50/p90 e2e latency over the dense path, with throughput/GPU unchanged. Also drops c128 from both TP2 and TP4 conc-lists for this recipe. Co-authored-by: Cursor <cursoragent@cursor.com>
Fangzhou-Ai
added a commit
to SemiAnalysisAI/InferenceX
that referenced
this pull request
Oct 1, 2026
…cludes #58671) Enables vllm-project/vllm#58671's ROCm paged MXFP4 sparse-logits indexer via --attention-config and --block-size 128, replacing the dense fp8 indexer path. A live A/B test (TP2 c16, matched 900s window) measured #58671 giving +14.7/+14.9% p50/p90 interactivity and -9.5/-10.8% p50/p90 e2e latency over the dense path, with throughput/GPU unchanged. Also drops c128 from both TP2 and TP4 conc-lists for this recipe. Co-authored-by: Cursor <cursoragent@cursor.com>
chunfangamd
added a commit
to SemiAnalysisAI/InferenceX
that referenced
this pull request
Oct 2, 2026
…er / DSv4.1-Flash vLLM:启用 ROCm 分页 MXFP4 稀疏 indexer (#3571) * Repin DSv4.1-Flash MI355X vLLM AgentX to nightly-rocm100-ac9126e5 (includes #58671) Enables vllm-project/vllm#58671's ROCm paged MXFP4 sparse-logits indexer via --attention-config and --block-size 128, replacing the dense fp8 indexer path. A live A/B test (TP2 c16, matched 900s window) measured #58671 giving +14.7/+14.9% p50/p90 interactivity and -9.5/-10.8% p50/p90 e2e latency over the dense path, with throughput/GPU unchanged. Also drops c128 from both TP2 and TP4 conc-lists for this recipe. Co-authored-by: Cursor <cursoragent@cursor.com> * Repin DSv4.1-Flash MI355X vLLM AgentX to nightly-rocm100-ac9126e5 (includes #58671) Enables vllm-project/vllm#58671's ROCm paged MXFP4 sparse-logits indexer via --attention-config and --block-size 128, replacing the dense fp8 indexer path. A live A/B test (TP2 c16, matched 900s window) measured #58671 giving +14.7/+14.9% p50/p90 interactivity and -9.5/-10.8% p50/p90 e2e latency over the dense path, with throughput/GPU unchanged. Also drops c128 from both TP2 and TP4 conc-lists for this recipe. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(dsv41flash): align MI355X vLLM changelog and docs with the recipe - perf-changelog: make the new entry English-only, as AGENTS.md now requires; record VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True, which this PR also sets; state why c128 is dropped; spell out the vLLM PR references. - amd-master.yaml and configuration-procedures (EN and ZH): replace the remaining c128 statements with the current c1-c64 range. No recipe or master-config value changes. 中文: - perf-changelog:按 AGENTS.md 的最新要求将新条目改为仅英文;记录本 PR 同时设置的 VLLM_ROCM_USE_AITER_TRITON_SPARSE_MLA=True;说明移除 c128 的原因;写全 vLLM PR 引用。 - amd-master.yaml 与 configuration-procedures(中英文):将残留的 c128 描述更新为当前的 c1-c64 范围。 配方与主配置的取值均无变化。 Co-authored-by: Cursor <cursoragent@cursor.com> * docs(dsv41flash): update MI355X GPU validation for the ac9126e5 pin Address review feedback: the GPU-validation paragraph still named the superseded nightly-rocm100-29468dde pin. Point it at nightly-rocm100-ac9126e5 and run 36824408313 (TP4 and TP2, concurrency 1-64; eval-only TP4 c64 GSM8K 0.9712 strict / 0.9704 flexible), list the earlier sweeps as superseded, and cite the merged upstream recipe vllm-project/recipes#1049. The English and Chinese pages change together. 中文:根据评审意见,GPU 验证段落仍写着已被取代的 nightly-rocm100-29468dde。 现改为 nightly-rocm100-ac9126e5 与运行 36824408313(TP4 与 TP2,并发 1-64; 仅评测的 TP4 c64 GSM8K 为 strict 0.9712 / flexible 0.9704),将此前的 sweep 标注为已被取代,并引用已合并的上游配方 vllm-project/recipes#1049。中英文页面 同步更新。 Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: Chun Fang <chun.fang@amd.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds a ROCm (gfx950) path for DeepSeek V4.1's indexer with an MXFP4 K cache, built on aiter's paged MXFP4 MQA-logits Gluon kernel. The CUDA path is unchanged. There are two modes:
indexer_kv_dtype="mxfp4"): dense logits over the whole context, read straight from the paged K cache.indexer_sparse_logits=True): the candidate source takes block maxima from the same pass that writes its logits, and the consumer layers score only the candidate pool instead of the whole context. Steps whose contexts are too short for the pool to pay off keep the regular pass, so it helps mainly at long contexts.Both modes:
k_norm_rope_mxfp4_cache, and Q is quantized byq_rope_mxfp4_quant.query_start_loc. Uniform spec-decode steps run asnext_n-row sequences, with the decode schedule built once per step.In sparse mode, the candidate pool is resolved once per step and shared by every consumer layer.
Requires aiter with the paged MXFP4 MQA-logits kernel and the cache ops (merged to aiter main) (ROCm/aiter#5834)
Tests
tests/kernels/attention/test_rocm_paged_mxfp4_indexer.pytests/v1/attention/test_rocm_paged_mxfp4_indexer_plan.pyPerformance
Benchmark: C=32, 128 prompts of 350K tokens, 350 out, spec. decoding = 5+1
Note, since this is done on artificial data, AL is lower (1.65) compared to more realistic 3.5.
8k/1k, NP=10*C, TP4, using mxfp4 dense path (too short kv for sparse), no spec decoding:
Accuracy
Long context accuracy check
/v1/completionswith the ids as the prompt,max_tokens=1,temperature=0,prompt_logprobs=0. For every position the server returns the log-probability the model gave to the token that actually comes next, given everything before it. Nothing is sampled; it is one teacher-forced pass over the whole book.Not a duplicate
#57517 and #56834 also add a ROCm MXFP4 indexer K cache, with dense scoring. The main
addition here is sparse MXFP4 (
indexer_sparse_logits=True): the candidate source takesblock maxima from the same pass that writes its logits, and the consumer layers score only
the candidate pool, read in place from the paged MXFP4 cache. Neither PR supports this
(#56834 rejects candidate blocks outright). For the dense path, this PR uses aiter's paged
MXFP4 MQA-logits kernel (ROCm/aiter#5834, in aiter v0.1.23) and leaves vLLM's shared Q/K
kernels untouched.
Follows the main idea from #56254 while going direct to the paged K for prefill.
AI assistance
Claude is used during development, and the PR author reviewed the AI driven changes.