Skip to content

[Bugfix][Attention] Reserve the sparse indexer's DCP top-k merge scratch through the workspace manager - #59367

Open
drakosha wants to merge 2 commits into
vllm-project:mainfrom
drakosha:fix-dcp-topk-merge-profiling
Open

drakosha wants to merge 2 commits into
vllm-project:mainfrom
drakosha:fix-dcp-topk-merge-profiling

Conversation

@drakosha

@drakosha drakosha commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Fixes #59317.

With DCP > 1 the sparse indexer merges each rank's local top-k into the global top-k in _merge_dcp_topk_global. The merge allocated three transients from the plain allocator per call: packed (T, K, 2) fp32, the all-gather output (dcp * T, K, 2) and the copy the communicator's all_gather makes to move the rank dimension inside. Peak (1 + 2 * dcp) * 8 * T * K bytes per rank, 1.13 GiB for T = 8192, K = 2048, dcp = 4, on top of the logits. The profile run takes the indexer's profiling branch, which never calls the merge, so the KV cache was sized as if the merge cost nothing and the engine found out on the first long prefill (torch.OutOfMemoryError ... in _merge_dcp_topk_global -> all_gather, 512 MiB is the gather output for those numbers).

This PR takes both routes proposed in the issue:

  • the merge walks the rows DCP_TOPK_MERGE_ROWS (1024) at a time, so the scratch is bounded at (1 + 2 * dcp) * 8 * 1024 * K bytes (144 MiB for dcp 4, K 2048) whatever the chunk size; the pack kernel takes row_starts per row and the selector is row-independent, so the result is unchanged;
  • the scratch comes from the workspace manager (_dcp_topk_merge_specs): the prefill path takes it in the same get_simultaneous call as the gathered K, which stays live across chunks and must not be overlapped; the decode path takes its own. The all-gather goes in place into the reserved rank-major buffer on the group's communicator (pynccl when enabled, dist.all_gather_into_tensor otherwise), and the rank dimension is moved with one strided copy_ into the row-major buffer the CuteDSL selector reads;
  • the indexer's profiling branch adds the same specs to profile_specs when dcp_world_size > 1, so the profile run reserves the scratch and the KV budget sees it.

Related: #59318 is the CUDA graph memory over-estimate that happened to cover this transient on our host (about 1 GiB per rank); its fix is #59368. Either fix alone changes the failure boundary, so they were validated together.

Not a duplicate: #47348 optimizes the CuteDSL merge kernels themselves and leaves the buffers as they are; #59211 is DCP for the kpool indexer, a different code path; #55132 is about the indexer reserving too much for the decode logits on ROCm, this is about a buffer the profiling run did not reserve at all.

Test Plan

CPU unit tests (no GPU), on vllm-openai:nightly af7f948 with the PR's sparse_attn_indexer.py mounted over the package:

python3 -m pytest -o addopts= tests/v1/attention/test_indexer_dcp_localize.py -q -k "merge_specs or walks_rows"

test_dcp_topk_merge_specs_bound_the_scratch checks the reservation shapes and the (1 + 2 * dcp) * 8 * rows * K bound; test_dcp_topk_merge_walks_rows_through_the_workspace runs the merge with faked kernels and all-gather over 2 * DCP_TOPK_MERGE_ROWS + 5 rows and checks the row slices, that every buffer is a contiguous view of the given workspace, the rank-major to row-major layout the selector receives, and the output. The existing GPU harness _merge_local_topks_global_with_fake_dcp (CUDA + CuteDSL) is updated to fake _dcp_all_gather_into chunk by chunk; those tests need a free GPU and were not run in this round.

Serving: GLM-5.3-NVFP4, 4x H200 NVL, TP4 DCP4 EP, MTP 3, fp8_ds_mla, max-model-len 786432, max-num-batched-tokens 8192, max-num-seqs 32, gpu-memory-utilization 0.945, on nightly af7f948 with this PR and the fix for #59318.

Test Result

Unit tests: 2 failed on the base image (no DCP_TOPK_MERGE_ROWS), 2 passed with the PR.

Serving, both fixes: the CUDA graph estimate is exact (0.91 GiB pool estimated and captured) and the KV cache is 13.25 GiB / 1,039,616 tokens per rank (13.32 GiB / 1,045,248 before, with the 1 GiB over-estimate that used to cover the merge). Fresh prefill 8x32768: 65.7 s, 5092 tok/s (5129 tok/s on the unpatched production image the same night); 2x131072: 69.7 s, 4941 tok/s (4988 before). Decode N=1 110.7, N=8 329.8, N=32 761.1 tok/s. The decode figures are single samples; across seven boots of the unpatched image that night N=8 ranged 263 to 317 tok/s and N=32 701 to 838 tok/s, so no decode change is claimed either way. Stress 3x600k: cold 3/3 in 373.8 s (370.8 s before), replay from CPU offload 3/3. Needle 200k at depth 0.1: retrieved, so the chunked merge selects the same tokens. The needle ran concurrently with the stress, so the replay wall time (47.3 s against 2.4 s before) is contention with the needle prefill, not a cost of the fix. GPU memory in use while serving: 141.0 GiB of 143.77 per GPU, against 142.4 GiB before (the merge transient no longer lives in the free memory).

Second boot, VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=0 with this fix only, the combination that OOMed on the first 8x32k prefill before: KV cache 15.34 GiB / 1,203,712 tokens, fresh prefill 8x32768 in 63.0 s (5103 tok/s), 2x131072 in 72.6 s, decode N=1 119.1, N=8 279.2, N=32 823.9 tok/s, no errors. 0.8 GiB free per GPU while serving, so that is the proof that the merge no longer needs the slack, not a recommended setting.

AI assistance

AI assistance (Claude Code) was used for the analysis, the patch, the tests and this description.

…tch through the workspace manager

_merge_dcp_topk_global() allocated (1 + 2 * dcp) * 8 * T * K bytes per
call from the plain allocator: the packed candidates, the all-gather
output and the copy that moves the rank dimension inside. The profile run
never calls the merge, so the KV cache was sized as if it cost nothing,
and with T = 8192, K = 2048, dcp = 4 the first long prefill needed
1.13 GiB per rank that was not there (OOM in the all-gather).

Merge DCP_TOPK_MERGE_ROWS rows at a time through buffers taken from the
workspace manager, gather in place on the group's communicator, and add
the buffers to the indexer's profiling specs so the profile run reserves
them. The scratch is bounded at 144 MiB per rank for the numbers above,
shares the prefill workspace allocation with the gathered K it must not
overlap, and the KV budget sees it.

Fixes vllm-project#59317

Co-Authored-By: Claude <noreply@anthropic.com>

Signed-off-by: Mikhail Kostryukov <mike@triptrack.net>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify

mergify Bot commented Oct 11, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @drakosha.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Oct 11, 2026
Resolves the conflict with vllm-project#54951 in sparse_attn_indexer.py: keep both the
merge workspace split and the TP row shard, and pass the narrowed
cu_seqlen_ks as row_starts. Row sharding requires dcp_world_size == 1, so
it never runs together with the DCP top-k merge.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Mikhail Kostryukov <mike@triptrack.net>
@mergify mergify Bot removed the needs-rebase label Oct 11, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Sparse-indexer DCP top-k merge allocates (1 + 2 * dcp) * 8 * T * K bytes per step that the profile run never sees, OOM on the first long prefill

1 participant