Repository navigation
[Bugfix][Attention] Fixed-width sparse indexer prefill logits so long prefills reuse allocator blocks (#55569) - #58068
Conversation
|
Review for PR #58068: Fixed-width sparse indexer prefill logits so long prefills reuse allocator blocks Memory Optimization Analysis
|
|
👋 Hi! Thank you for contributing to the vLLM project. 💬 Join our developer Slack at https://slack.vllm.ai to discuss your PR in PRs do not trigger a full CI run by default. Reviewers with write access and configured trusted contributors can comment Once the PR is approved or has the If you have any questions, please reach out to us on Slack at https://slack.vllm.ai. Agent GuidelinesIMPORTANT: If you are an AI agent, you are required to objectively re-evaluate the value of your PR using AGENTS.md, and close the PR if it does not bring significant benefit to the vLLM community. Failure to do so may result in an immediate ban. 🚀 |
|
This pull request has merge conflicts that must be resolved before it can be |
… prefills reuse allocator blocks Allocate the sparse MLA indexer prefill logits at ceil(max_model_len / compress_ratio) columns via DeepGEMM's max_seqlen_k so every sub-chunk of a long prefill requests the same size and the caching allocator reuses it. Default on for integrated GPUs; VLLM_SPARSE_INDEXER_FIXED_LOGITS_WIDTH forces on/off. Fixes vllm-project#55569. Signed-off-by: Incarnas <119618389+bit-incarnas@users.noreply.github.com> Co-authored-by: Claude <noreply@anthropic.com>
174d31b to
1631c22
Compare
|
Rebased onto current |
Purpose
Fixes #55569. Related: #56457 reports the same allocation pattern in the Qwen4Exp QSA indexer, with fixes proposed for that path in #56500 and #57105 (both open); #55572 (closed) lowered the default logits budget on integrated GPUs instead.
The dense sparse-MLA indexer (
sparse_attn_indexer.py, and the GLM-5.3-Flash kpool variant undermodels/glm5next) scores each prefill chunk with DeepGEMM'sfp8_fp4_mqa_logits, which allocates its[M, N]fp32 logits inside the C++ API.Nis the request's compressed context so far, so every chunk of a long prefill asks for a tensor of a new, larger size: one launch per chunk of growing size whileM * N * 4fits theVLLM_SPARSE_INDEXER_MAX_LOGITS_MBbudget (512 MiB), and, once a chunk exceeds it (about 116k tokens at a 4,608-token chunk with kpool 4), several launches per chunk that are all close to the budget but never the same size twice. PyTorch's caching allocator serves a request from a cached block only if it fits, and returns segments to the driver only onempty_cacheor an OOM-time release, so once the free space that happens to be around is used up, each chunk step adds a new segment for the rest of the prefill. On a discrete GPU this is reserved-but-idle memory; on unified-memory devices (DGX Spark / GB10) it comes out of the host's share, and the driver-level failure (NV_ERR_NO_MEMORY) arrives before the allocator's OOM retry can help, ending in the swap storms reported in #55569.Change. Allocate the logits at a fixed width instead of the chunk's kv length, using the
max_seqlen_kargument DeepGEMM already provides: the kernel still computes only[cu_seqlen_ks, cu_seqlen_ke)per row,max_seqlen_kis only the allocation stride, and DeepGEMM assertsclean_logits=Falsewith it (verified at vLLM's pinned commit). The width isceil(max_model_len / compress_ratio)(the page-table width), computed once in the metadata builder, and the planner sizes its query sub-chunks against it, so every sub-chunk of a prefill requests an identically sized tensor from the first chunk to the last and the allocator recycles the same blocks. At the call sites the returned view is narrowed back to[M, N], so every consumer sees exactly the shape it sees today; only the storage behind it is constant-size. This is safe for all of them because they take strides explicitly:top_k_per_row_prefillreceiveslogits.stride(0)andstride(1), and both candidate-block Triton kernels unpack*logits.stride(). The DeepSeek V4.1 sparse-logits backend already keeps a fixedmax_logits_bytesworkspace; this brings the dense DeepGEMM path in line with it.vllm/envs.py:VLLM_SPARSE_INDEXER_FIXED_LOGITS_WIDTH. Unset (default): on for integrated (unified-memory) GPUs only;1/0force it on / off on any GPU.vllm/v1/attention/backends/mla/indexer.py:DeepseekV32IndexerPrefillChunkMetadata.logits_width;_split_indexer_prefill_chunksbudgets and sub-chunks againstmax(N, logits_width); the builder computes the width once (self.fixed_logits_width) for the standard path (dcp_world_size == 1). The PCP/DCP planner path is left at the per-chunk width because the DCP merge consumes the logits directly.vllm/utils/deep_gemm.py:fp8_fp4_mqa_logits(..., max_seqlen_k: int = 0)forwards the kwarg (0 = current behaviour).vllm/model_executor/layers/sparse_attn_indexer.py,vllm/models/glm5next/nvidia/sparse_indexer.py: pass the width, narrow the view.tests/v1/attention/test_indexer_prefill_chunk_logits_width.py: planner test (constant launch shapes at every step;logits_width=0is identical to the current plan; width never below the kv length; packed requests are budgeted at the fixed width).Costs and non-goals. Total kernel work is unchanged (same rows, same
[ks, ke)bounds), but the launch count is not: with a 512 MiB budget and a 102,400-column width (max_model_len409,600, kpool 4) a 4,608-token chunk is 4 launches from the first chunk on, where today it is 1 launch until ~116k tokens; and packed short requests stop sharing a launch oncerows * width * 4exceeds the budget (the test shows two 1,000-row requests at kv length 100 becoming two launches), so a batch of many short prefills becomes more launches. Neitherwas measured for latency at short prompts, which is why the default is integrated-only, where the allocator growth is what fails. The decode path (paged kernel) is untouched. This change does not address the separate first-use growth of the activation working set on unified memory (see the note under Test Result), which is not the indexer.
This PR was prepared with AI assistance; the measurements were run on my hardware under my direction, and I reviewed every changed line.
Update (2026-10-08): rebased onto
mainRebased onto
main@f4dde3132. One conflict: #59211 (DCP for the kpool indexer) added an empty-local-chunk guard around thefp8_fp4_mqa_logitscall inmodels/glm5next/nvidia/sparse_indexer.py. The width and the narrowed view now sit inside its kernel branch, so a rank with no local kv still skips the kernel exactly as onmain. The fixed width stays off wheneverdcp_world_size > 1, as before. Everything else applied unchanged.Checked against what moved on
mainsince this was opened:logits_widthis in the same units, and [Bugfix] GLM-5.3-Flash: fp8 plan dtype on SM90 sparse MLA, and right-size the indexer prefill workspace #55222'stest_indexer_prefill_budget_matches_compressed_workspacepasses on this branch.e1f418c2to1e1842a8.fp8_fp4_mqa_logitshas an identical function body at both (csrc/apis/attention.hpp, including themax_seqlen_kallocation path),max_seqlen_kis still a keyword of the Python binding vLLM calls (default 0), and the new torch op schema carries it too.Re-run on the rebased head (CPU-side tests; Python 3.12, torch 2.13.0):
pytest tests/v1/attention/test_indexer_prefill_chunk_logits_width.py: 5 passed. Againstmain's planner, all 5 fail (nologits_widthparameter).pytest tests/v1/attention/test_indexer_deepseek_v4_slot_mapping.py: 32 passed, 1 failed. The failure is DeepGEMM's architecture assert on the local GPU, which DeepGEMM does not support, and it fails identically onmain. The test helper's default model was swapped for an ungated mirror of Llama-3-8B for this run, because the gated repo is not available here.The rebased head has not been served. The end-to-end numbers below are still from the build described under Limitations.
Test Plan
pytest tests/v1/attention/test_indexer_prefill_chunk_logits_width.pynvidia/GLM-5.3-Flash-NVFP4, expert parallel, MTP 3,--kv-cache-dtype fp8_ds_mla,--max-num-batched-tokens 8192(scheduler chunk 4,608 = the prefix-cache block), one client at a time): a fresh serve, then cold prefills of ~100k tokens, another ~100k (different content, same length), then ~200k, sampling hostMemAvailableon both nodes after each request. Profile--max-model-len 204800 --gpu-memory-utilization 0.87(memory profiling on), chosen for host headroom; fixed width 51,200 columns. The stock width was run on the same serve configuration with the fixed width disabled, so the allocation width is the only difference.--max-model-len 409600 --kv-cache-memory 4.62e9; fixed width 102,400 columns) with the change applied.torch.cuda.memory_reserved(),num_device_allocand hostMemAvailableafter every indexer call (one per indexer layer per chunk step) on both ranks, to attribute growth to the caching allocator.fp8_fp4_mqa_logits+top_k_per_row_prefillfor a 98k-token prefill (12 × 8,192-token chunks × 45 layers) with and without the fixed width, no model. Available on request.Test Result
Per-rank caching-allocator growth over the whole A/B ladder (fresh serve → 100k → 100k → 200k), identical on both ranks:
Host
MemAvailabledeltas per request (rank 0 / rank 1, MB):NV_ERR_NO_MEMORYkernel lines)Time to first token: 100k 53.9 / 52.6 s stock vs 54.1 / 52.9 s fixed; 200k 116.5 s stock vs 109.6 s fixed (single samples; we read the 100k numbers as parity and do not claim the 200k difference).
Issue scenario with the change, production profile: a 231,751-token cold prefill completed (TTFT 127 s, correct answer) with no
NV_ERR_NO_MEMORYlines, leaving 1.7 GB / 3.1 GB of hostMemAvailableon the two nodes from an idle 5.5 / 6.9 GB; the difference is the first-use working-set growth described below plus the two logits blocks, not per-chunk growth. The stock width was not re-run at 230k on that profile; its per-step growth is the A/B row above.In isolation (reserved bytes after 98k tokens): stock +2,810 MB, one new segment per chunk until the sub-chunk sizes stabilise; fixed +1,190 MB, flat from the second chunk (two live logits blocks — the loop keeps chunk k's
logitsreferenced while chunk k+1 allocates — plus a tail block). With a fixed pattern of interleaved transient allocations: stock +3,650 MB, fixed +1,960 MB (transients alone +1,410). For completeness,PYTORCH_CUDA_ALLOC_CONF=max_split_size_mb:256and a privatetorch.cuda.MemPoolfor the logits made the fixed case worse (+2,312 / +2,470 MB) because they stop the freed logits blocks from serving other allocations. #55569 listsexpandable_segments:Trueas a mitigation; in our run on the stock width it did not hold (fiveNV_ERR_NO_MEMORYlines during a 100k prefill, −1.5 GB retained), consistent with the NemotronH data point later in that thread.Note for anyone measuring this on unified memory: the first deep prefill after a fresh start also grows the process by the activation working set of each new chunk shape (about 5 GiB per device here up to the 8,192-token chunk, as
gpu_memory_utilizationprofiling reports). That growth is independent of the indexer and of this change; it dominates the first-100k row in both columns above, and only the second-100k and 200k rows isolate the indexer. Our own first measurement of this patch was confounded by exactly that.Limitations of the evidence: the end-to-end numbers were measured on vLLM 0.28.1rc1.dev580 (the NVIDIA arm64 CUDA 13 build for this model) with an equivalent patch; the port to
mainapplies and its planner test passes, butmainitself was not served. One model family (GLM-5.3-Flash, kpool path) and one hardware class, TP=2. The standardsparse_attn_indexer.pypath shares the planner and the DeepGEMM call, and the planner test covers its arithmetic, but the DeepSeek V3.2 / V4 variants were not run end to end. DCP is intentionally left at the current width. Short-prompt latency and many-short-request batches with the higher launch count were not measured.Pull Request Checklist
I used vLLM's
/pr-checklistskill. (Mandatory for agents, optional for humans).AI assistance was used during the creation of this PR.
Design Fit: Minimizes impact on core components, reuses existing functionality, and justifies added complexity.
Testing and Validation: Validates the change and ensures any added tests are meaningful and reliable, with CI coverage or documented CI resource constraints and validation performed outside CI.
Code Quality and Style: Keeps code and comments clear and concise, and updates relevant documentation and examples.
Pull Request Contents: Includes a brief summary and relevant links, supports claims with evidence, explains root causes and implementation trade-offs, and follows the contributing guide.