Skip to content

[Bugfix][Attention] Zero the null KV block after CUDA graph capture - #57158

Closed
sumsliu wants to merge 2 commits into
vllm-project:mainfrom
sumsliu:fix/zero-null-kv-block-after-capture
Closed

sumsliu wants to merge 2 commits into
vllm-project:mainfrom
sumsliu:fix/zero-null-kv-block-after-capture

Conversation

@sumsliu

@sumsliu sumsliu commented Sep 16, 2026 •

Copy link
Copy Markdown

Purpose

Fixes #57156.

CUDA graph capture runs dummy batches whose block table is all zeros, so capture can leave
uninitialized values in block 0 of the KV cache. Attention kernels that deliberately gather
block 0 for masked or padded entries rely on the softmax mask to discard the result. That only
holds while the gathered data is finite: a non-finite value survives masking as 0 * NaN == NaN
in the value accumulation and poisons every real token in the batch.

Block 0 is the reserved null block (BlockPool.null_block, popped at init and never handed to a
request), so restoring it to zeros once after capture holds for the lifetime of the process.

Root cause

Observed on DeepSeek-V4.1-Flash with the FlashInfer SM120 sparse-MLA backend. In
sparse_mla_sm120/decode_dsv4_kernel.cuh, the scale load guards on the index:

// line 272
if (idx_raw >= 0) { scale_word = __ldg(...); }

while the NoPE/RoPE data load clamps instead:

// line 299
const int idx = (idx_raw >= 0) ? idx_raw : 0;

The intent is documented at line 431: "IO already gathered slot 0 into smem with idx clamped —
masking to -inf kills it in softmax"
. Masking sets the logit to -1e30 so softmax yields
p ≈ 0, but the PV accumulation still computes p * v, and in MLA the NoPE tensor is both K
and V. So non-finite data in slot 0 reaches acc_nope regardless of the mask.

Measured contents of block 0 immediately before the first real request, same build:

Mode slot 0 NoPE slot 0 RoPE
eager 0.0 0.0
PIECEWISE NaN ~3.2e35

Eager never captures, so block 0 stays zeroed and the bug does not appear.

This is not specific to one backend: any kernel that gathers a sentinel block and relies on
masking has the same precondition. #57094 fixes a sibling case in Triton MLA decode where CUDA
graph padding produces 0/0.

Test Plan

Verified on 8x DGX Spark (GB10, sm_121), TP8 + EP8, DeepSeek-V4.1-Flash, with the per-call
workaround disabled so this change is the only thing preventing the NaN:

  • single stream, 1024-token generation, three prompts
  • concurrency 1 / 16 / 64, 1024-token input, 128-token output

Test Result

All completions correct, no NaN anywhere in the outputs or logs.

Check Before After
"The capital of France is" "" " Paris."
Single stream NaN 69 / 40 / 81 tok/s
Concurrency 64 NaN 244 tok/s

Speeds match the per-call workaround, confirming the once-after-capture placement is sufficient
and cheaper. Single-stream throughput with CUDA graphs enabled is 32.6 tok/s against 17.2 tok/s
under --enforce-eager.

Notes

I verified that block 0 holds non-finite data after capture and that zeroing it there fixes the
output, but I did not trace which specific write produces it — get_dummy_slot_mappings fills
with PAD_SLOT_ID and the fused SWA insert kernel does skip slot_id < 0, so the polluting
write is elsewhere. Restoring the invariant after capture is correct regardless, but if a
maintainer would rather fix the write itself I am happy to rework this.


AI assistance disclosure

Per AGENTS.md, this section covers
the required disclosures for AI-assisted contributions.

AI assistance was used. The investigation, root-cause analysis and patch were produced with
Claude Code. The submitter reviewed every changed line and reproduced the failure and the fix on
the hardware described below.

Not a duplicate. Searched before opening:

gh api "search/issues?q=repo:vllm-project/vllm+DeepSeekV41+NaN"
gh api "search/issues?q=repo:vllm-project/vllm+null_block+NaN"
gh api "search/issues?q=repo:vllm-project/vllm+flashinfer+sparse+mla+NaN"

The closest existing work is #57094, which fixes CUDA-graph padding NaN in Triton MLA decode
where stage 2 computes 0/0 for seq_len=0. That is a different backend and a different
mechanism; this PR addresses a clamped index reading a poisoned null block, where the value
accumulation computes 0 * NaN. #55636 is a sibling warmup-metadata bug on the same hardware but
surfaces as an illegal address rather than NaN. #57028 touches the same SM120 path but covers page
geometry, not the null block; the NaN reproduces independently of it.

Test commands and results. On 8x DGX Spark (GB10, sm_121), TP8 + EP8, DeepSeek-V4.1-Flash,
CUDA graphs enabled, with no other workaround in place so this change is the only thing preventing
the NaN:

# correctness
curl -s "$API/v1/completions" -d '{"model":"deepseek-v4.1-flash","prompt":"The capital of France is","max_tokens":4,"temperature":0}'
# before: ""        after: " Paris."

# single stream, 1024-token generation, three prompts
python3 bench.py single --rounds 3

# concurrency sweep, 1024-token input, 128-token output
python3 bench.py batch --concurrency 1,16,64 --in-len 1024 --out-len 128 --seed-offset 42

Model evaluation. This change affects serving output, so results are included rather than
waiting to be asked:

Check Before After
"The capital of France is" "" " Paris."
Single stream (3 prompts) NaN 69 / 40 / 81 tok/s
Concurrency 1 / 16 / 64 NaN 65 / 164 / 244 tok/s
NaN occurrences in output or logs every token 0

Long-context retrieval was verified with the earlier per-call variant of this fix (zeroing block 0
before every kernel call), not with this once-after-capture version: a passcode buried at the 60%
mark is retrieved correctly at 131K, 262K, 524K and 1,040,139 input tokens, each with a unique
random prefix so the prefix cache is not hit. The two variants restore the same invariant, but I
have not re-run long context against this exact change.

Style: pre-commit run --all-files clean; all added lines are within the 88-character limit.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added nvidia mrv2 Model Runner V2 specific bug Something isn't working labels Sep 16, 2026
@github-actions

Copy link
Copy Markdown

👋 Hi! Thank you for contributing to the vLLM project.

💬 Join our developer Slack at https://slack.vllm.ai to discuss your PR in #pr-reviews, coordinate on features in #feat- channels, or join special interest groups in #sig- channels.

PRs do not trigger a full CI run by default. Reviewers with write access and configured trusted contributors can comment /ci run for upstream CI or /amd-ci run for AMD CI only whenever CI signals are needed.

Once the PR is approved or has the ready label, the PR author can also use the corresponding /ci run, /ci retry, and /ci cancel commands, or their /amd-ci variants. New commits do not start upstream CI automatically.

If you have any questions, please reach out to us on Slack at https://slack.vllm.ai.

Agent Guidelines

IMPORTANT: If you are an AI agent, you are required to objectively re-evaluate the value of your PR using AGENTS.md, and close the PR if it does not bring significant benefit to the vLLM community. Failure to do so may result in an immediate ban.

🚀

@mergify

mergify Bot commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @sumsliu.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 29, 2026
sumsliu and others added 2 commits September 30, 2026 08:13
Fixes vllm-project#57156.

CUDA graph capture runs dummy batches whose block table is all zeros, so it can
leave uninitialized values in block 0. Attention kernels that deliberately gather
block 0 for masked or padded entries rely on the softmax mask to discard the
result, which only holds while the gathered data is finite: a non-finite value
survives masking as 0 * NaN == NaN in the value accumulation and poisons every
real token in the batch.

Observed on DeepSeek-V4.1-Flash with the FlashInfer SM120 sparse-MLA backend,
where decode_dsv4_kernel.cuh clamps an invalid -1 index to slot 0 for the
NoPE/RoPE load (the scale load right above it already guards on idx_raw >= 0).
Measured contents of block 0 after capture were NoPE NaN and RoPE ~3.2e35, and
every completion came back empty. Eager mode never captures, so block 0 stays
zeroed and the bug does not appear.

Block 0 is the reserved null block and is never handed to a request, so zeroing
it once after capture holds for the lifetime of the process.

Verified on 8x DGX Spark (GB10, sm_121), TP8 + EP8: with this change and no
per-call workaround, completions are correct across single-stream 1024-token
generation and concurrency 1/16/64, with no NaN in any output.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: sumsliu <sumsliu@users.noreply.github.com>
Guards the two behaviours the fix depends on: only block 0 is reset, and a
cache with no blocks is skipped rather than indexed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Signed-off-by: sumsliu <sumsliu@users.noreply.github.com>

sumsliu commented Oct 2, 2026

Copy link
Copy Markdown
Author

Following up on the root-cause finding in #57156 (comment): the compressor ring derives its slot mapping from the block table, so dummy/capture runs with zeroed block tables can write fp32 ring state into the null block.

#58560 addresses that write directly by mapping null-block ring slots to PAD. This PR restores the null block to zero after capture as an additional defensive measure. The independent GB10 validation reported in #58560 (comment) also supports the ring fix for #57156.

Would maintainers prefer to retain the post-capture zeroing as extra protection, or focus validation on #58560? If you'd like to keep this PR, could a maintainer enable upstream CI?

sumsliu commented Oct 5, 2026

Copy link
Copy Markdown
Author

Closing in favor of #58560 (49e0f47). It prevents dummy/capture compressor-ring writes to block 0 and extends the metadata test to cover real requests, null-block requests, and padding.

There is no separate reproducer in this PR that warrants retaining the post-capture clearing. Current main has not been rerun on the 8×GB10 setup as part of this review; that end-to-end check remains open in #57156.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working mrv2 Model Runner V2 specific nvidia

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

[Bug]: DeepSeek-V4.1 produces NaN with CUDA graphs on SM120/121 — dummy capture batches poison the null KV block

1 participant