Skip to content

[BugFix][Model Runner V2][Spec Decode] Fix decode instance's multi-layer MTP kv caches during P/D - #55055

Merged
WoosukKwon merged 1 commit into
vllm-project:mainfrom
TheEpicDolphin:bugfix/mrv2-multi-module-mtp-pd
Sep 16, 2026
Merged

WoosukKwon merged 1 commit into
vllm-project:mainfrom
TheEpicDolphin:bugfix/mrv2-multi-module-mtp-pd

Conversation

@TheEpicDolphin

@TheEpicDolphin TheEpicDolphin commented Sep 3, 2026 •

Copy link
Copy Markdown
Member

Context

During multi-module MTP with P/D, the prefill instance runs prefill on the full prompt and the last N MTP layers each generate a draft token. This is problematic on the last prefill chunk, because MTP layer i will embed its draft token into MTP layer i + 1. Those draft tokens are unverified, and remain in the KV cache for those last N-1 MTP layers. See the visualization below:
image

During standalone serving, the next decode step would receive the number of rejections, and would re-prefill those stale cache slots accordingly. During P/D, the decode instance doesn't receive this information, and thus the unverified draft tokens remain in the KV cache, hurting acceptance rates.

This PR

Generalizes the 1-token backoff used by Mamba to skip prefilling the last token on the prefill instance, and recompute that token on the decode instance, for multi-module MTP. Now, the prefill instance no longer pollutes the last N-1 MTP modules with unverified draft tokens, and the decode instance completes prefill for the last N-1 prompt tokens.

I created a helper method in VllmConfig called num_prefill_lookahead_tokens that is used to get the number of MTP modules, and use them for this prefill backoff. It is now called by several other callsites to reduce code duplication.

Benchmarks

Inkling-Small-NVFP4, 1P1D on one 4×GB200 node (prefill GPUs 0,1 / decode GPUs 2,3, TP=2 each),
{"method":"mtp","num_speculative_tokens":8}, SPEED-Bench throughput_16k/low_entropy,
2048 in / 2048 out, concurrency 64, 512 prompts. Acceptance read from the decode instance's /metrics.

Metric Before (ae71862c51) After (5ee6a77da2) Delta
acceptance_length 1.0891 3.5207 +2.4316
acceptance_rate 1.11% 31.51% +30.40
Total token throughput 3,520 tok/s 12,156 tok/s +245%
Mean TPOT 35.67 ms 9.99 ms −72%
Mean E2EL 73.6 s 20.5 s −72%
Benchmark wall time 634.8 s 204.4 s −68%
Position 0 1 2 3 4 5 6 7
Before .0259 .0179 .0132 .0095 .0074 .0060 .0050 .0042
After .6745 .4802 .3575 .2806 .2282 .1917 .1644 .1436

Both arms completed 512/512 requests with 0 failures on identical input (1,048,576 tokens).
Output token counts differ by 2.3%. An acceptance length of 1.089 means ~99% of drafts were
rejected. Speculative decoding was effectively inert under P/D before the fix.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@TheEpicDolphin TheEpicDolphin added the ready ONLY add when PR is ready to merge/full CI is needed label Sep 11, 2026
@TheEpicDolphin

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #88418 for commit b84a1b39b283.

…yer MTP kv caches during P/D

Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
@TheEpicDolphin
TheEpicDolphin force-pushed the bugfix/mrv2-multi-module-mtp-pd branch from b84a1b3 to 0beece0 Compare September 12, 2026 06:23
@TheEpicDolphin

Copy link
Copy Markdown
Member Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #88486 for commit 0beece0bd81d.

@WoosukKwon
WoosukKwon merged commit 85c1f58 into vllm-project:main Sep 16, 2026
150 of 151 checks passed
@TheEpicDolphin
TheEpicDolphin deleted the bugfix/mrv2-multi-module-mtp-pd branch September 16, 2026 00:46
MK-BK pushed a commit to MK-BK/vllm that referenced this pull request Sep 23, 2026
…yer MTP kv caches during P/D (vllm-project#55055)

Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working kv-cache-manager kv-connector mrv2 Model Runner V2 specific ready ONLY add when PR is ready to merge/full CI is needed scheduler

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants