Skip to content

[Spec Decode] Remove eager metadata rebuild during MTP fused multi-step decode - #58463

Open
TheEpicDolphin wants to merge 1 commit into
vllm-project:mainfrom
TheEpicDolphin:optimize-multi-step-decode-fusion
Open

TheEpicDolphin wants to merge 1 commit into
vllm-project:mainfrom
TheEpicDolphin:optimize-multi-step-decode-fusion

Conversation

@TheEpicDolphin

@TheEpicDolphin TheEpicDolphin commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Context

#46849 added the infrastructure to fuse last N-1 draft steps into a single cudagraph for models with attention backends that support it. Currently DeepSeek V4 satisfies the conditions to enable this, when using MTP > 1.

This PR

It turns out we can remove the first attention metadata build before drafting the last N-1 speculative tokens. In this PR, I make the necessary changes to DeepseekSparseSWAMetadataBuilder.update_draft_decode_metadata and refactor AutoRegressiveSpeculator so that the first draft metadata build is no longer necessary.

With these changes, DSV4 + MTP no longer rebuilds metadata during proposal, resulting in a correctness-preserving removal of about 1ms of host work per decode step. No significant throughput improvement was observed from these changes.

Profile

Captured while serving DeepSeek-V4-Flash + MTP=3 on 4xGB200:
Before
image
After
image

In the decode step above, propose's CPU time shrinks by an entire millisecond.

Benchmarks / Evals

Setup. DeepSeek-V4-Flash, MTP with num_speculative_tokens=3, 4x GB200 (TP4), --block-size 256 --kv-cache-dtype fp8 --moe-backend deep_gemm --no-enable-prefix-caching, VLLM_USE_V2_MODEL_RUNNER=1. Baseline is main at 34b069080d; PR is cfb93dedca. One server per arm; trace, benchmark and eval share it.

Speed-Bench 2K/2K (throughput_16k, low_entropy, temperature 0)

Prompts scaled 8x concurrency (min 16), seed 0. Greedy decoding keeps acceptance deterministic, so ITL p50 is the step-time metric.

Concurrency ITL p50 base (ms) ITL p50 PR (ms) Δ TPOT p50 base (ms) TPOT p50 PR (ms) Output tok/s base Output tok/s PR Δ Accept. len base Accept. len PR
1 8.97 8.97 0.0% 3.41 3.42 290.3 288.3 -0.7% 2.705 2.685
16 19.13 19.15 +0.1% 8.71 8.57 1806.7 1819.4 +0.7% 2.664 2.717
64 22.96 22.96 0.0% 14.52 14.58 4415.5 4380.6 -0.8% 2.693 2.693

Step time is unchanged at every concurrency. The ~1 ms of host work removed per step is not on the critical path for this model at TP4; throughput deltas track the small acceptance-length differences and are within noise.

AIME25 (MQA, 2 epochs x 30 problems, thinking on, temperature 1.0, top-p 0.95, max 98304 tokens, 64 concurrent)

Arm Exact match pass@2 Samples Request errors
base 0.9333 0.9667 60 0
PR 0.9333 0.9333 60 0

Mean accuracy identical. pass@2 differs by one problem in 30, which is expected sampling variance at temperature 1.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@TheEpicDolphin TheEpicDolphin added the ready ONLY add when PR is ready to merge/full CI is needed label Sep 23, 2026
@TheEpicDolphin

Copy link
Copy Markdown
Collaborator Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #90794 for commit 66746447d8f6.

@TheEpicDolphin

Copy link
Copy Markdown
Collaborator Author

/ci retry

@github-actions

Copy link
Copy Markdown

✅ Queued 2 failed job(s) for retry in Buildkite CI #90794.

…ep decode

Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
@TheEpicDolphin
TheEpicDolphin force-pushed the optimize-multi-step-decode-fusion branch from 6674644 to e0eaf5a Compare September 24, 2026 04:49
@TheEpicDolphin

Copy link
Copy Markdown
Collaborator Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #90856 for commit e0eaf5abff52.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

mrv2 Model Runner V2 specific ready ONLY add when PR is ready to merge/full CI is needed speculative-decoding

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant