[Spec Decode] Remove eager metadata rebuild during MTP fused multi-step decode - #58463
Open
TheEpicDolphin wants to merge 1 commit into
Open
TheEpicDolphin wants to merge 1 commit into
TheEpicDolphin wants to merge 1 commit into
Conversation
TheEpicDolphin
force-pushed
the
optimize-multi-step-decode-fusion
branch
from
September 23, 2026 22:36
cfb93de to
e230f58
Compare
TheEpicDolphin
marked this pull request as ready for review
September 23, 2026 22:41
TheEpicDolphin
requested review from
LucasWilkinson,
MatthewBonanni,
WoosukKwon,
alexm-redhat,
njhill,
pavanimajety,
yewentao256,
youkaichao and
zhuohan123
as code owners
September 23, 2026 22:41
Collaborator
Author
|
/ci run |
|
✅ Triggered Buildkite CI #90794 for commit |
Collaborator
Author
|
/ci retry |
|
✅ Queued 2 failed job(s) for retry in Buildkite CI #90794. |
…ep decode Signed-off-by: Giancarlo Delfin <gdelfin@inferact.ai>
TheEpicDolphin
force-pushed
the
optimize-multi-step-decode-fusion
branch
from
September 24, 2026 04:49
6674644 to
e0eaf5a
Compare
Collaborator
Author
|
/ci run |
|
✅ Triggered Buildkite CI #90856 for commit |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Context
#46849 added the infrastructure to fuse last N-1 draft steps into a single cudagraph for models with attention backends that support it. Currently DeepSeek V4 satisfies the conditions to enable this, when using MTP > 1.
This PR
It turns out we can remove the first attention metadata build before drafting the last N-1 speculative tokens. In this PR, I make the necessary changes to
DeepseekSparseSWAMetadataBuilder.update_draft_decode_metadataand refactorAutoRegressiveSpeculatorso that the first draft metadata build is no longer necessary.With these changes, DSV4 + MTP no longer rebuilds metadata during proposal, resulting in a correctness-preserving removal of about 1ms of host work per decode step. No significant throughput improvement was observed from these changes.
Profile
Captured while serving DeepSeek-V4-Flash + MTP=3 on 4xGB200:


Before
After
In the decode step above,
propose's CPU time shrinks by an entire millisecond.Benchmarks / Evals
Setup. DeepSeek-V4-Flash, MTP with
num_speculative_tokens=3, 4x GB200 (TP4),--block-size 256 --kv-cache-dtype fp8 --moe-backend deep_gemm --no-enable-prefix-caching,VLLM_USE_V2_MODEL_RUNNER=1. Baseline ismainat34b069080d; PR iscfb93dedca. One server per arm; trace, benchmark and eval share it.Speed-Bench 2K/2K (
throughput_16k,low_entropy, temperature 0)Prompts scaled 8x concurrency (min 16), seed 0. Greedy decoding keeps acceptance deterministic, so ITL p50 is the step-time metric.
Step time is unchanged at every concurrency. The ~1 ms of host work removed per step is not on the critical path for this model at TP4; throughput deltas track the small acceptance-length differences and are within noise.
AIME25 (MQA, 2 epochs x 30 problems, thinking on, temperature 1.0, top-p 0.95, max 98304 tokens, 64 concurrent)
Mean accuracy identical. pass@2 differs by one problem in 30, which is expected sampling variance at temperature 1.