Repository navigation
[Bugfix][KV Connector][Mooncake] Apply multi-module MTP prefill backoff in P/D - #58556
Open
waizuichougou wants to merge 1 commit into
Open
waizuichougou wants to merge 1 commit into
waizuichougou wants to merge 1 commit into
Conversation
waizuichougou
requested review from
ApostaC,
NickLucche,
ivanium,
orozery and
xuechendi
as code owners
September 24, 2026 12:54
Signed-off-by: waizuichougou <2082431897@qq.com>
waizuichougou
force-pushed
the
fix/mooncake-mtp-pd-prefill-backoff
branch
from
September 24, 2026 12:57
d731f91 to
3fc58ba
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Fix Model Runner V2 multi-module MTP with Mooncake P/D disaggregation.
The multi-module MTP scheduler exposes
VllmConfig.num_prefill_lookahead_tokensso the decoder can recompute the trailing lookahead window locally. The NIXL
connector already applies this backoff, but
MooncakeConnectorstill transfersthe full prompt for dense MTP requests and does not truncate the producer-side
prefill request.
For example, with a prompt of 10 tokens and a three-token MTP lookahead,
Mooncake currently transfers 10 tokens while the decoder-side MTP path expects
the last 2 tokens to be recomputed locally. This can leave unverified MTP draft
KV in the P/D transfer state and make the decoder reuse KV that it does not
rebuild.
This change:
prefill token count;
requests unchanged; and
This follows up on the multi-layer MTP P/D cache correction in #55055.
Test Plan
The regression tests cover the dense multi-module MTP P/D case, verify that
both sides use the same lookahead backoff, and confirm that ordinary dense
requests keep the existing full-prompt behavior. Existing hybrid Mamba tests
cover the unchanged one-token behavior.
Test Result
11 passedEssential Elements of an Effective PR Description Checklist