Skip to content

[Bugfix][KV Connector][Mooncake] Apply multi-module MTP prefill backoff in P/D - #58556

Open
waizuichougou wants to merge 1 commit into
vllm-project:mainfrom
waizuichougou:fix/mooncake-mtp-pd-prefill-backoff
Open

waizuichougou wants to merge 1 commit into
vllm-project:mainfrom
waizuichougou:fix/mooncake-mtp-pd-prefill-backoff

Conversation

@waizuichougou

@waizuichougou waizuichougou commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Fix Model Runner V2 multi-module MTP with Mooncake P/D disaggregation.

The multi-module MTP scheduler exposes VllmConfig.num_prefill_lookahead_tokens
so the decoder can recompute the trailing lookahead window locally. The NIXL
connector already applies this backoff, but MooncakeConnector still transfers
the full prompt for dense MTP requests and does not truncate the producer-side
prefill request.

For example, with a prompt of 10 tokens and a three-token MTP lookahead,
Mooncake currently transfers 10 tokens while the decoder-side MTP path expects
the last 2 tokens to be recomputed locally. This can leave unverified MTP draft
KV in the P/D transfer state and make the decoder reuse KV that it does not
rebuild.

This change:

  • applies the same prefill backoff as NIXL when calculating Mooncake's remote
    prefill token count;
  • truncates the same trailing tokens on the producer before prefix-cache lookup;
  • preserves the existing one-token Mamba backoff and leaves ordinary dense
    requests unchanged; and
  • keeps the truncation idempotent across preemption and rescheduling.

This follows up on the multi-layer MTP P/D cache correction in #55055.

Test Plan

pytest -q \
  tests/v1/kv_connector/unit/test_mooncake_connector_hybrid_mamba.py
python -m py_compile \
  vllm/distributed/kv_transfer/kv_connector/v1/mooncake/mooncake_connector.py \
  tests/v1/kv_connector/unit/test_mooncake_connector_hybrid_mamba.py

The regression tests cover the dense multi-module MTP P/D case, verify that
both sides use the same lookahead backoff, and confirm that ordinary dense
requests keep the existing full-prompt behavior. Existing hybrid Mamba tests
cover the unchanged one-token behavior.

Test Result

  • 11 passed
  • An end-to-end multi-module MTP P/D accuracy run has not been performed.

Essential Elements of an Effective PR Description Checklist
  • The purpose and affected behavior are described.
  • The test plan is included.
  • Test results are included.
  • No documentation or supported-model update is required.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added bug Something isn't working kv-connector labels Sep 24, 2026
Signed-off-by: waizuichougou <2082431897@qq.com>
@waizuichougou
waizuichougou force-pushed the fix/mooncake-mtp-pd-prefill-backoff branch from d731f91 to 3fc58ba Compare September 24, 2026 12:57
@mergify mergify Bot added the mooncake Mooncake KV-transfer / EC-transfer label Oct 1, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working kv-connector mooncake Mooncake KV-transfer / EC-transfer

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant