Skip to content

[Spec Decode] Support PCP-sharded MTP prefill - #53653

Closed
GirasoleY wants to merge 5 commits into
vllm-project:mainfrom
GirasoleY:review/pcp-prefill-mtp
Closed

GirasoleY wants to merge 5 commits into
vllm-project:mainfrom
GirasoleY:review/pcp-prefill-mtp

Conversation

@GirasoleY

@GirasoleY GirasoleY commented Aug 25, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Related RFC: #25749

Add single-module MTP support to the MRV2 PCP path while preserving PCP's long-prefill scaling:

  • keep both the target prefill and the first MTP draft pass PCP-sharded
  • restore target and draft outputs to global sequence order only at their sampling boundaries
  • run later MTP proposal steps as replicated decode queries
  • support pure prefill, replicated decode, mixed prefill/decode, and scheduled draft-token verification batches
  • preserve GPU-materialized token positions after speculative rejection instead of rebuilding them from a potentially stale CPU cursor
  • reuse the PCP manager's consumed rank-local input buffer and target hidden states for draft prefill, avoiding a doubled request buffer, metadata copies, and a redundant hidden-state all-gather
  • scope rank-local state to the exact global batch so dummy/profile and stale batches use the ordinary global draft path
  • distinguish PCP-expanded target slot mappings from ordinary replicated-draft mappings in sparse MLA
  • disable MTP top-k index sharing under PCP and wait for proposal cache writes before a KV producer publishes the result
  • account for PCP ranks in multi-node world-size and per-DP-rank placement
  • reject unsupported multi-module MTP and CUDA-graph configurations during startup

This relies only on the PCP implementation already present on vLLM main and is independent of the fused PCP norm/RoPE work.

Relationship to #53427

This is a materially different execution design from #53427. That PR keeps the target PCP-sharded but runs a fully replicated drafter on every PCP rank and currently excludes sparse MLA. This change shards the prompt-sized MTP draft pass across PCP ranks, supports the GLM-5.2 sparse-MLA path, and avoids repeating long-prefill draft compute and cache writes on every rank. It adopts the same GPU-cursor correctness principle for rank-local positions, but not the replicated-drafter architecture.

Scope

  • single-module MTP only; multi-module MTP and other speculative methods remain unsupported with PCP
  • CUDA graphs must be disabled for PCP+MTP
  • no model-routing change or new kernel

Test Plan

PYTHONDONTWRITEBYTECODE=1 python -m pytest -vv \
  tests/v1/worker/test_gpu_autoregressive_speculator.py \
  tests/v1/worker/test_gpu_pcp_manager.py
git diff --check

Before marking the PR ready, rerun an exact distributed PCP4/MTP3 GLM-5.2 model configuration on the rebased head.

Test Result

  • clean merge-tree against vLLM main at 7ca336929c169fee1210dd5293029d78811fba27
  • focused post-simplification tests: 24 passed, 2 CUDA-only tests skipped
  • commit-time checks passed, including Ruff, mypy, forbidden-import checks, configuration validation, SPDX, and sign-off
  • Python syntax compilation and git diff --check passed
  • pre-rebase CUDA validation: 21 focused tests passed
  • pre-rebase GLM-5.2 E2E validation used TP1/PCP8/EP8 with modular MTP5 on 8x B300, FP8 KV cache, an 8K batched-token limit, and a 4K long-prefill threshold:
    • restored decode and mixed spec/non-spec warmups completed
    • a 3,073-token prompt generated 32/32 requested tokens in the pure-decode probe
    • a sustained stream generated 128/128 requested tokens while two concurrent 3,073-token prefills each generated 8/8 requested tokens
    • the scheduler recorded a real mixed PCP batch with two fresh prefills and one decode; both node logs had zero errors
    • MTP metrics recorded 219 accepted draft tokens during the probe

The exact rebased and simplified head has not yet been rerun in a distributed PCP4/MTP3 model configuration, which is why this PR is being opened as a draft.

AI assistance

Codex assisted with the rebase audit, design comparison, simplification, regression tests, and PR description. The submitter reviewed the resulting changes and validation evidence.

GirasoleY and others added 5 commits August 24, 2026 15:53
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Size draft input buffers for the two local segments PCP can create per request. Reject unsupported multi-module MTP and CUDA-graph configurations during validation instead of failing later at runtime.

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Reuse the PCP manager's rank-local input buffer and target hidden states for the sharded draft prefill. Preserve GPU-materialized positions after speculative rejection and keep local-only sampling metadata disposable until the full batch is restored.

Make dummy and stale batches bypass rank-local PCP state, and derive MTP index sharing directly from the configured PCP size.

Co-authored-by: OpenAI Codex <codex@openai.com>

Signed-off-by: Summer Yang <girasoleyang@gmail.com>
@mergify

mergify Bot commented Aug 26, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @GirasoleY.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant