Repository navigation
Conversation
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Size draft input buffers for the two local segments PCP can create per request. Reject unsupported multi-module MTP and CUDA-graph configurations during validation instead of failing later at runtime. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Summer Yang <girasoleyang@gmail.com>
Reuse the PCP manager's rank-local input buffer and target hidden states for the sharded draft prefill. Preserve GPU-materialized positions after speculative rejection and keep local-only sampling metadata disposable until the full batch is restored. Make dummy and stale batches bypass rank-local PCP state, and derive MTP index sharing directly from the configured PCP size. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Summer Yang <girasoleyang@gmail.com>
3 of 9 tasks
1 task done
Contributor
|
This pull request has merge conflicts that must be resolved before it can be |
LucasWilkinson
pushed a commit
to neuralmagic/vllm
that referenced
this pull request
Aug 31, 2026
Signed-off-by: Summer Yang <girasoleyang@gmail.com>
This was referenced Sep 9, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Related RFC: #25749
Add single-module MTP support to the MRV2 PCP path while preserving PCP's long-prefill scaling:
This relies only on the PCP implementation already present on vLLM main and is independent of the fused PCP norm/RoPE work.
Relationship to #53427
This is a materially different execution design from #53427. That PR keeps the target PCP-sharded but runs a fully replicated drafter on every PCP rank and currently excludes sparse MLA. This change shards the prompt-sized MTP draft pass across PCP ranks, supports the GLM-5.2 sparse-MLA path, and avoids repeating long-prefill draft compute and cache writes on every rank. It adopts the same GPU-cursor correctness principle for rank-local positions, but not the replicated-drafter architecture.
Scope
Test Plan
Before marking the PR ready, rerun an exact distributed PCP4/MTP3 GLM-5.2 model configuration on the rebased head.
Test Result
7ca336929c169fee1210dd5293029d78811fba27git diff --checkpassedThe exact rebased and simplified head has not yet been rerun in a distributed PCP4/MTP3 model configuration, which is why this PR is being opened as a draft.
AI assistance
Codex assisted with the rebase audit, design comparison, simplification, regression tests, and PR description. The submitter reviewed the resulting changes and validation evidence.