Repository navigation
[Feat] Support PP with PCP in GPU Model Runner V2 - #59139
Merged
Merged
Conversation
pisceskkk
requested review from
WoosukKwon,
njhill and
yewentao256
as code owners
September 29, 2026 04:44
Assisted-by: OpenAI Codex Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
pisceskkk
force-pushed
the
codex/pcp-pp-main
branch
from
September 29, 2026 04:55
a5143bd to
343a3f6
Compare
Contributor
|
/ci run |
|
✅ Triggered Buildkite CI #91872 for commit |
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
Contributor
Author
|
/ci retry |
|
✅ Triggered Buildkite CI #91903 for commit |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Enable pipeline parallelism (PP) together with Prefill Context Parallelism (PCP) in GPU Model Runner V2 for MLA models.
The last PP stage restores PCP-local hidden states to global request order before sampling. Earlier PP stages use the global request batch when receiving sampled tokens and updating request state, while model forward continues on the PCP-local batch.
Related to #25749 and #50853. Open PRs #33403, #49109, #49246, and #53948 address other PCP or PP paths; none covers this Model Runner V2 request mapping for PP feedback.
AI assistance was used to prepare the implementation, tests, and validation.
Acceptance criteria
Test Plan
Run Ruff check and format on the three changed files. Compare PCP2 and PP2+PCP2 using a two-layer, randomly initialized DeepSeek-V3.2 FP8 model with one 607-token and one 7-token prompt, eight output tokens each, and prompt/output logprobs.
Evaluate DeepSeek-V2-Lite-Chat on all 1,319 GSM8K test questions with the same five-shot chat prompts, greedy decoding, and 256-token output limit in PCP2 and PP2+PCP2. Both runs use BF16, expert parallelism, FULL_AND_PIECEWISE CUDA graphs, and VLLM_MOE_SKIP_PADDING=0 on NVIDIA RTX 5090 GPUs.
Test Result
The focused test, Ruff check and format, and git diff --check passed on base aedaba8.
The DeepSeek-V3.2 functional check produced identical token IDs, text, and logprob counts in PCP2 and PP2+PCP2 for both requests.
DeepSeek-V2-Lite-Chat completed all 1,319 GSM8K questions in both topologies:
PP2+PCP2 scored 13 questions (0.99 percentage points) higher. Extracted answers matched on 1,070/1,319 questions; the cause of the cross-topology differences has not been isolated.