Repository navigation
Conversation
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: f31a8af8fd
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| attn_metadata = speculator._build_draft_attn_metadata( | ||
| num_reqs=num_reqs, | ||
| num_reqs_padded=num_reqs_padded, | ||
| num_tokens_padded=num_tokens_padded, | ||
| seq_lens_cpu_upper_bound=input_batch.seq_lens_cpu_upper_bound, |
There was a problem hiding this comment.
Refresh the draft max sequence length before metadata
For the standard AutoRegressiveSpeculator MTP path, this helper builds metadata before propose() updates draft_max_seq_len from the current batch, so _build_draft_attn_metadata() uses the preceding batch's value. If the longest request accepts all K drafts plus its bonus token, the next batch's sequence upper bound grows by K+1 while the saved bound only included K lookahead tokens, making max_seqlen_k one token smaller than seq_lens; FlashAttention MLA may then reject the launch or process invalid draft attention. Compute the current bound before this call or pass it directly rather than relying on mutable state from the previous proposal.
Useful? React with 👍 / 👎.
Enable MTP speculative decoding in the MRV2 GPU runner when prefill context parallelism is enabled. The target model stays PCP-sharded while the MTP drafter runs replicated over the global batch on every rank: - Drop PCP from the drafter's config copy so its attention groups, forward context, and cudagraphs are built without PCP. Since the drafter builds its own draft prefill attention metadata from the (global) input batch, no metadata plumbing is needed. - partition_batch supports multi-token decode rows: reuse the GPU request-state positions verbatim (the CPU num_computed_tokens is only an async upper bound after rejection), derive local seq_lens from the segment end positions, and take last-token logits indices directly instead of recombining sampled/draft tokens. - restore_hidden_states appends explicit zero rows when the global batch is graph-padded, and restore_hidden_state_buffer restores persistent max-token buffers (e.g. DeepSeek V4's pre-hc_head residual). Unsupported combinations (non-MTP methods, sparse MLA, DCP, adaptive verification) are rejected at config validation. Adapted from #53427. Co-authored-by: QiuChunshuo <qiuchunshuo@huawei.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Enable MTP speculative decoding in the MRV2 GPU runner when prefill context parallelism is enabled. The target model stays PCP-sharded while the MTP drafter runs replicated over the global batch on every rank: - Drop PCP from the drafter's config copy so its attention groups, forward context, and cudagraphs are built without PCP. Since the drafter builds its own draft prefill attention metadata from the (global) input batch, no metadata plumbing is needed. - partition_batch supports multi-token decode rows: reuse the GPU request-state positions verbatim (the CPU num_computed_tokens is only an async upper bound after rejection), derive local seq_lens from the segment end positions, and take last-token logits indices directly instead of recombining sampled/draft tokens. - restore_hidden_states appends explicit zero rows when the global batch is graph-padded, and restore_hidden_state_buffer restores persistent max-token buffers (e.g. DeepSeek V4's pre-hc_head residual). Unsupported combinations (non-MTP methods, sparse MLA, DCP, adaptive verification) are rejected at config validation. Adapted from #53427. Co-authored-by: QiuChunshuo <qiuchunshuo@huawei.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
|
Im thinking we should maybe do something more like pisceskkk#6 |
… glm-53-blog/prefiller
|
This pull request has merge conflicts that must be resolved before it can be |
…m-project#53427 merge 732f952 enabled PIECEWISE graphs for sparse MLA under PCP (with vllm-project#53515's persistent input buffers); the vllm#53427 merge brought the old guard back.
have modified the pr with pisceskkk#6 and resolved the merge conflicts. ptal when you have time, thanks! |
Enable MTP speculative decoding in the MRV2 GPU runner when prefill context parallelism is enabled. The target model stays PCP-sharded while the MTP drafter runs replicated over the global batch on every rank: - Drop PCP from the drafter's config copy so its attention groups, forward context, and cudagraphs are built without PCP. Since the drafter builds its own draft prefill attention metadata from the (global) input batch, no metadata plumbing is needed. - partition_batch supports multi-token decode rows: reuse the GPU request-state positions verbatim (the CPU num_computed_tokens is only an async upper bound after rejection), derive local seq_lens from the segment end positions, and take last-token logits indices directly instead of recombining sampled/draft tokens. - restore_hidden_states appends explicit zero rows when the global batch is graph-padded, and restore_hidden_state_buffer restores persistent max-token buffers (e.g. DeepSeek V4's pre-hc_head residual). Unsupported combinations (non-MTP methods, sparse MLA, DCP, adaptive verification) are rejected at config validation. Adapted from vllm-project#53427. Co-authored-by: QiuChunshuo <qiuchunshuo@huawei.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Enable MTP speculative decoding in the MRV2 GPU runner when prefill context parallelism is enabled. The target model stays PCP-sharded while the MTP drafter runs replicated over the global batch on every rank: - Drop PCP from the drafter's config copy so its attention groups, forward context, and cudagraphs are built without PCP. Since the drafter builds its own draft prefill attention metadata from the (global) input batch, no metadata plumbing is needed. - partition_batch supports multi-token decode rows: reuse the GPU request-state positions verbatim (the CPU num_computed_tokens is only an async upper bound after rejection), derive local seq_lens from the segment end positions, and take last-token logits indices directly instead of recombining sampled/draft tokens. - restore_hidden_states appends explicit zero rows when the global batch is graph-padded, and restore_hidden_state_buffer restores persistent max-token buffers (e.g. DeepSeek V4's pre-hc_head residual). Unsupported combinations (non-MTP methods, sparse MLA, DCP, adaptive verification) are rejected at config validation. Adapted from vllm-project#53427. Co-authored-by: QiuChunshuo <qiuchunshuo@huawei.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
…coding with PCP Enable MTP speculative decoding in the MRV2 GPU runner when prefill context parallelism is enabled. The target model stays PCP-sharded while the MTP drafter runs replicated over the global batch on every rank: - Drop PCP from the drafter's config copy so its attention groups, forward context, and cudagraphs are built without PCP. Since the drafter builds its own draft prefill attention metadata from the (global) input batch, no metadata plumbing is needed. - partition_batch supports multi-token decode rows: reuse the GPU request-state positions verbatim (the CPU num_computed_tokens is only an async upper bound after rejection), derive local seq_lens from the segment end positions, and take last-token logits indices directly instead of recombining sampled/draft tokens. - restore_hidden_states appends explicit zero rows when the global batch is graph-padded, and restore_hidden_state_buffer restores persistent max-token buffers (e.g. DeepSeek V4's pre-hc_head residual). Unsupported combinations (non-MTP methods, sparse MLA, DCP, adaptive verification) are rejected at config validation. Adapted from vllm-project#53427. Co-authored-by: QiuChunshuo <qiuchunshuo@huawei.com> Co-authored-by: Claude Fable 5 <noreply@anthropic.com> Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
|
This pull request has merge conflicts that must be resolved before it can be |
|
This pull request has merge conflicts that must be resolved before it can be |
Signed-off-by: QiuChunshuo <qiuchunshuo@huawei.com>
|
This pull request has merge conflicts that must be resolved before it can be |
Summary
Extend the PCP speculative-decoding support already merged in #56107 with backend-aware draft execution:
This allows PCP-capable and non-PCP-capable draft backends to coexist without disabling the upstream sharded implementation.
Implementation
Draft mode is resolved before model construction from the effective draft attention backend, including
attention_backendfrom the speculative config andbackend_per_kindprecedence:use_pcp=Trueand considers PCP-capable candidates.supports_pcp() == falsebuilds the drafter withprefill_context_parallel_size=1.The two runtime contracts remain separate:
For PIECEWISE graphs, replicated drafting captures with drafter-owned attention groups and buffers. Runtime graph-padding rows receive zero hidden states and the global
is_paddingmask.Warmup follows the same selection: sharded MTP stays rank-local, while replicated drafters restore the global batch before
propose().Validation
Current commit:
The available 5090 development container did not have a main-compatible Python/test environment, so model execution for this rebased commit remains to be rerun on the H20 environment.
Replicated-mode model evidence
The following full GSM8K results were collected on the earlier global-replicated implementation using 4x NVIDIA H20,
/home/weight/GLM-4.7-Flash,FLASH_ATTN_MLA, EP4, MTP3, all 1319 questions, 5-shot, temperature 0, seed 42, max output tokens 4096, and concurrency 32. They are retained as reference evidence for the replicated execution contract; they are not reported as a rerun of this rebased commit.Dynamic K used K=3 for batch sizes 1-8, K=2 for 9-16, and K=1 for 17-32. Its raw acceptance rate is not directly comparable with fixed K.
The fixed-K PCP and TP acceptance metrics were aligned (acceptance-rate delta 0.000771; acceptance-length delta 0.002314). PIECEWISE was also aligned with the eager PCP reference (acceptance-rate delta 0.000543; acceptance-length delta 0.001628).
Scope
AI assistance was used for implementation and validation. The human submitter reviewed the changes.