[Bugfix][MRV2][P/D] Preserve decode graph with recompute scheduler - #17128
realliujiaxu merged 1 commit into
Conversation
Signed-off-by: likailong <likailong5@huawei.com>
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [BugFix] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request addresses a bug where PD decode requests using the recompute scheduler were incorrectly identified as prefill, causing performance degradation by missing optimized decode graphs. By refining the reclassification of prompt-boundary recompute tokens, the system now correctly identifies these as decode operations, leading to significant throughput and latency improvements. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Ops][Feature] Support reclassifying PD tail in mixed decode batch for recompute schedulerSuggested PR Summary:
### What this PR does / why we need it?
This PR updates the `gather_batch_req_state` method in the Ascend V2 `GPUModelRunner` to support the PD decode recompute scheduler. When enabled, it identifies and reclassifies prefill requests that are actually part of a decode recompute as decode requests.
Feedback: A review comment suggests avoiding in-place mutation of `batch_state.is_prefilling_np` to prevent potential runtime errors or side-effects, recommending copying the array and using `_replace` instead.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Added unit tests in `tests/ut/worker/test_model_runner_v2.py` covering the reclassification logic, including mixed decode batches, non-matching prefills, and multi-token decode queries.…e) (#16967) ### What this PR does / why we need it? Enables `oproj_tensor_parallel_size` (fine-grained TP for the attention output projection) on model runner v2, scoped to its real deployment: a PD decode node (`tp == 1`, `dp > 1`, MoE) running a graph mode with the recompute scheduler. The o_proj exchange (`all_to_all` → partial GEMM → `reduce_scatter`) is an `all_to_all_single` without split sizes: every rank must forward the same number of tokens per step, or the collectives hang silently. In graph mode the upstream dispatch already guarantees that group-uniform count, so this PR adds no runner-side padding — it adds the contract that keeps the count uniform, and guardrails that turn a silent hang into an explicit failure: 1. Graph mode is required: `cudagraph_mode == NONE` (which covers `enforce_eager`) is rejected at config load. 2. Capture-bound check: the knob is kept only if the largest capture size covers the largest possible step (`min(max_num_batched_tokens, max_num_seqs * decode_query_len)`); otherwise it is disabled with a warning and the deployment still starts. 3. A guard at the top of `NPUModelRunner.prepare_inputs` — the only post-dispatch hook, reached by real steps only — raises an explicit `RuntimeError` if a step still dispatches to eager. Armed only for a real split (`> 1`). 4. PCP (`prefill_context_parallel_size > 1`) is rejected: its per-rank token recount breaks the uniform step size. 5. With `> 1`, the KV connector chain must carry `PreemptOffloadConnector` — startup fails without it — and an offload failure at preemption aborts the request instead of sending it back to P. Checked in `platform._check_ascend_config` because MultiConnector children re-derive the config from per-child copies that cannot see their siblings. Existing preconditions unchanged: PD decode node (`is_kv_consumer`), `recompute_scheduler_enable=true`, `tensor_parallel_size == 1`, MoE. Depends on #17128 (reclassifies the request-arrival boundary recompute as decode, so those steps stay on captured graphs); this PR is meant to land after it. ### Design conclusions (aligned with the SE) The scope and guardrails of this PR implement conclusions aligned with the SE: 1. **No MRV1 parity goal.** Fine-grained TP is adapted to model runner v2 only where necessary and reasonable; the bar is expected functional correctness, not parity with MRV1. 2. **Preconditions are enforced at config load.** `oproj_tensor_parallel_size` requires the recompute scheduler, preempt offload, a graph mode, and a pure-DP PD decode deployment (`tp == 1`, `dp > 1`). 3. **No eager fallback from an oversized step.** A step whose token count exceeds the largest captured size dispatches to eager in every graph mode; a conservative config-time check verifies the capture bound covers the largest possible step and disables the knob otherwise. Other fine-grained TP knobs can follow the same pattern. 4. **No return-to-P fallback under this feature.** Recompute preemption is paired with offload; the stock path sends the request back to P when the offload itself fails. This feature rejects that branch — recomputation on P cannot guarantee computation equivalence — and fails the request immediately with an explicit error instead. 5. **`PreemptOffloadConnector` is mandatory.** Config load rejects a real split whose KV connector chain does not carry it. ### Memory cost and expected savings The memory profile belongs to the feature, not to this PR — the validation and guard paths allocate no device memory. Measured on the test bed (DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers, `dp8 × tp1`, graph mode, identical connector config; per rank): | Item | Value | Basis | |---|---|---| | Fixed cost: HCCL buffers of the o_proj TP communication group | ~0.42 GB | measured (steady-state device delta; the freed weight is cancelled by the larger KV pool); allocated at graph-capture time | | Weight saving per layer, 1-byte (W8A8/FP8) checkpoint | 0.109 GiB × (1 − 1/N) | measured: 0.33 GiB @ `oproj=2`, 0.49 GiB @ `oproj=4` on 6 layers | | Full 61-layer extrapolation, W8A8/FP8, `oproj=8` | ~5.8 GB gross, ~5.4 GB net of the fixed cost | from the measured per-layer slope | | Net savings, DeepSeek-R1-W8A8, `oproj=8` | 5.8 GB | feature guide (`Fine_grained_TP.md`) | | DeepSeek-V3/R1 BF16, `oproj=8` | ~11.2 GiB net | from checkpoint dims (0.219 GiB per layer) | The freed weight memory returns to the KV cache, and the fixed cost is recovered after ~8 layers at `oproj=2` (~4 at `oproj=8`) on a 1-byte quantized checkpoint — trivially met by production-scale MoE models. ### Does this PR introduce _any_ user-facing change? Yes. - `oproj_tensor_parallel_size` now works on the V2 runner in graph-mode PD-decode deployments; previously it was V1-only. - New config-load rejections: `cudagraph_mode=none` / `enforce_eager`, `prefill_context_parallel_size > 1`, and — for a real split — a KV connector chain without `PreemptOffloadConnector`. - Capture-bound auto-disable with a warning instead of rejection; raise `max_cudagraph_capture_size` to use the knob at large `max_num_seqs`. - An eagerly dispatched step now fails the request with an explicit error instead of hanging; an offload failure at preemption aborts the request instead of returning it to P. ### How was this patch tested? **Unit tests**: new cases in `tests/ut/test_ascend_config.py` (graph-mode / PCP rejection, size-1 exemption, capture-bound matrix), `tests/ut/worker/test_model_runner_v2_finegrained_tp.py` (eager-step guard contract) and `tests/ut/core/test_recompute_scheduler.py` (offload-failure abort). CI on the current head: `pre-commit` (incl. mypy) and `cpu-ut` green (5718 passed, 67 skipped). **End-to-end** (single-node A3, 16 chips, PD disaggregation behind the LB proxy, 8-layer MiniMax-M3 with dummy weights — functional/consistency numbers, not real-weight performance): P = chip 0, `tp1` eager, Mooncake producer; D = chips 8-15, `dp8 × tp1`, `FULL_DECODE_ONLY` + recompute scheduler + `oproj_tensor_parallel_size=2`, Mooncake consumer; all traffic through the proxy. | Check | Result | |---|---| | config rejections | graph-mode and PD-scenario rejections fail startup with explicit messages; `--max-num-seqs 600` (600 > 512) auto-disables the knob with a warning and startup continues into graph capture | | continuous PD traffic, `oproj=2` | 100% of requests completed, guard fired 0 times, decode ran on captured FULL graphs, engine TPOT cc1/cc8 median 7.21 / 7.58 ms (990.7 tok/s at cc8) | | vs `oproj=0` baseline | TPOT +0.07 ms (~1%); outputs token-identical on 5 fixed prompts in every round | | negative control (same topology without #17128) | the first arrival step dispatches to eager → explicit `RuntimeError` within 3.7 s, no hang | | offload chain (`MultiConnector`: Mooncake + `PreemptOffloadConnector`) | starts, graph capture OK, cc8 7.67 ms / 943 tok/s, outputs identical, CPU offload pool active on 8/8 decode ranks | | tight-KV sweeps (`gpu-memory-utilization` 0.5 / 0.4 / 0.35, 100 MB CPU pool) | 100% of requests completed, outputs identical, no crash | **Left uncovered** (for honesty): the preemption/offload runtime path (KV save/restore and the capacity fallback) was not triggered — preemption is unreachable on the 8-layer dummy model — so the abort-instead-of-return-to-P branch is covered by unit tests only; and the without-connector startup failure is a config-load check that has not been reproduced e2e yet (it was observed on the test bed while the check still warned). - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: xuchi <xuchicolson@163.com>
### What this PR does / why we need it? Remove the obsolete vLLM 0.28.0 PCP+DP compatibility introduced by #16853 (`8f2e3fed73328148ea603ccfe5764542fa7cbbd2`) after #17004 upgraded the supported release to v0.29.0. Both v0.29.0 and the paired vLLM main provide native PCP-local token counting before dispatch. This is a conflict-resolved revert with targeted cleanup: - Remove the 0.28-only ParallelConfig validator shim, dispatch ContextVar/wrapper, the legacy dispatch block inside gather_batch_req_state, and PCP token-count override. - Preserve the independent v0.29.0 patch_parallel_config.py: the release still rejects PCP+DP without that patch. - Preserve #17128's PD decode recompute logic in gather_batch_req_state and its regression tests. - Preserve subsequent changes, including nullcontext and the v0.29.0 argument gates, generic PCP count/feature tests, and slot-buffer dtype coverage. - Remove obsolete adapter-specific tests and patch documentation. Remove imports made unused by the cleanup, including vllm_version_is in pcp_manager.py; that import predates #16853 but has no remaining references after this revert and would trigger Ruff F401. No new tests or unrelated formatting changes. The execute_model body is unindented only to remove its obsolete context wrapper. ### Does this PR introduce _any_ user-facing change? No intended behavior change on the supported v0.29.0 and fixed-main lanes. The retired v0.28.0 compatibility is removed; v0.29.0 PCP+DP configuration support remains provided by the independent validator patch. ### How was this patch tested? On rebased commit `d6b9f29e01ee2daa2d594b8c5c828e94dac0b128`: - Ruff lint/format, AST syntax, conflict-marker and git diff --check checks passed for the seven changed Python files. - Preserve the new upstream step_eplb_after import, PD decode recompute logic, and o_proj TP graph checks. Range-diff against 5034869 shows only upstream context changes; the cleanup scope is unchanged. - DCO Signed-off-by is retained. Fresh CI is pending; no local Ascend runtime tests were run on this revision. - Previous revision 5034869 passed [E2E run 35741354764](https://github.com/vllm-project/vllm-ascend/actions/runs/35741354764), including pre-commit, CPU UT, selected NPU jobs, upstream tests and ci-gate. Skipped jobs are not passes; these results do not establish validation of this rebased revision. Base: vllm-ascend main `972fcd5c974d55a7970ebe630329daa4b4049c49`. Release: vLLM v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`). vLLM main: vllm-project/vllm@84030bb - vLLM main: vllm-project/vllm@84030bb Signed-off-by: wzx0726 <278573478+wzx0726@users.noreply.github.com> Co-authored-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
…del runner v2 (#17361) ### What this PR does / why we need it? Extends fine-grained TP on model runner v2 beyond o_proj, on top of the framework landed in #16967 (and its prerequisite #17128). The four knobs split into two families with different contracts: 1. **MLP TP joins the o_proj (uniform-token) contract.** The MLP exchange (gate_up `all_gather` / down `reduce_scatter` over DP-axis groups) needs the same group-uniform token count per step as o_proj, so `mlp_tensor_parallel_size > 1` joins the graph-mode / PCP / `tp == 1` / PD-scenario checks, the capture-bound auto-disable (a miss disables both knobs together — the bound is a property of the step, not of a knob), the `PreemptOffloadConnector` gate, the offload-failure abort, and the eager-step guard (renamed `_check_finegrained_tp_graph_step`). 2. **Embedding TP is capacity-based and joins none of the uniform-token checks.** Its exchange pads every step to a fixed capacity inside the op, so it is step-shape-agnostic and cannot desync; the runtime capacity already floors at `max_num_batched_tokens`, so its one real failure mode fails fast on its own. 3. **The recompute-scheduler requirement covers all four knobs.** o_proj and embedding already required it on main; mlp joins as it enters the uniform-token contract, and lmhead joins so the whole `finegrained_tp_config` stays within the one configuration validated end-to-end. The first traffic test of the no-recompute path on the PD test bed produced deterministically wrong outputs with every knob off (a pre-existing defect unrelated to this PR, tracked separately), so that path stays outside the supported set until it is fixed. ### Does this PR introduce _any_ user-facing change? Yes. - `mlp_tensor_parallel_size` now works on the V2 runner in graph-mode PD-decode deployments, with the same preconditions as `oproj_tensor_parallel_size`. - All four `finegrained_tp_config` sizes require `recompute_scheduler_enable=true` (previously only o_proj and embedding did). - At runtime, an eagerly dispatched step under o_proj/MLP TP fails the request with an explicit error (the message now names both knobs). ### How was this patch tested? **Unit tests**: one new case in `tests/ut/test_ascend_config.py` (MLP capture-bound, incl. the both-knobs-disabled-together semantics); the remaining gates are shared conditions already covered by the existing o_proj cases. CI on the current head: `pre-commit` and `cpu-ut` green. **End-to-end** — single-node A3 (16 chips), PD disaggregation behind the LB proxy; DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers (real weights, `--quantization ascend`). P = chip 0 (`tp1`, eager, Mooncake producer); D = chips 8-15 (`dp8 × tp1`, `FULL_DECODE_ONLY`, capture size 512, recompute scheduler on, `MultiConnector` = Mooncake consumer + `PreemptOffloadConnector`); every request goes through the proxy. Each round: 8-concurrent × 256-token bench × 3 (6144 tokens) plus 5 fixed greedy prompts for consistency. Positive rounds (all: 8/8 startup, 100% of bench tokens, knob reported enabled, zero guard hits, zero auto-disables, zero collective errors): | Round | Engine TPOT | Output vs baseline | |---|---|---| | baseline | 7.12 ms | — | | `mlp=2` | 7.33 ms | 1/5 prompts identical, the rest diverge at the first token (see note) | | `oproj=2 + mlp=2` | 7.24 ms | same class as `mlp=2`; bit-identical to the three-knob round | | `embedding=2` | 7.32 ms | **5/5 bit-identical** | | all three knobs | 7.14 ms | bit-identical to the `oproj=2+mlp=2` round | Consistency note: re-sharding o_proj/MLP changes the reduction order, and the truncated 6-layer test model has a near-flat output distribution, so the greedy pick can flip at the first token — the same class of difference as changing the generic TP size. The embedding exchange is pure data movement and stays bit-identical, including when stacked with the other two knobs. Negative configs (rejected or disabled at startup, no hang): | Config | Behavior | |---|---| | `mlp=2` + `--enforce-eager` | rejected at config load ("only supported in graph mode") | | `mlp=2` without `PreemptOffloadConnector` | rejected with the explicit connector requirement | | `mlp=2` without the recompute scheduler | rejected at config load in ~15 s with the recompute-scheduler requirement (the all-knobs gate above) | | capture bound 512 < largest step 600 | both knobs auto-disabled with a warning; startup completes | Base controls: the rounds above ran on the base before #17159; since #17159 reworked the embedding exchange, the embedding-relevant cells were re-run on the current base — baseline and `embedding=2` outputs are **bit-identical** to the pre-#17159 results (TPOT 7.52 / 7.18 ms). A no-recompute control on the current base produced deterministically wrong outputs with **every knob off** (a pre-existing defect, tracked separately); `embedding=2` on that path is bit-identical to the no-knob run, i.e. fine-grained TP adds no effect — this is the defect behind the all-knobs recompute requirement. **Memory** (same bed): the family's fixed cost is one HCCL buffer set (~0.42 GB per rank, from #16967); MLP saves 0.55 GiB per rank at `mlp=2` on the 6-layer bed (dense-FFN layers × 1/2), o_proj 0.109 GiB per layer per rank on a 1-byte checkpoint (#16967). - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: xuchi <xuchicolson@163.com>
…llm-project#17128) ## Summary When PD decode uses `recompute_scheduler_enable=true`, the last prompt-token recompute is reported as prefill by the upstream batch state. This keeps `has_prefill` true and makes a mixed decode batch miss `FULL_DECODE_ONLY` graphs. Reclassify only the prompt-boundary recompute on a PD consumer, then refresh `has_prefill` and the uniform decode token count. This path does not require MTP and does not depend on vllm-project#17032. | Validation | Baseline | Fixed | |---|---:|---:| | Targeted UT | — | 4 passed | | Qwen3-32B-W8A8 requests | 128/128 | 128/128 | | Aggregate output throughput | 121.84 tok/s | 137.59 tok/s (+12.93%) | | Mean join decode latency | 3.2335 s | 2.3539 s (-27.20%) | Full `test_model_runner_v2.py`: 40 passed; its one remaining Spec-PP failure is also reproducible on the unmodified baseline. - vLLM main: vllm-project/vllm@84030bb Signed-off-by: likailong <likailong5@huawei.com>
…e) (vllm-project#16967) ### What this PR does / why we need it? Enables `oproj_tensor_parallel_size` (fine-grained TP for the attention output projection) on model runner v2, scoped to its real deployment: a PD decode node (`tp == 1`, `dp > 1`, MoE) running a graph mode with the recompute scheduler. The o_proj exchange (`all_to_all` → partial GEMM → `reduce_scatter`) is an `all_to_all_single` without split sizes: every rank must forward the same number of tokens per step, or the collectives hang silently. In graph mode the upstream dispatch already guarantees that group-uniform count, so this PR adds no runner-side padding — it adds the contract that keeps the count uniform, and guardrails that turn a silent hang into an explicit failure: 1. Graph mode is required: `cudagraph_mode == NONE` (which covers `enforce_eager`) is rejected at config load. 2. Capture-bound check: the knob is kept only if the largest capture size covers the largest possible step (`min(max_num_batched_tokens, max_num_seqs * decode_query_len)`); otherwise it is disabled with a warning and the deployment still starts. 3. A guard at the top of `NPUModelRunner.prepare_inputs` — the only post-dispatch hook, reached by real steps only — raises an explicit `RuntimeError` if a step still dispatches to eager. Armed only for a real split (`> 1`). 4. PCP (`prefill_context_parallel_size > 1`) is rejected: its per-rank token recount breaks the uniform step size. 5. With `> 1`, the KV connector chain must carry `PreemptOffloadConnector` — startup fails without it — and an offload failure at preemption aborts the request instead of sending it back to P. Checked in `platform._check_ascend_config` because MultiConnector children re-derive the config from per-child copies that cannot see their siblings. Existing preconditions unchanged: PD decode node (`is_kv_consumer`), `recompute_scheduler_enable=true`, `tensor_parallel_size == 1`, MoE. Depends on vllm-project#17128 (reclassifies the request-arrival boundary recompute as decode, so those steps stay on captured graphs); this PR is meant to land after it. ### Design conclusions (aligned with the SE) The scope and guardrails of this PR implement conclusions aligned with the SE: 1. **No MRV1 parity goal.** Fine-grained TP is adapted to model runner v2 only where necessary and reasonable; the bar is expected functional correctness, not parity with MRV1. 2. **Preconditions are enforced at config load.** `oproj_tensor_parallel_size` requires the recompute scheduler, preempt offload, a graph mode, and a pure-DP PD decode deployment (`tp == 1`, `dp > 1`). 3. **No eager fallback from an oversized step.** A step whose token count exceeds the largest captured size dispatches to eager in every graph mode; a conservative config-time check verifies the capture bound covers the largest possible step and disables the knob otherwise. Other fine-grained TP knobs can follow the same pattern. 4. **No return-to-P fallback under this feature.** Recompute preemption is paired with offload; the stock path sends the request back to P when the offload itself fails. This feature rejects that branch — recomputation on P cannot guarantee computation equivalence — and fails the request immediately with an explicit error instead. 5. **`PreemptOffloadConnector` is mandatory.** Config load rejects a real split whose KV connector chain does not carry it. ### Memory cost and expected savings The memory profile belongs to the feature, not to this PR — the validation and guard paths allocate no device memory. Measured on the test bed (DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers, `dp8 × tp1`, graph mode, identical connector config; per rank): | Item | Value | Basis | |---|---|---| | Fixed cost: HCCL buffers of the o_proj TP communication group | ~0.42 GB | measured (steady-state device delta; the freed weight is cancelled by the larger KV pool); allocated at graph-capture time | | Weight saving per layer, 1-byte (W8A8/FP8) checkpoint | 0.109 GiB × (1 − 1/N) | measured: 0.33 GiB @ `oproj=2`, 0.49 GiB @ `oproj=4` on 6 layers | | Full 61-layer extrapolation, W8A8/FP8, `oproj=8` | ~5.8 GB gross, ~5.4 GB net of the fixed cost | from the measured per-layer slope | | Net savings, DeepSeek-R1-W8A8, `oproj=8` | 5.8 GB | feature guide (`Fine_grained_TP.md`) | | DeepSeek-V3/R1 BF16, `oproj=8` | ~11.2 GiB net | from checkpoint dims (0.219 GiB per layer) | The freed weight memory returns to the KV cache, and the fixed cost is recovered after ~8 layers at `oproj=2` (~4 at `oproj=8`) on a 1-byte quantized checkpoint — trivially met by production-scale MoE models. ### Does this PR introduce _any_ user-facing change? Yes. - `oproj_tensor_parallel_size` now works on the V2 runner in graph-mode PD-decode deployments; previously it was V1-only. - New config-load rejections: `cudagraph_mode=none` / `enforce_eager`, `prefill_context_parallel_size > 1`, and — for a real split — a KV connector chain without `PreemptOffloadConnector`. - Capture-bound auto-disable with a warning instead of rejection; raise `max_cudagraph_capture_size` to use the knob at large `max_num_seqs`. - An eagerly dispatched step now fails the request with an explicit error instead of hanging; an offload failure at preemption aborts the request instead of returning it to P. ### How was this patch tested? **Unit tests**: new cases in `tests/ut/test_ascend_config.py` (graph-mode / PCP rejection, size-1 exemption, capture-bound matrix), `tests/ut/worker/test_model_runner_v2_finegrained_tp.py` (eager-step guard contract) and `tests/ut/core/test_recompute_scheduler.py` (offload-failure abort). CI on the current head: `pre-commit` (incl. mypy) and `cpu-ut` green (5718 passed, 67 skipped). **End-to-end** (single-node A3, 16 chips, PD disaggregation behind the LB proxy, 8-layer MiniMax-M3 with dummy weights — functional/consistency numbers, not real-weight performance): P = chip 0, `tp1` eager, Mooncake producer; D = chips 8-15, `dp8 × tp1`, `FULL_DECODE_ONLY` + recompute scheduler + `oproj_tensor_parallel_size=2`, Mooncake consumer; all traffic through the proxy. | Check | Result | |---|---| | config rejections | graph-mode and PD-scenario rejections fail startup with explicit messages; `--max-num-seqs 600` (600 > 512) auto-disables the knob with a warning and startup continues into graph capture | | continuous PD traffic, `oproj=2` | 100% of requests completed, guard fired 0 times, decode ran on captured FULL graphs, engine TPOT cc1/cc8 median 7.21 / 7.58 ms (990.7 tok/s at cc8) | | vs `oproj=0` baseline | TPOT +0.07 ms (~1%); outputs token-identical on 5 fixed prompts in every round | | negative control (same topology without vllm-project#17128) | the first arrival step dispatches to eager → explicit `RuntimeError` within 3.7 s, no hang | | offload chain (`MultiConnector`: Mooncake + `PreemptOffloadConnector`) | starts, graph capture OK, cc8 7.67 ms / 943 tok/s, outputs identical, CPU offload pool active on 8/8 decode ranks | | tight-KV sweeps (`gpu-memory-utilization` 0.5 / 0.4 / 0.35, 100 MB CPU pool) | 100% of requests completed, outputs identical, no crash | **Left uncovered** (for honesty): the preemption/offload runtime path (KV save/restore and the capacity fallback) was not triggered — preemption is unreachable on the 8-layer dummy model — so the abort-instead-of-return-to-P branch is covered by unit tests only; and the without-connector startup failure is a config-load check that has not been reproduced e2e yet (it was observed on the test bed while the check still warned). - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: xuchi <xuchicolson@163.com>
…17174) ### What this PR does / why we need it? Remove the obsolete vLLM 0.28.0 PCP+DP compatibility introduced by vllm-project#16853 (`8f2e3fed73328148ea603ccfe5764542fa7cbbd2`) after vllm-project#17004 upgraded the supported release to v0.29.0. Both v0.29.0 and the paired vLLM main provide native PCP-local token counting before dispatch. This is a conflict-resolved revert with targeted cleanup: - Remove the 0.28-only ParallelConfig validator shim, dispatch ContextVar/wrapper, the legacy dispatch block inside gather_batch_req_state, and PCP token-count override. - Preserve the independent v0.29.0 patch_parallel_config.py: the release still rejects PCP+DP without that patch. - Preserve vllm-project#17128's PD decode recompute logic in gather_batch_req_state and its regression tests. - Preserve subsequent changes, including nullcontext and the v0.29.0 argument gates, generic PCP count/feature tests, and slot-buffer dtype coverage. - Remove obsolete adapter-specific tests and patch documentation. Remove imports made unused by the cleanup, including vllm_version_is in pcp_manager.py; that import predates vllm-project#16853 but has no remaining references after this revert and would trigger Ruff F401. No new tests or unrelated formatting changes. The execute_model body is unindented only to remove its obsolete context wrapper. ### Does this PR introduce _any_ user-facing change? No intended behavior change on the supported v0.29.0 and fixed-main lanes. The retired v0.28.0 compatibility is removed; v0.29.0 PCP+DP configuration support remains provided by the independent validator patch. ### How was this patch tested? On rebased commit `d6b9f29e01ee2daa2d594b8c5c828e94dac0b128`: - Ruff lint/format, AST syntax, conflict-marker and git diff --check checks passed for the seven changed Python files. - Preserve the new upstream step_eplb_after import, PD decode recompute logic, and o_proj TP graph checks. Range-diff against 5034869 shows only upstream context changes; the cleanup scope is unchanged. - DCO Signed-off-by is retained. Fresh CI is pending; no local Ascend runtime tests were run on this revision. - Previous revision 5034869 passed [E2E run 35741354764](https://github.com/vllm-project/vllm-ascend/actions/runs/35741354764), including pre-commit, CPU UT, selected NPU jobs, upstream tests and ci-gate. Skipped jobs are not passes; these results do not establish validation of this rebased revision. Base: vllm-ascend main `972fcd5c974d55a7970ebe630329daa4b4049c49`. Release: vLLM v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`). vLLM main: vllm-project/vllm@84030bb - vLLM main: vllm-project/vllm@84030bb Signed-off-by: wzx0726 <278573478+wzx0726@users.noreply.github.com> Co-authored-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
…del runner v2 (vllm-project#17361) ### What this PR does / why we need it? Extends fine-grained TP on model runner v2 beyond o_proj, on top of the framework landed in vllm-project#16967 (and its prerequisite vllm-project#17128). The four knobs split into two families with different contracts: 1. **MLP TP joins the o_proj (uniform-token) contract.** The MLP exchange (gate_up `all_gather` / down `reduce_scatter` over DP-axis groups) needs the same group-uniform token count per step as o_proj, so `mlp_tensor_parallel_size > 1` joins the graph-mode / PCP / `tp == 1` / PD-scenario checks, the capture-bound auto-disable (a miss disables both knobs together — the bound is a property of the step, not of a knob), the `PreemptOffloadConnector` gate, the offload-failure abort, and the eager-step guard (renamed `_check_finegrained_tp_graph_step`). 2. **Embedding TP is capacity-based and joins none of the uniform-token checks.** Its exchange pads every step to a fixed capacity inside the op, so it is step-shape-agnostic and cannot desync; the runtime capacity already floors at `max_num_batched_tokens`, so its one real failure mode fails fast on its own. 3. **The recompute-scheduler requirement covers all four knobs.** o_proj and embedding already required it on main; mlp joins as it enters the uniform-token contract, and lmhead joins so the whole `finegrained_tp_config` stays within the one configuration validated end-to-end. The first traffic test of the no-recompute path on the PD test bed produced deterministically wrong outputs with every knob off (a pre-existing defect unrelated to this PR, tracked separately), so that path stays outside the supported set until it is fixed. ### Does this PR introduce _any_ user-facing change? Yes. - `mlp_tensor_parallel_size` now works on the V2 runner in graph-mode PD-decode deployments, with the same preconditions as `oproj_tensor_parallel_size`. - All four `finegrained_tp_config` sizes require `recompute_scheduler_enable=true` (previously only o_proj and embedding did). - At runtime, an eagerly dispatched step under o_proj/MLP TP fails the request with an explicit error (the message now names both knobs). ### How was this patch tested? **Unit tests**: one new case in `tests/ut/test_ascend_config.py` (MLP capture-bound, incl. the both-knobs-disabled-together semantics); the remaining gates are shared conditions already covered by the existing o_proj cases. CI on the current head: `pre-commit` and `cpu-ut` green. **End-to-end** — single-node A3 (16 chips), PD disaggregation behind the LB proxy; DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers (real weights, `--quantization ascend`). P = chip 0 (`tp1`, eager, Mooncake producer); D = chips 8-15 (`dp8 × tp1`, `FULL_DECODE_ONLY`, capture size 512, recompute scheduler on, `MultiConnector` = Mooncake consumer + `PreemptOffloadConnector`); every request goes through the proxy. Each round: 8-concurrent × 256-token bench × 3 (6144 tokens) plus 5 fixed greedy prompts for consistency. Positive rounds (all: 8/8 startup, 100% of bench tokens, knob reported enabled, zero guard hits, zero auto-disables, zero collective errors): | Round | Engine TPOT | Output vs baseline | |---|---|---| | baseline | 7.12 ms | — | | `mlp=2` | 7.33 ms | 1/5 prompts identical, the rest diverge at the first token (see note) | | `oproj=2 + mlp=2` | 7.24 ms | same class as `mlp=2`; bit-identical to the three-knob round | | `embedding=2` | 7.32 ms | **5/5 bit-identical** | | all three knobs | 7.14 ms | bit-identical to the `oproj=2+mlp=2` round | Consistency note: re-sharding o_proj/MLP changes the reduction order, and the truncated 6-layer test model has a near-flat output distribution, so the greedy pick can flip at the first token — the same class of difference as changing the generic TP size. The embedding exchange is pure data movement and stays bit-identical, including when stacked with the other two knobs. Negative configs (rejected or disabled at startup, no hang): | Config | Behavior | |---|---| | `mlp=2` + `--enforce-eager` | rejected at config load ("only supported in graph mode") | | `mlp=2` without `PreemptOffloadConnector` | rejected with the explicit connector requirement | | `mlp=2` without the recompute scheduler | rejected at config load in ~15 s with the recompute-scheduler requirement (the all-knobs gate above) | | capture bound 512 < largest step 600 | both knobs auto-disabled with a warning; startup completes | Base controls: the rounds above ran on the base before vllm-project#17159; since vllm-project#17159 reworked the embedding exchange, the embedding-relevant cells were re-run on the current base — baseline and `embedding=2` outputs are **bit-identical** to the pre-vllm-project#17159 results (TPOT 7.52 / 7.18 ms). A no-recompute control on the current base produced deterministically wrong outputs with **every knob off** (a pre-existing defect, tracked separately); `embedding=2` on that path is bit-identical to the no-knob run, i.e. fine-grained TP adds no effect — this is the defect behind the all-knobs recompute requirement. **Memory** (same bed): the family's fixed cost is one HCCL buffer set (~0.42 GB per rank, from vllm-project#16967); MLP saves 0.55 GiB per rank at `mlp=2` on the 6-layer bed (dense-FFN layers × 1/2), o_proj 0.109 GiB per layer per rank on a 1-byte checkpoint (vllm-project#16967). - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: xuchi <xuchicolson@163.com>
Summary
When PD decode uses
recompute_scheduler_enable=true, the last prompt-token recompute is reported as prefill by the upstream batch state. This keepshas_prefilltrue and makes a mixed decode batch missFULL_DECODE_ONLYgraphs. Reclassify only the prompt-boundary recompute on a PD consumer, then refreshhas_prefilland the uniform decode token count. This path does not require MTP and does not depend on #17032.Full
test_model_runner_v2.py: 40 passed; its one remaining Spec-PP failure is also reproducible on the unmodified baseline.