Skip to content

[Bugfix][MRV2][P/D] Preserve decode graph with recompute scheduler - #17128

Merged
realliujiaxu merged 1 commit into
vllm-project:mainfrom
iKeybot-code:fix/recompute-scheduler-no-mtp-decode-graph
Sep 22, 2026
Merged

realliujiaxu merged 1 commit into
vllm-project:mainfrom
iKeybot-code:fix/recompute-scheduler-no-mtp-decode-graph

Conversation

@iKeybot-code

@iKeybot-code iKeybot-code commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Summary

When PD decode uses recompute_scheduler_enable=true, the last prompt-token recompute is reported as prefill by the upstream batch state. This keeps has_prefill true and makes a mixed decode batch miss FULL_DECODE_ONLY graphs. Reclassify only the prompt-boundary recompute on a PD consumer, then refresh has_prefill and the uniform decode token count. This path does not require MTP and does not depend on #17032.

Validation Baseline Fixed
Targeted UT — 4 passed
Qwen3-32B-W8A8 requests 128/128 128/128
Aggregate output throughput 121.84 tok/s 137.59 tok/s (+12.93%)
Mean join decode latency 3.2335 s 2.3539 s (-27.20%)

Full test_model_runner_v2.py: 40 passed; its one remaining Spec-PP failure is also reproducible on the unmodified baseline.

Signed-off-by: likailong <likailong5@huawei.com>
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses a bug where PD decode requests using the recompute scheduler were incorrectly identified as prefill, causing performance degradation by missing optimized decode graphs. By refining the reclassification of prompt-boundary recompute tokens, the system now correctly identifies these as decode operations, leading to significant throughput and latency improvements.

Highlights

  • PD Decode Recompute Logic: Introduced logic to correctly reclassify prompt-boundary recompute tokens as decode tokens when the recompute scheduler is enabled, preventing them from being incorrectly flagged as prefill.
  • Batch State Refresh: Updated the batch state and uniform decode token count calculation to ensure mixed decode batches can properly utilize FULL_DECODE_ONLY graphs.
  • Performance Improvements: Achieved a 12.93% increase in aggregate output throughput and a 27.20% reduction in mean join decode latency for Qwen3-32B-W8A8 requests.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Support reclassifying PD tail in mixed decode batch for recompute scheduler

Suggested PR Summary:

### What this PR does / why we need it?
This PR updates the `gather_batch_req_state` method in the Ascend V2 `GPUModelRunner` to support the PD decode recompute scheduler. When enabled, it identifies and reclassifies prefill requests that are actually part of a decode recompute as decode requests.

Feedback: A review comment suggests avoiding in-place mutation of `batch_state.is_prefilling_np` to prevent potential runtime errors or side-effects, recommending copying the array and using `_replace` instead.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Added unit tests in `tests/ut/worker/test_model_runner_v2.py` covering the reclassification logic, including mixed decode batches, non-matching prefills, and multi-token decode queries.

Comment thread vllm_ascend/worker/v2/model_runner.py
@yiz-liu yiz-liu added the ready-precise run selected e2e test for pr label Sep 21, 2026
@iKeybot-code iKeybot-code changed the title [Bugfix] Preserve decode graph with recompute scheduler [Bugfix][MRV2] Preserve decode graph with recompute scheduler Sep 21, 2026
@iKeybot-code iKeybot-code changed the title [Bugfix][MRV2] Preserve decode graph with recompute scheduler [Bugfix][MRV2][P/D] Preserve decode graph with recompute scheduler Sep 21, 2026
@realliujiaxu
realliujiaxu merged commit dfb7219 into vllm-project:main Sep 22, 2026
43 of 44 checks passed
realliujiaxu pushed a commit that referenced this pull request Sep 23, 2026
…e) (#16967)

### What this PR does / why we need it?

Enables `oproj_tensor_parallel_size` (fine-grained TP for the attention
output projection) on model runner v2, scoped to its real deployment: a
PD decode node (`tp == 1`, `dp > 1`, MoE) running a graph mode with the
recompute scheduler.

The o_proj exchange (`all_to_all` → partial GEMM → `reduce_scatter`) is
an `all_to_all_single` without split sizes: every rank must forward the
same number of tokens per step, or the collectives hang silently. In
graph mode the upstream dispatch already guarantees that group-uniform
count, so this PR adds no runner-side padding — it adds the contract
that keeps the count uniform, and guardrails that turn a silent hang
into an explicit failure:

1. Graph mode is required: `cudagraph_mode == NONE` (which covers
`enforce_eager`) is rejected at config load.
2. Capture-bound check: the knob is kept only if the largest capture
size covers the largest possible step (`min(max_num_batched_tokens,
max_num_seqs * decode_query_len)`); otherwise it is disabled with a
warning and the deployment still starts.
3. A guard at the top of `NPUModelRunner.prepare_inputs` — the only
post-dispatch hook, reached by real steps only — raises an explicit
`RuntimeError` if a step still dispatches to eager. Armed only for a
real split (`> 1`).
4. PCP (`prefill_context_parallel_size > 1`) is rejected: its per-rank
token recount breaks the uniform step size.
5. With `> 1`, the KV connector chain must carry
`PreemptOffloadConnector` — startup fails without it — and an offload
failure at preemption aborts the request instead of sending it back to
P. Checked in `platform._check_ascend_config` because MultiConnector
children re-derive the config from per-child copies that cannot see
their siblings.

Existing preconditions unchanged: PD decode node (`is_kv_consumer`),
`recompute_scheduler_enable=true`, `tensor_parallel_size == 1`, MoE.

Depends on #17128 (reclassifies the request-arrival boundary recompute
as decode, so those steps stay on captured graphs); this PR is meant to
land after it.

### Design conclusions (aligned with the SE)

The scope and guardrails of this PR implement conclusions aligned with
the SE:

1. **No MRV1 parity goal.** Fine-grained TP is adapted to model runner
v2 only where necessary and reasonable; the bar is expected functional
correctness, not parity with MRV1.
2. **Preconditions are enforced at config load.**
`oproj_tensor_parallel_size` requires the recompute scheduler, preempt
offload, a graph mode, and a pure-DP PD decode deployment (`tp == 1`,
`dp > 1`).
3. **No eager fallback from an oversized step.** A step whose token
count exceeds the largest captured size dispatches to eager in every
graph mode; a conservative config-time check verifies the capture bound
covers the largest possible step and disables the knob otherwise. Other
fine-grained TP knobs can follow the same pattern.
4. **No return-to-P fallback under this feature.** Recompute preemption
is paired with offload; the stock path sends the request back to P when
the offload itself fails. This feature rejects that branch —
recomputation on P cannot guarantee computation equivalence — and fails
the request immediately with an explicit error instead.
5. **`PreemptOffloadConnector` is mandatory.** Config load rejects a
real split whose KV connector chain does not carry it.

### Memory cost and expected savings

The memory profile belongs to the feature, not to this PR — the
validation and guard paths allocate no device memory. Measured on the
test bed (DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers, `dp8
× tp1`, graph mode, identical connector config; per rank):

| Item | Value | Basis |
|---|---|---|
| Fixed cost: HCCL buffers of the o_proj TP communication group | ~0.42
GB | measured (steady-state device delta; the freed weight is cancelled
by the larger KV pool); allocated at graph-capture time |
| Weight saving per layer, 1-byte (W8A8/FP8) checkpoint | 0.109 GiB × (1
− 1/N) | measured: 0.33 GiB @ `oproj=2`, 0.49 GiB @ `oproj=4` on 6
layers |
| Full 61-layer extrapolation, W8A8/FP8, `oproj=8` | ~5.8 GB gross, ~5.4
GB net of the fixed cost | from the measured per-layer slope |
| Net savings, DeepSeek-R1-W8A8, `oproj=8` | 5.8 GB | feature guide
(`Fine_grained_TP.md`) |
| DeepSeek-V3/R1 BF16, `oproj=8` | ~11.2 GiB net | from checkpoint dims
(0.219 GiB per layer) |

The freed weight memory returns to the KV cache, and the fixed cost is
recovered after ~8 layers at `oproj=2` (~4 at `oproj=8`) on a 1-byte
quantized checkpoint — trivially met by production-scale MoE models.

### Does this PR introduce _any_ user-facing change?

Yes.

- `oproj_tensor_parallel_size` now works on the V2 runner in graph-mode
PD-decode deployments; previously it was V1-only.
- New config-load rejections: `cudagraph_mode=none` / `enforce_eager`,
`prefill_context_parallel_size > 1`, and — for a real split — a KV
connector chain without `PreemptOffloadConnector`.
- Capture-bound auto-disable with a warning instead of rejection; raise
`max_cudagraph_capture_size` to use the knob at large `max_num_seqs`.
- An eagerly dispatched step now fails the request with an explicit
error instead of hanging; an offload failure at preemption aborts the
request instead of returning it to P.

### How was this patch tested?

**Unit tests**: new cases in `tests/ut/test_ascend_config.py`
(graph-mode / PCP rejection, size-1 exemption, capture-bound matrix),
`tests/ut/worker/test_model_runner_v2_finegrained_tp.py` (eager-step
guard contract) and `tests/ut/core/test_recompute_scheduler.py`
(offload-failure abort). CI on the current head: `pre-commit` (incl.
mypy) and `cpu-ut` green (5718 passed, 67 skipped).

**End-to-end** (single-node A3, 16 chips, PD disaggregation behind the
LB proxy, 8-layer MiniMax-M3 with dummy weights — functional/consistency
numbers, not real-weight performance): P = chip 0, `tp1` eager, Mooncake
producer; D = chips 8-15, `dp8 × tp1`, `FULL_DECODE_ONLY` + recompute
scheduler + `oproj_tensor_parallel_size=2`, Mooncake consumer; all
traffic through the proxy.

| Check | Result |
|---|---|
| config rejections | graph-mode and PD-scenario rejections fail startup
with explicit messages; `--max-num-seqs 600` (600 > 512) auto-disables
the knob with a warning and startup continues into graph capture |
| continuous PD traffic, `oproj=2` | 100% of requests completed, guard
fired 0 times, decode ran on captured FULL graphs, engine TPOT cc1/cc8
median 7.21 / 7.58 ms (990.7 tok/s at cc8) |
| vs `oproj=0` baseline | TPOT +0.07 ms (~1%); outputs token-identical
on 5 fixed prompts in every round |
| negative control (same topology without #17128) | the first arrival
step dispatches to eager → explicit `RuntimeError` within 3.7 s, no hang
|
| offload chain (`MultiConnector`: Mooncake + `PreemptOffloadConnector`)
| starts, graph capture OK, cc8 7.67 ms / 943 tok/s, outputs identical,
CPU offload pool active on 8/8 decode ranks |
| tight-KV sweeps (`gpu-memory-utilization` 0.5 / 0.4 / 0.35, 100 MB CPU
pool) | 100% of requests completed, outputs identical, no crash |

**Left uncovered** (for honesty): the preemption/offload runtime path
(KV save/restore and the capacity fallback) was not triggered —
preemption is unreachable on the 8-layer dummy model — so the
abort-instead-of-return-to-P branch is covered by unit tests only; and
the without-connector startup failure is a config-load check that has
not been reproduced e2e yet (it was observed on the test bed while the
check still warned).

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: xuchi <xuchicolson@163.com>
realliujiaxu pushed a commit that referenced this pull request Sep 23, 2026
### What this PR does / why we need it?

Remove the obsolete vLLM 0.28.0 PCP+DP compatibility introduced by
#16853 (`8f2e3fed73328148ea603ccfe5764542fa7cbbd2`) after #17004
upgraded the supported release to v0.29.0. Both v0.29.0 and the paired
vLLM main provide native PCP-local token counting before dispatch.

This is a conflict-resolved revert with targeted cleanup:

- Remove the 0.28-only ParallelConfig validator shim, dispatch
ContextVar/wrapper, the legacy dispatch block inside
gather_batch_req_state, and PCP token-count override.
- Preserve the independent v0.29.0 patch_parallel_config.py: the release
still rejects PCP+DP without that patch.
- Preserve #17128's PD decode recompute logic in gather_batch_req_state
and its regression tests.
- Preserve subsequent changes, including nullcontext and the v0.29.0
argument gates, generic PCP count/feature tests, and slot-buffer dtype
coverage.
- Remove obsolete adapter-specific tests and patch documentation. Remove
imports made unused by the cleanup, including vllm_version_is in
pcp_manager.py; that import predates #16853 but has no remaining
references after this revert and would trigger Ruff F401.

No new tests or unrelated formatting changes. The execute_model body is
unindented only to remove its obsolete context wrapper.

### Does this PR introduce _any_ user-facing change?

No intended behavior change on the supported v0.29.0 and fixed-main
lanes. The retired v0.28.0 compatibility is removed; v0.29.0 PCP+DP
configuration support remains provided by the independent validator
patch.

### How was this patch tested?

On rebased commit `d6b9f29e01ee2daa2d594b8c5c828e94dac0b128`:

- Ruff lint/format, AST syntax, conflict-marker and git diff --check
checks passed for the seven changed Python files.
- Preserve the new upstream step_eplb_after import, PD decode recompute
logic, and o_proj TP graph checks. Range-diff against 5034869 shows
only upstream context changes; the cleanup scope is unchanged.
- DCO Signed-off-by is retained. Fresh CI is pending; no local Ascend
runtime tests were run on this revision.
- Previous revision 5034869 passed [E2E run
35741354764](https://github.com/vllm-project/vllm-ascend/actions/runs/35741354764),
including pre-commit, CPU UT, selected NPU jobs, upstream tests and
ci-gate. Skipped jobs are not passes; these results do not establish
validation of this rebased revision.

Base: vllm-ascend main `972fcd5c974d55a7970ebe630329daa4b4049c49`.
Release: vLLM v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`).
vLLM main:
vllm-project/vllm@84030bb



- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
Co-authored-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
weijinqian0 pushed a commit that referenced this pull request Sep 23, 2026
…del runner v2 (#17361)

### What this PR does / why we need it?

Extends fine-grained TP on model runner v2 beyond o_proj, on top of the
framework landed in #16967 (and its prerequisite #17128). The four knobs
split into two families with different contracts:

1. **MLP TP joins the o_proj (uniform-token) contract.** The MLP
exchange (gate_up `all_gather` / down `reduce_scatter` over DP-axis
groups) needs the same group-uniform token count per step as o_proj, so
`mlp_tensor_parallel_size > 1` joins the graph-mode / PCP / `tp == 1` /
PD-scenario checks, the capture-bound auto-disable (a miss disables both
knobs together — the bound is a property of the step, not of a knob),
the `PreemptOffloadConnector` gate, the offload-failure abort, and the
eager-step guard (renamed `_check_finegrained_tp_graph_step`).
2. **Embedding TP is capacity-based and joins none of the uniform-token
checks.** Its exchange pads every step to a fixed capacity inside the
op, so it is step-shape-agnostic and cannot desync; the runtime capacity
already floors at `max_num_batched_tokens`, so its one real failure mode
fails fast on its own.
3. **The recompute-scheduler requirement covers all four knobs.** o_proj
and embedding already required it on main; mlp joins as it enters the
uniform-token contract, and lmhead joins so the whole
`finegrained_tp_config` stays within the one configuration validated
end-to-end. The first traffic test of the no-recompute path on the PD
test bed produced deterministically wrong outputs with every knob off (a
pre-existing defect unrelated to this PR, tracked separately), so that
path stays outside the supported set until it is fixed.

### Does this PR introduce _any_ user-facing change?

Yes.

- `mlp_tensor_parallel_size` now works on the V2 runner in graph-mode
PD-decode deployments, with the same preconditions as
`oproj_tensor_parallel_size`.
- All four `finegrained_tp_config` sizes require
`recompute_scheduler_enable=true` (previously only o_proj and embedding
did).
- At runtime, an eagerly dispatched step under o_proj/MLP TP fails the
request with an explicit error (the message now names both knobs).

### How was this patch tested?

**Unit tests**: one new case in `tests/ut/test_ascend_config.py` (MLP
capture-bound, incl. the both-knobs-disabled-together semantics); the
remaining gates are shared conditions already covered by the existing
o_proj cases. CI on the current head: `pre-commit` and `cpu-ut` green.

**End-to-end** — single-node A3 (16 chips), PD disaggregation behind the
LB proxy; DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers (real
weights, `--quantization ascend`). P = chip 0 (`tp1`, eager, Mooncake
producer); D = chips 8-15 (`dp8 × tp1`, `FULL_DECODE_ONLY`, capture size
512, recompute scheduler on, `MultiConnector` = Mooncake consumer +
`PreemptOffloadConnector`); every request goes through the proxy. Each
round: 8-concurrent × 256-token bench × 3 (6144 tokens) plus 5 fixed
greedy prompts for consistency.

Positive rounds (all: 8/8 startup, 100% of bench tokens, knob reported
enabled, zero guard hits, zero auto-disables, zero collective errors):

| Round | Engine TPOT | Output vs baseline |
|---|---|---|
| baseline | 7.12 ms | — |
| `mlp=2` | 7.33 ms | 1/5 prompts identical, the rest diverge at the
first token (see note) |
| `oproj=2 + mlp=2` | 7.24 ms | same class as `mlp=2`; bit-identical to
the three-knob round |
| `embedding=2` | 7.32 ms | **5/5 bit-identical** |
| all three knobs | 7.14 ms | bit-identical to the `oproj=2+mlp=2` round
|

Consistency note: re-sharding o_proj/MLP changes the reduction order,
and the truncated 6-layer test model has a near-flat output
distribution, so the greedy pick can flip at the first token — the same
class of difference as changing the generic TP size. The embedding
exchange is pure data movement and stays bit-identical, including when
stacked with the other two knobs.

Negative configs (rejected or disabled at startup, no hang):

| Config | Behavior |
|---|---|
| `mlp=2` + `--enforce-eager` | rejected at config load ("only supported
in graph mode") |
| `mlp=2` without `PreemptOffloadConnector` | rejected with the explicit
connector requirement |
| `mlp=2` without the recompute scheduler | rejected at config load in
~15 s with the recompute-scheduler requirement (the all-knobs gate
above) |
| capture bound 512 < largest step 600 | both knobs auto-disabled with a
warning; startup completes |

Base controls: the rounds above ran on the base before #17159; since
#17159 reworked the embedding exchange, the embedding-relevant cells
were re-run on the current base — baseline and `embedding=2` outputs are
**bit-identical** to the pre-#17159 results (TPOT 7.52 / 7.18 ms). A
no-recompute control on the current base produced deterministically
wrong outputs with **every knob off** (a pre-existing defect, tracked
separately); `embedding=2` on that path is bit-identical to the no-knob
run, i.e. fine-grained TP adds no effect — this is the defect behind the
all-knobs recompute requirement.

**Memory** (same bed): the family's fixed cost is one HCCL buffer set
(~0.42 GB per rank, from #16967); MLP saves 0.55 GiB per rank at `mlp=2`
on the 6-layer bed (dense-FFN layers × 1/2), o_proj 0.109 GiB per layer
per rank on a 1-byte checkpoint (#16967).

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: xuchi <xuchicolson@163.com>
tangdafu pushed a commit to tangdafu/vllm-ascend that referenced this pull request Sep 23, 2026
…llm-project#17128)

## Summary

When PD decode uses `recompute_scheduler_enable=true`, the last
prompt-token recompute is reported as prefill by the upstream batch
state. This keeps `has_prefill` true and makes a mixed decode batch miss
`FULL_DECODE_ONLY` graphs. Reclassify only the prompt-boundary recompute
on a PD consumer, then refresh `has_prefill` and the uniform decode
token count. This path does not require MTP and does not depend on
vllm-project#17032.

| Validation | Baseline | Fixed |
|---|---:|---:|
| Targeted UT | — | 4 passed |
| Qwen3-32B-W8A8 requests | 128/128 | 128/128 |
| Aggregate output throughput | 121.84 tok/s | 137.59 tok/s (+12.93%) |
| Mean join decode latency | 3.2335 s | 2.3539 s (-27.20%) |

Full `test_model_runner_v2.py`: 40 passed; its one remaining Spec-PP
failure is also reproducible on the unmodified baseline.
- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: likailong <likailong5@huawei.com>
tangdafu pushed a commit to tangdafu/vllm-ascend that referenced this pull request Sep 23, 2026
…e) (vllm-project#16967)

### What this PR does / why we need it?

Enables `oproj_tensor_parallel_size` (fine-grained TP for the attention
output projection) on model runner v2, scoped to its real deployment: a
PD decode node (`tp == 1`, `dp > 1`, MoE) running a graph mode with the
recompute scheduler.

The o_proj exchange (`all_to_all` → partial GEMM → `reduce_scatter`) is
an `all_to_all_single` without split sizes: every rank must forward the
same number of tokens per step, or the collectives hang silently. In
graph mode the upstream dispatch already guarantees that group-uniform
count, so this PR adds no runner-side padding — it adds the contract
that keeps the count uniform, and guardrails that turn a silent hang
into an explicit failure:

1. Graph mode is required: `cudagraph_mode == NONE` (which covers
`enforce_eager`) is rejected at config load.
2. Capture-bound check: the knob is kept only if the largest capture
size covers the largest possible step (`min(max_num_batched_tokens,
max_num_seqs * decode_query_len)`); otherwise it is disabled with a
warning and the deployment still starts.
3. A guard at the top of `NPUModelRunner.prepare_inputs` — the only
post-dispatch hook, reached by real steps only — raises an explicit
`RuntimeError` if a step still dispatches to eager. Armed only for a
real split (`> 1`).
4. PCP (`prefill_context_parallel_size > 1`) is rejected: its per-rank
token recount breaks the uniform step size.
5. With `> 1`, the KV connector chain must carry
`PreemptOffloadConnector` — startup fails without it — and an offload
failure at preemption aborts the request instead of sending it back to
P. Checked in `platform._check_ascend_config` because MultiConnector
children re-derive the config from per-child copies that cannot see
their siblings.

Existing preconditions unchanged: PD decode node (`is_kv_consumer`),
`recompute_scheduler_enable=true`, `tensor_parallel_size == 1`, MoE.

Depends on vllm-project#17128 (reclassifies the request-arrival boundary recompute
as decode, so those steps stay on captured graphs); this PR is meant to
land after it.

### Design conclusions (aligned with the SE)

The scope and guardrails of this PR implement conclusions aligned with
the SE:

1. **No MRV1 parity goal.** Fine-grained TP is adapted to model runner
v2 only where necessary and reasonable; the bar is expected functional
correctness, not parity with MRV1.
2. **Preconditions are enforced at config load.**
`oproj_tensor_parallel_size` requires the recompute scheduler, preempt
offload, a graph mode, and a pure-DP PD decode deployment (`tp == 1`,
`dp > 1`).
3. **No eager fallback from an oversized step.** A step whose token
count exceeds the largest captured size dispatches to eager in every
graph mode; a conservative config-time check verifies the capture bound
covers the largest possible step and disables the knob otherwise. Other
fine-grained TP knobs can follow the same pattern.
4. **No return-to-P fallback under this feature.** Recompute preemption
is paired with offload; the stock path sends the request back to P when
the offload itself fails. This feature rejects that branch —
recomputation on P cannot guarantee computation equivalence — and fails
the request immediately with an explicit error instead.
5. **`PreemptOffloadConnector` is mandatory.** Config load rejects a
real split whose KV connector chain does not carry it.

### Memory cost and expected savings

The memory profile belongs to the feature, not to this PR — the
validation and guard paths allocate no device memory. Measured on the
test bed (DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers, `dp8
× tp1`, graph mode, identical connector config; per rank):

| Item | Value | Basis |
|---|---|---|
| Fixed cost: HCCL buffers of the o_proj TP communication group | ~0.42
GB | measured (steady-state device delta; the freed weight is cancelled
by the larger KV pool); allocated at graph-capture time |
| Weight saving per layer, 1-byte (W8A8/FP8) checkpoint | 0.109 GiB × (1
− 1/N) | measured: 0.33 GiB @ `oproj=2`, 0.49 GiB @ `oproj=4` on 6
layers |
| Full 61-layer extrapolation, W8A8/FP8, `oproj=8` | ~5.8 GB gross, ~5.4
GB net of the fixed cost | from the measured per-layer slope |
| Net savings, DeepSeek-R1-W8A8, `oproj=8` | 5.8 GB | feature guide
(`Fine_grained_TP.md`) |
| DeepSeek-V3/R1 BF16, `oproj=8` | ~11.2 GiB net | from checkpoint dims
(0.219 GiB per layer) |

The freed weight memory returns to the KV cache, and the fixed cost is
recovered after ~8 layers at `oproj=2` (~4 at `oproj=8`) on a 1-byte
quantized checkpoint — trivially met by production-scale MoE models.

### Does this PR introduce _any_ user-facing change?

Yes.

- `oproj_tensor_parallel_size` now works on the V2 runner in graph-mode
PD-decode deployments; previously it was V1-only.
- New config-load rejections: `cudagraph_mode=none` / `enforce_eager`,
`prefill_context_parallel_size > 1`, and — for a real split — a KV
connector chain without `PreemptOffloadConnector`.
- Capture-bound auto-disable with a warning instead of rejection; raise
`max_cudagraph_capture_size` to use the knob at large `max_num_seqs`.
- An eagerly dispatched step now fails the request with an explicit
error instead of hanging; an offload failure at preemption aborts the
request instead of returning it to P.

### How was this patch tested?

**Unit tests**: new cases in `tests/ut/test_ascend_config.py`
(graph-mode / PCP rejection, size-1 exemption, capture-bound matrix),
`tests/ut/worker/test_model_runner_v2_finegrained_tp.py` (eager-step
guard contract) and `tests/ut/core/test_recompute_scheduler.py`
(offload-failure abort). CI on the current head: `pre-commit` (incl.
mypy) and `cpu-ut` green (5718 passed, 67 skipped).

**End-to-end** (single-node A3, 16 chips, PD disaggregation behind the
LB proxy, 8-layer MiniMax-M3 with dummy weights — functional/consistency
numbers, not real-weight performance): P = chip 0, `tp1` eager, Mooncake
producer; D = chips 8-15, `dp8 × tp1`, `FULL_DECODE_ONLY` + recompute
scheduler + `oproj_tensor_parallel_size=2`, Mooncake consumer; all
traffic through the proxy.

| Check | Result |
|---|---|
| config rejections | graph-mode and PD-scenario rejections fail startup
with explicit messages; `--max-num-seqs 600` (600 > 512) auto-disables
the knob with a warning and startup continues into graph capture |
| continuous PD traffic, `oproj=2` | 100% of requests completed, guard
fired 0 times, decode ran on captured FULL graphs, engine TPOT cc1/cc8
median 7.21 / 7.58 ms (990.7 tok/s at cc8) |
| vs `oproj=0` baseline | TPOT +0.07 ms (~1%); outputs token-identical
on 5 fixed prompts in every round |
| negative control (same topology without vllm-project#17128) | the first arrival
step dispatches to eager → explicit `RuntimeError` within 3.7 s, no hang
|
| offload chain (`MultiConnector`: Mooncake + `PreemptOffloadConnector`)
| starts, graph capture OK, cc8 7.67 ms / 943 tok/s, outputs identical,
CPU offload pool active on 8/8 decode ranks |
| tight-KV sweeps (`gpu-memory-utilization` 0.5 / 0.4 / 0.35, 100 MB CPU
pool) | 100% of requests completed, outputs identical, no crash |

**Left uncovered** (for honesty): the preemption/offload runtime path
(KV save/restore and the capacity fallback) was not triggered —
preemption is unreachable on the 8-layer dummy model — so the
abort-instead-of-return-to-P branch is covered by unit tests only; and
the without-connector startup failure is a config-load check that has
not been reproduced e2e yet (it was observed on the test bed while the
check still warned).

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: xuchi <xuchicolson@163.com>
tangdafu pushed a commit to tangdafu/vllm-ascend that referenced this pull request Sep 23, 2026
…17174)

### What this PR does / why we need it?

Remove the obsolete vLLM 0.28.0 PCP+DP compatibility introduced by
vllm-project#16853 (`8f2e3fed73328148ea603ccfe5764542fa7cbbd2`) after vllm-project#17004
upgraded the supported release to v0.29.0. Both v0.29.0 and the paired
vLLM main provide native PCP-local token counting before dispatch.

This is a conflict-resolved revert with targeted cleanup:

- Remove the 0.28-only ParallelConfig validator shim, dispatch
ContextVar/wrapper, the legacy dispatch block inside
gather_batch_req_state, and PCP token-count override.
- Preserve the independent v0.29.0 patch_parallel_config.py: the release
still rejects PCP+DP without that patch.
- Preserve vllm-project#17128's PD decode recompute logic in gather_batch_req_state
and its regression tests.
- Preserve subsequent changes, including nullcontext and the v0.29.0
argument gates, generic PCP count/feature tests, and slot-buffer dtype
coverage.
- Remove obsolete adapter-specific tests and patch documentation. Remove
imports made unused by the cleanup, including vllm_version_is in
pcp_manager.py; that import predates vllm-project#16853 but has no remaining
references after this revert and would trigger Ruff F401.

No new tests or unrelated formatting changes. The execute_model body is
unindented only to remove its obsolete context wrapper.

### Does this PR introduce _any_ user-facing change?

No intended behavior change on the supported v0.29.0 and fixed-main
lanes. The retired v0.28.0 compatibility is removed; v0.29.0 PCP+DP
configuration support remains provided by the independent validator
patch.

### How was this patch tested?

On rebased commit `d6b9f29e01ee2daa2d594b8c5c828e94dac0b128`:

- Ruff lint/format, AST syntax, conflict-marker and git diff --check
checks passed for the seven changed Python files.
- Preserve the new upstream step_eplb_after import, PD decode recompute
logic, and o_proj TP graph checks. Range-diff against 5034869 shows
only upstream context changes; the cleanup scope is unchanged.
- DCO Signed-off-by is retained. Fresh CI is pending; no local Ascend
runtime tests were run on this revision.
- Previous revision 5034869 passed [E2E run
35741354764](https://github.com/vllm-project/vllm-ascend/actions/runs/35741354764),
including pre-commit, CPU UT, selected NPU jobs, upstream tests and
ci-gate. Skipped jobs are not passes; these results do not establish
validation of this rebased revision.

Base: vllm-ascend main `972fcd5c974d55a7970ebe630329daa4b4049c49`.
Release: vLLM v0.29.0 (`98dff2a81d747d1dba01a47f939f48c3526d4206`).
vLLM main:
vllm-project/vllm@84030bb



- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
Co-authored-by: wzx0726 <278573478+wzx0726@users.noreply.github.com>
hw-qianlingfeng pushed a commit to hw-qianlingfeng/vllm-ascend that referenced this pull request Sep 27, 2026
…del runner v2 (vllm-project#17361)

### What this PR does / why we need it?

Extends fine-grained TP on model runner v2 beyond o_proj, on top of the
framework landed in vllm-project#16967 (and its prerequisite vllm-project#17128). The four knobs
split into two families with different contracts:

1. **MLP TP joins the o_proj (uniform-token) contract.** The MLP
exchange (gate_up `all_gather` / down `reduce_scatter` over DP-axis
groups) needs the same group-uniform token count per step as o_proj, so
`mlp_tensor_parallel_size > 1` joins the graph-mode / PCP / `tp == 1` /
PD-scenario checks, the capture-bound auto-disable (a miss disables both
knobs together — the bound is a property of the step, not of a knob),
the `PreemptOffloadConnector` gate, the offload-failure abort, and the
eager-step guard (renamed `_check_finegrained_tp_graph_step`).
2. **Embedding TP is capacity-based and joins none of the uniform-token
checks.** Its exchange pads every step to a fixed capacity inside the
op, so it is step-shape-agnostic and cannot desync; the runtime capacity
already floors at `max_num_batched_tokens`, so its one real failure mode
fails fast on its own.
3. **The recompute-scheduler requirement covers all four knobs.** o_proj
and embedding already required it on main; mlp joins as it enters the
uniform-token contract, and lmhead joins so the whole
`finegrained_tp_config` stays within the one configuration validated
end-to-end. The first traffic test of the no-recompute path on the PD
test bed produced deterministically wrong outputs with every knob off (a
pre-existing defect unrelated to this PR, tracked separately), so that
path stays outside the supported set until it is fixed.

### Does this PR introduce _any_ user-facing change?

Yes.

- `mlp_tensor_parallel_size` now works on the V2 runner in graph-mode
PD-decode deployments, with the same preconditions as
`oproj_tensor_parallel_size`.
- All four `finegrained_tp_config` sizes require
`recompute_scheduler_enable=true` (previously only o_proj and embedding
did).
- At runtime, an eagerly dispatched step under o_proj/MLP TP fails the
request with an explicit error (the message now names both knobs).

### How was this patch tested?

**Unit tests**: one new case in `tests/ut/test_ascend_config.py` (MLP
capture-bound, incl. the both-knobs-disabled-together semantics); the
remaining gates are shared conditions already covered by the existing
o_proj cases. CI on the current head: `pre-commit` and `cpu-ut` green.

**End-to-end** — single-node A3 (16 chips), PD disaggregation behind the
LB proxy; DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers (real
weights, `--quantization ascend`). P = chip 0 (`tp1`, eager, Mooncake
producer); D = chips 8-15 (`dp8 × tp1`, `FULL_DECODE_ONLY`, capture size
512, recompute scheduler on, `MultiConnector` = Mooncake consumer +
`PreemptOffloadConnector`); every request goes through the proxy. Each
round: 8-concurrent × 256-token bench × 3 (6144 tokens) plus 5 fixed
greedy prompts for consistency.

Positive rounds (all: 8/8 startup, 100% of bench tokens, knob reported
enabled, zero guard hits, zero auto-disables, zero collective errors):

| Round | Engine TPOT | Output vs baseline |
|---|---|---|
| baseline | 7.12 ms | — |
| `mlp=2` | 7.33 ms | 1/5 prompts identical, the rest diverge at the
first token (see note) |
| `oproj=2 + mlp=2` | 7.24 ms | same class as `mlp=2`; bit-identical to
the three-knob round |
| `embedding=2` | 7.32 ms | **5/5 bit-identical** |
| all three knobs | 7.14 ms | bit-identical to the `oproj=2+mlp=2` round
|

Consistency note: re-sharding o_proj/MLP changes the reduction order,
and the truncated 6-layer test model has a near-flat output
distribution, so the greedy pick can flip at the first token — the same
class of difference as changing the generic TP size. The embedding
exchange is pure data movement and stays bit-identical, including when
stacked with the other two knobs.

Negative configs (rejected or disabled at startup, no hang):

| Config | Behavior |
|---|---|
| `mlp=2` + `--enforce-eager` | rejected at config load ("only supported
in graph mode") |
| `mlp=2` without `PreemptOffloadConnector` | rejected with the explicit
connector requirement |
| `mlp=2` without the recompute scheduler | rejected at config load in
~15 s with the recompute-scheduler requirement (the all-knobs gate
above) |
| capture bound 512 < largest step 600 | both knobs auto-disabled with a
warning; startup completes |

Base controls: the rounds above ran on the base before vllm-project#17159; since
vllm-project#17159 reworked the embedding exchange, the embedding-relevant cells
were re-run on the current base — baseline and `embedding=2` outputs are
**bit-identical** to the pre-vllm-project#17159 results (TPOT 7.52 / 7.18 ms). A
no-recompute control on the current base produced deterministically
wrong outputs with **every knob off** (a pre-existing defect, tracked
separately); `embedding=2` on that path is bit-identical to the no-knob
run, i.e. fine-grained TP adds no effect — this is the defect behind the
all-knobs recompute requirement.

**Memory** (same bed): the family's fixed cost is one HCCL buffer set
(~0.42 GB per rank, from vllm-project#16967); MLP saves 0.55 GiB per rank at `mlp=2`
on the 6-layer bed (dense-FFN layers × 1/2), o_proj 0.109 GiB per layer
per rank on a 1-byte checkpoint (vllm-project#16967).

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: xuchi <xuchicolson@163.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants