Skip to content

[Feature][MRV2] Support o_proj TP in model runner v2 (graph-mode scope) - #16967

Merged
realliujiaxu merged 2 commits into
vllm-project:mainfrom
xuchi-0808:a00333-oproj-tp-mrv2-graphmode
Sep 23, 2026
Merged

realliujiaxu merged 2 commits into
vllm-project:mainfrom
xuchi-0808:a00333-oproj-tp-mrv2-graphmode

Conversation

@xuchi-0808

@xuchi-0808 xuchi-0808 commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Enables oproj_tensor_parallel_size (fine-grained TP for the attention output projection) on model runner v2, scoped to its real deployment: a PD decode node (tp == 1, dp > 1, MoE) running a graph mode with the recompute scheduler.

The o_proj exchange (all_to_all → partial GEMM → reduce_scatter) is an all_to_all_single without split sizes: every rank must forward the same number of tokens per step, or the collectives hang silently. In graph mode the upstream dispatch already guarantees that group-uniform count, so this PR adds no runner-side padding — it adds the contract that keeps the count uniform, and guardrails that turn a silent hang into an explicit failure:

  1. Graph mode is required: cudagraph_mode == NONE (which covers enforce_eager) is rejected at config load.
  2. Capture-bound check: the knob is kept only if the largest capture size covers the largest possible step (min(max_num_batched_tokens, max_num_seqs * decode_query_len)); otherwise it is disabled with a warning and the deployment still starts.
  3. A guard at the top of NPUModelRunner.prepare_inputs — the only post-dispatch hook, reached by real steps only — raises an explicit RuntimeError if a step still dispatches to eager. Armed only for a real split (> 1).
  4. PCP (prefill_context_parallel_size > 1) is rejected: its per-rank token recount breaks the uniform step size.
  5. With > 1, the KV connector chain must carry PreemptOffloadConnector — startup fails without it — and an offload failure at preemption aborts the request instead of sending it back to P. Checked in platform._check_ascend_config because MultiConnector children re-derive the config from per-child copies that cannot see their siblings.

Existing preconditions unchanged: PD decode node (is_kv_consumer), recompute_scheduler_enable=true, tensor_parallel_size == 1, MoE.

Depends on #17128 (reclassifies the request-arrival boundary recompute as decode, so those steps stay on captured graphs); this PR is meant to land after it.

Design conclusions (aligned with the SE)

The scope and guardrails of this PR implement conclusions aligned with the SE:

  1. No MRV1 parity goal. Fine-grained TP is adapted to model runner v2 only where necessary and reasonable; the bar is expected functional correctness, not parity with MRV1.
  2. Preconditions are enforced at config load. oproj_tensor_parallel_size requires the recompute scheduler, preempt offload, a graph mode, and a pure-DP PD decode deployment (tp == 1, dp > 1).
  3. No eager fallback from an oversized step. A step whose token count exceeds the largest captured size dispatches to eager in every graph mode; a conservative config-time check verifies the capture bound covers the largest possible step and disables the knob otherwise. Other fine-grained TP knobs can follow the same pattern.
  4. No return-to-P fallback under this feature. Recompute preemption is paired with offload; the stock path sends the request back to P when the offload itself fails. This feature rejects that branch — recomputation on P cannot guarantee computation equivalence — and fails the request immediately with an explicit error instead.
  5. PreemptOffloadConnector is mandatory. Config load rejects a real split whose KV connector chain does not carry it.

Memory cost and expected savings

The memory profile belongs to the feature, not to this PR — the validation and guard paths allocate no device memory. Measured on the test bed (DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers, dp8 × tp1, graph mode, identical connector config; per rank):

Item Value Basis
Fixed cost: HCCL buffers of the o_proj TP communication group ~0.42 GB measured (steady-state device delta; the freed weight is cancelled by the larger KV pool); allocated at graph-capture time
Weight saving per layer, 1-byte (W8A8/FP8) checkpoint 0.109 GiB × (1 − 1/N) measured: 0.33 GiB @ oproj=2, 0.49 GiB @ oproj=4 on 6 layers
Full 61-layer extrapolation, W8A8/FP8, oproj=8 ~5.8 GB gross, ~5.4 GB net of the fixed cost from the measured per-layer slope
Net savings, DeepSeek-R1-W8A8, oproj=8 5.8 GB feature guide (Fine_grained_TP.md)
DeepSeek-V3/R1 BF16, oproj=8 ~11.2 GiB net from checkpoint dims (0.219 GiB per layer)

The freed weight memory returns to the KV cache, and the fixed cost is recovered after ~8 layers at oproj=2 (~4 at oproj=8) on a 1-byte quantized checkpoint — trivially met by production-scale MoE models.

Does this PR introduce any user-facing change?

Yes.

  • oproj_tensor_parallel_size now works on the V2 runner in graph-mode PD-decode deployments; previously it was V1-only.
  • New config-load rejections: cudagraph_mode=none / enforce_eager, prefill_context_parallel_size > 1, and — for a real split — a KV connector chain without PreemptOffloadConnector.
  • Capture-bound auto-disable with a warning instead of rejection; raise max_cudagraph_capture_size to use the knob at large max_num_seqs.
  • An eagerly dispatched step now fails the request with an explicit error instead of hanging; an offload failure at preemption aborts the request instead of returning it to P.

How was this patch tested?

Unit tests: new cases in tests/ut/test_ascend_config.py (graph-mode / PCP rejection, size-1 exemption, capture-bound matrix), tests/ut/worker/test_model_runner_v2_finegrained_tp.py (eager-step guard contract) and tests/ut/core/test_recompute_scheduler.py (offload-failure abort). CI on the current head: pre-commit (incl. mypy) and cpu-ut green (5718 passed, 67 skipped).

End-to-end (single-node A3, 16 chips, PD disaggregation behind the LB proxy, 8-layer MiniMax-M3 with dummy weights — functional/consistency numbers, not real-weight performance): P = chip 0, tp1 eager, Mooncake producer; D = chips 8-15, dp8 × tp1, FULL_DECODE_ONLY + recompute scheduler + oproj_tensor_parallel_size=2, Mooncake consumer; all traffic through the proxy.

Check Result
config rejections graph-mode and PD-scenario rejections fail startup with explicit messages; --max-num-seqs 600 (600 > 512) auto-disables the knob with a warning and startup continues into graph capture
continuous PD traffic, oproj=2 100% of requests completed, guard fired 0 times, decode ran on captured FULL graphs, engine TPOT cc1/cc8 median 7.21 / 7.58 ms (990.7 tok/s at cc8)
vs oproj=0 baseline TPOT +0.07 ms (~1%); outputs token-identical on 5 fixed prompts in every round
negative control (same topology without #17128) the first arrival step dispatches to eager → explicit RuntimeError within 3.7 s, no hang
offload chain (MultiConnector: Mooncake + PreemptOffloadConnector) starts, graph capture OK, cc8 7.67 ms / 943 tok/s, outputs identical, CPU offload pool active on 8/8 decode ranks
tight-KV sweeps (gpu-memory-utilization 0.5 / 0.4 / 0.35, 100 MB CPU pool) 100% of requests completed, outputs identical, no crash

Left uncovered (for honesty): the preemption/offload runtime path (KV save/restore and the capacity fallback) was not triggered — preemption is unreachable on the 8-layer dummy model — so the abort-instead-of-return-to-P branch is covered by unit tests only; and the without-connector startup failure is a config-load check that has not been reproduced e2e yet (it was observed on the test bed while the check still warned).

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request enables fine-grained tensor parallelism for the o_proj layer on the vLLM V2 model runner. The implementation focuses on graph-mode execution to ensure uniform token counts across DP ranks, which is critical for maintaining synchronization in cross-DP collectives. It introduces robust configuration validation and a runtime safety mechanism to prevent potential deadlocks, ensuring that any deviation from the required graph-mode execution results in an explicit error rather than a silent hang.

Highlights

  • o_proj TP Support: Enabled fine-grained tensor parallelism for o_proj on the V2 model runner, specifically for graph-mode PD-decode deployments.
  • Configuration Validation: Added strict config-load checks: rejects eager mode and PCP combinations, and auto-disables the feature with a warning if the cudagraph capture bound is insufficient for the max possible step.
  • Runtime Safety: Implemented a fail-fast guard in NPUModelRunner.prepare_inputs that raises a RuntimeError if an eager step is dispatched, preventing silent HCCL collective hangs.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:\n\nmarkdown\n[Ops][Feature] Add validation, fallback, and runtime guards for o_proj TP\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\nThis PR introduces validation, automatic fallback, and runtime guards for `o_proj` tensor parallelism (TP) on Ascend. Specifically:\n- It ensures `o_proj` TP is only used in graph mode and rejects combinations with Prefill Context Parallelism (PCP) at config load.\n- It automatically disables `o_proj` TP with a warning if the maximum CUDA graph capture size is insufficient to cover the largest possible step, preventing hangs in cross-DP collectives.\n- It adds a runtime guard in `NPUModelRunner` to raise an explicit error if a step unexpectedly dispatches to eager mode.\n- It updates the documentation and adds comprehensive unit tests to verify these preconditions and behaviors.\n\n### Does this PR introduce _any_ user-facing change?\nYes, it adds configuration validation that may reject invalid configurations (e.g., combining `o_proj` TP with PCP or eager mode) or print warnings and disable `o_proj` TP if the CUDA graph capture size is too small.\n\n### How was this patch tested?\nNew unit tests were added in `tests/ut/test_ascend_config.py` and `tests/ut/worker/test_model_runner_v2_finegrained_tp.py` to cover the validation logic, fallback behavior, and runtime guards.\n\n\nNo review comments were provided to assess.

@github-actions github-actions Bot added documentation Improvements or additions to documentation module:tests module:core labels Sep 20, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [Feature] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@xuchi-0808
xuchi-0808 force-pushed the a00333-oproj-tp-mrv2-graphmode branch 3 times, most recently from 8bbdefa to fe83212 Compare September 21, 2026 08:20
@github-actions github-actions Bot removed the documentation Improvements or additions to documentation label Sep 21, 2026
@xuchi-0808
xuchi-0808 force-pushed the a00333-oproj-tp-mrv2-graphmode branch 16 times, most recently from f155cb7 to 9beaa45 Compare September 22, 2026 08:48
@CXY-Katrina CXY-Katrina added the ready-precise run selected e2e test for pr label Sep 22, 2026
@xuchi-0808
xuchi-0808 force-pushed the a00333-oproj-tp-mrv2-graphmode branch from 9beaa45 to 2a06b44 Compare September 22, 2026 09:38
@xuchi-0808
xuchi-0808 force-pushed the a00333-oproj-tp-mrv2-graphmode branch from 2a06b44 to d425eed Compare September 22, 2026 09:49
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@xuchi-0808
xuchi-0808 force-pushed the a00333-oproj-tp-mrv2-graphmode branch from d425eed to 29708a4 Compare September 23, 2026 02:24
…ture-bound risk

Require a graph mode (cudagraph_mode != NONE, covering enforce_eager), reject
PCP, and keep the knob only when the largest cudagraph capture size covers the
largest possible step (min(max_num_batched_tokens, max_num_seqs *
decode_query_len)); otherwise disable it with a warning instead of rejecting
the deployment — an oversized step dispatches to eager and hangs the cross-DP
HCCL collectives.

Also require PreemptOffloadConnector in the KV connector chain (checked
driver-side, where sibling connectors are visible): a preempted request must
not return to the prefill node, whose KV recomputation loses precision.

Signed-off-by: xuchi <xuchicolson@163.com>
…recomputation

Eager dispatch keeps per-rank token counts, so an eagerly dispatched step
would hang the cross-DP HCCL collectives without an error. Raise an explicit
RuntimeError in prepare_inputs (reached by real steps only; profiling and
capture run through the dummy path).

Fail closed at runtime as well: when KV offload fails before decode-side
preemption under o_proj TP, abort the request instead of returning it to P
for recomputation.

Signed-off-by: xuchi <xuchicolson@163.com>
@xuchi-0808
xuchi-0808 force-pushed the a00333-oproj-tp-mrv2-graphmode branch from 29708a4 to bfe296a Compare September 23, 2026 03:05
@xuchi-0808

Copy link
Copy Markdown
Contributor Author

Since the last approvals the branch was only rebased onto latest main. The only content change is the conflict resolution in test_model_runner_v2_finegrained_tp.py (from #17233's test refactor — upstream's refactored test adopted, our guard test kept); everything else is byte-for-byte identical and no production code changed.

@realliujiaxu
realliujiaxu merged commit ab730ad into vllm-project:main Sep 23, 2026
30 checks passed
weijinqian0 pushed a commit that referenced this pull request Sep 23, 2026
…del runner v2 (#17361)

### What this PR does / why we need it?

Extends fine-grained TP on model runner v2 beyond o_proj, on top of the
framework landed in #16967 (and its prerequisite #17128). The four knobs
split into two families with different contracts:

1. **MLP TP joins the o_proj (uniform-token) contract.** The MLP
exchange (gate_up `all_gather` / down `reduce_scatter` over DP-axis
groups) needs the same group-uniform token count per step as o_proj, so
`mlp_tensor_parallel_size > 1` joins the graph-mode / PCP / `tp == 1` /
PD-scenario checks, the capture-bound auto-disable (a miss disables both
knobs together — the bound is a property of the step, not of a knob),
the `PreemptOffloadConnector` gate, the offload-failure abort, and the
eager-step guard (renamed `_check_finegrained_tp_graph_step`).
2. **Embedding TP is capacity-based and joins none of the uniform-token
checks.** Its exchange pads every step to a fixed capacity inside the
op, so it is step-shape-agnostic and cannot desync; the runtime capacity
already floors at `max_num_batched_tokens`, so its one real failure mode
fails fast on its own.
3. **The recompute-scheduler requirement covers all four knobs.** o_proj
and embedding already required it on main; mlp joins as it enters the
uniform-token contract, and lmhead joins so the whole
`finegrained_tp_config` stays within the one configuration validated
end-to-end. The first traffic test of the no-recompute path on the PD
test bed produced deterministically wrong outputs with every knob off (a
pre-existing defect unrelated to this PR, tracked separately), so that
path stays outside the supported set until it is fixed.

### Does this PR introduce _any_ user-facing change?

Yes.

- `mlp_tensor_parallel_size` now works on the V2 runner in graph-mode
PD-decode deployments, with the same preconditions as
`oproj_tensor_parallel_size`.
- All four `finegrained_tp_config` sizes require
`recompute_scheduler_enable=true` (previously only o_proj and embedding
did).
- At runtime, an eagerly dispatched step under o_proj/MLP TP fails the
request with an explicit error (the message now names both knobs).

### How was this patch tested?

**Unit tests**: one new case in `tests/ut/test_ascend_config.py` (MLP
capture-bound, incl. the both-knobs-disabled-together semantics); the
remaining gates are shared conditions already covered by the existing
o_proj cases. CI on the current head: `pre-commit` and `cpu-ut` green.

**End-to-end** — single-node A3 (16 chips), PD disaggregation behind the
LB proxy; DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers (real
weights, `--quantization ascend`). P = chip 0 (`tp1`, eager, Mooncake
producer); D = chips 8-15 (`dp8 × tp1`, `FULL_DECODE_ONLY`, capture size
512, recompute scheduler on, `MultiConnector` = Mooncake consumer +
`PreemptOffloadConnector`); every request goes through the proxy. Each
round: 8-concurrent × 256-token bench × 3 (6144 tokens) plus 5 fixed
greedy prompts for consistency.

Positive rounds (all: 8/8 startup, 100% of bench tokens, knob reported
enabled, zero guard hits, zero auto-disables, zero collective errors):

| Round | Engine TPOT | Output vs baseline |
|---|---|---|
| baseline | 7.12 ms | — |
| `mlp=2` | 7.33 ms | 1/5 prompts identical, the rest diverge at the
first token (see note) |
| `oproj=2 + mlp=2` | 7.24 ms | same class as `mlp=2`; bit-identical to
the three-knob round |
| `embedding=2` | 7.32 ms | **5/5 bit-identical** |
| all three knobs | 7.14 ms | bit-identical to the `oproj=2+mlp=2` round
|

Consistency note: re-sharding o_proj/MLP changes the reduction order,
and the truncated 6-layer test model has a near-flat output
distribution, so the greedy pick can flip at the first token — the same
class of difference as changing the generic TP size. The embedding
exchange is pure data movement and stays bit-identical, including when
stacked with the other two knobs.

Negative configs (rejected or disabled at startup, no hang):

| Config | Behavior |
|---|---|
| `mlp=2` + `--enforce-eager` | rejected at config load ("only supported
in graph mode") |
| `mlp=2` without `PreemptOffloadConnector` | rejected with the explicit
connector requirement |
| `mlp=2` without the recompute scheduler | rejected at config load in
~15 s with the recompute-scheduler requirement (the all-knobs gate
above) |
| capture bound 512 < largest step 600 | both knobs auto-disabled with a
warning; startup completes |

Base controls: the rounds above ran on the base before #17159; since
#17159 reworked the embedding exchange, the embedding-relevant cells
were re-run on the current base — baseline and `embedding=2` outputs are
**bit-identical** to the pre-#17159 results (TPOT 7.52 / 7.18 ms). A
no-recompute control on the current base produced deterministically
wrong outputs with **every knob off** (a pre-existing defect, tracked
separately); `embedding=2` on that path is bit-identical to the no-knob
run, i.e. fine-grained TP adds no effect — this is the defect behind the
all-knobs recompute requirement.

**Memory** (same bed): the family's fixed cost is one HCCL buffer set
(~0.42 GB per rank, from #16967); MLP saves 0.55 GiB per rank at `mlp=2`
on the 6-layer bed (dense-FFN layers × 1/2), o_proj 0.109 GiB per layer
per rank on a 1-byte checkpoint (#16967).

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: xuchi <xuchicolson@163.com>
tangdafu pushed a commit to tangdafu/vllm-ascend that referenced this pull request Sep 23, 2026
…e) (vllm-project#16967)

### What this PR does / why we need it?

Enables `oproj_tensor_parallel_size` (fine-grained TP for the attention
output projection) on model runner v2, scoped to its real deployment: a
PD decode node (`tp == 1`, `dp > 1`, MoE) running a graph mode with the
recompute scheduler.

The o_proj exchange (`all_to_all` → partial GEMM → `reduce_scatter`) is
an `all_to_all_single` without split sizes: every rank must forward the
same number of tokens per step, or the collectives hang silently. In
graph mode the upstream dispatch already guarantees that group-uniform
count, so this PR adds no runner-side padding — it adds the contract
that keeps the count uniform, and guardrails that turn a silent hang
into an explicit failure:

1. Graph mode is required: `cudagraph_mode == NONE` (which covers
`enforce_eager`) is rejected at config load.
2. Capture-bound check: the knob is kept only if the largest capture
size covers the largest possible step (`min(max_num_batched_tokens,
max_num_seqs * decode_query_len)`); otherwise it is disabled with a
warning and the deployment still starts.
3. A guard at the top of `NPUModelRunner.prepare_inputs` — the only
post-dispatch hook, reached by real steps only — raises an explicit
`RuntimeError` if a step still dispatches to eager. Armed only for a
real split (`> 1`).
4. PCP (`prefill_context_parallel_size > 1`) is rejected: its per-rank
token recount breaks the uniform step size.
5. With `> 1`, the KV connector chain must carry
`PreemptOffloadConnector` — startup fails without it — and an offload
failure at preemption aborts the request instead of sending it back to
P. Checked in `platform._check_ascend_config` because MultiConnector
children re-derive the config from per-child copies that cannot see
their siblings.

Existing preconditions unchanged: PD decode node (`is_kv_consumer`),
`recompute_scheduler_enable=true`, `tensor_parallel_size == 1`, MoE.

Depends on vllm-project#17128 (reclassifies the request-arrival boundary recompute
as decode, so those steps stay on captured graphs); this PR is meant to
land after it.

### Design conclusions (aligned with the SE)

The scope and guardrails of this PR implement conclusions aligned with
the SE:

1. **No MRV1 parity goal.** Fine-grained TP is adapted to model runner
v2 only where necessary and reasonable; the bar is expected functional
correctness, not parity with MRV1.
2. **Preconditions are enforced at config load.**
`oproj_tensor_parallel_size` requires the recompute scheduler, preempt
offload, a graph mode, and a pure-DP PD decode deployment (`tp == 1`,
`dp > 1`).
3. **No eager fallback from an oversized step.** A step whose token
count exceeds the largest captured size dispatches to eager in every
graph mode; a conservative config-time check verifies the capture bound
covers the largest possible step and disables the knob otherwise. Other
fine-grained TP knobs can follow the same pattern.
4. **No return-to-P fallback under this feature.** Recompute preemption
is paired with offload; the stock path sends the request back to P when
the offload itself fails. This feature rejects that branch —
recomputation on P cannot guarantee computation equivalence — and fails
the request immediately with an explicit error instead.
5. **`PreemptOffloadConnector` is mandatory.** Config load rejects a
real split whose KV connector chain does not carry it.

### Memory cost and expected savings

The memory profile belongs to the feature, not to this PR — the
validation and guard paths allocate no device memory. Measured on the
test bed (DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers, `dp8
× tp1`, graph mode, identical connector config; per rank):

| Item | Value | Basis |
|---|---|---|
| Fixed cost: HCCL buffers of the o_proj TP communication group | ~0.42
GB | measured (steady-state device delta; the freed weight is cancelled
by the larger KV pool); allocated at graph-capture time |
| Weight saving per layer, 1-byte (W8A8/FP8) checkpoint | 0.109 GiB × (1
− 1/N) | measured: 0.33 GiB @ `oproj=2`, 0.49 GiB @ `oproj=4` on 6
layers |
| Full 61-layer extrapolation, W8A8/FP8, `oproj=8` | ~5.8 GB gross, ~5.4
GB net of the fixed cost | from the measured per-layer slope |
| Net savings, DeepSeek-R1-W8A8, `oproj=8` | 5.8 GB | feature guide
(`Fine_grained_TP.md`) |
| DeepSeek-V3/R1 BF16, `oproj=8` | ~11.2 GiB net | from checkpoint dims
(0.219 GiB per layer) |

The freed weight memory returns to the KV cache, and the fixed cost is
recovered after ~8 layers at `oproj=2` (~4 at `oproj=8`) on a 1-byte
quantized checkpoint — trivially met by production-scale MoE models.

### Does this PR introduce _any_ user-facing change?

Yes.

- `oproj_tensor_parallel_size` now works on the V2 runner in graph-mode
PD-decode deployments; previously it was V1-only.
- New config-load rejections: `cudagraph_mode=none` / `enforce_eager`,
`prefill_context_parallel_size > 1`, and — for a real split — a KV
connector chain without `PreemptOffloadConnector`.
- Capture-bound auto-disable with a warning instead of rejection; raise
`max_cudagraph_capture_size` to use the knob at large `max_num_seqs`.
- An eagerly dispatched step now fails the request with an explicit
error instead of hanging; an offload failure at preemption aborts the
request instead of returning it to P.

### How was this patch tested?

**Unit tests**: new cases in `tests/ut/test_ascend_config.py`
(graph-mode / PCP rejection, size-1 exemption, capture-bound matrix),
`tests/ut/worker/test_model_runner_v2_finegrained_tp.py` (eager-step
guard contract) and `tests/ut/core/test_recompute_scheduler.py`
(offload-failure abort). CI on the current head: `pre-commit` (incl.
mypy) and `cpu-ut` green (5718 passed, 67 skipped).

**End-to-end** (single-node A3, 16 chips, PD disaggregation behind the
LB proxy, 8-layer MiniMax-M3 with dummy weights — functional/consistency
numbers, not real-weight performance): P = chip 0, `tp1` eager, Mooncake
producer; D = chips 8-15, `dp8 × tp1`, `FULL_DECODE_ONLY` + recompute
scheduler + `oproj_tensor_parallel_size=2`, Mooncake consumer; all
traffic through the proxy.

| Check | Result |
|---|---|
| config rejections | graph-mode and PD-scenario rejections fail startup
with explicit messages; `--max-num-seqs 600` (600 > 512) auto-disables
the knob with a warning and startup continues into graph capture |
| continuous PD traffic, `oproj=2` | 100% of requests completed, guard
fired 0 times, decode ran on captured FULL graphs, engine TPOT cc1/cc8
median 7.21 / 7.58 ms (990.7 tok/s at cc8) |
| vs `oproj=0` baseline | TPOT +0.07 ms (~1%); outputs token-identical
on 5 fixed prompts in every round |
| negative control (same topology without vllm-project#17128) | the first arrival
step dispatches to eager → explicit `RuntimeError` within 3.7 s, no hang
|
| offload chain (`MultiConnector`: Mooncake + `PreemptOffloadConnector`)
| starts, graph capture OK, cc8 7.67 ms / 943 tok/s, outputs identical,
CPU offload pool active on 8/8 decode ranks |
| tight-KV sweeps (`gpu-memory-utilization` 0.5 / 0.4 / 0.35, 100 MB CPU
pool) | 100% of requests completed, outputs identical, no crash |

**Left uncovered** (for honesty): the preemption/offload runtime path
(KV save/restore and the capacity fallback) was not triggered —
preemption is unreachable on the 8-layer dummy model — so the
abort-instead-of-return-to-P branch is covered by unit tests only; and
the without-connector startup failure is a config-load check that has
not been reproduced e2e yet (it was observed on the test bed while the
check still warned).

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: xuchi <xuchicolson@163.com>
hw-qianlingfeng pushed a commit to hw-qianlingfeng/vllm-ascend that referenced this pull request Sep 27, 2026
…del runner v2 (vllm-project#17361)

### What this PR does / why we need it?

Extends fine-grained TP on model runner v2 beyond o_proj, on top of the
framework landed in vllm-project#16967 (and its prerequisite vllm-project#17128). The four knobs
split into two families with different contracts:

1. **MLP TP joins the o_proj (uniform-token) contract.** The MLP
exchange (gate_up `all_gather` / down `reduce_scatter` over DP-axis
groups) needs the same group-uniform token count per step as o_proj, so
`mlp_tensor_parallel_size > 1` joins the graph-mode / PCP / `tp == 1` /
PD-scenario checks, the capture-bound auto-disable (a miss disables both
knobs together — the bound is a property of the step, not of a knob),
the `PreemptOffloadConnector` gate, the offload-failure abort, and the
eager-step guard (renamed `_check_finegrained_tp_graph_step`).
2. **Embedding TP is capacity-based and joins none of the uniform-token
checks.** Its exchange pads every step to a fixed capacity inside the
op, so it is step-shape-agnostic and cannot desync; the runtime capacity
already floors at `max_num_batched_tokens`, so its one real failure mode
fails fast on its own.
3. **The recompute-scheduler requirement covers all four knobs.** o_proj
and embedding already required it on main; mlp joins as it enters the
uniform-token contract, and lmhead joins so the whole
`finegrained_tp_config` stays within the one configuration validated
end-to-end. The first traffic test of the no-recompute path on the PD
test bed produced deterministically wrong outputs with every knob off (a
pre-existing defect unrelated to this PR, tracked separately), so that
path stays outside the supported set until it is fixed.

### Does this PR introduce _any_ user-facing change?

Yes.

- `mlp_tensor_parallel_size` now works on the V2 runner in graph-mode
PD-decode deployments, with the same preconditions as
`oproj_tensor_parallel_size`.
- All four `finegrained_tp_config` sizes require
`recompute_scheduler_enable=true` (previously only o_proj and embedding
did).
- At runtime, an eagerly dispatched step under o_proj/MLP TP fails the
request with an explicit error (the message now names both knobs).

### How was this patch tested?

**Unit tests**: one new case in `tests/ut/test_ascend_config.py` (MLP
capture-bound, incl. the both-knobs-disabled-together semantics); the
remaining gates are shared conditions already covered by the existing
o_proj cases. CI on the current head: `pre-commit` and `cpu-ut` green.

**End-to-end** — single-node A3 (16 chips), PD disaggregation behind the
LB proxy; DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers (real
weights, `--quantization ascend`). P = chip 0 (`tp1`, eager, Mooncake
producer); D = chips 8-15 (`dp8 × tp1`, `FULL_DECODE_ONLY`, capture size
512, recompute scheduler on, `MultiConnector` = Mooncake consumer +
`PreemptOffloadConnector`); every request goes through the proxy. Each
round: 8-concurrent × 256-token bench × 3 (6144 tokens) plus 5 fixed
greedy prompts for consistency.

Positive rounds (all: 8/8 startup, 100% of bench tokens, knob reported
enabled, zero guard hits, zero auto-disables, zero collective errors):

| Round | Engine TPOT | Output vs baseline |
|---|---|---|
| baseline | 7.12 ms | — |
| `mlp=2` | 7.33 ms | 1/5 prompts identical, the rest diverge at the
first token (see note) |
| `oproj=2 + mlp=2` | 7.24 ms | same class as `mlp=2`; bit-identical to
the three-knob round |
| `embedding=2` | 7.32 ms | **5/5 bit-identical** |
| all three knobs | 7.14 ms | bit-identical to the `oproj=2+mlp=2` round
|

Consistency note: re-sharding o_proj/MLP changes the reduction order,
and the truncated 6-layer test model has a near-flat output
distribution, so the greedy pick can flip at the first token — the same
class of difference as changing the generic TP size. The embedding
exchange is pure data movement and stays bit-identical, including when
stacked with the other two knobs.

Negative configs (rejected or disabled at startup, no hang):

| Config | Behavior |
|---|---|
| `mlp=2` + `--enforce-eager` | rejected at config load ("only supported
in graph mode") |
| `mlp=2` without `PreemptOffloadConnector` | rejected with the explicit
connector requirement |
| `mlp=2` without the recompute scheduler | rejected at config load in
~15 s with the recompute-scheduler requirement (the all-knobs gate
above) |
| capture bound 512 < largest step 600 | both knobs auto-disabled with a
warning; startup completes |

Base controls: the rounds above ran on the base before vllm-project#17159; since
vllm-project#17159 reworked the embedding exchange, the embedding-relevant cells
were re-run on the current base — baseline and `embedding=2` outputs are
**bit-identical** to the pre-vllm-project#17159 results (TPOT 7.52 / 7.18 ms). A
no-recompute control on the current base produced deterministically
wrong outputs with **every knob off** (a pre-existing defect, tracked
separately); `embedding=2` on that path is bit-identical to the no-knob
run, i.e. fine-grained TP adds no effect — this is the defect behind the
all-knobs recompute requirement.

**Memory** (same bed): the family's fixed cost is one HCCL buffer set
(~0.42 GB per rank, from vllm-project#16967); MLP saves 0.55 GiB per rank at `mlp=2`
on the 6-layer bed (dense-FFN layers × 1/2), o_proj 0.109 GiB per layer
per rank on a 1-byte checkpoint (vllm-project#16967).

- vLLM main:
vllm-project/vllm@84030bb

---------

Signed-off-by: xuchi <xuchicolson@163.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants