[Feature][MRV2] Support o_proj TP in model runner v2 (graph-mode scope) - #16967
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request enables fine-grained tensor parallelism for the o_proj layer on the vLLM V2 model runner. The implementation focuses on graph-mode execution to ensure uniform token counts across DP ranks, which is critical for maintaining synchronization in cross-DP collectives. It introduces robust configuration validation and a runtime safety mechanism to prevent potential deadlocks, ensuring that any deviation from the required graph-mode execution results in an explicit error rather than a silent hang. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:\n\nmarkdown\n[Ops][Feature] Add validation, fallback, and runtime guards for o_proj TP\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\nThis PR introduces validation, automatic fallback, and runtime guards for `o_proj` tensor parallelism (TP) on Ascend. Specifically:\n- It ensures `o_proj` TP is only used in graph mode and rejects combinations with Prefill Context Parallelism (PCP) at config load.\n- It automatically disables `o_proj` TP with a warning if the maximum CUDA graph capture size is insufficient to cover the largest possible step, preventing hangs in cross-DP collectives.\n- It adds a runtime guard in `NPUModelRunner` to raise an explicit error if a step unexpectedly dispatches to eager mode.\n- It updates the documentation and adds comprehensive unit tests to verify these preconditions and behaviors.\n\n### Does this PR introduce _any_ user-facing change?\nYes, it adds configuration validation that may reject invalid configurations (e.g., combining `o_proj` TP with PCP or eager mode) or print warnings and disable `o_proj` TP if the CUDA graph capture size is too small.\n\n### How was this patch tested?\nNew unit tests were added in `tests/ut/test_ascend_config.py` and `tests/ut/worker/test_model_runner_v2_finegrained_tp.py` to cover the validation logic, fallback behavior, and runtime guards.\n\n\nNo review comments were provided to assess.
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. Tip 💡 Consider Linking a Related Issue or RFCYour PR title contains the [Feature] tag, indicating a bug fix or new feature. Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:
🙏 Thanks for helping us keep the project well-organized! |
8bbdefa to
fe83212
Compare
f155cb7 to
9beaa45
Compare
9beaa45 to
2a06b44
Compare
2a06b44 to
d425eed
Compare
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
d425eed to
29708a4
Compare
…ture-bound risk Require a graph mode (cudagraph_mode != NONE, covering enforce_eager), reject PCP, and keep the knob only when the largest cudagraph capture size covers the largest possible step (min(max_num_batched_tokens, max_num_seqs * decode_query_len)); otherwise disable it with a warning instead of rejecting the deployment — an oversized step dispatches to eager and hangs the cross-DP HCCL collectives. Also require PreemptOffloadConnector in the KV connector chain (checked driver-side, where sibling connectors are visible): a preempted request must not return to the prefill node, whose KV recomputation loses precision. Signed-off-by: xuchi <xuchicolson@163.com>
…recomputation Eager dispatch keeps per-rank token counts, so an eagerly dispatched step would hang the cross-DP HCCL collectives without an error. Raise an explicit RuntimeError in prepare_inputs (reached by real steps only; profiling and capture run through the dummy path). Fail closed at runtime as well: when KV offload fails before decode-side preemption under o_proj TP, abort the request instead of returning it to P for recomputation. Signed-off-by: xuchi <xuchicolson@163.com>
29708a4 to
bfe296a
Compare
|
Since the last approvals the branch was only rebased onto latest main. The only content change is the conflict resolution in test_model_runner_v2_finegrained_tp.py (from #17233's test refactor — upstream's refactored test adopted, our guard test kept); everything else is byte-for-byte identical and no production code changed. |
…del runner v2 (#17361) ### What this PR does / why we need it? Extends fine-grained TP on model runner v2 beyond o_proj, on top of the framework landed in #16967 (and its prerequisite #17128). The four knobs split into two families with different contracts: 1. **MLP TP joins the o_proj (uniform-token) contract.** The MLP exchange (gate_up `all_gather` / down `reduce_scatter` over DP-axis groups) needs the same group-uniform token count per step as o_proj, so `mlp_tensor_parallel_size > 1` joins the graph-mode / PCP / `tp == 1` / PD-scenario checks, the capture-bound auto-disable (a miss disables both knobs together — the bound is a property of the step, not of a knob), the `PreemptOffloadConnector` gate, the offload-failure abort, and the eager-step guard (renamed `_check_finegrained_tp_graph_step`). 2. **Embedding TP is capacity-based and joins none of the uniform-token checks.** Its exchange pads every step to a fixed capacity inside the op, so it is step-shape-agnostic and cannot desync; the runtime capacity already floors at `max_num_batched_tokens`, so its one real failure mode fails fast on its own. 3. **The recompute-scheduler requirement covers all four knobs.** o_proj and embedding already required it on main; mlp joins as it enters the uniform-token contract, and lmhead joins so the whole `finegrained_tp_config` stays within the one configuration validated end-to-end. The first traffic test of the no-recompute path on the PD test bed produced deterministically wrong outputs with every knob off (a pre-existing defect unrelated to this PR, tracked separately), so that path stays outside the supported set until it is fixed. ### Does this PR introduce _any_ user-facing change? Yes. - `mlp_tensor_parallel_size` now works on the V2 runner in graph-mode PD-decode deployments, with the same preconditions as `oproj_tensor_parallel_size`. - All four `finegrained_tp_config` sizes require `recompute_scheduler_enable=true` (previously only o_proj and embedding did). - At runtime, an eagerly dispatched step under o_proj/MLP TP fails the request with an explicit error (the message now names both knobs). ### How was this patch tested? **Unit tests**: one new case in `tests/ut/test_ascend_config.py` (MLP capture-bound, incl. the both-knobs-disabled-together semantics); the remaining gates are shared conditions already covered by the existing o_proj cases. CI on the current head: `pre-commit` and `cpu-ut` green. **End-to-end** — single-node A3 (16 chips), PD disaggregation behind the LB proxy; DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers (real weights, `--quantization ascend`). P = chip 0 (`tp1`, eager, Mooncake producer); D = chips 8-15 (`dp8 × tp1`, `FULL_DECODE_ONLY`, capture size 512, recompute scheduler on, `MultiConnector` = Mooncake consumer + `PreemptOffloadConnector`); every request goes through the proxy. Each round: 8-concurrent × 256-token bench × 3 (6144 tokens) plus 5 fixed greedy prompts for consistency. Positive rounds (all: 8/8 startup, 100% of bench tokens, knob reported enabled, zero guard hits, zero auto-disables, zero collective errors): | Round | Engine TPOT | Output vs baseline | |---|---|---| | baseline | 7.12 ms | — | | `mlp=2` | 7.33 ms | 1/5 prompts identical, the rest diverge at the first token (see note) | | `oproj=2 + mlp=2` | 7.24 ms | same class as `mlp=2`; bit-identical to the three-knob round | | `embedding=2` | 7.32 ms | **5/5 bit-identical** | | all three knobs | 7.14 ms | bit-identical to the `oproj=2+mlp=2` round | Consistency note: re-sharding o_proj/MLP changes the reduction order, and the truncated 6-layer test model has a near-flat output distribution, so the greedy pick can flip at the first token — the same class of difference as changing the generic TP size. The embedding exchange is pure data movement and stays bit-identical, including when stacked with the other two knobs. Negative configs (rejected or disabled at startup, no hang): | Config | Behavior | |---|---| | `mlp=2` + `--enforce-eager` | rejected at config load ("only supported in graph mode") | | `mlp=2` without `PreemptOffloadConnector` | rejected with the explicit connector requirement | | `mlp=2` without the recompute scheduler | rejected at config load in ~15 s with the recompute-scheduler requirement (the all-knobs gate above) | | capture bound 512 < largest step 600 | both knobs auto-disabled with a warning; startup completes | Base controls: the rounds above ran on the base before #17159; since #17159 reworked the embedding exchange, the embedding-relevant cells were re-run on the current base — baseline and `embedding=2` outputs are **bit-identical** to the pre-#17159 results (TPOT 7.52 / 7.18 ms). A no-recompute control on the current base produced deterministically wrong outputs with **every knob off** (a pre-existing defect, tracked separately); `embedding=2` on that path is bit-identical to the no-knob run, i.e. fine-grained TP adds no effect — this is the defect behind the all-knobs recompute requirement. **Memory** (same bed): the family's fixed cost is one HCCL buffer set (~0.42 GB per rank, from #16967); MLP saves 0.55 GiB per rank at `mlp=2` on the 6-layer bed (dense-FFN layers × 1/2), o_proj 0.109 GiB per layer per rank on a 1-byte checkpoint (#16967). - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: xuchi <xuchicolson@163.com>
…e) (vllm-project#16967) ### What this PR does / why we need it? Enables `oproj_tensor_parallel_size` (fine-grained TP for the attention output projection) on model runner v2, scoped to its real deployment: a PD decode node (`tp == 1`, `dp > 1`, MoE) running a graph mode with the recompute scheduler. The o_proj exchange (`all_to_all` → partial GEMM → `reduce_scatter`) is an `all_to_all_single` without split sizes: every rank must forward the same number of tokens per step, or the collectives hang silently. In graph mode the upstream dispatch already guarantees that group-uniform count, so this PR adds no runner-side padding — it adds the contract that keeps the count uniform, and guardrails that turn a silent hang into an explicit failure: 1. Graph mode is required: `cudagraph_mode == NONE` (which covers `enforce_eager`) is rejected at config load. 2. Capture-bound check: the knob is kept only if the largest capture size covers the largest possible step (`min(max_num_batched_tokens, max_num_seqs * decode_query_len)`); otherwise it is disabled with a warning and the deployment still starts. 3. A guard at the top of `NPUModelRunner.prepare_inputs` — the only post-dispatch hook, reached by real steps only — raises an explicit `RuntimeError` if a step still dispatches to eager. Armed only for a real split (`> 1`). 4. PCP (`prefill_context_parallel_size > 1`) is rejected: its per-rank token recount breaks the uniform step size. 5. With `> 1`, the KV connector chain must carry `PreemptOffloadConnector` — startup fails without it — and an offload failure at preemption aborts the request instead of sending it back to P. Checked in `platform._check_ascend_config` because MultiConnector children re-derive the config from per-child copies that cannot see their siblings. Existing preconditions unchanged: PD decode node (`is_kv_consumer`), `recompute_scheduler_enable=true`, `tensor_parallel_size == 1`, MoE. Depends on vllm-project#17128 (reclassifies the request-arrival boundary recompute as decode, so those steps stay on captured graphs); this PR is meant to land after it. ### Design conclusions (aligned with the SE) The scope and guardrails of this PR implement conclusions aligned with the SE: 1. **No MRV1 parity goal.** Fine-grained TP is adapted to model runner v2 only where necessary and reasonable; the bar is expected functional correctness, not parity with MRV1. 2. **Preconditions are enforced at config load.** `oproj_tensor_parallel_size` requires the recompute scheduler, preempt offload, a graph mode, and a pure-DP PD decode deployment (`tp == 1`, `dp > 1`). 3. **No eager fallback from an oversized step.** A step whose token count exceeds the largest captured size dispatches to eager in every graph mode; a conservative config-time check verifies the capture bound covers the largest possible step and disables the knob otherwise. Other fine-grained TP knobs can follow the same pattern. 4. **No return-to-P fallback under this feature.** Recompute preemption is paired with offload; the stock path sends the request back to P when the offload itself fails. This feature rejects that branch — recomputation on P cannot guarantee computation equivalence — and fails the request immediately with an explicit error instead. 5. **`PreemptOffloadConnector` is mandatory.** Config load rejects a real split whose KV connector chain does not carry it. ### Memory cost and expected savings The memory profile belongs to the feature, not to this PR — the validation and guard paths allocate no device memory. Measured on the test bed (DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers, `dp8 × tp1`, graph mode, identical connector config; per rank): | Item | Value | Basis | |---|---|---| | Fixed cost: HCCL buffers of the o_proj TP communication group | ~0.42 GB | measured (steady-state device delta; the freed weight is cancelled by the larger KV pool); allocated at graph-capture time | | Weight saving per layer, 1-byte (W8A8/FP8) checkpoint | 0.109 GiB × (1 − 1/N) | measured: 0.33 GiB @ `oproj=2`, 0.49 GiB @ `oproj=4` on 6 layers | | Full 61-layer extrapolation, W8A8/FP8, `oproj=8` | ~5.8 GB gross, ~5.4 GB net of the fixed cost | from the measured per-layer slope | | Net savings, DeepSeek-R1-W8A8, `oproj=8` | 5.8 GB | feature guide (`Fine_grained_TP.md`) | | DeepSeek-V3/R1 BF16, `oproj=8` | ~11.2 GiB net | from checkpoint dims (0.219 GiB per layer) | The freed weight memory returns to the KV cache, and the fixed cost is recovered after ~8 layers at `oproj=2` (~4 at `oproj=8`) on a 1-byte quantized checkpoint — trivially met by production-scale MoE models. ### Does this PR introduce _any_ user-facing change? Yes. - `oproj_tensor_parallel_size` now works on the V2 runner in graph-mode PD-decode deployments; previously it was V1-only. - New config-load rejections: `cudagraph_mode=none` / `enforce_eager`, `prefill_context_parallel_size > 1`, and — for a real split — a KV connector chain without `PreemptOffloadConnector`. - Capture-bound auto-disable with a warning instead of rejection; raise `max_cudagraph_capture_size` to use the knob at large `max_num_seqs`. - An eagerly dispatched step now fails the request with an explicit error instead of hanging; an offload failure at preemption aborts the request instead of returning it to P. ### How was this patch tested? **Unit tests**: new cases in `tests/ut/test_ascend_config.py` (graph-mode / PCP rejection, size-1 exemption, capture-bound matrix), `tests/ut/worker/test_model_runner_v2_finegrained_tp.py` (eager-step guard contract) and `tests/ut/core/test_recompute_scheduler.py` (offload-failure abort). CI on the current head: `pre-commit` (incl. mypy) and `cpu-ut` green (5718 passed, 67 skipped). **End-to-end** (single-node A3, 16 chips, PD disaggregation behind the LB proxy, 8-layer MiniMax-M3 with dummy weights — functional/consistency numbers, not real-weight performance): P = chip 0, `tp1` eager, Mooncake producer; D = chips 8-15, `dp8 × tp1`, `FULL_DECODE_ONLY` + recompute scheduler + `oproj_tensor_parallel_size=2`, Mooncake consumer; all traffic through the proxy. | Check | Result | |---|---| | config rejections | graph-mode and PD-scenario rejections fail startup with explicit messages; `--max-num-seqs 600` (600 > 512) auto-disables the knob with a warning and startup continues into graph capture | | continuous PD traffic, `oproj=2` | 100% of requests completed, guard fired 0 times, decode ran on captured FULL graphs, engine TPOT cc1/cc8 median 7.21 / 7.58 ms (990.7 tok/s at cc8) | | vs `oproj=0` baseline | TPOT +0.07 ms (~1%); outputs token-identical on 5 fixed prompts in every round | | negative control (same topology without vllm-project#17128) | the first arrival step dispatches to eager → explicit `RuntimeError` within 3.7 s, no hang | | offload chain (`MultiConnector`: Mooncake + `PreemptOffloadConnector`) | starts, graph capture OK, cc8 7.67 ms / 943 tok/s, outputs identical, CPU offload pool active on 8/8 decode ranks | | tight-KV sweeps (`gpu-memory-utilization` 0.5 / 0.4 / 0.35, 100 MB CPU pool) | 100% of requests completed, outputs identical, no crash | **Left uncovered** (for honesty): the preemption/offload runtime path (KV save/restore and the capacity fallback) was not triggered — preemption is unreachable on the 8-layer dummy model — so the abort-instead-of-return-to-P branch is covered by unit tests only; and the without-connector startup failure is a config-load check that has not been reproduced e2e yet (it was observed on the test bed while the check still warned). - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: xuchi <xuchicolson@163.com>
…del runner v2 (vllm-project#17361) ### What this PR does / why we need it? Extends fine-grained TP on model runner v2 beyond o_proj, on top of the framework landed in vllm-project#16967 (and its prerequisite vllm-project#17128). The four knobs split into two families with different contracts: 1. **MLP TP joins the o_proj (uniform-token) contract.** The MLP exchange (gate_up `all_gather` / down `reduce_scatter` over DP-axis groups) needs the same group-uniform token count per step as o_proj, so `mlp_tensor_parallel_size > 1` joins the graph-mode / PCP / `tp == 1` / PD-scenario checks, the capture-bound auto-disable (a miss disables both knobs together — the bound is a property of the step, not of a knob), the `PreemptOffloadConnector` gate, the offload-failure abort, and the eager-step guard (renamed `_check_finegrained_tp_graph_step`). 2. **Embedding TP is capacity-based and joins none of the uniform-token checks.** Its exchange pads every step to a fixed capacity inside the op, so it is step-shape-agnostic and cannot desync; the runtime capacity already floors at `max_num_batched_tokens`, so its one real failure mode fails fast on its own. 3. **The recompute-scheduler requirement covers all four knobs.** o_proj and embedding already required it on main; mlp joins as it enters the uniform-token contract, and lmhead joins so the whole `finegrained_tp_config` stays within the one configuration validated end-to-end. The first traffic test of the no-recompute path on the PD test bed produced deterministically wrong outputs with every knob off (a pre-existing defect unrelated to this PR, tracked separately), so that path stays outside the supported set until it is fixed. ### Does this PR introduce _any_ user-facing change? Yes. - `mlp_tensor_parallel_size` now works on the V2 runner in graph-mode PD-decode deployments, with the same preconditions as `oproj_tensor_parallel_size`. - All four `finegrained_tp_config` sizes require `recompute_scheduler_enable=true` (previously only o_proj and embedding did). - At runtime, an eagerly dispatched step under o_proj/MLP TP fails the request with an explicit error (the message now names both knobs). ### How was this patch tested? **Unit tests**: one new case in `tests/ut/test_ascend_config.py` (MLP capture-bound, incl. the both-knobs-disabled-together semantics); the remaining gates are shared conditions already covered by the existing o_proj cases. CI on the current head: `pre-commit` and `cpu-ut` green. **End-to-end** — single-node A3 (16 chips), PD disaggregation behind the LB proxy; DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers (real weights, `--quantization ascend`). P = chip 0 (`tp1`, eager, Mooncake producer); D = chips 8-15 (`dp8 × tp1`, `FULL_DECODE_ONLY`, capture size 512, recompute scheduler on, `MultiConnector` = Mooncake consumer + `PreemptOffloadConnector`); every request goes through the proxy. Each round: 8-concurrent × 256-token bench × 3 (6144 tokens) plus 5 fixed greedy prompts for consistency. Positive rounds (all: 8/8 startup, 100% of bench tokens, knob reported enabled, zero guard hits, zero auto-disables, zero collective errors): | Round | Engine TPOT | Output vs baseline | |---|---|---| | baseline | 7.12 ms | — | | `mlp=2` | 7.33 ms | 1/5 prompts identical, the rest diverge at the first token (see note) | | `oproj=2 + mlp=2` | 7.24 ms | same class as `mlp=2`; bit-identical to the three-knob round | | `embedding=2` | 7.32 ms | **5/5 bit-identical** | | all three knobs | 7.14 ms | bit-identical to the `oproj=2+mlp=2` round | Consistency note: re-sharding o_proj/MLP changes the reduction order, and the truncated 6-layer test model has a near-flat output distribution, so the greedy pick can flip at the first token — the same class of difference as changing the generic TP size. The embedding exchange is pure data movement and stays bit-identical, including when stacked with the other two knobs. Negative configs (rejected or disabled at startup, no hang): | Config | Behavior | |---|---| | `mlp=2` + `--enforce-eager` | rejected at config load ("only supported in graph mode") | | `mlp=2` without `PreemptOffloadConnector` | rejected with the explicit connector requirement | | `mlp=2` without the recompute scheduler | rejected at config load in ~15 s with the recompute-scheduler requirement (the all-knobs gate above) | | capture bound 512 < largest step 600 | both knobs auto-disabled with a warning; startup completes | Base controls: the rounds above ran on the base before vllm-project#17159; since vllm-project#17159 reworked the embedding exchange, the embedding-relevant cells were re-run on the current base — baseline and `embedding=2` outputs are **bit-identical** to the pre-vllm-project#17159 results (TPOT 7.52 / 7.18 ms). A no-recompute control on the current base produced deterministically wrong outputs with **every knob off** (a pre-existing defect, tracked separately); `embedding=2` on that path is bit-identical to the no-knob run, i.e. fine-grained TP adds no effect — this is the defect behind the all-knobs recompute requirement. **Memory** (same bed): the family's fixed cost is one HCCL buffer set (~0.42 GB per rank, from vllm-project#16967); MLP saves 0.55 GiB per rank at `mlp=2` on the 6-layer bed (dense-FFN layers × 1/2), o_proj 0.109 GiB per layer per rank on a 1-byte checkpoint (vllm-project#16967). - vLLM main: vllm-project/vllm@84030bb --------- Signed-off-by: xuchi <xuchicolson@163.com>
What this PR does / why we need it?
Enables
oproj_tensor_parallel_size(fine-grained TP for the attention output projection) on model runner v2, scoped to its real deployment: a PD decode node (tp == 1,dp > 1, MoE) running a graph mode with the recompute scheduler.The o_proj exchange (
all_to_all→ partial GEMM →reduce_scatter) is anall_to_all_singlewithout split sizes: every rank must forward the same number of tokens per step, or the collectives hang silently. In graph mode the upstream dispatch already guarantees that group-uniform count, so this PR adds no runner-side padding — it adds the contract that keeps the count uniform, and guardrails that turn a silent hang into an explicit failure:cudagraph_mode == NONE(which coversenforce_eager) is rejected at config load.min(max_num_batched_tokens, max_num_seqs * decode_query_len)); otherwise it is disabled with a warning and the deployment still starts.NPUModelRunner.prepare_inputs— the only post-dispatch hook, reached by real steps only — raises an explicitRuntimeErrorif a step still dispatches to eager. Armed only for a real split (> 1).prefill_context_parallel_size > 1) is rejected: its per-rank token recount breaks the uniform step size.> 1, the KV connector chain must carryPreemptOffloadConnector— startup fails without it — and an offload failure at preemption aborts the request instead of sending it back to P. Checked inplatform._check_ascend_configbecause MultiConnector children re-derive the config from per-child copies that cannot see their siblings.Existing preconditions unchanged: PD decode node (
is_kv_consumer),recompute_scheduler_enable=true,tensor_parallel_size == 1, MoE.Depends on #17128 (reclassifies the request-arrival boundary recompute as decode, so those steps stay on captured graphs); this PR is meant to land after it.
Design conclusions (aligned with the SE)
The scope and guardrails of this PR implement conclusions aligned with the SE:
oproj_tensor_parallel_sizerequires the recompute scheduler, preempt offload, a graph mode, and a pure-DP PD decode deployment (tp == 1,dp > 1).PreemptOffloadConnectoris mandatory. Config load rejects a real split whose KV connector chain does not carry it.Memory cost and expected savings
The memory profile belongs to the feature, not to this PR — the validation and guard paths allocate no device memory. Measured on the test bed (DeepSeek-V3.1-arch W8A8 checkpoint truncated to 6 layers,
dp8 × tp1, graph mode, identical connector config; per rank):oproj=2, 0.49 GiB @oproj=4on 6 layersoproj=8oproj=8Fine_grained_TP.md)oproj=8The freed weight memory returns to the KV cache, and the fixed cost is recovered after ~8 layers at
oproj=2(~4 atoproj=8) on a 1-byte quantized checkpoint — trivially met by production-scale MoE models.Does this PR introduce any user-facing change?
Yes.
oproj_tensor_parallel_sizenow works on the V2 runner in graph-mode PD-decode deployments; previously it was V1-only.cudagraph_mode=none/enforce_eager,prefill_context_parallel_size > 1, and — for a real split — a KV connector chain withoutPreemptOffloadConnector.max_cudagraph_capture_sizeto use the knob at largemax_num_seqs.How was this patch tested?
Unit tests: new cases in
tests/ut/test_ascend_config.py(graph-mode / PCP rejection, size-1 exemption, capture-bound matrix),tests/ut/worker/test_model_runner_v2_finegrained_tp.py(eager-step guard contract) andtests/ut/core/test_recompute_scheduler.py(offload-failure abort). CI on the current head:pre-commit(incl. mypy) andcpu-utgreen (5718 passed, 67 skipped).End-to-end (single-node A3, 16 chips, PD disaggregation behind the LB proxy, 8-layer MiniMax-M3 with dummy weights — functional/consistency numbers, not real-weight performance): P = chip 0,
tp1eager, Mooncake producer; D = chips 8-15,dp8 × tp1,FULL_DECODE_ONLY+ recompute scheduler +oproj_tensor_parallel_size=2, Mooncake consumer; all traffic through the proxy.--max-num-seqs 600(600 > 512) auto-disables the knob with a warning and startup continues into graph captureoproj=2oproj=0baselineRuntimeErrorwithin 3.7 s, no hangMultiConnector: Mooncake +PreemptOffloadConnector)gpu-memory-utilization0.5 / 0.4 / 0.35, 100 MB CPU pool)Left uncovered (for honesty): the preemption/offload runtime path (KV save/restore and the capacity fallback) was not triggered — preemption is unreachable on the 8-layer dummy model — so the abort-instead-of-return-to-P branch is covered by unit tests only; and the without-connector startup failure is a config-load check that has not been reproduced e2e yet (it was observed on the test bed while the check still warned).