Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions docs/source/user_guide/configuration/additional_config.md
Original file line number Diff line number Diff line change
Expand Up @@ -22,6 +22,7 @@ Starting from [PR #9064](https://github.com/vllm-project/vllm-ascend/pull/9064),
| `VLLM_ASCEND_ENABLE_NZ` | `weight_nz_mode` | Integer (unchanged, field name changed) |
| `VLLM_ASCEND_ENABLE_FUSED_MC2` | `enable_fused_mc2` | Integer (unchanged) |
| `VLLM_ASCEND_FUSION_OP_TRANSPOSE_KV_CACHE_BY_BLOCK` | `enable_transpose_kv_cache_by_block` | `"1"` → `true`, `"0"` → `false` |
| `VLLM_ASCEND_ENABLE_FLASHCOMM1` | `enable_flashcomm1` | `"1"` → `true`, `"0"` → `false` |

## How to use

Expand Down Expand Up @@ -73,6 +74,7 @@ The following table lists additional configuration options available in vLLM Asc
| `enable_fused_mc2` | int | `0` | Fused MC2 configuration. `0` disables the fused path and `1` enables it when the model and parallel configuration support it. The legacy `VLLM_ASCEND_ENABLE_FUSED_MC2` environment variable is no longer supported. |
| `enable_transpose_kv_cache_by_block`| bool | `True` | Whether to enable transpose KV cache by block. The legacy `VLLM_ASCEND_FUSION_OP_TRANSPOSE_KV_CACHE_BY_BLOCK` environment variable is no longer supported. |
| `enable_dsa_cp` | bool | `False` | Whether to enable dsa_cp for DeepSeek V3.2, DeepSeek V4, and other models with the same architecture. This feature requires sequence parallelism to be enabled. Enabling it automatically enables FlashComm.|
| `enable_flashcomm1` | bool | `False` | Whether to enable SP MoE. The legacy `VLLM_ASCEND_ENABLE_FLASHCOMM1` environment variable is kept for compatibility. See [Sequence Parallelism](../feature_guide/sequence_parallelism.md). |
| `enable_pcp_o_proj_weight_sharding` | bool | `False` | Whether SFA-PCP shards the O-proj weight across the PCP group at load time and switches between PCP-local and gathered weight views at runtime. This option does not affect DSA-CP, whose original policy remains fixed: prefill gathers the full O-proj weight and decode uses the local weight. This option must be set when the server starts. |
| `rejection_sampler_config` | dict | `{}` | Configuration options for rejection sampler (block verify and entropy verify). |
| `dynamic_spec_config` | dict | `{}` | Configuration options for Dynamic Speculative Decoding. See [Dynamic Speculative Decoding](../feature_guide/speculative_decoding.md#dynamic-speculative-decoding). |
Expand Down
72 changes: 66 additions & 6 deletions docs/source/user_guide/feature_guide/sequence_parallelism.md
Original file line number Diff line number Diff line change
Expand Up @@ -3,15 +3,75 @@
## Overview

Sequence Parallelism (SP) shards the token dimension across tensor-parallel
ranks around the communication boundaries of transformer layers. vLLM owns
ranks around the communication boundaries of transformer layers.

## How to use
On vLLM Ascend, SP currently covers the MoE path (SP MoE). The attention
`o_proj` ends with a TP all-reduce, so its inputs are replicated on every TP
rank. Feeding those replicated tokens directly into the experts duplicates
compute and communication under expert parallelism. SP MoE keeps the expert
inputs sharded by sequence and restores the expected layout at the MoE output
boundary instead.
Comment on lines +8 to +13

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The attention o_proj is a RowParallelLinear layer. Its inputs are sharded across TP ranks, and its outputs are all-reduced (replicated) across TP ranks. Therefore, the inputs to the subsequent MoE layer (which are the outputs of o_proj) are replicated on every TP rank. Saying that o_proj's inputs are replicated is technically incorrect and confusing.\n\nSuggested PR Title:\n\nmarkdown\n[Doc][Misc] Complete SP MoE guide and temporary FlashComm switch\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\n\nCompletes the `sequence_parallelism.md` feature guide, which currently only has\nan Overview skeleton plus a stale \"FlashComm is deprecated\" note. The new content\ncovers SP MoE end to end:\n\n- Principle: upstream `ParallelConfig.use_sequence_parallel_moe` activation\n conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable `all2all_backend`),\n the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP all-gather\n flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`,\n `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior.\n- How to use: upstream serve flags, constraints (TP>1, EP required for MoE,\n TP-multiple capture sizes, PCP incompatibility).\n- The temporary Ascend-only FlashComm switch: by default the platform forces\n `all2all_backend=flashinfer_all2allv` (SP MoE off); setting\n `additional_config.enable_flashcomm1` (preferred) or\n `VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents that the\n switch is temporary/deprecated and will be removed once SP is supported.\n\n### Does this PR introduce _any_ user-facing change?\n\nDocumentation only. No code behavior change.\n\n### How was this patch tested?\n\n- `markdownlint docs/source/user_guide/feature_guide/sequence_parallelism.md` passes.\n- No code changed, so no unit/e2e tests apply.\n

Suggested change
On vLLM Ascend, SP currently covers the MoE path (SP MoE). The attention
`o_proj` ends with a TP all-reduce, so its inputs are replicated on every TP
rank. Feeding those replicated tokens directly into the experts duplicates
compute and communication under expert parallelism. SP MoE keeps the expert
inputs sharded by sequence and restores the expected layout at the MoE output
boundary instead.
On vLLM Ascend, SP currently covers the MoE path (SP MoE). The attention\n`o_proj` ends with a TP all-reduce, so its outputs (which are the inputs to the MoE layer) are replicated on every TP\nrank. Feeding those replicated tokens directly into the experts duplicates\ncompute and communication under expert parallelism. SP MoE keeps the expert\ninputs sharded by sequence and restores the expected layout at the MoE output\nboundary instead.
References
  1. PR reviews must output suggested PR Title and PR Summary in markdown code blocks. (link)


**The original flashcomm feature overlapped functionally with the SP feature and has been deprecated since v0.27.1.**

## Principle

SP MoE shards the input along the token dimension in each
Transformer layer. Different TP ranks therefore process different tokens,
avoiding duplicate expert computation for the same tokens.

Automatically enabled when DP>1, TP>1 are set, and specific all2all backend with moe model.
The former FlashComm settings are removed.
The main data flow of an MoE layer is:

```text
FlashComm is deprecated
Sequence-parallel input sharding
-> TP all-gather: collect tokens from all ranks
-> attention
-> TP reduce scatter
-> RMS Norm
-> Router
-> all-to-all
-> Moe
```

Different DP ranks may have different numbers of valid tokens. Therefore, the
buffer after all-gather cannot be treated as a contiguous sequence of valid
tokens; it must be unpadded and zero-padded according to each rank's local
token size. This keeps tokens sequence-sharded during expert computation and
reduces duplicate computation and unnecessary communication.

## How to use

Steps to follow to enable SP currently:

- `tensor_parallel_size > 1` and `data_parallel_size > 1`.
- `enable_expert_parallel` is set (MoE models only).
- `--additional-config '{"enable_flashcomm1": true}'` set `flashcomm1`

### Temporary FlashComm switch (Ascend only)

Until SP support is fully validated, vLLM Ascend keeps SP MoE option by original flashcomm option.

To opt into upstream SP MoE, set one of the following (the
`additional_config` form is preferred):

```bash
# Preferred.
vllm serve <moe-model> \
--data-parallel-size 2 \
--tensor-parallel-size 2 \
--enable-expert-parallel \
--additional-config '{"enable_flashcomm1": true}'
```

```bash
# Kept for compatibility.
VLLM_ASCEND_ENABLE_FLASHCOMM1=1 vllm serve <moe-model> \
--data-parallel-size 2 \
--tensor-parallel-size 2 \
--enable-expert-parallel
```

Use the upstream SP configuration shown above instead.
This switch is temporary and deprecated. Referencing either form logs a
`FlashComm is deprecated` warning from `init_ascend_config`, and the override
carries a `TODO` to remove it once SP is supported — after that, the upstream
configuration above takes effect directly.
8 changes: 4 additions & 4 deletions vllm_ascend/ops/register_custom_ops.py
Original file line number Diff line number Diff line change
Expand Up @@ -41,15 +41,15 @@ def _pad_to_ep_local_size(x: torch.Tensor, max_local_size: int) -> torch.Tensor:


def _maybe_all_gather_and_maybe_unpad_impl(x: torch.Tensor) -> torch.Tensor:
"""仅用于 EP 通信场景:EP all_gather + 按 DP token 分布 unpad。"""
"""EP communication only: EP all_gather followed by unpad according to the DP token distribution."""
forward_context = get_forward_context()
dp_metadata = forward_context.dp_metadata
ep_group = get_ep_group()
local_sizes = _get_ep_local_sizes(dp_metadata, ep_group)
if local_sizes is not None:
max_local_size = max(local_sizes)
# all_gather 要求各 rank 输入等长:先 pad 到 max_local_size,
# gather 后再按各 rank 真实的 local_sizes 截回。
# all_gather requires equal-length inputs on every rank: pad to
# max_local_size first, then trim back to each rank's real local size.
x = _pad_to_ep_local_size(x, max_local_size)
# need to unpad from ep size
x = ep_group.all_gather(x, 0)
Expand All @@ -73,7 +73,7 @@ def _maybe_all_gather_and_maybe_unpad_impl(x: torch.Tensor) -> torch.Tensor:


def _maybe_pad_and_reduce_impl(x: torch.Tensor) -> torch.Tensor:
"""仅用于 EP 通信场景:按 DP token 分布 pad 后做 EP reduce_scatter。"""
"""EP communication only: pad according to the DP token distribution, then EP reduce_scatter."""
forward_context = get_forward_context()

if _EXTRA_CTX.is_draft_model and is_vl_model():
Expand Down
Loading