diff --git a/docs/source/user_guide/configuration/additional_config.md b/docs/source/user_guide/configuration/additional_config.md index 8b87a50bce98..3c80a52f2f09 100644 --- a/docs/source/user_guide/configuration/additional_config.md +++ b/docs/source/user_guide/configuration/additional_config.md @@ -22,6 +22,7 @@ Starting from [PR #9064](https://github.com/vllm-project/vllm-ascend/pull/9064), | `VLLM_ASCEND_ENABLE_NZ` | `weight_nz_mode` | Integer (unchanged, field name changed) | | `VLLM_ASCEND_ENABLE_FUSED_MC2` | `enable_fused_mc2` | Integer (unchanged) | | `VLLM_ASCEND_FUSION_OP_TRANSPOSE_KV_CACHE_BY_BLOCK` | `enable_transpose_kv_cache_by_block` | `"1"` → `true`, `"0"` → `false` | +| `VLLM_ASCEND_ENABLE_FLASHCOMM1` | `enable_flashcomm1` | `"1"` → `true`, `"0"` → `false` | ## How to use @@ -73,6 +74,7 @@ The following table lists additional configuration options available in vLLM Asc | `enable_fused_mc2` | int | `0` | Fused MC2 configuration. `0` disables the fused path and `1` enables it when the model and parallel configuration support it. The legacy `VLLM_ASCEND_ENABLE_FUSED_MC2` environment variable is no longer supported. | | `enable_transpose_kv_cache_by_block`| bool | `True` | Whether to enable transpose KV cache by block. The legacy `VLLM_ASCEND_FUSION_OP_TRANSPOSE_KV_CACHE_BY_BLOCK` environment variable is no longer supported. | | `enable_dsa_cp` | bool | `False` | Whether to enable dsa_cp for DeepSeek V3.2, DeepSeek V4, and other models with the same architecture. This feature requires sequence parallelism to be enabled. Enabling it automatically enables FlashComm.| +| `enable_flashcomm1` | bool | `False` | Whether to enable SP MoE. The legacy `VLLM_ASCEND_ENABLE_FLASHCOMM1` environment variable is kept for compatibility. See [Sequence Parallelism](../feature_guide/sequence_parallelism.md). | | `enable_pcp_o_proj_weight_sharding` | bool | `False` | Whether SFA-PCP shards the O-proj weight across the PCP group at load time and switches between PCP-local and gathered weight views at runtime. This option does not affect DSA-CP, whose original policy remains fixed: prefill gathers the full O-proj weight and decode uses the local weight. This option must be set when the server starts. | | `rejection_sampler_config` | dict | `{}` | Configuration options for rejection sampler (block verify and entropy verify). | | `dynamic_spec_config` | dict | `{}` | Configuration options for Dynamic Speculative Decoding. See [Dynamic Speculative Decoding](../feature_guide/speculative_decoding.md#dynamic-speculative-decoding). | diff --git a/docs/source/user_guide/feature_guide/sequence_parallelism.md b/docs/source/user_guide/feature_guide/sequence_parallelism.md index 0e6b8d5b72a0..0c6d2b5519d3 100644 --- a/docs/source/user_guide/feature_guide/sequence_parallelism.md +++ b/docs/source/user_guide/feature_guide/sequence_parallelism.md @@ -3,15 +3,75 @@ ## Overview Sequence Parallelism (SP) shards the token dimension across tensor-parallel -ranks around the communication boundaries of transformer layers. vLLM owns +ranks around the communication boundaries of transformer layers. -## How to use +On vLLM Ascend, SP currently covers the MoE path (SP MoE). The attention +`o_proj` ends with a TP all-reduce, so its inputs are replicated on every TP +rank. Feeding those replicated tokens directly into the experts duplicates +compute and communication under expert parallelism. SP MoE keeps the expert +inputs sharded by sequence and restores the expected layout at the MoE output +boundary instead. + +**The original flashcomm feature overlapped functionally with the SP feature and has been deprecated since v0.27.1.** + +## Principle + +SP MoE shards the input along the token dimension in each +Transformer layer. Different TP ranks therefore process different tokens, +avoiding duplicate expert computation for the same tokens. -Automatically enabled when DP>1, TP>1 are set, and specific all2all backend with moe model. -The former FlashComm settings are removed. +The main data flow of an MoE layer is: ```text -FlashComm is deprecated +Sequence-parallel input sharding + -> TP all-gather: collect tokens from all ranks + -> attention + -> TP reduce scatter + -> RMS Norm + -> Router + -> all-to-all + -> Moe +``` + +Different DP ranks may have different numbers of valid tokens. Therefore, the +buffer after all-gather cannot be treated as a contiguous sequence of valid +tokens; it must be unpadded and zero-padded according to each rank's local +token size. This keeps tokens sequence-sharded during expert computation and +reduces duplicate computation and unnecessary communication. + +## How to use + +Steps to follow to enable SP currently: + +- `tensor_parallel_size > 1` and `data_parallel_size > 1`. +- `enable_expert_parallel` is set (MoE models only). +- `--additional-config '{"enable_flashcomm1": true}'` set `flashcomm1` + +### Temporary FlashComm switch (Ascend only) + +Until SP support is fully validated, vLLM Ascend keeps SP MoE option by original flashcomm option. + +To opt into upstream SP MoE, set one of the following (the +`additional_config` form is preferred): + +```bash +# Preferred. +vllm serve \ + --data-parallel-size 2 \ + --tensor-parallel-size 2 \ + --enable-expert-parallel \ + --additional-config '{"enable_flashcomm1": true}' +``` + +```bash +# Kept for compatibility. +VLLM_ASCEND_ENABLE_FLASHCOMM1=1 vllm serve \ + --data-parallel-size 2 \ + --tensor-parallel-size 2 \ + --enable-expert-parallel ``` -Use the upstream SP configuration shown above instead. +This switch is temporary and deprecated. Referencing either form logs a +`FlashComm is deprecated` warning from `init_ascend_config`, and the override +carries a `TODO` to remove it once SP is supported — after that, the upstream +configuration above takes effect directly. diff --git a/vllm_ascend/ops/register_custom_ops.py b/vllm_ascend/ops/register_custom_ops.py index 4f1df8f97ac3..c1cf4b68a434 100644 --- a/vllm_ascend/ops/register_custom_ops.py +++ b/vllm_ascend/ops/register_custom_ops.py @@ -41,15 +41,15 @@ def _pad_to_ep_local_size(x: torch.Tensor, max_local_size: int) -> torch.Tensor: def _maybe_all_gather_and_maybe_unpad_impl(x: torch.Tensor) -> torch.Tensor: - """仅用于 EP 通信场景:EP all_gather + 按 DP token 分布 unpad。""" + """EP communication only: EP all_gather followed by unpad according to the DP token distribution.""" forward_context = get_forward_context() dp_metadata = forward_context.dp_metadata ep_group = get_ep_group() local_sizes = _get_ep_local_sizes(dp_metadata, ep_group) if local_sizes is not None: max_local_size = max(local_sizes) - # all_gather 要求各 rank 输入等长:先 pad 到 max_local_size, - # gather 后再按各 rank 真实的 local_sizes 截回。 + # all_gather requires equal-length inputs on every rank: pad to + # max_local_size first, then trim back to each rank's real local size. x = _pad_to_ep_local_size(x, max_local_size) # need to unpad from ep size x = ep_group.all_gather(x, 0) @@ -73,7 +73,7 @@ def _maybe_all_gather_and_maybe_unpad_impl(x: torch.Tensor) -> torch.Tensor: def _maybe_pad_and_reduce_impl(x: torch.Tensor) -> torch.Tensor: - """仅用于 EP 通信场景:按 DP token 分布 pad 后做 EP reduce_scatter。""" + """EP communication only: pad according to the DP token distribution, then EP reduce_scatter.""" forward_context = get_forward_context() if _EXTRA_CTX.is_draft_model and is_vl_model():