Repository navigation
[Doc][Misc] Complete SP MoE guide and temporary FlashComm switch - #15737
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request updates the documentation for Sequence Parallelism (SP) on vLLM Ascend, specifically focusing on the MoE path. It provides a thorough guide on the principles, configuration requirements, and constraints for enabling SP MoE, while clarifying the temporary nature of the current Ascend-specific FlashComm switch. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request updates the Sequence Parallelism documentation for vLLM Ascend, adding comprehensive details on the SP MoE path, its activation principles, communication flow, usage instructions, constraints, and the temporary FlashComm switch. The reviewer identified a technical inaccuracy in the explanation of the o_proj layer's inputs and outputs, providing a code suggestion to clarify that the outputs of o_proj (which feed into the MoE layer) are replicated rather than its inputs.
| On vLLM Ascend, SP currently covers the MoE path (SP MoE). The attention | ||
| `o_proj` ends with a TP all-reduce, so its inputs are replicated on every TP | ||
| rank. Feeding those replicated tokens directly into the experts duplicates | ||
| compute and communication under expert parallelism. SP MoE keeps the expert | ||
| inputs sharded by sequence and restores the expected layout at the MoE output | ||
| boundary instead. |
There was a problem hiding this comment.
The attention o_proj is a RowParallelLinear layer. Its inputs are sharded across TP ranks, and its outputs are all-reduced (replicated) across TP ranks. Therefore, the inputs to the subsequent MoE layer (which are the outputs of o_proj) are replicated on every TP rank. Saying that o_proj's inputs are replicated is technically incorrect and confusing.\n\nSuggested PR Title:\n\nmarkdown\n[Doc][Misc] Complete SP MoE guide and temporary FlashComm switch\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\n\nCompletes the `sequence_parallelism.md` feature guide, which currently only has\nan Overview skeleton plus a stale \"FlashComm is deprecated\" note. The new content\ncovers SP MoE end to end:\n\n- Principle: upstream `ParallelConfig.use_sequence_parallel_moe` activation\n conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable `all2all_backend`),\n the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP all-gather\n flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`,\n `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior.\n- How to use: upstream serve flags, constraints (TP>1, EP required for MoE,\n TP-multiple capture sizes, PCP incompatibility).\n- The temporary Ascend-only FlashComm switch: by default the platform forces\n `all2all_backend=flashinfer_all2allv` (SP MoE off); setting\n `additional_config.enable_flashcomm1` (preferred) or\n `VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents that the\n switch is temporary/deprecated and will be removed once SP is supported.\n\n### Does this PR introduce _any_ user-facing change?\n\nDocumentation only. No code behavior change.\n\n### How was this patch tested?\n\n- `markdownlint docs/source/user_guide/feature_guide/sequence_parallelism.md` passes.\n- No code changed, so no unit/e2e tests apply.\n
| On vLLM Ascend, SP currently covers the MoE path (SP MoE). The attention | |
| `o_proj` ends with a TP all-reduce, so its inputs are replicated on every TP | |
| rank. Feeding those replicated tokens directly into the experts duplicates | |
| compute and communication under expert parallelism. SP MoE keeps the expert | |
| inputs sharded by sequence and restores the expected layout at the MoE output | |
| boundary instead. | |
| On vLLM Ascend, SP currently covers the MoE path (SP MoE). The attention\n`o_proj` ends with a TP all-reduce, so its outputs (which are the inputs to the MoE layer) are replicated on every TP\nrank. Feeding those replicated tokens directly into the experts duplicates\ncompute and communication under expert parallelism. SP MoE keeps the expert\ninputs sharded by sequence and restores the expected layout at the MoE output\nboundary instead. |
References
- PR reviews must output suggested PR Title and PR Summary in markdown code blocks. (link)
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
|
explain the relationship between SP and flashcomm1, and why control sp with {"flashcomm1": true" |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Fix MD022/MD032 blank lines around headings and lists, MD009 trailing spaces, and MD047 missing trailing newline. Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
5cdc194 to
be290cf
Compare
…m-project#15737) ### What this PR does / why we need it? Completes the `sequence_parallelism.md` feature guide, which currently only has an Overview skeleton plus a stale "FlashComm is deprecated" note. The new content covers SP MoE end to end: - Principle: upstream `ParallelConfig.use_sequence_parallel_moe` activation conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable `all2all_backend`), the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP all-gather flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`, `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior. - How to use: upstream serve flags, constraints (TP>1, EP required for MoE, TP-multiple capture sizes, PCP incompatibility). - The temporary Ascend-only FlashComm switch: by default the platform forces `all2all_backend=flashinfer_all2allv` (SP MoE off); setting `additional_config.enable_flashcomm1` (preferred) or `VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents that the switch is temporary/deprecated and will be removed once SP is supported. ### Does this PR introduce _any_ user-facing change? Documentation only. No code behavior change. ### How was this patch tested? - `markdownlint docs/source/user_guide/feature_guide/sequence_parallelism.md` passes. - No code changed, so no unit/e2e tests apply. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: XuRongSheng <1843167357@qq.com>
…m-project#15737) ### What this PR does / why we need it? Completes the `sequence_parallelism.md` feature guide, which currently only has an Overview skeleton plus a stale "FlashComm is deprecated" note. The new content covers SP MoE end to end: - Principle: upstream `ParallelConfig.use_sequence_parallel_moe` activation conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable `all2all_backend`), the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP all-gather flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`, `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior. - How to use: upstream serve flags, constraints (TP>1, EP required for MoE, TP-multiple capture sizes, PCP incompatibility). - The temporary Ascend-only FlashComm switch: by default the platform forces `all2all_backend=flashinfer_all2allv` (SP MoE off); setting `additional_config.enable_flashcomm1` (preferred) or `VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents that the switch is temporary/deprecated and will be removed once SP is supported. ### Does this PR introduce _any_ user-facing change? Documentation only. No code behavior change. ### How was this patch tested? - `markdownlint docs/source/user_guide/feature_guide/sequence_parallelism.md` passes. - No code changed, so no unit/e2e tests apply. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: XuRongSheng <1843167357@qq.com>
…m-project#15737) ### What this PR does / why we need it? Completes the `sequence_parallelism.md` feature guide, which currently only has an Overview skeleton plus a stale "FlashComm is deprecated" note. The new content covers SP MoE end to end: - Principle: upstream `ParallelConfig.use_sequence_parallel_moe` activation conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable `all2all_backend`), the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP all-gather flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`, `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior. - How to use: upstream serve flags, constraints (TP>1, EP required for MoE, TP-multiple capture sizes, PCP incompatibility). - The temporary Ascend-only FlashComm switch: by default the platform forces `all2all_backend=flashinfer_all2allv` (SP MoE off); setting `additional_config.enable_flashcomm1` (preferred) or `VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents that the switch is temporary/deprecated and will be removed once SP is supported. ### Does this PR introduce _any_ user-facing change? Documentation only. No code behavior change. ### How was this patch tested? - `markdownlint docs/source/user_guide/feature_guide/sequence_parallelism.md` passes. - No code changed, so no unit/e2e tests apply. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: XuRongSheng <1843167357@qq.com> Signed-off-by: tianming2009 <13246728590@163.com>
…m-project#15737) ### What this PR does / why we need it? Completes the `sequence_parallelism.md` feature guide, which currently only has an Overview skeleton plus a stale "FlashComm is deprecated" note. The new content covers SP MoE end to end: - Principle: upstream `ParallelConfig.use_sequence_parallel_moe` activation conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable `all2all_backend`), the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP all-gather flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`, `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior. - How to use: upstream serve flags, constraints (TP>1, EP required for MoE, TP-multiple capture sizes, PCP incompatibility). - The temporary Ascend-only FlashComm switch: by default the platform forces `all2all_backend=flashinfer_all2allv` (SP MoE off); setting `additional_config.enable_flashcomm1` (preferred) or `VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents that the switch is temporary/deprecated and will be removed once SP is supported. ### Does this PR introduce _any_ user-facing change? Documentation only. No code behavior change. ### How was this patch tested? - `markdownlint docs/source/user_guide/feature_guide/sequence_parallelism.md` passes. - No code changed, so no unit/e2e tests apply. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: XuRongSheng <1843167357@qq.com> Signed-off-by: like-0517 <ithwlike@126.com>
What this PR does / why we need it?
Completes the
sequence_parallelism.mdfeature guide, which currently only hasan Overview skeleton plus a stale "FlashComm is deprecated" note. The new content
covers SP MoE end to end:
ParallelConfig.use_sequence_parallel_moeactivationconditions (TP>1, DP>1,
enable_expert_parallel, SP-capableall2all_backend),the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP all-gather
flow, and the compile-time (
pass_config.enable_sp,sp_min_token_num,enable_sp_by_pass) / run-time (TP-aligned cudagraph sizes) behavior.TP-multiple capture sizes, PCP incompatibility).
all2all_backend=flashinfer_all2allv(SP MoE off); settingadditional_config.enable_flashcomm1(preferred) orVLLM_ASCEND_ENABLE_FLASHCOMM1=1opts into upstream SP MoE. Documents that theswitch is temporary/deprecated and will be removed once SP is supported.
Does this PR introduce any user-facing change?
Documentation only. No code behavior change.
How was this patch tested?
markdownlint docs/source/user_guide/feature_guide/sequence_parallelism.mdpasses.No code changed, so no unit/e2e tests apply.
vLLM main: vllm-project/vllm@b2f6858