Skip to content

[Doc][Misc] Complete SP MoE guide and temporary FlashComm switch - #15737

Merged
realliujiaxu merged 9 commits into
vllm-project:mainfrom
MmMmaru:agent/sequence-parallelism-doc
Sep 10, 2026
Merged

realliujiaxu merged 9 commits into
vllm-project:mainfrom
MmMmaru:agent/sequence-parallelism-doc

Conversation

@MmMmaru

@MmMmaru MmMmaru commented Sep 4, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Completes the sequence_parallelism.md feature guide, which currently only has
an Overview skeleton plus a stale "FlashComm is deprecated" note. The new content
covers SP MoE end to end:

  • Principle: upstream ParallelConfig.use_sequence_parallel_moe activation
    conditions (TP>1, DP>1, enable_expert_parallel, SP-capable all2all_backend),
    the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP all-gather
    flow, and the compile-time (pass_config.enable_sp, sp_min_token_num,
    enable_sp_by_pass) / run-time (TP-aligned cudagraph sizes) behavior.
  • How to use: upstream serve flags, constraints (TP>1, EP required for MoE,
    TP-multiple capture sizes, PCP incompatibility).
  • The temporary Ascend-only FlashComm switch: by default the platform forces
    all2all_backend=flashinfer_all2allv (SP MoE off); setting
    additional_config.enable_flashcomm1 (preferred) or
    VLLM_ASCEND_ENABLE_FLASHCOMM1=1 opts into upstream SP MoE. Documents that the
    switch is temporary/deprecated and will be removed once SP is supported.

Does this PR introduce any user-facing change?

Documentation only. No code behavior change.

How was this patch tested?

  • markdownlint docs/source/user_guide/feature_guide/sequence_parallelism.md passes.

  • No code changed, so no unit/e2e tests apply.

  • vLLM main: vllm-project/vllm@b2f6858

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request updates the documentation for Sequence Parallelism (SP) on vLLM Ascend, specifically focusing on the MoE path. It provides a thorough guide on the principles, configuration requirements, and constraints for enabling SP MoE, while clarifying the temporary nature of the current Ascend-specific FlashComm switch.

Highlights

  • Documentation Update: Completed the sequence_parallelism.md guide to include comprehensive details on Sequence Parallelism (SP) for Mixture of Experts (MoE) models.
  • Technical Principles: Added detailed explanations of the SP MoE flow, including EP all-gather/unpad, MoE compute, and TP-aligned graph capture requirements.
  • Temporary Configuration Switch: Documented the temporary Ascend-only FlashComm switch used to opt into SP MoE, noting its deprecated status and future removal.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the Sequence Parallelism documentation for vLLM Ascend, adding comprehensive details on the SP MoE path, its activation principles, communication flow, usage instructions, constraints, and the temporary FlashComm switch. The reviewer identified a technical inaccuracy in the explanation of the o_proj layer's inputs and outputs, providing a code suggestion to clarify that the outputs of o_proj (which feed into the MoE layer) are replicated rather than its inputs.

Comment on lines +8 to +13
On vLLM Ascend, SP currently covers the MoE path (SP MoE). The attention
`o_proj` ends with a TP all-reduce, so its inputs are replicated on every TP
rank. Feeding those replicated tokens directly into the experts duplicates
compute and communication under expert parallelism. SP MoE keeps the expert
inputs sharded by sequence and restores the expected layout at the MoE output
boundary instead.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The attention o_proj is a RowParallelLinear layer. Its inputs are sharded across TP ranks, and its outputs are all-reduced (replicated) across TP ranks. Therefore, the inputs to the subsequent MoE layer (which are the outputs of o_proj) are replicated on every TP rank. Saying that o_proj's inputs are replicated is technically incorrect and confusing.\n\nSuggested PR Title:\n\nmarkdown\n[Doc][Misc] Complete SP MoE guide and temporary FlashComm switch\n\n\nSuggested PR Summary:\n\nmarkdown\n### What this PR does / why we need it?\n\nCompletes the `sequence_parallelism.md` feature guide, which currently only has\nan Overview skeleton plus a stale \"FlashComm is deprecated\" note. The new content\ncovers SP MoE end to end:\n\n- Principle: upstream `ParallelConfig.use_sequence_parallel_moe` activation\n conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable `all2all_backend`),\n the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP all-gather\n flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`,\n `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior.\n- How to use: upstream serve flags, constraints (TP>1, EP required for MoE,\n TP-multiple capture sizes, PCP incompatibility).\n- The temporary Ascend-only FlashComm switch: by default the platform forces\n `all2all_backend=flashinfer_all2allv` (SP MoE off); setting\n `additional_config.enable_flashcomm1` (preferred) or\n `VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents that the\n switch is temporary/deprecated and will be removed once SP is supported.\n\n### Does this PR introduce _any_ user-facing change?\n\nDocumentation only. No code behavior change.\n\n### How was this patch tested?\n\n- `markdownlint docs/source/user_guide/feature_guide/sequence_parallelism.md` passes.\n- No code changed, so no unit/e2e tests apply.\n

Suggested change
On vLLM Ascend, SP currently covers the MoE path (SP MoE). The attention
`o_proj` ends with a TP all-reduce, so its inputs are replicated on every TP
rank. Feeding those replicated tokens directly into the experts duplicates
compute and communication under expert parallelism. SP MoE keeps the expert
inputs sharded by sequence and restores the expected layout at the MoE output
boundary instead.
On vLLM Ascend, SP currently covers the MoE path (SP MoE). The attention\n`o_proj` ends with a TP all-reduce, so its outputs (which are the inputs to the MoE layer) are replicated on every TP\nrank. Feeding those replicated tokens directly into the experts duplicates\ncompute and communication under expert parallelism. SP MoE keeps the expert\ninputs sharded by sequence and restores the expected layout at the MoE output\nboundary instead.
References
  1. PR reviews must output suggested PR Title and PR Summary in markdown code blocks. (link)

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 4, 2026
@github-actions

github-actions Bot commented Sep 4, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@MmMmaru
MmMmaru marked this pull request as ready for review September 4, 2026 06:14
@realliujiaxu

Copy link
Copy Markdown
Collaborator

explain the relationship between SP and flashcomm1, and why control sp with {"flashcomm1": true"

@realliujiaxu realliujiaxu added the ready-precise run selected e2e test for pr label Sep 9, 2026
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Fix MD022/MD032 blank lines around headings and lists, MD009 trailing spaces, and MD047 missing trailing newline.

Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: XuRongSheng <1843167357@qq.com>
@MmMmaru
MmMmaru force-pushed the agent/sequence-parallelism-doc branch from 5cdc194 to be290cf Compare September 10, 2026 03:41
@realliujiaxu
realliujiaxu merged commit 4980809 into vllm-project:main Sep 10, 2026
17 checks passed
chen-commits pushed a commit to chen-commits/vllm-ascend that referenced this pull request Sep 10, 2026
…m-project#15737)

### What this PR does / why we need it?

Completes the `sequence_parallelism.md` feature guide, which currently
only has
an Overview skeleton plus a stale "FlashComm is deprecated" note. The
new content
covers SP MoE end to end:

- Principle: upstream `ParallelConfig.use_sequence_parallel_moe`
activation
conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable
`all2all_backend`),
the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP
all-gather
flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`,
  `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior.
- How to use: upstream serve flags, constraints (TP>1, EP required for
MoE,
  TP-multiple capture sizes, PCP incompatibility).
- The temporary Ascend-only FlashComm switch: by default the platform
forces
  `all2all_backend=flashinfer_all2allv` (SP MoE off); setting
  `additional_config.enable_flashcomm1` (preferred) or
`VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents
that the
switch is temporary/deprecated and will be removed once SP is supported.

### Does this PR introduce _any_ user-facing change?

Documentation only. No code behavior change.

### How was this patch tested?

- `markdownlint
docs/source/user_guide/feature_guide/sequence_parallelism.md` passes.
- No code changed, so no unit/e2e tests apply.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: XuRongSheng <1843167357@qq.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
…m-project#15737)

### What this PR does / why we need it?

Completes the `sequence_parallelism.md` feature guide, which currently
only has
an Overview skeleton plus a stale "FlashComm is deprecated" note. The
new content
covers SP MoE end to end:

- Principle: upstream `ParallelConfig.use_sequence_parallel_moe`
activation
conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable
`all2all_backend`),
the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP
all-gather
flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`,
  `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior.
- How to use: upstream serve flags, constraints (TP>1, EP required for
MoE,
  TP-multiple capture sizes, PCP incompatibility).
- The temporary Ascend-only FlashComm switch: by default the platform
forces
  `all2all_backend=flashinfer_all2allv` (SP MoE off); setting
  `additional_config.enable_flashcomm1` (preferred) or
`VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents
that the
switch is temporary/deprecated and will be removed once SP is supported.

### Does this PR introduce _any_ user-facing change?

Documentation only. No code behavior change.

### How was this patch tested?

- `markdownlint
docs/source/user_guide/feature_guide/sequence_parallelism.md` passes.
- No code changed, so no unit/e2e tests apply.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: XuRongSheng <1843167357@qq.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
…m-project#15737)

### What this PR does / why we need it?

Completes the `sequence_parallelism.md` feature guide, which currently
only has
an Overview skeleton plus a stale "FlashComm is deprecated" note. The
new content
covers SP MoE end to end:

- Principle: upstream `ParallelConfig.use_sequence_parallel_moe`
activation
conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable
`all2all_backend`),
the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP
all-gather
flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`,
  `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior.
- How to use: upstream serve flags, constraints (TP>1, EP required for
MoE,
  TP-multiple capture sizes, PCP incompatibility).
- The temporary Ascend-only FlashComm switch: by default the platform
forces
  `all2all_backend=flashinfer_all2allv` (SP MoE off); setting
  `additional_config.enable_flashcomm1` (preferred) or
`VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents
that the
switch is temporary/deprecated and will be removed once SP is supported.

### Does this PR introduce _any_ user-facing change?

Documentation only. No code behavior change.

### How was this patch tested?

- `markdownlint
docs/source/user_guide/feature_guide/sequence_parallelism.md` passes.
- No code changed, so no unit/e2e tests apply.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
…m-project#15737)

### What this PR does / why we need it?

Completes the `sequence_parallelism.md` feature guide, which currently
only has
an Overview skeleton plus a stale "FlashComm is deprecated" note. The
new content
covers SP MoE end to end:

- Principle: upstream `ParallelConfig.use_sequence_parallel_moe`
activation
conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable
`all2all_backend`),
the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP
all-gather
flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`,
  `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior.
- How to use: upstream serve flags, constraints (TP>1, EP required for
MoE,
  TP-multiple capture sizes, PCP incompatibility).
- The temporary Ascend-only FlashComm switch: by default the platform
forces
  `all2all_backend=flashinfer_all2allv` (SP MoE off); setting
  `additional_config.enable_flashcomm1` (preferred) or
`VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents
that the
switch is temporary/deprecated and will be removed once SP is supported.

### Does this PR introduce _any_ user-facing change?

Documentation only. No code behavior change.

### How was this patch tested?

- `markdownlint
docs/source/user_guide/feature_guide/sequence_parallelism.md` passes.
- No code changed, so no unit/e2e tests apply.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: XuRongSheng <1843167357@qq.com>
Signed-off-by: like-0517 <ithwlike@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation module:ops ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants