Skip to content

[Doc] Backport SP MoE guide and temporary FlashComm switch to v0.27.1rc - #16324

Closed
MmMmaru wants to merge 3 commits into
vllm-project:releases/v0.27.1rcfrom
MmMmaru:agent/backport-sp-moe-guide-v0271
Closed

MmMmaru wants to merge 3 commits into
vllm-project:releases/v0.27.1rcfrom
MmMmaru:agent/backport-sp-moe-guide-v0271

Conversation

@MmMmaru

@MmMmaru MmMmaru commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

Backports the SP MoE feature guide to the releases/v0.27.1rc branch
(cherry-pick of #15737 from main).

The sequence_parallelism.md guide on the release branch is still the
old Overview skeleton with a stale "FlashComm is deprecated" note. This
brings over the completed end-to-end content:

  • Principle: ParallelConfig.use_sequence_parallel_moe activation
    conditions, per-layer EP all-gather/unpad → MoE → pad/reduce-scatter
    → TP all-gather flow, compile-time and run-time behavior.
  • How to use: serve flags, constraints (TP>1, EP required,
    TP-multiple capture sizes, PCP incompatibility).
  • The temporary Ascend-only FlashComm switch
    (additional_config.enable_flashcomm1 preferred,
    VLLM_ASCEND_ENABLE_FLASHCOMM1=1 for compatibility), documented as
    temporary/deprecated.
  • Comment-only English translation in register_custom_ops.py
    (no behavior change).

Conflict note: additional_config.md had a release-side enable_dsa_cp
row; resolved by taking the newer main wording (adds the
auto-enables-FlashComm sentence) plus the two new option rows.

Does this PR introduce any user-facing change?

Documentation only. No code behavior change.

How was this patch tested?

  • markdownlint --config .markdownlint.yaml passes on both touched
    markdown files.

  • ruff check passes on register_custom_ops.py (comment-only change).

  • No code changed, so no unit/e2e tests apply.

  • vLLM main: vllm-project/vllm@ba07e4a

…m-project#15737)

Completes the `sequence_parallelism.md` feature guide, which currently
only has
an Overview skeleton plus a stale "FlashComm is deprecated" note. The
new content
covers SP MoE end to end:

- Principle: upstream `ParallelConfig.use_sequence_parallel_moe`
activation
conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable
`all2all_backend`),
the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP
all-gather
flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`,
  `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior.
- How to use: upstream serve flags, constraints (TP>1, EP required for
MoE,
  TP-multiple capture sizes, PCP incompatibility).
- The temporary Ascend-only FlashComm switch: by default the platform
forces
  `all2all_backend=flashinfer_all2allv` (SP MoE off); setting
  `additional_config.enable_flashcomm1` (preferred) or
`VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents
that the
switch is temporary/deprecated and will be removed once SP is supported.

Documentation only. No code behavior change.

- `markdownlint
docs/source/user_guide/feature_guide/sequence_parallelism.md` passes.
- No code changed, so no unit/e2e tests apply.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: XuRongSheng <1843167357@qq.com>
(cherry picked from commit 4980809)
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request backports essential documentation regarding Sequence Parallelism (SP) for Mixture-of-Experts (MoE) models to the v0.27.1rc release branch. It clarifies the operational principles and configuration requirements for SP MoE on Ascend hardware, while providing clear guidance on the temporary, deprecated use of the FlashComm switch for compatibility. The changes are strictly limited to documentation and code comments, ensuring no functional changes to the runtime behavior.

Highlights

  • Documentation Update: Backported the comprehensive Sequence Parallelism (SP) MoE guide to the v0.27.1rc release branch, providing detailed explanations of the data flow and activation conditions.
  • Temporary FlashComm Configuration: Documented the temporary Ascend-specific FlashComm switch (enable_flashcomm1) as a deprecated mechanism for enabling SP MoE until full upstream support is validated.
  • Codebase Localization: Updated comments in register_custom_ops.py to English to improve accessibility for international contributors.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Doc][Misc] Update Sequence Parallelism and FlashComm documentation and translate comments

Suggested PR Summary:

### What this PR does / why we need it?
This PR updates the documentation for Sequence Parallelism (SP) and FlashComm on Ascend, providing a detailed explanation of SP MoE, its principle, data flow, and usage. It documents new configuration options including `enable_flashcomm1` and `enable_pcp_o_proj_weight_sharding`. Additionally, it translates Chinese code comments to English in `vllm_ascend/ops/register_custom_ops.py`.

### Does this PR introduce _any_ user-facing change?
Yes, documentation has been updated to reflect new configuration options and the deprecation of FlashComm.

### How was this patch tested?
No functional code changes were introduced; only documentation and comments were updated.

I have no review comments to evaluate.

@github-actions github-actions Bot added documentation Improvements or additions to documentation module:ops labels Sep 11, 2026
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

Signed-off-by: XuRongSheng <1843167357@qq.com>
@MmMmaru
MmMmaru force-pushed the agent/backport-sp-moe-guide-v0271 branch from a789214 to 4dd8eab Compare September 11, 2026 03:37
@MmMmaru
MmMmaru marked this pull request as ready for review September 11, 2026 03:38
Signed-off-by: Xu Rongsheng <1843167357@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation merge-conflicts module:ops

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant