Skip to content

[Feat][SP] Suport SP for VL MoE models - #7044

Merged
realliujiaxu merged 23 commits into
vllm-project:mainfrom
realliujiaxu:moe-sp-allgather
Mar 24, 2026
Merged

realliujiaxu merged 23 commits into
vllm-project:mainfrom
realliujiaxu:moe-sp-allgather

Conversation

@realliujiaxu

@realliujiaxu realliujiaxu commented Mar 6, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

2nd PR for #5712, extend SP to VL MoE models.

Does this PR introduce any user-facing change?

remove sp_threshold in additional config and reuse sp_min_token_num from vLLM.

How was this patch tested?

  • Model: Qwen3-VL-30B-A3B,
  • TP4 DP2
  • 100 reqs
  • max concurrency 1
Seq length Mean TTFT (ms) main Mean TTFT (ms) this PR
4k 429.40 323.3
16k 1297.01 911.74

@github-actions

github-actions Bot commented Mar 6, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly refactors the sequence parallelism (SP) compilation pass architecture to enhance modularity and maintainability. It introduces a two-stage pass system, separating AllReduce and AllGather pattern applications, and integrates a new pass for eliminating redundant graph operations. Concurrently, the documentation for sequence parallelism has been comprehensively updated to reflect these architectural changes, providing clearer explanations and pattern details for users and developers. These changes streamline the SP implementation without altering user-facing APIs or behavior, focusing on internal code quality and clarity.

Highlights

  • Pass Architecture Refactor: The sequence parallelism (SP) pass architecture has been refactored into two distinct passes: SequenceParallelismPass and SequenceParallelismAllgatherEpPass. The SequenceParallelismPass now focuses on AllReduce-based patterns, while the new SequenceParallelismAllgatherEpPass handles AllGather-related patterns, improving modularity and clarity.
  • Pattern Reorganization and Naming Convention: AllGather patterns, including those for middle and last layers and Qwen3-VL, have been moved into the new sequence_parallelism_moe.py file. The AllGatherChunkNoOpCleanupPass logic has been merged into SequenceParallelismAllgatherEpPass as the AllGatherChunkNoOpPattern. Additionally, the 'Ascend' prefix has been removed from pass and pattern names for better generalization.
  • No-Op Elimination Pass: A new NoOpEliminationPass has been introduced to remove redundant view-like operations (e.g., view, reshape) from the computation graph, optimizing the graph before applying SP patterns.
  • Enhanced Documentation: The sequence_parallelism.md user guide has been significantly updated. It now includes a detailed 'Pass Design' section with match/replacement tables for all patterns, an explanation for Qwen3-VL's special handling, and a dedicated section for SequenceParallelismAllgatherEpPass with an overview and diagram.
  • Conditional FlashComm V1 and SP Enablement: Logic has been updated in custom operations and platform configuration to conditionally enable FlashComm V1 and handle SP-related communication based on the new pass-based SP enablement flag (enable_sp_by_pass), ensuring correct behavior during graph compilation.

🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Changelog
  • docs/source/user_guide/feature_guide/sequence_parallelism.md
    • Added a 'Pass Design' section detailing the new SequenceParallelismPass and SequenceParallelismAllgatherEpPass.
    • Included match and replacement tables for various SP patterns.
    • Provided an explanation for the special handling required by Qwen3-VL models.
    • Added an overview and diagram for the AllGather EP computation graph and cleanup.
  • vllm_ascend/compilation/graph_fusion_pass_manager.py
    • Updated the configure method to append SequenceParallelismPass and SequenceParallelismAllgatherEpPass when sequence parallelism is enabled.
  • vllm_ascend/compilation/passes/allgather_chunk_noop_pass.py
    • Added a new pass file AllGatherChunkNoOpCleanupPass which defines a pattern to fold redundant all_gather and chunk operations, though this pattern is now integrated into SequenceParallelismAllgatherEpPass.
  • vllm_ascend/compilation/passes/noop_elimination.py
    • Added a new pass NoOpEliminationPass to remove no-op view/reshape nodes from the computation graph.
  • vllm_ascend/compilation/passes/sequence_parallelism.py
    • Renamed AscendMiddleAllReduceRMSNormPattern to MiddleAllReduceRMSNormPattern.
    • Renamed AscendLastAllReduceRMSNormPattern to LastAllReduceRMSNormPattern.
    • Renamed AscendQwen3VLMiddleAllReduceRMSNormPattern to Qwen3VLMiddleAllReduceRMSNormPattern.
    • Renamed AscendSequenceParallelismPass to SequenceParallelismPass.
    • Imported and integrated NoOpEliminationPass to run before applying SP patterns.
  • vllm_ascend/compilation/passes/sequence_parallelism_moe.py
    • Added a new file to define SequenceParallelismAllgatherEpPass.
    • Defined MiddleLayerAllgatherAddRMSNormPattern to postpone all_gather for middle layers.
    • Defined LastLayerAllgatherRMSNormPattern for last layers with all_gather + RMSNorm fusion.
    • Defined Qwen3VLMiddleLayerAllgatherAddRMSNormPattern for Qwen3-VL specific all_gather paths.
    • Defined AllGatherChunkNoOpPattern to fold all_gather + sequence_parallel_chunk_impl into an identity operation.
  • vllm_ascend/ops/fused_moe/prepare_finalize.py
    • Modified prepare and finalize methods to check enable_sp_by_pass() in addition to enable_sp() for conditional logic.
  • vllm_ascend/ops/register_custom_ops.py
    • Updated _maybe_all_gather_and_maybe_unpad_impl to consider enable_sp_by_pass() for FlashComm V1 enablement and added a TODO for unpadding.
    • Updated _maybe_pad_and_reduce_impl to consider enable_sp_by_pass() for FlashComm V1 enablement and added a TODO for padding.
    • Updated _maybe_pad_and_reduce_fake to consider enable_sp_by_pass().
  • vllm_ascend/patch/worker/patch_distributed.py
    • Added an all_gather method to the GroupCoordinatorPatch class.
  • vllm_ascend/platform.py
    • Modified check_and_update_config to conditionally set parallel_config.all2all_backend based on whether SP is enabled via passes.
  • vllm_ascend/utils.py
    • Added a global variable _ENABLE_SP_BY_PASS for caching the SP enablement status.
    • Refactored enable_sp_by_pass function for proper initialization, caching, and handling of None for vllm_config.
  • vllm_ascend/worker/worker.py
    • Modified the profile method signature to include an optional profile_prefix parameter.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@realliujiaxu
realliujiaxu marked this pull request as draft March 6, 2026 08:58

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request refactors the sequence parallelism (SP) pass architecture by splitting it into two passes, SequenceParallelismPass and SequenceParallelismAllgatherEpPass, and reorganizes the associated patterns. The documentation for SP is also significantly improved. The changes are mostly well-structured, but there are a few critical issues. I've found some incomplete logic marked with TODO comments in custom ops that could lead to bugs, an unused parameter in a method, and an obsolete file that should be removed. Please address these points to ensure the correctness and quality of the code.

I am having trouble creating individual review comments. Click here to see my feedback.

vllm_ascend/ops/register_custom_ops.py (56-57)

critical

The TODO: do unpad comment indicates that the unpadding logic is missing when enable_sp_by_pass() is true. This will cause the function to return a padded tensor, which can lead to shape mismatches and incorrect results in downstream operations. This should be implemented to ensure correctness.

vllm_ascend/ops/register_custom_ops.py (93-94)

critical

The TODO: do pad comment indicates that the padding logic is missing when enable_sp_by_pass() is true. The function performs reduce_scatter without padding the input tensor x. This can lead to incorrect behavior or errors if the downstream operations expect a padded tensor. The padding logic should be implemented here.

vllm_ascend/compilation/passes/allgather_chunk_noop_pass.py (1-40)

high

This file seems to be obsolete. The PR description states that AllGatherChunkNoOpCleanupPass is merged into SequenceParallelismAllgatherEpPass as AllGatherChunkNoOpPattern. The new file vllm_ascend/compilation/passes/sequence_parallelism_moe.py already contains AllGatherChunkNoOpPattern. This file allgather_chunk_noop_pass.py defining AllGatherChunkNoOpCleanupPass appears to be a leftover and should be removed to avoid confusion and dead code.

vllm_ascend/worker/worker.py (514)

high

The newly added parameter profile_prefix is not used within the profile method. It should either be used or removed to avoid confusion and maintain clean code.

    def profile(self, is_start: bool = True):

@realliujiaxu realliujiaxu changed the title [Refactor][SP] Refactor sequence parallelism pass architecture and improve docs [Feat][SP] Suport SP for VL MoE models Mar 9, 2026
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@realliujiaxu
realliujiaxu force-pushed the moe-sp-allgather branch 2 times, most recently from e3a3eb2 to 07aa792 Compare March 12, 2026 03:48
@realliujiaxu
realliujiaxu force-pushed the moe-sp-allgather branch 2 times, most recently from a140561 to 0d39195 Compare March 12, 2026 04:55
@realliujiaxu
realliujiaxu marked this pull request as ready for review March 12, 2026 11:43
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Made-with: Cursor
Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Made-with: Cursor
Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Made-with: Cursor
Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Keep the three sequence parallelism MoE cases as separate pytest items while reusing a single distributed worker setup. This cuts repeated initialization cost and updates CI timing to reflect the longer multicard runtime.

Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Made-with: Cursor
Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Comment thread vllm_ascend/patch/worker/patch_distributed.py

@wxsIcey wxsIcey left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Not familiar with the details of the SP, but from pass perspective, approve.

Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Signed-off-by: realliujiaxu <realliujiaxu@163.com>
@realliujiaxu
realliujiaxu merged commit 5d12446 into vllm-project:main Mar 24, 2026
54 of 55 checks passed
starmountain1997 pushed a commit to starmountain1997/vllm-ascend that referenced this pull request Mar 25, 2026
### What this PR does / why we need it?

2nd PR for vllm-project#5712,
extend SP to VL MoE models.


### Does this PR introduce _any_ user-facing change?
remove `sp_threshold` in additional config and reuse `sp_min_token_num`
from vLLM.


### How was this patch tested?
- Model: Qwen3-VL-30B-A3B, 
- TP4 DP2
- 100 reqs
- max concurrency 1

| Seq length | Mean TTFT (ms) main | Mean TTFT (ms) this PR |
|------------|---------------------|------------------------|
| 4k         | 429.40               | 323.3                  |
| 16k        | 1297.01              | 911.74                |

- vLLM version: v0.16.0
- vLLM main:
vllm-project/vllm@4034c3d

---------

Signed-off-by: realliujiaxu <realliujiaxu@163.com>
845473182 pushed a commit to 845473182/vllm-ascend that referenced this pull request Mar 25, 2026
…to qwen3next_graph

* 'main' of https://github.com/vllm-project/vllm-ascend: (94 commits)
  [bugfix] Fixed the error issue when overlaying MTP and full decode on DSV3.1 C8. (vllm-project#7571)
  [eagle3][pcp] fix acceptance rate for eagle3 and pcp enabled (vllm-project#7549)
  [bugfix][CI] fix '_OpNamespace' 'vllm' object has no attribute 'qkv_rmsnorm_rope' (vllm-project#7620)
  [Nightly] Nightly pre-build image (vllm-project#7388)
  [Bugfix]Fix deepseek 3.2 C8  precision by rotary tensor (vllm-project#7537)
  adapt to main2main for model runner v2 (vllm-project#7578)
  [Patch] Fix balance scheduling (vllm-project#7611)
  [310P]fused recurrent gated delta rule pytorch core and ut (vllm-project#7398)
  [CI] refine issue triage rules, wan regex and update stale setting (vllm-project#7531)
  [Lint]Add lint hooks for clang-format, shellcheck, forbidden imports, and boolean context manager checks (vllm-project#7511)
  [doc] add enable_sparse_c8 option in configuration options (vllm-project#7600)
  lower log level in PD Disaggregation (vllm-project#7589)
  [model_runner_v2]:optimize the performance of the _compute_slot_mappings_kernel (vllm-project#7575)
  [Feat][SP] Suport SP for VL MoE models (vllm-project#7044)
  Fix  Qwen3Next CI Config (vllm-project#7561)
  [Feat] Add npugraph_ex enablement logging (vllm-project#7574)
  [UT] Align input arguments with Ascend(Yarn)RotaryEmbedding with vLLM and add ut (vllm-project#7358)
  [P/D] Check wildcard  address for layerwise connector (vllm-project#7389)
  [P/D] [Bugfix] fix mooncake layerconnector dead when update_decoder_info fail (vllm-project#7514)
  [BugFix][P/D] fix padding error on FullGraph mode && fix layerwise connector mamba accuracy (vllm-project#7506)
  ...
lihaokun-2026 pushed a commit to lihaokun-2026/vllm-ascend that referenced this pull request Mar 29, 2026
### What this PR does / why we need it?

2nd PR for vllm-project#5712,
extend SP to VL MoE models.


### Does this PR introduce _any_ user-facing change?
remove `sp_threshold` in additional config and reuse `sp_min_token_num`
from vLLM.


### How was this patch tested?
- Model: Qwen3-VL-30B-A3B, 
- TP4 DP2
- 100 reqs
- max concurrency 1

| Seq length | Mean TTFT (ms) main | Mean TTFT (ms) this PR |
|------------|---------------------|------------------------|
| 4k         | 429.40               | 323.3                  |
| 16k        | 1297.01              | 911.74                |

- vLLM version: v0.16.0
- vLLM main:
vllm-project/vllm@4034c3d

---------

Signed-off-by: realliujiaxu <realliujiaxu@163.com>
chenchuw886 pushed a commit to chenchuw886/vllm-ascend that referenced this pull request Apr 1, 2026
### What this PR does / why we need it?

2nd PR for vllm-project#5712,
extend SP to VL MoE models.


### Does this PR introduce _any_ user-facing change?
remove `sp_threshold` in additional config and reuse `sp_min_token_num`
from vLLM.


### How was this patch tested?
- Model: Qwen3-VL-30B-A3B, 
- TP4 DP2
- 100 reqs
- max concurrency 1

| Seq length | Mean TTFT (ms) main | Mean TTFT (ms) this PR |
|------------|---------------------|------------------------|
| 4k         | 429.40               | 323.3                  |
| 16k        | 1297.01              | 911.74                |

- vLLM version: v0.16.0
- vLLM main:
vllm-project/vllm@4034c3d

---------

Signed-off-by: realliujiaxu <realliujiaxu@163.com>
zouyida2052 pushed a commit to zouyida2052/vllm-ascend that referenced this pull request Apr 28, 2026
### What this PR does / why we need it?

2nd PR for vllm-project#5712,
extend SP to VL MoE models.

### Does this PR introduce _any_ user-facing change?
remove `sp_threshold` in additional config and reuse `sp_min_token_num`
from vLLM.

### How was this patch tested?
- Model: Qwen3-VL-30B-A3B,
- TP4 DP2
- 100 reqs
- max concurrency 1

| Seq length | Mean TTFT (ms) main | Mean TTFT (ms) this PR |
|------------|---------------------|------------------------|
| 4k         | 429.40               | 323.3                  |
| 16k        | 1297.01              | 911.74                |

- vLLM version: v0.16.0
- vLLM main:
vllm-project/vllm@4034c3d

---------

Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Signed-off-by: zouyida2052 <zouyida2002@gmail.com>
yangzhe-2026 pushed a commit to yangzhe-2026/vllm-ascend that referenced this pull request May 6, 2026
### What this PR does / why we need it?

2nd PR for vllm-project#5712,
extend SP to VL MoE models.


### Does this PR introduce _any_ user-facing change?
remove `sp_threshold` in additional config and reuse `sp_min_token_num`
from vLLM.


### How was this patch tested?
- Model: Qwen3-VL-30B-A3B, 
- TP4 DP2
- 100 reqs
- max concurrency 1

| Seq length | Mean TTFT (ms) main | Mean TTFT (ms) this PR |
|------------|---------------------|------------------------|
| 4k         | 429.40               | 323.3                  |
| 16k        | 1297.01              | 911.74                |

- vLLM version: v0.16.0
- vLLM main:
vllm-project/vllm@4034c3d

---------

Signed-off-by: realliujiaxu <realliujiaxu@163.com>
nanxingMy pushed a commit to nanxingMy/vllm-ascend that referenced this pull request May 15, 2026
### What this PR does / why we need it?

2nd PR for vllm-project#5712,
extend SP to VL MoE models.


### Does this PR introduce _any_ user-facing change?
remove `sp_threshold` in additional config and reuse `sp_min_token_num`
from vLLM.


### How was this patch tested?
- Model: Qwen3-VL-30B-A3B, 
- TP4 DP2
- 100 reqs
- max concurrency 1

| Seq length | Mean TTFT (ms) main | Mean TTFT (ms) this PR |
|------------|---------------------|------------------------|
| 4k         | 429.40               | 323.3                  |
| 16k        | 1297.01              | 911.74                |

- vLLM version: v0.16.0
- vLLM main:
vllm-project/vllm@4034c3d

---------

Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Signed-off-by: nanxing <1014662416@qq.com>
ader47 pushed a commit to ader47/vllm-ascend that referenced this pull request Jun 18, 2026
### What this PR does / why we need it?

2nd PR for vllm-project#5712,
extend SP to VL MoE models.


### Does this PR introduce _any_ user-facing change?
remove `sp_threshold` in additional config and reuse `sp_min_token_num`
from vLLM.


### How was this patch tested?
- Model: Qwen3-VL-30B-A3B, 
- TP4 DP2
- 100 reqs
- max concurrency 1

| Seq length | Mean TTFT (ms) main | Mean TTFT (ms) this PR |
|------------|---------------------|------------------------|
| 4k         | 429.40               | 323.3                  |
| 16k        | 1297.01              | 911.74                |

- vLLM version: v0.16.0
- vLLM main:
vllm-project/vllm@4034c3d

---------

Signed-off-by: realliujiaxu <realliujiaxu@163.com>
CXY-Katrina pushed a commit to CXY-Katrina/vllm-ascend that referenced this pull request Jun 27, 2026
### What this PR does / why we need it?

2nd PR for vllm-project#5712,
extend SP to VL MoE models.


### Does this PR introduce _any_ user-facing change?
remove `sp_threshold` in additional config and reuse `sp_min_token_num`
from vLLM.


### How was this patch tested?
- Model: Qwen3-VL-30B-A3B, 
- TP4 DP2
- 100 reqs
- max concurrency 1

| Seq length | Mean TTFT (ms) main | Mean TTFT (ms) this PR |
|------------|---------------------|------------------------|
| 4k         | 429.40               | 323.3                  |
| 16k        | 1297.01              | 911.74                |

- vLLM version: v0.16.0
- vLLM main:
vllm-project/vllm@4034c3d

---------

Signed-off-by: realliujiaxu <realliujiaxu@163.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation module:core module:ops

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants