Repository navigation
Conversation
…m-project#15737) Completes the `sequence_parallelism.md` feature guide, which currently only has an Overview skeleton plus a stale "FlashComm is deprecated" note. The new content covers SP MoE end to end: - Principle: upstream `ParallelConfig.use_sequence_parallel_moe` activation conditions (TP>1, DP>1, `enable_expert_parallel`, SP-capable `all2all_backend`), the per-layer EP all-gather/unpad -> MoE -> pad/reduce-scatter -> TP all-gather flow, and the compile-time (`pass_config.enable_sp`, `sp_min_token_num`, `enable_sp_by_pass`) / run-time (TP-aligned cudagraph sizes) behavior. - How to use: upstream serve flags, constraints (TP>1, EP required for MoE, TP-multiple capture sizes, PCP incompatibility). - The temporary Ascend-only FlashComm switch: by default the platform forces `all2all_backend=flashinfer_all2allv` (SP MoE off); setting `additional_config.enable_flashcomm1` (preferred) or `VLLM_ASCEND_ENABLE_FLASHCOMM1=1` opts into upstream SP MoE. Documents that the switch is temporary/deprecated and will be removed once SP is supported. Documentation only. No code behavior change. - `markdownlint docs/source/user_guide/feature_guide/sequence_parallelism.md` passes. - No code changed, so no unit/e2e tests apply. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: XuRongSheng <1843167357@qq.com> (cherry picked from commit 4980809)
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request backports essential documentation regarding Sequence Parallelism (SP) for Mixture-of-Experts (MoE) models to the v0.27.1rc release branch. It clarifies the operational principles and configuration requirements for SP MoE on Ascend hardware, while providing clear guidance on the temporary, deprecated use of the FlashComm switch for compatibility. The changes are strictly limited to documentation and code comments, ensuring no functional changes to the runtime behavior. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Doc][Misc] Update Sequence Parallelism and FlashComm documentation and translate commentsSuggested PR Summary:
### What this PR does / why we need it?
This PR updates the documentation for Sequence Parallelism (SP) and FlashComm on Ascend, providing a detailed explanation of SP MoE, its principle, data flow, and usage. It documents new configuration options including `enable_flashcomm1` and `enable_pcp_o_proj_weight_sharding`. Additionally, it translates Chinese code comments to English in `vllm_ascend/ops/register_custom_ops.py`.
### Does this PR introduce _any_ user-facing change?
Yes, documentation has been updated to reflect new configuration options and the deprecation of FlashComm.
### How was this patch tested?
No functional code changes were introduced; only documentation and comments were updated.I have no review comments to evaluate.
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
Signed-off-by: XuRongSheng <1843167357@qq.com>
a789214 to
4dd8eab
Compare
Signed-off-by: Xu Rongsheng <1843167357@qq.com>
What this PR does / why we need it?
Backports the SP MoE feature guide to the
releases/v0.27.1rcbranch(cherry-pick of #15737 from
main).The
sequence_parallelism.mdguide on the release branch is still theold Overview skeleton with a stale "FlashComm is deprecated" note. This
brings over the completed end-to-end content:
ParallelConfig.use_sequence_parallel_moeactivationconditions, per-layer EP all-gather/unpad → MoE → pad/reduce-scatter
→ TP all-gather flow, compile-time and run-time behavior.
TP-multiple capture sizes, PCP incompatibility).
(
additional_config.enable_flashcomm1preferred,VLLM_ASCEND_ENABLE_FLASHCOMM1=1for compatibility), documented astemporary/deprecated.
register_custom_ops.py(no behavior change).
Conflict note:
additional_config.mdhad a release-sideenable_dsa_cprow; resolved by taking the newer
mainwording (adds theauto-enables-FlashComm sentence) plus the two new option rows.
Does this PR introduce any user-facing change?
Documentation only. No code behavior change.
How was this patch tested?
markdownlint --config .markdownlint.yamlpasses on both touchedmarkdown files.
ruff checkpasses onregister_custom_ops.py(comment-only change).No code changed, so no unit/e2e tests apply.
vLLM main: vllm-project/vllm@ba07e4a