Skip to content

[BugFix]keep padded layout through EP SP gather/reduce custom ops … - #17759

Open
Xuyzhen wants to merge 2 commits into
vllm-project:mainfrom
Xuyzhen:dts/6867/fix_prefill_v2
Open

Xuyzhen wants to merge 2 commits into
vllm-project:mainfrom
Xuyzhen:dts/6867/fix_prefill_v2

Conversation

@Xuyzhen

@Xuyzhen Xuyzhen commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

…(rc1 sp_by_pass restore)

What this PR does / why we need it?

Restores the rc1 sp_by_pass semantics for the EP/SP MoE dispatch/finalize custom ops, fixing a prefill-side performance regression.

Problem. In _maybe_all_gather_and_maybe_unpad_impl, after padding each rank's input to max_local_size and all-gathering, the SP path unpacks the gathered tensor with a per-shard python slicing loop + torch.cat. In _maybe_pad_and_reduce_impl, the finalize path re-packs the tensor with new_zeros + a copy loop before reduce_scatter. This pack/unpack pair runs on the MoE hot path (3 gather calls per MoE layer, ~60 layers per prefill) and is pure host-side overhead. Measured on DeepSeek V3.1-terminated W4A8, TP8+SP, 16K/1K mixed workload (Atlas A2, ~100s profile window):

  • ~9.8s extra host time per window (10% of one host CPU)
  • ~32% of it converts to wall time (~3.1s) by congesting the host queue and delaying collective submission of every following op
  • end-to-end: -4.1% overall TPS, TTFT +6.8%, first-prefill window +10.6%
  • victim ops with unchanged call counts inflate 7-28% on average (e.g. aclnnMatmul 6.9x, aclnnKvRmsNormRope 8.9x avg-duration inflation)

Fix. Once the input is padded to max_local_size for the all_gather, keep the padded layout all the way to the paired reduce_scatter:

  • gather: pad, all_gather, return the gathered tensor as-is (no unpad)
  • reduce: reduce_scatter(x.view(-1, ...)) directly (zero-alloc, zero-copy), then slice the pad tail off this rank's shard
  • the gather fake shape now matches len(local_sizes) * max(local_sizes)

The extra pad rows computed by fast ranks are wall-clock free (they already wait for the slow rank at the collective), and the numerical output is unchanged: pad rows are dropped before the result leaves the reduce op.

Consumer audit (why the padded layout is safe downstream):

  • main's all_gather_input_ids() gathers input_ids through the same op, so its rows stay row-aligned with the padded hidden states; hash routing is row-independent and masks pad rows (input_ids == -1)
  • shared-expert gather slicing uses _EXTRA_CTX.num_tokens (attention-side metadata), not the prepare-side token count, so it is unaffected
  • dispatcher shapes are computed from live tensor shapes
  • the DP (non-SP) unpad/repack paths are untouched

Does this PR introduce any user-facing change?

No.

How was this patch tested?

  • tests/ut/ops/test_register_custom_ops.py updated to the padded-layout semantics, including the draft-model EP-layout regression test and the PCP (DP x PCP x TP) token-order round-trip test
  • NPU benchmark (DeepSeek V3.1-terminated W4A8, TP8+SP, 16K/1K mixed, averaged over repeated runs): 323.15 TPS before vs 337.9 TPS after, against an rc1 baseline of ~333-338 TPS (data collected on the v0.27.1-based equivalent of this patch, commit 833c691; the ported diff is identical in the three touched functions)

Rebase note

Originally targeted v0.27.1 and rebased onto current main to resolve conflicts. Main-side changes audited during the port:

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request optimizes the Expert Parallelism (EP) communication path by adopting a consistent padded layout for gather and reduce operations. By keeping tensors in a padded state (ep_size, max_local_size) throughout the process, the implementation removes redundant host-side operations like slicing and concatenation. While this approach may introduce minor redundant computations on faster ranks, it significantly improves overall performance by eliminating overhead in the hot path of MoE layers.

Highlights

  • EP/SP Communication Optimization: Refactored Expert Parallelism (EP) communication to maintain a padded layout throughout the gather/reduce process, eliminating costly per-shard Python slicing and concatenation operations.
  • Performance Improvement: By avoiding unpadding during gather and padding during reduce, the implementation reduces host-side overhead for MoE layers, particularly beneficial for prefill paths.
  • Test Suite Updates: Updated unit tests in test_register_custom_ops.py to reflect the new padded-layout semantics and ensure correct handling of local token chunks.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.


Tip

💡 Consider Linking a Related Issue or RFC

Your PR title contains the [BugFix] tag, indicating a bug fix or new feature.

Linking a related issue or RFC in the PR description is strongly encouraged — it gives reviewers helpful context and speeds up the review. You can use any of these keywords:

  • Fixes #<issue_number>
  • Closes #<issue_number>
  • Resolves #<issue_number>
  • Refs #<rfc_or_issue_number> (for RFCs)

🙏 Thanks for helping us keep the project well-organized!

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Optimize EP communication by keeping padded layout in MoE

Suggested PR Summary:

### What this PR does / why we need it?
This PR optimizes the Expert Parallel (EP) communication path in MoE by keeping the tensors in a padded layout throughout the MoE execution (v2 padded-layout semantics). Instead of unpadding after `all_gather` and repadding before `reduce_scatter`, the implementation now keeps the padded layout. This completely eliminates the host overhead associated with per-shard Python slicing and concatenation loops, which is particularly beneficial for deep networks (e.g., 61 layers in prefill hot paths).

Additionally, two critical issues were identified during the review:
1. Potential runtime errors or silent data corruption in `all_gather` if the padded tensor is non-contiguous. It is recommended to call `.contiguous()` on the padded tensor.
2. Potential `RuntimeError` when calling `.view()` on a non-contiguous tensor `x` in `_maybe_pad_and_reduce_impl`. It is recommended to use `.reshape(...).contiguous()` before passing the tensor to `reduce_scatter`.

### Does this PR introduce _any_ user-facing change?
No. This is an internal performance optimization for EP communication.

### How was this patch tested?
The changes were verified by updating the existing unit tests in `tests/ut/ops/test_register_custom_ops.py` to match the new padded-layout shapes and assertions.

# gather 后再按各 rank 真实的 local_sizes 截回。
# all_gather 要求各 rank 输入等长: pad 到 max_local_size。
# 等分片时 _pad_to_ep_local_size 零开销原样返回。
x = _pad_to_ep_local_size(x, max_local_size)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

If x is already of size max_local_size, _pad_to_ep_local_size returns x directly. If x is non-contiguous (e.g., from sequence parallel slicing), passing it directly to all_gather can cause silent data corruption or runtime errors. Ensuring contiguity here is highly recommended.

Suggested change
x = _pad_to_ep_local_size(x, max_local_size)
x = _pad_to_ep_local_size(x, max_local_size).contiguous()

reduced = ep_group.reduce_scatter(padded_x.view(-1, *x.shape[1:]), 0)
# 与 v2 gather 配对: x 行数 = len(local_sizes) * max(local_sizes),
# 每 rank 的真实 token 位于其分片头部, pad 尾部会被 slice 丢弃。
reduced = ep_group.reduce_scatter(x.view(-1, *x.shape[1:]), 0)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Since x is the output of MoE layers, it may not be contiguous. Calling .view() on a non-contiguous tensor will raise a RuntimeError. Furthermore, collective communication operations like reduce_scatter require contiguous input tensors to avoid silent data corruption or runtime crashes. Use .reshape() and ensure contiguity by calling .contiguous() before passing to reduce_scatter.

Suggested change
reduced = ep_group.reduce_scatter(x.view(-1, *x.shape[1:]), 0)
reduced = ep_group.reduce_scatter(x.reshape(-1, *x.shape[1:]).contiguous(), 0)

…tom ops

The EP/SP dispatch path unpads the all_gather output with a per-shard
python slicing loop + torch.cat, and the finalize path re-pads it with
new_zeros + a copy loop before reduce_scatter. On the MoE prefill hot
path (3 gather calls per MoE layer) this host-side packing/unpacking
costs ~10% extra host time per forward and congests the collective
queue (measured on DeepSeek-V3.1 W4A8 TP8/SP: -4.1% overall TPS, TTFT
+6.8%).

Restore the rc1 sp_by_pass semantics: once the input is padded to
max_local_size for the all_gather, keep the padded layout all the way
to the paired reduce_scatter, which then runs zero-alloc / zero-copy
and only slices the pad tail off its own shard. Extra pad rows on fast
ranks are wall-clock free (they wait for the slow rank at the
collective anyway).

Notes:
- the gather fake shape now matches len(local_sizes) * max(local_sizes)
- draft models and PCP-expanded layouts keep the same code path
- the DP (non-SP) unpad/repack paths are untouched

Signed-off-by: Xuyzhen <958522639@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant