Skip to content

Revert "GMM custom operator optimization in small batch scenarios (vllm-project#7100)" - #7557

Merged
MengqingCao merged 1 commit into
vllm-project:mainfrom
SunnyLee151064:revert_7100
Mar 24, 2026
Merged

MengqingCao merged 1 commit into
vllm-project:mainfrom
SunnyLee151064:revert_7100

Conversation

@SunnyLee151064

@SunnyLee151064 SunnyLee151064 commented Mar 23, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

This reverts commit 42bcad7. The commit cause accuracy decrease of qwen3Next, 150 items of gsm8k, 98 -> 91.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request performs a full revert of a prior commit that aimed to optimize Grouped Matrix Multiplication (GMM) custom operators for small batch sizes. The revert removes the specialized logic and operator calls that were conditionally applied based on batch size, restoring the system to its state before the optimization was introduced. This change ensures consistent behavior across all batch sizes by removing the small-batch specific implementations.

Highlights

  • Reverted Optimization: Reverted a previous commit that introduced an optimization for Grouped Matrix Multiplication (GMM) custom operators specifically for small batch scenarios.
  • Conditional Logic Removal: Removed conditional logic that previously selected different MoE routing and grouped matmul implementations based on the batch size and quantization mode.
  • Operator Restoration: Restored the use of aclnnMoeGroupedMatmulWeightNz in the C++ binding and torch_npu.npu_grouped_matmul in Python for MoE operations, undoing the small batch specific operator calls.

🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@SunnyLee151064 SunnyLee151064 changed the title Revert "GMM custom operator optimization in small batch scenarios (#7… Revert PR 7100 for qwen3Next nightly accuracy problem Mar 23, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request reverts commit 42bcad7e9b68cf92d9b98a9072eff10edc8d6f33, which introduced an optimization for GMM custom operators in small batch scenarios. The changes remove the specialized code paths for small batches across C++ and Python files.

While the code changes correctly perform the revert, the pull request description is currently empty. This violates the repository's style guide (lines 12-40), which requires a clear explanation of the changes. For a revert, it is critical to document why the original changes are being undone (e.g., correctness issues, performance regressions). Additionally, the pull request title does not follow the prescribed format (lines 41-61). Please update the description and title to provide this crucial context for reviewers and future reference.

Suggested PR Title:

[Ops][BugFix] Revert GMM custom operator optimization in small batch scenarios

Suggested PR Summary:

### What this PR does / why we need it?

This PR reverts commit 42bcad7e9b68cf92d9b98a9072eff10edc8d6f33.

The reverted commit introduced an optimization for the GMM custom operator in small batch scenarios. This revert is necessary because **[PLEASE FILL IN THE REASON FOR THE REVERT, e.g., it caused correctness issues under certain conditions or led to performance regressions in other scenarios.]**

This change removes the specialized code paths for small batches and restores the previous, stable implementation.

### Does this PR introduce _any_ user-facing change?

No. This change fixes an issue with a previous optimization and is not expected to introduce any user-facing changes, other than restoring correct behavior and/or performance.

### How was this patch tested?

CI passed. As this is a revert to a previously validated state, existing tests are sufficient to ensure correctness.

…lm-project#7100)"

This reverts commit 42bcad7.

Signed-off-by: Your Name <you@example.com>
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@MengqingCao MengqingCao changed the title Revert PR 7100 for qwen3Next nightly accuracy problem Revert "GMM custom operator optimization in small batch scenarios (vllm-project#7100)" Mar 24, 2026
@MengqingCao
MengqingCao merged commit 475b4b0 into vllm-project:main Mar 24, 2026
36 checks passed
starmountain1997 pushed a commit to starmountain1997/vllm-ascend that referenced this pull request Mar 25, 2026
…lm-project#7100)" (vllm-project#7557)

### What this PR does / why we need it?
This reverts commit 42bcad7. The commit
cause accuracy decrease of qwen3Next, 150 items of gsm8k, 98 -> 91.

- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@6a9cceb

Signed-off-by: Your Name <you@example.com>
Co-authored-by: Your Name <you@example.com>
lihaokun-2026 pushed a commit to lihaokun-2026/vllm-ascend that referenced this pull request Mar 29, 2026
…lm-project#7100)" (vllm-project#7557)

### What this PR does / why we need it?
This reverts commit 42bcad7. The commit
cause accuracy decrease of qwen3Next, 150 items of gsm8k, 98 -> 91.

- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@6a9cceb

Signed-off-by: Your Name <you@example.com>
Co-authored-by: Your Name <you@example.com>
chenchuw886 pushed a commit to chenchuw886/vllm-ascend that referenced this pull request Apr 1, 2026
…lm-project#7100)" (vllm-project#7557)

### What this PR does / why we need it?
This reverts commit 42bcad7. The commit
cause accuracy decrease of qwen3Next, 150 items of gsm8k, 98 -> 91.

- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@6a9cceb

Signed-off-by: Your Name <you@example.com>
Co-authored-by: Your Name <you@example.com>
yangzhe-2026 pushed a commit to yangzhe-2026/vllm-ascend that referenced this pull request May 6, 2026
…lm-project#7100)" (vllm-project#7557)

### What this PR does / why we need it?
This reverts commit 42bcad7. The commit
cause accuracy decrease of qwen3Next, 150 items of gsm8k, 98 -> 91.

- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@6a9cceb

Signed-off-by: Your Name <you@example.com>
Co-authored-by: Your Name <you@example.com>
nanxingMy pushed a commit to nanxingMy/vllm-ascend that referenced this pull request May 15, 2026
…lm-project#7100)" (vllm-project#7557)

### What this PR does / why we need it?
This reverts commit 42bcad7. The commit
cause accuracy decrease of qwen3Next, 150 items of gsm8k, 98 -> 91.

- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@6a9cceb

Signed-off-by: Your Name <you@example.com>
Co-authored-by: Your Name <you@example.com>
Signed-off-by: nanxing <1014662416@qq.com>
ader47 pushed a commit to ader47/vllm-ascend that referenced this pull request Jun 18, 2026
…lm-project#7100)" (vllm-project#7557)

### What this PR does / why we need it?
This reverts commit 42bcad7. The commit
cause accuracy decrease of qwen3Next, 150 items of gsm8k, 98 -> 91.

- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@6a9cceb

Signed-off-by: Your Name <you@example.com>
Co-authored-by: Your Name <you@example.com>
CXY-Katrina pushed a commit to CXY-Katrina/vllm-ascend that referenced this pull request Jun 27, 2026
…lm-project#7100)" (vllm-project#7557)

### What this PR does / why we need it?
This reverts commit fa4f8f6. The commit
cause accuracy decrease of qwen3Next, 150 items of gsm8k, 98 -> 91.

- vLLM version: v0.18.0
- vLLM main:
vllm-project/vllm@6a9cceb

Signed-off-by: Your Name <you@example.com>
Co-authored-by: Your Name <you@example.com>
ZT-AIA pushed a commit that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

The op has had no callers since #7557 reverted the small-batch GMM
optimization introduced in #7100 (qwen3-next gsm8k 98 -> 91). All
grouped-matmul paths run on torch_npu.npu_grouped_matmul.

- delete csrc/moe/moe_grouped_matmul/ (op_host, op_kernel, op_api)
- drop schema/impl registration from csrc/torch_binding.cpp
- drop meta registration from csrc/torch_binding_meta.cpp
- drop the op from csrc/build_aclnn.sh build lists (ascend910b/910_93)

### Does this PR introduce _any_ user-facing change?
Yes, the `moe_grouped_matmul` operator is no longer available in the
PyTorch bindings.

### How was this patch tested?
No new tests were added as this is a code removal PR. Existing CI tests
should pass.

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: TangPeng <85704592@qq.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
### What this PR does / why we need it?

The op has had no callers since vllm-project#7557 reverted the small-batch GMM
optimization introduced in vllm-project#7100 (qwen3-next gsm8k 98 -> 91). All
grouped-matmul paths run on torch_npu.npu_grouped_matmul.

- delete csrc/moe/moe_grouped_matmul/ (op_host, op_kernel, op_api)
- drop schema/impl registration from csrc/torch_binding.cpp
- drop meta registration from csrc/torch_binding_meta.cpp
- drop the op from csrc/build_aclnn.sh build lists (ascend910b/910_93)

### Does this PR introduce _any_ user-facing change?
Yes, the `moe_grouped_matmul` operator is no longer available in the
PyTorch bindings.

### How was this patch tested?
No new tests were added as this is a code removal PR. Existing CI tests
should pass.

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: TangPeng <85704592@qq.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
### What this PR does / why we need it?

The op has had no callers since vllm-project#7557 reverted the small-batch GMM
optimization introduced in vllm-project#7100 (qwen3-next gsm8k 98 -> 91). All
grouped-matmul paths run on torch_npu.npu_grouped_matmul.

- delete csrc/moe/moe_grouped_matmul/ (op_host, op_kernel, op_api)
- drop schema/impl registration from csrc/torch_binding.cpp
- drop meta registration from csrc/torch_binding_meta.cpp
- drop the op from csrc/build_aclnn.sh build lists (ascend910b/910_93)

### Does this PR introduce _any_ user-facing change?
Yes, the `moe_grouped_matmul` operator is no longer available in the
PyTorch bindings.

### How was this patch tested?
No new tests were added as this is a code removal PR. Existing CI tests
should pass.

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: TangPeng <85704592@qq.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
### What this PR does / why we need it?

The op has had no callers since vllm-project#7557 reverted the small-batch GMM
optimization introduced in vllm-project#7100 (qwen3-next gsm8k 98 -> 91). All
grouped-matmul paths run on torch_npu.npu_grouped_matmul.

- delete csrc/moe/moe_grouped_matmul/ (op_host, op_kernel, op_api)
- drop schema/impl registration from csrc/torch_binding.cpp
- drop meta registration from csrc/torch_binding_meta.cpp
- drop the op from csrc/build_aclnn.sh build lists (ascend910b/910_93)

### Does this PR introduce _any_ user-facing change?
Yes, the `moe_grouped_matmul` operator is no longer available in the
PyTorch bindings.

### How was this patch tested?
No new tests were added as this is a code removal PR. Existing CI tests
should pass.

- vLLM main:
vllm-project/vllm@b2f6858

Signed-off-by: TangPeng <85704592@qq.com>
Signed-off-by: like-0517 <ithwlike@126.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants