Skip to content

[Feat.][310P] addrmsnorm for 300I DUO - #6704

Merged
wangxiyuan merged 4 commits into
vllm-project:mainfrom
Tflowers-0129:rmsnorm
Feb 13, 2026
Merged

wangxiyuan merged 4 commits into
vllm-project:mainfrom
Tflowers-0129:rmsnorm

Conversation

@Tflowers-0129

@Tflowers-0129 Tflowers-0129 commented Feb 11, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

This PR integrates the npu_add_rms_norm fused kernel for RMSNorm operations with residual connections on 310P devices. This change optimizes the computation by replacing a two-step process (manual residual addition followed by RMSNorm) with a single, more efficient fused operation. This is needed to improve the performance of models utilizing RMSNorm with residual connections on the 310P architecture.

Fixes #

Does this PR introduce any user-facing change?

No, this PR introduces an internal optimization and does not change any user-facing APIs or behaviors.

How was this patch tested?

This patch was tested with updated unit tests (test_RMSNorm_forward_310p) that mock the npu_add_rms_norm operation to verify the correctness of the fused kernel integration.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @Tflowers-0129, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces an optimization for the RMSNorm layer specifically targeting 310P Ascend devices. By integrating the torch_npu.npu_add_rms_norm operation, the system can now perform residual addition and RMS normalization in a single, more efficient NPU-accelerated step when a residual connection is present. This change aims to enhance performance for models utilizing RMSNorm with residual connections on the specified hardware, with corresponding updates to unit tests to validate the new implementation.

Highlights

  • Optimized RMSNorm for 310P: The RMSNorm layer for 310P devices now utilizes torch_npu.npu_add_rms_norm to efficiently combine residual addition and RMS normalization into a single NPU operation.
  • Improved Residual Handling: The forward_oot method in the 310P layernorm implementation was refactored to directly use the optimized npu_add_rms_norm when a residual tensor is provided, streamlining the computation and adding bias handling.
  • Test Case Updates: The unit tests for RMSNorm on 310P hardware (test_RMSNorm_forward_310p) were updated to patch and assert against the new npu_add_rms_norm function, ensuring correct behavior of the optimized path.
  • Code Cleanup: An unused import (AscendDeviceType) was removed from the test_layernorm.py file, and minor formatting adjustments were made to a pytest.mark.skip decorator.

🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Changelog
  • tests/ut/ops/test_layernorm.py
    • Removed an unused import statement for AscendDeviceType.
    • Adjusted the formatting of a pytest.mark.skip decorator.
    • Modified test_RMSNorm_forward_310p to patch torch_npu.npu_add_rms_norm and updated assertions to reflect the combined operation.
  • vllm_ascend/_310p/ops/layernorm.py
    • Replaced the manual residual addition and subsequent torch_npu.npu_rms_norm call with a single call to torch_npu.npu_add_rms_norm when a residual is present.
    • Added logic to apply bias after npu_add_rms_norm if self.bias is not None.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a significant optimization for RMSNorm operations on 310P devices by leveraging the npu_add_rms_norm fused kernel. This change replaces a manual addition of residual and a separate RMSNorm call with a single, more efficient fused operation, which should improve performance. The corresponding unit tests have been updated to reflect this new behavior and correctly mock the fused operation. Additionally, the pull request title and summary can be improved to adhere to the repository's style guide for better clarity and consistency.

Suggested PR Title:

[Ops][Feature] Add fused RMSNorm for 310P DUO

Suggested PR Summary:

### What this PR does / why we need it?
This PR integrates the `npu_add_rms_norm` fused kernel for RMSNorm operations with residual connections on 310P devices. This change optimizes the computation by replacing a two-step process (manual residual addition followed by RMSNorm) with a single, more efficient fused operation. This is needed to improve the performance of models utilizing RMSNorm with residual connections on the 310P architecture.

Fixes #

### Does this PR introduce _any_ user-facing change?
No, this PR introduces an internal optimization and does not change any user-facing APIs or behaviors.

### How was this patch tested?
This patch was tested with updated unit tests (`test_RMSNorm_forward_310p`) that mock the `npu_add_rms_norm` operation to verify the correctness of the fused kernel integration.

Comment thread tests/ut/ops/test_layernorm.py
Comment thread tests/ut/ops/test_layernorm.py
Comment thread vllm_ascend/_310p/ops/layernorm.py
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
@wangxiyuan
wangxiyuan merged commit f40256b into vllm-project:main Feb 13, 2026
25 checks passed
845473182 pushed a commit to 845473182/vllm-ascend that referenced this pull request Feb 24, 2026
…ascend into qwen3next_rebase

* 'qwen3next_rebase' of https://github.com/845473182/vllm-ascend:
  [Bugfix][DispatchFFNCombine] resolve vec error caused by unaligned UB access (vllm-project#6707)
  [Lint] Adapt lint tools for windows (vllm-project#6727)
  [main][Docs] Fix typos across documentation (vllm-project#6728)
  [Feat.][310P]: weightNZ feature with quant or unquant. (vllm-project#6705)
  [Feat.][310P] addrmsnorm for 300I DUO (vllm-project#6704)
  [Graph][Fusion] Integrating inductor pass and npugraph ex pass (vllm-project#6354)
  [bugfix] adapt bugfix for norm_quant_fusion_pass to npugraph_ex (vllm-project#6726)
  [doc] add A2 series doc for GLM5.md (vllm-project#6717)
@wangxiyuan wangxiyuan mentioned this pull request Feb 24, 2026
HF-001 pushed a commit to HF-001/vllm-ascend that referenced this pull request Feb 25, 2026
### What this PR does / why we need it?
This PR integrates the `npu_add_rms_norm` fused kernel for RMSNorm
operations with residual connections on 310P devices. This change
optimizes the computation by replacing a two-step process (manual
residual addition followed by RMSNorm) with a single, more efficient
fused operation. This is needed to improve the performance of models
utilizing RMSNorm with residual connections on the 310P architecture.

Fixes #

### Does this PR introduce _any_ user-facing change?
No, this PR introduces an internal optimization and does not change any
user-facing APIs or behaviors.

### How was this patch tested?
This patch was tested with updated unit tests
(`test_RMSNorm_forward_310p`) that mock the `npu_add_rms_norm` operation
to verify the correctness of the fused kernel integration.

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
banxiaduhuo pushed a commit to banxiaduhuo/vllm-ascend that referenced this pull request Feb 26, 2026
### What this PR does / why we need it?
This PR integrates the `npu_add_rms_norm` fused kernel for RMSNorm
operations with residual connections on 310P devices. This change
optimizes the computation by replacing a two-step process (manual
residual addition followed by RMSNorm) with a single, more efficient
fused operation. This is needed to improve the performance of models
utilizing RMSNorm with residual connections on the 310P architecture.

Fixes #

### Does this PR introduce _any_ user-facing change?
No, this PR introduces an internal optimization and does not change any
user-facing APIs or behaviors.

### How was this patch tested?
This patch was tested with updated unit tests
(`test_RMSNorm_forward_310p`) that mock the `npu_add_rms_norm` operation
to verify the correctness of the fused kernel integration.

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
maoxx241 pushed a commit to maoxx241/vllm-ascend that referenced this pull request Mar 2, 2026
### What this PR does / why we need it?
This PR integrates the `npu_add_rms_norm` fused kernel for RMSNorm
operations with residual connections on 310P devices. This change
optimizes the computation by replacing a two-step process (manual
residual addition followed by RMSNorm) with a single, more efficient
fused operation. This is needed to improve the performance of models
utilizing RMSNorm with residual connections on the 310P architecture.

Fixes #

### Does this PR introduce _any_ user-facing change?
No, this PR introduces an internal optimization and does not change any
user-facing APIs or behaviors.

### How was this patch tested?
This patch was tested with updated unit tests
(`test_RMSNorm_forward_310p`) that mock the `npu_add_rms_norm` operation
to verify the correctness of the fused kernel integration.

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
LCAIZJ pushed a commit to LCAIZJ/vllm-ascend that referenced this pull request Mar 7, 2026
### What this PR does / why we need it?
This PR integrates the `npu_add_rms_norm` fused kernel for RMSNorm
operations with residual connections on 310P devices. This change
optimizes the computation by replacing a two-step process (manual
residual addition followed by RMSNorm) with a single, more efficient
fused operation. This is needed to improve the performance of models
utilizing RMSNorm with residual connections on the 310P architecture.

Fixes #

### Does this PR introduce _any_ user-facing change?
No, this PR introduces an internal optimization and does not change any
user-facing APIs or behaviors.

### How was this patch tested?
This patch was tested with updated unit tests
(`test_RMSNorm_forward_310p`) that mock the `npu_add_rms_norm` operation
to verify the correctness of the fused kernel integration.

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
yangzhe-2026 pushed a commit to yangzhe-2026/vllm-ascend that referenced this pull request May 6, 2026
### What this PR does / why we need it?
This PR integrates the `npu_add_rms_norm` fused kernel for RMSNorm
operations with residual connections on 310P devices. This change
optimizes the computation by replacing a two-step process (manual
residual addition followed by RMSNorm) with a single, more efficient
fused operation. This is needed to improve the performance of models
utilizing RMSNorm with residual connections on the 310P architecture.

Fixes #

### Does this PR introduce _any_ user-facing change?
No, this PR introduces an internal optimization and does not change any
user-facing APIs or behaviors.

### How was this patch tested?
This patch was tested with updated unit tests
(`test_RMSNorm_forward_310p`) that mock the `npu_add_rms_norm` operation
to verify the correctness of the fused kernel integration.

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
nanxingMy pushed a commit to nanxingMy/vllm-ascend that referenced this pull request May 15, 2026
### What this PR does / why we need it?
This PR integrates the `npu_add_rms_norm` fused kernel for RMSNorm
operations with residual connections on 310P devices. This change
optimizes the computation by replacing a two-step process (manual
residual addition followed by RMSNorm) with a single, more efficient
fused operation. This is needed to improve the performance of models
utilizing RMSNorm with residual connections on the 310P architecture.

Fixes #

### Does this PR introduce _any_ user-facing change?
No, this PR introduces an internal optimization and does not change any
user-facing APIs or behaviors.

### How was this patch tested?
This patch was tested with updated unit tests
(`test_RMSNorm_forward_310p`) that mock the `npu_add_rms_norm` operation
to verify the correctness of the fused kernel integration.

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: nanxing <1014662416@qq.com>
ader47 pushed a commit to ader47/vllm-ascend that referenced this pull request Jun 18, 2026
### What this PR does / why we need it?
This PR integrates the `npu_add_rms_norm` fused kernel for RMSNorm
operations with residual connections on 310P devices. This change
optimizes the computation by replacing a two-step process (manual
residual addition followed by RMSNorm) with a single, more efficient
fused operation. This is needed to improve the performance of models
utilizing RMSNorm with residual connections on the 310P architecture.

Fixes #

### Does this PR introduce _any_ user-facing change?
No, this PR introduces an internal optimization and does not change any
user-facing APIs or behaviors.

### How was this patch tested?
This patch was tested with updated unit tests
(`test_RMSNorm_forward_310p`) that mock the `npu_add_rms_norm` operation
to verify the correctness of the fused kernel integration.

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
CXY-Katrina pushed a commit to CXY-Katrina/vllm-ascend that referenced this pull request Jun 27, 2026
### What this PR does / why we need it?
This PR integrates the `npu_add_rms_norm` fused kernel for RMSNorm
operations with residual connections on 310P devices. This change
optimizes the computation by replacing a two-step process (manual
residual addition followed by RMSNorm) with a single, more efficient
fused operation. This is needed to improve the performance of models
utilizing RMSNorm with residual connections on the 310P architecture.

Fixes #

### Does this PR introduce _any_ user-facing change?
No, this PR introduces an internal optimization and does not change any
user-facing APIs or behaviors.

### How was this patch tested?
This patch was tested with updated unit tests
(`test_RMSNorm_forward_310p`) that mock the `npu_add_rms_norm` operation
to verify the correctness of the fused kernel integration.

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants