Skip to content

Intergration test - #7972

Closed
Tflowers-0129 wants to merge 32 commits into
vllm-project:mainfrom
Tflowers-0129:intergration-test
Closed

Tflowers-0129 wants to merge 32 commits into
vllm-project:mainfrom
Tflowers-0129:intergration-test

Conversation

@Tflowers-0129

@Tflowers-0129 Tflowers-0129 commented Apr 3, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

Does this PR introduce any user-facing change?

How was this patch tested?

pu-zhe and others added 30 commits March 28, 2026 09:10
Signed-off-by: pu-zhe <zpuaa@outlook.com>
Signed-off-by: pu-zhe <zpuaa@outlook.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: pu-zhe <zpuaa@outlook.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: pu-zhe <zpuaa@outlook.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: pu-zhe <zpuaa@outlook.com>
Signed-off-by: pu-zhe <zpuaa@outlook.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces significant performance optimizations for the 310P backend by integrating native NPU kernels for core operations like gated delta rules and causal convolutions. It also standardizes state tensor layouts and improves quantization efficiency through NZ format casting, ensuring better alignment with NPU hardware capabilities. These changes collectively aim to reduce latency and improve throughput for supported models.

Highlights

  • Integration of NPU-optimized kernels: Integrated optimized NPU kernels for recurrent gated delta rule and causal convolution, replacing previous Python-based implementations for improved performance.
  • State layout and quantization updates: Updated state tensor layouts to [V, K] and implemented NZ format casting for W8A8 dynamic quantization to enhance inference speed on 310P hardware.
  • MoE and Routing improvements: Enabled NPU-native MoE gating for softmax routing and removed unsupported Expert Parallelism for 310P, while updating token dispatch logic.

🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces 310P platform optimizations, including NPU-specific fused kernels for MoE gating and token dispatching, W8A8 dynamic quantization for linear layers, and updated state layouts for Gated Delta Rule operators. Feedback focuses on correcting the order of transpose operations for NZ format conversion, fixing potential shape mismatches in quantization logic, and ensuring the correct token count mode in the dispatcher. Additionally, the PR title and summary must be updated to comply with the repository's style guide.

Comment on lines +198 to +201
if pertoken_scale.dim() == 2:
need_unsqz = True
quantized_x = quantized_x.squeeze(dim=1)
pertoken_scale = pertoken_scale.squeeze(dim=1)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The logic for squeezing dimensions assumes that if pertoken_scale is 2D, the sequence dimension (dim 1) is of size 1. If the input x has a sequence length greater than 1 (e.g., during prefill), squeeze(dim=1) will not remove the dimension if its size is not 1, or it will incorrectly collapse it if it is. More importantly, the unconditional unsqueeze(dim=1) at line 214 will result in a 4D output if the input was 3D with L > 1, which will likely cause shape mismatch errors downstream. It is safer to check the shape explicitly.

Suggested change
if pertoken_scale.dim() == 2:
need_unsqz = True
quantized_x = quantized_x.squeeze(dim=1)
pertoken_scale = pertoken_scale.squeeze(dim=1)
if x.dim() == 3 and x.shape[1] == 1:
need_unsqz = True
quantized_x = quantized_x.squeeze(dim=1)
pertoken_scale = pertoken_scale.squeeze(dim=1)


def process_weights_after_loading(self, layer: torch.nn.Module) -> None:
# cast quantized weight tensors in NZ format for higher inference speed
layer.weight.data = maybe_trans_nz(layer.weight.data).transpose(0, 1)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Transposing a tensor after it has been converted to the "NZ" (Fractal) format using maybe_trans_nz is incorrect. The NZ layout is memory-sensitive, and a standard transpose operation will either trigger an expensive layout conversion back to ND or result in an invalid memory layout for the npu_quant_matmul kernel. The transpose should be performed before the NZ conversion to ensure the kernel receives the expected [K, N] layout in NZ format.

Suggested change
layer.weight.data = maybe_trans_nz(layer.weight.data).transpose(0, 1)
layer.weight.data = maybe_trans_nz(layer.weight.data.transpose(0, 1).contiguous())

Comment on lines +155 to +156
layer.w13_weight.data = maybe_trans_nz(layer.w13_weight.data)
layer.w2_weight.data = maybe_trans_nz(layer.w2_weight.data)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The MoE weights w13_weight and w2_weight are typically stored in [E, 2*I, H] and [E, H, I] layouts. However, the Ascend MoE kernels usually expect the expert weights to be transposed in the hidden dimension (e.g., [E, H, 2*I]) and converted to NZ format. The current implementation is missing the transpose before maybe_trans_nz, which is inconsistent with the unquantized implementation in vllm_ascend/_310p/fused_moe/fused_moe.py.

Suggested change
layer.w13_weight.data = maybe_trans_nz(layer.w13_weight.data)
layer.w2_weight.data = maybe_trans_nz(layer.w2_weight.data)
layer.w13_weight.data = maybe_trans_nz(layer.w13_weight.data.transpose(1, 2).contiguous())
layer.w2_weight.data = maybe_trans_nz(layer.w2_weight.data.transpose(1, 2).contiguous())


expert_tokens = expert_tokens.to(torch.int64)
group_list_type = 1 # `count` mode
group_list_type = 0 # `cumsum` mode

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

npu_moe_init_routing_v2 returns expert_tokens as a tensor of token counts per expert. Setting group_list_type = 0 (which indicates cumsum mode) without actually performing a cumulative sum on expert_tokens will cause the downstream MoE kernels to misinterpret the token distribution, leading to incorrect results.

Suggested change
group_list_type = 0 # `cumsum` mode
group_list_type = 1 # `count` mode

@@ -1,12 +1,13 @@
import gc

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The Pull Request title and summary do not follow the repository style guide. Please update them to adhere to the required format.

Suggested PR Title:

[Ops][Feature] Add 310P optimizations and integration tests for GDN and W8A8

Suggested PR Summary:

### What this PR does / why we need it?
This PR introduces several optimizations and fixes for the 310P Ascend platform, including:
- Optimized expert selection using `npu_moe_gating_top_k_softmax`.
- Improved token dispatching using `npu_moe_init_routing_v2`.
- Support for W8A8 dynamic quantization for linear layers on 310P.
- Updated GDN (Gated Delta Net) operators and state layout for better performance and compatibility.
- Added integration tests for chunked and recurrent gated delta rules.

### Does this PR introduce _any_ user-facing change?
No.

### How was this patch tested?
Added new integration tests in `tests/e2e/nightly/single_node/ops/singlecard_ops/triton/` and unit tests in `tests/ut/_310p/`.
References
  1. The PR title and summary must follow the specific format defined in the repository style guide. (link)

@github-actions

github-actions Bot commented Apr 3, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@github-actions

github-actions Bot commented Apr 7, 2026

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@Tflowers-0129

Copy link
Copy Markdown
Collaborator Author

ONLY FOR INTERGRATION TEST

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants