Intergration test - #7972
Intergration test#7972Tflowers-0129 wants to merge 32 commits into
Conversation
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: pu-zhe <zpuaa@outlook.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: YangShuai52 <yangshuai153@huawei.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces significant performance optimizations for the 310P backend by integrating native NPU kernels for core operations like gated delta rules and causal convolutions. It also standardizes state tensor layouts and improves quantization efficiency through NZ format casting, ensuring better alignment with NPU hardware capabilities. These changes collectively aim to reduce latency and improve throughput for supported models. Highlights
🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request introduces 310P platform optimizations, including NPU-specific fused kernels for MoE gating and token dispatching, W8A8 dynamic quantization for linear layers, and updated state layouts for Gated Delta Rule operators. Feedback focuses on correcting the order of transpose operations for NZ format conversion, fixing potential shape mismatches in quantization logic, and ensuring the correct token count mode in the dispatcher. Additionally, the PR title and summary must be updated to comply with the repository's style guide.
| if pertoken_scale.dim() == 2: | ||
| need_unsqz = True | ||
| quantized_x = quantized_x.squeeze(dim=1) | ||
| pertoken_scale = pertoken_scale.squeeze(dim=1) |
There was a problem hiding this comment.
The logic for squeezing dimensions assumes that if pertoken_scale is 2D, the sequence dimension (dim 1) is of size 1. If the input x has a sequence length greater than 1 (e.g., during prefill), squeeze(dim=1) will not remove the dimension if its size is not 1, or it will incorrectly collapse it if it is. More importantly, the unconditional unsqueeze(dim=1) at line 214 will result in a 4D output if the input was 3D with L > 1, which will likely cause shape mismatch errors downstream. It is safer to check the shape explicitly.
| if pertoken_scale.dim() == 2: | |
| need_unsqz = True | |
| quantized_x = quantized_x.squeeze(dim=1) | |
| pertoken_scale = pertoken_scale.squeeze(dim=1) | |
| if x.dim() == 3 and x.shape[1] == 1: | |
| need_unsqz = True | |
| quantized_x = quantized_x.squeeze(dim=1) | |
| pertoken_scale = pertoken_scale.squeeze(dim=1) |
|
|
||
| def process_weights_after_loading(self, layer: torch.nn.Module) -> None: | ||
| # cast quantized weight tensors in NZ format for higher inference speed | ||
| layer.weight.data = maybe_trans_nz(layer.weight.data).transpose(0, 1) |
There was a problem hiding this comment.
Transposing a tensor after it has been converted to the "NZ" (Fractal) format using maybe_trans_nz is incorrect. The NZ layout is memory-sensitive, and a standard transpose operation will either trigger an expensive layout conversion back to ND or result in an invalid memory layout for the npu_quant_matmul kernel. The transpose should be performed before the NZ conversion to ensure the kernel receives the expected [K, N] layout in NZ format.
| layer.weight.data = maybe_trans_nz(layer.weight.data).transpose(0, 1) | |
| layer.weight.data = maybe_trans_nz(layer.weight.data.transpose(0, 1).contiguous()) |
| layer.w13_weight.data = maybe_trans_nz(layer.w13_weight.data) | ||
| layer.w2_weight.data = maybe_trans_nz(layer.w2_weight.data) |
There was a problem hiding this comment.
The MoE weights w13_weight and w2_weight are typically stored in [E, 2*I, H] and [E, H, I] layouts. However, the Ascend MoE kernels usually expect the expert weights to be transposed in the hidden dimension (e.g., [E, H, 2*I]) and converted to NZ format. The current implementation is missing the transpose before maybe_trans_nz, which is inconsistent with the unquantized implementation in vllm_ascend/_310p/fused_moe/fused_moe.py.
| layer.w13_weight.data = maybe_trans_nz(layer.w13_weight.data) | |
| layer.w2_weight.data = maybe_trans_nz(layer.w2_weight.data) | |
| layer.w13_weight.data = maybe_trans_nz(layer.w13_weight.data.transpose(1, 2).contiguous()) | |
| layer.w2_weight.data = maybe_trans_nz(layer.w2_weight.data.transpose(1, 2).contiguous()) |
|
|
||
| expert_tokens = expert_tokens.to(torch.int64) | ||
| group_list_type = 1 # `count` mode | ||
| group_list_type = 0 # `cumsum` mode |
There was a problem hiding this comment.
npu_moe_init_routing_v2 returns expert_tokens as a tensor of token counts per expert. Setting group_list_type = 0 (which indicates cumsum mode) without actually performing a cumulative sum on expert_tokens will cause the downstream MoE kernels to misinterpret the token distribution, leading to incorrect results.
| group_list_type = 0 # `cumsum` mode | |
| group_list_type = 1 # `count` mode |
| @@ -1,12 +1,13 @@ | |||
| import gc | |||
There was a problem hiding this comment.
The Pull Request title and summary do not follow the repository style guide. Please update them to adhere to the required format.
Suggested PR Title:
[Ops][Feature] Add 310P optimizations and integration tests for GDN and W8A8Suggested PR Summary:
### What this PR does / why we need it?
This PR introduces several optimizations and fixes for the 310P Ascend platform, including:
- Optimized expert selection using `npu_moe_gating_top_k_softmax`.
- Improved token dispatching using `npu_moe_init_routing_v2`.
- Support for W8A8 dynamic quantization for linear layers on 310P.
- Updated GDN (Gated Delta Net) operators and state layout for better performance and compatibility.
- Added integration tests for chunked and recurrent gated delta rules.
### Does this PR introduce _any_ user-facing change?
No.
### How was this patch tested?
Added new integration tests in `tests/e2e/nightly/single_node/ops/singlecard_ops/triton/` and unit tests in `tests/ut/_310p/`.References
- The PR title and summary must follow the specific format defined in the repository style guide. (link)
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
|
This pull request has conflicts, please resolve those before we can evaluate the pull request. |
|
ONLY FOR INTERGRATION TEST |
What this PR does / why we need it?
Does this PR introduce any user-facing change?
How was this patch tested?