Skip to content

[310P]: add torch chunk gated delta rule and 910b parity ut - #7594

Merged
wangxiyuan merged 5 commits into
vllm-project:mainfrom
Tflowers-0129:feat/310p-chunk-delta-transformers-ref
Mar 25, 2026
Merged

wangxiyuan merged 5 commits into
vllm-project:mainfrom
Tflowers-0129:feat/310p-chunk-delta-transformers-ref

Conversation

@Tflowers-0129

@Tflowers-0129 Tflowers-0129 commented Mar 24, 2026 •

Copy link
Copy Markdown
Collaborator

What this PR does / why we need it?

RFC #7394
Add a PyTorch implementation of the chunk gated delta rule on 310P.

Does this PR introduce any user-facing change?

NO

How was this patch tested?

UT

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces a PyTorch-based implementation of the chunked gated delta rule, specifically designed for the Ascend 310P environment. This provides a robust fallback mechanism for this critical operation within the vllm-ascend project, ensuring functional correctness across different input configurations through newly added unit tests.

Highlights

  • New Feature Implementation: Implemented a PyTorch-based chunk_gated_delta_rule function, intended as a fallback or reference for the Ascend 310P platform.
  • Module Integration: Integrated the new chunk_gated_delta_rule_pytorch function into the vllm_ascend._310p.ops.fla module.
  • Comprehensive Testing: Added extensive unit tests for the chunk_gated_delta_rule_pytorch function, covering output shape and data type, variable-length sequence handling, and TND (Total Number of Tokens, Number of Heads, Dimension) input format.

🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a PyTorch fallback implementation for chunk_gated_delta_rule, intended for the Ascend 310P platform. The implementation supports various input formats and is accompanied by unit tests. My review has identified one high-severity issue regarding the use of DeprecationWarning for an unsupported parameter, which should be changed to an error to prevent silent failures.

As per the repository's style guide, here are suggestions for the pull request title and description:

Suggested PR Title:

[Feat/310p][Ops][Feature] Add PyTorch fallback for chunk_gated_delta_rule

Suggested PR Summary:

### What this PR does / why we need it?
This pull request introduces a PyTorch-based fallback implementation for the `chunk_gated_delta_rule` operation, specifically for the Ascend 310P platform. This is necessary to provide a functional equivalent for environments where a specialized hardware kernel is not available, ensuring broader compatibility and enabling testing/development on different platforms.

The implementation is aligned with the `torch_chunk_gated_delta_rule` from the `transformers` library (for Qwen3-Next models) and supports both standard (BTHD) and variable-length (TND with `cu_seqlens`) inputs.

### Does this PR introduce _any_ user-facing change?
Yes, this PR adds a new operator `chunk_gated_delta_rule_pytorch` under `vllm_ascend._310p.ops.fla`. This is an internal-facing change for developers working on model support, but does not affect end-users of the vLLM API directly.

### How was this patch tested?
The patch was tested by adding a new unit test file (`tests/ut/_310p/ops/test_chunk_gated_delta_rule_310.py`). The tests cover:
- Correctness of output shapes and dtypes.
- The variable-length input path using `cu_seqlens`.
- Equivalence between TND (token-major) and BTHD (batch-major) input formats for variable-length sequences.
All tests pass, ensuring the implementation is correct for the tested scenarios.

Comment thread vllm_ascend/_310p/ops/fla/chunk_gated_delta_rule.py
@github-actions

Copy link
Copy Markdown
Contributor

This pull request has conflicts, please resolve those before we can evaluate the pull request.

@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@Tflowers-0129 Tflowers-0129 changed the title Feat/310p chunk delta transformers ref [310p] chunk delta ruler Mar 24, 2026
- add _310p chunk_gated_delta_rule_pytorch with vLLM-compatible interface
- align internal chunk recurrence math with transformers torch reference flow
- add CPU UT for shape/dtype and varlen interface path

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: Tflowers-0129 <2906339855@qq.com>
- normalize 3D TND inputs to 4D internal layout when cu_seqlens is provided
- keep existing BTHD interface unchanged
- return TND output when TND input is used
- add UT to verify TND and BTHD varlen parity

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
@Tflowers-0129
Tflowers-0129 force-pushed the feat/310p-chunk-delta-transformers-ref branch from f44c5e6 to a892e12 Compare March 25, 2026 01:09
@Tflowers-0129 Tflowers-0129 changed the title [310p] chunk delta ruler [310P]: add torch chunk gated delta rule and 910b parity ut Mar 25, 2026
@wangxiyuan
wangxiyuan merged commit e0e585a into vllm-project:main Mar 25, 2026
35 checks passed
lihaokun-2026 pushed a commit to lihaokun-2026/vllm-ascend that referenced this pull request Mar 29, 2026
…ject#7594)

### What this PR does / why we need it?
RFC vllm-project#7394
Add a PyTorch implementation of the  chunk gated delta rule on 310P.
### Does this PR introduce _any_ user-facing change?
NO
### How was this patch tested?
UT

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
HF-001 pushed a commit to HF-001/vllm-ascend that referenced this pull request Mar 31, 2026
…ject#7594)

### What this PR does / why we need it?
RFC vllm-project#7394
Add a PyTorch implementation of the  chunk gated delta rule on 310P.
### Does this PR introduce _any_ user-facing change?
NO
### How was this patch tested?
UT

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: 01267596 <xiongkai123@cmbchina.com>
chenchuw886 pushed a commit to chenchuw886/vllm-ascend that referenced this pull request Apr 1, 2026
…ject#7594)

### What this PR does / why we need it?
RFC vllm-project#7394
Add a PyTorch implementation of the  chunk gated delta rule on 310P.
### Does this PR introduce _any_ user-facing change?
NO
### How was this patch tested?
UT

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
zouyida2052 pushed a commit to zouyida2052/vllm-ascend that referenced this pull request Apr 28, 2026
…ject#7594)

### What this PR does / why we need it?
RFC vllm-project#7394
Add a PyTorch implementation of the  chunk gated delta rule on 310P.
### Does this PR introduce _any_ user-facing change?
NO
### How was this patch tested?
UT

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: zouyida2052 <zouyida2002@gmail.com>
yangzhe-2026 pushed a commit to yangzhe-2026/vllm-ascend that referenced this pull request May 6, 2026
…ject#7594)

### What this PR does / why we need it?
RFC vllm-project#7394
Add a PyTorch implementation of the  chunk gated delta rule on 310P.
### Does this PR introduce _any_ user-facing change?
NO
### How was this patch tested?
UT

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
nanxingMy pushed a commit to nanxingMy/vllm-ascend that referenced this pull request May 15, 2026
…ject#7594)

### What this PR does / why we need it?
RFC vllm-project#7394
Add a PyTorch implementation of the  chunk gated delta rule on 310P.
### Does this PR introduce _any_ user-facing change?
NO
### How was this patch tested?
UT

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Signed-off-by: nanxing <1014662416@qq.com>
ader47 pushed a commit to ader47/vllm-ascend that referenced this pull request Jun 18, 2026
…ject#7594)

### What this PR does / why we need it?
RFC vllm-project#7394
Add a PyTorch implementation of the  chunk gated delta rule on 310P.
### Does this PR introduce _any_ user-facing change?
NO
### How was this patch tested?
UT

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
CXY-Katrina pushed a commit to CXY-Katrina/vllm-ascend that referenced this pull request Jun 27, 2026
…ject#7594)

### What this PR does / why we need it?
RFC vllm-project#7394
Add a PyTorch implementation of the  chunk gated delta rule on 310P.
### Does this PR introduce _any_ user-facing change?
NO
### How was this patch tested?
UT

---------

Signed-off-by: Tflowers-0129 <2906339855@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants