Skip to content

[Ops][Feature] Support Ascend 950 and upgrade batch invariant ops to 2.0.0 - #15956

Merged
weijinqian0 merged 7 commits into
vllm-project:mainfrom
wangx700:feat/batch-invariance-a5-2.0.0
Sep 10, 2026
Merged

weijinqian0 merged 7 commits into
vllm-project:mainfrom
wangx700:feat/batch-invariance-a5-2.0.0

Conversation

@wangx700

@wangx700 wangx700 commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

What this PR does / why we need it?

ref to #10072

Upgrade the batch invariance operator run package and PyTorch extension archive to 2.0.0. Map ascend950 variants to the 950 package and enable operator installation during both Ubuntu and openEuler A5 image builds.

Document Ascend 950 installation and graph support. Remove the forced PIECEWISE configuration from the online/offline examples and the TP4 batch invariance test, preserving its capture sizes and updating the graph coverage annotation.

Restrict the custom reduce-sum operator to FP16, FP32, and BF16. Other dtypes fall back to the saved native torch.sum implementation, avoiding recursive calls through the patched Tensor.sum. Add parameterized regression tests for supported and unsupported NPU dtypes.

Does this PR introduce any user-facing change?

A5 source-built images install the batch invariance operators. Users still set VLLM_BATCH_INVARIANT=1 at inference time. Examples and the TP4 consistency test use the default graph configuration.

How was this patch tested?

  • pytest -sv tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py
======================================================================= warnings summary ========================================================================
../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings
  /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111
  /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.model(

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
=========================================================== 1 passed, 15 warnings in 94.14s (0:01:34) ===========================================================
[root@C02A09-OS1 vllm-ascend]# /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
  • VLLM_BATCH_INVARIANT=1 pip install -v -e . --no-build-isolation --no-deps --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple --trusted-host triton-ascend.osinfra.cn
    Enable VLLM_BATCH_INVARIANT=1 on A5, install vllm-ascend, and after installation run test_batch_invariant_tp4.py to verify success.

  • Build the image with Dockerfile, and verification is OK.

  • vLLM main: vllm-project/vllm@b2f6858

Signed-off-by: wangx700 <wangxin700@huawei.com>
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces support for the Ascend 950 (Atlas A5) hardware within the batch invariance feature set. It upgrades the underlying operator packages to version 2.0.0, which unlocks advanced graph execution modes. Additionally, the changes simplify the configuration requirements for users and tests by removing mandatory 'PIECEWISE' constraints, reflecting improved capabilities in the updated operator stack.

Highlights

  • Ascend A5 Support: Added support for Ascend 950 (Atlas A5) hardware, including updated build configurations and operator installation for both Ubuntu and openEuler images.
  • Operator Upgrade: Upgraded the batch invariance operator run package and PyTorch extension archive to version 2.0.0, enabling support for 'FULL' and 'FULL_DECODE_ONLY' cudagraph modes.
  • Configuration Cleanup: Removed forced 'PIECEWISE' cudagraph mode configurations from examples and tests, allowing for default graph configuration usage.
New Features

🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@github-actions github-actions Bot added documentation Improvements or additions to documentation module:tests labels Sep 7, 2026
@github-actions

github-actions Bot commented Sep 7, 2026

Copy link
Copy Markdown
Contributor

👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:‌‌

  • A PR should do only one thing, smaller PRs enable faster reviews.
  • Every PR should include unit tests and end-to-end tests ‌to ensure it works and is not broken by other future PRs.
  • Write the commit message by fulfilling the PR description to help reviewer and future developers understand.

If CI fails, you can run linting and testing checks locally according Contributing and Testing.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

Suggested PR Title:

[Ops][Feature] Support Ascend 950 and upgrade batch invariant ops to 2.0.0

Suggested PR Summary:

### What this PR does / why we need it?
This PR adds support for Ascend 950 (Atlas A5) NPUs and upgrades the batch invariant operators to version 2.0.0. This upgrade enables graph execution, including `FULL` and `FULL_DECODE_ONLY` cudagraph modes, which allows removing the piecewise-only restriction in documentation and tests. Additionally, the Dockerfiles for A5 are updated to build with `VLLM_BATCH_INVARIANT=1`.

Feedback: In both `Dockerfile.a5` and `Dockerfile.a5.openEuler`, the `COMPILE_CUSTOM_KERNELS` build argument is not exported during the `RUN` step, which might cause custom kernel compilation to be skipped. It should be explicitly exported.

### Does this PR introduce _any_ user-facing change?
Yes, users can now use batch invariance on Ascend 950 (Atlas A5) NPUs and utilize `FULL` and `FULL_DECODE_ONLY` cudagraph modes with batch invariance enabled.

### How was this patch tested?
The patch updates the end-to-end consistency test `test_batch_invariant_tp4.py` to run with `graph_mode="full_and_piecewise"` and removes the explicit `PIECEWISE` compilation configuration override.

Comment thread Dockerfile.a5
COPY . /vllm-workspace/vllm-ascend/

RUN export PIP_EXTRA_INDEX_URL="${ASCEND_INDEX_URL}" && \
export VLLM_BATCH_INVARIANT=1 && \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The COMPILE_CUSTOM_KERNELS variable is defined as a build argument (ARG COMPILE_CUSTOM_KERNELS=1), but it is not exported as an environment variable during the RUN step. Consequently, the custom kernel compilation might be skipped during pip install. Exporting COMPILE_CUSTOM_KERNELS ensures that the custom kernels (including the batch invariant operators) are properly compiled during the Docker image build.

    export VLLM_BATCH_INVARIANT=1 && \
    export COMPILE_CUSTOM_KERNELS=${COMPILE_CUSTOM_KERNELS} && \

@wangx700 wangx700 Sep 9, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ARG COMPILE_CUSTOM_KERNELS=1 is declared within the build stage, so Docker passes it to subsequent RUN instructions as a build-time environment variable. The pip install subprocess can read it through os.getenv, and an additional export is unnecessary. The project also defaults this setting to 1 when unset.

Comment thread Dockerfile.a5.openEuler
COPY . /vllm-workspace/vllm-ascend/

RUN export PIP_EXTRA_INDEX_URL="${ASCEND_INDEX_URL}" && \
export VLLM_BATCH_INVARIANT=1 && \

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The COMPILE_CUSTOM_KERNELS variable is defined as a build argument (ARG COMPILE_CUSTOM_KERNELS=1), but it is not exported as an environment variable during the RUN step. Consequently, the custom kernel compilation might be skipped during pip install. Exporting COMPILE_CUSTOM_KERNELS ensures that the custom kernels (including the batch invariant operators) are properly compiled during the Docker image build.

    export VLLM_BATCH_INVARIANT=1 && \
    export COMPILE_CUSTOM_KERNELS=${COMPILE_CUSTOM_KERNELS} && \

@wangx700 wangx700 Sep 9, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

ARG COMPILE_CUSTOM_KERNELS=1 is declared within the build stage, so Docker passes it to subsequent RUN instructions as a build-time environment variable. The pip install subprocess can read it through os.getenv, and an additional export is unnecessary. The project also defaults this setting to 1 when unset.

Signed-off-by: wangx700 <wangxin700@huawei.com>
@wangx700 wangx700 changed the title [Feat][Batch Invariance] Support Ascend A5 and upgrade operators to 2.0.0 [Feat][Batch Invariance] Support Ascend A5 and upgrade operators version to 2.0.0 Sep 7, 2026
@wangx700 wangx700 changed the title [Feat][Batch Invariance] Support Ascend A5 and upgrade operators version to 2.0.0 [Ops][Batch Invariance] Support Ascend A5 and upgrade operators version to 2.0.0 Sep 7, 2026
Signed-off-by: wangx700 <wangxin700@huawei.com>
@wangx700 wangx700 changed the title [Ops][Batch Invariance] Support Ascend A5 and upgrade operators version to 2.0.0 [Ops][Feature] Support Ascend 950 and upgrade batch invariant ops to 2.0.0 Sep 8, 2026
@wangx700

wangx700 commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor Author

Signed-off-by: wangx700 <wangxin700@huawei.com>
@wangx700
wangx700 marked this pull request as ready for review September 9, 2026 06:17
Signed-off-by: wangx700 <wangxin700@huawei.com>
Signed-off-by: wangx700 <wangxin700@huawei.com>

Batch invariance currently requires Ascend Atlas A2 and A3 inference products NPUs.
We will support Ascend 950 Products and other NPUs in the future.
Batch invariance supports Ascend Atlas A2, A3, and Ascend 950 NPUs.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ascend 950 Products

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK, Ascend 950 NPUs was changed to Ascend 950 Products


Batch invariance currently requires Ascend Atlas A2 and A3 inference products NPUs.
We will support Ascend 950 Products and other NPUs in the future.
Batch invariance supports Ascend Atlas A2, A3, and Ascend 950 NPUs.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ascend Atlas A2,A3。去掉Ascend

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

OK, Ascend was removed.

Signed-off-by: wangx700 <wangxin700@huawei.com>

@Ronald1995 Ronald1995 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@Ronald1995 Ronald1995 added the ready-precise run selected e2e test for pr label Sep 9, 2026
@weijinqian0
weijinqian0 merged commit 984b1b6 into vllm-project:main Sep 10, 2026
44 of 45 checks passed
chen-commits pushed a commit to chen-commits/vllm-ascend that referenced this pull request Sep 10, 2026
…2.0.0 (vllm-project#15956)

### What this PR does / why we need it?

ref to vllm-project#10072

Upgrade the batch invariance operator run package and PyTorch extension
archive to 2.0.0. Map ascend950 variants to the 950 package and enable
operator installation during both Ubuntu and openEuler A5 image builds.

Document Ascend 950 installation and graph support. Remove the forced
PIECEWISE configuration from the online/offline examples and the TP4
batch invariance test, preserving its capture sizes and updating the
graph coverage annotation.

Restrict the custom reduce-sum operator to FP16, FP32, and BF16. Other
dtypes fall back to the saved native torch.sum implementation, avoiding
recursive calls through the patched Tensor.sum. Add parameterized
regression tests for supported and unsupported NPU dtypes.

### Does this PR introduce _any_ user-facing change?
A5 source-built images install the batch invariance operators. Users
still set VLLM_BATCH_INVARIANT=1 at inference time. Examples and the TP4
consistency test use the default graph configuration.

### How was this patch tested?
- pytest -sv
tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py
```
======================================================================= warnings summary ========================================================================
../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings
  /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111
  /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.model(

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
=========================================================== 1 passed, 15 warnings in 94.14s (0:01:34) ===========================================================
[root@C02A09-OS1 vllm-ascend]# /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
```
- VLLM_BATCH_INVARIANT=1 pip install -v -e . --no-build-isolation
--no-deps --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple
--trusted-host triton-ascend.osinfra.cn
Enable VLLM_BATCH_INVARIANT=1 on A5, install vllm-ascend, and after
installation run test_batch_invariant_tp4.py to verify success.
- Build the image with Dockerfile, and verification is OK.


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wangx700 <wangxin700@huawei.com>
sunny-rain-63 pushed a commit to sunny-rain-63/vllm-ascend that referenced this pull request Sep 12, 2026
…2.0.0 (vllm-project#15956)

### What this PR does / why we need it?

ref to vllm-project#10072

Upgrade the batch invariance operator run package and PyTorch extension
archive to 2.0.0. Map ascend950 variants to the 950 package and enable
operator installation during both Ubuntu and openEuler A5 image builds.

Document Ascend 950 installation and graph support. Remove the forced
PIECEWISE configuration from the online/offline examples and the TP4
batch invariance test, preserving its capture sizes and updating the
graph coverage annotation.

Restrict the custom reduce-sum operator to FP16, FP32, and BF16. Other
dtypes fall back to the saved native torch.sum implementation, avoiding
recursive calls through the patched Tensor.sum. Add parameterized
regression tests for supported and unsupported NPU dtypes.

### Does this PR introduce _any_ user-facing change?
A5 source-built images install the batch invariance operators. Users
still set VLLM_BATCH_INVARIANT=1 at inference time. Examples and the TP4
consistency test use the default graph configuration.

### How was this patch tested?
- pytest -sv
tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py
```
======================================================================= warnings summary ========================================================================
../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings
  /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111
  /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.model(

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
=========================================================== 1 passed, 15 warnings in 94.14s (0:01:34) ===========================================================
[root@C02A09-OS1 vllm-ascend]# /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
```
- VLLM_BATCH_INVARIANT=1 pip install -v -e . --no-build-isolation
--no-deps --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple
--trusted-host triton-ascend.osinfra.cn
Enable VLLM_BATCH_INVARIANT=1 on A5, install vllm-ascend, and after
installation run test_batch_invariant_tp4.py to verify success.
- Build the image with Dockerfile, and verification is OK.


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wangx700 <wangxin700@huawei.com>
johnnysluckydays pushed a commit to johnnysluckydays/vllm-ascend that referenced this pull request Sep 14, 2026
…2.0.0 (vllm-project#15956)

### What this PR does / why we need it?

ref to vllm-project#10072

Upgrade the batch invariance operator run package and PyTorch extension
archive to 2.0.0. Map ascend950 variants to the 950 package and enable
operator installation during both Ubuntu and openEuler A5 image builds.

Document Ascend 950 installation and graph support. Remove the forced
PIECEWISE configuration from the online/offline examples and the TP4
batch invariance test, preserving its capture sizes and updating the
graph coverage annotation.

Restrict the custom reduce-sum operator to FP16, FP32, and BF16. Other
dtypes fall back to the saved native torch.sum implementation, avoiding
recursive calls through the patched Tensor.sum. Add parameterized
regression tests for supported and unsupported NPU dtypes.

### Does this PR introduce _any_ user-facing change?
A5 source-built images install the batch invariance operators. Users
still set VLLM_BATCH_INVARIANT=1 at inference time. Examples and the TP4
consistency test use the default graph configuration.

### How was this patch tested?
- pytest -sv
tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py
```
======================================================================= warnings summary ========================================================================
../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings
  /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111
  /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.model(

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
=========================================================== 1 passed, 15 warnings in 94.14s (0:01:34) ===========================================================
[root@C02A09-OS1 vllm-ascend]# /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
```
- VLLM_BATCH_INVARIANT=1 pip install -v -e . --no-build-isolation
--no-deps --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple
--trusted-host triton-ascend.osinfra.cn
Enable VLLM_BATCH_INVARIANT=1 on A5, install vllm-ascend, and after
installation run test_batch_invariant_tp4.py to verify success.
- Build the image with Dockerfile, and verification is OK.


- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wangx700 <wangxin700@huawei.com>
Signed-off-by: tianming2009 <13246728590@163.com>
like-0517 pushed a commit to like-0517/vllm-ascend that referenced this pull request Sep 15, 2026
…2.0.0 (vllm-project#15956)

### What this PR does / why we need it?

ref to vllm-project#10072

Upgrade the batch invariance operator run package and PyTorch extension
archive to 2.0.0. Map ascend950 variants to the 950 package and enable
operator installation during both Ubuntu and openEuler A5 image builds.

Document Ascend 950 installation and graph support. Remove the forced
PIECEWISE configuration from the online/offline examples and the TP4
batch invariance test, preserving its capture sizes and updating the
graph coverage annotation.

Restrict the custom reduce-sum operator to FP16, FP32, and BF16. Other
dtypes fall back to the saved native torch.sum implementation, avoiding
recursive calls through the patched Tensor.sum. Add parameterized
regression tests for supported and unsupported NPU dtypes.

### Does this PR introduce _any_ user-facing change?
A5 source-built images install the batch invariance operators. Users
still set VLLM_BATCH_INVARIANT=1 at inference time. Examples and the TP4
consistency test use the default graph configuration.

### How was this patch tested?
- pytest -sv
tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py
```
======================================================================= warnings summary ========================================================================
../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings
  /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`.
    warnings.warn(

tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111
  /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo?  You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html
    @pytest.mark.model(

-- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html
=========================================================== 1 passed, 15 warnings in 94.14s (0:01:34) ===========================================================
[root@C02A09-OS1 vllm-ascend]# /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown
  warnings.warn('resource_tracker: There appear to be %d '
```
- VLLM_BATCH_INVARIANT=1 pip install -v -e . --no-build-isolation
--no-deps --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple
--trusted-host triton-ascend.osinfra.cn
Enable VLLM_BATCH_INVARIANT=1 on A5, install vllm-ascend, and after
installation run test_batch_invariant_tp4.py to verify success.
- Build the image with Dockerfile, and verification is OK.

- vLLM main:
vllm-project/vllm@b2f6858

---------

Signed-off-by: wangx700 <wangxin700@huawei.com>
Signed-off-by: like-0517 <ithwlike@126.com>
linfeng-yuan pushed a commit that referenced this pull request Sep 17, 2026
…ons (#16413)

### What this PR does / why we need it?

`vllm_ascend/batch_invariant.py:reduce_sum` unconditionally forwards
every NPU
`torch.sum(x, dim, keepdim)` (for supported dtypes) to
`torch.ops.batch_invariant_ops.npu_reduce_sum_batch_invariant`. But the
`aclnnReduceSumBatchInvariant` kernel **only supports reducing the
last dimension**; for any other `dim` it raises:

```
AclNN_Parameter_Error(EZ1001): Provided dim only support last dim
```

Because `enable_batch_invariant_mode()` monkey-patches `torch.sum` /
`torch.Tensor.sum` globally, **any** non-last-dim `.sum()` on NPU kills
all TP
workers.

To fix it, this PR only takes the batch-invariant path for a **genuine
last-dim reduction**:

1. The last dim can only be spelled as `-1` or `x.dim() - 1`, so the
guard is
   simply `dim == -1 or dim == x.dim() - 1`, and the caller's `dim` is
   forwarded to the kernel unchanged. No dimension arithmetic is needed,
and equality comparison is safe for every value of `dim`: a tuple, for
example, simply fails both comparisons and falls back instead of
raising.
   The kernel is therefore used only when all of the following hold: NPU
tensor, a genuine last-dim reduction, and dtype in `{fp16, fp32, bf16}`.
Forwarding `dim` unchanged means every previously-working last-dim call
keeps
the exact behavior it had before this PR, and the existing unit test
stays
   valid.
2. Everything else - non-last-dim, tuple `dim`, full reduction (`dim is
None`
   on >1-D), CPU tensors, unsupported dtypes - falls back to the **saved
native `torch_sum` reference** captured at module import, so the
fallback
   cannot recurse into the patched entry points.

Every input that now falls back previously crashed, so the change
strictly
widens the set of inputs that run; there is no working behavior being
changed.

Why this does not affect batch invariance?

1. **The invariance-critical reductions are all last-dim, and they are
unchanged.** The reductions whose results must be batch-invariant
(softmax
   denominators, RMSNorm variance - reductions over the hidden/vocab
   dimension) are last-dim reductions on `[num_tokens, hidden]` tensors.
Native kernels pick their reduction strategy from the total tensor size,
so the per-row accumulation order can differ between BS=1 and BS=N -
that
   is exactly what `npu_reduce_sum_batch_invariant` prevents with a
   row-count-independent accumulation order. These calls still take the
   batch-invariant kernel; their behavior is byte-for-byte identical to
   before.

2. **The newly-fallback paths have no batch-shape-dependent input.** The
   non-last-dim sums that now fall back (e.g. `embeds.sum(dim=0)` in
multimodal position-embedding preprocessing) are per-request
computations:
the input tensor is built from one request's own patches, so its shape
and
values are identical in a BS=1 run and in a BS=N run. Native `torch.sum`
is deterministic for a given tensor (same shape ? same kernel and tiling
?
   same accumulation order), so identical input yields bitwise-identical
   output across batch sizes. Invariance is preserved.

3. **No regression is possible.** Every path that now falls back
previously
raised EZ1001 (non-last-dim int dims) or a schema type error (tuple
dims).
   There was no working behavior to preserve.

4. **It follows the established fallback design.** #15956 introduced the
same
pattern for the dtype axis: unsupported dtypes "fall back to the saved
   native torch.sum implementation, avoiding recursive calls through the
   patched Tensor.sum". This PR extends that pattern from dtype to dim.

### Does this PR introduce _any_ user-facing change?

No

### How was this patch tested?

- 10-path NPU unit test on Ascend910 (dim=0 fallback, dim=2 / dim=-1
equivalence
  vs native, keepdim, 1-D `.sum()`, tuple dim, all-dims reduction, fp16,
  functional `torch.sum`, fp32) - all pass.
- End-to-end batch-invariance logprobs test (BS=1 vs BS=N bitwise) runs
through
  `profile_run` and generation with no EZ1001.


vllm-project/vllm@b2f6858

---------

Signed-off-by: EliasKaslan <3248569436@qq.com>

- vLLM main:
vllm-project/vllm@84030bb

Signed-off-by: EliasKaslan <3248569436@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation module:core module:tests ready-precise run selected e2e test for pr

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants