Repository navigation
[Ops][Feature] Support Ascend 950 and upgrade batch invariant ops to 2.0.0 - #15956
Conversation
Signed-off-by: wangx700 <wangxin700@huawei.com>
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces support for the Ascend 950 (Atlas A5) hardware within the batch invariance feature set. It upgrades the underlying operator packages to version 2.0.0, which unlocks advanced graph execution modes. Additionally, the changes simplify the configuration requirements for users and tests by removing mandatory 'PIECEWISE' constraints, reflecting improved capabilities in the updated operator stack. Highlights
New Features🧠 You can now enable Memory (public preview) to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console. Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize the Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counterproductive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
👋 Hi! Thank you for contributing to the vLLM Ascend project. The following points will speed up your PR merge:
If CI fails, you can run linting and testing checks locally according Contributing and Testing. |
There was a problem hiding this comment.
Code Review
Suggested PR Title:
[Ops][Feature] Support Ascend 950 and upgrade batch invariant ops to 2.0.0Suggested PR Summary:
### What this PR does / why we need it?
This PR adds support for Ascend 950 (Atlas A5) NPUs and upgrades the batch invariant operators to version 2.0.0. This upgrade enables graph execution, including `FULL` and `FULL_DECODE_ONLY` cudagraph modes, which allows removing the piecewise-only restriction in documentation and tests. Additionally, the Dockerfiles for A5 are updated to build with `VLLM_BATCH_INVARIANT=1`.
Feedback: In both `Dockerfile.a5` and `Dockerfile.a5.openEuler`, the `COMPILE_CUSTOM_KERNELS` build argument is not exported during the `RUN` step, which might cause custom kernel compilation to be skipped. It should be explicitly exported.
### Does this PR introduce _any_ user-facing change?
Yes, users can now use batch invariance on Ascend 950 (Atlas A5) NPUs and utilize `FULL` and `FULL_DECODE_ONLY` cudagraph modes with batch invariance enabled.
### How was this patch tested?
The patch updates the end-to-end consistency test `test_batch_invariant_tp4.py` to run with `graph_mode="full_and_piecewise"` and removes the explicit `PIECEWISE` compilation configuration override.| COPY . /vllm-workspace/vllm-ascend/ | ||
|
|
||
| RUN export PIP_EXTRA_INDEX_URL="${ASCEND_INDEX_URL}" && \ | ||
| export VLLM_BATCH_INVARIANT=1 && \ |
There was a problem hiding this comment.
The COMPILE_CUSTOM_KERNELS variable is defined as a build argument (ARG COMPILE_CUSTOM_KERNELS=1), but it is not exported as an environment variable during the RUN step. Consequently, the custom kernel compilation might be skipped during pip install. Exporting COMPILE_CUSTOM_KERNELS ensures that the custom kernels (including the batch invariant operators) are properly compiled during the Docker image build.
export VLLM_BATCH_INVARIANT=1 && \
export COMPILE_CUSTOM_KERNELS=${COMPILE_CUSTOM_KERNELS} && \
There was a problem hiding this comment.
ARG COMPILE_CUSTOM_KERNELS=1 is declared within the build stage, so Docker passes it to subsequent RUN instructions as a build-time environment variable. The pip install subprocess can read it through os.getenv, and an additional export is unnecessary. The project also defaults this setting to 1 when unset.
| COPY . /vllm-workspace/vllm-ascend/ | ||
|
|
||
| RUN export PIP_EXTRA_INDEX_URL="${ASCEND_INDEX_URL}" && \ | ||
| export VLLM_BATCH_INVARIANT=1 && \ |
There was a problem hiding this comment.
The COMPILE_CUSTOM_KERNELS variable is defined as a build argument (ARG COMPILE_CUSTOM_KERNELS=1), but it is not exported as an environment variable during the RUN step. Consequently, the custom kernel compilation might be skipped during pip install. Exporting COMPILE_CUSTOM_KERNELS ensures that the custom kernels (including the batch invariant operators) are properly compiled during the Docker image build.
export VLLM_BATCH_INVARIANT=1 && \
export COMPILE_CUSTOM_KERNELS=${COMPILE_CUSTOM_KERNELS} && \
There was a problem hiding this comment.
ARG COMPILE_CUSTOM_KERNELS=1 is declared within the build stage, so Docker passes it to subsequent RUN instructions as a build-time environment variable. The pip install subprocess can read it through os.getenv, and an additional export is unnecessary. The project also defaults this setting to 1 when unset.
Signed-off-by: wangx700 <wangxin700@huawei.com>
Signed-off-by: wangx700 <wangxin700@huawei.com>
Signed-off-by: wangx700 <wangxin700@huawei.com>
Signed-off-by: wangx700 <wangxin700@huawei.com>
Signed-off-by: wangx700 <wangxin700@huawei.com>
|
|
||
| Batch invariance currently requires Ascend Atlas A2 and A3 inference products NPUs. | ||
| We will support Ascend 950 Products and other NPUs in the future. | ||
| Batch invariance supports Ascend Atlas A2, A3, and Ascend 950 NPUs. |
There was a problem hiding this comment.
OK, Ascend 950 NPUs was changed to Ascend 950 Products
|
|
||
| Batch invariance currently requires Ascend Atlas A2 and A3 inference products NPUs. | ||
| We will support Ascend 950 Products and other NPUs in the future. | ||
| Batch invariance supports Ascend Atlas A2, A3, and Ascend 950 NPUs. |
There was a problem hiding this comment.
Ascend Atlas A2,A3。去掉Ascend
There was a problem hiding this comment.
OK, Ascend was removed.
Signed-off-by: wangx700 <wangxin700@huawei.com>
…2.0.0 (vllm-project#15956) ### What this PR does / why we need it? ref to vllm-project#10072 Upgrade the batch invariance operator run package and PyTorch extension archive to 2.0.0. Map ascend950 variants to the 950 package and enable operator installation during both Ubuntu and openEuler A5 image builds. Document Ascend 950 installation and graph support. Remove the forced PIECEWISE configuration from the online/offline examples and the TP4 batch invariance test, preserving its capture sizes and updating the graph coverage annotation. Restrict the custom reduce-sum operator to FP16, FP32, and BF16. Other dtypes fall back to the saved native torch.sum implementation, avoiding recursive calls through the patched Tensor.sum. Add parameterized regression tests for supported and unsupported NPU dtypes. ### Does this PR introduce _any_ user-facing change? A5 source-built images install the batch invariance operators. Users still set VLLM_BATCH_INVARIANT=1 at inference time. Examples and the TP4 consistency test use the default graph configuration. ### How was this patch tested? - pytest -sv tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py ``` ======================================================================= warnings summary ======================================================================== ../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`. warnings.warn( tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111 /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo? You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html @pytest.mark.model( -- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html =========================================================== 1 passed, 15 warnings in 94.14s (0:01:34) =========================================================== [root@C02A09-OS1 vllm-ascend]# /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' ``` - VLLM_BATCH_INVARIANT=1 pip install -v -e . --no-build-isolation --no-deps --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple --trusted-host triton-ascend.osinfra.cn Enable VLLM_BATCH_INVARIANT=1 on A5, install vllm-ascend, and after installation run test_batch_invariant_tp4.py to verify success. - Build the image with Dockerfile, and verification is OK. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: wangx700 <wangxin700@huawei.com>
…2.0.0 (vllm-project#15956) ### What this PR does / why we need it? ref to vllm-project#10072 Upgrade the batch invariance operator run package and PyTorch extension archive to 2.0.0. Map ascend950 variants to the 950 package and enable operator installation during both Ubuntu and openEuler A5 image builds. Document Ascend 950 installation and graph support. Remove the forced PIECEWISE configuration from the online/offline examples and the TP4 batch invariance test, preserving its capture sizes and updating the graph coverage annotation. Restrict the custom reduce-sum operator to FP16, FP32, and BF16. Other dtypes fall back to the saved native torch.sum implementation, avoiding recursive calls through the patched Tensor.sum. Add parameterized regression tests for supported and unsupported NPU dtypes. ### Does this PR introduce _any_ user-facing change? A5 source-built images install the batch invariance operators. Users still set VLLM_BATCH_INVARIANT=1 at inference time. Examples and the TP4 consistency test use the default graph configuration. ### How was this patch tested? - pytest -sv tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py ``` ======================================================================= warnings summary ======================================================================== ../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`. warnings.warn( tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111 /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo? You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html @pytest.mark.model( -- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html =========================================================== 1 passed, 15 warnings in 94.14s (0:01:34) =========================================================== [root@C02A09-OS1 vllm-ascend]# /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' ``` - VLLM_BATCH_INVARIANT=1 pip install -v -e . --no-build-isolation --no-deps --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple --trusted-host triton-ascend.osinfra.cn Enable VLLM_BATCH_INVARIANT=1 on A5, install vllm-ascend, and after installation run test_batch_invariant_tp4.py to verify success. - Build the image with Dockerfile, and verification is OK. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: wangx700 <wangxin700@huawei.com>
…2.0.0 (vllm-project#15956) ### What this PR does / why we need it? ref to vllm-project#10072 Upgrade the batch invariance operator run package and PyTorch extension archive to 2.0.0. Map ascend950 variants to the 950 package and enable operator installation during both Ubuntu and openEuler A5 image builds. Document Ascend 950 installation and graph support. Remove the forced PIECEWISE configuration from the online/offline examples and the TP4 batch invariance test, preserving its capture sizes and updating the graph coverage annotation. Restrict the custom reduce-sum operator to FP16, FP32, and BF16. Other dtypes fall back to the saved native torch.sum implementation, avoiding recursive calls through the patched Tensor.sum. Add parameterized regression tests for supported and unsupported NPU dtypes. ### Does this PR introduce _any_ user-facing change? A5 source-built images install the batch invariance operators. Users still set VLLM_BATCH_INVARIANT=1 at inference time. Examples and the TP4 consistency test use the default graph configuration. ### How was this patch tested? - pytest -sv tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py ``` ======================================================================= warnings summary ======================================================================== ../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`. warnings.warn( tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111 /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo? You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html @pytest.mark.model( -- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html =========================================================== 1 passed, 15 warnings in 94.14s (0:01:34) =========================================================== [root@C02A09-OS1 vllm-ascend]# /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' ``` - VLLM_BATCH_INVARIANT=1 pip install -v -e . --no-build-isolation --no-deps --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple --trusted-host triton-ascend.osinfra.cn Enable VLLM_BATCH_INVARIANT=1 on A5, install vllm-ascend, and after installation run test_batch_invariant_tp4.py to verify success. - Build the image with Dockerfile, and verification is OK. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: wangx700 <wangxin700@huawei.com> Signed-off-by: tianming2009 <13246728590@163.com>
…2.0.0 (vllm-project#15956) ### What this PR does / why we need it? ref to vllm-project#10072 Upgrade the batch invariance operator run package and PyTorch extension archive to 2.0.0. Map ascend950 variants to the 950 package and enable operator installation during both Ubuntu and openEuler A5 image builds. Document Ascend 950 installation and graph support. Remove the forced PIECEWISE configuration from the online/offline examples and the TP4 batch invariance test, preserving its capture sizes and updating the graph coverage annotation. Restrict the custom reduce-sum operator to FP16, FP32, and BF16. Other dtypes fall back to the saved native torch.sum implementation, avoiding recursive calls through the patched Tensor.sum. Add parameterized regression tests for supported and unsupported NPU dtypes. ### Does this PR introduce _any_ user-facing change? A5 source-built images install the batch invariance operators. Users still set VLLM_BATCH_INVARIANT=1 at inference time. Examples and the TP4 consistency test use the default graph configuration. ### How was this patch tested? - pytest -sv tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py ``` ======================================================================= warnings summary ======================================================================== ../../../../../usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: 14 warnings /usr/local/python3.11.10/lib/python3.11/site-packages/torch/jit/_script.py:362: DeprecationWarning: `torch.jit.script_method` is deprecated. Please switch to `torch.compile` or `torch.export`. warnings.warn( tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111 /mnt/share/w00899129/vllm_code_0908/vllm-ascend/tests/e2e/pull_request/four_card/rlhf/consistency/test_batch_invariant_tp4.py:111: PytestUnknownMarkWarning: Unknown pytest.mark.model - is this a typo? You can register custom marks to avoid this warning - for details, see https://docs.pytest.org/en/stable/how-to/mark.html @pytest.mark.model( -- Docs: https://docs.pytest.org/en/stable/how-to/capture-warnings.html =========================================================== 1 passed, 15 warnings in 94.14s (0:01:34) =========================================================== [root@C02A09-OS1 vllm-ascend]# /usr/local/python3.11.10/lib/python3.11/multiprocessing/resource_tracker.py:254: UserWarning: resource_tracker: There appear to be 1 leaked shared_memory objects to clean up at shutdown warnings.warn('resource_tracker: There appear to be %d ' ``` - VLLM_BATCH_INVARIANT=1 pip install -v -e . --no-build-isolation --no-deps --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple --trusted-host triton-ascend.osinfra.cn Enable VLLM_BATCH_INVARIANT=1 on A5, install vllm-ascend, and after installation run test_batch_invariant_tp4.py to verify success. - Build the image with Dockerfile, and verification is OK. - vLLM main: vllm-project/vllm@b2f6858 --------- Signed-off-by: wangx700 <wangxin700@huawei.com> Signed-off-by: like-0517 <ithwlike@126.com>
…ons (#16413) ### What this PR does / why we need it? `vllm_ascend/batch_invariant.py:reduce_sum` unconditionally forwards every NPU `torch.sum(x, dim, keepdim)` (for supported dtypes) to `torch.ops.batch_invariant_ops.npu_reduce_sum_batch_invariant`. But the `aclnnReduceSumBatchInvariant` kernel **only supports reducing the last dimension**; for any other `dim` it raises: ``` AclNN_Parameter_Error(EZ1001): Provided dim only support last dim ``` Because `enable_batch_invariant_mode()` monkey-patches `torch.sum` / `torch.Tensor.sum` globally, **any** non-last-dim `.sum()` on NPU kills all TP workers. To fix it, this PR only takes the batch-invariant path for a **genuine last-dim reduction**: 1. The last dim can only be spelled as `-1` or `x.dim() - 1`, so the guard is simply `dim == -1 or dim == x.dim() - 1`, and the caller's `dim` is forwarded to the kernel unchanged. No dimension arithmetic is needed, and equality comparison is safe for every value of `dim`: a tuple, for example, simply fails both comparisons and falls back instead of raising. The kernel is therefore used only when all of the following hold: NPU tensor, a genuine last-dim reduction, and dtype in `{fp16, fp32, bf16}`. Forwarding `dim` unchanged means every previously-working last-dim call keeps the exact behavior it had before this PR, and the existing unit test stays valid. 2. Everything else - non-last-dim, tuple `dim`, full reduction (`dim is None` on >1-D), CPU tensors, unsupported dtypes - falls back to the **saved native `torch_sum` reference** captured at module import, so the fallback cannot recurse into the patched entry points. Every input that now falls back previously crashed, so the change strictly widens the set of inputs that run; there is no working behavior being changed. Why this does not affect batch invariance? 1. **The invariance-critical reductions are all last-dim, and they are unchanged.** The reductions whose results must be batch-invariant (softmax denominators, RMSNorm variance - reductions over the hidden/vocab dimension) are last-dim reductions on `[num_tokens, hidden]` tensors. Native kernels pick their reduction strategy from the total tensor size, so the per-row accumulation order can differ between BS=1 and BS=N - that is exactly what `npu_reduce_sum_batch_invariant` prevents with a row-count-independent accumulation order. These calls still take the batch-invariant kernel; their behavior is byte-for-byte identical to before. 2. **The newly-fallback paths have no batch-shape-dependent input.** The non-last-dim sums that now fall back (e.g. `embeds.sum(dim=0)` in multimodal position-embedding preprocessing) are per-request computations: the input tensor is built from one request's own patches, so its shape and values are identical in a BS=1 run and in a BS=N run. Native `torch.sum` is deterministic for a given tensor (same shape ? same kernel and tiling ? same accumulation order), so identical input yields bitwise-identical output across batch sizes. Invariance is preserved. 3. **No regression is possible.** Every path that now falls back previously raised EZ1001 (non-last-dim int dims) or a schema type error (tuple dims). There was no working behavior to preserve. 4. **It follows the established fallback design.** #15956 introduced the same pattern for the dtype axis: unsupported dtypes "fall back to the saved native torch.sum implementation, avoiding recursive calls through the patched Tensor.sum". This PR extends that pattern from dtype to dim. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? - 10-path NPU unit test on Ascend910 (dim=0 fallback, dim=2 / dim=-1 equivalence vs native, keepdim, 1-D `.sum()`, tuple dim, all-dims reduction, fp16, functional `torch.sum`, fp32) - all pass. - End-to-end batch-invariance logprobs test (BS=1 vs BS=N bitwise) runs through `profile_run` and generation with no EZ1001. vllm-project/vllm@b2f6858 --------- Signed-off-by: EliasKaslan <3248569436@qq.com> - vLLM main: vllm-project/vllm@84030bb Signed-off-by: EliasKaslan <3248569436@qq.com>
What this PR does / why we need it?
ref to #10072
Upgrade the batch invariance operator run package and PyTorch extension archive to 2.0.0. Map ascend950 variants to the 950 package and enable operator installation during both Ubuntu and openEuler A5 image builds.
Document Ascend 950 installation and graph support. Remove the forced PIECEWISE configuration from the online/offline examples and the TP4 batch invariance test, preserving its capture sizes and updating the graph coverage annotation.
Restrict the custom reduce-sum operator to FP16, FP32, and BF16. Other dtypes fall back to the saved native torch.sum implementation, avoiding recursive calls through the patched Tensor.sum. Add parameterized regression tests for supported and unsupported NPU dtypes.
Does this PR introduce any user-facing change?
A5 source-built images install the batch invariance operators. Users still set VLLM_BATCH_INVARIANT=1 at inference time. Examples and the TP4 consistency test use the default graph configuration.
How was this patch tested?
VLLM_BATCH_INVARIANT=1 pip install -v -e . --no-build-isolation --no-deps --extra-index-url=https://triton-ascend.osinfra.cn/pypi/simple --trusted-host triton-ascend.osinfra.cn
Enable VLLM_BATCH_INVARIANT=1 on A5, install vllm-ascend, and after installation run test_batch_invariant_tp4.py to verify success.
Build the image with Dockerfile, and verification is OK.
vLLM main: vllm-project/vllm@b2f6858