Repository navigation
Conversation
CLA Signature PassMcZyWu, thanks for your pull request. All authors of the commits have signed the CLA. 👍 |
CLA Signature PassMcZyWu, thanks for your pull request. All authors of the commits have signed the CLA. 👍 |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Run the existing Qwen3-235B-A22B W8A8 performance case from
testcaseson both CANN 9.0.0 and 9.1.0 A3 images, ensuring that the runtime includes sgl-project/sglang#40814. The images may predate that fix.Modifications
.github/workflows/single-test-npu.ymland retain its pull-request trigger againsttestcases.linux-aarch64-a3-16, using:swr.cn-southwest-2.myhuaweicloud.com/base_image/dockerhub/lmsysorg/sglang:main-cann9.0.0-a3swr.cn-southwest-2.myhuaweicloud.com/base_image/dockerhub/lmsysorg/sglang:main-cann9.1.0-a3e98ffc4cc699d7953f17e751ce669043a7e17bcbdirectly to/sgl-workspace/sglang, preserving the image's CANN, PyTorch and kernel dependencies. Accept an already-applied patch only when the reverse check succeeds; fail if the patch cannot be applied or verified. Verify that Python resolves the image's SGLang package and thatgraph.replay()precedesupdate_future.result().test/registered/npu/performance/qwen3_235b_a22b/test_npu_qwen3_235b_w8a8_8p_in3k5_out1k5_50ms.py, with the testcase and Ascend test helpers from thistestcases-based checkout. Despite the8pfilename, its existing arguments are--nnodes 1 --tp 16 --dp-size 16.--cuda-graph-bs-decodeoption in the testcase to avoid the ambiguous--cuda-graph-bsprefix now shared by decode and prefill options. Keep the existing batch-size values.Accuracy Tests
No model or accuracy-threshold changes. The testcase only changes the decode graph argument name. Hardware execution is delegated to the PR's NPU CI jobs.
Speed Tests and Profiling
The selected case retains its existing 3,500-input / 1,500-output token workload, concurrency 432 and performance thresholds. CANN 9.0.0 and 9.1.0 results are pending the two CI jobs.
Local validation:
git diff --check: passed.e98ffc4cc699d7953f17e751ce669043a7e17bcb. Verified the already-applied and incompatible-source checks.python3alias). Both underlying checks were then run directly with Python and passed.CI States
Latest PR Test (Base): ❌ Run #36369201551
Latest PR Test (Extra): ✅ Run #36369209152
Latest PR Test (AMD ROCm 7.2): ❌ Run #36369201498