Repository navigation
🐛 [NPU] Fix Qwen3.5 dense serving on Ascend 950 - #32745
TallMessiWu wants to merge 23 commits into
Conversation
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
065135b to
2ec8c88
Compare
The lint gate rejects new files under test/registered/kernels/ unless they register a *-kernel-* suite, and no such suite runs on CPU: the workflows define base-b-kernel-unit-test-* on GPU runners only. This test mocks sgl_kernel_npu and asserts registry dispatch and error messages, so it is a CPU unit test rather than a kernel test. The rejection failed lint, which gates pr-gate, so every test job in the matrix was skipped -- including base-a-test-cpu, where this test runs. Path and suite now agree: test/registered/unit/npu/, next to the other NPU CPU unit tests, keeping register_cpu_ci(suite="base-a-test-cpu"). All 11 tests still pass.
|
/tag-and-rerun-ci |
10158fe to
4f35f3e
Compare
Drop the graph-rebind fix: upstream sgl-project#39589 fixes the same race, and more completely. Both versions order the input rebind before the replay -- ours by doing it on the calling thread, upstream's by blocking on the future -- but upstream reuses one device-bound worker whose executor initializer calls set_device, instead of creating a thread per replay. Take upstream's file whole; nothing of ours is left to carry. The other ten files merged without conflicts.
|
/rerun-failed-ci |
|
I checked the current CI failures against the actual CI merge commit
Two multimodal cases remain inconclusive: MiniMax-H3 exceeds the latency baseline (about 471 s versus 155 s), and GLM-Image receives HTTP 400 from its internal AR All 13 |
Dependency and current validation
Updated on 2026-10-10 with ordinary merges; base remains
main. Branch head:cf48b477db. Includes official mainc55790d38b.Requires a target wheel containing sgl-project/sgl-kernel-npu#638 for the Gemma API on Ascend 950. This branch is the shared-fix source for #32601 (then #32602) and #34387. #36426 is a separate vision-tower validation dependency; it is not merged into this branch.
Current-head checks: 8 isolated Gemma API/registry tests passed (3 SRT cases excluded), and real CPU post-load/scale-expression probes passed for K=32,64,96,4304. Applicable pre-commit hooks passed. These are CPU/static checks, not 910/950 operator, model-load, warmup, accuracy, or performance validation. Full SRT test collection was unavailable in this environment because its serving dependencies are not installed. Target hardware validation remains required.
Earlier validation and measurements below predate this refresh.
Motivation
Two separate defects stop Qwen3.5 dense models from serving on Ascend 950. Each section below covers one; they are independent and can be reviewed separately.
This PR also carried a fix for the NPU graph rebind ordering in
NPUCudaGraphBackend.replay_with_input_update, where the replay was issued while the input rebind was still in flight andrtModelExecutefailed intermittently. #39589 has since fixed the same race upstream, and more completely -- it blocks on the update future before replaying and reuses one device-bound worker instead of creating a thread per replay -- so that fix is dropped here and the file now matches main.Qwen3.5 uses Gemma RMSNorm. Ascend 950 has no registered kernel for
torch_npu.npu_gemma_rms_norm, so the previous direct NPU path prevents Qwen3.5 from loading and serving on Ascend 950.This PR is the SGLang half of the paired fix with sgl-project/sgl-kernel-npu#638.
sgl-kernel-npuowns build-time SoC dispatch and exposes stable APIs; SGLang does not detect the SoC or choose providers.Modifications
KernelBackend.SGL_KERNEL_NPUas the only Gemma RMSNorm backend on AscendTORCH_NPUregistration and the duplicated native forward pathnpu_gemma_rms_normfromnorm.gemma_rmsnormfor the ordinary path and preserve the existingnorm.add_rmsnorm_bias.add_gemma_rms_normresidual APIsrtimport fallback totorch_npu.npu_gemma_rms_normfor pre-provider 910 wheels, while an explicitly selectedSGL_KERNEL_NPUbackend does not silently fall back when its stable API is unavailableGemmaRMSNormandGemma3RMSNormPaired architecture
For Ascend support, this PR requires a target-specific package built from sgl-project/sgl-kernel-npu#638; its Gemma-specific provider staging has since been generalized into generic target-provider staging by sgl-project/sgl-kernel-npu#734, with the stable operator API unchanged. Provider selection occurs only while building that wheel;
910is the unified logical target for both A2 and A3.Additional fix: MXFP8 scale placeholder for partial blocks
ModelSlimMXFP8Scheme.create_weightssizes the weight-scale placeholder asinput_size_per_partition // 32, which drops the trailing partial block wheneverKis not a multiple of 32, andNPUMXFP8LinearMethod.process_weights_after_loadingthen pairs the scales withreshape(n, k // 2, 2), which cannot consume an odd scale count.Both are reachable on a real checkpoint. A ModelSlim
W8A8_MXFP8export of Qwen3.5-27B stores the vision MLP down projection as:ceil(4304 / 32) = 135, while the placeholder allocates 134, so the parameter fails to load; and 135 is odd, so the pair reshape fails as well. The placeholder now rounds up and an odd scale count is padded before the pair reshape. Layers whoseKis a multiple of 32 are unaffected (the same block'slinear_fc1hasK=1152and 36 scales).Reaching the vision tower at all additionally requires #36426, which forwards the ModelSlim config to the Qwen3-VL vision encoder; that PR is independent and not included here. The rounding is a correctness fix on its own: any MXFP8 linear with
K % 32 != 0hits it.Accuracy Tests
SGL_KERNEL_NPU, preserve the genericout=contract, and fail clearly on missing/incompatible packagesnpu_gemma_rms_norm(...)->(output, rstd)andadd_gemma_rms_norm(...)->(norm_output, residual_sum)contracts forGemmaRMSNormandGemma3RMSNormtest/registered/unit/layers/quantization/test_modelslim_mxfp8.py: the scale placeholder rounds up forK=4304and the post-load pair layout pads the odd 135th scalepy_compile, Black, isort, andgit diff --checkpassed on Windowsatol=2e-2, rtol=2e-2Speed Tests and Profiling
Both provider comparisons used Ascend 950PR_958b, 50 warmups, 200 synchronized p50 iterations, FP16/BF16, rows
1/16/128/512, and hidden sizes256/2048/4096/5120.npu_rms_norm(input, 1 + weight): -20.25%; worst regression 82.34%Triton / ACLNNp50 latency ratio: 1.7765xTherefore, sgl-kernel-npu #638 uses target-specific ACLNN providers for the ordinary Gemma RMSNorm path: native
npu_gemma_rms_norm(input, weight, eps)on the 910 target (A2/A3), and standardnpu_rms_norm(input, 1 + weight, eps)on the 950 target. The residual-add path usesnpu_add_rms_norm(..., 1 + weight, eps)on all supported Ascend targets. SGLang contains no shape- or SoC-based runtime dispatch.CI States
Latest PR Test (Base): ⏳ Run #38040873295
Latest PR Test (Extra): ❌ Run #38040873070
Latest PR Test (AMD ROCm 10): ❌ Run #38040873274