Skip to content

🐛 [NPU] Fix Qwen3.5 dense serving on Ascend 950 - #32745

Open
TallMessiWu wants to merge 23 commits into
sgl-project:mainfrom
TallMessiWu:junlin_qwen3.5_dense_w8a8
Open

TallMessiWu wants to merge 23 commits into
sgl-project:mainfrom
TallMessiWu:junlin_qwen3.5_dense_w8a8

Conversation

@TallMessiWu

@TallMessiWu TallMessiWu commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

Dependency and current validation

Updated on 2026-10-10 with ordinary merges; base remains main. Branch head: cf48b477db. Includes official main c55790d38b.

Requires a target wheel containing sgl-project/sgl-kernel-npu#638 for the Gemma API on Ascend 950. This branch is the shared-fix source for #32601 (then #32602) and #34387. #36426 is a separate vision-tower validation dependency; it is not merged into this branch.

Current-head checks: 8 isolated Gemma API/registry tests passed (3 SRT cases excluded), and real CPU post-load/scale-expression probes passed for K=32,64,96,4304. Applicable pre-commit hooks passed. These are CPU/static checks, not 910/950 operator, model-load, warmup, accuracy, or performance validation. Full SRT test collection was unavailable in this environment because its serving dependencies are not installed. Target hardware validation remains required.

Earlier validation and measurements below predate this refresh.

Motivation

Two separate defects stop Qwen3.5 dense models from serving on Ascend 950. Each section below covers one; they are independent and can be reviewed separately.

This PR also carried a fix for the NPU graph rebind ordering in NPUCudaGraphBackend.replay_with_input_update, where the replay was issued while the input rebind was still in flight and rtModelExecute failed intermittently. #39589 has since fixed the same race upstream, and more completely -- it blocks on the update future before replaying and reuses one device-bound worker instead of creating a thread per replay -- so that fix is dropped here and the file now matches main.

Qwen3.5 uses Gemma RMSNorm. Ascend 950 has no registered kernel for torch_npu.npu_gemma_rms_norm, so the previous direct NPU path prevents Qwen3.5 from loading and serving on Ascend 950.

This PR is the SGLang half of the paired fix with sgl-project/sgl-kernel-npu#638. sgl-kernel-npu owns build-time SoC dispatch and exposes stable APIs; SGLang does not detect the SoC or choose providers.

Modifications

  • keep KernelBackend.SGL_KERNEL_NPU as the only Gemma RMSNorm backend on Ascend
  • remove Gemma TORCH_NPU registration and the duplicated native forward path
  • import npu_gemma_rms_norm from norm.gemma_rmsnorm for the ordinary path and preserve the existing norm.add_rmsnorm_bias.add_gemma_rms_norm residual API
  • keep SoC/provider probing out of SGLang; preserve the legacy srt import fallback to torch_npu.npu_gemma_rms_norm for pre-provider 910 wheels, while an explicitly selected SGL_KERNEL_NPU backend does not silently fall back when its stable API is unavailable
  • report an actionable package/version error if the stable API is unavailable
  • preserve the established ordinary and residual call signatures in GemmaRMSNorm and Gemma3RMSNorm
  • preserve the explicit pure-Torch debugging switch

Paired architecture

SGLang GemmaRMSNorm
  -> ordinary: sgl-kernel-npu npu_gemma_rms_norm
       -> 910 wheel (A2/A3): native Gemma ACLNN
       -> 950 wheel: standard RMSNorm ACLNN with 1 + weight
  -> residual: sgl-kernel-npu add_gemma_rms_norm
       -> all target wheels: add RMSNorm ACLNN with 1 + weight

For Ascend support, this PR requires a target-specific package built from sgl-project/sgl-kernel-npu#638; its Gemma-specific provider staging has since been generalized into generic target-provider staging by sgl-project/sgl-kernel-npu#734, with the stable operator API unchanged. Provider selection occurs only while building that wheel; 910 is the unified logical target for both A2 and A3.

Additional fix: MXFP8 scale placeholder for partial blocks

ModelSlimMXFP8Scheme.create_weights sizes the weight-scale placeholder as input_size_per_partition // 32, which drops the trailing partial block whenever K is not a multiple of 32, and NPUMXFP8LinearMethod.process_weights_after_loading then pairs the scales with reshape(n, k // 2, 2), which cannot consume an odd scale count.

Both are reachable on a real checkpoint. A ModelSlim W8A8_MXFP8 export of Qwen3.5-27B stores the vision MLP down projection as:

model.visual.blocks.0.mlp.linear_fc2.weight        [1152, 4304]   W8A8_MXFP8
model.visual.blocks.0.mlp.linear_fc2.weight_scale  [1152,  135]   W8A8_MXFP8

ceil(4304 / 32) = 135, while the placeholder allocates 134, so the parameter fails to load; and 135 is odd, so the pair reshape fails as well. The placeholder now rounds up and an odd scale count is padded before the pair reshape. Layers whose K is a multiple of 32 are unaffected (the same block's linear_fc1 has K=1152 and 36 scales).

Reaching the vision tower at all additionally requires #36426, which forwards the ModelSlim config to the Qwen3-VL vision encoder; that PR is independent and not included here. The rounding is a correctness fix on its own: any MXFP8 linear with K % 32 != 0 hits it.

Accuracy Tests

  • registry/mock coverage verifies that Ascend Gemma operations register only SGL_KERNEL_NPU, preserve the generic out= contract, and fail clearly on missing/incompatible packages
  • delegation tests cover the established npu_gemma_rms_norm(...)->(output, rstd) and add_gemma_rms_norm(...)->(norm_output, residual_sum) contracts for GemmaRMSNorm and Gemma3RMSNorm
  • test/registered/unit/layers/quantization/test_modelslim_mxfp8.py: the scale placeholder rounds up for K=4304 and the post-load pair layout pads the odd 135th scale
  • py_compile, Black, isort, and git diff --check passed on Windows
  • Ascend 950PR_958b ordinary-path correctness: 13 cases passed
  • Ascend 950PR_958b residual-add correctness: 32/32 FP16/BF16 cases passed; residual sums matched exactly and normalized outputs passed atol=2e-2, rtol=2e-2

Speed Tests and Profiling

Both provider comparisons used Ascend 950PR_958b, 50 warmups, 200 synchronized p50 iterations, FP16/BF16, rows 1/16/128/512, and hidden sizes 256/2048/4096/5120.

  • ordinary Triton geometric speedup versus npu_rms_norm(input, 1 + weight): -20.25%; worst regression 82.34%
  • residual-add geometric Triton / ACLNN p50 latency ratio: 1.7765x
  • residual-add ACLNN p50 latency reduction relative to Triton: 43.7%
  • both Triton performance gates: FAIL

Therefore, sgl-kernel-npu #638 uses target-specific ACLNN providers for the ordinary Gemma RMSNorm path: native npu_gemma_rms_norm(input, weight, eps) on the 910 target (A2/A3), and standard npu_rms_norm(input, 1 + weight, eps) on the 950 target. The residual-add path uses npu_add_rms_norm(..., 1 + weight, eps) on all supported Ascend targets. SGLang contains no shape- or SoC-based runtime dispatch.


CI States

Latest PR Test (Base): ⏳ Run #38040873295
Latest PR Test (Extra): ❌ Run #38040873070
Latest PR Test (AMD ROCm 10): ❌ Run #38040873274

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@TallMessiWu
TallMessiWu marked this pull request as ready for review July 29, 2026 08:39
@TallMessiWu
TallMessiWu requested a review from merrymercy as a code owner July 29, 2026 08:39
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

Comment thread python/sglang/srt/utils/common.py Outdated
The lint gate rejects new files under test/registered/kernels/ unless
they register a *-kernel-* suite, and no such suite runs on CPU: the
workflows define base-b-kernel-unit-test-* on GPU runners only. This test
mocks sgl_kernel_npu and asserts registry dispatch and error messages, so
it is a CPU unit test rather than a kernel test.

The rejection failed lint, which gates pr-gate, so every test job in the
matrix was skipped -- including base-a-test-cpu, where this test runs.

Path and suite now agree: test/registered/unit/npu/, next to the other
NPU CPU unit tests, keeping register_cpu_ci(suite="base-a-test-cpu").
All 11 tests still pass.
(cherry picked from commit fc9cd5b)
(cherry picked from commit 2a248ca)
@ping1jing2

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 19, 2026
@TallMessiWu
TallMessiWu force-pushed the junlin_qwen3.5_dense_w8a8 branch from 10158fe to 4f35f3e Compare September 20, 2026 06:25
Drop the graph-rebind fix: upstream sgl-project#39589 fixes the same race, and more
completely. Both versions order the input rebind before the replay -- ours by
doing it on the calling thread, upstream's by blocking on the future -- but
upstream reuses one device-bound worker whose executor initializer calls
set_device, instead of creating a thread per replay. Take upstream's file
whole; nothing of ours is left to carry.

The other ten files merged without conflicts.
@TallMessiWu

Copy link
Copy Markdown
Contributor Author

/rerun-failed-ci

@TallMessiWu

Copy link
Copy Markdown
Contributor Author

I checked the current CI failures against the actual CI merge commit 436baecb3ac, its upstream parent 5c5f593110e, and PR head 54b0865665. The failures investigated so far do not show a regression caused by this PR. The clearest cases are an existing upstream HiCache bug and CI dependency/download failures; the multimodal failures still have unresolved causes.

  • NPU HiCache (1-NPU and 16-NPU): existing upstream code issue. The failure is RuntimeError: Boolean value of Tensor with more than one value is ambiguous at pool_host/mha.py:161, from getattr(pool, side, None) or (). The entire file is byte-for-byte identical between the upstream parent and the CI merge commit. A minimal CPU probe using the expression extracted from each version reproduces the same exception when the buffer is a Tensor. The failing Qwen HiCache path does not use the Gemma provider or ModelSlim MXFP8 changes in this PR. Failure log · Upstream source

  • NPU MoE: missing DeepEP operator/library. The stack ends in deep_ep's get_dispatch_layout, with aclnnDispatchLayout or aclnnDispatchLayoutGetWorkspaceSize not in libopapi.so, or libopapi.so not found. The model is Qwen3-30B-A3B-Instruct-2507 with quantization=None; the failing path does not enter the Gemma or ModelSlim MXFP8 changes. This points to operator/library availability in the CI environment. Failure log

  • Base GPU and Xeon: Hugging Face HTTP 429 rate limits. The GPU root failure occurs while fetching the EAGLE3 model; the Xeon failure occurs while fetching THUDM/LongBench-v2. The GPU server exit follows the download failure, and many subsequent red jobs only report the fail-fast dependency gate. These are not evidence of a regression in the PR's runtime changes. GPU failure log

  • AMD and XPU: failures in separate, unchanged paths. The inspected AMD failures include FP8 KV-cache mismatch, an AITER MXFP8 result containing NaN, Triton shared-memory exhaustion, accuracy/performance thresholds, and tracing assertions. In particular, the AMD MXFP8 test calls fp8_hip/AITER, not the NPU ModelSlim MXFP8 implementation modified here. The XPU failure is a mismatch between the expected and actual GPTQ error message. None of these failing implementations is changed by this PR; the added Gemma backend is registered for NPU only. This is a code-path assessment, not a hardware A/B validation. AMD MXFP8 failure log

  • CI gates/artifacts: one 4-NPU job passed Run test and failed only at Upload test logs; Extra CI is blocked by the missing run-ci-extra label. Artifact-upload failure

Two multimodal cases remain inconclusive: MiniMax-H3 exceeds the latency baseline (about 471 s versus 155 s), and GLM-Image receives HTTP 400 from its internal AR /generate service. The inspected paths show no direct overlap with this PR's changes, but the latency case needs an upstream baseline run on the same hardware/environment, and the GLM case needs the AR rejection details before its root cause can be established. GLM failure log

All 13 base-a-test-cpu shards and base-a-test-npu passed on the current head. The earlier failure caused by the new test missing a __main__ entry point was fixed and did not recur. These passing checks do not establish full Ascend 950 end-to-end correctness, but the current failure evidence does not justify changing the Gemma provider or ModelSlim MXFP8 fix to address these red jobs.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation jit-kernel memory-pool npu run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants