Repository navigation
Conversation
…are mamba scatter - ModelSlim MoE: return None for FLOAT experts, let FusedMoE fall back to NPUUnquantMoEMethod - qwen3_5_mtp: detect all-FLOAT mtp.* in _mtp_quant_config and drop quant_config entirely - ascend_hybrid_linear_attn: compute h_block_size from UB budget to avoid Bisheng ub overflow
|
LGTM |
TamirBaydasov
left a comment
There was a problem hiding this comment.
Thanks for applying requested changes, now modelslim.py looks much better.
|
Hi @OrangeRedeng — gentle follow-up on this PR. The requested ModelSlim changes have been addressed, both review threads are resolved, and TamirBaydasov has approved that part. I also synced the latest Could you please take another look or add the |
|
Hi, @w1ida! Could you please fix Lint? |
|
@OrangeRedeng I missed the lint issue earlier, but it’s fixed and passing now. Thanks~ |
|
/tag-and-rerun-ci |
|
Independent repro of the UB overflow from a different ModelSlim checkpoint (Qwen3.8-27B-W8A8, GDN hybrid + MTP, TP2, Ascend 910 / 9362, sgl-kernel-npu 2026.6.1, sglang 0.5.17.dev):
Happy to add this repro to the PR description if useful. |
|
Hi @iforgetmyname could you please help take a look at the merge status of this NPU PR? The review comments have been addressed, the relevant Codeowner reviews are approved, and NPU CI is passing. The latest Base CI was cancelled before completion, so the PR is still blocked. Could you help move it toward merge, or advise if any remaining check needs to be addressed? Thanks! |
|
/tag-and-rerun-ci |
|
@whybeyoung @iforgetmyname Could we proceed with review/merge based on the current CI evidence? I’ve rerun the NPU CI multiple times, and the remaining failures look unrelated to this PR rather than regressions introduced here:
Given that the relevant NPU path for this PR is passing and the remaining failures match known/unrelated CI issues, could we treat them as non-blocking and proceed with review/merge? Thanks! |
close Issue #34211
Motivation
Fixes startup failure when serving ModelSlim-quantized Qwen3.5 NEXTN checkpoints on NPU, where the MTP (draft) module is stored unquantized (all
mtp.*entries inquant_model_description.jsonareFLOAT) while the main model isW8A8_DYNAMIC.Additionally fixes a runtime crash on the first inference request due to Ascend UB overflow in the mamba state scatter kernel.
Modifications
1. ModelSlim MoE unquantized fallback (
modelslim.py)get_quant_method: Useis_layer_skippedbeforeget_moe_schemeand returnNoneonly when every expert projection is explicitly markedFLOATget_moe_schemevalidation andValueErrorhandling for missing, mixed, or unsupported schemes2. MTP whole-model unquant detection (
qwen3_5_mtp.py)_mtp_quant_confighelper that detects when allmtp.*entries areFLOATand dropsquant_configentirelyforward()(SGLANG_DEEPEP_BF16_DISPATCH,DEEP_NORMAL_MODE_USE_INT8_QUANT) are triggered3. UB-aware mamba scatter (
ascend_hybrid_linear_attn_backend.py)h_block_sizedynamically from Ascend UB budget (192KB) instead of hardcodingh_block_size=2h_block_size=1brings the tile to 64KB single / 128KB double, fitting within budgetAccuracy Tests
Before:
ValueError: Unsupported ModelSlim MoE schemes for layer mtp.layers.0.mlp.experts: W13='FLOAT', W2='FLOAT'--speculative-draft-model-quantization unquant, first request crashes witherror: ub overflow, requires 2097152 bits while 1572864 bits available!After:
move_intermediate_cachewithh_block_size=1: numerically exact vs reference (max abs diff = 0.0)Known limitation:
Speculative decoding output quality degrades (garbled/repeated text) compared to no-spec baseline on this checkpoint. This is a separate pre-existing issue in the NPU verify/state-commit path, not caused by these fixes. The corruption persists across graph/radix toggles and draft depths, pointing to the SSM state scatter indexing or conv-state management. Baseline (no spec) output is correct.
Speed Tests and Profiling
Metrics
Checklist
CI States
Latest PR Test (Base): ✅ Run #36371641895
Latest PR Test (Extra): ❌ Run #36371641899
Latest PR Test (AMD ROCm 10): ❌ Run #36371641912