Skip to content

[NPU] Fix ModelSlim NEXTN startup: unquantized MoE fallback and UB-aware mamba scatter - #34353

Open
w1ida wants to merge 15 commits into
sgl-project:mainfrom
w1ida:fix/npu-modelslim-mtp-nextn
Open

w1ida wants to merge 15 commits into
sgl-project:mainfrom
w1ida:fix/npu-modelslim-mtp-nextn

Conversation

@w1ida

@w1ida w1ida commented Aug 11, 2026 •

Copy link
Copy Markdown

close Issue #34211

Motivation

Fixes startup failure when serving ModelSlim-quantized Qwen3.5 NEXTN checkpoints on NPU, where the MTP (draft) module is stored unquantized (all mtp.* entries in quant_model_description.json are FLOAT) while the main model is W8A8_DYNAMIC.

Additionally fixes a runtime crash on the first inference request due to Ascend UB overflow in the mamba state scatter kernel.

Modifications

1. ModelSlim MoE unquantized fallback (modelslim.py)

  • get_quant_method: Use is_layer_skipped before get_moe_scheme and return None only when every expert projection is explicitly marked FLOAT
  • Preserve get_moe_scheme validation and ValueError handling for missing, mixed, or unsupported schemes

2. MTP whole-model unquant detection (qwen3_5_mtp.py)

  • Add ModelSlim branch to _mtp_quant_config helper that detects when all mtp.* entries are FLOAT and drops quant_config entirely
  • This ensures the bf16 MoE dispatch overrides in forward() (SGLANG_DEEPEP_BF16_DISPATCH, DEEP_NORMAL_MODE_USE_INT8_QUANT) are triggered

3. UB-aware mamba scatter (ascend_hybrid_linear_attn_backend.py)

  • Compute h_block_size dynamically from Ascend UB budget (192KB) instead of hardcoding h_block_size=2
  • For this model (V=128, K=128, fp32), the tile is 128KB single-buffered but Bisheng multi-buffer pass doubles it to 256KB, exceeding 192KB UB
  • The computed h_block_size=1 brings the tile to 64KB single / 128KB double, fitting within budget
  • Add warning when a single (V,K) plane already overflows UB

Accuracy Tests

Before:

  • Startup fails with ValueError: Unsupported ModelSlim MoE schemes for layer mtp.layers.0.mlp.experts: W13='FLOAT', W2='FLOAT'
  • After workaround with --speculative-draft-model-quantization unquant, first request crashes with error: ub overflow, requires 2097152 bits while 1572864 bits available!

After:

  • Startup succeeds with original command line (no extra flags needed)
  • MTP loads in bf16 (3.90 GB as expected)
  • First request completes successfully
  • Unit-tested move_intermediate_cache with h_block_size=1: numerically exact vs reference (max abs diff = 0.0)

Known limitation:
Speculative decoding output quality degrades (garbled/repeated text) compared to no-spec baseline on this checkpoint. This is a separate pre-existing issue in the NPU verify/state-commit path, not caused by these fixes. The corruption persists across graph/radix toggles and draft depths, pointing to the SSM state scatter indexing or conv-state management. Baseline (no spec) output is correct.

Speed Tests and Profiling

python3 -m sglang.benchmark.serving \
  --backend sglang \
  --base-url http://127.0.0.1:8898 \
  --model /workspace/user_data/Qwen3.6-35B-A3B-w8a8 \
  --served-model-name Qwen3.6-35B-A3B \
  --dataset-name random-ids \
  --num-prompts 100 \
  --random-input-len 1024 \
  --random-output-len 100 \
  --random-range-ratio 1.0 \
  --request-rate inf \
  --max-concurrency <1|5|10> \
  --tokenize-prompt \
  --output-file /tmp/bench_conc_<N>.jsonl \
  --disable-tqdm

Metrics

Metric Concurrency 1 Concurrency 5 Concurrency 10
Benchmark Duration (s) 181.1 87.2 56.0
Successful Requests 100 100 100
Request Throughput (req/s) 0.55 1.15 1.78
Input Throughput (tok/s) 565.4 1174.5 1827.2
Output Throughput (tok/s) 55.2 114.7 178.4
Total Throughput (tok/s) 620.6 1289.2 2005.7
Mean E2E Latency (ms) 1810 4339 5419
Mean TTFT (ms) 324 1634 1288
Mean TPOT (ms/tok) 15.0 27.3 41.7

Checklist

  • Format your code according to the Format code with pre-commit
  • Add unit tests
  • Update documentation
  • Provide accuracy and speed benchmark results
  • Follow the SGLang code style guidance

CI States

Latest PR Test (Base): ✅ Run #36371641895
Latest PR Test (Extra): ❌ Run #36371641899
Latest PR Test (AMD ROCm 10): ❌ Run #36371641912

…are mamba scatter

- ModelSlim MoE: return None for FLOAT experts, let FusedMoE fall back to NPUUnquantMoEMethod
- qwen3_5_mtp: detect all-FLOAT mtp.* in _mtp_quant_config and drop quant_config entirely
- ascend_hybrid_linear_attn: compute h_block_size from UB budget to avoid Bisheng ub overflow
Comment thread python/sglang/srt/layers/quantization/modelslim/modelslim.py Outdated
Comment thread python/sglang/srt/layers/quantization/modelslim/modelslim.py Outdated
@TamirBaydasov

Copy link
Copy Markdown
Contributor

LGTM

@TamirBaydasov TamirBaydasov left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for applying requested changes, now modelslim.py looks much better.

@w1ida

w1ida commented Aug 28, 2026

Copy link
Copy Markdown
Author

Hi @OrangeRedeng — gentle follow-up on this PR.

The requested ModelSlim changes have been addressed, both review threads are resolved, and TamirBaydasov has approved that part. I also synced the latest main today. Lint is currently running, while the remaining CI jobs are gated by the missing run-ci label.

Could you please take another look or add the run-ci label when convenient? If another maintainer should review the remaining NPU changes, a redirect would also be greatly appreciated. Thanks!

@OrangeRedeng

Copy link
Copy Markdown
Collaborator

Hi, @w1ida! Could you please fix Lint?

@w1ida

w1ida commented Aug 29, 2026

Copy link
Copy Markdown
Author

@OrangeRedeng I missed the lint issue earlier, but it’s fixed and passing now. Thanks~

@ping1jing2

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@ping1jing2 ping1jing2 self-assigned this Aug 29, 2026
@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Aug 29, 2026
@Mr-qiji

Mr-qiji commented Sep 3, 2026

Copy link
Copy Markdown

Independent repro of the UB overflow from a different ModelSlim checkpoint (Qwen3.8-27B-W8A8, GDN hybrid + MTP, TP2, Ascend 910 / 9362, sgl-kernel-npu 2026.6.1, sglang 0.5.17.dev):

  • Server crashes during target-verify graph capture with the identical error: ub overflow, requires 2097152 bits while 1572864 bits available! in move_cache_dynamic_last_kernel_h_block (via update_mamba_state_after_mtp_verify).
  • For this model's intermediate cache shape your _mamba_scatter_h_block_size computes h_block_size=1; we independently verified h_block_size=1: it compiles, and a 512-token MTP generation is fully coherent (matches the no-spec baseline), so the commit at h=1 is numerically correct, not just compile-able.
  • Note: in sgl-kernel-npu the move_intermediate_cache(...) default is still h_block_size=2 (kernel PR Improve benchmark scripts & rename some scripts #477 was closed without merging), so until this PR lands any ModelSlim + MTP NPU deployment hits this at startup.

Happy to add this repro to the PR description if useful.

@w1ida

w1ida commented Sep 18, 2026

Copy link
Copy Markdown
Author

Hi @iforgetmyname could you please help take a look at the merge status of this NPU PR?

The review comments have been addressed, the relevant Codeowner reviews are approved, and NPU CI is passing. The latest Base CI was cancelled before completion, so the PR is still blocked.

Could you help move it toward merge, or advise if any remaining check needs to be addressed? Thanks!

@ping1jing2

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

w1ida commented Sep 29, 2026

Copy link
Copy Markdown
Author

@whybeyoung @iforgetmyname Could we proceed with review/merge based on the current CI evidence?

I’ve rerun the NPU CI multiple times, and the remaining failures look unrelated to this PR rather than regressions introduced here:

Given that the relevant NPU path for this PR is passing and the remaining failures match known/unrelated CI issues, could we treat them as non-blocking and proceed with review/merge? Thanks!

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

npu run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants