Skip to content

[Model] Support Qwen3.6 ModelOpt mixed NVFP4 - #27906

Merged
Fridge003 merged 18 commits into
mainfrom
mmangkad/support-qwen36-nvfp4-modelopt
Jul 6, 2026
Merged

Fridge003 merged 18 commits into
mainfrom
mmangkad/support-qwen36-nvfp4-modelopt

Conversation

@mmangkad

@mmangkad mmangkad commented Jun 11, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Support nvidia/Qwen3.6-35B-A3B-NVFP4. Closes #26967.

This checkpoint is a little different from the regular modelopt_fp4 NVFP4 path. Its config.json reports ModelOpt MIXED_PRECISION and lists per-layer FP8 / W4A16 NVFP4 entries under wrapper prefixes like model.language_model.*, including lm_head as W4A16 NVFP4. The FP8 KV-cache setting is only present in hf_quant_config.json, so SGLang needs to read that file instead of relying only on config.json.

FWIW, this official checkpoint does not seem to give any accuracy or serving advantage over a regular NVFP4 export in the validation below. This PR is mainly about supporting the official checkpoint format.

Modifications

  • Handle this checkpoint as ModelOpt mixed NVFP4 instead of routing it through the w4afp8 fallback path.
    • Detect MIXED_PRECISION configs with actual NVFP4 / W4A16_NVFP4 layer entries.
    • Read hf_quant_config.json when config.json is missing runtime metadata such as KV-cache quantization.
  • Resolve the checkpoint's quantized-layer names to SGLang module names.
    • Handle wrapper prefixes such as model.language_model.* and language_model.model.*.
    • Resolve nested/top-level lm_head quantization metadata.
  • Support the quantized dense pieces used by this checkpoint.
    • Treat ModelOpt quantized ParallelLMHead like a quantized linear layer and use its quantized apply path for logits.
    • Add W4A16 NVFP4 dense loading/execution for packed NVFP4 weights with fp16/bf16 activations.
    • Accept scalar ModelOpt scale tensors when loading one-element scale parameters.
  • Support MTP with the checkpoint's quantized target LM head.
    • Keep the Qwen3.5 MTP body unquantized for modelopt_fp4 and modelopt_mixed checkpoints because the tested checkpoint excludes MTP layers from quantization.
    • Share quantized LM-head parameters and runtime metadata from the target model to the draft model for speculative serving.

Accuracy Tests

Validated concurrently on a single GB300 node, with each validation variant pinned to one GPU.

Commands used for the four validation variants

Official checkpoint, no MTP:

CUDA_VISIBLE_DEVICES=0 sglang serve \
  --model-path nvidia/Qwen3.6-35B-A3B-NVFP4 \
  --disable-radix-cache \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --trust-remote-code \
  --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 3}' \
  --max-running-requests 256 \
  --mem-fraction-static 0.85 \
  --port 30000

python -m sglang.test.run_eval \
  --base-url http://127.0.0.1:30000 \
  --model nvidia/Qwen3.6-35B-A3B-NVFP4 \
  --eval-name gsm8k \
  --max-tokens 32768 \
  --num-threads 256 \
  --temperature 1.0 \
  --top-p 0.95

Official checkpoint, MTP:

CUDA_VISIBLE_DEVICES=1 sglang serve \
  --model-path nvidia/Qwen3.6-35B-A3B-NVFP4 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --trust-remote-code \
  --speculative-algorithm EAGLE \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4 \
  --mamba-scheduler-strategy extra_buffer \
  --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 3}' \
  --max-running-requests 256 \
  --mem-fraction-static 0.85 \
  --port 30001

python -m sglang.test.run_eval \
  --base-url http://127.0.0.1:30001 \
  --model nvidia/Qwen3.6-35B-A3B-NVFP4 \
  --eval-name gsm8k \
  --max-tokens 32768 \
  --num-threads 256 \
  --temperature 1.0 \
  --top-p 0.95

Regular NVFP4, no MTP:

CUDA_VISIBLE_DEVICES=2 sglang serve \
  --model-path mmangkad/Qwen3.6-35B-A3B-NVFP4 \
  --disable-radix-cache \
  --quantization modelopt_fp4 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --trust-remote-code \
  --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 1}' \
  --max-running-requests 256 \
  --mem-fraction-static 0.85 \
  --port 30002

python -m sglang.test.run_eval \
  --base-url http://127.0.0.1:30002 \
  --model mmangkad/Qwen3.6-35B-A3B-NVFP4 \
  --eval-name gsm8k \
  --max-tokens 32768 \
  --num-threads 256 \
  --temperature 1.0 \
  --top-p 0.95

Regular NVFP4, MTP:

CUDA_VISIBLE_DEVICES=3 sglang serve \
  --model-path mmangkad/Qwen3.6-35B-A3B-NVFP4 \
  --quantization modelopt_fp4 \
  --reasoning-parser qwen3 \
  --tool-call-parser qwen3_coder \
  --trust-remote-code \
  --speculative-algorithm EAGLE \
  --speculative-num-steps 3 \
  --speculative-eagle-topk 1 \
  --speculative-num-draft-tokens 4 \
  --mamba-scheduler-strategy extra_buffer \
  --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 1}' \
  --max-running-requests 256 \
  --mem-fraction-static 0.85 \
  --port 30003

python -m sglang.test.run_eval \
  --base-url http://127.0.0.1:30003 \
  --model mmangkad/Qwen3.6-35B-A3B-NVFP4 \
  --eval-name gsm8k \
  --max-tokens 32768 \
  --num-threads 256 \
  --temperature 1.0 \
  --top-p 0.95

GSM8K results:

Model Mode GSM8K score
nvidia/Qwen3.6-35B-A3B-NVFP4 no MTP 0.973
nvidia/Qwen3.6-35B-A3B-NVFP4 MTP 0.972
mmangkad/Qwen3.6-35B-A3B-NVFP4 no MTP 0.973
mmangkad/Qwen3.6-35B-A3B-NVFP4 MTP 0.971

Speed Tests and Profiling

Synthetic serving benchmark:

Benchmark commands

Official checkpoint, no MTP:

python -m sglang.bench_serving \
  --port 30000 \
  --model nvidia/Qwen3.6-35B-A3B-NVFP4 \
  --dataset-name random \
  --random-input-len 1024 \
  --random-output-len 1024 \
  --max-concurrency 1 \
  --num-prompts 8 \
  --random-range-ratio 1.0

Official checkpoint, MTP:

python -m sglang.bench_serving \
  --port 30001 \
  --model nvidia/Qwen3.6-35B-A3B-NVFP4 \
  --dataset-name random \
  --random-input-len 1024 \
  --random-output-len 1024 \
  --max-concurrency 1 \
  --num-prompts 8 \
  --random-range-ratio 1.0

Regular NVFP4, no MTP:

python -m sglang.bench_serving \
  --port 30002 \
  --model mmangkad/Qwen3.6-35B-A3B-NVFP4 \
  --dataset-name random \
  --random-input-len 1024 \
  --random-output-len 1024 \
  --max-concurrency 1 \
  --num-prompts 8 \
  --random-range-ratio 1.0

Regular NVFP4, MTP:

python -m sglang.bench_serving \
  --port 30003 \
  --model mmangkad/Qwen3.6-35B-A3B-NVFP4 \
  --dataset-name random \
  --random-input-len 1024 \
  --random-output-len 1024 \
  --max-concurrency 1 \
  --num-prompts 8 \
  --random-range-ratio 1.0

Regular NVFP4 versus official checkpoint:

Mode Output tok/s Mean E2E latency Mean TTFT Mean TPOT
no MTP 1.97x faster 49.1% lower 5.1% higher 50.3% lower
MTP 1.55x faster 35.3% lower 8.2% higher 36.8% lower

Raw results:

Model Mode Output tok/s Mean E2E ms Mean TTFT ms Mean TPOT ms
nvidia/Qwen3.6-35B-A3B-NVFP4 no MTP 163.50 6261.41 125.78 6.00
nvidia/Qwen3.6-35B-A3B-NVFP4 MTP 331.92 3083.37 101.95 2.91
mmangkad/Qwen3.6-35B-A3B-NVFP4 no MTP 321.35 3185.31 132.21 2.98
mmangkad/Qwen3.6-35B-A3B-NVFP4 MTP 512.99 1994.67 110.28 1.84

CI States

Latest PR Test (Base): ✅ Run #28743427486
Latest PR Test (Extra): ✅ Run #28767943010

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces support for the W4A16_NVFP4 quantization method in ModelOpt, updates prefix resolution for quantized layers, and enables sharing of lm_head quantization attributes for speculative draft models (such as Qwen 3.5 MTP and EAGLE). The review feedback identifies two important issues: a potential AttributeError in model_config.py if quantized_layers is explicitly set to null in the configuration, and a bug in logits_processor.py where GGUF models might be incorrectly intercepted by the new quantization check, bypassing the standard FP32 handling.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread python/sglang/srt/layers/logits_processor.py Outdated
Comment thread python/sglang/srt/configs/model_config.py Outdated
@mmangkad

Copy link
Copy Markdown
Collaborator Author

/rerun-test test/registered/4-gpu-models/test_nvidia_nemotron_3_super_nvfp4.py test/registered/gb300/test_qwen35_nvfp4.py test/registered/models_e2e/test_qwen35_fp4_flashinfer.py test/registered/models_e2e/test_qwen35_fp4_mtp.py

@github-actions

github-actions Bot commented Jun 11, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/4-gpu-models/test_nvidia_nemotron_3_super_nvfp4.py test/registered/gb300/test_qwen35_nvfp4.py test/registered/models_e2e/test_qwen35_fp4_flashinfer.py test/registered/models_e2e/test_qwen35_fp4_mtp.py:

🚀 4-gpu-b200 (3 tests): ✅ View workflow run

cd test/ && python3 registered/4-gpu-models/test_nvidia_nemotron_3_super_nvfp4.py
cd test/ && python3 registered/models_e2e/test_qwen35_fp4_flashinfer.py
cd test/ && python3 registered/models_e2e/test_qwen35_fp4_mtp.py

⛔ test/registered/gb300/test_qwen35_nvfp4.py: Suite nightly-4-gpu-gb300 in test/registered/gb300/test_qwen35_nvfp4.py is not dispatchable via /rerun-test. It has no entry in _LEGACY_SUITE_TO_RUNNER_CONFIG — either it is a non-CUDA suite (npu/amd) or it runs on hardware with no matching runner_config in scripts/ci/runner_configs.yml.

@mmangkad

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Jun 11, 2026
Comment thread python/sglang/srt/model_loader/weight_utils.py
@mmangkad
mmangkad enabled auto-merge (squash) June 12, 2026 03:30
@mmangkad
mmangkad disabled auto-merge June 15, 2026 23:25
@Fridge003 Fridge003 added the run-ci-extra CI: also run the extra suite (requires run-ci) label Jul 6, 2026
@Fridge003
Fridge003 merged commit b1942fc into main Jul 6, 2026
624 of 671 checks passed
@Fridge003
Fridge003 deleted the mmangkad/support-qwen36-nvfp4-modelopt branch July 6, 2026 04:31
@Fridge003 Fridge003 added the release-highlight Candidate PR for release note highlight label Jul 6, 2026
@mmangkad mmangkad mentioned this pull request Jul 6, 2026
4 of 5 tasks
@robbiemu

robbiemu commented Jul 6, 2026

Copy link
Copy Markdown

Can confirm this works for nvidia/Qwen3.6-27B-NVFP4 as well — EAGLE/MTP (NEXTN) on the native head, ~61% acceptance, ~2.3× throughput. The 27B dense checkpoint has the same mtp.*-excluded NVFP4 layout as the 35B-A3B.

Chronostasys pushed a commit to MindLab-Research/sglang that referenced this pull request Aug 24, 2026
twu3202 added a commit to twu3202/sglang that referenced this pull request Sep 29, 2026
A compressed-tensors pack-quantized lm_head stores its weight as
weight_packed / weight_scale / weight_shape and has no .weight. The target
serves fine, but NEXTN died in get_embed_and_head and the DFlash2 selector
refused the head, because should_apply_lm_head_quant_method rejected any
head without .weight (a guard for the ModelOpt dtype reads, from sgl-project#27906).

- Admit heads without .weight through the gate; only the three ModelOpt
  methods still need the tensor for their layout checks.
- _compute_lm_head: such heads take the quant_method.apply branch and keep
  the fp32 activation cast the old fallback gave them; the fallback becomes
  an explicit error.
- qwen3_5 / qwen3_5_text accessors hand out None for a packed head; the
  eagle worker shares the module through set_lm_head_from_target and fails
  clearly when it cannot (lm_head_is_packed in spec_utils). Since sgl-project#39643 a
  pipeline stage runs the same init_lm_head. The multi-layer EAGLE worker,
  which shares the tensor only, fails clearly on a packed head.
- Tests: the DFlash2 selector cases are parametrized over both packed
  layouts; a new CPU file covers the gate, the logits path, the accessors
  and init_lm_head.
twu3202 added a commit to twu3202/sglang that referenced this pull request Oct 3, 2026
A compressed-tensors pack-quantized lm_head stores its weight as
weight_packed / weight_scale / weight_shape and has no .weight. The target
serves fine, but NEXTN died in get_embed_and_head and the DFlash2 selector
refused the head, because should_apply_lm_head_quant_method rejected any
head without .weight (a guard for the ModelOpt dtype reads, from sgl-project#27906).

- Admit heads without .weight through the gate; only the three ModelOpt
  methods still need the tensor for their layout checks.
- _compute_lm_head: such heads take the quant_method.apply branch and keep
  the fp32 activation cast the old fallback gave them; the fallback becomes
  an explicit error.
- qwen3_5 / qwen3_5_text accessors hand out None for a packed head; the
  eagle worker shares the module through set_lm_head_from_target and fails
  clearly when it cannot (lm_head_is_packed in spec_utils). Since sgl-project#39643 a
  pipeline stage runs the same init_lm_head. The multi-layer EAGLE worker,
  which shares the tensor only, fails clearly on a packed head.
- Tests: the DFlash2 selector cases are parametrized over both packed
  layouts; a new CPU file covers the gate, the logits path, the accessors
  and init_lm_head.
coconut49 added a commit to coconut49/sglang that referenced this pull request Oct 9, 2026
sgl-project#27906 added _has_lm_head_runtime_attrs and should_apply_lm_head_quant_method,
sgl-project#34158 moved them to the bottom of the file, and the Qwen3.8 rebase (sgl-project#35758)
pasted the old copies back at the top. Python keeps the later definitions, so
the top copies were dead and had already drifted (sgl-project#35120 only touched the live
one). No behavior change.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

quant LLM Quantization release-highlight Candidate PR for release note highlight run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature] Support nvidia/Qwen3.6-35B-A3B-NVFP4

4 participants