Repository navigation
[Model] Support Qwen3.6 ModelOpt mixed NVFP4 - #27906
Conversation
There was a problem hiding this comment.
Code Review
This pull request introduces support for the W4A16_NVFP4 quantization method in ModelOpt, updates prefix resolution for quantized layers, and enables sharing of lm_head quantization attributes for speculative draft models (such as Qwen 3.5 MTP and EAGLE). The review feedback identifies two important issues: a potential AttributeError in model_config.py if quantized_layers is explicitly set to null in the configuration, and a bug in logits_processor.py where GGUF models might be incorrectly intercepted by the new quantization check, bypassing the standard FP32 handling.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
|
/rerun-test test/registered/4-gpu-models/test_nvidia_nemotron_3_super_nvfp4.py test/registered/gb300/test_qwen35_nvfp4.py test/registered/models_e2e/test_qwen35_fp4_flashinfer.py test/registered/models_e2e/test_qwen35_fp4_mtp.py |
|
Results for 🚀 ⛔ |
|
/tag-and-rerun-ci |
…modelopt # Conflicts: # python/sglang/srt/layers/quantization/modelopt_quant.py
…delopt' into support-qwen36-nvfp4-modelopt
|
Can confirm this works for nvidia/Qwen3.6-27B-NVFP4 as well — EAGLE/MTP (NEXTN) on the native head, ~61% acceptance, ~2.3× throughput. The 27B dense checkpoint has the same mtp.*-excluded NVFP4 layout as the 35B-A3B. |
A compressed-tensors pack-quantized lm_head stores its weight as weight_packed / weight_scale / weight_shape and has no .weight. The target serves fine, but NEXTN died in get_embed_and_head and the DFlash2 selector refused the head, because should_apply_lm_head_quant_method rejected any head without .weight (a guard for the ModelOpt dtype reads, from sgl-project#27906). - Admit heads without .weight through the gate; only the three ModelOpt methods still need the tensor for their layout checks. - _compute_lm_head: such heads take the quant_method.apply branch and keep the fp32 activation cast the old fallback gave them; the fallback becomes an explicit error. - qwen3_5 / qwen3_5_text accessors hand out None for a packed head; the eagle worker shares the module through set_lm_head_from_target and fails clearly when it cannot (lm_head_is_packed in spec_utils). Since sgl-project#39643 a pipeline stage runs the same init_lm_head. The multi-layer EAGLE worker, which shares the tensor only, fails clearly on a packed head. - Tests: the DFlash2 selector cases are parametrized over both packed layouts; a new CPU file covers the gate, the logits path, the accessors and init_lm_head.
A compressed-tensors pack-quantized lm_head stores its weight as weight_packed / weight_scale / weight_shape and has no .weight. The target serves fine, but NEXTN died in get_embed_and_head and the DFlash2 selector refused the head, because should_apply_lm_head_quant_method rejected any head without .weight (a guard for the ModelOpt dtype reads, from sgl-project#27906). - Admit heads without .weight through the gate; only the three ModelOpt methods still need the tensor for their layout checks. - _compute_lm_head: such heads take the quant_method.apply branch and keep the fp32 activation cast the old fallback gave them; the fallback becomes an explicit error. - qwen3_5 / qwen3_5_text accessors hand out None for a packed head; the eagle worker shares the module through set_lm_head_from_target and fails clearly when it cannot (lm_head_is_packed in spec_utils). Since sgl-project#39643 a pipeline stage runs the same init_lm_head. The multi-layer EAGLE worker, which shares the tensor only, fails clearly on a packed head. - Tests: the DFlash2 selector cases are parametrized over both packed layouts; a new CPU file covers the gate, the logits path, the accessors and init_lm_head.
sgl-project#27906 added _has_lm_head_runtime_attrs and should_apply_lm_head_quant_method, sgl-project#34158 moved them to the bottom of the file, and the Qwen3.8 rebase (sgl-project#35758) pasted the old copies back at the top. Python keeps the later definitions, so the top copies were dead and had already drifted (sgl-project#35120 only touched the live one). No behavior change.
Motivation
Support
nvidia/Qwen3.6-35B-A3B-NVFP4. Closes #26967.This checkpoint is a little different from the regular
modelopt_fp4NVFP4 path. Itsconfig.jsonreports ModelOptMIXED_PRECISIONand lists per-layer FP8 / W4A16 NVFP4 entries under wrapper prefixes likemodel.language_model.*, includinglm_headas W4A16 NVFP4. The FP8 KV-cache setting is only present inhf_quant_config.json, so SGLang needs to read that file instead of relying only onconfig.json.FWIW, this official checkpoint does not seem to give any accuracy or serving advantage over a regular NVFP4 export in the validation below. This PR is mainly about supporting the official checkpoint format.
Modifications
w4afp8fallback path.MIXED_PRECISIONconfigs with actualNVFP4/W4A16_NVFP4layer entries.hf_quant_config.jsonwhenconfig.jsonis missing runtime metadata such as KV-cache quantization.model.language_model.*andlanguage_model.model.*.lm_headquantization metadata.ParallelLMHeadlike a quantized linear layer and use its quantized apply path for logits.modelopt_fp4andmodelopt_mixedcheckpoints because the tested checkpoint excludes MTP layers from quantization.Accuracy Tests
Validated concurrently on a single GB300 node, with each validation variant pinned to one GPU.
Commands used for the four validation variants
Official checkpoint, no MTP:
CUDA_VISIBLE_DEVICES=0 sglang serve \ --model-path nvidia/Qwen3.6-35B-A3B-NVFP4 \ --disable-radix-cache \ --reasoning-parser qwen3 \ --tool-call-parser qwen3_coder \ --trust-remote-code \ --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 3}' \ --max-running-requests 256 \ --mem-fraction-static 0.85 \ --port 30000 python -m sglang.test.run_eval \ --base-url http://127.0.0.1:30000 \ --model nvidia/Qwen3.6-35B-A3B-NVFP4 \ --eval-name gsm8k \ --max-tokens 32768 \ --num-threads 256 \ --temperature 1.0 \ --top-p 0.95Official checkpoint, MTP:
CUDA_VISIBLE_DEVICES=1 sglang serve \ --model-path nvidia/Qwen3.6-35B-A3B-NVFP4 \ --reasoning-parser qwen3 \ --tool-call-parser qwen3_coder \ --trust-remote-code \ --speculative-algorithm EAGLE \ --speculative-num-steps 3 \ --speculative-eagle-topk 1 \ --speculative-num-draft-tokens 4 \ --mamba-scheduler-strategy extra_buffer \ --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 3}' \ --max-running-requests 256 \ --mem-fraction-static 0.85 \ --port 30001 python -m sglang.test.run_eval \ --base-url http://127.0.0.1:30001 \ --model nvidia/Qwen3.6-35B-A3B-NVFP4 \ --eval-name gsm8k \ --max-tokens 32768 \ --num-threads 256 \ --temperature 1.0 \ --top-p 0.95Regular NVFP4, no MTP:
CUDA_VISIBLE_DEVICES=2 sglang serve \ --model-path mmangkad/Qwen3.6-35B-A3B-NVFP4 \ --disable-radix-cache \ --quantization modelopt_fp4 \ --reasoning-parser qwen3 \ --tool-call-parser qwen3_coder \ --trust-remote-code \ --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 1}' \ --max-running-requests 256 \ --mem-fraction-static 0.85 \ --port 30002 python -m sglang.test.run_eval \ --base-url http://127.0.0.1:30002 \ --model mmangkad/Qwen3.6-35B-A3B-NVFP4 \ --eval-name gsm8k \ --max-tokens 32768 \ --num-threads 256 \ --temperature 1.0 \ --top-p 0.95Regular NVFP4, MTP:
CUDA_VISIBLE_DEVICES=3 sglang serve \ --model-path mmangkad/Qwen3.6-35B-A3B-NVFP4 \ --quantization modelopt_fp4 \ --reasoning-parser qwen3 \ --tool-call-parser qwen3_coder \ --trust-remote-code \ --speculative-algorithm EAGLE \ --speculative-num-steps 3 \ --speculative-eagle-topk 1 \ --speculative-num-draft-tokens 4 \ --mamba-scheduler-strategy extra_buffer \ --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 1}' \ --max-running-requests 256 \ --mem-fraction-static 0.85 \ --port 30003 python -m sglang.test.run_eval \ --base-url http://127.0.0.1:30003 \ --model mmangkad/Qwen3.6-35B-A3B-NVFP4 \ --eval-name gsm8k \ --max-tokens 32768 \ --num-threads 256 \ --temperature 1.0 \ --top-p 0.95GSM8K results:
nvidia/Qwen3.6-35B-A3B-NVFP4nvidia/Qwen3.6-35B-A3B-NVFP4mmangkad/Qwen3.6-35B-A3B-NVFP4mmangkad/Qwen3.6-35B-A3B-NVFP4Speed Tests and Profiling
Synthetic serving benchmark:
Benchmark commands
Official checkpoint, no MTP:
Official checkpoint, MTP:
Regular NVFP4, no MTP:
Regular NVFP4, MTP:
Regular NVFP4 versus official checkpoint:
Raw results:
nvidia/Qwen3.6-35B-A3B-NVFP4nvidia/Qwen3.6-35B-A3B-NVFP4mmangkad/Qwen3.6-35B-A3B-NVFP4mmangkad/Qwen3.6-35B-A3B-NVFP4CI States
Latest PR Test (Base): ✅ Run #28743427486
Latest PR Test (Extra): ✅ Run #28767943010