Repository navigation
Conversation
Honor per-module ModelOpt quantization rules, including NVIDIA W4A16_NVFP4 labels and quantized LM heads. Avoid destructive Qwen weight-name mapping and accept equivalent scalar scale shapes.
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
|
Hi @robbiemu, Thanks for the PR — Root causeTwo issues:
FixConfig layer (standalone, no dependency on this PR)
This fix is based on HEAD ( DFlash layer (depends on this PR)
This fix is based on VerificationBoth patches tested and verified — model loads correctly, DFlash speculative decoding works, no dimension or dtype errors. Recommendation
We plan to submit a separate PR for the DFlash layer fix based on your branch. Happy to discuss or coordinate if needed. |
|
Hi @robbiemu, Thanks for the PR — We noticed that in our DFlash speculative decoding setup, the We've submitted a fix PR: #30119 that addresses both the config layer and DFlash layer. Happy to discuss or coordinate if needed! |
Based on work from PR sgl-project#30078 by @robbiemu which fixed modelopt_mixed model support in SGLang. That PR introduced a side effect: the --speculative-draft-model-quantization unquant parameter no longer works, causing the draft model to inherit the main model's modelopt_mixed quantization. After fixing the unquant parameter (config layer changes below), we discovered a bug in the DFlash worker: it reuses the main model's lm_head to save memory, but performs raw torch.matmul on the FP4-quantized weight without calling quant_method.apply to unpack it. Since the FP4 weight is packed 2:1 (2560 vs 5120), this causes a dimension mismatch crash at runtime. Config layer fix: - base_config.py + modelopt_quant.py: override_quantization_method returns None when user specifies non-modelopt quant (unquant, fp8) - model_config.py: convert unquant to None in from_server_args (same as main model path in server_args.py:2590) DFlash layer fix: - dflash_worker_v2.py: when lm_head has quantization, call quant_method.apply instead of raw torch.matmul - Cast draft model output to model_config.dtype (bf16/fp16) before passing to quant_method.apply (draft outputs float32 which fp4_quantize rejects) - Slice full_logits[:, :num_org] for vocab truncation Tested with modelopt_mixed main model + DFlash draft model. Model loads correctly, speculative decoding works, no dimension or dtype errors. Related: PR sgl-project#30078
|
Superseded by #27906 |
|
@mmangkad hi Mohammed, Im glad you merged the earlier pr. I have to say that something is wrong with github unless that pr name was just changed, because I thoroughly searched, for more than half an hour, before starting my branch (and its pretty clear in the history that several others did not find it either). Just in case there is some behavior that lead to a name change and that submission, and it is not a problem with github, is there anything learnable there? Does sglang keep prs in a searchable way that we have missed? Edit: Actually, I see the issue we had. That PR still does not list this model or architecture. |
Motivation
Enable the official
nvidia/Qwen3.6-27B-NVFP4checkpoint to load and generate correctly through SGLang’s ModelOpt mixed-precision path.The checkpoint declares
MIXED_PRECISIONwith per-modulequantized_layers, includingW4A16_NVFP4modules and a quantized LM head. Several independent issues prevented correct execution:NVFP4label.ParallelLMHeaddid not receive its ModelOpt quantization method.[]and[1]) failed loading.The destructive name-mapping behavior is also discussed in #23687. This PR does not claim to fully resolve that issue’s separate FP8 checkpoint.
Modifications
MIXED_PRECISIONcheckpoints containingquantized_layersthroughmodelopt_mixed.NVFP4, includingW4A16_NVFP4.ParallelLMHead.Accuracy Tests
Focused unit tests:
Tested classes:
End-to-end validation used:
The patched runtime completed:
Speed Tests and Profiling
No speed claims are made by this PR.
Before the fix, the quantized LM head could not produce valid logits through the dense matmul path, so there was no valid before/after performance baseline. This change selects the checkpoint-declared quantized execution path.
Checklist
black --check; the complete pre-commit suite was not available locally.-- BTW, I apologize if I'm new to sglang and not following normal protocol, frankly Im unclear what that is inre reserved issues. I posted this, and did this work, because I needed it locally and thought others might appreciate it.
also, #29857 still applies with this model and fix sadly.
fixes #26967
CI States
Latest PR Test (Base): ❌ Run #28686595548
Latest PR Test (Extra): ❌ Run #28686595441