Skip to content

Fix ModelOpt mixed-precision NVFP4 routing - #30078

Closed
robbiemu wants to merge 2 commits into
sgl-project:mainfrom
robbiemu:fix/modelopt-mixed-module-routing
Closed

robbiemu wants to merge 2 commits into
sgl-project:mainfrom
robbiemu:fix/modelopt-mixed-module-routing

Conversation

@robbiemu

@robbiemu robbiemu commented Jul 3, 2026 •

Copy link
Copy Markdown

Motivation

Enable the official nvidia/Qwen3.6-27B-NVFP4 checkpoint to load and generate correctly through SGLang’s ModelOpt mixed-precision path.

The checkpoint declares MIXED_PRECISION with per-module quantized_layers, including W4A16_NVFP4 modules and a quantized LM head. Several independent issues prevented correct execution:

  • Automatic routing selected the uniform W4AFP8 path instead of preserving per-module rules.
  • ModelOpt recognized only the exact NVFP4 label.
  • ParallelLMHead did not receive its ModelOpt quantization method.
  • Logit computation bypassed quantized LM-head methods.
  • Equivalent scalar scale shapes ([] and [1]) failed loading.
  • Qwen weight-name mapping mutated names across unsuccessful mapping attempts.

The destructive name-mapping behavior is also discussed in #23687. This PR does not claim to fully resolve that issue’s separate FP8 checkpoint.

Modifications

  • Route MIXED_PRECISION checkpoints containing quantized_layers through modelopt_mixed.
  • Preserve the existing W4AFP8 fallback for checkpoints without per-module rules.
  • Accept ModelOpt algorithm labels ending in NVFP4, including W4A16_NVFP4.
  • Apply mixed-precision quantization methods to ParallelLMHead.
  • Execute quantized LM heads through their quantization method in the logits processor.
  • Accept equivalent one-element scalar scale representations.
  • Avoid destructive Qwen parameter-name mutation during stacked-parameter mapping.
  • Add unit coverage for routing, W4AFP8 fallback, NVIDIA’s NVFP4 label, quantized LM heads, and scalar scale loading.

Accuracy Tests

Focused unit tests:

10 passed, 8 subtests passed

Tested classes:

test_modelopt_loader.py::TestParseQuantHfConfig
test_modelopt_loader.py::TestModelOptMixedPrecisionConfig
test_logits_processor_quantized_lm_head.py

End-to-end validation used:

Checkpoint: nvidia/Qwen3.6-27B-NVFP4
Revision: 0893e1606ff3d5f97a441f405d5fc541a6bdf404
Hardware: NVIDIA DGX Spark / GB10

The patched runtime completed:

  • Clean model loading and coherent autoregressive generation.
  • Visible chat completion.
  • Parsed function/tool calling.
  • Grammar-constrained strict JSON-schema output.
  • Three simultaneous approximately 250k-token prompts, approximately 750k aggregate, against an 827,041-token shared KV pool.
  • Cancellation followed by successful recovery, without a container restart.

Speed Tests and Profiling

No speed claims are made by this PR.

Before the fix, the quantized LM head could not produce valid logits through the dense matmul path, so there was no valid before/after performance baseline. This change selects the checkpoint-declared quantized execution path.

Checklist

-- BTW, I apologize if I'm new to sglang and not following normal protocol, frankly Im unclear what that is inre reserved issues. I posted this, and did this work, because I needed it locally and thought others might appreciate it.

also, #29857 still applies with this model and fix sadly.

fixes #26967


CI States

Latest PR Test (Base): ❌ Run #28686595548
Latest PR Test (Extra): ❌ Run #28686595441

Honor per-module ModelOpt quantization rules, including NVIDIA W4A16_NVFP4 labels and quantized LM heads. Avoid destructive Qwen weight-name mapping and accept equivalent scalar scale shapes.
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@github-actions github-actions Bot added the quant LLM Quantization label Jul 3, 2026
@HH1162

HH1162 commented Jul 4, 2026

Copy link
Copy Markdown

Hi @robbiemu,

Thanks for the PR — modelopt_mixed is now working. We found a side effect: the --speculative-draft-model-quantization unquant parameter no longer works, causing the draft model to inherit the main model's modelopt_mixed quantization. This leads to a crash at runtime because DFlash reuses the main model's FP4-quantized lm_head (packed 2:1) while the draft model's hidden states are full precision — the dimension mismatch (5120 vs 2560) causes torch.matmul to fail.

Root cause

Two issues:

  1. Config override: override_quantization_method in base_config.py and modelopt_quant.py returns "modelopt_mixed" even when user explicitly set unquant for the draft model.
  2. DFlash logits: DFlash reuses the main model's FP4-quantized lm_head but does raw torch.matmul instead of quant_method.apply, causing dimension mismatch (5120 vs 2560).

Fix

Config layer (standalone, no dependency on this PR)

  • base_config.py + modelopt_quant.py: override_quantization_method returns None when user specifies non-modelopt quant (unquant, fp8)
  • model_config.py: Convert "unquant" to None in from_server_args (same as main model path)

This fix is based on HEAD (104d6bfb3b) and can be applied independently.

DFlash layer (depends on this PR)

  • Import UnquantizedEmbeddingMethod
  • In _greedy_sample_from_vocab_parallel_head (fast path + general path):
    • Check if lm_head has quantization method
    • Cast draft model output to model_config.dtype (bf16/fp16) — draft model outputs float32 which fp4_quantize rejects
    • Call quant_method.apply(lm_head, hs, None) instead of raw torch.matmul
    • Slice full_logits[:, :num_org] for vocab truncation

This fix is based on 604e18bb56 (your branch).

Verification

Both patches tested and verified — model loads correctly, DFlash speculative decoding works, no dimension or dtype errors.

Recommendation

  • Config layer fix should be merged separately (no dependency on this PR)
  • DFlash layer fix should be incorporated into this PR or merged as a follow-up

We plan to submit a separate PR for the DFlash layer fix based on your branch. Happy to discuss or coordinate if needed.

@HH1162

HH1162 commented Jul 4, 2026 •

Copy link
Copy Markdown

Hi @robbiemu,

Thanks for the PR — modelopt_mixed is now working, great work!

We noticed that in our DFlash speculative decoding setup, the --speculative-draft-model-quantization unquant parameter stopped working after this change. It seems the draft model now inherits the main model's modelopt_mixed quantization, which causes a dimension mismatch when DFlash reuses the main model's FP4-quantized lm_head.

We've submitted a fix PR: #30119 that addresses both the config layer and DFlash layer. Happy to discuss or coordinate if needed!

HH1162 added a commit to HH1162/sglang that referenced this pull request Jul 4, 2026
Based on work from PR sgl-project#30078 by @robbiemu which fixed modelopt_mixed
model support in SGLang. That PR introduced a side effect: the
--speculative-draft-model-quantization unquant parameter no longer works,
causing the draft model to inherit the main model's modelopt_mixed
quantization.

After fixing the unquant parameter (config layer changes below), we
discovered a bug in the DFlash worker: it reuses the main model's lm_head
to save memory, but performs raw torch.matmul on the FP4-quantized weight
without calling quant_method.apply to unpack it. Since the FP4 weight is
packed 2:1 (2560 vs 5120), this causes a dimension mismatch crash at
runtime.

Config layer fix:
- base_config.py + modelopt_quant.py: override_quantization_method returns
  None when user specifies non-modelopt quant (unquant, fp8)
- model_config.py: convert unquant to None in from_server_args
  (same as main model path in server_args.py:2590)

DFlash layer fix:
- dflash_worker_v2.py: when lm_head has quantization, call
  quant_method.apply instead of raw torch.matmul
- Cast draft model output to model_config.dtype (bf16/fp16) before
  passing to quant_method.apply (draft outputs float32 which
  fp4_quantize rejects)
- Slice full_logits[:, :num_org] for vocab truncation

Tested with modelopt_mixed main model + DFlash draft model.
Model loads correctly, speculative decoding works, no dimension or
dtype errors.

Related: PR sgl-project#30078
@robbiemu

robbiemu commented Jul 4, 2026 •

Copy link
Copy Markdown
Author

We've submitted a fix PR: #30119 that addresses both the config layer and DFlash layer.

Glad to contribute! It’s a holiday and I’m just on my phone, but does this address the MTP issue you can see at the bottom of #29857 as well?

@HH1162

HH1162 commented Jul 4, 2026

Copy link
Copy Markdown

Hi @robbiemu, happy holidays!

Not yet — since we use DFlash, we only looked at DFlash-related issues so far. We haven't checked the MTP issue in #29857 yet. We'll take a look when we get to it.

@mmangkad

mmangkad commented Jul 6, 2026

Copy link
Copy Markdown
Collaborator

Superseded by #27906

@mmangkad mmangkad closed this Jul 6, 2026
@robbiemu

robbiemu commented Jul 6, 2026 •

Copy link
Copy Markdown
Author

@mmangkad hi Mohammed, Im glad you merged the earlier pr. I have to say that something is wrong with github unless that pr name was just changed, because I thoroughly searched, for more than half an hour, before starting my branch (and its pretty clear in the history that several others did not find it either). Just in case there is some behavior that lead to a name change and that submission, and it is not a problem with github, is there anything learnable there? Does sglang keep prs in a searchable way that we have missed?

Edit: Actually, I see the issue we had. That PR still does not list this model or architecture.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

quant LLM Quantization

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Feature] Support nvidia/Qwen3.6-35B-A3B-NVFP4

3 participants