Skip to content

[Bugfix][Spec Decode] Defer disposable DeepseekV32 MTP head for modelopt nvfp4 - #58209

Open
cjackal wants to merge 3 commits into
vllm-project:mainfrom
cjackal:inferact-glm53
Open

cjackal wants to merge 3 commits into
vllm-project:mainfrom
cjackal:inferact-glm53

Conversation

@cjackal

@cjackal cjackal commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

#55442 deferred the disposable ParallelLMHead in Glm5Next MTP, but GLM-5.3(glm_moe_dsa) which needs the same fix does not take that path.

SharedHead appends head to the layer prefix (model.layers.N.head), while modelopt checkpoints name it model.layers.N.shared_head.
The ignore entry never matches in ModelOptQuantConfigBase.is_layer_excluded, the head is allocated as NVFP4, and the weight_scale NaN-sentinel check added in #52501 aborts get_model() before the proposer can swap in the target lm_head:

...
(Worker_TP7_DCP7_EP7 pid=1656) ERROR 09-23 00:30:38 [multiproc_executor.py:943] RuntimeError: NVFP4 weight_scale for layer 'parallel_lm_head' was never loaded (still NaN). The checkpoint likely stores this layer as BF16 (not FP4). Fix: pass quant_config=None when constructing this layer, or add it to the quantization ignore list.
(EngineCore pid=606) INFO 09-23 00:30:38 [multiproc_executor.py:472] [shutdown] Executor: waiting for worker exit count=8
...

Repro: Inferact/GLM-5.3-NVFP4 with --speculative-config.method mtp.

Test Plan

Run Inferact/GLM-5.3-NVFP4 with --speculative-config.method mtp and check that model weights are loaded successfully.

Test Result

...
(Worker_TP0_DCP0_EP0 pid=1697)
(Worker_TP0_DCP0_EP0 pid=1697) INFO 09-23 01:55:28 [default_loader.py:430] Loading weights took 232.85 seconds
(Worker_TP0_DCP0_EP0 pid=1697) INFO 09-23 01:55:31 [unquantized.py:498] Using MoEPrepareAndFinalizeNoDPEPMonolithic
...

Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan, such as providing test command.
  • The test results, such as pasting the results comparison before and after, or e2e results
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model.

…opt NVFP4

Signed-off-by: cjackal <44624812+cjackal@users.noreply.github.com>

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify mergify Bot added deepseek Related to DeepSeek models bug Something isn't working labels Sep 22, 2026
@drakosha

Copy link
Copy Markdown
Contributor

Hit the same failure with Inferact/GLM-5.3-NVFP4 and MTP on 4x H200 (main af7f9488c),
tracked as part 1 of #59306. This is the right place for the fix: since #52861 the GLM DSA draft is
built by DeepseekV32MultiTokenPredictorLayer, and #55442 only covered the Glm5Next path.

Two notes from checking the diff:

  1. Head deferral on GLM: validated on our side with an equivalent change (head skipped when the
    target is glm_moe_dsa); the model loads and serves with num_speculative_tokens=3. For GLM
    checkpoints the .shared_head.head. skip in load_weights is a no-op (zai-org/GLM-5.2 and
    the Inferact NVFP4 quant carry only model.layers.78.shared_head.norm.weight for the MTP
    layer), so that run covers this diff as far as GLM goes. deepseek-ai/DeepSeek-V3.2 does ship
    model.layers.61.shared_head.head.weight, so for V3.2 the skip is what keeps loading working;
    I have not run a V3.2 checkpoint against it. ROCm I cannot test.

  2. [Bugfix][Model][Spec Decode] Defer disposable GLM MTP head #55442 came with a unit test for the Glm5Next layer. The same test for this layer fails on
    main and passes with this PR, in case you want to add it to tests/v1/spec_decode/test_mtp.py:

@pytest.mark.parametrize(
    "module_name",
    ["vllm.models.deepseek_v32.nvidia.mtp", "vllm.models.deepseek_v32.amd.mtp"],
)
def test_deepseek_v32_mtp_defers_lm_head(default_vllm_config, module_name):
    """The DSA draft (GLM-5.2/5.3 and DeepSeek-V3.2) must not build a head.

    The placeholder is replaced by the target ``lm_head`` after loading; on
    quantized GLM checkpoints, which ship no ``shared_head.head``, building it
    trips the NVFP4 unloaded-scale check first. The draft config reaches the
    layer with ``model_type`` already rewritten to ``deepseek_mtp``.
    """
    from vllm.model_executor.models import deepseek_mtp

    mtp = importlib.import_module(module_name)

    config = mock.MagicMock(
        hidden_size=16,
        rms_norm_eps=1e-5,
        index_topk=8,
        model_type="deepseek_mtp",
    )
    vllm_config = mock.MagicMock()
    vllm_config.speculative_config.draft_model_config.hf_config = config
    vllm_config.speculative_config.num_speculative_tokens = 1
    vllm_config.scheduler_config.max_num_batched_tokens = 4
    vllm_config.scheduler_config.max_num_seqs = 1

    with (
        mock.patch.object(mtp, "DeepseekV32DecoderLayer", return_value=nn.Identity()),
        mock.patch.object(deepseek_mtp, "ParallelLMHead") as parallel_lm_head,
        mock.patch.object(mtp.current_platform, "device_type", "cpu"),
    ):
        layer = mtp.DeepseekV32MultiTokenPredictorLayer(vllm_config, "model.layers.1")

    parallel_lm_head.assert_not_called()
    assert layer.shared_head.head is None

(import importlib at the top of the file.) On main: 2 failed, 1 passed, with ParallelLMHead
built under prefix model.layers.1.head, which is also why the modelopt ignore list never
matches. With this PR: 3 passed. Run inside vllm/vllm-openai:nightly on CPU with
VLLM_TARGET_DEVICE=cpu.

Separate from this PR: the self._eh_plan = build_glm52_plan(...) if config.model_type == "glm_moe_dsa" else None line a few lines up (#49791) tests the draft config, whose model_type
is already deepseek_mtp after SpeculativeConfig.hf_config_override, so the GLM low-latency
plan for eh_proj has never been built. Noted in #59306.

AI assistance was used for this check.

cjackal and others added 2 commits October 1, 2026 01:35
Co-Authored-By: Mikhail Kostryukov <mike@triptrack.net>
Signed-off-by: cjackal <44624812+cjackal@users.noreply.github.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working deepseek Related to DeepSeek models speculative-decoding

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants