Repository navigation
Conversation
…opt NVFP4 Signed-off-by: cjackal <44624812+cjackal@users.noreply.github.com>
|
Hit the same failure with Two notes from checking the diff:
@pytest.mark.parametrize(
"module_name",
["vllm.models.deepseek_v32.nvidia.mtp", "vllm.models.deepseek_v32.amd.mtp"],
)
def test_deepseek_v32_mtp_defers_lm_head(default_vllm_config, module_name):
"""The DSA draft (GLM-5.2/5.3 and DeepSeek-V3.2) must not build a head.
The placeholder is replaced by the target ``lm_head`` after loading; on
quantized GLM checkpoints, which ship no ``shared_head.head``, building it
trips the NVFP4 unloaded-scale check first. The draft config reaches the
layer with ``model_type`` already rewritten to ``deepseek_mtp``.
"""
from vllm.model_executor.models import deepseek_mtp
mtp = importlib.import_module(module_name)
config = mock.MagicMock(
hidden_size=16,
rms_norm_eps=1e-5,
index_topk=8,
model_type="deepseek_mtp",
)
vllm_config = mock.MagicMock()
vllm_config.speculative_config.draft_model_config.hf_config = config
vllm_config.speculative_config.num_speculative_tokens = 1
vllm_config.scheduler_config.max_num_batched_tokens = 4
vllm_config.scheduler_config.max_num_seqs = 1
with (
mock.patch.object(mtp, "DeepseekV32DecoderLayer", return_value=nn.Identity()),
mock.patch.object(deepseek_mtp, "ParallelLMHead") as parallel_lm_head,
mock.patch.object(mtp.current_platform, "device_type", "cpu"),
):
layer = mtp.DeepseekV32MultiTokenPredictorLayer(vllm_config, "model.layers.1")
parallel_lm_head.assert_not_called()
assert layer.shared_head.head is None( Separate from this PR: the AI assistance was used for this check. |
Co-Authored-By: Mikhail Kostryukov <mike@triptrack.net> Signed-off-by: cjackal <44624812+cjackal@users.noreply.github.com>
Purpose
#55442 deferred the disposable
ParallelLMHeadin Glm5Next MTP, but GLM-5.3(glm_moe_dsa) which needs the same fix does not take that path.SharedHeadappendsheadto the layer prefix (model.layers.N.head), while modelopt checkpoints name itmodel.layers.N.shared_head.The ignore entry never matches in
ModelOptQuantConfigBase.is_layer_excluded, the head is allocated as NVFP4, and theweight_scaleNaN-sentinel check added in #52501 abortsget_model()before the proposer can swap in the targetlm_head:Repro:
Inferact/GLM-5.3-NVFP4with--speculative-config.method mtp.Test Plan
Run
Inferact/GLM-5.3-NVFP4with--speculative-config.method mtpand check that model weights are loaded successfully.Test Result
Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model.