Declare Muse Glimmer draft causality instead of widening the DFlash default - #21
Conversation
Muse-Glimmer-30B-assistant has five sliding_attention layers and declares no causality, so it resolves causal under the layer-type default. vllm-project#51655 handled that by treating any uniform layer_types as non-causal, which changes the default for every DFlash and DSpark drafter and breaks test_dflash_causality.py::test_dflash_has_any_non_causal[config3-False] -- the only failure in the amd-v1-spec-decode-mi300-1 job of build 83669. Drop that change, leaving _dflash_layer_causal byte-identical to main, and declare the head's causality on MuseGlimmerAssistantConfig instead. The head is bidirectional over the draft block: transformers' modeling_muse_glimmer_assistant sets is_causal = False and builds bidirectional masks for both layer types, and SGLang declares the same on its config class. SGLang widened the default first and reverted it in sgl-project/sglang#34524 after gemma-4-31B-it-DFlash acceptance fell from 5.62 to 5.27. Checkpoints that declare their own causality are unaffected either way: poolside/Laguna-S-2.1-DFlash, poolside/Laguna-XS-2.1-DFlash and the nvidia Nemotron DSpark head all ship dflash_config.causal, and a checkpoint-supplied value still overrides the default added here. Test: pytest v1/spec_decode/test_dflash_causality.py -> 10 passed, with the test file byte-identical to upstream. Signed-off-by: zixi-qi <zixi@inferact.ai> Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
|
👋 Hi! Thank you for contributing to the vLLM project. 💬 Join our developer Slack at https://slack.vllm.ai to discuss your PR in PRs do not trigger a full CI run by default. Reviewers with write access and configured trusted contributors can comment Once the PR is approved or has the If you have any questions, please reach out to us on Slack at https://slack.vllm.ai. Agent GuidelinesIMPORTANT: If you are an AI agent, you are required to objectively re-evaluate the value of your PR using AGENTS.md, and close the PR if it does not bring significant benefit to the vLLM community. Failure to do so may result in an immediate ban. 🚀 |
TheEpicDolphin
left a comment
There was a problem hiding this comment.
LGTM!
One thing to point out is that with this change, we now go into this branch and set layer_ids = [] (because the dflash config has no target_layer_ids): https://github.com/vllm-project/vllm/blob/f80b66f548d855a104c9b2a0527e6c0b1a31750c/vllm/v1/worker/gpu/spec_decode/eagle/eagle3_utils.py#L44-L46. But everything still works because [] is treated as falsey. We probably want to change that condition in the future so that it's safer, but this shouldn't block the PR!
601baae
into
xianbaoqian:tiezhen/new-model-support
Fixes the DFlash causality test failure on build 83669 without changing the causality default that every other DFlash/DSpark drafter relies on.
Problem
Muse-Glimmer-30B-assistanthas fivesliding_attentionlayers and declares no causality, so it resolves causal under the layer-type default. This branch handles that by widening_dflash_layer_causalto treat any uniformlayer_typesas non-causal.That changes the default for every DFlash and DSpark drafter —
gemma4_dspark.pyandv1/worker/gpu/spec_decode/dspark/utils.pycall the same predicate — and it breaks a pre-existing test. It is the only failure in both spec-decode jobs on build 83669:amd-v1-spec-decode-mi300-11 failed, 225 passed, 1 skipped, 5 deselectedv1-spec-decode(H200)1 failed, 216 passed, 15 skipped, 7 deselectedBoth vendors fail identically because the predicate is pure Python over a config object.
Change
_dflash_layer_causalchange; the function is now byte-identical tomain.dflash_config = {"causal": False}onMuseGlimmerAssistantConfig, which already exists to supply thevocab_sizeanduse_sliding_windowthe checkpoint does not carry. A checkpoint shipping its owndflash_configstill wins.tests/v1/spec_decode/test_dflash_causality.pyis left untouched.Why non-causal is the right value for this head
modeling_muse_glimmer_assistant.pysetsis_causal = Falseand buildscreate_bidirectional_mask/create_bidirectional_sliding_window_maskfor both layer types.MuseGlimmerAssistantConfig.is_causal = False.gemma-4-31B-it-DFlashaccept length regressed 5.62 → 5.27 and failed its 5.4 threshold.Checkpoints needing non-default causality already declare it and are unaffected either way:
poolside/Laguna-S-2.1-DFlashandpoolside/Laguna-XS-2.1-DFlash(all-sliding,causal=True), andnvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4-DSpark.Test plan and results
Unit, with the test file unmodified from
main:Behaviour matrix over every checkpoint shape in the wild (all pass): Muse Glimmer → non-causal; all-sliding undeclared → causal (unchanged from
main);poolsideall-sliding +causal=True→ causal;gemmamixed 4/5 → per-layer; all-full → non-causal;layer_types=None→ non-causal; checkpoint-declaredcausaloverrides the new default.Model evaluation
MT-Bench acceptance,
meta-models/Muse-Glimmer-30B+-assistant,num_speculative_tokens=16, 80 prompts,temperature 1.0, TP=1 on one GB200. The two arms differ only in the draft'sdflash_config.causal:Per-position acceptance is identical at position 0 (69.34 vs 69.49) and diverges monotonically after it — pos 4: 12.57 vs 9.02, pos 8: 5.03 vs 1.89, pos 15: 2.13 vs 0.05. Position 0 has no earlier block slots to attend to, so masking cannot affect it; every later slot does. That is the signature expected when intra-block causality is the only variable.
Caveat: one run per arm at
temperature 1.0, so the individual figures carry sampling noise. The direction and the per-position structure are well outside it.Notes
Not a duplicate: this targets the
tiezhen/new-model-supportbranch behind vllm-project#51655 and only adjusts how that branch reaches non-causal drafting for Muse Glimmer.AI assistance (Claude Code) was used to produce this change and the measurements above.