Repository navigation
[Bugfix][MiMo-Audio] Restore make_empty_intermediate_tensors on the talker for vLLM 0.28 - #6803
Conversation
Self-reviewWhat I checked
Note for anyone reproducing: the one-line repro in #6790 cannot reach the crash. Not included: a unit-level regression test. |
|
This PR appears to belong to: docs/design/module/model_integration.md. Module owners: @tzhouam @gcanlin @rk9595, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
…alker for vLLM 0.28 vLLM 0.27 declared SupportsPP.make_empty_intermediate_tensors as a method on the Protocol class, so subclasses using SupportsPP as a base inherited a stub and the attribute always resolved. vLLM 0.28 changed it to a bare annotation, which creates no class attribute. MiMoAudioLLMForConditionalGeneration declares SupportsPP but never assigned the attribute, and MiMoAudioForConditionalGeneration.__init__ reads it off that inner model (mimo_audio.py:586), so loading MiMo-Audio now fails with AttributeError during initialize_model. Delegate to the inner Qwen2ForCausalLM, matching glm_tts, fish_speech_slow_ar and voxcpm2_talker. Closes vllm-project#6790 Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
fdb9c86 to
6046dcf
Compare
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
…alker for vLLM 0.28 (vllm-project#6803) Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
… attribute MiMoAudioToken2WavForConditionalGenerationVLLM and CovoAudioCode2WavForConditionalGeneration declare SupportsPP but never provide make_empty_intermediate_tensors. Since vLLM 0.28 turned that member from a method on the Protocol class into a bare annotation, nothing supplies it, so vllm.model_executor.models.interfaces.supports_pp() reports True for both while any read of the attribute raises AttributeError -- the failure fixed for the talker in vllm-project#6803. Neither stage implements pipeline parallelism: both are vocoders with no inner LM to delegate to, and their forward methods ignore intermediate_tensors (mypy already flagged both signatures as incompatible with the supertype). Drop the declaration rather than inventing a PP implementation they do not have. Add a static contract test over every vllm_omni class declaring SupportsPP, asserting it defines the attribute or assigns it in __init__. The check is AST based because models are too heavy to construct in a unit test and the repo convention is to assign on the instance, which no class-level hasattr sees. Reverting the vllm-project#6803 fix makes this test fail on that class. Closes vllm-project#6859 Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
…alker for vLLM 0.28 (vllm-project#6803) Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Closes #6790
What broke
After the vLLM 0.28.0 rebase (#6606),
XiaomiMiMo/MiMo-Audio-7B-Instructno longer loads — engine-core startup fails while loading the stage-0 model:The same model loads and runs on vLLM 0.27.0, the pre-rebase pin.
Repro
(The one-line repro in the issue stops earlier, at that tokenizer check, before reaching the crash.)
Environment: vllm-omni
main@8be601ed, vLLM 0.28.0, RTX A6000 48 GB, driver 580.173.02,dtype=torch.bfloat16,enforce_eager=True.Root cause
SupportsPPchanged shape between the two pins:SupportsPPis used as a base class here, so on 0.27 every subclass inherited that stub and the attribute always resolved even when the subclass never assigned it. On 0.28 there is nothing to inherit.MiMoAudioLLMForConditionalGeneration(mimo_audio_llm.py:489) declaresSupportsPPand — pergit log -S— has never assignedmake_empty_intermediate_tensors. The outerMiMoAudioForConditionalGeneration.__init__reads it off that inner model:which is why the
AttributeErrornames the inner class but is raised inside the outer model's__init__, underinitialize_model→load_model.Fix
Delegate to the inner LM, which is a
Qwen2ForCausalLMand sets the attribute in its own__init__(itself by delegation). This matches the pattern the other omni talkers already use —glm_tts.py:828,fish_speech_slow_ar.py:225,voxcpm2_talker.py:852.The SPDX header in the diff was added by the repo's own
check_spdx_headerpre-commit hook when the file was touched; it is not a manual change.Scope
Two other
SupportsPPclasses never assign the attribute either —MiMoAudioToken2WavForConditionalGenerationVLLM(mimo_audio_code2wav.py:428) andCovoAudioCode2WavForConditionalGeneration(covo_audio_code2wav.py:18). Neither is reachable the same way: the only other consumer is the model runner's profiling path, guarded bynot get_pp_group().is_first_rank, so they would surface only under PP > 1, and neither has an inner LM to delegate to. Left out of this PR deliberately — happy to file a separate issue.Verification
Both runs on one RTX A6000 48 GB box, same command, same cached weights.
Clean
main@8be601ed— the traceback confirms the read site and that the lookup falls through tonn.Module.__getattr__:With this patch: no
AttributeError, both stages initialize, engine reaches ready and shuts down cleanly (LOAD OK).No unit test: the class cannot be constructed without weights (its
__init__builds tensor-parallel linears and two transformer stacks), so a load-level check is the meaningful regression signal. Happy to add a construction-mocking test if you'd prefer one.