Skip to content

[Bugfix][MiMo-Audio] Restore make_empty_intermediate_tensors on the talker for vLLM 0.28 - #6803

Merged
hsliuustc0106 merged 1 commit into
vllm-project:mainfrom
rk9595:fix/mimo-audio-supportspp-0.28
Aug 31, 2026
Merged

hsliuustc0106 merged 1 commit into
vllm-project:mainfrom
rk9595:fix/mimo-audio-supportspp-0.28

Conversation

@rk9595

@rk9595 rk9595 commented Aug 30, 2026

Copy link
Copy Markdown
Contributor

Closes #6790

What broke

After the vLLM 0.28.0 rebase (#6606), XiaomiMiMo/MiMo-Audio-7B-Instruct no longer loads — engine-core startup fails while loading the stage-0 model:

File ".../vllm_omni/worker/gpu_model_runner.py", line 180, in load_model
File ".../vllm/v1/worker/gpu_model_runner.py", line 5435, in load_model
  self.model = model_loader.load_model(
AttributeError: 'MiMoAudioLLMForConditionalGeneration' object has no attribute 'make_empty_intermediate_tensors'

The same model loads and runs on vLLM 0.27.0, the pre-rebase pin.

Repro

# MIMO_AUDIO_TOKENIZER_PATH is mandatory and is checked with os.path.exists()
# (mimo_audio.py:291), so it has to be a local directory — an HF repo id fails.
hf download XiaomiMiMo/MiMo-Audio-Tokenizer --local-dir /root/mimo-audio-tokenizer
export MIMO_AUDIO_TOKENIZER_PATH=/root/mimo-audio-tokenizer

python -c "from vllm_omni import Omni; Omni(model='XiaomiMiMo/MiMo-Audio-7B-Instruct', deploy_config='vllm_omni/deploy/mimo_audio.yaml')"

(The one-line repro in the issue stops earlier, at that tokenizer check, before reaching the crash.)

Environment: vllm-omni main @ 8be601ed, vLLM 0.28.0, RTX A6000 48 GB, driver 580.173.02, dtype=torch.bfloat16, enforce_eager=True.

Root cause

SupportsPP changed shape between the two pins:

# vllm 0.27.0, model_executor/models/interfaces.py:655 — a real method on the Protocol class
class SupportsPP(Protocol):
    def make_empty_intermediate_tensors(self, batch_size, dtype, device) -> "IntermediateTensors":
        ...

# vllm 0.28.0 — a bare annotation, which creates no class attribute
class SupportsPP(Protocol):
    make_empty_intermediate_tensors: _MakeEmptyIntermediateTensors

SupportsPP is used as a base class here, so on 0.27 every subclass inherited that stub and the attribute always resolved even when the subclass never assigned it. On 0.28 there is nothing to inherit.

MiMoAudioLLMForConditionalGeneration (mimo_audio_llm.py:489) declares SupportsPP and — per git log -S — has never assigned make_empty_intermediate_tensors. The outer MiMoAudioForConditionalGeneration.__init__ reads it off that inner model:

# mimo_audio.py:585-589 — self.fused_thinker_talker is built at :560 via
# init_vllm_registered_model(architectures=["MiMoAudioLLMModel"]), which
# registry.py:239 maps to MiMoAudioLLMForConditionalGeneration
self.make_empty_intermediate_tensors = (
    (self.fused_thinker_talker.make_empty_intermediate_tensors)   # :586, raises here
    if self.model_stage == "fused_thinker_talker"
    else lambda: None
)

which is why the AttributeError names the inner class but is raised inside the outer model's __init__, under initialize_model → load_model.

Fix

Delegate to the inner LM, which is a Qwen2ForCausalLM and sets the attribute in its own __init__ (itself by delegation). This matches the pattern the other omni talkers already use — glm_tts.py:828, fish_speech_slow_ar.py:225, voxcpm2_talker.py:852.

The SPDX header in the diff was added by the repo's own check_spdx_header pre-commit hook when the file was touched; it is not a manual change.

Scope

Two other SupportsPP classes never assign the attribute either — MiMoAudioToken2WavForConditionalGenerationVLLM (mimo_audio_code2wav.py:428) and CovoAudioCode2WavForConditionalGeneration (covo_audio_code2wav.py:18). Neither is reachable the same way: the only other consumer is the model runner's profiling path, guarded by not get_pp_group().is_first_rank, so they would surface only under PP > 1, and neither has an inner LM to delegate to. Left out of this PR deliberately — happy to file a separate issue.

Verification

Both runs on one RTX A6000 48 GB box, same command, same cached weights.

Clean main @ 8be601ed — the traceback confirms the read site and that the lookup falls through to nn.Module.__getattr__:

File ".../vllm/model_executor/model_loader/utils.py", line 58, in initialize_model
  model = model_class(vllm_config=vllm_config, prefix=prefix)
File ".../vllm_omni/model_executor/models/mimo_audio/mimo_audio.py", line 586, in __init__
  (self.fused_thinker_talker.make_empty_intermediate_tensors)
File ".../torch/nn/modules/module.py", line 1967, in __getattr__
  raise AttributeError(
AttributeError: 'MiMoAudioLLMForConditionalGeneration' object has no attribute 'make_empty_intermediate_tensors'

With this patch: no AttributeError, both stages initialize, engine reaches ready and shuts down cleanly (LOAD OK).

No unit test: the class cannot be constructed without weights (its __init__ builds tensor-parallel linears and two transformer stacks), so a load-level check is the meaningful regression signal. Happy to add a construction-mocking test if you'd prefer one.

@rk9595

rk9595 commented Aug 30, 2026

Copy link
Copy Markdown
Contributor Author

Self-review

What I checked

  • Root cause is the vLLM 0.28 SupportsPP change, not the mimo code. Diffed model_executor/models/interfaces.py between the v0.27.0 and v0.28.0 tags: make_empty_intermediate_tensors went from a method on the Protocol class (inherited by every subclass) to a bare annotation (inherited by none). git log -S confirms MiMoAudioLLMForConditionalGeneration never assigned it — it only ever worked through that inherited stub.

  • The failing read is in the outer model, not the inner one. The runtime traceback ends in torch/nn/modules/module.py:1967 __getattr__, reached from mimo_audio.py:586, under initialize_model. So the fix belongs on the inner class that must expose the attribute, which is what this PR does.

  • The delegation target is valid on 0.28. Qwen2ForCausalLM.__init__ sets make_empty_intermediate_tensors as an instance attribute (itself delegating to its inner model), so self.model.make_empty_intermediate_tensors resolves at the point this PR reads it — after init_vllm_registered_model returns.

  • Verified both directions on real hardware (RTX A6000 48 GB, driver 580.173.02, vLLM 0.28.0, main @ 8be601ed): unpatched fails with the reported AttributeError; patched reaches LOAD OK with both stages initialized and a clean shutdown. Same box, same command, same cached weights.

  • Swept the whole bug class. Every vllm_omni class inheriting SupportsPP was checked for whether it assigns the attribute. Two others do not — MiMoAudioToken2WavForConditionalGenerationVLLM and CovoAudioCode2WavForConditionalGeneration — but neither is reachable this way (the only other consumer is the runner's profiling path, guarded by not get_pp_group().is_first_rank) and neither has an inner LM to delegate to. Deliberately out of scope here; happy to file separately.

  • Local gates: ruff check and ruff format --check pass on the changed file; mypy output is unchanged from clean main. The SPDX header in the diff was inserted by the repo's own check_spdx_header pre-commit hook when the file was touched.

Note for anyone reproducing: the one-line repro in #6790 cannot reach the crash. MIMO_AUDIO_TOKENIZER_PATH is mandatory and is checked with os.path.exists (mimo_audio.py:291), so it must point at a downloaded directory — an HF repo id fails with Audio tokenizer not exists. The PR description has the working sequence.

Not included: a unit-level regression test. MiMoAudioLLMForConditionalGeneration.__init__ builds tensor-parallel linears and two transformer stacks, so constructing it without weights needs heavy mocking for a one-line assertion. If you'd like coverage, I'd suggest a contract test over the model registry asserting that every SupportsPP class assigns make_empty_intermediate_tensors — that guards the whole class of breakage, including the two latent cases above, and runs on CPU. Glad to add it in this PR if you prefer.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/model_integration.md.

Module owners: @tzhouam @gcanlin

@rk9595, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@rk9595

rk9595 commented Aug 30, 2026

Copy link
Copy Markdown
Contributor Author

@vllm-omni-review-bot

…alker for vLLM 0.28

vLLM 0.27 declared SupportsPP.make_empty_intermediate_tensors as a method on
the Protocol class, so subclasses using SupportsPP as a base inherited a stub
and the attribute always resolved. vLLM 0.28 changed it to a bare annotation,
which creates no class attribute.

MiMoAudioLLMForConditionalGeneration declares SupportsPP but never assigned
the attribute, and MiMoAudioForConditionalGeneration.__init__ reads it off
that inner model (mimo_audio.py:586), so loading MiMo-Audio now fails with
AttributeError during initialize_model.

Delegate to the inner Qwen2ForCausalLM, matching glm_tts, fish_speech_slow_ar
and voxcpm2_talker.

Closes vllm-project#6790

Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
@rk9595
rk9595 force-pushed the fix/mimo-audio-supportspp-0.28 branch from fdb9c86 to 6046dcf Compare August 30, 2026 06:35
@linyueqian linyueqian added the ready label to trigger buildkite CI label Aug 31, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 6046dcfecaec produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Aug 31, 2026
@hsliuustc0106
hsliuustc0106 merged commit b864374 into vllm-project:main Aug 31, 2026
7 of 9 checks passed
JoseCarlosGarcia95 pushed a commit to valendra-tech/vllm-omni that referenced this pull request Sep 5, 2026
…alker for vLLM 0.28 (vllm-project#6803)

Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
rk9595 added a commit to rk9595/vllm-omni that referenced this pull request Sep 16, 2026
… attribute

MiMoAudioToken2WavForConditionalGenerationVLLM and
CovoAudioCode2WavForConditionalGeneration declare SupportsPP but never provide
make_empty_intermediate_tensors. Since vLLM 0.28 turned that member from a
method on the Protocol class into a bare annotation, nothing supplies it, so
vllm.model_executor.models.interfaces.supports_pp() reports True for both while
any read of the attribute raises AttributeError -- the failure fixed for the
talker in vllm-project#6803.

Neither stage implements pipeline parallelism: both are vocoders with no inner
LM to delegate to, and their forward methods ignore intermediate_tensors (mypy
already flagged both signatures as incompatible with the supertype). Drop the
declaration rather than inventing a PP implementation they do not have.

Add a static contract test over every vllm_omni class declaring SupportsPP,
asserting it defines the attribute or assigns it in __init__. The check is AST
based because models are too heavy to construct in a unit test and the repo
convention is to assign on the instance, which no class-level hasattr sees.
Reverting the vllm-project#6803 fix makes this test fail on that class.

Closes vllm-project#6859

Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…alker for vLLM 0.28 (vllm-project#6803)

Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: MiMo Audio fails to load on vLLM 0.28.0 (missing make_empty_intermediate_tensors)

4 participants