Skip to content

[Bugfix] Stop vocoder stages from declaring SupportsPP without make_empty_intermediate_tensors - #6860

Open
rk9595 wants to merge 4 commits into
vllm-project:mainfrom
rk9595:fix/vocoder-supportspp
Open

rk9595 wants to merge 4 commits into
vllm-project:mainfrom
rk9595:fix/vocoder-supportspp

Conversation

@rk9595

@rk9595 rk9595 commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

Closes #6859

Follow-up to #6803, which fixed the MiMo-Audio talker. This is the remaining half of the same sweep.

What broke

MiMoAudioToken2WavForConditionalGenerationVLLM (mimo_audio_code2wav.py) and CovoAudioCode2WavForConditionalGeneration (covo_audio_code2wav.py) declare SupportsPP but never provide make_empty_intermediate_tensors.

vLLM 0.28 turned that member from a method on the SupportsPP Protocol class into a bare annotation, so declaring the interface no longer supplies one. Verified on main @ fe4a2a74 with vLLM 0.28.0:

from vllm.model_executor.models.interfaces import supports_pp
supports_pp(MiMoAudioToken2WavForConditionalGenerationVLLM)  # True
supports_pp(CovoAudioCode2WavForConditionalGeneration)       # True
hasattr(cls, "make_empty_intermediate_tensors")              # False for both

supports_pp() is _supports_pp_attributes(model) and _supports_pp_inspect(model). The first is True purely because SupportsPP sits in the MRO. The second is True only incidentally — both forward methods take **kwargs, and supports_kw counts a VAR_KEYWORD parameter as accepting intermediate_tensors, which neither forward actually uses.

The flag is recorded as _ModelInfo.supports_pp (registry.py:858) and surfaced as ModelConfig.is_pp_supported, so both stages are advertised as pipeline-parallel capable. Under PP>1 a non-first rank calls self.model.make_empty_intermediate_tensors(...) in the runner's profiling path and raises AttributeError — the same failure as #6790.

Fix

Drop SupportsPP from both, rather than inventing a PP implementation they do not have:

  • Neither has an inner LM to delegate to — both are vocoder stages, unlike the talker in [Bugfix][MiMo-Audio] Restore make_empty_intermediate_tensors on the talker for vLLM 0.28 #6803 which delegates to Qwen2ForCausalLM.
  • Neither forward accepts or uses intermediate_tensors.
  • mypy already flagged this. Before this PR it reported Signature of "forward" incompatible with supertype "SupportsPP" on both classes; those two errors disappear with the declaration removed (25 → 23 errors on these two files, every remaining one pre-existing and unrelated).
  • Nothing in vllm_omni reads supports_pp or SupportsPP outside the model classes themselves, so no in-repo behaviour depends on the claim.

Regression test

tests/model_executor/models/test_supports_pp_contract.py walks every vllm_omni class declaring SupportsPP and asserts it either defines make_empty_intermediate_tensors or assigns it in __init__.

The check is static (AST) for two reasons: these models pull in torch and vLLM layers and build tensor-parallel linears, so constructing one in a unit test is not viable; and the repo-wide convention — glm_tts.py:828, fish_speech_slow_ar.py:225, voxcpm2_talker.py:852, and mimo_audio_llm.py after #6803 — is to assign on the instance inside __init__, mirroring upstream Qwen2ForCausalLM, which no class-level hasattr would see.

It guards the whole class of breakage, not just these two classes:

CPU-only, no model construction, runs in ~1s.

Verification

  • New test: fails on main listing both classes, passes here (checked both directions).
  • tests/model_executor/models/mimo_audio + the new test: 14 passed.
  • ruff check and ruff format --check clean on all three files.
  • mypy: two errors fixed, no new ones.

The SPDX headers in the diff were added by the repo's own check_spdx_header pre-commit hook when the files were touched; they are not manual edits.

@rk9595

rk9595 commented Aug 31, 2026

Copy link
Copy Markdown
Contributor Author

Self-review

What I checked

  • Confirmed the capability claim is live, not theoretical. Ran supports_pp() from vllm.model_executor.models.interfaces against both classes on main @ fe4a2a74: both return True while hasattr(cls, "make_empty_intermediate_tensors") is False. Traced why — _supports_pp_attributes is True from the MRO, and _supports_pp_inspect is True only because supports_kw accepts a **kwargs signature as taking intermediate_tensors.

  • Checked the blast radius before removing the declaration. grep for supports_pp / SupportsPP across vllm_omni/ and tests/ finds no reader outside the model classes themselves. Upstream consumes it via _ModelInfo.supports_pp → ModelConfig.is_pp_supported, which is a capability report, so removing an untrue claim narrows what vLLM advertises and changes nothing else.

  • Chose removal over implementation deliberately. The [Bugfix][MiMo-Audio] Restore make_empty_intermediate_tensors on the talker for vLLM 0.28 #6803 fix delegated to an inner Qwen2ForCausalLM; these two have no inner LM, and their forward methods neither accept nor use intermediate_tensors. Giving them a synthetic make_empty_intermediate_tensors would make the false claim harder to detect rather than fixing it.

  • mypy corroborates. Both classes previously produced Signature of "forward" incompatible with supertype "SupportsPP". Those two errors are gone; I diffed full mypy output before and after (25 → 23 on these files) and every remaining error is pre-existing and untouched, with only line numbers shifted by the SPDX header.

  • Tested the regression test in both directions, since a test that only passes proves nothing: it fails on main naming both classes, passes on this branch, and — after temporarily reverting the [Bugfix][MiMo-Audio] Restore make_empty_intermediate_tensors on the talker for vLLM 0.28 #6803 one-liner — fails on MiMoAudioLLMForConditionalGeneration, so it genuinely guards the original bug too.

  • Local gates: ruff check and ruff format --check clean; tests/model_executor/models/mimo_audio plus the new test give 14 passed in ~1s, CPU only.

Open question for reviewers: if you would rather keep SupportsPP on the vocoder stages and give them a real make_empty_intermediate_tensors, say so and I will switch — the contract test is written to accept either resolution, so it stands on its own regardless of which way you prefer.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/model_integration.md.

Module owners: @tzhouam @gcanlin

@rk9595, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Aug 31, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 353c43deba5d produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@hsliuustc0106 hsliuustc0106 added bug Something isn't working tts code related to tts models labels Sep 1, 2026
@rk9595

rk9595 commented Sep 2, 2026

Copy link
Copy Markdown
Contributor Author

@tzhouam @gcanlin — gentle ping on this one when you have a moment.

It's the remaining half of the sweep started in #6803 (merged): same SupportsPP-without-make_empty_intermediate_tensors defect, applied to the two vocoder stages. Self-review is above, and all checks are green on 1e56ffe (build 3.11/3.12, pre-commit, DCO, docs).

Happy to rebase or split it differently if you'd prefer a different shape.

for target in child.targets:
if isinstance(target, ast.Attribute) and target.attr == ATTR:
return True
if isinstance(child, ast.AnnAssign):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reject bare annotations; they still leave the SupportsPP attribute missing at runtime.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch — fixed in f1142e33.

You're right that the AnnAssign branch was wrong in exactly the way this PR exists to guard against: a bare annotation is precisely how vLLM 0.28 lost make_empty_intermediate_tensors, so accepting one would have let the checker green-light the original bug. AnnAssign now only counts when child.value is not None.

While there I also made the Assign branch match a Name target (make_empty_intermediate_tensors = staticmethod(...) at class level), which it previously missed — the two branches now agree on what counts as a binding. Happy to drop that half if you'd rather keep the diff to the one thing you flagged.

Verified in both directions. Added parametrized tests over _provides_attr itself: 5 shapes that bind at runtime (method, self.x = ..., self.x: T = ..., class-level assign, class-level annotated assign) and 3 that don't (bare class annotation, bare self.x: T, unrelated method). Plus an end-to-end check — rewriting qwen2_5_omni_token2wav.py:1447 from self.make_empty_intermediate_tensors = _empty_intermediate_tensors to self.make_empty_intermediate_tensors: object now fails the scan naming that class; before this commit it passed. Reverted, of course.

No class in vllm_omni/model_executor/models/ uses a bare annotation today, so the tightened check is green on main. 9 passed, ~1s, CPU only; ruff check / ruff format --check clean on 0.14.10 (the pinned pre-commit rev).

Also rebased onto current main (ab561485) — the branch was still sitting on fe4a2a74. mimo_audio_code2wav.py still needs its SPDX header; covo_audio_code2wav.py picked one up upstream, so that hunk shrank.

@lishunyang12 could you add the ready label? No Buildkite lane has run on this PR yet, and I'd like CI on it before it merges.

… attribute

MiMoAudioToken2WavForConditionalGenerationVLLM and
CovoAudioCode2WavForConditionalGeneration declare SupportsPP but never provide
make_empty_intermediate_tensors. Since vLLM 0.28 turned that member from a
method on the Protocol class into a bare annotation, nothing supplies it, so
vllm.model_executor.models.interfaces.supports_pp() reports True for both while
any read of the attribute raises AttributeError -- the failure fixed for the
talker in vllm-project#6803.

Neither stage implements pipeline parallelism: both are vocoders with no inner
LM to delegate to, and their forward methods ignore intermediate_tensors (mypy
already flagged both signatures as incompatible with the supertype). Drop the
declaration rather than inventing a PP implementation they do not have.

Add a static contract test over every vllm_omni class declaring SupportsPP,
asserting it defines the attribute or assigns it in __init__. The check is AST
based because models are too heavy to construct in a unit test and the repo
convention is to assign on the instance, which no class-level hasattr sees.
Reverting the vllm-project#6803 fix makes this test fail on that class.

Closes vllm-project#6859

Signed-off-by: rk9595 <rakesh.kariya@somaiya.edu>
A bare annotation binds nothing at runtime, which is exactly how
make_empty_intermediate_tensors went missing in vLLM 0.28, so only
count an AnnAssign that carries a value. Also count class-level
assignment, which the Assign branch previously missed.

Adds bidirectional tests over the checker itself.

Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
@rk9595
rk9595 force-pushed the fix/vocoder-supportspp branch from 1e56ffe to f1142e3 Compare September 16, 2026 05:16
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 7 days

@rk9595 this pull request has had no human commit, comment or review since 2026-09-16. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

Bring the branch up to date with main (vLLM 0.30.0 rebase, CoVo-Audio prompt fix). No conflicts; the contract test still passes and finds no new SupportsPP violators.

Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
@rk9595

rk9595 commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

Current plan and status.

The open review thread from @lishunyang12 is resolved — bare SupportsPP annotations are now rejected, pushed in f1142e33 on 2026-09-16 with a reply on the thread. Nothing is outstanding on my side.

Since the branch had fallen 172 commits behind (including the vLLM 0.30.0 rebase in #7820 and the CoVo-Audio fix in #7909), I have just merged current main in as 80f3851f. The merge was conflict-free, both vocoder changes are intact, and the contract test still passes at 9/9 against the updated tree — importantly it finds no new SupportsPP violators among the models added upstream in the meantime.

The change is unchanged in scope: 3 files, +140/-4. Two vocoder stages stop declaring SupportsPP (which they never satisfied), plus a static AST test that keeps the regression from returning.

Two things are needed to move this forward, neither of which I can do myself:

  1. A re-review from @lishunyang12 on the resolved thread.
  2. The ready label, so the full test suite runs — only pre-commit, build (3.11), build (3.12) and DCO run without it, and they are green. @Gaohan123 could you add it when you get a chance?

Happy to rebase again if it drifts; I would just rather not keep re-merging while it waits, since each update restarts the clock.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 7 days

@rk9595 this pull request has had no human commit, comment or review since 2026-09-23. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

Bring the branch up to date with main. Conflict-free; the contract test still passes 9/9 and finds no new SupportsPP violators among models added upstream.

Signed-off-by: Rakesh Kariya <rakesh.kariya@somaiya.edu>
@rk9595

rk9595 commented Oct 2, 2026

Copy link
Copy Markdown
Contributor Author

Plan unchanged, and the PR is current again.

The branch had fallen 139 commits behind, so I have merged main in as 353c43de. Conflict-free, scope still 3 files / +140-4, and the contract test passes 9/9 — it finds no new SupportsPP violators among the models added upstream since 2026-09-23. Neither vocoder file has been touched upstream.

Review state: the one blocker @lishunyang12 raised was resolved on 2026-09-16 (f1142e33 — bare annotations are now rejected) and answered on the thread. There is nothing open on my side.

This is now waiting on two maintainer actions, neither of which I can perform:

  1. The ready label. Without it only pre-commit, build (3.11), build (3.12) and DCO run; all four are green. mergeable_state stays blocked until the gated suite runs, and essentially every recently merged PR here carries ready ([Misc] remove legacy full_duplex alias for auto_response #8024, [New Model] Ming-Image (inclusionAI/Ming-Image-0.1-Design) Support #8021, [Bugfix] Fix Qwen3-TTS prefill probe on MRv2 runner #8065, [BugFix][Diffusion] Scope Wan RMSNorm patch to NPU #8047, [optimization] Added Wan 2.2 VAE decode optimizations #7056).
  2. A re-review on the resolved thread.

@Gaohan123 @hsliuustc0106 could one of you add ready? Tagging the code owners for the touched paths as well: @tzhouam @gcanlin for vllm_omni/model_executor/models/, @yenuo26 @NickCao for tests/.

If the direction here is no longer wanted, I would rather hear that and close it than keep it on the stale-bot rotation — this is the second 7-day nudge. Either outcome is fine; I would just like a decision.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on zcode (GLM-5.3-Flash) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 1f8a1753-d749-4bdd-8d33-65e8ea0fdc54) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 120s (try ).

1 similar comment
@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 1f8a1753-d749-4bdd-8d33-65e8ea0fdc54) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 120s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 03c8bcb8-0e7d-46db-bb32-2e38d1d7d000) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 600s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: da4199db-f6f9-45e8-b936-98a8c0ef249f) — check zcode login and the CLI version; falling back to direct/cursor/auto).

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

Changes since the previous review

  • 0 new inline finding(s); 0 finding(s) below.

CI at 353c43deba5d (2026-10-10T03:28:05.600746+00:00): required check(s) blocking: buildkite/vllm-omni (missing).

Note: The assigned review arm strict/zcode/GLM-5.3-Flash could not complete this review, so it was produced by the fallback arm direct/cursor/auto. It is excluded from the routing experiment.

Full review analysis

PR description

The MiMo-Audio token-to-wav stage and the CoVo-Audio code2wav stage no longer subclass vLLM's SupportsPP. Their forward methods do not read pipeline intermediate tensors, and neither class defines make_empty_intermediate_tensors, so the base class was advertising pipeline-parallel support they do not implement. A new CPU test parses model classes that still list SupportsPP and fails unless that attribute is actually bound, including rejecting a bare annotation with no value.

Change flow

flowchart LR
  vocoders["[CHANGED] MiMo and CoVo code2wav drop SupportsPP"]:::changed
  flag["[EXISTING] supports_pp to is_pp_supported"]:::existing
  profile["[CHANGED] Non-first PP rank no longer expects the missing method"]:::changed
  scan["[NEW] AST SupportsPP contract test"]:::new
  vocoders --> flag
  flag --> profile
  scan --> vocoders
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working tts code related to tts models

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Vocoder stages declare SupportsPP without make_empty_intermediate_tensors (PP>1 would hit the #6790 failure)

4 participants