Skip to content

[Bugfix][Model] Fix CoVo-Audio prompt processing and dummy loading - #7909

Merged
linyueqian merged 3 commits into
vllm-project:mainfrom
jeffaa729:bugfix/covo-audio-prompt-and-dummy-loading
Sep 22, 2026
Merged

linyueqian merged 3 commits into
vllm-project:mainfrom
jeffaa729:bugfix/covo-audio-prompt-and-dummy-loading

Conversation

@jeffaa729

Copy link
Copy Markdown
Contributor

Purpose

Fixes #7619.
Fixes #7906.

This also unblocks the CoVo-Audio baseline validation for #7452.

CoVo-Audio could fail in three consecutive startup and input-processing paths:

  1. Token2WavDecoder replaced PyTorch's standard state_dict() and load_state_dict() methods with incompatible signatures. During a core_model run, vLLM's dummy loader recursively called state_dict(destination=..., prefix=..., keep_vars=...), causing:

    TypeError: Token2WavDecoder.state_dict() got an unexpected keyword argument 'destination'
    

    This change removes those overrides and restores the standard nn.Module state-dict contract. The partial vendor checkpoint is loaded with strict=False, preserving the previous behavior of loading the available wavegan and token2latent weights while allowing model-created state that is absent from the checkpoint.

  2. The CoVo audio PromptReplacement used the string "<|cAUDIO|>" as its target. The pinned vLLM prompt-update interface expects a sequence of integer token IDs, so text matching attempted to decode a string as a token-ID vector and raised:

    TypeError: argument 'ids': Can't extract `str` to `Vec`
    

    The replacement target now uses [audio_token_id].

  3. CovoAudioDummyInputsBuilder constructed ProcessorInputs with an untokenized string prompt. Profiling expects prompt to be a list of integer token IDs, which caused integer token matching against a string:

    TypeError: must be str, not int
    

    The custom override is removed so the base dummy-input builder tokenizes the text before constructing ProcessorInputs.

Test Plan

vLLM Version: 0.29.0

vLLM-Omni Commit: e41a670a3

Hardware: 1× NVIDIA H100 80 GB

Python: 3.12.3

pytest -s -v tests/e2e/offline_inference/test_covo_audio_expansion.py::test_audio_to_audio --run-level core_model

pytest -s -v tests/e2e/offline_inference/test_covo_audio_expansion.py::test_audio_to_audio --run-level full_model

pytest -s -v tests/e2e/online_serving/test_covo_audio_expansion.py::test_audio_to_audio_001 --run-level core_model

pytest -s -v tests/e2e/online_serving/test_covo_audio_expansion.py::test_audio_to_audio_001 --run-level full_model

pre-commit run --files \
  vllm_omni/model_executor/models/covo_audio/covo_audio.py \
  vllm_omni/model_executor/models/covo_audio/covo_audio_code2wav.py \
  vllm_omni/model_executor/models/covo_audio/token2wav.py

Test Result

Serving path Run level Result
Offline core_model Passed
Offline full_model Passed
Online core_model Passed
Online full_model Passed

All applicable pre-commit hooks and git diff --check passed.

AI Assistance

AI assistance: Used Codex to analyze tracebacks, develop and review the fixes, and draft this PR description. I reviewed the changes and ran the reported validation locally.

Signed-off-by: jeffaa729 <hoiwanglo@gmail.com>
Signed-off-by: jeffaa729 <hoiwanglo@gmail.com>
Signed-off-by: jeffaa729 <hoiwanglo@gmail.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/model_integration.md, docs/design/module/ar_runtime.md.

Module owners: @gcanlin @tzhouam @Sy0307

Routing: @gcanlin via module of the changed files, semantic router, CODEOWNERS; @tzhouam via module of the changed files, semantic router, CODEOWNERS; @Sy0307 via module of the changed files, CODEOWNERS

@jeffaa729, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@jeffaa729

Copy link
Copy Markdown
Contributor Author

@vllm-omni-review-bot

@jeffaa729

Copy link
Copy Markdown
Contributor Author

Self-review:

  • removing the custom Token2WavDecoder.state_dict() and load_state_dict() restores the standard PyTorch module contract, while strict=False preserves the previous partial-checkpoint loading behavior.
  • PromptReplacement now uses integer token IDs and that the base dummy-input builder tokenizes the CoVo prompt before profiling.
  • Ran the existing CoVo offline and online E2E tests with both core_model and full_model on an NVIDIA H100 80 GB; all four configurations passed.
  • Confirmed that the changes are limited to CoVo-Audio internals and introduce no user-facing API or configuration changes.

@hsliuustc0106 hsliuustc0106 added bug Something isn't working tts code related to tts models labels Sep 22, 2026
@linyueqian linyueqian added the ready label to trigger buildkite CI label Sep 22, 2026

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All three fixes are correct and minimal. Deleting the state_dict and load_state_dict overrides restores the nn.Module contract that vLLM's dummy loader relies on (destination, prefix, keep_vars), and strict=False on the whole decoder reproduces what the old per-submodule loader did for the wavegan. and token2latent. prefixes. The PromptReplacement target becomes [audio_token_id], which the pinned vLLM prompt-update API requires, and audio_token_id is resolved once from the tokenizer vocabulary in the same function. Dropping the custom get_dummy_processor_inputs lets the base builder tokenize the dummy text, which is what the untokenized ProcessorInputs was tripping over. One inline suggestion on keeping the partial-load visible.

Static read at e41a670a; fork head, no PR code executed. pre-commit and DCO are green; no ready label, so a general lane is still needed before merge.

weights_only=True,
)
self.decoder.load_state_dict(ckpt)
self.decoder.load_state_dict(ckpt, strict=False)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] strict=False on the whole decoder now also hides a genuinely missing wavegan. or token2latent. weight, which the old loader at least confined to those two submodules. Capturing the IncompatibleKeys result and logging (or asserting empty) the missing_keys restricted to those two prefixes keeps the partial-checkpoint behaviour without losing the signal when a vendor checkpoint changes shape.

@linyueqian
linyueqian merged commit 5152b9b into vllm-project:main Sep 22, 2026
7 of 9 checks passed
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
…llm-project#7909)

Signed-off-by: jeffaa729 <hoiwanglo@gmail.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working ready label to trigger buildkite CI tts code related to tts models

Projects

None yet

4 participants