Repository navigation
[Bugfix][Model] Fix CoVo-Audio prompt processing and dummy loading - #7909
linyueqian merged 3 commits into
Conversation
Signed-off-by: jeffaa729 <hoiwanglo@gmail.com>
Signed-off-by: jeffaa729 <hoiwanglo@gmail.com>
Signed-off-by: jeffaa729 <hoiwanglo@gmail.com>
|
This PR appears to belong to: docs/design/module/model_integration.md, docs/design/module/ar_runtime.md. Module owners: @gcanlin @tzhouam @Sy0307 Routing: @gcanlin via module of the changed files, semantic router, CODEOWNERS; @tzhouam via module of the changed files, semantic router, CODEOWNERS; @Sy0307 via module of the changed files, CODEOWNERS @jeffaa729, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
|
Self-review:
|
linyueqian
left a comment
There was a problem hiding this comment.
All three fixes are correct and minimal. Deleting the state_dict and load_state_dict overrides restores the nn.Module contract that vLLM's dummy loader relies on (destination, prefix, keep_vars), and strict=False on the whole decoder reproduces what the old per-submodule loader did for the wavegan. and token2latent. prefixes. The PromptReplacement target becomes [audio_token_id], which the pinned vLLM prompt-update API requires, and audio_token_id is resolved once from the tokenizer vocabulary in the same function. Dropping the custom get_dummy_processor_inputs lets the base builder tokenize the dummy text, which is what the untokenized ProcessorInputs was tripping over. One inline suggestion on keeping the partial-load visible.
Static read at e41a670a; fork head, no PR code executed. pre-commit and DCO are green; no ready label, so a general lane is still needed before merge.
| weights_only=True, | ||
| ) | ||
| self.decoder.load_state_dict(ckpt) | ||
| self.decoder.load_state_dict(ckpt, strict=False) |
There was a problem hiding this comment.
[suggestion] strict=False on the whole decoder now also hides a genuinely missing wavegan. or token2latent. weight, which the old loader at least confined to those two submodules. Capturing the IncompatibleKeys result and logging (or asserting empty) the missing_keys restricted to those two prefixes keeps the partial-checkpoint behaviour without losing the signal when a vendor checkpoint changes shape.
…llm-project#7909) Signed-off-by: jeffaa729 <hoiwanglo@gmail.com> Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
…llm-project#7909) Signed-off-by: jeffaa729 <hoiwanglo@gmail.com>
Purpose
Fixes #7619.
Fixes #7906.
This also unblocks the CoVo-Audio baseline validation for #7452.
CoVo-Audio could fail in three consecutive startup and input-processing paths:
Token2WavDecoderreplaced PyTorch's standardstate_dict()andload_state_dict()methods with incompatible signatures. During acore_modelrun, vLLM's dummy loader recursively calledstate_dict(destination=..., prefix=..., keep_vars=...), causing:This change removes those overrides and restores the standard
nn.Modulestate-dict contract. The partial vendor checkpoint is loaded withstrict=False, preserving the previous behavior of loading the availablewaveganandtoken2latentweights while allowing model-created state that is absent from the checkpoint.The CoVo audio
PromptReplacementused the string"<|cAUDIO|>"as its target. The pinned vLLM prompt-update interface expects a sequence of integer token IDs, so text matching attempted to decode a string as a token-ID vector and raised:The replacement target now uses
[audio_token_id].CovoAudioDummyInputsBuilderconstructedProcessorInputswith an untokenized string prompt. Profiling expectspromptto be a list of integer token IDs, which caused integer token matching against a string:The custom override is removed so the base dummy-input builder tokenizes the text before constructing
ProcessorInputs.Test Plan
vLLM Version:
0.29.0vLLM-Omni Commit:
e41a670a3Hardware: 1× NVIDIA H100 80 GB
Python:
3.12.3Test Result
core_modelfull_modelcore_modelfull_modelAll applicable pre-commit hooks and
git diff --checkpassed.AI Assistance
AI assistance: Used Codex to analyze tracebacks, develop and review the fixes, and draft this PR description. I reviewed the changes and ran the reported validation locally.