Repository navigation
[Model] Aura support - Non async chunk path - #4257
Conversation
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
ef67708 to
93449bb
Compare
Model Integration Review: Aura Omni PipelineVerdict: COMMENT with suggestions for improvement SummaryThis PR adds support for the AURA Omni pipeline, a 4-stage multimodal pipeline (ASR → AURA/Qwen3-VL → Qwen3-TTS Talker → Code2Wav). The implementation includes proper stage bridging logic, non-async chunk silent path handling, and dual TTS mode support (Base voice-clone + CustomVoice speaker-driven). ✅ Strengths
📋 PR Description Checklist ItemsThe PR description is missing the following required items (from the checklist in the template):
🔍 Code Quality Observations
🤔 Questions for Consideration
💡 Suggestions
ConclusionThe implementation is technically sound with good code quality and comprehensive unit tests. The main gap is the lack of test results in the PR description, which would help reviewers and users understand the pipeline's behavior and performance. Adding those would strengthen the PR significantly. Note: This is not a diffusion model PR—AURA is a Vision-Language model (Qwen3-VL) in a multimodal pipeline with ASR and TTS stages. Diffusion-specific requirements do not apply. |
|
A few initial question: in the test files, should we use English-only fixture strings (e.g. transcripts/ref text) for consistency with most unit tests, or keep Chinese since AURA defaults to Chinese responses? |
they default to Chinese response, but I will change all examples later with english only |
| resumable: bool = False, | ||
| ) -> Any: | ||
| next_pool = self.stage_pools[next_stage_id] | ||
| if self._next_stage_input_is_tokens(next_input): |
There was a problem hiding this comment.
I don't understand the logics in this branching, consider add some docs here
There was a problem hiding this comment.
For other models, after the first stage reads text, subsequent stages can directly use tokens as input. However, aura_omni first uses ASR to convert speech to text, and the input for the second stage is text, not tokens. I am checking if the vocabulary of qwen3-asr is consistent with that of qwen3-vl. If they are same, I may be able to directly use the sampled tokens as input for Aura. If so, I will delete this part modification.
There was a problem hiding this comment.
I tested passing tokens directly and found a precision issue. I would like to keep the current implementation in this PR.
| request.external_req_id = request.request_id | ||
| return request | ||
|
|
||
| processor = self._get_stage_input_processor(next_stage_id) |
There was a problem hiding this comment.
We placed the input processor in AsyncOmniEngine, why here still require one?
There was a problem hiding this comment.
Same explanation as above
| stage_submit_ts=submit_ts, | ||
| ) | ||
| ) | ||
| await self._cleanup_request_ids([req_id, *self._cfg_tracker.cleanup_parent(req_id)]) |
There was a problem hiding this comment.
this is the logic for AURA? If so, don't need to consider cfg tracker here
There was a problem hiding this comment.
function here is general for all similar models. Aura determines whether to respond based on the user's input. If the model does not respond (silent) in this round and the inference has ended, it will clear the request. If other similar models also support silent, this section can be reused.
| *, | ||
| final_output_type: str | None, | ||
| ) -> RequestOutput: | ||
| """Build a terminal empty output when no downstream stage input exists.""" |
There was a problem hiding this comment.
Does CompletionOutput have default value in its parameters? This seems just return a default CompletionOutput, could be replace by something like CompletionOutput()?
There was a problem hiding this comment.
vllm.output.CompletionOutput don't have default value
8979060 to
efd60d0
Compare
|
|
||
| return conversation, [engine_prompt] | ||
|
|
||
| def _is_aura_omni_pipeline(self) -> bool: |
There was a problem hiding this comment.
should this function be placed in aura_omni/pipeline.py ?
There was a problem hiding this comment.
This function should be kept in the serving_chat.py. The function name here is not quite accurate, I have changed it to a more generalized one. All cascaded models need to perform similar judgments and multimodal data defer operations.
| stage_names.add(str(model_stage)) | ||
| return {"asr", "aura", "code2wav"}.issubset(stage_names) | ||
|
|
||
| async def _strip_aura_videos_for_asr( |
There was a problem hiding this comment.
There are many other mm data exraction functions in serving chat, maybe we should make these exaction processes as a part of pipeline abstraction sometime later
There was a problem hiding this comment.
The orchestrator silent-path change reads correct and well-scoped: empty next_inputs while not-finished waits for more outputs; while finished, it emits a terminal empty output from the final stage and cleans up the request + CFG-companion ids. This only fires when a stage finishes producing no downstream input (which previously hung the request), so it's effectively a lifecycle fix with no regression for pipelines that always produce inputs.
Two things worth confirming before merge:
- Please confirm
_needs_multistage_multimodal_split()returnsFalsefor existing single-stage / Qwen3-Omni pipelines, so_preprocess_chatstays a no-op for them. - Nits:
sr=24000is hardcoded in_build_terminal_empty_output(could be sourced from the final stage client for non-24kHz vocoders); theadditional_informationinit block inserving_chat.pylooks duplicated.
No streaming-video here (deferred to the #4424-based follow-up), so this is mergeable as model integration.
|
ci failure is not related to this PR |
hsliuustc0106
left a comment
There was a problem hiding this comment.
Thanks for the AURA wiring. I found a couple of blocking issues in the user-visible paths: streaming can hand only a delta/empty string to TTS, and the default text,audio response shape can expose both the ASR transcript and AURA answer as final text choices. I also left a smaller docs/client mismatch around the served model name.
|
|
||
| def _extract_text(source_output: Any) -> str: | ||
| output = _extract_output(source_output) | ||
| text = getattr(output, "text", None) |
There was a problem hiding this comment.
In streaming chat completions, AR sampling params are coerced to DELTA output, and the output processor attaches the full generated text as cumulative_text on the final output. This helper returns output.text first, which can be empty or just the last delta, so aura2tts() can skip a non-silent AURA reply or synthesize only a suffix. Please prefer non-empty cumulative_text before text here, and add a regression using the existing _source_delta_final_output() test helper.
| model_stage="asr", | ||
| execution_type=StageExecutionType.LLM_AR, | ||
| input_sources=(), | ||
| final_output=True, |
There was a problem hiding this comment.
With both stage 0 and stage 1 marked as final_output=True and final_output_type="text", the default modalities=["text", "audio"] path will append choices from both final text stages. That means the OpenAI response can expose the ASR transcript as one text choice and the AURA answer as another, even though examples document AURA as the text response. Please make ASR internal-only for this pipeline, or explicitly filter stage-0 text from chat responses.
| `pipeline: aura_omni`, so the four-stage topology is used even if the | ||
| command-line `--model` points at one of the component checkpoints. | ||
|
|
||
| Send requests with `"model": "aura_omni"`. The ASR, AURA, and Qwen3-TTS |
There was a problem hiding this comment.
This request model name does not match the server command above, which sets --served-model-name aurateam/AURA. The curl script defaults to MODEL=aura_omni while the Python client uses aurateam/AURA, so following the docs/examples as written can produce an unknown-model error. Please make the served model name and clients consistent.
Add an aura_omni pipeline that composes Qwen3-ASR, AURA/Qwen3-VL, and the native Qwen3-TTS Talker -> Code2Wav stages. Include deploy config, ASR-to-AURA and AURA-to-TTS stage processors, docs, and tests for the new multi-stage topology. Signed-off-by: Rein Yang <ruiruyang2@gmail.com>
Adds offline, online, curl, and Gradio examples for the ASR -> AURA -> Qwen3-TTS pipeline, including docs and runnable launch scripts. Also adapts AURA’s Qwen3-VL stage to load its remote-config checkpoint safely and filters ASR audio before forwarding multimodal data to the VLM stage. Signed-off-by: Rein Yang <ruiruyang2@gmail.com>
…age bridging and tuning deploy config Complete native AURA Omni stage bridging, then update to disable prefix caching on AURA/Talker and switch Talker+Code2Wav to the CustomVoice checkpoint, reducing cross-turn carryover and improving multi-request TTS consistency. Signed-off-by: Rein Yang <ruiruyang2@gmail.com>
Always pass AURA assistant token ids to Qwen3-TTS via PRECOMPUTED_TEXT_IDS_KEY instead of re-tokenizing response text, teach Qwen3-TTS talker to accept precomputed text ids. Signed-off-by: Rein Yang <ruiruyang2@gmail.com>
…dio code length Signed-off-by: Rein Yang <ruiruyang2@gmail.com>
Handle empty next-stage inputs in orchestrator non-async forwarding: return a terminal empty output (text/audio by final_output_type) and clean up request state when upstream output is finished, and only debug-return when unfinished. This unblocks AURA silent responses where aura2tts emits no downstream TTS request. Signed-off-by: Rein Yang <ruiruyang2@gmail.com>
Signed-off-by: R2-Y <ruiruyang2@gmail.com>
Replace Aura-specific ASR video stripping with stage-aware deferred multimodal handoff for cascaded pipelines, and make Qwen3-TTS token passthrough opt-in so the default path sends text. Add regression coverage for deferred modalities and TTS passthrough behavior. Signed-off-by: R2-Y <ruiruyang2@gmail.com>
Signed-off-by: R2-Y <ruiruyang2@gmail.com>
2f18be1 to
a8fac61
Compare
Signed-off-by: R2-Y <ruiruyang2@gmail.com>
a8fac61 to
7618ce0
Compare
|
|
@hsliuustc0106 Ready to merge |
Merging main pulled in new call sites that used the pre-rename output queue API, which this branch had already renamed output_async_queue -> output_sync_queue (put_nowait on the sync side). These were semantic conflicts: git auto-merged them without a textual conflict, so they landed referencing an attribute the constructor no longer sets. - orchestrator.py:1251 (from vllm-project#4079, diffusion request-level batching) and orchestrator.py:1419 (from vllm-project#4257, Aura non-async-chunk path) still called `await self.output_async_queue.put(...)`, which would raise AttributeError on those terminal-output error/edge paths. Convert to `self.output_sync_queue.put_nowait(...)`. - tests/engine/test_orchestrator_stage_input_bridge.py (new file from vllm-project#4257, marked core_model/cpu) constructed Orchestrator with the old `output_async_queue=` kwarg and failed against the new signature. Update to `output_sync_queue=output_q.sync_q`. Tested on vLLM 0.24.0: engine + orchestrator unit tests (41 passed) and a Qwen3-TTS-12Hz-0.6B streaming /v1/audio/speech smoke test. Signed-off-by: zeningc <zening.chen@yahoo.com>
Signed-off-by: Rein Yang <ruiruyang2@gmail.com> Signed-off-by: R2-Y <ruiruyang2@gmail.com>
Before submitting: run the precheck-pr skill with code agent for a self-check against project conventions — catches dead code, missing benchmarks, and title format issues.
PLEASE FILL IN THE PR DESCRIPTION HERE ENSURING ALL CHECKLIST ITEMS (AT THE BOTTOM) HAVE BEEN CONSIDERED.
Purpose
Topology: ASR -> AURA (Qwen3-VL) -> Qwen3-TTS Talker -> Qwen3-TTS Code2Wav
Adds pipeline/deploy wiring and stage processor integration for this topology.
Add stage bridging logic
asr2aura: converts ASR output + visual input into AURA-consumable prompt format.aura2tts: converts AURA output into TTS-stage input with task-specific metadata.Support two TTS operating modes
Uses PRECOMPUTED_TEXT_IDS_KEY for direct token passthrough from AURA to TTS.
Avoids re-tokenizing generated text and improves consistency with model-side conditioning.
Test Plan
Test Result
Online inference:
send single openai request with audio question & video:
Gradio:

Test on 1xH200 with command (TTS use base voice clone)
Here I used random-mm to provide a performance baseline, but the audio and video content are not strongly related in this datasets, aura is very likely to return silent.

We need to introduce video input + realted audio QA datasets in the future to obtain more reliable performance data.
Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model. Please runmkdocs serveto sync the documentation editions to./docs.BEFORE SUBMITTING, please read https://github.com/vllm-project/vllm-omni/blob/main/CONTRIBUTING.md
(anything written below this line will be removed by GitHub Actions)