Skip to content

[Model] Aura support - Non async chunk path - #4257

Merged
hsliuustc0106 merged 11 commits into
vllm-project:mainfrom
R2-Y:aura_support
Jun 21, 2026
Merged

hsliuustc0106 merged 11 commits into
vllm-project:mainfrom
R2-Y:aura_support

Conversation

@R2-Y

@R2-Y R2-Y commented Jun 8, 2026 •

Copy link
Copy Markdown
Contributor

Before submitting: run the precheck-pr skill with code agent for a self-check against project conventions — catches dead code, missing benchmarks, and title format issues.

PLEASE FILL IN THE PR DESCRIPTION HERE ENSURING ALL CHECKLIST ITEMS (AT THE BOTTOM) HAVE BEEN CONSIDERED.

Purpose

  1. support 4-stage AURA pipeline

Topology: ASR -> AURA (Qwen3-VL) -> Qwen3-TTS Talker -> Qwen3-TTS Code2Wav
Adds pipeline/deploy wiring and stage processor integration for this topology.

  1. Add stage bridging logic
    asr2aura: converts ASR output + visual input into AURA-consumable prompt format.
    aura2tts: converts AURA output into TTS-stage input with task-specific metadata.

  2. Support two TTS operating modes

  • Base voice-clone mode (reference audio + reference text)
  • CustomVoice speaker-driven mode
  1. Add AURA-to-TTS token passthrough

Uses PRECOMPUTED_TEXT_IDS_KEY for direct token passthrough from AURA to TTS.
Avoids re-tokenizing generated text and improves consistency with model-side conditioning.

  1. Non-async chunk silent-path
  • Handles next_inputs=[] in orchestrator forwarding:
    • If finished: emit terminal empty output (text/audio by final_output_type) and cleanup request state.
    • If not finished: debug-return and wait for subsequent outputs.
  • Prevents silent responses from hanging the request lifecycle.
image

Test Plan

  1. Functional correctness test (offline, online, Gradio)
  2. Performance test

Test Result

  1. Functional correctness test
    Online inference:
    send single openai request with audio question & video:
image

Gradio:
image

  1. Performance test
    Test on 1xH200 with command (TTS use base voice clone)
  vllm bench serve \
    --omni \
    --host 127.0.0.1 \
    --port 8666 \
    --model /data/models/AURA \
    --backend openai-chat-omni \
    --endpoint /v1/chat/completions \
    --dataset-name random-mm \
    --random-input-len 64 \
    --random-output-len 32 \
    --random-range-ratio 0.0 \
    --ignore-eos \
    --num-prompts 64 \
    --num-warmups 1 \
    --max-concurrency ${c} \
    --request-rate inf \
    --random-mm-base-items-per-request 2 \
    --random-mm-num-mm-items-range-ratio 0.0 \
    --random-mm-limit-mm-per-prompt '{"audio":1,"video":1}' \
    --random-mm-bucket-config '{"(0, 3, 1)":0.5,"(224, 224, 2)":0.5}' \
    --percentile-metrics ttft,tpot,itl,e2el,audio_ttfp,audio_rtf,audio_duration \
    --save-result \
    --result-dir ./aura_bench \
    --result-filename "aura_c${c}.json"
metrics con1 con16
Completed / Failed 32 / 0 64 / 0
Total (s) 6.14 2.55
Request Throughput (RPS) 5.21 25.13
Output Token Throughput (tok/s) 333.65 1608.30
Total Token Throughput (tok/s) 615.17 2965.29
Max Output Tokens/s 187 904
TTFT mean (ms) 50.19 346.19
TPOT mean (ms) 1.68 2.85
ITL mean (ms) 3.31 5.61
E2EL mean (ms) 191.43 592.20

Here I used random-mm to provide a performance baseline, but the audio and video content are not strongly related in this datasets, aura is very likely to return silent.
image
We need to introduce video input + realted audio QA datasets in the future to obtain more reliable performance data.


Essential Elements of an Effective PR Description Checklist
  • The purpose of the PR, such as "Fix some issue (link existing issues this PR will resolve)".
  • The test plan. Please provide the test scripts & test commands. Please state the reasons if your codes don't require additional test scripts. For test file guidelines, please check the test style doc
  • The test results. Please paste the results comparison before and after, or the e2e results.
  • (Optional) The necessary documentation update, such as updating supported_models.md and examples for a new model. Please run mkdocs serve to sync the documentation editions to ./docs.
  • (Optional) Release notes update. If your change is user-facing, please update the release notes draft.

BEFORE SUBMITTING, please read https://github.com/vllm-project/vllm-omni/blob/main/CONTRIBUTING.md
(anything written below this line will be removed by GitHub Actions)

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@R2-Y R2-Y mentioned this pull request Jun 8, 2026
1 task done
@R2-Y R2-Y changed the title Aura support - Non async chunk path [Model] Aura support - Non async chunk path Jun 8, 2026
@amy-why-3459
amy-why-3459 force-pushed the aura_support branch 3 times, most recently from ef67708 to 93449bb Compare June 8, 2026 12:38
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

Model Integration Review: Aura Omni Pipeline

Verdict: COMMENT with suggestions for improvement

Summary

This PR adds support for the AURA Omni pipeline, a 4-stage multimodal pipeline (ASR → AURA/Qwen3-VL → Qwen3-TTS Talker → Code2Wav). The implementation includes proper stage bridging logic, non-async chunk silent path handling, and dual TTS mode support (Base voice-clone + CustomVoice speaker-driven).

✅ Strengths

  1. Good documentation: User-facing docs in docs/user_guide/examples/online_serving/aura_omni.md and example README
  2. Comprehensive tests: Unit tests for stage processors (asr2aura, aura2tts), orchestrator stage bridging, model architecture registration, and deploy config validation
  3. Proper pipeline integration: Correctly uses aura_omni pipeline with 4-stage topology and stage input processors
  4. Silent path handling: Implements _build_terminal_empty_output() for cases where AURA emits <|silent|>, preventing request lifecycle hangs
  5. Token passthrough: Uses PRECOMPUTED_TEXT_IDS_KEY to pass AURA-generated tokens directly to TTS without re-tokenization, improving consistency
  6. Dual TTS modes: Properly supports both Base voice-clone and CustomVoice speaker modes with appropriate parameter handling

📋 PR Description Checklist Items

The PR description is missing the following required items (from the checklist in the template):

  • Test Plan: No test scripts or test commands provided. Please add:

    • Example commands to run the pipeline in offline mode
    • Example commands to test the online serving endpoint
    • Expected behavior for each test case
  • Test Results: No actual test output showing the pipeline working. Please add:

    • Sample input (audio + video)
    • Sample output from AURA stage (text response)
    • Sample audio output from TTS stages
    • Performance metrics if available (latency per stage, memory usage)

🔍 Code Quality Observations

  1. Model wrapper rationale is clear: The AuraQwen3VLForConditionalGeneration wrapper properly addresses the remote config issue with AURA's custom configuration_qwen3_vl module. The fallback logic in get_hf_config() is robust.

  2. Video handling: The _strip_aura_videos_for_asr() method correctly removes video from ASR input and stashes it for downstream AURA consumption. This prevents unnecessary video processing at the ASR stage.

  3. Stage input processor tests: The unit tests for asr2aura and aura2tts are comprehensive, covering:

    • Video payload preservation across stages
    • Audio dropping before AURA stage
    • Silent response filtering
    • Token ID passthrough to TTS
    • Both TTS modes (Base and CustomVoice)
  4. Orchestrator changes: The silent path handling logic in _forward_to_next_stage() properly distinguishes between "finished with empty output" vs "waiting for more outputs", preventing request hangs.

🤔 Questions for Consideration

  1. Model architecture: Is AURA based on Qwen3-VL-Audio (similar to Qwen3-Omni)? The code treats it as a vision-language model without explicit audio modality handling at the AURA stage. Is this correct?

  2. Performance characteristics: What are the expected latencies for each stage? Are there any known bottlenecks, especially for video processing?

  3. Model availability: Is aurateam/AURA publicly accessible on Hugging Face? If not, please note that users will need to provide local checkpoint paths in the deploy config.

  4. Streaming behavior: How does the pipeline handle streaming inputs? The implementation appears focused on non-async chunk paths—please clarify if streaming is supported.

  5. Memory usage: With 4 stages and video + audio modalities, what are the GPU memory requirements? Should the docs include hardware recommendations?

💡 Suggestions

  1. Add test results: Populate the "Test Plan" and "Test Result" sections of the PR description with concrete examples and metrics.

  2. Consider integration test: Add a smoke test that loads the pipeline and runs a minimal example (similar to the existing example, but automated and checked).

  3. Document hardware requirements: Add recommended GPU setup for the 4-stage pipeline (e.g., "1x A100 80GB" or similar).

  4. Clarify model availability: If the model is not public, add a note in the docs about obtaining access.

Conclusion

The implementation is technically sound with good code quality and comprehensive unit tests. The main gap is the lack of test results in the PR description, which would help reviewers and users understand the pipeline's behavior and performance. Adding those would strengthen the PR significantly.


Note: This is not a diffusion model PR—AURA is a Vision-Language model (Qwen3-VL) in a multimodal pipeline with ASR and TTS stages. Diffusion-specific requirements do not apply.

@NumberWan

NumberWan commented Jun 9, 2026 •

Copy link
Copy Markdown
Contributor

A few initial question: in the test files, should we use English-only fixture strings (e.g. transcripts/ref text) for consistency with most unit tests, or keep Chinese since AURA defaults to Chinese responses?

@R2-Y

R2-Y commented Jun 9, 2026

Copy link
Copy Markdown
Contributor Author

A few initial question: in the test files, should we use English-only fixture strings (e.g. transcripts/ref text) for consistency with most unit tests, or keep Chinese since AURA defaults to Chinese responses?

they default to Chinese response, but I will change all examples later with english only

resumable: bool = False,
) -> Any:
next_pool = self.stage_pools[next_stage_id]
if self._next_stage_input_is_tokens(next_input):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I don't understand the logics in this branching, consider add some docs here

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For other models, after the first stage reads text, subsequent stages can directly use tokens as input. However, aura_omni first uses ASR to convert speech to text, and the input for the second stage is text, not tokens. I am checking if the vocabulary of qwen3-asr is consistent with that of qwen3-vl. If they are same, I may be able to directly use the sampled tokens as input for Aura. If so, I will delete this part modification.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I tested passing tokens directly and found a precision issue. I would like to keep the current implementation in this PR.

request.external_req_id = request.request_id
return request

processor = self._get_stage_input_processor(next_stage_id)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We placed the input processor in AsyncOmniEngine, why here still require one?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same explanation as above

stage_submit_ts=submit_ts,
)
)
await self._cleanup_request_ids([req_id, *self._cfg_tracker.cleanup_parent(req_id)])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this is the logic for AURA? If so, don't need to consider cfg tracker here

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

function here is general for all similar models. Aura determines whether to respond based on the user's input. If the model does not respond (silent) in this round and the inference has ended, it will clear the request. If other similar models also support silent, this section can be reused.

*,
final_output_type: str | None,
) -> RequestOutput:
"""Build a terminal empty output when no downstream stage input exists."""

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does CompletionOutput have default value in its parameters? This seems just return a default CompletionOutput, could be replace by something like CompletionOutput()?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

vllm.output.CompletionOutput don't have default value

@amy-why-3459
amy-why-3459 force-pushed the aura_support branch 5 times, most recently from 8979060 to efd60d0 Compare June 10, 2026 06:40

return conversation, [engine_prompt]

def _is_aura_omni_pipeline(self) -> bool:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

should this function be placed in aura_omni/pipeline.py ?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This function should be kept in the serving_chat.py. The function name here is not quite accurate, I have changed it to a more generalized one. All cascaded models need to perform similar judgments and multimodal data defer operations.

stage_names.add(str(model_stage))
return {"asr", "aura", "code2wav"}.issubset(stage_names)

async def _strip_aura_videos_for_asr(

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

There are many other mm data exraction functions in serving chat, maybe we should make these exaction processes as a part of pipeline abstraction sometime later

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

same as above

@R2-Y R2-Y mentioned this pull request Jun 12, 2026
1 task done

@david6666666 david6666666 left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@lishunyang12 lishunyang12 left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The orchestrator silent-path change reads correct and well-scoped: empty next_inputs while not-finished waits for more outputs; while finished, it emits a terminal empty output from the final stage and cleans up the request + CFG-companion ids. This only fires when a stage finishes producing no downstream input (which previously hung the request), so it's effectively a lifecycle fix with no regression for pipelines that always produce inputs.

Two things worth confirming before merge:

  1. Please confirm _needs_multistage_multimodal_split() returns False for existing single-stage / Qwen3-Omni pipelines, so _preprocess_chat stays a no-op for them.
  2. Nits: sr=24000 is hardcoded in _build_terminal_empty_output (could be sourced from the final stage client for non-24kHz vocoders); the additional_information init block in serving_chat.py looks duplicated.

No streaming-video here (deferred to the #4424-based follow-up), so this is mergeable as model integration.

@lishunyang12 lishunyang12 added the ready label to trigger buildkite CI label Jun 17, 2026

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@NumberWan NumberWan left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

ci failure is not related to this PR

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the AURA wiring. I found a couple of blocking issues in the user-visible paths: streaming can hand only a delta/empty string to TTS, and the default text,audio response shape can expose both the ASR transcript and AURA answer as final text choices. I also left a smaller docs/client mismatch around the served model name.


def _extract_text(source_output: Any) -> str:
output = _extract_output(source_output)
text = getattr(output, "text", None)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In streaming chat completions, AR sampling params are coerced to DELTA output, and the output processor attaches the full generated text as cumulative_text on the final output. This helper returns output.text first, which can be empty or just the last delta, so aura2tts() can skip a non-silent AURA reply or synthesize only a suffix. Please prefer non-empty cumulative_text before text here, and add a regression using the existing _source_delta_final_output() test helper.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed

model_stage="asr",
execution_type=StageExecutionType.LLM_AR,
input_sources=(),
final_output=True,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

With both stage 0 and stage 1 marked as final_output=True and final_output_type="text", the default modalities=["text", "audio"] path will append choices from both final text stages. That means the OpenAI response can expose the ASR transcript as one text choice and the AURA answer as another, even though examples document AURA as the text response. Please make ASR internal-only for this pipeline, or explicitly filter stage-0 text from chat responses.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed

`pipeline: aura_omni`, so the four-stage topology is used even if the
command-line `--model` points at one of the component checkpoints.

Send requests with `"model": "aura_omni"`. The ASR, AURA, and Qwen3-TTS

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This request model name does not match the server command above, which sets --served-model-name aurateam/AURA. The curl script defaults to MODEL=aura_omni while the Python client uses aurateam/AURA, so following the docs/examples as written can produce an unknown-model error. Please make the served model name and clients consistent.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed

R2-Y added 9 commits June 19, 2026 21:10
Add an aura_omni pipeline that composes Qwen3-ASR, AURA/Qwen3-VL, and the native Qwen3-TTS Talker -> Code2Wav stages. Include deploy config, ASR-to-AURA and AURA-to-TTS stage processors, docs, and tests for the new multi-stage topology.

Signed-off-by: Rein Yang <ruiruyang2@gmail.com>
Adds offline, online, curl, and Gradio examples for the ASR -> AURA -> Qwen3-TTS pipeline, including docs and runnable launch scripts. Also adapts AURA’s Qwen3-VL stage to load its remote-config checkpoint safely and filters ASR audio before forwarding multimodal data to the VLM stage.

Signed-off-by: Rein Yang <ruiruyang2@gmail.com>
…age bridging and tuning deploy config

Complete native AURA Omni stage bridging, then update  to disable prefix caching on AURA/Talker and switch Talker+Code2Wav to the CustomVoice checkpoint, reducing cross-turn carryover and improving multi-request TTS consistency.

Signed-off-by: Rein Yang <ruiruyang2@gmail.com>
Always pass AURA assistant token ids to Qwen3-TTS via PRECOMPUTED_TEXT_IDS_KEY instead of re-tokenizing response text, teach Qwen3-TTS talker to accept precomputed text ids.

Signed-off-by: Rein Yang <ruiruyang2@gmail.com>
…dio code length

Signed-off-by: Rein Yang <ruiruyang2@gmail.com>
Handle empty next-stage inputs in orchestrator non-async forwarding: return a terminal empty output (text/audio by final_output_type) and clean up request state when upstream output is finished, and only debug-return when unfinished. This unblocks AURA silent responses where aura2tts emits no downstream TTS request.

Signed-off-by: Rein Yang <ruiruyang2@gmail.com>
Signed-off-by: R2-Y <ruiruyang2@gmail.com>
Replace Aura-specific ASR video stripping with stage-aware deferred multimodal handoff for cascaded pipelines, and make Qwen3-TTS token passthrough opt-in so the default path sends text. Add regression coverage for deferred modalities and TTS passthrough behavior.

Signed-off-by: R2-Y <ruiruyang2@gmail.com>
Signed-off-by: R2-Y <ruiruyang2@gmail.com>
Signed-off-by: R2-Y <ruiruyang2@gmail.com>
@R2-Y

R2-Y commented Jun 19, 2026

Copy link
Copy Markdown
Contributor Author

The orchestrator silent-path change reads correct and well-scoped: empty next_inputs while not-finished waits for more outputs; while finished, it emits a terminal empty output from the final stage and cleans up the request + CFG-companion ids. This only fires when a stage finishes producing no downstream input (which previously hung the request), so it's effectively a lifecycle fix with no regression for pipelines that always produce inputs.

Two things worth confirming before merge:

  1. Please confirm _needs_multistage_multimodal_split() returns False for existing single-stage / Qwen3-Omni pipelines, so _preprocess_chat stays a no-op for them.
  2. Nits: sr=24000 is hardcoded in _build_terminal_empty_output (could be sourced from the final stage client for non-24kHz vocoders); the additional_information init block in serving_chat.py looks duplicated.

No streaming-video here (deferred to the #4424-based follow-up), so this is mergeable as model integration.

  1. Confirmed and added regression coverage: single-stage returns False, and Qwen3-Omni returns False because the first thinker stage already owns the multimodal inputs while talker/code2wav do not introduce downstream-only modalities. AURA ASR -> AURA remains the positive case.

  2. terminal empty audio output now sources the sample rate from the final stage client/config when available and only falls back to 24kHz. I also deduplicated the additional_information initialization through a small helper while keeping it lazily initialized.

@R2-Y

R2-Y commented Jun 19, 2026

Copy link
Copy Markdown
Contributor Author

@hsliuustc0106 Ready to merge

@Gaohan123 Gaohan123 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Jun 21, 2026
@hsliuustc0106
hsliuustc0106 merged commit afbcc51 into vllm-project:main Jun 21, 2026
7 of 8 checks passed
@R2-Y
R2-Y deleted the aura_support branch June 23, 2026 07:23
zeningc added a commit to zeningc/vllm-omni that referenced this pull request Jul 1, 2026
Merging main pulled in new call sites that used the pre-rename output
queue API, which this branch had already renamed output_async_queue ->
output_sync_queue (put_nowait on the sync side). These were semantic
conflicts: git auto-merged them without a textual conflict, so they
landed referencing an attribute the constructor no longer sets.

- orchestrator.py:1251 (from vllm-project#4079, diffusion request-level batching)
  and orchestrator.py:1419 (from vllm-project#4257, Aura non-async-chunk path) still
  called `await self.output_async_queue.put(...)`, which would raise
  AttributeError on those terminal-output error/edge paths. Convert to
  `self.output_sync_queue.put_nowait(...)`.
- tests/engine/test_orchestrator_stage_input_bridge.py (new file from
  vllm-project#4257, marked core_model/cpu) constructed Orchestrator with the old
  `output_async_queue=` kwarg and failed against the new signature.
  Update to `output_sync_queue=output_q.sync_q`.

Tested on vLLM 0.24.0: engine + orchestrator unit tests (41 passed) and
a Qwen3-TTS-12Hz-0.6B streaming /v1/audio/speech smoke test.

Signed-off-by: zeningc <zening.chen@yahoo.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Signed-off-by: Rein Yang <ruiruyang2@gmail.com>
Signed-off-by: R2-Y <ruiruyang2@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

9 participants