Repository navigation
[Model] Add Qwen3-Omni duplex plugin and pipeline variant - #7691
mjZhaoElaine wants to merge 8 commits into
Conversation
|
This PR appears to belong to: docs/design/module/engine_orchestration.md, docs/design/module/model_integration.md, docs/design/module/ar_runtime.md. Module owners: @tzhouam @fake0fan @Gaohan123 Routing: @tzhouam via module of the changed files, CODEOWNERS; @fake0fan via module of the changed files; @Gaohan123 via module of the changed files @mjZhaoElaine, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
|
Thanks for pushing Qwen3-Omni onto the unified duplex path. I ran an independent real-GPU validation of the current PR head To get the path running, I needed two validation-only adaptations:
With those in place, single-turn Thinker -> Talker -> Code2Wav works end to end. The main issue I then saw is cross-turn semantic history. Using one WebSocket/session, Server VAD, and no client Turn 1 completed with: After full However, Turn 2 completed normally with: This looks consistent with the current Qwen plugin shape: each committed turn creates a fresh non-resumable Stage0 request, while So the session/history ownership itself appears to be working; the missing seam seems to be canonical history -> Qwen Stage0 prompt construction. I have a small two-turn realtime acceptance client for this flow (Server VAD, no client commit/create, event/id checks, optional semantic assertion) and can contribute/adapt it as a regression test if that would be useful. |
linyueqian
left a comment
There was a problem hiding this comment.
The plugin shape is the right way to bring Qwen3-Omni onto the unified duplex runtime (#7413), and the opt-in deployment variant keeps the turn-based path untouched. The integration is not finished, though: the plan hands the runtime a text-plus-audio prompt that the stage port cannot submit, each turn starts a fresh conversation, completed ephemeral requests are never released, and two negotiated session settings are ignored. All items are inline; two independent reviewers confirmed the six and added the data-plane audio key mismatch and the text-only turn failure, which are folded in. The branch also conflicts with main, so a rebase is needed regardless.
Static read at 50ceef31 against merge-base 63032c91, all 28 files including the stacked framework changes; fork head, no PR code executed; pre-commit and DCO green, no ready label so no general lane has run, and the description already says GPU validation was deferred. The submission path was traced: DuplexOrchestrator.submit passes dict(submission.prompt) straight into build_engine_core_request_from_tokens, which reads prompt["prompt_token_ids"] on its first line; the MiniCPM-o and Nemotron plugins emit token ids in their plans for exactly that reason.
| "is_speech": bool(payload.get("is_speech", True)), | ||
| } | ||
| prompt: dict[str, object] = { | ||
| "prompt": _thinker_prompt(system_prompt=system_prompt, user_text=user_text), |
There was a problem hiding this comment.
[blocking] This plan omits prompt_token_ids on purpose, but DuplexOrchestrator.submit passes the plan's prompt straight to build_engine_core_request_from_tokens, whose first line is prompt_token_ids = prompt["prompt_token_ids"], so every committed utterance raises KeyError before the Thinker sees it. That builder also does not process multi_modal_data, so adding token ids alone would still drop the audio features. Run the prompt through the stage input processor (or tokenize and preprocess in the plugin the way the MiniCPM-o and Nemotron plans do) and add a test that drives the real orchestrator submission path rather than the plan object alone.
There was a problem hiding this comment.
Fixed. Committed turns now go through the Stage-0 input processor so prompt_token_ids and audio features survive build_engine_core_request_from_tokens. Covered by the real submit-path tests and the GPU first turn (no KeyError). See ccf29ca21.
| extra = user_text if user_text else "" | ||
| return ( | ||
| f"<|im_start|>system\n{system_prompt}<|im_end|>\n" | ||
| "<|im_start|>user\n<|audio_start|><|audio_pad|><|audio_end|>" |
There was a problem hiding this comment.
[important] This template renders only the system message and the current utterance, and each turn is submitted under a fresh request id with resumable=False, so the previous model context is discarded every time. A second VAD turn that refers to the previous answer starts from nothing even though the session holds committed history. Render the relevant prior user audio/text and assistant turns into the prompt (bounded, as the turn-based path does) and assert the contents of the second submitted prompt in a test.
There was a problem hiding this comment.
Fixed. plan_append now renders bounded prior user/assistant content into the Thinker prompt. L1/L2 assert the second-turn prompt contents; the GPU Server VAD case requires turn 2 to recall Orchid Falcon. See ccf29ca21 and d139b5706.
| # decision short-circuited the pipeline — closes the stream (and | ||
| # offers the model another silence unit). | ||
| response_completes_here = item.stage_id >= item.context.final_stage_id or item.decision is not None | ||
| if finished and emitted_response and not self._out.auto_responds() and response_completes_here: |
There was a problem hiding this comment.
[important] Completion here only closes the projector stream, and the orchestrator intentionally skips automatic cleanup for session-owned requests. With a new request id per turn, request states, pool bindings, session resource entries and projector state accumulate until the session closes, and the running-request counter is incremented per turn without ever being decremented. Add a terminal cleanup path for non-resumable session requests and a test that repeated completed turns leave bounded bookkeeping.
There was a problem hiding this comment.
Fixed. Completed non-resumable ephemeral session requests are released; L2 asserts the Stage-0 bind does not remain after the turn finishes. See ccf29ca21.
| updated["qwen3_system_prompt"] = str(extra["qwen3_system_prompt"]) | ||
| if config.instructions: | ||
| updated["instructions"] = config.instructions | ||
| updated.setdefault("qwen3_system_prompt", config.instructions) |
There was a problem hiding this comment.
[important] prepare_runtime_config always seeds qwen3_system_prompt, so the setdefault here never applies a later instructions update: open a session, update instructions before the first audio commit, and session.updated echoes the new text while plan_append keeps building with the original prompt. Replace the effective system prompt when an instructions update is accepted and no explicit qwen3_system_prompt override is set.
There was a problem hiding this comment.
Fixed. An accepted instructions update replaces the effective system prompt unless qwen3_system_prompt is set. L1/L2 cover open-then-update-before-first-commit. See ccf29ca21.
| runtime_config: dict[str, object], | ||
| defaults: tuple[object, ...], | ||
| ) -> tuple[object, ...]: | ||
| del runtime_config |
There was a problem hiding this comment.
[important] Neither runtime-config method carries temperature or max_tokens, and this method ignores runtime_config entirely, so a session negotiating max_response_output_tokens=16 still submits the Thinker with the deployment's max_tokens=2048 and its temperature is ignored. Propagate the negotiated sampling settings on session creation and on updates; as written the client's generation limits have no effect on latency or resource use.
There was a problem hiding this comment.
Fixed. Negotiated temperature and max_tokens / max_response_output_tokens are applied on session create and update and reach Thinker sampling. See ccf29ca21.
| Keep Qwen's default `session_mode: turn`; `duplex_session` enables the Realtime handler without changing scheduler | ||
| semantics. Bare `/v1/realtime` selects that handler. The client uses `?duplex=0` only for the existing non-Server-VAD | ||
| wire flow; `?duplex=1` remains a supported compatibility alias. | ||
| Bare `/v1/realtime` selects the duplex handler. Use `?duplex=0` for the |
There was a problem hiding this comment.
[suggestion] With the duplex YAML shown just above, duplex startup sets openai_serving_realtime=None, so ?duplex=0 answers "Realtime API is not available" and closes the socket rather than falling back to the turn-based handler. Say here that the legacy flow needs the stock turn-based deployment.
There was a problem hiding this comment.
Fixed. The README now states that ?duplex=0 on the duplex YAML is not a turn-based fallback and that the legacy flow needs the stock turn-based deploy. See 52f1b3df9.
|
|
||
| def _audio_payload(metadata: Mapping[str, object]) -> object | None: | ||
| """Code2Wav PCM under the ``audio`` key.""" | ||
| return metadata.get("audio") |
There was a problem hiding this comment.
[important] _audio_payload reads metadata.get("audio"), but the Qwen3-Omni Code2Wav stage emits its waveform under model_outputs (with sr beside it), which is why the turn-based realtime path selects "audio" if "audio" in mm else "model_outputs". With the key as written, real Code2Wav audio never projects to the client, and the new tests do not notice because both of them hard-code audio in their fake stage output. Accept model_outputs here (or normalise the key at the stage boundary) and feed a real-shaped Code2Wav payload through the test.
There was a problem hiding this comment.
Fixed. The data plane accepts Code2Wav model_outputs (with sr); L1 feeds a real-shaped payload rather than a hard-coded audio key. GPU e2e asserts response.output_audio.delta. See ccf29ca21 and d139b5706.
| audio_b64 = payload.get("audio") | ||
| has_audio = isinstance(audio_b64, str) and bool(audio_b64) | ||
| if not has_audio: | ||
| raise ValueError("Qwen3-Omni duplex commit requires audio") |
There was a problem hiding this comment.
[important] A text-only turn is a valid Realtime flow (conversation.item.create with text, then response.create), and this unconditional check raises for it, which the runtime turns into a session failure rather than a rejected command. Either accept text-only commits (the Thinker prompt already renders user_text) or reject them with a protocol error that leaves the session open.
There was a problem hiding this comment.
Fixed. A text-only conversation.item.create plus response.create submits without failing the session. See ccf29ca21.
Omni ReviewBot: no human activity for 7 days@mjZhaoElaine this pull request has had no human commit, comment or review since 2026-09-22. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
50ceef3 to
4ec3b2c
Compare
Omni ReviewBot: supersededThe CI failure noted on |
|
@JoeLoveSally thanks for the GPU notes on
Your two-turn Server VAD client shape is what that e2e now follows. Barge-in / talking over TTS stays out of scope for this half-duplex path. |
Register an opt-in qwen3_omni_moe_duplex pipeline that commits one ephemeral Thinker→Talker→Code2Wav request per Server VAD turn. Signed-off-by: Mengjie Zhao <zmj0129@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Encode Code2Wav deltas with the session response format, drain incremental chunks, and fail closed on non-final plan_append so the CPU gate can prove the ephemeral Stage0 re-bind. Signed-off-by: Mengjie Zhao <zmj0129@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com> Signed-off-by: Mengjie Zhao <zmj0129@gmail.com>
Match existing README/YAML comment style: say which YAML to pass, drop migration history. Signed-off-by: Mengjie Zhao <zmj0129@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Render bounded conversation history, apply session instructions and sampling, project Code2Wav model_outputs, accept interrupt_response false, and keep text-only Realtime turns from failing the session. Signed-off-by: Mengjie Zhao <zmj0129@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Retarget the websocket Server VAD case onto qwen3_omni_moe_duplex, require turn-2 recall of Orchid Falcon, and run it from a dedicated merge step at --run-level advanced_model. Signed-off-by: Mengjie Zhao <zmj0129@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
The GPU client was dumping leftover repeating synth after VAD endpoint, which looked like a second user utterance during TTS. Signed-off-by: Mengjie Zhao <zmj0129@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Match the stock plugin so Stage-0 finished is intermediate, and only Code2Wav stage ids can mark end_of_turn. Signed-off-by: Mengjie Zhao <zmj0129@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
e82f08d to
deaed8b
Compare
CI ruff-format wants the import string on one line. Signed-off-by: Mengjie Zhao <zmj0129@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Omni ReviewBot: no human activity for 7 days@mjZhaoElaine this pull request has had no human commit, comment or review since 2026-09-29. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
Purpose
Toward #7636 (Qwen3-Omni Server VAD restore, Issue 5 and Issue 18).
Adds an opt-in Qwen3-Omni duplex serving path. Stock
qwen3_omni_moestays turn-based. Operators selectpipeline: qwen3_omni_moe_duplexwithsession_mode: duplex(bundledvllm_omni/deploy/qwen3_omni_moe_duplex.yaml) to load a turn-commit plugin that submits one ephemeral Thinker→Talker→Code2Wav request per committed Server VAD utterance, projects Thinker text throughobserve_stage_outputasresponse.output_audio_transcript.delta, and leaves Talker/Code2Wav to produce audio.The plugin encodes Code2Wav deltas with the session
response_format, drains incremental audio chunks, and rejects non-finalplan_appendinstead of submitting a dummy prompt. Bounded conversation history, sessioninstructions/ sampling, Code2Wavmodel_outputs, text-only Realtime turns, andinterrupt_response: false(half-duplex / barge-in off) are included. Thinker stage-0 output is intermediate: only a Code2Wav stage id can markend_of_turn. Existing MiniCPM-o / PersonaPlex / Nemotron duplex deployments are unchanged. Legacy YAMLs that combinedsession_mode: turnwith aduplex_session:block must migrate to the new variant.This branch is rebased onto current
main(AURA #7633 is already merged). The turn-commit plugin lives induplex/moe_plugin.pyso it coexists with stockduplex/plugin.py.Issue 18 is in this PR:
test_server_vad_multi_turn_without_client_commitis retargeted topipeline: qwen3_omni_moe_duplex+session_mode: duplex, requires turn 2 to recall turn 1 (Orchid Falcon), and is listed as a dedicated Buildkite merge step at--run-level advanced_model. The GPU helper acks completed playback and stops sending after the firstinput_audio_buffer.speech_stoppedso leftover synthetic PCM is not treated as a second utterance. Barge-in / talking over TTS is out of scope for this half-duplex path.Test Plan
vLLM Version: 0.30.0
vLLM-Omni Commit: deaed8b
model_outputsdraininterrupt_response: false/v1/realtimeServer VAD e2e: two Server VAD turns on Instruct weights (load_format=auto), half-duplex helper acks playback and sends one utterance per turn, turn 2 recalls Orchid FalconTest Result
CPU gate green on a RunPod CPU session (vLLM 0.30.0):
Extra CPU sweep of
test_duplex_orchestrator.py+test_vad_backend.py: 77 passed, 1 skipped.Testing Done
GPU
/v1/realtimeServer VAD e2e passed on 2× NVIDIA RTX PRO 6000 Blackwell, vLLM 0.30.0,--run-level advanced_model(Instruct weights,load_format=auto). One case: two Server VAD turns, no client commit,interrupt_response: false.Recipe:
vo-s02-qwen3-omni-duplex --only "00 40"runningtests/entrypoints/openai_api/test_qwen3_omni_realtime_websocket.py::TestQwen3OmniRealtimeWebSocket::test_server_vad_multi_turn_without_client_commitMade with Cursor