Skip to content

[Model] Add Qwen3-Omni duplex plugin and pipeline variant - #7691

Open
mjZhaoElaine wants to merge 8 commits into
vllm-project:mainfrom
mjZhaoElaine:VO-20260916-01-S02-port-qwen3-omni-duplex-plugin
Open

mjZhaoElaine wants to merge 8 commits into
vllm-project:mainfrom
mjZhaoElaine:VO-20260916-01-S02-port-qwen3-omni-duplex-plugin

Conversation

@mjZhaoElaine

@mjZhaoElaine mjZhaoElaine commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Toward #7636 (Qwen3-Omni Server VAD restore, Issue 5 and Issue 18).

Adds an opt-in Qwen3-Omni duplex serving path. Stock qwen3_omni_moe stays turn-based. Operators select pipeline: qwen3_omni_moe_duplex with session_mode: duplex (bundled vllm_omni/deploy/qwen3_omni_moe_duplex.yaml) to load a turn-commit plugin that submits one ephemeral Thinker→Talker→Code2Wav request per committed Server VAD utterance, projects Thinker text through observe_stage_output as response.output_audio_transcript.delta, and leaves Talker/Code2Wav to produce audio.

The plugin encodes Code2Wav deltas with the session response_format, drains incremental audio chunks, and rejects non-final plan_append instead of submitting a dummy prompt. Bounded conversation history, session instructions / sampling, Code2Wav model_outputs, text-only Realtime turns, and interrupt_response: false (half-duplex / barge-in off) are included. Thinker stage-0 output is intermediate: only a Code2Wav stage id can mark end_of_turn. Existing MiniCPM-o / PersonaPlex / Nemotron duplex deployments are unchanged. Legacy YAMLs that combined session_mode: turn with a duplex_session: block must migrate to the new variant.

This branch is rebased onto current main (AURA #7633 is already merged). The turn-commit plugin lives in duplex/moe_plugin.py so it coexists with stock duplex/plugin.py.

Issue 18 is in this PR: test_server_vad_multi_turn_without_client_commit is retargeted to pipeline: qwen3_omni_moe_duplex + session_mode: duplex, requires turn 2 to recall turn 1 (Orchid Falcon), and is listed as a dedicated Buildkite merge step at --run-level advanced_model. The GPU helper acks completed playback and stops sending after the first input_audio_buffer.speech_stopped so leftover synthetic PCM is not treated as a second utterance. Barge-in / talking over TTS is out of scope for this half-duplex path.

Test Plan

vLLM Version: 0.30.0

vLLM-Omni Commit: deaed8b

  • CPU L1: plugin contract (history, instructions, sampling, text-only, intermediate Thinker output), PCM commit buffer, pipeline-variant isolation, data-plane model_outputs drain
  • CPU L2: Server VAD turns, second-turn history contents, bounded bookkeeping, instructions/sampling on submit, text-only session stays open, interrupt_response: false
  • GPU /v1/realtime Server VAD e2e: two Server VAD turns on Instruct weights (load_format=auto), half-duplex helper acks playback and sends one utterance per turn, turn 2 recalls Orchid Falcon
pytest -sv tests/model_executor/models/qwen3_omni/duplex tests/config/test_config_factory.py::TestPipelineDiscovery::test_registry_has_known_models -m 'core_model and cpu'
pytest -sv tests/engine/duplex/test_qwen3_omni_server_vad_turn.py tests/engine/duplex/test_vad_backend.py tests/engine/test_duplex_orchestrator.py -m 'core_model and cpu'
pytest -s -v tests/entrypoints/openai_api/test_qwen3_omni_realtime_websocket.py::TestQwen3OmniRealtimeWebSocket::test_server_vad_multi_turn_without_client_commit -m 'advanced_model and cuda and H100' --run-level advanced_model

Test Result

CPU gate green on a RunPod CPU session (vLLM 0.30.0):

id	name	exit	elapsed	verdict
00	preflight	0	0m13s	PASS
10	plugin-l1	0	0m16s	PASS(35 passed)
20	server-vad-turn	0	0m16s	PASS(9 passed)

Extra CPU sweep of test_duplex_orchestrator.py + test_vad_backend.py: 77 passed, 1 skipped.

Testing Done

GPU /v1/realtime Server VAD e2e passed on 2× NVIDIA RTX PRO 6000 Blackwell, vLLM 0.30.0, --run-level advanced_model (Instruct weights, load_format=auto). One case: two Server VAD turns, no client commit, interrupt_response: false.

Recipe: vo-s02-qwen3-omni-duplex --only "00 40" running

tests/entrypoints/openai_api/test_qwen3_omni_realtime_websocket.py::TestQwen3OmniRealtimeWebSocket::test_server_vad_multi_turn_without_client_commit

id  name             exit  elapsed  verdict
00  preflight        0     0m11s    PASS
40  gpu-server-vad   0     2m16s    PASS(1 passed)
GATE_VERDICT=PASS
DONE 0
[VAD_DIAG H1] sent_bytes=156800/171486 stopped_on_speech_stopped=True
[VAD_DIAG H2] turn=1 counts started=1 stopped=1 committed=1 created=1 done=1 overlap=0
[VAD_DIAG H1] sent_bytes=102400/219656 stopped_on_speech_stopped=True
[VAD_DIAG H2] turn=2 counts started=1 stopped=1 committed=1 created=1 done=1 overlap=0
Initializing a V1 LLM engine (v0.30.0) with config: model='Qwen/Qwen3-Omni-30B-A3B-Instruct' ... load_format=auto
================== 1 passed, 17 warnings in 124.06s (0:02:04) ==================

Made with Cursor

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/engine_orchestration.md, docs/design/module/model_integration.md, docs/design/module/ar_runtime.md.

Module owners: @tzhouam @fake0fan @Gaohan123

Routing: @tzhouam via module of the changed files, CODEOWNERS; @fake0fan via module of the changed files; @Gaohan123 via module of the changed files

@mjZhaoElaine, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 17, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit e7a57b9d6a44 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@hsliuustc0106 hsliuustc0106 added high priority high priority issue, needs to be done asap omni code related to omni models enhancement New feature or request labels Sep 17, 2026

Copy link
Copy Markdown

Thanks for pushing Qwen3-Omni onto the unified duplex path. I ran an independent real-GPU validation of the current PR head 50ceef31cc10fbeb0c94d803b74a9fa66c6baaf5 on DGX Spark / GB10 (ARM64), mainly to exercise the GPU /v1/realtime path that is deferred to the follow-up.

To get the path running, I needed two validation-only adaptations:

  • client: interrupt_response=false -> true, since the same-commit server rejects false;
  • local runtime bridge: run the raw Qwen multimodal prompt through the existing Stage0 InputProcessor, otherwise the first committed turn reaches build_engine_core_request_from_tokens() and raises KeyError: 'prompt_token_ids'.

With those in place, single-turn Thinker -> Talker -> Code2Wav works end to end.

The main issue I then saw is cross-turn semantic history. Using one WebSocket/session, Server VAD, and no client input_audio_buffer.commit / response.create:

Turn 1: "My code name is Orchid Falcon."
Turn 2: "What is my code name?"

Turn 1 completed with:

Understood. I will remember the codename "Orchid Falcon".

After full playback.ack, the server reported history_committed: true. When Turn 2 was committed, the same session reported history_len: 3, and the new user item correctly referenced the Turn 1 assistant item as previous_item_id.

However, Turn 2 completed normally with:

You did not ask me to remember a codename.

This looks consistent with the current Qwen plugin shape: each committed turn creates a fresh non-resumable Stage0 request, while Qwen3OmniDuplexPlugin.plan_append() builds the Thinker prompt from the system prompt + current audio (+ optional initial_user_text) without rendering the existing canonical session history.

So the session/history ownership itself appears to be working; the missing seam seems to be canonical history -> Qwen Stage0 prompt construction.

I have a small two-turn realtime acceptance client for this flow (Server VAD, no client commit/create, event/id checks, optional semantic assertion) and can contribute/adapt it as a regression test if that would be useful.

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The plugin shape is the right way to bring Qwen3-Omni onto the unified duplex runtime (#7413), and the opt-in deployment variant keeps the turn-based path untouched. The integration is not finished, though: the plan hands the runtime a text-plus-audio prompt that the stage port cannot submit, each turn starts a fresh conversation, completed ephemeral requests are never released, and two negotiated session settings are ignored. All items are inline; two independent reviewers confirmed the six and added the data-plane audio key mismatch and the text-only turn failure, which are folded in. The branch also conflicts with main, so a rebase is needed regardless.

Static read at 50ceef31 against merge-base 63032c91, all 28 files including the stacked framework changes; fork head, no PR code executed; pre-commit and DCO green, no ready label so no general lane has run, and the description already says GPU validation was deferred. The submission path was traced: DuplexOrchestrator.submit passes dict(submission.prompt) straight into build_engine_core_request_from_tokens, which reads prompt["prompt_token_ids"] on its first line; the MiniCPM-o and Nemotron plugins emit token ids in their plans for exactly that reason.

"is_speech": bool(payload.get("is_speech", True)),
}
prompt: dict[str, object] = {
"prompt": _thinker_prompt(system_prompt=system_prompt, user_text=user_text),

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[blocking] This plan omits prompt_token_ids on purpose, but DuplexOrchestrator.submit passes the plan's prompt straight to build_engine_core_request_from_tokens, whose first line is prompt_token_ids = prompt["prompt_token_ids"], so every committed utterance raises KeyError before the Thinker sees it. That builder also does not process multi_modal_data, so adding token ids alone would still drop the audio features. Run the prompt through the stage input processor (or tokenize and preprocess in the plugin the way the MiniCPM-o and Nemotron plans do) and add a test that drives the real orchestrator submission path rather than the plan object alone.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. Committed turns now go through the Stage-0 input processor so prompt_token_ids and audio features survive build_engine_core_request_from_tokens. Covered by the real submit-path tests and the GPU first turn (no KeyError). See ccf29ca21.

extra = user_text if user_text else ""
return (
f"<|im_start|>system\n{system_prompt}<|im_end|>\n"
"<|im_start|>user\n<|audio_start|><|audio_pad|><|audio_end|>"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[important] This template renders only the system message and the current utterance, and each turn is submitted under a fresh request id with resumable=False, so the previous model context is discarded every time. A second VAD turn that refers to the previous answer starts from nothing even though the session holds committed history. Render the relevant prior user audio/text and assistant turns into the prompt (bounded, as the turn-based path does) and assert the contents of the second submitted prompt in a test.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. plan_append now renders bounded prior user/assistant content into the Thinker prompt. L1/L2 assert the second-turn prompt contents; the GPU Server VAD case requires turn 2 to recall Orchid Falcon. See ccf29ca21 and d139b5706.

# decision short-circuited the pipeline — closes the stream (and
# offers the model another silence unit).
response_completes_here = item.stage_id >= item.context.final_stage_id or item.decision is not None
if finished and emitted_response and not self._out.auto_responds() and response_completes_here:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[important] Completion here only closes the projector stream, and the orchestrator intentionally skips automatic cleanup for session-owned requests. With a new request id per turn, request states, pool bindings, session resource entries and projector state accumulate until the session closes, and the running-request counter is incremented per turn without ever being decremented. Add a terminal cleanup path for non-resumable session requests and a test that repeated completed turns leave bounded bookkeeping.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. Completed non-resumable ephemeral session requests are released; L2 asserts the Stage-0 bind does not remain after the turn finishes. See ccf29ca21.

updated["qwen3_system_prompt"] = str(extra["qwen3_system_prompt"])
if config.instructions:
updated["instructions"] = config.instructions
updated.setdefault("qwen3_system_prompt", config.instructions)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[important] prepare_runtime_config always seeds qwen3_system_prompt, so the setdefault here never applies a later instructions update: open a session, update instructions before the first audio commit, and session.updated echoes the new text while plan_append keeps building with the original prompt. Replace the effective system prompt when an instructions update is accepted and no explicit qwen3_system_prompt override is set.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. An accepted instructions update replaces the effective system prompt unless qwen3_system_prompt is set. L1/L2 cover open-then-update-before-first-commit. See ccf29ca21.

runtime_config: dict[str, object],
defaults: tuple[object, ...],
) -> tuple[object, ...]:
del runtime_config

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[important] Neither runtime-config method carries temperature or max_tokens, and this method ignores runtime_config entirely, so a session negotiating max_response_output_tokens=16 still submits the Thinker with the deployment's max_tokens=2048 and its temperature is ignored. Propagate the negotiated sampling settings on session creation and on updates; as written the client's generation limits have no effect on latency or resource use.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. Negotiated temperature and max_tokens / max_response_output_tokens are applied on session create and update and reach Thinker sampling. See ccf29ca21.

Keep Qwen's default `session_mode: turn`; `duplex_session` enables the Realtime handler without changing scheduler
semantics. Bare `/v1/realtime` selects that handler. The client uses `?duplex=0` only for the existing non-Server-VAD
wire flow; `?duplex=1` remains a supported compatibility alias.
Bare `/v1/realtime` selects the duplex handler. Use `?duplex=0` for the

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] With the duplex YAML shown just above, duplex startup sets openai_serving_realtime=None, so ?duplex=0 answers "Realtime API is not available" and closes the socket rather than falling back to the turn-based handler. Say here that the legacy flow needs the stock turn-based deployment.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. The README now states that ?duplex=0 on the duplex YAML is not a turn-based fallback and that the legacy flow needs the stock turn-based deploy. See 52f1b3df9.


def _audio_payload(metadata: Mapping[str, object]) -> object | None:
"""Code2Wav PCM under the ``audio`` key."""
return metadata.get("audio")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[important] _audio_payload reads metadata.get("audio"), but the Qwen3-Omni Code2Wav stage emits its waveform under model_outputs (with sr beside it), which is why the turn-based realtime path selects "audio" if "audio" in mm else "model_outputs". With the key as written, real Code2Wav audio never projects to the client, and the new tests do not notice because both of them hard-code audio in their fake stage output. Accept model_outputs here (or normalise the key at the stage boundary) and feed a real-shaped Code2Wav payload through the test.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. The data plane accepts Code2Wav model_outputs (with sr); L1 feeds a real-shaped payload rather than a hard-coded audio key. GPU e2e asserts response.output_audio.delta. See ccf29ca21 and d139b5706.

audio_b64 = payload.get("audio")
has_audio = isinstance(audio_b64, str) and bool(audio_b64)
if not has_audio:
raise ValueError("Qwen3-Omni duplex commit requires audio")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[important] A text-only turn is a valid Realtime flow (conversation.item.create with text, then response.create), and this unconditional check raises for it, which the runtime turns into a session failure rather than a rejected command. Either accept text-only commits (the Thinker prompt already renders user_text) or reject them with a protocol error that leaves the session open.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed. A text-only conversation.item.create plus response.create submits without failing the session. See ccf29ca21.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 7 days

@mjZhaoElaine this pull request has had no human commit, comment or review since 2026-09-22. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@mjZhaoElaine
mjZhaoElaine force-pushed the VO-20260916-01-S02-port-qwen3-omni-duplex-plugin branch from 50ceef3 to 4ec3b2c Compare September 29, 2026 18:50
@mjZhaoElaine mjZhaoElaine changed the title [Model] Add Qwen3-Omni duplex plugin and pipeline variant [WIP][Model] Add Qwen3-Omni duplex plugin and pipeline variant Sep 29, 2026
@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Omni ReviewBot: superseded

The CI failure noted on 4ec3b2c7415a refers to an earlier head; the pull request now points at d139b570692c.

@mjZhaoElaine mjZhaoElaine changed the title [WIP][Model] Add Qwen3-Omni duplex plugin and pipeline variant [Model] Add Qwen3-Omni duplex plugin and pipeline variant Sep 29, 2026
@mjZhaoElaine

Copy link
Copy Markdown
Contributor Author

@JoeLoveSally thanks for the GPU notes on 50ceef31. Those three seams are in this head:

  • interrupt_response: false is accepted as half-duplex (listen_only); the same-commit reject is gone (ccf29ca21).
  • Committed turns go through the Stage-0 input processor, so the first turn no longer hits KeyError: prompt_token_ids.
  • plan_append renders bounded session history. The re-enabled GPU case (test_server_vad_multi_turn_without_client_commit) requires turn 2 to recall Orchid Falcon; that passed on Instruct weights with load_format=auto (d139b5706).

Your two-turn Server VAD client shape is what that e2e now follows. Barge-in / talking over TTS stays out of scope for this half-duplex path.

mjZhaoElaine and others added 2 commits September 29, 2026 16:40
Register an opt-in qwen3_omni_moe_duplex pipeline that commits one
ephemeral Thinker→Talker→Code2Wav request per Server VAD turn.

Signed-off-by: Mengjie Zhao <zmj0129@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Encode Code2Wav deltas with the session response format, drain
incremental chunks, and fail closed on non-final plan_append so the
CPU gate can prove the ephemeral Stage0 re-bind.

Signed-off-by: Mengjie Zhao <zmj0129@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Mengjie Zhao <zmj0129@gmail.com>
mjZhaoElaine and others added 5 commits September 29, 2026 16:40
Match existing README/YAML comment style: say which YAML to pass, drop migration history.

Signed-off-by: Mengjie Zhao <zmj0129@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Render bounded conversation history, apply session instructions and
sampling, project Code2Wav model_outputs, accept interrupt_response
false, and keep text-only Realtime turns from failing the session.

Signed-off-by: Mengjie Zhao <zmj0129@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Retarget the websocket Server VAD case onto qwen3_omni_moe_duplex,
require turn-2 recall of Orchid Falcon, and run it from a dedicated
merge step at --run-level advanced_model.

Signed-off-by: Mengjie Zhao <zmj0129@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
The GPU client was dumping leftover repeating synth after VAD endpoint,
which looked like a second user utterance during TTS.

Signed-off-by: Mengjie Zhao <zmj0129@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Match the stock plugin so Stage-0 finished is intermediate, and only
Code2Wav stage ids can mark end_of_turn.

Signed-off-by: Mengjie Zhao <zmj0129@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@mjZhaoElaine
mjZhaoElaine force-pushed the VO-20260916-01-S02-port-qwen3-omni-duplex-plugin branch from e82f08d to deaed8b Compare September 29, 2026 23:41
CI ruff-format wants the import string on one line.

Signed-off-by: Mengjie Zhao <zmj0129@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 7 days

@mjZhaoElaine this pull request has had no human commit, comment or review since 2026-09-29. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request high priority high priority issue, needs to be done asap omni code related to omni models

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants