Skip to content

[Core] Graduate MiniCPM-o 4.5 and PersonaPlex full-duplex serving out of experimental - #6196

Merged
Gaohan123 merged 47 commits into
vllm-project:mainfrom
chickeyton:fullduplex3
Sep 5, 2026
Merged

Gaohan123 merged 47 commits into
vllm-project:mainfrom
chickeyton:fullduplex3

Conversation

@chickeyton

@chickeyton chickeyton commented Aug 14, 2026 •

Copy link
Copy Markdown
Contributor

As the PR 1 of [RFC]: Moving MiniCPM-o 4.5 and PersonaPlex out of Experimental

The full-duplex stack currently lives under vllm_omni/experimental/fullduplex/. It was placed there when MiniCPM-o 4.5 duplex serving landed (#3907), but the situation has changed since then:

  • The code is no longer experimental in practice. It backs the /v1/realtime?duplex=1 endpoint, has unit and E2E coverage wired into CI, and has been validated end to end on GPUs.
  • PersonaPlex ([Feature] PersonaPlex (Moshi-based full-duplex S2S): native vLLM port + duplex serving #4771) became the second model to plug into the same duplex runtime, and it did so without any change to the shared code. Two independent models (turn-based MiniCPM-o vs. a continuous Moshi-style model) running through the same seams is a good signal that the architecture is stable enough to graduate.
  • Most of what sits under experimental/fullduplex/ is not model code at all — it is a model-neutral engine control plane and a model-neutral serving stack. Keeping generic infrastructure under an experimental namespace sends the wrong signal to users and makes the import paths awkward for the stable modules that already depend on it.

The joyvl/ demo and the small core/ scaffold it uses are genuinely demo-only and are not part of this proposal; they stay in experimental/ for now.

Understanding of the current code structure: slides

Purpose

Pure Relocation

No behavior change

What From experimental/fullduplex/ To Level
Duplex engine kernel (control plane, sessions, leases, fences, messages) engine/ vllm_omni/engine/duplex/ Orchestrator
Generic duplex serving stack (WebSocket handler, session runner, Realtime projection) openai/ vllm_omni/entrypoints/duplex/ Serving
MiniCPM-o 4.5 duplex adapter minicpmo45/ vllm_omni/model_executor/models/minicpmo_4_5/duplex/ Serving, Orchestrator , Worker
PersonaPlex duplex adapter personaplex/ (production glue only) vllm_omni/model_executor/models/personaplex/duplex/ Serving, Orchestrator , Worker
Small shared pieces output.py, model_executor.py, request_client.py vllm_omni/outputs/duplex.py, vllm_omni/model_executor/duplex_sampling.py, vllm_omni/entrypoints/duplex_request_client.py Serving & Orchestrator, Worker, Serving

Existing stable modules that need touching — in every case the change is an import-path or config-string rewrite, not a logic change:

Module Change
vllm_omni/engine/orchestrator.py duplex kernel imports (control plane, session, lease, messages, output tagging) point at engine/duplex/
vllm_omni/engine/async_omni_engine.py control-client / lease / runtime-loader imports
vllm_omni/entrypoints/async_omni.py duplex request-client and lease/message imports
vllm_omni/entrypoints/openai/api_server.py duplex session-handler import
vllm_omni/worker/gpu_ar_model_runner.py duplex sampling-hook import
vllm_omni/model_executor/stage_input_processors/minicpmo_4_5_omni.py stage-transfer schema and PCM-framing imports
vllm_omni/model_executor/models/minicpmo_4_5/ (pipeline + model files) duplex_runtime_extension / duplex_serving_adapter dotted strings retargeted; model-file imports updated
vllm_omni/model_executor/models/personaplex/ (pipeline + talker) same dotted-string retarget; the talker's lazy Stage 0 runtime import
.buildkite/cuda/test-ready.yml, test-merge.yml duplex source-filter paths replaced
MiniCPM demo/e2e scripts (examples/, tests/e2e/online_serving/) demo WS client import path

Points worth calling out:

  • The existing plugin boundary is kept as-is: model adapters are still loaded only through the duplex_runtime_extension / duplex_serving_adapter dotted-string paths in each model's pipeline config. The generic serving and engine code never imports model code — this is enforced by import-boundary tests that move along with everything else.
  • PersonaPlex's demo tier is dropped rather than moved: the single-process frame engine and the standalone Moshi-web-client server duplicate what the unified /v1/duplex stack already does in production. Only the 8 production glue files graduate.
  • Tests are relocated to mirror the new layout, and the Buildkite path filters are updated accordingly. The relocated unit tests are all core_model and cpu, so they keep running in the existing CPU blocks without CI command changes.
  • The two design documents move to docs/design/ with paths updated.

A working branch with the full move exists and I am happy to open the PR(s) once there is agreement on the target layout.

New DuplexClient API — public interface description

Add DuplexClient python API for the /v1/realtime?duplex=1
WebSocket endpoint. The package is client-side only and is never imported by
runtime/serving code (enforced by tests/engine/test_duplex_import_boundary.py
and the AST guard in tests/entrypoints/openai_api/test_duplex_capability.py).

vllm_omni.clients.duplex.DuplexClient

class DuplexClient:
    """Async client for one duplex session over ``/v1/realtime?duplex=1``.

    Example:
        from vllm_omni.clients.duplex import DuplexClient, audio_data_url, read_pcm16_wav
        from vllm_omni.clients.minicpmo_4_5 import create_duplex_session_config

        cfg: SessionConfig = create_duplex_session_config(ref_audio=audio_data_url(ref_wav_path))
        async with DuplexClient("ws://localhost:8099/v1/realtime",
                                model="openbmb/MiniCPM-o-4_5", config=cfg) as client:
            await client.stream_pcm(read_pcm16_wav(question_wav))  # paced realtime chunks
            await client.commit(create_response=False)             # model decides listen/speak
            async for response in client.responses():
                if response.decision == "listen":
                    continue                                       # model kept listening
                async for chunk in response.audio():               # 24 kHz PCM16 deltas
                    play(chunk)
                await client.ack_playback(response.played_ms,
                                          response_id=response.response_id)
                break
    """

    def __init__(
        self,
        url: str,                                   # ws://host:port/v1/realtime (duplex=1 added automatically)
        *,
        model: str,                                 # served model name
        config: SessionConfig | None = None,        # session payload; defaults to SessionConfig()
        session_id: str | None = None,              # explicit session id (resume/takeover)
        reconnect: ReconnectPolicy | None = ReconnectPolicy(),  # None disables auto-resume
        heartbeat_interval_s: float | None = 30.0,  # None disables heartbeats
        handshake_timeout_s: float = 30.0,
        connect: ConnectFn | None = None,           # custom transport factory (tests/benchmarks)
    ) -> None: ...

    # Populated after the handshake:
    session_id: str | None
    session_info: dict[str, object]
    incarnation: int
    resume_token: str | None

    # -- lifecycle -----------------------------------------------------------

    async def __aenter__(self) -> DuplexClient:
        """Connect, send ``session.update``, wait for ``session.created``."""

    async def __aexit__(self, exc_type, exc, traceback) -> None:
        """Graceful ``session.close`` on clean exit (best-effort close on the
        error path so the server frees the session slot), then teardown."""

    async def close(self, *, timeout_s: float = 20.0) -> None:
        """Send ``session.close`` and wait for the server to confirm."""

    # -- input ----------------------------------------------------------------

    async def append_audio(
        self,
        pcm: bytes,
        *,
        is_speech: bool | None = None,
        video_frames: Sequence[str] | None = None,  # base64 JPEG/PNG, one per ~1 s unit
    ) -> None:
        """Append one chunk of input audio in the session's input format.

        A turn is ended with ``commit``, never by a flag on the append.
        ``video_frames`` entries are bare base64; data URLs are accepted and
        their prefix stripped (the wire contract carries bare base64)."""

    async def stream_pcm(
        self,
        pcm: bytes,
        *,
        chunk_ms: int = 200,
        realtime: bool = True,
        is_speech: bool | None = None,
        video_frames: Sequence[str] | None = None,        # one base frame per model unit
        stacked_video_frames: Sequence[str | None] | None = None,  # per-unit composites
    ) -> int:
        """Slice ``pcm`` into ``chunk_ms`` chunks, pace, append; interleave
        camera frames on the appends that close model units. Returns the
        number of base frames actually sent."""

    async def commit(self, *, final: bool = True, create_response: bool | None = None) -> None:
        """``input_audio_buffer.commit``; ``create_response=False`` leaves the
        listen/speak decision to the model (native duplex)."""

    async def ack_playback(
        self,
        played_ms: float,
        *,
        response_id: str | None = None,
        item_id: str | None = None,
        committed_ms: float | None = None,
    ) -> None:
        """Report cumulative playback progress; call periodically while playing."""

    async def cancel_response(self, response_id: str | None = None) -> None:
        """Cancel the active (or a specific) response (client-forced interruption)."""

    async def clear_input(self) -> None:
        """Drop un-committed input audio; composes with ``cancel_response``
        for client-forced barge-in on models that support it."""

    async def send(self, event: dict[str, object]) -> str:
        """Escape hatch: send one raw client event; returns the stamped event_id."""

    # -- output ----------------------------------------------------------------

    def __aiter__(self) -> AsyncIterator[DuplexEvent]: ...

    async def events(self) -> AsyncIterator[DuplexEvent]:
        """Iterate over every server event from now on (typed, multi-subscriber).
        Delivery is non-blocking with drop-oldest backpressure: a consumer
        that stops draining loses its oldest buffered events instead of
        stalling the client's reader."""

    async def responses(self) -> AsyncIterator[ResponseHandle]:
        """Iterate over response lifecycles (single consumer). A handle is
        yielded once its decision is knowable — at its first output, or
        already finished by a terminal listen (``decision == "listen"``,
        no audio). Raises ``DuplexProtocolError`` when the server rejects a
        request with an ``error`` event while waiting."""

    async def wait_for(self, *types: str, timeout_s: float) -> DuplexEvent:
        """Wait for the next event whose type is in ``types``."""

And add 2 utility functions for creating session config:

vllm_omni.clients.minicpmo_4_5.create_duplex_session_config
vllm_omni.clients.personaplex.create_duplex_session_config

Benchmarking Update

  • Use the new DuplexClient instand of the experimental RealtimeDupleClient
  • Fix the playback ack, make it acknowledge the audio played continuously:

The benchmark's playback reporting did not behave like a real listener: it
acknowledged each response's audio once, after the fact, instead of
reporting progress continuously while the audio played. In multi-turn
duplex sessions the conversation moves on while earlier audio is still
playing, so an after-the-fact acknowledgement can arrive after a later
user turn has been committed — which the server correctly refuses,
because accepting it could reorder conversation history. A benchmark must
exercise the protocol the way a real client does, so playback
acknowledgement was changed to report cumulative progress live,
checkpointing each response as soon as its audio starts. This makes
multi-turn cases complete under the server's history-ordering contract
instead of tripping over it, and measures the serving stack under
realistic client behavior.

Test Plan

vLLM Version: 0.28.0

1. Unittests

pytest -q \
  tests/engine/duplex/ \
  tests/entrypoints/duplex/ \
  tests/model_executor/models/minicpmo_4_5/duplex/ \
  tests/model_executor/models/personaplex/ \
  tests/engine/test_duplex_import_boundary.py \
  tests/entrypoints/openai_api/test_duplex_capability.py \
  tests/entrypoints/openai_api/test_duplex_handler.py \
  tests/entrypoints/openai/test_duplex_protocol.py \
  tests/entrypoints/openai/test_duplex_session_attachment.py \
  tests/entrypoints/test_async_omni_duplex.py \
  tests/entrypoints/test_duplex_fence_propagation.py \
  tests/worker/test_native_duplex_hooks.py \
  tests/engine/test_orchestrator.py \
  tests/engine/test_orchestrator_stage_input_bridge.py \
  tests/engine/test_async_omni_engine_outputs.py \
  tests/model_executor/models/minicpmo_4_5/test_pipeline.py \
  tests/benchmarks/patch/test_patch.py \
  tests/e2e/features/fullduplex/ \
  tests/clients/ \
  tests/model_executor/models/nemotron_voicechat/ \
  tests/model_executor/models/test_nemotron_voicechat_registration.py \
  tests/benchmarks/test_omniinteract.py \
  tests/benchmarks/duplex/ \
  tests/e2e/features/fullduplex/test_omni_duplex_eval_cli.py \
  tests/config/test_config_factory.py \
  tests/examples/test_minicpmo_realtime_duplex_simple_demo.py \
  tests/examples/test_minicpmo_realtime_web_static.py \
  tests/e2e/online_serving/test_minicpmo_realtime_duplex_drivers.py

2. MiniCPM-o 4.5 barge in test

sequenceDiagram
    autonumber
    participant U as user (client script)
    participant A as assistant (server + model)

    Note over U,A: GREET
    U->>A: start a voice session (with the reference voice)
    A-->>U: ready

    Note over U,A: ASK
    U->>A: ask a question (speak the question WAV)
    Note right of A: decides on its own to answer<br/>(no VAD — the model chooses)
    A-->>U: starts answering aloud

    Note over U,A: INTERRUPT
    Note over U: listens for ~2 seconds
    par both talk at once
        U->>A: interrupts with a follow-up
    and
        A-->>U: still speaking the first answer
    end

    alt the interruption is substantial
        Note right of A: stops mid-sentence (barge-in),<br/>first answer is cut off
    else it was just a short remark
        Note right of A: finishes the first answer,<br/>then takes the follow-up
    end

    Note over U,A: ANSWER AGAIN
    A-->>U: answers the follow-up aloud

    Note over U,A: HANG UP
    U->>A: goodbye (close the session)
    Note over U: saves one recording per answer<br/>plus a summary of what happened
Loading

start server

# defaults: MODEL=openbmb/MiniCPM-o-4_5, PORT=8099 — override via env
PORT=8098 examples/online_serving/minicpmo/barge_in_serve.sh
# which execs:
#   vllm-omni serve openbmb/MiniCPM-o-4_5 --omni \
#       --deploy-config vllm_omni/deploy/minicpmo_4_5.yaml \
#       --trust-remote-code --host 0.0.0.0 --port 8098

client

python examples/online_serving/barge_in_client.py \
    --url ws://127.0.0.1:8098/v1/realtime \
    --model openbmb/MiniCPM-o-4_5 \
    --ref-audio "$HF_HOME/hub/models--openbmb--MiniCPM-o-4_5/snapshots/<rev>/assets/HT_ref_audio.wav" \
    --question-wav tests/assets/minicpmo_4_5/response_required_16k.wav \
    --interrupt-wav tests/assets/minicpmo_4_5/soft_interrupt_16k.wav \
    --output-dir ./barge_in_out

Test Result

1. Unittests

At tip dae63b7a (after merging upstream main 376b28ac), on the vLLM
0.28.0 venv (torch 2.13.0+cu130), 4× L20X host:

  • duplex sweep (one pytest session): 768 passed, 0 failed (8m0s)
  • additional suites (nemotron, omniinteract, duplex benchmarks, drivers,
    demos, config factory): 518 passed, 0 failed
  • ruff check / ruff format --check, forbidden-imports, markdownlint,
    and typos hooks clean on every changed file

Live e2e (validated at 03f8d8a5, 1× L20X per run):

  • MiniCPM-o duplex core tier: 1 passed (test_duplex_websocket_protocol_smoke)
  • MiniCPM-o duplex advanced tier: 3 passed (response_required,
    video_input, resume_and_takeover)
  • live public-client probe test_duplex_client_live.py (handshake, speak
    response via ResponseHandle, playback acks, forced transport drop with
    automatic session.resume, 5 s heartbeat, clean close): 1 passed

2. MiniCPM-o 4.5 barge in test

barge_in_test_files.zip

PASS

barge-in flow:

  handshake:  input_ref_audio.wav        (voice sample only)
  user:       input_question_16k.wav  →  response_1_completed.wav
  user:       input_interrupt_16k.wav   (starts DURING response_1 = barge-in)
                                      →  response_2_completed.wav  (四大发明)
                                      (soft interrupted in the middle of  input_interrupt_16k, new question: 1+1 = ?)
                                      →  response_3_completed.wav  (好的)
                                      →  response_4_completed.wav  (1+1=2)

Inputs (what the client streamed):

Outputs (what the server produced):

  • response_1_completed.wav … response_4_completed.wav — the four model responses (response 4 is the follow-up answer, "1加1等于 2。")
  • summary.json — per-response text, audio duration, and status
  • events.jsonl — the full WebSocket wire log of the run

Duplex E2E benchmark tables (PR #6522 format)

OmniInteract realtime benchmark, one deterministic case per subset,
--max-concurrency 1 --num-warmups 0, exactly the setup of PR #6522's
"E2E Benchmark Result" section. Both runs: MiniCPM-o 4.5
(vllm_omni/deploy/minicpmo_4_5.yaml), one NVIDIA L20X on test server 1,
.venv-fd3-v028 (vLLM 0.28.0), dataset lucky-lance/OmniInteract
(pre-extracted), ref audio MiniCPM-o-4_5/assets/HT_ref_audio.wav,
2026-09-03. The per-subset deterministic selection drew the same cases in
both runs (1q1a/0028.mp4, 1q1a_math/0013.mp4,
1qna/Blender_Banana_Pancakes-21_44.mp4), so the tables are directly
comparable.

fullduplex3 @ dae63b7a

Subset Requests E2EL (ms) TTFT (ms) TPOT (ms) Output tok/s Audio TTFP (ms) Audio RTF Generated audio (s) Official manifest
1q1a 1/1 310070.51 0.04 10.24 1.43 0.02 0.56 87.28 included
1q1a_math 1/1 304958.92 0.04 10.37 2.25 0.02 1.10 129.44 included
1qna 1/1 305081.22 0.04 9.00 0.60 0.02 0.98 38.76 included

All three runs completed with Successful requests: 1,
Failed requests: 0, streaming continuity OK rate 100 %, and a per-case
.done artifact; every case is included in official_eval_manifest.jsonl.
The benchmark runner on this branch drives the public
vllm_omni.clients.duplex.DuplexClient.

upstream main @ d1a5a7c1

Subset Requests E2EL (ms) TTFT (ms) TPOT (ms) Output tok/s Audio TTFP (ms) Audio RTF Generated audio (s) Official manifest
1q1a 1/1 309531.21 0.28 10.30 1.42 0.17 0.57 87.28 included
1q1a_math 1/1 304918.80 0.48 10.23 2.25 0.14 1.12 129.44 included
1qna 1/1 305175.98 0.34 9.01 0.60 0.17 0.98 38.76 included

All three runs completed with Successful requests: 1,
Failed requests: 0, streaming continuity OK rate 100 %, and a per-case
.done artifact; every case is included in official_eval_manifest.jsonl.
This run used the upstream benchmark runner on its legacy probe client
(the venv's editable install was re-pointed at the upstream checkout for
the whole pass and verified via the import path, then restored).

Why the ~8× difference in TTFT and Audio TTFP ?: Due to implementation difference in measurement, timestamping location. Main's legacy probe client stamps each frame inside its WebSocket recv loop — approximating true wire inter-arrival spacing (~0.1–0.3 ms). Fullduplex3's benchmark stamps in EventCollector.add when the consumer task drains the subscriber queue — when both frames are already queued, they get stamped back-to-back in one wakeup (~20–50 µs), so the queue hop compresses the gap. Fullduplex3 under-measures an interval main measures more faithfully; neither interval carries model-latency information.

Conclusion: There is no performance difference by this PR

Pre-check report for fullduplex3 @ dae63b7a (PR #6196) — re-run

Run 2026-09-03 (second pass, after the title/description fixes were applied
on GitHub) with the repo's precheck-pr skill, full mode, diffed
against the upstream-main merge-base 376b28ac (178 files). The branch tip
is unchanged since the first pass, so the code dimensions were re-swept and
match; the GitHub-side blockers were re-verified against the live PR.
Nothing was posted to GitHub; this report is local only.

  • Mode: full
  • Type: general ∪ new-model (relocated model adapters) — performance
    claims present (benchmark tables in the PR body)

Cleared since the first pass

Previous blocker Status now
PR title missing [Prefix] ✓ fixed — title is now [Core] Graduate MiniCPM-o 4.5 and PersonaPlex full-duplex serving out of experimental
Stale/missing description content ✓ fixed — verified against the live body: the DuplexClient snippet is current (no append_audio(final=…), no "data-URL JPEGs"; stacked_video_frames and the responses() error contract present), the Behavior notes section exists with both the duplex_output_decision key rename and the native_duplex wire-name generalization, the Test Plan/Result read vLLM 0.28.0 / 768 passed / minicpmo_4_5.yaml (no 0.27.0 / 671 / deleted-yaml references), and the embedded 2026-08-17 pre-check block is gone

Dimension results

Dimension Result
PR title format ✓ [Core] prefix, no WIP/Draft; commit subjects all correctly prefixed
Code quality ⚠ 6 new broad-except sites, all in vllm_omni/clients/duplex.py transport/cleanup boundaries: 3 re-raise as typed errors (send, _default_connect, __aenter__), 1 is the read-loop's resume boundary, 2 are best-effort cleanup swallows annotated # noqa: BLE001. None on fail-fast paths — deliberate shapes, discussed in review. 0 kwargs string-lookup plumbing in production, 0 Any hints, 0 hot-path clone/deepcopy, 0 event-loop blocking, 0 new torch.cuda call sites
Examples policy ⚠ 1 new Python example, examples/online_serving/barge_in_client.py — model-neutral entrypoint (model via --model, behavior via vllm_omni.clients presets), permitted by the policy; advisory: the per-preset glue in-script (_PRESET_DEFAULT_CHUNK_MS, _convert_to_session_format) would sit better on the preset modules
Simplification ⚠ 1 candidate: omniinteract._RealtimeSession and patch._RealtimeTTSProbe are two thin DuplexClient+EventCollector wrappers with overlapping shape — a shared probe-session helper in vllm_omni.clients would remove the duplication (advisory follow-up)
PR desc integrity ✓ description matches the diff: current API snippet, both behavior notes, current test numbers, benchmark tables with commit ids for both code bases
Registry/config ✓ no registry changes; duplex plugin dotted paths resolve (config-factory, registration, and pipeline suites green)
Pre-commit gates ✓ over the full diff on the test server: check-spdx-header, check-mark, check-torch-cuda-call, check-forbidden-imports, check-buildkite all Passed; ruff check/format, markdownlint-cli2, typos clean; DCO sign-off on all commits; no allowed_files / MAX_MODEL_TYPE_BRANCHES growth. ⚠ mypy-3.10 (a manual-stage hook, not a default gate) remains red across the relocated duplex serving files — the same untyped legacy patterns that moved; mypy's exclude list only covers model trees, so these files were equally red at their old experimental paths
Dead code ✓ all recent additions referenced; function-body imports are the documented, boundary-test-enforced CLI lazy imports; ruff/isort clean
Rebase/mergeable ✓ merge-base 376b28ac (upstream main has since moved 6 commits, none touching the duplex path); GitHub reports mergeable: true. ⚠ mergeable_state: unstable — status checks need a fresh CI run on head dae63b7a (matches the reviewer's "re-run CI" note; not something the contributor fixes in the diff)
Accuracy ✓ behavioral barge-in test with artifacts in the PR body; live e2e core 1/1 + advanced 3/3 + public-client probe 1/1; benchmark outputs validated (transcripts, WAVs, official-eval manifest)
Benchmark ✓ PR-#6522-format tables in the body with hardware, software versions, commit ids for both code bases, and an equal-performance conclusion vs main

Verdict

0 blocking | 4 warnings — ready for review.

All four warnings are advisory, none require pre-merge code changes:

  1. The broad-except shapes in the client are deliberate transport/cleanup
    boundaries already discussed in review.
  2. Moving the barge-in example's per-preset glue onto the preset modules is
    a reasonable follow-up.
  3. Consolidating the two DuplexClient probe wrappers
    (_RealtimeSession / _RealtimeTTSProbe) is a reasonable follow-up.
  4. Manual-stage mypy redness on the relocated files is the repo's existing
    ratchet posture, inherited by relocation.

The only remaining action item is external: trigger a fresh CI run on head
dae63b7a to clear mergeable_state: unstable.

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@chickeyton chickeyton changed the title Moving MiniCPM-o 4.5 and PersonaPlex out of Experimental [WIP]Moving MiniCPM-o 4.5 and PersonaPlex out of Experimental Aug 14, 2026
@chickeyton

Copy link
Copy Markdown
Contributor Author

@Sy0307 Please review, see if that compactible with nemo-voicechat-11B

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to be related to model: minicpm.

Model owners: @y-null

@chickeyton, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@chickeyton
chickeyton force-pushed the fullduplex3 branch 2 times, most recently from fd52837 to 52cacaf Compare August 14, 2026 09:31

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Ran the relocated suites and the duplex e2e on an H200-class GPU server (vLLM 0.27.0, torch 2.13.0+cu130, Python 3.12) against 52cacaf07.

Batch Result
Relocated duplex unit batch (18 suites) 622 passed, 0 failed, on both vLLM 0.26.0 and 0.27.0
tests/e2e/online_serving/test_minicpmo_4_5_duplex.py 3 passed, 308s, -m cuda --run-level advanced_model
tests/e2e/online_serving/personaplex_realtime_duplex.py passed, errors: []

The MiniCPM run covers all three tests rather than only the core_model one CI selects. The PersonaPlex driver returned two concurrent sessions with identical capability payloads, resource_exhausted on the third, a recycled slot, and non-silent 24 kHz mono output on all three sessions (7.28 / 14.88 / 7.28 s, RMS around 0.03, frame coverage 0.968 to 0.989).

One environment note that is not caused by this PR: on vLLM 0.26.0 every staged server dies at startup with TypeError: GPUModelRunner.maybe_randomize_inputs() got an unexpected keyword argument 'randomize_inputs'. I reproduced that byte for byte on a clean main worktree at the merge base, so it is a version floor rather than a regression here. Stage 2 also needs s3tokenizer and step-audio2 / hyperpyyaml, which live in the dev extra rather than requirements/cuda.txt.

The relocation itself checks out. All 5,693 vllm_omni.* import statements at this head resolve statically, no experimental.fullduplex production references remain (including in the tree merged against current main), all four plugin dotted strings resolve, and the lazy to eager import change introduces no cycles.

Findings below, three of them inline.

[Medium] Module design metadata not updated by the move.

docs/design/module/engine_orchestration.md:46 still lists tests/e2e/features/fullduplex/engine/** in validation_paths, which this PR empties. Neither that page nor docs/design/module/entrypoints.md claims the new locations: vllm_omni/engine/duplex/**, vllm_omni/entrypoints/duplex/**, vllm_omni/entrypoints/duplex_request_client.py. Worth noting entrypoints/openai/duplex_capability.py was covered by the vllm_omni/entrypoints/openai/** glob before the move and now sits outside every declared path. Everything else lands inside an existing glob, so this is the full list.

[Low] Two public symbols removed.

DuplexCapabilities.minicpmo45_native() becomes the module function minicpmo45_native_capabilities(), and entrypoints/duplex/serving.py drops _minicpmo_session_state / _minicpmo_sessions / _minicpmo_data_plane, which carried a comment describing them as accessors for downstream tests and extensions. Both are good changes and no in-repo caller is left. Flagging only because a caller that follows the stated relocation contract and updates just the import path gets an AttributeError.

[Low] Title. [WIP] is not one of the prefixes listed in docs/contributing/README.md, and the PR title becomes the squash headline on main. Both commits already carry [Core] subjects.

On the scope claim.

The description says a pure relocation with no behavior change, but the PersonaPlex demo tier is deleted rather than moved: the standalone server, the elastic batching from #4771, the offline example, the headless client, and run_server.sh. The README and recipe do say so, so this is a description accuracy point rather than a silent removal. Two things would help reviewers who only read the body: state the deletion there, and say what replaces the --batch-size concurrency, since vllm_omni/deploy/personaplex.yaml ships max_sessions: 2 and the unified path admits one scheduler request per session rather than advancing B slots on a shared 80 ms tick.

Separately, both PersonaPlex docs this PR rewrites still tell users to run --stage-configs-path, which the CLI now rejects (tests/entrypoints/test_serve.py:60 asserts the rejection). --deploy-config is the working flag and is what I used. That is repo-wide debt across several recipes rather than something introduced here, but this PR removes run_server.sh, so the non-working command becomes the only documented way to serve PersonaPlex.

Comment thread vllm_omni/benchmarks/patch/patch.py Outdated
from vllm_omni.experimental.fullduplex.client import (
from vllm_omni.metrics import definitions as defs
from vllm_omni.metrics.utils import coerce_positive_int_scalar
from vllm_omni.model_executor.models.minicpmo_4_5.duplex.client import (

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[High] This is generic benchmark plumbing importing a model-private module.

client.py is a generic Realtime WebSocket client with three non-MiniCPM consumers: this file, tests/e2e/online_serving/personaplex_realtime_duplex.py:21, and the MiniCPM demos. Filing it under models/minicpmo_4_5/duplex/ means importing it executes vllm_omni/model_executor/models/minicpmo_4_5/duplex/__init__.py, which eagerly loads adapter, input, policy, and stage0. So vllm-omni bench serve (vllm_omni/benchmarks/serve.py:11 -> vllm_omni/entrypoints/cli/benchmark/serve.py:13) now pulls in the MiniCPM duplex serving adapter and Stage-0 runtime, and a PersonaPlex-only validation path depends on MiniCPM's package.

vllm_omni/entrypoints/duplex/client.py would sit next to the existing consumers and keep the plugin boundary this PR is establishing.

@chickeyton chickeyton Aug 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Resolved


forbidden_prefixes = (
"vllm_omni.model_executor.models.minicpmo_4_5.duplex",
"vllm_omni.model_executor.models.minicpmo_4_5.duplex.client",

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[High] The probe above imports only vllm_omni.engine.async_omni_engine, vllm_omni.engine.orchestrator, and vllm_omni.entrypoints.async_omni, so this forbidden list does not cover vllm_omni/benchmarks/patch/patch.py:61, which imports vllm_omni.model_executor.models.minicpmo_4_5.duplex.client unconditionally. The invariant reads as enforced, but the one shipped module that violates it is outside the scope of the check.

This entry is also redundant: the match is name == prefix or name.startswith(prefix + "."), so line 51 already covers ...duplex.client.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Resolved

from vllm_omni.outputs import OmniRequestOutput

DUPLEX_OUTPUT_DECISION_KEY = "_vllm_omni.experimental.fullduplex.duplex_output_decision"
DUPLEX_OUTPUT_DECISION_KEY = "_vllm_omni.duplex.duplex_output_decision"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Low] The value changed here, from _vllm_omni.experimental.fullduplex.duplex_output_decision.

The writer (vllm_omni/engine/orchestrator.py:1622) and the readers (vllm_omni/entrypoints/duplex_request_client.py:359, vllm_omni/model_executor/models/minicpmo_4_5/duplex/data_plane.py:151) all move together, so nothing breaks in-tree. Calling it out because this key rides in OmniRequestOutput._custom_output across the engine to client boundary, which makes it a value change under a "no behavior change" description.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

the rename is intentional, and you're right that it deserved a callout under the PR description rather than riding silently. I've added a behavior note to the description.

Rationale for keeping the new value rather than preserving the old string: the key rides in the underscore-private OmniRequestOutput._custom_output, which isn't a public wire contract, and the writer (engine/orchestrator.py) and all readers (entrypoints/duplex_request_client.py, models/minicpmo_4_5/duplex/data_plane.py) move atomically in this PR — there is no supported mixed-version engine/frontend deployment that could observe the old key. Keeping _vllm_omni.experimental.fullduplex.duplex_output_decision would embed a permanent experimental.fullduplex reference in the stable tree, which is exactly what this migration removes.

@hsliuustc0106 hsliuustc0106 added refactor refactoring for better code scalability and quality omni code related to omni models labels Aug 16, 2026
@Gaohan123 Gaohan123 added this to the v0.28.0 milestone Aug 16, 2026
@chickeyton
chickeyton force-pushed the fullduplex3 branch 6 times, most recently from 441d725 to ba2c1af Compare August 17, 2026 08:51
@hsliuustc0106 hsliuustc0106 added the ready label to trigger buildkite CI label Sep 4, 2026
handshake_queue = self._add_subscriber()
try:
self._reader_task = asyncio.create_task(self._read_loop(), name="duplex-client-reader")
await self.send(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[High] session_id alone cannot resume or take over an existing session.

Entering the client always sends session.update, while the server only enters its resume path when the first event is session.resume, which additionally requires incarnation, resume_token, and last_received_server_event_seq. Therefore, constructing a new DuplexClient with a known session_id will attempt to create that session (and may collide with the existing attachment) rather than resume it, contrary to the public API description and realtime_duplex_api.md.

Please either add explicit resume credentials and select the session.resume handshake on initial connection, or narrow the documentation to say that only automatic reconnect within the same DuplexClient instance is supported and that session_id merely chooses the ID for a new session.

@chickeyton chickeyton Sep 5, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right — the entry handshake is always session.update, so a fresh
DuplexClient given a known session_id would try to create that session, not
resume it. Fixed in 458a342 by narrowing the claim on both surfaces:

  • realtime_duplex_api.md now says session_id only names the session the
    client creates, that resume is supported solely as automatic reconnect
    within the same DuplexClient instance (via ReconnectPolicy), and that
    re-attaching from a new client or process requires the wire-level
    session.resume handshake (resume_token, incarnation,
    last_received_server_event_seq), which this client does not expose.
  • The DuplexClient class docstring states the same contract.

and, I will propose to drop incarnation and client side chose session id in the follow up refactor RFC, so the resume logics will be refactored later

for queue in list(self._subscribers):
await queue.put(event)

if isinstance(event, SessionClosed):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please resolve it. @chickeyton

Comment thread vllm_omni/clients/duplex.py Outdated
if request_started_at_s is not None:
# Lazily import so this module stays importable without the
# vllm_omni metrics stack (the client is otherwise standalone).
from vllm_omni.metrics.definitions import compute_audio_rtf

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[Medium] This import breaks the lightweight dependency boundary advertised for the public client.

Importing vllm_omni.metrics.definitions first executes vllm_omni.metrics.__init__, which pulls in the Prometheus/statistics stack and dependencies such as prometheus_client and prettytable. In an environment containing only the documented client dependencies (pybase64 and websockets), the client lifecycle tests pass, but the three tests that call EventCollector.timing_summary() fail on these server-side dependencies.

Since compute_audio_rtf() is only a guarded division of generation time by audio duration, please calculate it locally or move the calculation to a genuinely dependency-free helper so that vllm_omni.clients.duplex remains usable with its documented dependency set.

@chickeyton chickeyton Sep 5, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

vllm_omni.metrics import was removed, now the function outputs audio_generation_ms and audio_duration_ms which are sufficient for computing the metric Audio RTF by the caller with compute_audio_rtf()

chickeyton and others added 4 commits September 5, 2026 11:29
…sume

Review fixes for the DuplexClient library (PR vllm-project#6196, round 24):

- EventCollector.timing_summary no longer imports
  vllm_omni.metrics.definitions for the RTF value: that import executes
  vllm_omni.metrics.__init__, which pulls the Prometheus/statistics stack
  (prometheus_client, prettytable) into an environment that only carries
  the documented client dependencies (pybase64 + websockets). The RTF is
  the same guarded division, now computed locally with a comment pinning
  it to the server-side metric definition. A new import-boundary probe
  exercises timing_summary in a subprocess and fails if importing the
  client package pulls any vllm_omni module outside vllm_omni.clients.

- The constructor's session_id only names the session to create; it
  cannot resume or take over an existing one (the entry handshake is
  always session.update, while the server-side resume path additionally
  needs resume_token, incarnation and last_received_server_event_seq).
  The class docstring and realtime_duplex_api.md now say exactly that:
  resume is automatic reconnect within the same client instance, and
  cross-client takeover requires the wire-level session.resume handshake,
  which this client does not expose.

Signed-off-by: chickeyton <ngton2014@gmail.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mmoj3yYiiBZAdLFDiE4fKp
…dary probe

timing_summary() takes after_s/response_id as keywords; the new client
import-boundary probe passed response_id positionally and failed on
TypeError instead of exercising the RTF path.

Signed-off-by: chickeyton <ngton2014@gmail.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mmoj3yYiiBZAdLFDiE4fKp
… in callers

Follow-up to the round-24 review fix for the client dependency boundary:
instead of duplicating the RTF formula inside the client,
EventCollector.timing_summary now reports raw measurements only
(ttft_ms, ttfp_ms, audio_generation_ms, audio_duration_ms) and leaves
derived metrics to the caller. The benchmarks (omniinteract and the
Seed-TTS realtime path in the benchmark patch) derive the RTF from
those raw fields through a shared helper that calls the canonical
vllm_omni.metrics.definitions.compute_audio_rtf, so the metric has one
definition in the tree and the client keeps its documented dependency
set (pybase64 + websockets). summarize_session_request_metrics is
unchanged: it averages whatever caller-assembled dicts provide, and
its docstring now says rtf is caller-computed.

Signed-off-by: chickeyton <ngton2014@gmail.com>
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Mmoj3yYiiBZAdLFDiE4fKp
@chickeyton
chickeyton requested a review from Gaohan123 September 5, 2026 04:32
@Sy0307 Sy0307 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 5, 2026

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
@Gaohan123 Gaohan123 removed the ready label to trigger buildkite CI label Sep 5, 2026
@Gaohan123 Gaohan123 added the ready label to trigger buildkite CI label Sep 5, 2026
@Gaohan123
Gaohan123 enabled auto-merge (squash) September 5, 2026 07:47
@Gaohan123
Gaohan123 merged commit 5d7fdf9 into vllm-project:main Sep 5, 2026
8 of 9 checks passed
guozhihao-224 added a commit to guozhihao-224/vllm-omni that referenced this pull request Sep 5, 2026
Keep turn-metrics on the graduated duplex_request_client path after vllm-project#6196.

Co-authored-by: Cursor <cursoragent@cursor.com>

Signed-off-by: guozhihao-224 <guozhihaoemail@gmail.com>
JoseCarlosGarcia95 pushed a commit to valendra-tech/vllm-omni that referenced this pull request Sep 5, 2026
… of experimental (vllm-project#6196)

Signed-off-by: chickeyton <ngton2014@gmail.com>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
ZhengWG pushed a commit to ZhengWG/vllm-omni that referenced this pull request Sep 8, 2026
… of experimental (vllm-project#6196)

Signed-off-by: chickeyton <ngton2014@gmail.com>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
Signed-off-by: ZhengWG <zwg0606@gmail.com>
NickCao added a commit to NickCao/vllm-omni that referenced this pull request Sep 9, 2026
The /v1/realtime duplex output projector emitted the deprecated beta
Realtime API names (response.audio.delta, response.audio_transcript.delta,
response.audio.done, response.audio_transcript.done). The official openai
Python SDK's non-beta openai.types.realtime module, and clients built
against it such as livekit-plugins-openai's non-Azure code path, only
recognize the current names (response.output_audio.delta, etc.) -- so an
unmodified modern OpenAI Realtime client connected to any duplex model
(PersonaPlex, MiniCPM-o 4.5, Nemotron VoiceChat) never saw any audio or
transcript output at all.

docs/serving/realtime_duplex_api.md claimed this dialect "deliberately
emits and accepts both the beta and the GA spellings at once" for these
events, alongside the genuinely-dual conversation.item.added/.created.
Checked history: at 5d7fdf9 (vllm-project#6196), which first wrote that claim, the
code already only ever emitted the beta name -- the doc was wrong from
the start, not a later regression.

Switch the four event names to their current spelling and update every
in-repo consumer (the first-party DuplexClient, benchmark/example clients,
and tests) to match, and correct the doc claim. Session-field and
conversation.item.* dual beta/GA spelling support is unaffected -- only
these four output event types were beta-only.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
NickCao added a commit to NickCao/vllm-omni that referenced this pull request Sep 9, 2026
The /v1/realtime duplex output projector emitted the deprecated beta
Realtime API names (response.audio.delta, response.audio_transcript.delta,
response.audio.done, response.audio_transcript.done). The official openai
Python SDK's non-beta openai.types.realtime module, and clients built
against it such as livekit-plugins-openai's non-Azure code path, only
recognize the current names (response.output_audio.delta, etc.) -- so an
unmodified modern OpenAI Realtime client connected to any duplex model
(PersonaPlex, MiniCPM-o 4.5, Nemotron VoiceChat) never saw any audio or
transcript output at all.

docs/serving/realtime_duplex_api.md claimed this dialect "deliberately
emits and accepts both the beta and the GA spellings at once" for these
events, alongside the genuinely-dual conversation.item.added/.created.
Checked history: at 5d7fdf9 (vllm-project#6196), which first wrote that claim, the
code already only ever emitted the beta name -- the doc was wrong from
the start, not a later regression.

Switch the four event names to their current spelling and update every
in-repo consumer: the first-party DuplexClient and its example
(barge_in_client.py), benchmark clients, the MiniCPM demo script and its
browser JS, the server's own duplex warmup probe (api_server.py, which
would otherwise time out on every startup waiting for an event that no
longer arrives), design/serving docs, and tests across the duplex,
PersonaPlex, and MiniCPM suites. Session-field and conversation.item.*
dual beta/GA spelling support is unaffected -- only these four output
event types were beta-only. Left the separate legacy (pre-duplex)
Qwen3-Omni realtime fallback and the unrelated video-stream protocol
untouched; they don't share this code or its OpenAI SDK type contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
NickCao added a commit to NickCao/vllm-omni that referenced this pull request Sep 9, 2026
The /v1/realtime duplex output projector emitted the deprecated beta
Realtime API names (response.audio.delta, response.audio_transcript.delta,
response.audio.done, response.audio_transcript.done). The official openai
Python SDK's non-beta openai.types.realtime module, and clients built
against it such as livekit-plugins-openai's non-Azure code path, only
recognize the current names (response.output_audio.delta, etc.) -- so an
unmodified modern OpenAI Realtime client connected to any duplex model
(PersonaPlex, MiniCPM-o 4.5, Nemotron VoiceChat) never saw any audio or
transcript output at all.

docs/serving/realtime_duplex_api.md claimed this dialect "deliberately
emits and accepts both the beta and the GA spellings at once" for these
events, alongside the genuinely-dual conversation.item.added/.created.
Checked history: at 5d7fdf9 (vllm-project#6196), which first wrote that claim, the
code already only ever emitted the beta name -- the doc was wrong from
the start, not a later regression.

Switch the four event names to their current spelling and update every
in-repo consumer: the first-party DuplexClient and its example
(barge_in_client.py), benchmark clients, the MiniCPM demo script and its
browser JS, the server's own duplex warmup probe (api_server.py, which
would otherwise time out on every startup waiting for an event that no
longer arrives), design/serving docs, and tests across the duplex,
PersonaPlex, and MiniCPM suites. Session-field and conversation.item.*
dual beta/GA spelling support is unaffected -- only these four output
event types were beta-only. Left the separate legacy (pre-duplex)
Qwen3-Omni realtime fallback and the unrelated video-stream protocol
untouched; they don't share this code or its OpenAI SDK type contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Nick Cao <ncao@redhat.com>
NickCao added a commit to NickCao/vllm-omni that referenced this pull request Sep 9, 2026
The /v1/realtime duplex output projector emitted the deprecated beta
Realtime API names (response.audio.delta, response.audio_transcript.delta,
response.audio.done, response.audio_transcript.done). The official openai
Python SDK's non-beta openai.types.realtime module, and clients built
against it such as livekit-plugins-openai's non-Azure code path, only
recognize the current names (response.output_audio.delta, etc.) -- so an
unmodified modern OpenAI Realtime client connected to any duplex model
(PersonaPlex, MiniCPM-o 4.5, Nemotron VoiceChat) never saw any audio or
transcript output at all.

docs/serving/realtime_duplex_api.md claimed this dialect "deliberately
emits and accepts both the beta and the GA spellings at once" for these
events, alongside the genuinely-dual conversation.item.added/.created.
Checked history: at 5d7fdf9 (vllm-project#6196), which first wrote that claim, the
code already only ever emitted the beta name -- the doc was wrong from
the start, not a later regression.

Switch the four event names to their current spelling and update every
in-repo consumer: the first-party DuplexClient and its example
(barge_in_client.py), benchmark clients, the MiniCPM demo script and its
browser JS, the server's own duplex warmup probe (api_server.py, which
would otherwise time out on every startup waiting for an event that no
longer arrives), design/serving docs, and tests across the duplex,
PersonaPlex, and MiniCPM suites. Session-field and conversation.item.*
dual beta/GA spelling support is unaffected -- only these four output
event types were beta-only. Left the separate legacy (pre-duplex)
Qwen3-Omni realtime fallback and the unrelated video-stream protocol
untouched; they don't share this code or its OpenAI SDK type contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Nick Cao <ncao@redhat.com>
NickCao added a commit to NickCao/vllm-omni that referenced this pull request Sep 9, 2026
The /v1/realtime duplex output projector emitted the deprecated beta
Realtime API names (response.audio.delta, response.audio_transcript.delta,
response.audio.done, response.audio_transcript.done). The official openai
Python SDK's non-beta openai.types.realtime module, and clients built
against it such as livekit-plugins-openai's non-Azure code path, only
recognize the current names (response.output_audio.delta, etc.) -- so an
unmodified modern OpenAI Realtime client connected to any duplex model
(PersonaPlex, MiniCPM-o 4.5, Nemotron VoiceChat) never saw any audio or
transcript output at all.

docs/serving/realtime_duplex_api.md claimed this dialect "deliberately
emits and accepts both the beta and the GA spellings at once" for these
events, alongside the genuinely-dual conversation.item.added/.created.
Checked history: at 5d7fdf9 (vllm-project#6196), which first wrote that claim, the
code already only ever emitted the beta name -- the doc was wrong from
the start, not a later regression.

Switch the four event names to their current spelling and update every
in-repo consumer: the first-party DuplexClient and its example
(barge_in_client.py), benchmark clients, the MiniCPM demo script and its
browser JS, the server's own duplex warmup probe (api_server.py, which
would otherwise time out on every startup waiting for an event that no
longer arrives), design/serving docs, and tests across the duplex,
PersonaPlex, and MiniCPM suites. Session-field and conversation.item.*
dual beta/GA spelling support is unaffected -- only these four output
event types were beta-only. Left the separate legacy (pre-duplex)
Qwen3-Omni realtime fallback and the unrelated video-stream protocol
untouched; they don't share this code or its OpenAI SDK type contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Nick Cao <ncao@redhat.com>
NickCao added a commit to NickCao/vllm-omni that referenced this pull request Sep 10, 2026
The /v1/realtime duplex output projector emitted the deprecated beta
Realtime API names (response.audio.delta, response.audio_transcript.delta,
response.audio.done, response.audio_transcript.done). The official openai
Python SDK's non-beta openai.types.realtime module, and clients built
against it such as livekit-plugins-openai's non-Azure code path, only
recognize the current names (response.output_audio.delta, etc.) -- so an
unmodified modern OpenAI Realtime client connected to any duplex model
(PersonaPlex, MiniCPM-o 4.5, Nemotron VoiceChat) never saw any audio or
transcript output at all.

docs/serving/realtime_duplex_api.md claimed this dialect "deliberately
emits and accepts both the beta and the GA spellings at once" for these
events, alongside the genuinely-dual conversation.item.added/.created.
Checked history: at 5d7fdf9 (vllm-project#6196), which first wrote that claim, the
code already only ever emitted the beta name -- the doc was wrong from
the start, not a later regression.

Switch the four event names to their current spelling and update every
in-repo consumer: the first-party DuplexClient and its example
(barge_in_client.py), benchmark clients, the MiniCPM demo script and its
browser JS, the server's own duplex warmup probe (api_server.py, which
would otherwise time out on every startup waiting for an event that no
longer arrives), design/serving docs, and tests across the duplex,
PersonaPlex, and MiniCPM suites. Session-field and conversation.item.*
dual beta/GA spelling support is unaffected -- only these four output
event types were beta-only. Left the separate legacy (pre-duplex)
Qwen3-Omni realtime fallback and the unrelated video-stream protocol
untouched; they don't share this code or its OpenAI SDK type contract.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Signed-off-by: Nick Cao <ncao@redhat.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
… of experimental (vllm-project#6196)

Signed-off-by: chickeyton <ngton2014@gmail.com>
Signed-off-by: Gao Han <hgaoaf@connect.ust.hk>
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Co-authored-by: Gao Han <hgaoaf@connect.ust.hk>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

frontend code related to entrypoint high priority high priority issue, needs to be done asap omni code related to omni models ready label to trigger buildkite CI refactor refactoring for better code scalability and quality

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants