Skip to content

[Core][Frontend] Unified Full-duplex Framework - #7413

Merged
Gaohan123 merged 96 commits into
vllm-project:mainfrom
chickeyton:duplex_refactor2
Sep 16, 2026
Merged

Gaohan123 merged 96 commits into
vllm-project:mainfrom
chickeyton:duplex_refactor2

Conversation

@chickeyton

@chickeyton chickeyton commented Sep 11, 2026 •

Copy link
Copy Markdown
Contributor

PR1 of RFC #7181

Purpose

Changes

  1. Make duplex serving layer a thin layer, mirgrate the duplex logics like DuplexSession, and session runner into engine layer
    Seperate the turn based API and engine logics and the duplex API and engine logics by adding DuplexOmni, DuplexOmniEngine and DuplexOrchestrator. Mirgate the duplex API and logics from AsyncOmni,AsyncOmniEngine and Orchestrator to the newly added classes.
  2. Make /v1/duplex an alias of /v1/realtime?duplex=1 and drop the native_duplex parameter in new session message. A duplex model keeps /v1/chat/completions: DuplexOmni extends AsyncOmni, so the ordinary chat service runs on the same engine and the same weights, gated on the model declaring DuplexCapabilities.supports_chat_completions. Only /v1/realtime?duplex=1 (/v1/duplex), /v1/chat/completions, /v1/models and /health are supported by this Full-duplex framework; every other turn-based route reports "not available". This amends RFC decision D9 ("duplex-only API server").
  3. Add InlineDuplexClient, same usage as DuplexClient but relay on DuplexOmni object.
  4. Drop the concept of incarnation, make the session id the new session be always allocated by server, simplify the session creation and session resume logics.
  5. Fit MiniCPM-o 4.5 into the framework

Temporary Changes

  1. Since Qwen3-Omni, PersonalPlex and Nemotron VoiceChat are to be fitted into the Unified Full-duplex Framework in PR5, PR2 and PR4 in the begining RFC, their tests are skipped by adding:
test/model_executor/models/personaplex/duplex/conftest.py
test/model_executor/models/nemotron_voicechat/duplex/conftest.py

the above files are going to be removed after PR2 an PR4

and

test/entrypoints/openai_api/test_qwen3_omni_realtime_websocket.py

Other bugfixes

Three defects this benchmark found, all fixed here

  1. The terminal-event guard rejected our own close.
  2. Stage-0 engine metrics were dropped, so TPOT could not be computed.
  3. 1q1a_math deadlocked, and the model appeared to generate 10x too much audio.

Benefits

  1. A duplex session now has exactly one owner, owned by engine layer, more compact and easly to read code.

  2. The turn-base (e.g. HunyuanImage model) modules no longer carries duplex code, AsyncOmni-AsyncOmniEngine-Orchestrator stack become less bulky.

  3. Simplified usage and documenations by dropping all the unnecessary fallback modes, leaving the two surfaces a duplex session actually backs: v1/realtime?duplex=1 and v1/chat/completions.

  4. Advantages for AURA's intergration (may also true to other duplex models):

  • One in-process session object that can own AURA's history, instead of
    today's registry that is written in the API server process and read again
    in the Stage 1 worker process.
  • Interrupt, cancel, barge-in, resume and multi-session with less to wire
    and debug.
    The baseline duplex framework implements the same behaviours
    and also binds forwarded stage requests to the session, so AURA would
    inherit the same feature list on either tree once it is ported off the
    streaming-video handler (which has none of them). The difference is where
    they live: one handler in the runner reaching all four stages, versus the
    serving mixins, the RPC boundary and the engine control plane with fence
    re-validation on both sides.
  • A single, validated plugin seam (the one DuplexModelPlugin the engine
    loads and hands to the orchestrator) where AURA's turn, silent and tool
    rules live in AURA's own file, instead of the baseline's engine extension
    plus serving adapter pair.

New Architecture

              Python users                     WebSocket clients              HTTP clients
        ┌──────────────────────┐        ┌────────────────────────┐    ┌──────────────────────┐
        │  InlineDuplexClient  │        │      DuplexClient      │    │   OpenAI chat SDK    │
        │  (DuplexClientBase)  │        │   (DuplexClientBase)   │    │ /v1/chat/completions │
        └──────────┬───────────┘        └───────────┬────────────┘    └──────────┬───────────┘
                   │   DuplexSessionHandle          │ /v1/realtime?duplex=1 (JSON)│
                   │                                ▼                             │
                   │              ┌──────────────────────────────────────┐   thin layer
                   │              │ OmniDuplexSessionHandler             │        │
                   │              │  websocket I/O, DuplexCommand.from_  │        ▼
                   │              │  realtime(), event.to_realtime(),    │  OmniOpenAIServingChat
                   │              │  attachment / resume tokens / replay │  (unchanged; gated on
                   │              └──────────────────┬───────────────────┘  supports_chat_completions)
                   ▼                                 ▼   DuplexSessionHandle        │
        ┌──────────────────────────────────────────────────────────────────────────────┐
        │ DuplexOmni(AsyncOmni)                      [entrypoints/duplex_omni.py]       │  thin layer
        │   open_session / get_session / resume_session / close_session                │
        │   DuplexSessionHandle: submit(DuplexCommand) / events()                      │
        │   inherited from AsyncOmni: generate(), EngineClient — the chat service runs  │
        │   on this same object, one engine and one set of weights                     │
        └──────────────────────────────────────┬───────────────────────────────────────┘
                                               │ engine.open/close/resume/touch_session (RPC), submit_command (one-way)
                                               ▼
        ┌──────────────────────────────────────────────────────────────────────────────┐   one single Duplex Session owner
        │ DuplexOmniEngine(AsyncOmniEngine) [engine/duplex_omni_engine.py, thin]        │
        │   message construction + queue put/RPC; creates DuplexOrchestrator            │
        │   (OmniEngineBase: stage init, orchestrator thread, queues, RPC router)       │
        │  └─ DuplexOrchestrator(Orchestrator)            [engine/duplex_orchestrator.py]│
        │       stage forwarding/prewarm/cleanup with session-owned policy              │
        │       _dispatch_message: session messages here, everything else falls through │
        │                          to Orchestrator, so one engine also admits requests  │
        │       ├─ DuplexSessionManager      admission, lease reaper, open/close/resume, │
        │       │                            dispatch of commands to runners            │
        │       └─ DuplexSessionRunner ×N    ONE session state (DuplexEngineSession),   │
        │                                    ordered mailbox, append planning + stage   │
        │                                    submit, data-plane decisions, response /   │
        │                                    playback / history, VAD, typed events      │
        │          [engine/duplex/session/]  runner + manager + the collaborators it     │
        │                                    delegates to: SessionControl, SessionEmitter│
        │                                    ModelChannel, AppendAttempt, helpers        │
        │       plugin: DuplexModelPlugin    (pipeline.duplex_plugin, one class/model)  │
        └──────────────────────────────────────────────────────────────────────────────┘

Class relationships:

OmniBase                        OmniEngineBase                       OrchestratorBase
 ├─ Omni        (sync, turn)     └─ AsyncOmniEngine  (turn requests)  └─ Orchestrator        (turn-based admission)
 └─ AsyncOmniBase                    └─ DuplexOmniEngine (+session)       └─ DuplexOrchestrator  (hosts DuplexSessionManager)
     └─ AsyncOmni   (EngineClient)  _create_engine -> AsyncOmniEngine   _create_orchestrator -> Orchestrator
         └─ DuplexOmni              _create_engine -> DuplexOmniEngine  _create_orchestrator -> DuplexOrchestrator

The duplex classes extend their turn-based counterparts rather than sitting
beside them, so one engine serves a duplex session and an ordinary request at the
same time. The dependency runs duplex -> turn-based and never the reverse, so
Orchestrator, AsyncOmniEngine and AsyncOmni still carry no duplex code and
tests/engine/test_duplex_import_boundary.py still holds. Two tests pin the
relationship itself (test_the_duplex_stack_extends_the_turn_based_one,
test_an_unrecognised_message_reaches_the_turn_based_handler), because a later
tidy-up back into siblings would silently stop a duplex server serving chat.

Verified on hardware rather than argued: a websocket duplex session held open and
mid-turn from 0.07 s to 4.31 s, with two /v1/chat/completions answering inside
that window, on one process and one set of weights.

Test Plan

vLLM Version: 0.29.0 · Hardware: 1× NVIDIA L20X (143 GB) per run, test server 1 · Venvs: .venv-dr2-v029 (CPU sweeps), .venv-fd3-v029 (GPU serving, torch 2.13.0+cu130, cloned from the 0.28.0 GPU venv with vLLM upgraded — the branch needs 0.29.0 since the #7230 rebase)

1. Unittests — refactored duplex framework

pytest -q -m "not gpu" \
  tests/engine/duplex/ \
  tests/entrypoints/duplex/ \
  tests/entrypoints/openai_api/test_duplex_api_server.py \
  tests/entrypoints/openai/test_duplex_session_attachment.py \
  tests/entrypoints/test_duplex_omni.py \
  tests/entrypoints/test_omni_entrypoints.py \
  tests/entrypoints/test_omni_base_profiler.py \
  tests/engine/test_duplex_orchestrator.py \
  tests/engine/test_duplex_omni_engine.py \
  tests/engine/test_duplex_import_boundary.py \
  tests/clients/ \
  tests/worker/test_native_duplex_hooks.py \
  tests/benchmarks/patch/test_patch.py \
  tests/benchmarks/test_omniinteract.py

1b. Everything else under tests/engine, tests/entrypoints, tests/clients

The lists above enumerate paths, so a test this refactor moved out from under
stays unrun. This sweep is what catches that, and it is compared against the
same run on main rather than against zero — the tree has pre-existing
failures that are not this PR's to fix.

pytest -q -m "not gpu and not cuda" tests/engine/ tests/entrypoints/ tests/clients/
# and the same command in a checkout of main, for the baseline

2. Unittests — MiniCPM-o 4.5

pytest -q -m "not gpu and not cuda" \
  tests/model_executor/models/minicpmo_4_5/ \
  tests/e2e/online_serving/test_minicpmo_realtime_duplex_drivers.py

3. MiniCPM-o 4.5 E2E and barge-in test

E2E suite — two CI tiers, each starting its own three-stage single-GPU server:

pytest -q tests/e2e/online_serving/test_minicpmo_4_5_duplex.py \
  -m "core_model and cuda"     --run-level core_model
pytest -q tests/e2e/online_serving/test_minicpmo_4_5_duplex.py \
  -m "advanced_model and cuda" --run-level advanced_model

Barge-in — the model must stop mid-sentence when the user talks over it:

sequenceDiagram
    autonumber
    participant U as user (client script)
    participant A as assistant (server + model)

    Note over U,A: GREET
    U->>A: start a voice session (with the reference voice)
    A-->>U: ready

    Note over U,A: ASK
    U->>A: ask a question (speak the question WAV)
    Note right of A: decides on its own to answer<br/>(no VAD — the model chooses)
    A-->>U: starts answering aloud

    Note over U,A: INTERRUPT
    Note over U: listens for ~2 seconds
    par both talk at once
        U->>A: interrupts with a follow-up
    and
        A-->>U: still speaking the first answer
    end

    alt the interruption is substantial
        Note right of A: stops mid-sentence (barge-in),<br/>first answer is cut off
    else it was just a short remark
        Note right of A: finishes the first answer,<br/>then takes the follow-up
    end

    Note over U,A: ANSWER AGAIN
    A-->>U: answers the follow-up aloud

    Note over U,A: HANG UP
    U->>A: goodbye (close the session)
    Note over U: saves one recording per answer<br/>plus a summary of what happened
Loading

start server

# defaults: MODEL=openbmb/MiniCPM-o-4_5, PORT=8099 — override via env
PORT=8098 examples/online_serving/minicpmo/barge_in_serve.sh
# which execs:
#   vllm-omni serve openbmb/MiniCPM-o-4_5 --omni \
#       --deploy-config vllm_omni/deploy/minicpmo_4_5.yaml \
#       --trust-remote-code --host 0.0.0.0 --port 8098

client

python examples/online_serving/barge_in_client.py \
    --url ws://127.0.0.1:8098/v1/realtime \
    --model openbmb/MiniCPM-o-4_5 \
    --ref-audio "$HF_HOME/hub/models--openbmb--MiniCPM-o-4_5/snapshots/<rev>/assets/HT_ref_audio.wav" \
    --question-wav tests/assets/minicpmo_4_5/response_required_16k.wav \
    --interrupt-wav tests/assets/minicpmo_4_5/soft_interrupt_16k.wav \
    --output-dir ./barge_in_out

This PR amends RFC decision D9: a duplex model now serves /v1/chat/completions
as well, on the same engine, so the same server is probed for the route
alongside the barge-in run:

curl -s -o /dev/null -w "%{http_code}\n" -X POST http://127.0.0.1:8098/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"openbmb/MiniCPM-o-4_5","messages":[{"role":"user","content":"hi"}]}'

4. E2E benchmark with MiniCPM-o 4.5

OmniInteract realtime benchmark, one deterministic case per subset,
--num-prompts 1 --max-concurrency 1 --num-warmups 0 — the setup of PR #6522's
"E2E Benchmark Result" section. Run twice: once on this branch and once on its
merge base main @ 767cc797, same box and venv, so the two tables are a
like-for-like comparison rather than a quote from an older run.

vllm-omni bench serve --omni --trust-remote-code \
  --backend openai-realtime-duplex --dataset-name omniinteract \
  --dataset-path "$OMNIINTERACT_ROOT" --model openbmb/MiniCPM-o-4_5 \
  --base-url http://127.0.0.1:8133 --endpoint /v1/realtime \
  --omniinteract-subsets "$SUBSET" \
  --omniinteract-ref-audio ".../MiniCPM-o-4_5/assets/HT_ref_audio.wav" \
  --omniinteract-output-dir ./artifacts --num-prompts 1 \
  --max-concurrency 1 --num-warmups 0 --save-result --result-dir ./result

Test Result

Head 0bb92c89 (main merged up to 21d86ec9), merge base main @ 21d86ec9. vLLM 0.29.0, 1× NVIDIA
L20X, test server 1, .venv-dr2-v029 (CPU) / .venv-fd3-v029 (GPU serving).

Every number below names the commit it was measured at. Sections 0, 1d and 3
were re-run after CI builds 15206, 15224, 15291, 15306 and 15319; section 4's
benchmark was not re-run after the last merge and is labelled accordingly.

0. What CI will actually run

The lanes .buildkite/cuda/test-merge.yml triggers for this PR's files, run
verbatim at 5d17f455:

Gate lane Command Result
:29 pytest -sv tests/entrypoints tests/engine -m 'core_model and cpu' 2700 passed, 4 skipped, 74 deselected, 1 xfailed (362.25 s)
:35 pytest -sv tests/ -m 'core_model and cpu' --ignore=diffusion,model_executor,entrypoints,engine 24 collection errors, all missing optional packages in this venv (torchmetrics ×18, mistune ×3, openpyxl ×2, ray ×1); none reference vllm_omni
:131 pytest -s -v tests/e2e/online_serving/test_minicpmo_4_5_duplex.py -m 'advanced_model and cuda' --run-level advanced_model 11 passed (462.53 s)
:23 pytest -sv tests/diffusion -m 'core_model and cpu' 5377 passed, 47 skipped at 014350a8 (two torchmetrics files excluded locally; CI has it). The one real failure, an AST source test that looks for _absorb_diffusion_metrics on Orchestrator after this PR moved it to OrchestratorBase, is fixed in c1dff45d (12 passed)
:44 pytest -sv tests/diffusion/lora/ -m 'advanced_model and cuda' 1 passed at 014350a8
ready · duplex pytest -s -v test_minicpmo_4_5_duplex.py test_duplex_client_live.py -m 'core_model and cuda' 2 passed at c737b7c0 (test_duplex_client_live_session 1 passed in 777.68 s at 8ae42c6f; section 3 covers the two failures CI build 15206 hit here)
nightly · duplex pytest -s -v test_minicpmo_4_5_duplex_expansion.py -m 'full_model and cuda' 3 passed at c737b7c0 — admission_and_expiry_reaper, soft_interrupt, server_vad_hard_interrupt (the last two failed in build 15224; section 3)
ready · online (restored) pytest -s -v tests/e2e/online_serving/test_minicpmo_4_5.py -m 'advanced_model and cuda' 4 passed at 1fbe5013 — the /v1/chat/completions suite this PR had deleted, back unchanged from main (section 3)
:69 entrypoint H100 pytest -s -v tests/entrypoints/ -m 'advanced_model and cuda and H100' build 15291: 10 passed, 1 failed — the turn-based server-VAD realtime test, deselected in 0a63ceb9 and deferred to RFC PR 5 (section 0 note)
ready · perf pytest -s -v tests/dfx/perf/scripts/run_benchmark.py --test-config-file tests/dfx/perf/tests/test_minicpmo_4_5_ready.json -m 'core_model and cuda' 1 passed in 745 s at 77f03ddb — Seed-TTS 1×10: 10/10, 4×40: 40/40, 0 failed (build 15306 had 39/40; section 0 note)
ready · perf (3× sample) same config with num_prompts: [120], max_concurrency: [4] 120/120, 0 failed, 173 s at 6f7ebd25; no turn needed more than the usual 12 s of silence (build 15319 had 39/40; section 0 note)

Lane :35 originally reported 25 errors, one of which was this PR's: a stale
DuplexFence import in tests/e2e/features/fullduplex/test_core_contracts.py,
still pointing at engine.duplex.messages after the symbol moved to
engine.duplex.contracts. Being a collection error it interrupted the whole
lane rather than failing one test. Fixed in 5d17f455.

CI build 15206 failed on four things, all fixed since: the DuplexFence
import above (5d17f455), two AsyncEventResolver imports that aborted lane
:23 at collection so it ran nothing (fixed before 0c6e0747), the protocol
smoke close race (117e7e03) and the live-client hang (8ae42c6f); the
diffusion source test (c1dff45d) was found by re-running lane :23.

CI build 15224 (nightly, at 8ae42c6f) had 10 red jobs: 8 were the agent
being lost mid-run (exit −1, logs stop during model load or at 40–81 % of a
perf run), the duplex lane's test_duplex_server_vad_hard_interrupt (fixed in
c737b7c0, section 3), and the MiniMax-H3 FP8 quality test, which fails
identically on main (its test omits the trust_remote_code that vae.py
requires; nothing on that path is this PR's). The other pipelines at nearby
heads: AMD's one failure was the metrics source test (c1dff45d); Intel's was
No space left on device during git clone; NPU's perf lanes failed at
seed-tts concurrency 8 against max_sessions: 4 — every Seed-TTS utterance
is now a duplex session, so the benchmark probe now queues for a slot on the
retryable resource_exhausted (3dc07acf); NPU's accuracy lane fails a
count check that is independent of this PR (--no-oversample, a 1088-row
Seed-TTS en set, num_prompts=2000; WER itself passed at 0.0208).

CI build 15291 (entrypoint H100 lane, at 93f53f1f) failed one test:
test_server_vad_multi_turn_without_client_commit — a turn-based
Qwen3-Omni session with turn_detection: server_vad and no client commits.
This PR removed the serving-side server VAD that endpointed such sessions, so
a turn-based model now reaches vLLM's realtime handler, which has no server
VAD and rejects the OpenAI-style session.update. That is the P0 raised on
this PR; restoring turn-based server VAD on the transport is RFC PR 5, and
the test is deselected in the lane meanwhile (0a63ceb9) — the file and its
three other realtime tests keep running.

CI build 15319 (ready pipeline at 070b480c, 23 of 26 lanes green) had three
red lanes. Other Test: test_patch.py's fake realtime client did not accept
the until argument stream_silence gained in 77f03ddb (6f7ebd25). Perf:
1/40 Seed-TTS requests timed out waiting for response.done — a model-native
session may choose to listen on the first silence units, and 12 s of silence
then a 180 s wait with no more audio is exactly that stall; the silence budget
is 30 s now, spent only on a turn the model is slow to take since the silence
stops at response.done, a turn needing more than 12 s is logged, and a turn
that never settles is reported with what the model did (6f7ebd25; 3× sample
proof above). Engine&Entrypoints: 16 test_output_metadata_snapshots.py
failures that are main's own — #7608 changed consolidate_tensors to take
the output modality without updating #7448's tests, and main's lane has been
red since build 15307; the tests now pass OutputModality.AUDIO (0bb92c89,
37 passed with test_output_modality.py). The same two failures were AMD's
only red lanes; Intel's was No space left on device during git clone.

CI build 15306 (MiniCPM-o 4.5 perf lane, at 0a63ceb9) failed 1 of 40
Seed-TTS requests at concurrency 4: expected one audio response, got 2.
main drove each Seed-TTS turn with response.create, which yields exactly
one response; the port to the duplex route seeds the text into a model-native
session and streamed the whole 12 s silence cap in real time before waiting
for response.done, so a model that reached turn_eos after a few seconds
kept hearing silence and occasionally spoke again on it. The probe now stops
the silence the moment the turn settles and measures the first response that
carried audio (77f03ddb; a later one is logged as the model's own speech).
Re-run of the same test on the branch server: 50/50 requests, no second
response at all. The same log's touch_duplex_session failed ERROR traceback
was a lease touch racing the session's own close; it is a debug-level result
now.

The two full_model E2E cases that fail (test_duplex_soft_interrupt,
test_duplex_server_vad_hard_interrupt) are not in any gate lane — there is
no full_model lane in test-merge.yml; they run in test-nightly.yml on
H100/B200. Their driver reaches the server through vllm_omni/experimental/.

1. Unittests — refactored duplex framework

742 passed, 0 failed (344.91 s, at 5d17f455).

Suite Collected What it drives
tests/engine/duplex/ 340+ the real DuplexSessionManager / DuplexSessionRunner / DuplexEngineSession against a fake DuplexStagePort — admission, leases, resume/takeover, close-vs-cleanup ordering, the typed command/event contract, the seeded text turn, and the server-VAD overlap decisions
tests/worker/test_native_duplex_hooks.py 65 worker-side sampling hooks
tests/entrypoints/test_omni_entrypoints.py 64 shared entrypoint bases
tests/clients/ 57 DuplexClient, model session configs
tests/benchmarks/ 87 benchmark backend and patches
tests/entrypoints/openai_api/test_duplex_api_server.py 21 which routes exist on a duplex server, which are refused, and that chat is wired only when the model declares it
tests/entrypoints/duplex/ 17 serving as pure transport (journal, resume, idle timeout, takeover)
tests/engine/test_duplex_orchestrator.py 16 the DuplexOrchestrator, the hierarchy the two surfaces depend on, and replica-death cleanup
tests/entrypoints/openai/test_duplex_session_attachment.py 18 attachment registry, detach generations, resume rollback
tests/entrypoints/test_duplex_omni.py 10 DuplexOmni handle lifecycle
tests/engine/test_duplex_import_boundary.py 10 the turn-based stack never imports duplex code
tests/engine/test_duplex_omni_engine.py 9 the DuplexOmniEngine layer

Cases added by review feedback in this round, each verified to fail when its fix
is reverted:

  • test_server_vad_speech_started_reaches_a_barge_in_decision (+2) — server_vad
    with barge_in_on_speech reached an UnboundLocalError on every append
  • test_registry_resume_cancelled_mid_delivery_rolls_back_like_a_failure[activation|replay]
    — CancelledError bypassed the resume rollback
  • test_a_resume_that_fails_to_activate_does_not_detach_the_live_attachment — a
    losing resume could detach the winner
  • test_a_dead_replica_closes_the_sessions_it_was_serving — replica death left
    the owning session and its admission slot alive
  • test_raw_wire_hints_cannot_override_the_normalized_typed_fields (+1)
  • test_an_item_only_response_create_is_refused_not_left_hanging
  • test_the_duplex_stack_extends_the_turn_based_one,
    test_an_unrecognised_message_reaches_the_turn_based_handler

1b. Everything else under tests/engine, tests/entrypoints, tests/clients

2760 passed, 1 failed, 42 skipped, 1 xfailed (401.87 s, at 5d17f455).

The single failure reproduces on main @ 1b6cd282 (2796 passed, 1 failed,
same test), so it is not this PR's:

Failure Also fails on main
test_stage_engine_core_proc.py::test_preprocess_add_request_preserves_omni_fields yes

1c. Turn-based regression — AsyncOmni / AsyncOmniEngine / Orchestrator

941 passed, 1 failed (the same pre-existing test above), 39 skipped
(115.28 s, at 5d17f455).

This is the suite that matters for the re-parenting: DuplexOrchestrator now
extends Orchestrator, DuplexOmniEngine extends AsyncOmniEngine, and
DuplexOmni extends AsyncOmni. Covers test_orchestrator.py,
test_orchestrator_error_handling.py, test_async_omni_engine_stage_init.py,
test_stage_engine_core_proc.py, the four test_async_omni*.py files,
test_omni_sleep_mode.py, test_pd_disaggregation.py,
test_stream_finish_reason.py and tests/config/.

The structural half is test_duplex_import_boundary.py, which spawns a fresh
interpreter, imports the turn-based modules and fails if any duplex module
appears in sys.modules. It passes: the dependency runs duplex -> turn-based
and never the reverse.

1d. Diffusion (DiT) lanes — the turn-based path end to end

The CPU suites above never start a DiT stage, so the re-parented
Orchestrator / AsyncOmniEngine / AsyncOmni stack was also run through
CI's diffusion lanes, verbatim, at 014350a8:

Lane Result
merge gate :23 tests/diffusion cpu 5377 passed, 47 skipped; the one real failure fixed in c1dff45d
merge gate :44 tests/diffusion/lora cuda 1 passed
ready tests/diffusion/cache cuda 8 passed
ready tests/diffusion/batching cuda 6 passed
ready tests/diffusion/offloader cuda 5 passed
ready tests/diffusion/distributed cuda, cards_1 2 passed (297 deselected: CI's L4 marker does not apply on L20X)

accelerate had to be installed in the venv first — without it the cache and
batching lanes fail at pipeline init on both main and this branch — and the
two torchmetrics quality files are excluded locally; CI carries both.

2. Unittests — MiniCPM-o 4.5

515 passed, 2 failed, 3 skipped, 18 deselected (38.50 s, at 5d17f455).

Both failures are test_code2wav_batching.py ([fp32-cache], [bf16-cache])
and reproduce on main @ 1b6cd282 (427 passed, 2 failed, same two ids).
Run in the CPU venv, which does not carry the optional cosyvoice2 package.

3. MiniCPM-o 4.5 E2E and barge-in test

Tier Result
advanced_model (gate lane :131) 11 passed in 462.53 s, at 1dabb4cd
core_model · test_duplex_client_live_session 1 passed in 777.68 s, at 8ae42c6f
core_model · protocol smoke 1 failed at 1dabb4cd; cause fixed in 117e7e03, regression test in test_duplex_serving.py

Both core-tier failures in CI build 15206 were real, not flakes:

  • Protocol smoke (ConnectionClosedError: no close frame received or sent,
    0.63 s after the socket opened): DuplexSessionHandle._deliver marks the
    handle closed in the same step that queues session.closed, and the serving
    read loop exits on that flag — so the endpoint could return, and the ASGI
    server tear the socket down, before the pump had written the terminal event.
    Racy, which is why it passed on re-run. The endpoint now waits for the pump
    to drain a closed session (117e7e03).
  • Live client (900 s TimeoutError draining response.audio(); main
    passes the same test in 577 s): ModelChannel._build_stage_output built the
    decision output with from_stage_output(..., finished=True), whose
    source-copy overwrites finished — a resumable stage request never finishes,
    only its segment does — so the MiniCPM-o projector, which recognises a listen
    only on a finished output, projected the model's listen to nothing: no
    response.listen, no next silence unit, no response.done. The trace of a
    failing run: 69 TTS chunks, 64 silence units (the whole auto budget), one
    stage-0 listen, zero projected results, then 600 s of silence until the
    stall net fired. main was covered by its orchestrator queue re-asserting
    finished on every direct-response message. finished is now set after
    construction (8ae42c6f); two harness tests use the real finished=False
    shape and fail with the fix reverted.
  • Server-VAD hard interrupt (test_duplex_server_vad_hard_interrupt,
    nightly duplex lane; TimeoutError waiting for the follow-up response,
    main passes in 543 s): the runner's new _maybe_schedule_vad_commit
    committed the turn at the detector's speech_stopped for every session.
    main did that only for turn-based server VAD; a model-native session's
    server VAD was interrupt-only. On a native session the early commit cut
    the utterance short: the model listened on that final unit, the trailing
    silence became a second near-empty turn, and no follow-up response could
    ever appear (like-for-like event streams: main one commit then
    response.created; branch two commits, two listens, silence). The
    auto-commit is skipped for auto-response sessions (c737b7c0); two
    harness tests with a scripted detector pin both contracts.

/v1/chat/completions — restored examples and tests

The route is the ordinary chat service on the duplex engine, so the examples
and tests deleted with the old adapter are back unchanged from main
(1fbe5013, wired into CI again in bcbf17ef). Run against the branch's
duplex server: test_text_to_text_001, test_text_to_audio_with_reference_audio,
test_text_to_audio_with_default_reference, test_mix_to_text_audio_001 —
4 passed (the three audio tests need pyttsx3/espeak-ng and opencc,
which CI carries).

Barge-in

PASS — 4 responses, artifacts in barge_in_out/.

  handshake:  HT_ref_audio.wav            (voice sample only)
  user:       response_required_16k.wav →  response_1_completed.wav  6.40 s
  user:       soft_interrupt_16k.wav      (starts DURING response_1 = barge-in)
                                       →  response_2_completed.wav  7.04 s
                                       →  response_3_completed.wav  3.12 s
                                       →  response_4_completed.wav  2.80 s

/v1/chat/completions on the same server

200 in 2.10 s. The other turn-based routes stay off: /v1/completions
answers 501, /v1/models 200.

Both surfaces exercised at the same time on one process: a websocket duplex
session held open and mid-turn from 0.07 s to 4.31 s, with two chat completions
answering inside that window (3.29 s and 3.81 s) and the session answering
afterwards.

4. E2E benchmark with MiniCPM-o 4.5

Measured at 2bad104c against main @ 1b6cd282 and not re-run since;
the commits added after it are review fixes and tests, none touching the
generation path. Both tables were run back to back in one GPU reservation, same
box, same venv, same dataset and reference audio.

Branch duplex_refactor2 @ 2bad104c

Subset Requests E2EL (ms) TTFT (ms) TPOT (ms) Output tok/s Audio TTFP (ms) Audio RTF Generated audio (s) Responses
1q1a 1/1 308583.36 1.96 10.52 2.51 1.94 0.64 86.88 17
1q1a_math 1/1 303651.53 0.23 10.67 2.55 0.22 0.99 131.00 9
1qna 1/1 303738.58 0.04 8.97 0.62 0.02 0.98 38.84 1

Reference: main @ 1b6cd282

Subset Requests E2EL (ms) TTFT (ms) TPOT (ms) Output tok/s Audio TTFP (ms) Audio RTF Generated audio (s) Responses
1q1a 1/1 309604.28 0.04 10.08 1.20 0.02 0.65 86.88 17
1q1a_math 1/1 305239.48 0.04 9.78 2.25 0.02 1.07 131.00 9
1qna 1/1 305285.54 0.04 8.66 0.60 0.02 0.98 38.84 1

Generated audio and response counts match exactly on all three subsets, which is
the property that matters: the refactor changes who owns the session, not what
the model produces. E2EL is dominated by the ~300 s of input per clip, so the
sub-1 % differences are wall-clock noise. TPOT is within 0.5 ms on two subsets
and 0.9 ms on 1q1a_math, in the branch's disfavour — stated rather than
rounded away, though it sits inside this benchmark's run-to-run spread.

5. Lint and pre-commit

At 1dabb4cd: ruff check clean, ruff format --check clean (2648 files).
The later commits (117e7e03, c1dff45d, 8ae42c6f, c737b7c0, 3dc07acf,
e88f0dab, 77f03ddb, 6f7ebd25, 0bb92c89) each passed both on the files they
touch; the restored files in
1fbe5013 are main's verbatim; pre-commit was not re-run after 1dabb4cd.

pre-commit run --all-files fails 5 hooks — markdownlint-cli2, mypy-3.10,
Ensure test files have CI marks, Lint shell scripts, Check SPDX headers.
All five fail identically on main, and every file they name is untouched by
this PR. Net new hook failures: 0. (typos was the one this PR broke; fixed
in 98c5c0dc.) Lint shell scripts is unknown rather than passing —
shellcheck is not installed on the test box, so it evaluated nothing.

6. Pre-check

precheck-pr, full mode, at 1dabb4cd: 0 blocking, 3 warnings.

Dimension Result
PR title format ✓ (was ✗ — retitled [Core][Frontend] …)
Examples policy ✓ no new model-specific Python example
Simplification / dead code ✓ (was ⚠ — TurnDetectionConfig.as_realtime removed in a24d8d8f)
SPDX, forbidden imports, new torch.cuda sites ✓ / 0 / 0
Test CI marks ✓ every added test file carries a level and hardware mark
Code quality ⚠ 70 broad except Exception added; 60 of 81 log, re-raise, return or emit within two lines. Worth a narrowing pass later
PR description integrity ⚠ resolved by this update
Pre-commit gates ⚠ see section 5 — 0 net new

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)

…ared bases with duplex siblings

Part 1 of RFC vllm-project#7181 (unified full-duplex framework). Extract the shared base
classes and add the duplex siblings next to the turn-based classes:

- OmniBase -> AsyncOmniBase -> {AsyncOmni, DuplexOmni}
- OmniEngineBase -> {AsyncOmniEngine, DuplexOmniEngine}
- OrchestratorBase -> {Orchestrator, DuplexOrchestrator}

Collaborators are constructed explicitly (_create_engine / _create_orchestrator);
the generic bases carry no duplex vocabulary. DuplexOmni is a thin Python API
over the engine: open_session() returns a DuplexSessionHandle with a
single-consumer events() stream, session ids are always allocated server-side
as duplex-<uuid4 hex>, and an engine crash ends every handle through the new
_on_engine_dead seam.

PipelineConfig gains duplex_plugin, which is what makes a pipeline a duplex
model. The pre-framework duplex_runtime_extension / duplex_serving_adapter /
duplex_control_enabled fields stay as unread fields so that the pipelines of
the models that are not ported yet still construct; each one goes away with
the follow-up PR that ports its model.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
…one model plugin

Sessions now live inside DuplexOrchestrator: one DuplexEngineSession per
session, owned by a DuplexSessionRunner on the orchestrator loop and admitted /
reaped / dispatched by DuplexSessionManager. The serving-side control plane,
runtime bridge, request client, cross-boundary fence protocol and the API-side
session registry are deleted.

The wire vocabulary becomes a typed contract: DuplexCommand dataclasses
(command_from_realtime) in, DuplexEvent dataclasses (to_realtime) out, carried
over the in-process engine queue as typed messages; control failures are
DuplexSessionError. Server-side VAD, the commit policy and the audio helpers
move from entrypoints/duplex to engine/duplex.

Session teardown has one owner: the manager emits session.closed and
session.expired only after the stage requests are aborted, so the event also
means the admission slot is free; admission holds the slot across the plugin
await; a closing session refuses further commands and control operations.

Model integration is a single DuplexModelPlugin ABC selected by
PipelineConfig.duplex_plugin, with MiniCPM-o 4.5 as its one implementation.
The extra_body.native_duplex opt-in leaves the model side: a duplex model is
model-native and appends audio chunks.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
…duplex clients

OmniDuplexSessionHandler only bridges a Realtime websocket to a
DuplexSessionHandle (RealtimeEnvelope, DuplexSessionAttachmentRegistry for
resume and takeover); every session decision is made in the engine. A send
failure or a read-idle timeout is treated as a disconnect of that socket, and
a socket that lost a takeover never acts on the session again.

vllm-omni serve runs a duplex model through DuplexOmni and serves
/v1/realtime?duplex=1 (alias /v1/duplex), /v1/models and /health only.
MiniCPM-o 4.5 servers no longer serve /v1/chat/completions (breaking, RFC
decision D9). The extra_body.native_duplex opt-in is removed from the client,
server and benchmark code. The startup warmup reads the engine plugin's
silence unit and uses the canonical Realtime URL builder.

DuplexClientBase is an ABC with two implementations: DuplexClient over the
websocket and the new InlineDuplexClient over an in-process DuplexOmni, both
translating client events with the session's declared audio defaults. Clients
can no longer choose session ids.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
New CPU suites for the session runner, session manager, engine session, plugin
loading, typed commands and events, DuplexOrchestrator, DuplexOmniEngine,
DuplexOmni and the serving transport, including the teardown ordering, the
admission race, engine death and the transport-failure paths. The suites of the
deleted control plane, handler, runtime adapter and protocol modules are
removed. Existing orchestrator, engine, client, entrypoint and e2e driver tests
follow the new module layout, the server-allocated session ids and the plugin
contract. The MiniCPM-o online chat e2e and invalid-parameter tests go away
with the chat route.

The duplex tests of the models that are not ported yet are left untouched but
not collected (one conftest per duplex test directory): they import the
pre-framework runtime and serving adapter this PR removes, and an import error
would abort the whole pytest session.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
Rewrite docs/design/fullduplex.md for the engine-resident design and update the
Realtime duplex API guide (server-allocated ids, /v1/duplex alias,
InlineDuplexClient, Python API section, corrected event catalogue), the module
ownership docs, the stage config docs, the deploy profile comments and the
MiniCPM-o recipe and example README for the duplex-only server. Both pages now
name MiniCPM-o 4.5 as the only model served over the framework in this PR.
Remove the MiniCPM-o chat-completions examples and Gradio demo, add --inline to
the barge-in client and drop client-chosen session ids from the Realtime demo.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
…ion paths

A transport send failure could not detach its session: the pump called
DuplexSessionAttachmentRegistry.detach() without the required
attachment_generation, so every failed send raised TypeError inside the pump,
killed it, and left the session attached to a dead socket with no disconnect
grace until its idle TTL. detach() now takes attachment_generation=None for
callers that only know the session (the pump outlives any one connection), and
answers False for an already detached session so the disconnect that follows a
failed send cannot restart the engine's grace window.

The test that claimed to cover this passed against the broken code: it closed
the socket, which also ended the read loop, so the read-loop disconnect path
did the detach it asserted. It now breaks only the socket's write half.

DuplexSessionRunner.expire emits session.expired even when a close already
deferred the terminal event, so a wire session.close racing a stage failure --
which takes the runner away from the manager -- cannot end a session with no
terminal event at all.

DuplexSessionManager.open clears _admitting only for the call that set it; a
duplicate open of a live id used to release the admission slot the first open
was holding. Unreachable while ids are server-allocated, so this is a guard.

Also: the resume credentials stamped onto session.created are a typed
dataclass instead of a dict splatted into dataclasses.replace, reap_expired no
longer binds two different cleanup dataclasses to one loop variable, the
"requires audio" error names input_audio_buffer.append, and the OmniInteract
session wrapper drops the session_id it ignores now that ids are
server-allocated.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
…cases

tests/entrypoints/openai_api/test_duplex_api_server.py locks the duplex-only
server (RFC vllm-project#7181 D9), which had no coverage: the pipeline probe that decides
DuplexOmni vs AsyncOmni, the app.state snapshot (the session handler is wired
and every turn-based service stays None, which is what makes the turn-based
routes answer "not available"), both duplex websocket routes reaching the
handler while /v1/realtime without the flag still reports the turn-based
Realtime API unavailable, and the startup-warmup gate.

Two e2e cases replace modality coverage the deleted online chat suite had.
test_duplex_single_session_still_image_input attaches one JPEG frame, which the
driver repeats on every unit -- a camera track advances, a still does not.
test_duplex_text_only_response_create_is_rejected pins what a text prompt gets
now that every session is model-native: conversation.item.create with an
input_text part plus response.create is refused with
response_create_without_input, because the chat-completion fallback that used
to answer it is gone. Both were validated on a GPU against the bundled deploy
profile.

tests/e2e/online_serving/test_minicpmo_4_5_expansion.py goes away with its
sibling test_minicpmo_4_5.py: all five of its tests drive
/v1/chat/completions, which a duplex-only MiniCPM-o server does not serve.

The CI wiring is brought back in step with the refactor. Four pipelines still
ran or watched tests/e2e/online_serving/test_minicpmo_4_5.py, deleted with the
chat route, so the ready pipeline's "MiniCPM-o 4.5 Online Test" had no test
file left to run and the merge and NPU jobs passed a missing path to pytest.
The duplex jobs watched the deleted duplex_request_client.py and
openai/duplex_capability.py and none of the new framework modules, so a change
to duplex_omni.py, duplex_orchestrator.py, duplex_omni_engine.py or the new
bases triggered no duplex job.

tools/pre_commit/check_forbidden_imports.py still allow-listed the moved and
deleted duplex files for "import base64", so the hook failed on the six files
that now import it. Five files this PR already touches gain the SPDX header
the check_spdx_header hook requires.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
…-o 4.5

docs/serving/full_duplex_api.md was the one duplex page the refactor missed. It
listed PersonaPlex and Nemotron VoiceChat as current unified-runtime
integrations and used PersonaPlex for its capability example; it now carries
the same PR1 scope statement as fullduplex.md and realtime_duplex_api.md and a
model-neutral capability paragraph. Its enablement section also said a model is
duplex when the pipeline declares duplex_plugin *and* the deploy config sets
session_mode: duplex; the plugin alone decides, and a duplex model whose deploy
config omits session_mode fails at startup rather than falling back.

realtime_duplex_api.md records how text reaches a model-native session: the
OpenAI text-prompt shape does not drive a response, and text-to-speech goes
through extra_body.duplex_initial_user_text plus audio units, which a two-arm
GPU probe confirmed still produces the seeded speech. The alias paragraph names
the aliases that exist and points at the translator instead of a table that
does not exist in fullduplex.md, and states that the pre-Realtime WAV-append
aliases push_chunk and input.audio.append are not accepted. The admission
bullet no longer reads "resource_exhausted or resource_exhausted" after
duplex_session_busy was removed.

fullduplex.md described the pre-fix close ordering (the manager emits
session.closed after the stage cleanup, not the runner before it) and omitted
audio_encoding.py from its layout.

vllm_omni/experimental/fullduplex/README.md pointed model authors at the
duplex_serving_adapter / duplex_runtime_extension dotted strings this PR
replaced, and its "stable tree" listing named the deleted request client.

Also: recipes/README.md describes MiniCPM-o online serving as full-duplex,
supported_models.md links its recipe, and the barge-in flow doc covers the
--inline mode.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
…he knob

Review feedback on RFC vllm-project#7181.

The "ordered event stream" bullet in the duplex design doc was read as a
delivery guarantee. It is not: the engine output queue and each
DuplexSessionHandle outbox are unbounded FIFOs, so a cancellation is delivered
behind whatever audio was already emitted for the response it cancels. The
bullet now separates order from latency, points a client that must stop
quickly at its own playback cancellation, and names bounded buffers plus
invalidation-skip as follow-up rather than claiming them.

duplex_session.completed_append_cache_size is no longer read. It bounded the
per-session completed-append table that made a retried append RPC submit once;
the framework carries appends as one-way commands on the runner's ordered
mailbox, so there is no client-visible retry to deduplicate. The field stays
because the deploy configs of the models that are not ported yet set it and
DuplexSessionRuntimeConfig would reject an unknown key, so it is commented as
unread and the existing "no framework source reads the pre-framework fields"
guard is extended to cover it.

No behaviour change: prose, one comment and one assertion.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
…nits

Stage1 reports a cumulative snapshot of every sample it has generated for a
request, and slice_cumulative_audio keeps a per-request cursor so only the new
tail reaches the client. A continuation unit arrives as a one-sample snapshot.
Because that is smaller than the cursor, it was read as "the producer restarted
its buffer" and the cursor was rewound to 1; the next real snapshot was then
sliced from the start and the whole session's audio went out again as a single
delta:

  slice ... prev=637440 num=1        <- continuation placeholder, cursor rewound
  slice ... prev=1      num=649920   <- 27.08s re-sent instead of 0.52s

Once per response, that turned a 300s OmniInteract case into 1373s of emitted
audio. The client's serialized playback clock then ran hundreds of seconds
behind the wire, its progress acks kept the session from ever settling, and the
case timed out at 1201s with Stage2 still waiting for a chunk. 1q1a_math now
completes in 303.7s with 131.00s of audio, against 304.96s and 129.44s before
the refactor.

A snapshot below the cursor is genuinely ambiguous from that snapshot alone --
a restarted producer and a unit carrying no cumulative audio look identical --
so the decision is deferred by one snapshot instead of guessed. A dip no longer
moves the cursor; only a second, larger below-cursor snapshot is accepted as a
restart, which is what a restarted producer actually looks like as it grows.
Anything above the cursor proves the stream never restarted and is sliced as
usual.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
DuplexSessionRunner.on_stage_output returns early for a Stage0 output that
carries no decision: it is plain text for the TTS stage and must be forwarded
rather than consumed. That early return also dropped the StageRequestStats
delivered with the output. Before sessions moved into the engine the
orchestrator published those separately as a StageMetricsMessage on the output
queue, a branch that req_state.session_owned now bypasses, so nothing reported
them at all -- every emitted event carried "stage_metrics": {} and no client
could see an engine-side token count.

The visible cost is the benchmark: with no stage0_tokens block, OmniInteract
cannot compute TPOT or ITL and falls back to tokenizing the transcript for its
output-token count. TPOT is now measured again at 10.57 / 10.47 / 9.07 ms
across the three subsets, against 10.24 / 10.37 / 9.00 before the refactor.

The metrics go to the session on the runner's mailbox rather than straight from
the orchestrator loop, so they stay ordered with the session's other work. A
stage that hands its output to the next stage reports before the response
carrying those tokens exists, so the session holds the snapshot and folds it
into the first response that opens after it. That attributes a turn the model
never spoke to the next response it does speak, which keeps the session total
whole; dropping the snapshot instead would lose the tokens outright.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
A client session.close is now answered with session.closed carrying
reason: "client_close"; before the refactor that event carried no reason at
all. realtime_duplex_api.md documents the reason as part of the contract, and
it is what lets a client tell "I closed this" from "this died on me", so the
consumer is what needs updating.

_raise_if_session_terminated used the absence of a reason as its only way to
recognise the close it had just requested, so every OmniInteract case failed
with "Unexpected session.closed: client_close" after running to completion --
all three benchmark subsets, on a session that had done nothing wrong. It now
accepts the reason the server stamps on our own close and still fails the case
for any other one, so a disconnect, a lease timeout or a shutdown is caught
exactly as before.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
Seed-TTS had two entry points for MiniCPM-o and this PR removed both. The
openai-chat-omni entries drove /v1/chat/completions, which a duplex model no
longer serves (RFC decision D9). The openai-realtime-tts backend asked for
speech the OpenAI way -- a text conversation item followed by response.create
-- which a model-native duplex session refuses with
response_create_without_input, because it generates from audio units and takes
its text once, in the session context.

_RealtimeTTSProbe now drives the shape the model actually has: one session per
utterance, opened with extra_body.duplex_initial_user_text carrying the target
text, then silence appended in realtime so the clock advances and the model
answers the seeded turn. A two-arm GPU probe confirms the seeded arm speaks the
seeded sentence while the control arm produces unrelated output.

The cost is multi-turn. The backend existed so that
--seed-tts-turns-per-session could group several target texts into one session;
with one injection point at session open that cannot be expressed, and
preserving it needs a framework capability this PR does not have -- a way to
inject a user text turn mid-session. The flag now groups texts into one
benchmark request and its help says each turn runs in its own session.

Because a session open per utterance re-encodes the reference audio and
rebuilds the context, the recorded baselines are not stale so much as measuring
something else, so they are removed rather than carried forward. The four perf
configs move to the realtime backend and /v1/realtime; new baselines have to be
taken from the ported path.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
The turn-based e2e suite that this PR deletes with /v1/chat/completions was the
only coverage for several MiniCPM-o modalities. Five duplex cases take that
back: text to audio in both locales, text to text (the one duplex session that
opens without a reference voice, since ref_audio is required only for audio
output), a long audio answer, long-form generation, and isolation between
sequential sessions. The seeded cases also guard the Seed-TTS shape, so the
benchmark path and the modality are covered by the same test.

Four session-manager cases cover admission and terminal-event edges that the
close-vs-cleanup ordering made reachable: a duplicate open arriving while the
first is still in flight must not release the slot the first one holds; an
expiry must still emit a terminal event when a close deferred it; a terminal
event is emitted once even if expiry runs twice; and a wire session.close on a
session that bound no stage request still answers with session.closed, which is
the protocol smoke test's exact shape and the only thing that close returns.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01KfPFmE6NdUkftRMD9WDTW8

Signed-off-by: chickeyton <ngton2014@gmail.com>
@chickeyton chickeyton changed the title Unified Full-duplex Framework [WIP] Unified Full-duplex Framework Sep 11, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/observability.md, docs/design/module/engine_orchestration.md.

Module owners: @lishunyang12 @vraiti @tzhouam

Routing: @lishunyang12 via module of the changed files; @vraiti via module of the changed files; @tzhouam via semantic router, CODEOWNERS

@chickeyton, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@0z5a 0z5a mentioned this pull request Sep 20, 2026
24 tasks
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner frontend code related to entrypoint high priority high priority issue, needs to be done asap omni code related to omni models ready label to trigger buildkite CI refactor refactoring for better code scalability and quality

Projects

None yet

Development

Successfully merging this pull request may close these issues.