Skip to content

[Core][Frontend] Bound duplex output delivery and reject cancelled audio - #7644

Merged
linyueqian merged 6 commits into
vllm-project:mainfrom
NolenLiang:feat/duplex-output-delivery
Oct 4, 2026
Merged

linyueqian merged 6 commits into
vllm-project:mainfrom
NolenLiang:feat/duplex-output-delivery

Conversation

@NolenLiang

@NolenLiang NolenLiang commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Follow up on merged #7413 with bounded per-session output and rejection of cancelled audio before delivery. This is an independent change against main, rebased onto b308e196, with head 5dae6122; upstream changes are not part of this diff. The rebase preserves #7785's cancelled-open rollback, append compensation, and resume/lease-generation handling. Output-buffer ownership is released synchronously in the unified open rollback, before shutdown can be cancelled again. This alignment preserves the server-side request-start metrics from #7714 alongside the existing closing guards, AURA request-list aborts, draining-response ownership, and completed-turn audio filtering. Cancellation now invalidates every cancelled draining response before awaiting stage abort, including silent session close. The overflow policy and creation-event snapshot fix are unchanged.

This implements the output-delivery work discussed in the proposal and the author's response. Real-model coverage is limited to MiniCPM-o 4.5; the buffer applies to the shared session layer. Real Qwen model validation and general playback/history policy are outside this change's scope.

Problem and fix

The shared engine output queue and per-handle queues are unbounded. Cancelling generation does not remove audio already waiting for delivery, so old audio can arrive before cancellation notifications. Blocking the shared pump on a slow consumer would also delay other sessions.

  • Add one bounded output buffer per session, shared by the engine and handle; public audio bypasses the shared queue.
  • Invalidate matching queued audio and an already-dequeued event for both the active response and cancelled draining responses. WebSocket delivery checks again under the connection lock, before sequence allocation, replay recording, or on_accepted. Local Python iteration uses the same validity state.
  • Keep valid events ordered: normal completion cannot overtake valid audio, and final session closure has independent space.
  • On overflow, preserve a pending ending's identity while making it failed, without duplicating an ending already accepted. Do not retain partial failed-response history; preserve request bindings needed for engine cleanup.
  • Keep the growing session item separate from the three response-creation events, so their queued payloads cannot grow after admission. This fixes the accounting gap identified in review without copying audio or all event payloads.

Behavior and limits

  • Defaults: 2 MiB of compact JSON including base64 audio and 512 events per session. Error/response-ending reserve: 64 KiB and eight events. One held event and a small final closure notification are additional. These are not bounds on total Python/process memory.
  • Proposed overflow policy: fail the active response with output_backpressure and close only the affected session. Keeping that session open after stopping just its response is not implemented. The duplex-owner policy discussion remains open; the latest reviewer update explicitly says it does not block this PR.
  • After overflow, suppress newly emitted ordinary events, including text/audio and content-part/output-item completion markers. Previously queued non-audio events retain order. Accepted cancellation, unlike overflow, removes only matching audio; already-sequenced replay entries are unchanged.
  • If the error/ending reserve also fills, those notifications may be omitted. Closure retains its own slot; oversized details become close_details_exceed_output_limit. This guarantees local closure capacity, not network delivery or successful engine-resource release.
  • Failed opens close abandoned handles. Late events cannot reopen closed handles, and non-error events cannot fall back to the shared queue after buffer removal.
  • Already-yielded or transmitted audio cannot be recalled; clients still need to stop playback. The runner's raw stage-output mailbox and replay storage are outside this limit. Delivery uses the framework's same-process, separate-thread arrangement.
  • Native MiniCPM still accounts audio before emission and does not pass on_accepted. Callback coverage does not establish that native sent_ms measures delivered or physically played audio.
  • The creation-event fix does not solve every mutable-item case in #7636 Issue 11; later truncation can still change shared content in completion/retrieval events. That broader work remains separate.

Test Plan

vLLM Version: Official vLLM 0.30.0 ARM64 environment, Transformers 5.14.1, and model-specific speech dependencies.

vLLM-Omni Commit: Current fixed baseline b308e196; candidate 5dae6122. The current rebase has CPU contract coverage in the official vLLM 0.30.0 ARM64 runtime. Real-model results below remain attributed to their recorded revisions, including 75570ee3 / 4b4e157; they are not new-head GPU measurements. The earlier regression-only revision 879f8b4 records the AURA assertions before their integration fixes on the previous baseline.

Real-model runs use openbmb/MiniCPM-o-4_5 at 503e754, one GB200 per service and the matching MiniCPM deploy configuration, without overriding its compilation settings. WebSocket runs use the production duplex handler in a small loopback application, not the complete OpenAI application.

Current fixed scenarios:

  1. Baseline/candidate WebSocket pair: same short input, hold a send before sequencing, pause six seconds, submit cancellation, then release after two seconds. Observe the other original response through natural completion, with the existing timeout retained.
  2. Candidate WebSocket overflow: hold one session's server-side outbound lock before sequencing, with a 768 KiB / 512-event limit; observe its failure/closure, the other response across closure, and audio from a replacement session. This injects a server-side send pause; the client continues reading, so it is not a physically slow network/client experiment.
  3. Candidate local Python pair: pause one handle, cancel its response, continue reading for late audio, and observe the other session's output before closing both handles.
  4. Review regression: generate creation events with the real projector, enqueue them, then generate text/transcript growth. Verify the session item grows while the queued payloads and their admitted byte counts remain unchanged.
  5. Protocol-move regression: extend the existing error-code case to verify output_backpressure still serializes as rate_limit_error through the shared duplex codec.
  6. AURA integration regressions: pause the existing test stage port's abort operation and verify that held and queued audio for a cancelled draining response is already invalid, for both barge-in and silent session close. Retain the existing request/resource/terminal assertions.

Related checks, in the matching runtime with repository test dependencies:

python -m pytest -q tests/engine/duplex/ \
  tests/engine/test_duplex_orchestrator.py tests/engine/test_duplex_omni_engine.py \
  tests/engine/test_duplex_import_boundary.py tests/entrypoints/test_duplex_omni.py \
  tests/entrypoints/duplex/ tests/entrypoints/openai/test_duplex_session_attachment.py \
  tests/clients/ tests/protocol/ \
  tests/model_executor/models/qwen3_omni/test_duplex_plugin.py \
  tests/model_executor/models/aura_omni/duplex/ \
  tests/model_executor/stage_input_processors/test_aura_omni_duplex_history.py
pre-commit run --from-ref b308e196 --to-ref 5dae6122

Test Result

Feedback is requested on the overflow policy and completion behavior; neither policy changed in the review follow-up.

Cancellation-compensation rebase at 5dae612

Rebased onto fixed main b308e196 after #7785. The four textual conflicts were resolved without replacing the upstream cancellation-compensation paths. The existing AURA draining-response routing and pre-abort audio invalidation are retained. Two upstream cancelled-open tests now receive the session output buffer, with assertions that its ownership is removed on cancellation and before a second cancellation during rollback.

  • Related CPU checks in official vLLM 0.30.0 ARM64: 899 passed, 1 skipped, 16 warnings; job/runtime exit status 0.
  • The same changed-file pre-commit selection on candidate and fixed baseline reports 13 passed, 1 failed, 5 inapplicable skipped. Mypy is the only failure; normalized diagnostics are identical: 8 errors / 7 notes. Both working trees remain clean. This is not an all-green local typechecking result.
  • Independent comparison of the old and rebased patch found no additional integration issue. No real-model run was performed for this rebase; previous MiniCPM measurements retain their original revisions and limits.
  • Hosted checks on 5dae6122 are now successful: DCO, pre-commit, Python 3.11 / 3.12 builds, and Read the Docs (verified 2026-09-25 UTC). This does not change the baseline mypy caveat or add model coverage.

Metrics integration at 4b4e157

This revision retains the merged #7714 timing behavior and #7992 prewarm fix. The earlier new_epoch_requires_first_chunk failure is also tracked in #7979. Upstream selected #7992; #7995 closed without merging after its author confirmed the narrower fix. No codec recovery workaround is added here.

  • Full applicable pre-commit on candidate and fixed baseline used the same 26-path selection, omitting only the two files absent from baseline. Both report 13 passed, 1 failed, 5 inapplicable skipped. Mypy is the sole failure: 10 errors / 7 notes, with identical diagnostic multisets after normalizing line/column numbers; nine errors are in production files and one is in an upstream test helper. Neither working tree changed. This is not an all-green local typechecking result.
  • Related checks on 4b4e157: 870 passed, 1 skipped, 16 warnings, with zero job/runtime exit status. This includes CPU AURA/Qwen contract coverage, not real AURA/Qwen model validation.
  • All four fresh MiniCPM runs completed with zero job/runtime/probe-and-shutdown exit status. The WebSocket negative control reproduced 9 old audio chunks / 432,000 decoded bytes after accepted cancellation; the candidate delivered 0. The held 48,000-byte event was journaled and delivered only on baseline. In each run, the other original response produced nine chunks during the pause and two after the cancelled ending, then completed naturally. Client/server event identity, sequence, audio length and hash matched in the raw records. This is one controlled pair, not a repeatability or latency claim.
  • Candidate overflow passed all ten checks: output_backpressure error, exactly one failed/uncommitted original-response ending, then closure of that session. Held audio was absent from journal/send/receive. The other original response produced three chunks while held and five after closure; a distinct replacement produced 40,320 bytes of real audio and restored two active sessions. Normal and replacement sessions were explicitly closed after observation, without waiting for natural completion.
  • Candidate local handles delivered 0 old audio chunks / 0 bytes after cancellation. The other original response produced eleven chunks during the pause and two after the cancelled ending, then completed naturally. This covers direct Python handles, not a complete InlineDuplexClient application.
  • The previous new_epoch_requires_first_chunk failure did not recur in this cohort on the new upstream baseline. Earlier failed runs remain failed evidence. Shared-memory shutdown warnings remain: 86 objects in each WebSocket comparison run, 26 in overflow, and 32 in local handles. These tracker counts are not measurements of unreclaimed bytes after process exit. This validation does not establish leak-free cleanup, physical playback behavior, or engine-stop latency.
  • Source, observers, launchers and model metadata were frozen and checksum-verified. The selected runtime and installed versions were recorded, and overlays were mounted read-only during each run; external image/overlay contents were not fully content-addressed or locked against all host writers.

The previous AURA and main-alignment measurements used vLLM 0.29.0. Older evidence remains scoped to its recorded runtime and revisions; it is not vLLM 0.30 validation. Historical outcomes and limits remain unchanged.

Previous AURA integration validation at 7ed2bc0

AURA integration at 7ed2bc0

  • Final related checks at 7ed2bc0: 849 passed, 1 skipped, 16 warnings, with zero outer/runtime exit status. The selection includes AURA engine, plugin, and history tests as well as the existing Qwen plugin tests. This is CPU contract coverage, not real AURA/Qwen model coverage.
  • Full applicable pre-commit on candidate and baseline used the same 26-path selection, omitting only the two new files absent from baseline. Both report 13 passed, 1 failed, 5 inapplicable skipped. Mypy is the sole failure: 9 errors / 7 notes with identical diagnostic multisets after normalizing line/column numbers. Test typechecking passes, and neither working tree changed. This is not an all-green pre-commit result.
  • The new real MiniCPM WebSocket baseline/candidate pair both failed: stage 2 raised MiniCPMO45Code2WavBatchError with reason new_epoch_requires_first_chunk, followed by request_cleanup closures. Both have nonzero outer/runtime/probe status. Ten of eleven scenario checks passed; the no-errors/no-unexpected-close check failed. Before that failure, old audio after cancellation was 7 chunks / 336,000 decoded bytes on baseline versus 0 on candidate; the held event entered replay/delivery only on baseline, and the other original response completed naturally in both. These are partial observations from failed runs, not a successful pair or evidence that the stage-2 failure is fixed.
  • Candidate WebSocket overflow passed all ten scenario checks, with zero outer/runtime/probe exit status: output_backpressure, one failed original-response ending, then closure of the affected session. Held audio was absent from replay/send/receive records. The other original response produced three chunks while held and six after the affected client's closure; a distinct replacement session produced audio and restored two active sessions. The normal and replacement sessions were explicitly closed after observation, without waiting for natural completion.
  • Candidate local-handle delivery also completed with zero outer/runtime/probe exit status. The cancelled response delivered 0 old audio chunks / 0 bytes after cancellation submission and after the cancelled ending. The other original response produced eight chunks while the slow reader was paused and three after the cancelled ending. Both handles were explicitly closed; this exercises direct Python handles, not a complete InlineDuplexClient application.
  • Shared-memory shutdown warnings remain: 33 objects plus one semaphore in each failed WebSocket comparison run, 29 objects in overflow, and 30 in the local-handle run. Successful probe exits do not establish leak-free cleanup, physical playback behavior, or engine-stop latency. The shared stage-2 failure in the WebSocket comparison remains unresolved.

The regression-only revision reported 843 passed, 6 failed, 1 skipped. Both new cancellation cases failed at the held-audio validity assertion before the fix. The other four failures identified upstream-added call sites needing adaptation: two AURA test contexts lacked closing state, a session-manager test omitted awaiting the buffer-based event reader, and the new input-modality error path passed a single event to the list-based emitter. The latter test now also reads the same bounded output buffer as the session. Original assertions are retained.

The earlier Qwen advanced-E2E tracking issue #7732 is now closed. The CPU AURA/Qwen tests and real MiniCPM observations here do not establish real AURA or Qwen model coverage.

Previous main-alignment validation at bf04024

Validation at bf04024

Four real-model runs completed on 2ab5d17 / 245942b: the baseline/candidate WebSocket pair, candidate overflow, and candidate local handles. All four had zero outer/runtime/probe exit status. The final head only adapts existing tests; these are not claimed as model runs on bf04024.

  • Final related checks at bf04024: 697 passed, 1 skipped, 16 warnings, with zero outer/runtime exit status. The selection now also includes the existing Qwen plugin test file; this is not real Qwen model or advanced E2E coverage.
  • The initial check run at 245942b reported 673 passed, 1 failed, 1 skipped. The failure was an upstream-added test constructing a session-open message without the required output buffer. Both newly added call sites are now adapted, including the Qwen test harness sharing that same buffer; no new tests or assertions were added. The final run above includes the formerly failing case.
  • Full applicable pre-commit hooks ran on candidate and baseline with the candidate's changed-file selection, omitting only the two new files absent from baseline. All applicable hooks except mypy passed. Both report 11 errors / 7 notes; normalizing line/column numbers leaves identical diagnostic multisets, including duplicate counts. Neither run is all green, and both working trees remained unchanged.
WebSocket short-input observation Baseline 2ab5d17 Candidate 245942b
Old audio after cancellation submission 8 chunks / 384,000 decoded bytes 0
Held audio sequenced, journaled and delivered Yes No
Other original response during pause / after cancelled ending 8 / 2 audio chunks 8 / 3 audio chunks
Other original response completed naturally Yes Yes

Both WebSocket runs passed all eleven scenario checks. Counts exclude subsequent response IDs, and the other original response completed before explicit cleanup. This is one pair, not a repeatability or latency claim.

Candidate overflow passed all ten conditions: output_backpressure error, one failed original response ending, then closure of only the affected session. The held audio was neither journaled nor delivered. The other original response produced four chunks while held and five after the affected client's closure; a distinct replacement session produced audio and restored two active sessions. Normal and replacement sessions were explicitly closed, not observed to finish naturally.

Candidate local-handle delivery received 0 old audio chunks / 0 bytes after cancellation. The other original response produced ten chunks while the slow reader was paused and three after the cancelled ending; this time it completed naturally before explicit handle cleanup. This exercises direct handles, not a complete InlineDuplexClient application.

Shared-memory cleanup warnings remain in all four runs. The deliberate server-side hold is not a physically slow network experiment, and the earlier long-input timeouts below remain unresolved. These MiniCPM observations do not validate the advanced Qwen E2E cases in #7732; that merge restriction remains subject to maintainer resolution or explicit approval.

Previous protocol-move validation at f259bb1

Rebased validation at f259bb1

The new baseline/candidate WebSocket pair, candidate overflow, candidate local-handle run, and related checks each ran once in the runtime above. All five completed with zero outer/runtime exit status; all four model probes also exited successfully.

  • Related checks: 652 passed, 1 skipped, 16 warnings. This now includes tests/protocol/, so the count is not directly comparable to the earlier 502-case selection.
  • Full applicable pre-commit hooks ran on candidate and baseline using the candidate's changed-file selection; only the baseline selection omitted the two newly added files. All applicable hooks except mypy passed. Candidate: 50 errors / 35 notes; baseline: 51 errors / 35 notes. After normalizing line numbers, there are no new diagnostics; one existing assignment error is absent. Neither run is all green, and neither working tree changed.
WebSocket short-input observation Baseline e3be42e Candidate f259bb1
Old audio after cancellation submission 8 chunks / 384,000 decoded bytes 0
Held audio sequenced, journaled and delivered Yes No
Other original response during pause / after cancelled ending 8 / 3 audio chunks 7 / 3 audio chunks
Other original response completed naturally Yes Yes

Both runs passed all eleven scenario checks without errors or unexpected closure. The other original responses completed before explicit cleanup; later response IDs are excluded from these counts. This is one new pair, not a repeatability or latency claim.

Candidate WebSocket overflow passed all ten conditions: output_backpressure error, one failed original response ending, then closure of only the affected session. The held audio was absent from replay/send/receive records. The other original response produced four chunks while the send was held and five after the affected client's closure; a distinct replacement session produced audio and restored two active sessions. The other and replacement sessions were explicitly closed, not observed to finish naturally.

Candidate local-handle delivery received 0 old audio chunks / 0 bytes after cancellation. The other session produced nine chunks during the pause and three after the cancelled ending. Both handles were explicitly closed; this is direct-handle coverage, not a complete InlineDuplexClient application.

Shared-memory cleanup warnings remain in all four model runs. Successful exit does not establish leak-free cleanup, physical playback behavior, or engine-stop latency. The deliberate server-side hold is not a physically slow network experiment. Earlier long-input timeouts below remain unresolved.

Validation before the protocol-move rebase

Review follow-up at b2bb100

  • The new regression failed both text/transcript cases against the unchanged d0b5699 production code. On b2bb100, the related-check command without the subsequently added tests/protocol/ selection completed with 502 passed, 1 skipped, 16 warnings, including both cases. To run only the regression: python -m pytest -q tests/engine/duplex/test_delivery.py -k queued_response_creation.
  • Full applicable pre-commit hooks ran on candidate and baseline using the same changed-file selection. All applicable hooks except mypy passed. Candidate: 50 mypy errors / 35 notes; baseline: 51 errors / 35 notes. No new diagnostic remained after normalizing line numbers; one existing assignment error disappeared. This selection now includes realtime_events.py, so these totals are not comparable to the earlier 28-error selection. Neither run is all green, and neither working tree was changed.
  • One final-head real-model WebSocket run passed all eleven short-input conditions, with zero outer/runtime/probe exit status. Old audio after cancellation was 0 chunks / 0 bytes at both server and client; the held event was neither journaled nor delivered. The other original response produced 7 chunks during the pause, 3 after the cancelled ending, and completed naturally with 14 chunks / 677,760 decoded bytes before explicit cleanup. Later response IDs were excluded; server/client audio records matched by identity, size and content digest.
  • This was one candidate run, not a new baseline pair. Overflow and local-handle scenarios were not rerun on b2bb100. Shared-memory cleanup still warned about 84 objects; successful exit does not establish leak-free cleanup. The same loopback, deliberate-pause and playback limitations below apply.

Original post-merge validation at d0b5699

All five original fixed runs completed with zero outer/runtime exit status: four real-model scenarios and one related-check run, each submitted once. The following results belong to d0b5699, not the follow-up commit.

  • Full applicable pre-commit hooks ran on candidate and baseline: Markdownlint, Ruff, formatting and other applicable checks passed. Both reported the same 28 mypy errors and 12 notes after normalizing line numbers; neither run is all green. The baseline comparison excluded the two newly added files absent from baseline. Neither working tree was modified by the checks.
  • Related checks on d0b5699: 500 passed, 1 skipped, 16 warnings.
  • The fixed observer snapshots and reproduction guide are now published separately; they do not add files to this production diff. Executable content matches the measured scripts, with only missing license headers added. The guide records the runtime dependencies and limitations; it is not a complete environment installer.
WebSocket short-input observation Baseline e78a5d0 Candidate d0b5699
Old audio after cancellation submission 8 chunks / 384,000 decoded bytes 0
Held audio sequenced, journaled and delivered Yes No
Other original response during pause / after cancelled ending 8 / 3 audio chunks 8 / 3 audio chunks
Other original response completed naturally Yes, 14 total chunks Yes, 14 total chunks

Both runs satisfied all eleven scenario checks. Their other original responses completed before explicit cleanup; later response IDs were excluded from those counts. Client audio records matched server records by event/response identity, size and content digest. This is one pair, not a repeatability claim.

Candidate WebSocket overflow satisfied all ten conditions: output_backpressure error, one failed original response ending, then session closure; held audio absent from replay/send/receive records; the other original response continued across closure; a distinct replacement session produced 40,320 decoded audio bytes and restored two active sessions. Normal and replacement sessions were explicitly closed afterward, not observed to finish naturally in this scenario.

Candidate local-handle observation received no old audio after cancellation; the other session continued with nine chunks during the pause and three after the cancelled ending. Both handles were explicitly closed. This exercised direct handles, not a complete InlineDuplexClient application.

Shared-memory cleanup warnings remained in all four model runs. Successful exits do not establish leak-free cleanup.

Earlier fixed-checkpoint evidence and unresolved observations

On 8ae42c6 versus cf30cd2, the short-input WebSocket pair observed eight old audio chunks / 384,000 decoded bytes after cancellation submission on baseline versus zero on candidate. The held event was recorded for replay on baseline and not on candidate. Both other original responses completed naturally. This used three input repetitions and does not explain earlier failures with twelve repetitions.

On c1dff45 versus cd36fe4, a separate four-run long-input cohort completed only the candidate/cancellation condition; baseline/cancellation and both no-cancellation conditions timed out while the other response continued producing audio without a recorded natural ending. All failures remain failures. These observations do not identify the cause or exclude an effect from this change.

An earlier cd36fe4 overflow observer passed its ten checks, but its outer launcher exited with a shell error: that overall run was not clean. The current launch chain and observers are frozen together to avoid changes during execution. Shared-memory cleanup warnings in earlier model runs remain unresolved.

These are scenario observations, not performance comparisons. ASGI send return is not physical playback; deliberate pauses do not measure engine-stop latency. Sustained memory bounds, leak-free cleanup, speech quality, and reliable long-response completion are not established.

AI assistance: Codex helped with implementation, validation, and writing.

@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-09-16T12:16:35.346875Z 19ddd69 Draft marked ready
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/observability.md, docs/design/module/engine_orchestration.md.

Module owners: @lishunyang12 @vraiti @tzhouam

Routing: @lishunyang12 via module of the changed files, CODEOWNERS; @vraiti via module of the changed files; @tzhouam via semantic router, CODEOWNERS

@NolenLiang, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 5dae6122f22a produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

NolenLiang commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor Author

Self-review, assisted by Codex: checked cancellation filtering before sequencing, replay recording and on_accepted; shared validity rules for WebSocket/local handles; overflow endings and closure; and the published evidence and limitations.

The MiniCPM-o GB200 comparison observed eight late audio chunks on baseline versus zero with this change; the other original response completed in both. Separate overflow and local-handle runs are documented in the PR. Related checks: 500 passed, one skipped. The published and measured commits have identical source trees.

Feedback is requested on closing the affected session after overflow and allowing clients to finish cleanup from closure when completion markers are omitted. Earlier long-input timeouts and cleanup warnings remain unresolved. This does not establish GPU-stop latency or physical playback behavior.

2026-09-22 UTC: AURA rebase follow-up

Rebased onto c7cd37b and pushed 7ed2bc0. The conflict resolution preserves AURA request-list aborts, draining-response ownership, and completed-turn filtering alongside the closing guards. New regression coverage also exposed cancelled draining audio remaining valid while stage abort was awaited, including silent close; all cancelled draining responses are now invalidated before that await. The overflow policy is unchanged.

Related checks on the published head: 849 passed, 1 skipped. Both new held-audio regression cases failed before the fix. Full applicable pre-commit passes except mypy; candidate and fixed baseline have identical normalized diagnostics (9 errors / 7 notes), so this is not an all-green result.

Real MiniCPM overflow passed all ten scenario checks; the local-handle run received no old audio after cancellation while the other original response continued. Both runs have zero outer/runtime/probe exit status. The new WebSocket baseline/candidate comparison failed on both revisions with stage-2 new_epoch_requires_first_chunk and subsequent request_cleanup closures. Before the failure, cancellation left seven old audio chunks on baseline versus zero on candidate, and the other original response completed in both. These partial observations do not make the pair successful or resolve the stage-2 failure. Shared-memory cleanup warnings remain.

The PR body records the fixed scenarios, results, and limits, with earlier evidence retained separately. Hosted CI is still pending at this update; DCO has passed. Real AURA/Qwen model validation remains outside these measurements.

2026-09-23 UTC: current-main and vLLM 0.30 follow-up

Published 4b4e157 on fixed main 75570ee3, retaining the merged #7714 request-start metrics alongside the output closing guards. This baseline includes #7992; #7995 closed without merging after its author confirmed #7992 addressed the observed MiniCPM failures. No codec recovery workaround is added here.

Related checks on the published head: 870 passed, 1 skipped. Full applicable pre-commit passes except mypy; candidate and fixed baseline have identical normalized diagnostics (10 errors / 7 notes). This remains a baseline typechecking failure, not an all-green result.

The fresh MiniCPM cohort uses official vLLM 0.30.0 ARM64 with Transformers 5.14.1. All four model runs exited successfully. The WebSocket negative control delivered 9 old chunks / 432,000 bytes after accepted cancellation, versus 0 on the candidate; only baseline journaled and delivered the held event. The other original response completed naturally in both runs. Overflow passed all ten checks and a replacement session produced real audio. Direct local handles delivered no old cancelled audio while the other original response completed naturally.

The former stage-2 failure did not recur in this cohort; earlier failed runs remain documented as failed. Shutdown SHM warnings remain (86 objects on each WS side, 26 in overflow, 32 in local handles). This is one controlled pair, not a repeatability, engine-stop latency, physical playback, or leak-free cleanup claim. Real AURA/Qwen model coverage remains separate. The PR body records the exact scenarios and evidence limits; hosted CI for this new head is checked separately.

2026-09-23 UTC: shutdown SHM attribution

A separate instrumented MiniCPM rerun on the same 4b4e157 source completed successfully. The 86 tracker warnings corresponded to 86 physically present, unread Stage 0 connector objects: 116 bytes each, 9,976 bytes total. Creation was traced through _send_single_request_for_generation and SharedMemoryConnector.put; none had a successful reader attach in the observed process tree. They remained present across application shutdown and were subsequently unlinked by Python's resource tracker. All recorded tracker unlinks and process-exit records preceded the final namespace snapshot, which found none of these names remaining.

This attributes this cohort's warnings to unread connector objects, not merely stale registration for already-deleted names. It does not establish a cleanup fix, retained-mapping/FD release, or bounded process RSS. Payloads were not read, and the 116-byte size alone does not identify their contents. The existing #7245/#4349 creation and cleanup work needs to be compared with the observed producer/consumer boundary before proposing another deletion path. No cleanup code or broader claim has been added to this output-delivery PR.

2026-09-24 UTC: hosted checks ready for re-review

The current head 4b4e157 now has successful DCO, pre-commit, Python 3.11 and 3.12 builds, and Read the Docs. The current commit's combined status is successful. This updates the previously pending hosted-check status for the re-review requested above. The real-model results and their limits remain as recorded in the preceding entries; this check-status update does not add new model coverage.

2026-09-25 UTC: cancellation-compensation rebase

Rebased onto fixed main b308e196 and published 5dae6122 after #7785 introduced conflicts in the session manager, Python entrypoint, manager tests and design document. The resolution keeps the upstream cancelled-open rollback, append compensation, resume replay and lease-generation fencing. Output-buffer ownership is removed inside the unified open rollback, before its first await; the two existing cancelled-open tests now cover that ownership, including a second cancellation during shutdown. Existing AURA draining-response routing and pre-abort audio invalidation are preserved.

Related CPU checks in official vLLM 0.30.0 ARM64: 899 passed, 1 skipped, 16 warnings; job/runtime exit status 0. Full applicable pre-commit passes except mypy; candidate and fixed baseline have identical normalized diagnostics (8 errors / 7 notes), with clean working trees. This remains a baseline typechecking failure, not an all-green local result.

No new real-model measurement was performed for this rebase. The MiniCPM observations above remain scoped to their original revisions; CPU AURA/Qwen coverage does not establish real AURA/Qwen model coverage. The prior head was retained in a backup branch. Hosted checks on 5dae6122 are now successful: DCO, pre-commit, Python 3.11 / 3.12 builds, and Read the Docs (verified 2026-09-25 UTC). This updates the prior pending status without adding model coverage. AI assistance: Codex helped resolve the conflicts, verify integration and prepare this update.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 16, 2026
@NolenLiang
NolenLiang force-pushed the feat/duplex-output-delivery branch from 19ddd69 to d0b5699 Compare September 16, 2026 15:32
@linyueqian linyueqian added the ready label to trigger buildkite CI label Sep 16, 2026

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at d0b56997. The buffer itself is sound: admission and release are symmetric for both the ordinary budget and the termination reserve, the reserve is reachable only by ErrorEvent and ResponseDone, session closure keeps its own single slot so a full reserve cannot block it, get() releases the previous result's tracking and holds exactly one dequeued event, and invalidate() removes only queued audio for the cancelled response up to the cancelled epoch while leaving every other event in order. The consumer side does what the description promises: the WebSocket path re-checks the held event under guard() before journaling and sequencing and releases the guard before network I/O, so an invalidated chunk is never recorded for replay and at most one already-committed chunk can precede response.cancelled; the local iterator re-checks with is_valid() after get(). AudioDelta is measured from the base64 length without re-serialising the payload, and to_realtime() builds a fresh dict, so blanking delta on the measured copy is safe. Overflow in _emit_raw marks the session failed once, converts the batch's own ending instead of projecting a second one, lets only errors and failed or cancelled endings through the reserve afterwards, and the runner closes just that session.

Two items inline. The [important] one is an accounting gap rather than a crash: the three item events the projector emits for a new response share one mutable item dict that the projector keeps mutating as transcript and text deltas arrive, so their measured size at admission understates what is pending, and the bound this PR introduces is softer than its number for exactly the events that grow. The [question] is the overflow policy, which the description itself marks as awaiting maintainer agreement; I am leaving the verdict at comment rather than approve until that is settled, since closing the session is a behaviour the duplex owners should choose deliberately.

Validation: static read of the worktree against origin/main plus one external model (codex; grok timed out); no PR code executed. I added ready and the general lane is running at this head.

# directly without serializing the large string or decoding PCM.
audio_size = len(event.delta)
payload["delta"] = ""
size = audio_size + len(json.dumps(payload, ensure_ascii=False, separators=(",", ":")).encode("utf-8"))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[important] The size is measured once at admission, but for the item events it is measured against a dict that keeps changing while the event sits in the queue. realtime_events.py builds one item dict per response, stores it in state.conversation_items, and passes the same object to ItemAdded, ItemCreated and OutputItemAdded (lines 270, 271 and 805); _refresh_in_progress_response_item then mutates that dict in place on every transcript and text delta (called at 833, 838 and 867). So three queued payloads grow together while _bytes still holds their admission size, and the consumer serialises the grown version at delivery time. With ordinary transcripts the drift is kilobytes against a 2 MiB budget, so this is not a correctness failure today, but it means the bound is not the bound this PR documents, and a long text-only response can push real pending bytes well past the limit without ever raising DuplexOutputOverflowError. Two cheap fixes, either of which keeps the accounting honest: snapshot the item at emit time for these three events (the projector already does dict(item) for OutputItemAdded/OutputItemDone at 1062 and 1065, so this is the same pattern applied consistently), or measure again in get() and reconcile the delta before handing the event over. The snapshot also removes a pre-existing oddity where a delayed conversation.item.added carries text the client has not been sent yet.

@NolenLiang NolenLiang Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks—fixed in b2bb100. Both regression cases failed before the fix; the related suite now reports 502 passed, 1 skipped. One real MiniCPM-o WebSocket rerun passed all 11 checks. Overflow/inline were not rerun on this head, and cleanup warnings remain; details are in the PR description.

AMD and Intel CI are red on this head. AMD's blocking Engine&Entrypoints step includes our duplex tests; Intel's configured tests do not select the changed test files. Could someone with log access share the failing step and first error? We cannot yet attribute either failure.

"""Whether ``session.closed`` / ``session.expired`` already left this runner."""
return self.run.closed_emitted

def fail_output(self, message: str, *, response_id: str | None, terminal: ResponseDone | None) -> list[DuplexEvent]:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[question] The description calls this the proposed overflow policy and says maintainer agreement is needed for closing the whole session rather than only failing the active response. The mechanics here are right (fail once, keep the request binding for engine cleanup, close only this session), so my question is to the duplex owners rather than the author: is session closure the behaviour you want on output_backpressure, or should a slow consumer lose the response and keep its session? The answer changes what clients have to implement for recovery, so it is worth deciding before this merges rather than after.

@NolenLiang NolenLiang Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@chickeyton @Sy0307 could you confirm the overflow policy for this PR? It currently fails the response and closes only the affected session; clients may need to finish cleanup from session.closed without ordinary completion markers. Is this acceptable, or should overflow stop only the response and preserve the session?

If response-only recovery is preferred, we will assess the scope and follow up before committing to a change.

@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 17, 2026

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving at b2bb1007. The one [important] item from my first pass is fixed where it belongs: the projector now stores a separate dict in state.conversation_items for the in-progress response item, so the three response-creation events (conversation.item.added, conversation.item.created, response.output_item.added) keep the payload that was measured when they were admitted to the bounded buffer, while later transcript and text deltas grow only the state copy. That closes the accounting gap without adding a copy to audio delivery, and it also removes the older oddity where a delayed creation event carried text the client had not been sent yet. The buffer itself and the consumer-side re-check were already sound in my first read, so the code side of this PR is done from my point of view.

The overflow policy question stands as a product decision rather than a defect: the description still marks closing the whole session on output_backpressure as proposed pending maintainer agreement. I am approving the implementation, and the tower will hold the merge until a duplex owner acknowledges that policy on the thread, since it changes what clients must implement for recovery.

Validation: static comparison of the patch at d0b56997 against this head plus a read of the changed projector path; no PR code executed. I re-fired the general lane at this head after the push.

@linyueqian

Copy link
Copy Markdown
Collaborator

#7640 (the Realtime wire codec move) just landed on main as e3be42e0, and a trial merge of this branch onto it conflicts in vllm_omni/engine/duplex/events.py and docs/design/fullduplex.md: the event classes this PR touches now live in vllm_omni/protocol/duplex/events.py with engine/duplex/events.py reduced to re-exports, so the SessionClosed / ErrorEvent changes here need to move over, and the design-doc paragraphs need re-basing on the new text. A rebase is needed before this can merge; the approval stands on the content, and I will re-approve the rebased head once its lane is green. The overflow-policy question to the duplex owners is still the other open item.

@NolenLiang
NolenLiang force-pushed the feat/duplex-output-delivery branch from b2bb100 to f259bb1 Compare September 17, 2026 14:56
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 17, 2026

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-approving at f259bb13. This is the rebase over the codec move: the branch now sits on e3be42e0, and the patch is byte-identical to the one approved at b2bb1007 except for the mechanical port of the one-line error-code mapping (output_backpressure -> rate_limit_error) from the old engine/duplex/events.py into protocol/duplex/errors.py, with the codec single-source test parametrized to cover it. DCO is green on the new head. Two things still gate the merge, neither of them about this code: the duplex owners' answer on session closure versus response failure on output_backpressure (the question you relayed to chickeyton and Sy0307), and #7732, since main's Qwen3-Omni advanced e2e is red and this touches the duplex session runner; #7059 also landed on runner.py and model_channel.py after your base, so a final rebase onto current main before merging would keep the tested tree and the merged tree identical.

@NolenLiang
NolenLiang force-pushed the feat/duplex-output-delivery branch from f259bb1 to bf04024 Compare September 18, 2026 03:37
@linyueqian

Copy link
Copy Markdown
Collaborator

@NolenLiang one more rebase is needed: #7633 (AURA in the duplex runtime) landed on main as 62b15142 and this branch now conflicts in vllm_omni/engine/duplex/session/model_channel.py in three places, all in the model-output routing path this PR touches. main now aborts through _abort_request([data_plane_request_id], ...) with a list, routes text of a draining response through session.append_draining_assistant_text(target_id, ...) and session.playback_for_response(target_id), and rejects late audio of a completed model turn; your closing early-returns and the _end_active_response_before_future_model_turn call need to sit on top of that. The approval stands on the content; once the rebased head has a green lane I will re-approve and merge. The overflow-policy question stays with the duplex owners and does not block this.

@NolenLiang
NolenLiang force-pushed the feat/duplex-output-delivery branch from bf04024 to 7ed2bc0 Compare September 22, 2026 06:05
@hsliuustc0106 hsliuustc0106 added core related to core module: cache, scheduler, engine, worker, modelrunner frontend code related to entrypoint labels Sep 23, 2026
@NolenLiang
NolenLiang force-pushed the feat/duplex-output-delivery branch from 7ed2bc0 to 4b4e157 Compare September 23, 2026 03:43
Keep pending output bounded per session, reject invalidated audio before delivery, and isolate slow-consumer overflow from other sessions.

Signed-off-by: Nolen Liang <nliang@nvidia.com>
Separate the growing conversation item from the three response-creation events so later text and transcript deltas cannot outgrow their admitted byte count. Cover the real projector-to-buffer path without adding copying to audio delivery. This is limited to response creation, not the wider item snapshot work in vllm-project#7636.

Signed-off-by: Nolen Liang <nliang@nvidia.com>
Signed-off-by: Nolen Liang <nliang@nvidia.com>
Pass the existing output buffer at the upstream-added session-open call sites and share it with the runner harness. Preserve the tests and assertions unchanged.

Signed-off-by: Nolen Liang <nliang@nvidia.com>
Invalidate all cancelled draining responses before awaiting stage abort, including silent session close. Adapt the upstream-added modality error call and AURA test contexts to the bounded-output interface. Keep upstream assertions and cover held and queued draining audio across a paused abort.

Signed-off-by: Nolen Liang <nliang@nvidia.com>
Signed-off-by: bcsdhjew <nliang@nvidia.com>
@NolenLiang
NolenLiang force-pushed the feat/duplex-output-delivery branch from 4b4e157 to 5dae612 Compare September 25, 2026 04:19
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 10 days

@NolenLiang this pull request has had no human commit, comment or review since 2026-09-22. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@Mike001-wq

Copy link
Copy Markdown

Hi @linyueqian, following up on your note about re-approving and merging after the rebase and a green lane.
Current head is 5dae612, incorporating the AURA and cancellation-compensation changes. All five publicly visible checks are successful. The latest rebase validation reports 899 passed and 1 skipped; the evidence and limitations are documented in the PR.
Could you re-review this head, confirm or trigger any required Buildkite lane, and merge once the remaining checks and review threads are cleared? Per your last update, the overflow-policy discussion remains non-blocking.
Thanks!

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have re-reviewed the changes at 5dae6122. The rebase preserves AURA response ownership, list-form aborts, and upstream cancellation compensation. Active and draining audio are invalidated before awaiting stage abort, including silent close, and the cancelled-open rollback removes output-buffer ownership before its first await. The earlier finding on creation-payload accounting remains fixed. I did not find any material issues in this delta, and the overflow-policy discussion remains non-blocking as previously noted.

Validation consisted of a static comparison of the old and new PR footprints, affected call sites, and regression assertions, along with a passing git diff --check. No PR code or tests were executed, and the reported CPU results and historical model measurements were not independently reproduced.

@linyueqian
linyueqian merged commit f781f7b into vllm-project:main Oct 4, 2026
6 of 9 checks passed
@linyueqian

Copy link
Copy Markdown
Collaborator

The post-merge main CUDA build 16680 is failing in Engine & Entrypoints, with eight PersonaPlex test failures and four Nemotron fixture errors that all report a missing OpenDuplexSessionMessage.output_buffer. The reviewed PR build 16676 passed, but these two model-specific harness files landed on main after this PR's base and were absent from its tested head. I missed that merge interaction.

The repair is confined to tests/engine/duplex/test_session_runner_personaplex.py and tests/engine/duplex/test_session_runner_nemotron_voicechat.py, requiring us to create a DuplexOutputBuffer using the same runtime limits as the manager and then pass that buffer to both the open message and the shared Harness. Since the production constructor already passes it, the required API should stay intact. I have the two-file patch prepared and local pre-commit passes, but the affected tests and main CUDA still need to run with it, so I am holding the remaining merge queue until this is repaired and main is green.

NumberWan added a commit to NumberWan/vllm-omni that referenced this pull request Oct 5, 2026
Keep session.updated default application and pass the session handle into
_send_event, which vllm-project#7644 made required.

Signed-off-by: NumberWan <wantszkin2003@gmail.com>
chickeyton added a commit to chickeyton/vllm-omni that referenced this pull request Oct 5, 2026
Brings in 177 upstream commits (through vllm-project#8489). Conflicts were all the
same shape: upstream imports from engine.duplex.events / .commands /
.realtime_events, which this branch removed, so the merged files import
the same names from protocol.duplex.events / protocol.duplex.commands /
engine.duplex.projection, and to_realtime() reads to_wire().

- engine/duplex/commands.py stays deleted. Upstream's vllm-project#7985 change there
  (an explicit empty __slots__ on the engine DuplexCommand mixin for
  Python 3.10) has no counterpart here because DuplexCommand is the wire
  base itself; test_duplex_commands.py pins the single-slotted-base layout
  instead of the removed mixin test.
- The new delivery.py (vllm-project#7644) and its tests are repointed the same way.
- test_realtime_codec_single_source.py takes upstream's parametrized
  error-type test and drops the engine-commands subclass test, whose
  subject no longer exists.
- docs/design/fullduplex.md keeps this branch's module tree plus the new
  delivery.py entry.

Signed-off-by: chickeyton <ngton2014@gmail.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Zssan12 added a commit to Zssan12/vllm-omni that referenced this pull request Oct 9, 2026
Session-owned duplex events now go to the per-session DuplexOutputBuffer (vllm-project#7644) rather than the orchestrator output queue. Open the session with a buffer and read the events from it, as the other duplex orchestrator tests do.

Signed-off-by: zxy <1335108150@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner frontend code related to entrypoint high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants