Repository navigation
[Core][Frontend] Bound duplex output delivery and reject cancelled audio - #7644
Conversation
5592f8c to
19ddd69
Compare
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
This PR appears to belong to: docs/design/module/observability.md, docs/design/module/engine_orchestration.md. Module owners: @lishunyang12 @vraiti @tzhouam Routing: @lishunyang12 via module of the changed files, CODEOWNERS; @vraiti via module of the changed files; @tzhouam via semantic router, CODEOWNERS @NolenLiang, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
|
Self-review, assisted by Codex: checked cancellation filtering before sequencing, replay recording and The MiniCPM-o GB200 comparison observed eight late audio chunks on baseline versus zero with this change; the other original response completed in both. Separate overflow and local-handle runs are documented in the PR. Related checks: 500 passed, one skipped. The published and measured commits have identical source trees. Feedback is requested on closing the affected session after overflow and allowing clients to finish cleanup from closure when completion markers are omitted. Earlier long-input timeouts and cleanup warnings remain unresolved. This does not establish GPU-stop latency or physical playback behavior. 2026-09-22 UTC: AURA rebase follow-upRebased onto Related checks on the published head: 849 passed, 1 skipped. Both new held-audio regression cases failed before the fix. Full applicable pre-commit passes except mypy; candidate and fixed baseline have identical normalized diagnostics (9 errors / 7 notes), so this is not an all-green result. Real MiniCPM overflow passed all ten scenario checks; the local-handle run received no old audio after cancellation while the other original response continued. Both runs have zero outer/runtime/probe exit status. The new WebSocket baseline/candidate comparison failed on both revisions with stage-2 The PR body records the fixed scenarios, results, and limits, with earlier evidence retained separately. Hosted CI is still pending at this update; DCO has passed. Real AURA/Qwen model validation remains outside these measurements. 2026-09-23 UTC: current-main and vLLM 0.30 follow-upPublished Related checks on the published head: 870 passed, 1 skipped. Full applicable pre-commit passes except mypy; candidate and fixed baseline have identical normalized diagnostics (10 errors / 7 notes). This remains a baseline typechecking failure, not an all-green result. The fresh MiniCPM cohort uses official vLLM 0.30.0 ARM64 with Transformers 5.14.1. All four model runs exited successfully. The WebSocket negative control delivered 9 old chunks / 432,000 bytes after accepted cancellation, versus 0 on the candidate; only baseline journaled and delivered the held event. The other original response completed naturally in both runs. Overflow passed all ten checks and a replacement session produced real audio. Direct local handles delivered no old cancelled audio while the other original response completed naturally. The former stage-2 failure did not recur in this cohort; earlier failed runs remain documented as failed. Shutdown SHM warnings remain (86 objects on each WS side, 26 in overflow, 32 in local handles). This is one controlled pair, not a repeatability, engine-stop latency, physical playback, or leak-free cleanup claim. Real AURA/Qwen model coverage remains separate. The PR body records the exact scenarios and evidence limits; hosted CI for this new head is checked separately. 2026-09-23 UTC: shutdown SHM attributionA separate instrumented MiniCPM rerun on the same This attributes this cohort's warnings to unread connector objects, not merely stale registration for already-deleted names. It does not establish a cleanup fix, retained-mapping/FD release, or bounded process RSS. Payloads were not read, and the 116-byte size alone does not identify their contents. The existing #7245/#4349 creation and cleanup work needs to be compared with the observed producer/consumer boundary before proposing another deletion path. No cleanup code or broader claim has been added to this output-delivery PR. 2026-09-24 UTC: hosted checks ready for re-reviewThe current head 2026-09-25 UTC: cancellation-compensation rebaseRebased onto fixed main Related CPU checks in official vLLM 0.30.0 ARM64: 899 passed, 1 skipped, 16 warnings; job/runtime exit status 0. Full applicable pre-commit passes except mypy; candidate and fixed baseline have identical normalized diagnostics (8 errors / 7 notes), with clean working trees. This remains a baseline typechecking failure, not an all-green local result. No new real-model measurement was performed for this rebase. The MiniCPM observations above remain scoped to their original revisions; CPU AURA/Qwen coverage does not establish real AURA/Qwen model coverage. The prior head was retained in a backup branch. Hosted checks on |
19ddd69 to
d0b5699
Compare
linyueqian
left a comment
There was a problem hiding this comment.
Reviewed at d0b56997. The buffer itself is sound: admission and release are symmetric for both the ordinary budget and the termination reserve, the reserve is reachable only by ErrorEvent and ResponseDone, session closure keeps its own single slot so a full reserve cannot block it, get() releases the previous result's tracking and holds exactly one dequeued event, and invalidate() removes only queued audio for the cancelled response up to the cancelled epoch while leaving every other event in order. The consumer side does what the description promises: the WebSocket path re-checks the held event under guard() before journaling and sequencing and releases the guard before network I/O, so an invalidated chunk is never recorded for replay and at most one already-committed chunk can precede response.cancelled; the local iterator re-checks with is_valid() after get(). AudioDelta is measured from the base64 length without re-serialising the payload, and to_realtime() builds a fresh dict, so blanking delta on the measured copy is safe. Overflow in _emit_raw marks the session failed once, converts the batch's own ending instead of projecting a second one, lets only errors and failed or cancelled endings through the reserve afterwards, and the runner closes just that session.
Two items inline. The [important] one is an accounting gap rather than a crash: the three item events the projector emits for a new response share one mutable item dict that the projector keeps mutating as transcript and text deltas arrive, so their measured size at admission understates what is pending, and the bound this PR introduces is softer than its number for exactly the events that grow. The [question] is the overflow policy, which the description itself marks as awaiting maintainer agreement; I am leaving the verdict at comment rather than approve until that is settled, since closing the session is a behaviour the duplex owners should choose deliberately.
Validation: static read of the worktree against origin/main plus one external model (codex; grok timed out); no PR code executed. I added ready and the general lane is running at this head.
| # directly without serializing the large string or decoding PCM. | ||
| audio_size = len(event.delta) | ||
| payload["delta"] = "" | ||
| size = audio_size + len(json.dumps(payload, ensure_ascii=False, separators=(",", ":")).encode("utf-8")) |
There was a problem hiding this comment.
[important] The size is measured once at admission, but for the item events it is measured against a dict that keeps changing while the event sits in the queue. realtime_events.py builds one item dict per response, stores it in state.conversation_items, and passes the same object to ItemAdded, ItemCreated and OutputItemAdded (lines 270, 271 and 805); _refresh_in_progress_response_item then mutates that dict in place on every transcript and text delta (called at 833, 838 and 867). So three queued payloads grow together while _bytes still holds their admission size, and the consumer serialises the grown version at delivery time. With ordinary transcripts the drift is kilobytes against a 2 MiB budget, so this is not a correctness failure today, but it means the bound is not the bound this PR documents, and a long text-only response can push real pending bytes well past the limit without ever raising DuplexOutputOverflowError. Two cheap fixes, either of which keeps the accounting honest: snapshot the item at emit time for these three events (the projector already does dict(item) for OutputItemAdded/OutputItemDone at 1062 and 1065, so this is the same pattern applied consistently), or measure again in get() and reconcile the delta before handing the event over. The snapshot also removes a pre-existing oddity where a delayed conversation.item.added carries text the client has not been sent yet.
There was a problem hiding this comment.
Thanks—fixed in b2bb100. Both regression cases failed before the fix; the related suite now reports 502 passed, 1 skipped. One real MiniCPM-o WebSocket rerun passed all 11 checks. Overflow/inline were not rerun on this head, and cleanup warnings remain; details are in the PR description.
AMD and Intel CI are red on this head. AMD's blocking Engine&Entrypoints step includes our duplex tests; Intel's configured tests do not select the changed test files. Could someone with log access share the failing step and first error? We cannot yet attribute either failure.
| """Whether ``session.closed`` / ``session.expired`` already left this runner.""" | ||
| return self.run.closed_emitted | ||
|
|
||
| def fail_output(self, message: str, *, response_id: str | None, terminal: ResponseDone | None) -> list[DuplexEvent]: |
There was a problem hiding this comment.
[question] The description calls this the proposed overflow policy and says maintainer agreement is needed for closing the whole session rather than only failing the active response. The mechanics here are right (fail once, keep the request binding for engine cleanup, close only this session), so my question is to the duplex owners rather than the author: is session closure the behaviour you want on output_backpressure, or should a slow consumer lose the response and keep its session? The answer changes what clients have to implement for recovery, so it is worth deciding before this merges rather than after.
There was a problem hiding this comment.
@chickeyton @Sy0307 could you confirm the overflow policy for this PR? It currently fails the response and closes only the affected session; clients may need to finish cleanup from session.closed without ordinary completion markers. Is this acceptable, or should overflow stop only the response and preserve the session?
If response-only recovery is preferred, we will assess the scope and follow up before committing to a change.
linyueqian
left a comment
There was a problem hiding this comment.
Approving at b2bb1007. The one [important] item from my first pass is fixed where it belongs: the projector now stores a separate dict in state.conversation_items for the in-progress response item, so the three response-creation events (conversation.item.added, conversation.item.created, response.output_item.added) keep the payload that was measured when they were admitted to the bounded buffer, while later transcript and text deltas grow only the state copy. That closes the accounting gap without adding a copy to audio delivery, and it also removes the older oddity where a delayed creation event carried text the client had not been sent yet. The buffer itself and the consumer-side re-check were already sound in my first read, so the code side of this PR is done from my point of view.
The overflow policy question stands as a product decision rather than a defect: the description still marks closing the whole session on output_backpressure as proposed pending maintainer agreement. I am approving the implementation, and the tower will hold the merge until a duplex owner acknowledges that policy on the thread, since it changes what clients must implement for recovery.
Validation: static comparison of the patch at d0b56997 against this head plus a read of the changed projector path; no PR code executed. I re-fired the general lane at this head after the push.
|
#7640 (the Realtime wire codec move) just landed on main as |
b2bb100 to
f259bb1
Compare
linyueqian
left a comment
There was a problem hiding this comment.
Re-approving at f259bb13. This is the rebase over the codec move: the branch now sits on e3be42e0, and the patch is byte-identical to the one approved at b2bb1007 except for the mechanical port of the one-line error-code mapping (output_backpressure -> rate_limit_error) from the old engine/duplex/events.py into protocol/duplex/errors.py, with the codec single-source test parametrized to cover it. DCO is green on the new head. Two things still gate the merge, neither of them about this code: the duplex owners' answer on session closure versus response failure on output_backpressure (the question you relayed to chickeyton and Sy0307), and #7732, since main's Qwen3-Omni advanced e2e is red and this touches the duplex session runner; #7059 also landed on runner.py and model_channel.py after your base, so a final rebase onto current main before merging would keep the tested tree and the merged tree identical.
f259bb1 to
bf04024
Compare
|
@NolenLiang one more rebase is needed: #7633 (AURA in the duplex runtime) landed on |
bf04024 to
7ed2bc0
Compare
7ed2bc0 to
4b4e157
Compare
Keep pending output bounded per session, reject invalidated audio before delivery, and isolate slow-consumer overflow from other sessions. Signed-off-by: Nolen Liang <nliang@nvidia.com>
Separate the growing conversation item from the three response-creation events so later text and transcript deltas cannot outgrow their admitted byte count. Cover the real projector-to-buffer path without adding copying to audio delivery. This is limited to response creation, not the wider item snapshot work in vllm-project#7636. Signed-off-by: Nolen Liang <nliang@nvidia.com>
Signed-off-by: Nolen Liang <nliang@nvidia.com>
Pass the existing output buffer at the upstream-added session-open call sites and share it with the runner harness. Preserve the tests and assertions unchanged. Signed-off-by: Nolen Liang <nliang@nvidia.com>
Invalidate all cancelled draining responses before awaiting stage abort, including silent session close. Adapt the upstream-added modality error call and AURA test contexts to the bounded-output interface. Keep upstream assertions and cover held and queued draining audio across a paused abort. Signed-off-by: Nolen Liang <nliang@nvidia.com>
Signed-off-by: bcsdhjew <nliang@nvidia.com>
4b4e157 to
5dae612
Compare
Omni ReviewBot: no human activity for 10 days@NolenLiang this pull request has had no human commit, comment or review since 2026-09-22. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
|
Hi @linyueqian, following up on your note about re-approving and merging after the rebase and a green lane. |
linyueqian
left a comment
There was a problem hiding this comment.
I have re-reviewed the changes at 5dae6122. The rebase preserves AURA response ownership, list-form aborts, and upstream cancellation compensation. Active and draining audio are invalidated before awaiting stage abort, including silent close, and the cancelled-open rollback removes output-buffer ownership before its first await. The earlier finding on creation-payload accounting remains fixed. I did not find any material issues in this delta, and the overflow-policy discussion remains non-blocking as previously noted.
Validation consisted of a static comparison of the old and new PR footprints, affected call sites, and regression assertions, along with a passing git diff --check. No PR code or tests were executed, and the reported CPU results and historical model measurements were not independently reproduced.
|
The post-merge main CUDA build 16680 is failing in Engine & Entrypoints, with eight PersonaPlex test failures and four Nemotron fixture errors that all report a missing The repair is confined to |
Keep session.updated default application and pass the session handle into _send_event, which vllm-project#7644 made required. Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Brings in 177 upstream commits (through vllm-project#8489). Conflicts were all the same shape: upstream imports from engine.duplex.events / .commands / .realtime_events, which this branch removed, so the merged files import the same names from protocol.duplex.events / protocol.duplex.commands / engine.duplex.projection, and to_realtime() reads to_wire(). - engine/duplex/commands.py stays deleted. Upstream's vllm-project#7985 change there (an explicit empty __slots__ on the engine DuplexCommand mixin for Python 3.10) has no counterpart here because DuplexCommand is the wire base itself; test_duplex_commands.py pins the single-slotted-base layout instead of the removed mixin test. - The new delivery.py (vllm-project#7644) and its tests are repointed the same way. - test_realtime_codec_single_source.py takes upstream's parametrized error-type test and drops the engine-commands subclass test, whose subject no longer exists. - docs/design/fullduplex.md keeps this branch's module tree plus the new delivery.py entry. Signed-off-by: chickeyton <ngton2014@gmail.com> Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Session-owned duplex events now go to the per-session DuplexOutputBuffer (vllm-project#7644) rather than the orchestrator output queue. Open the session with a buffer and read the events from it, as the other duplex orchestrator tests do. Signed-off-by: zxy <1335108150@qq.com>
Purpose
Follow up on merged #7413 with bounded per-session output and rejection of cancelled audio before delivery. This is an independent change against
main, rebased ontob308e196, with head5dae6122; upstream changes are not part of this diff. The rebase preserves #7785's cancelled-open rollback, append compensation, and resume/lease-generation handling. Output-buffer ownership is released synchronously in the unified open rollback, before shutdown can be cancelled again. This alignment preserves the server-side request-start metrics from #7714 alongside the existing closing guards, AURA request-list aborts, draining-response ownership, and completed-turn audio filtering. Cancellation now invalidates every cancelled draining response before awaiting stage abort, including silent session close. The overflow policy and creation-event snapshot fix are unchanged.This implements the output-delivery work discussed in the proposal and the author's response. Real-model coverage is limited to MiniCPM-o 4.5; the buffer applies to the shared session layer. Real Qwen model validation and general playback/history policy are outside this change's scope.
Problem and fix
The shared engine output queue and per-handle queues are unbounded. Cancelling generation does not remove audio already waiting for delivery, so old audio can arrive before cancellation notifications. Blocking the shared pump on a slow consumer would also delay other sessions.
on_accepted. Local Python iteration uses the same validity state.Behavior and limits
output_backpressureand close only the affected session. Keeping that session open after stopping just its response is not implemented. The duplex-owner policy discussion remains open; the latest reviewer update explicitly says it does not block this PR.close_details_exceed_output_limit. This guarantees local closure capacity, not network delivery or successful engine-resource release.on_accepted. Callback coverage does not establish that nativesent_msmeasures delivered or physically played audio.Test Plan
vLLM Version: Official vLLM 0.30.0 ARM64 environment, Transformers 5.14.1, and model-specific speech dependencies.
vLLM-Omni Commit: Current fixed baseline
b308e196; candidate5dae6122. The current rebase has CPU contract coverage in the official vLLM 0.30.0 ARM64 runtime. Real-model results below remain attributed to their recorded revisions, including75570ee3/4b4e157; they are not new-head GPU measurements. The earlier regression-only revision879f8b4records the AURA assertions before their integration fixes on the previous baseline.Real-model runs use
openbmb/MiniCPM-o-4_5at503e754, one GB200 per service and the matching MiniCPM deploy configuration, without overriding its compilation settings. WebSocket runs use the production duplex handler in a small loopback application, not the complete OpenAI application.Current fixed scenarios:
output_backpressurestill serializes asrate_limit_errorthrough the shared duplex codec.Related checks, in the matching runtime with repository test dependencies:
Test Result
Feedback is requested on the overflow policy and completion behavior; neither policy changed in the review follow-up.
Cancellation-compensation rebase at 5dae612
Rebased onto fixed main
b308e196after #7785. The four textual conflicts were resolved without replacing the upstream cancellation-compensation paths. The existing AURA draining-response routing and pre-abort audio invalidation are retained. Two upstream cancelled-open tests now receive the session output buffer, with assertions that its ownership is removed on cancellation and before a second cancellation during rollback.5dae6122are now successful: DCO, pre-commit, Python 3.11 / 3.12 builds, and Read the Docs (verified 2026-09-25 UTC). This does not change the baseline mypy caveat or add model coverage.Metrics integration at 4b4e157
This revision retains the merged #7714 timing behavior and #7992 prewarm fix. The earlier
new_epoch_requires_first_chunkfailure is also tracked in #7979. Upstream selected #7992; #7995 closed without merging after its author confirmed the narrower fix. No codec recovery workaround is added here.4b4e157: 870 passed, 1 skipped, 16 warnings, with zero job/runtime exit status. This includes CPU AURA/Qwen contract coverage, not real AURA/Qwen model validation.output_backpressureerror, exactly one failed/uncommitted original-response ending, then closure of that session. Held audio was absent from journal/send/receive. The other original response produced three chunks while held and five after closure; a distinct replacement produced 40,320 bytes of real audio and restored two active sessions. Normal and replacement sessions were explicitly closed after observation, without waiting for natural completion.new_epoch_requires_first_chunkfailure did not recur in this cohort on the new upstream baseline. Earlier failed runs remain failed evidence. Shared-memory shutdown warnings remain: 86 objects in each WebSocket comparison run, 26 in overflow, and 32 in local handles. These tracker counts are not measurements of unreclaimed bytes after process exit. This validation does not establish leak-free cleanup, physical playback behavior, or engine-stop latency.The previous AURA and main-alignment measurements used vLLM 0.29.0. Older evidence remains scoped to its recorded runtime and revisions; it is not vLLM 0.30 validation. Historical outcomes and limits remain unchanged.
Previous AURA integration validation at 7ed2bc0
AURA integration at 7ed2bc0
7ed2bc0: 849 passed, 1 skipped, 16 warnings, with zero outer/runtime exit status. The selection includes AURA engine, plugin, and history tests as well as the existing Qwen plugin tests. This is CPU contract coverage, not real AURA/Qwen model coverage.MiniCPMO45Code2WavBatchErrorwith reasonnew_epoch_requires_first_chunk, followed byrequest_cleanupclosures. Both have nonzero outer/runtime/probe status. Ten of eleven scenario checks passed; the no-errors/no-unexpected-close check failed. Before that failure, old audio after cancellation was 7 chunks / 336,000 decoded bytes on baseline versus 0 on candidate; the held event entered replay/delivery only on baseline, and the other original response completed naturally in both. These are partial observations from failed runs, not a successful pair or evidence that the stage-2 failure is fixed.output_backpressure, one failed original-response ending, then closure of the affected session. Held audio was absent from replay/send/receive records. The other original response produced three chunks while held and six after the affected client's closure; a distinct replacement session produced audio and restored two active sessions. The normal and replacement sessions were explicitly closed after observation, without waiting for natural completion.The regression-only revision reported 843 passed, 6 failed, 1 skipped. Both new cancellation cases failed at the held-audio validity assertion before the fix. The other four failures identified upstream-added call sites needing adaptation: two AURA test contexts lacked closing state, a session-manager test omitted awaiting the buffer-based event reader, and the new input-modality error path passed a single event to the list-based emitter. The latter test now also reads the same bounded output buffer as the session. Original assertions are retained.
The earlier Qwen advanced-E2E tracking issue #7732 is now closed. The CPU AURA/Qwen tests and real MiniCPM observations here do not establish real AURA or Qwen model coverage.
Previous main-alignment validation at bf04024
Validation at bf04024
Four real-model runs completed on
2ab5d17/245942b: the baseline/candidate WebSocket pair, candidate overflow, and candidate local handles. All four had zero outer/runtime/probe exit status. The final head only adapts existing tests; these are not claimed as model runs onbf04024.bf04024: 697 passed, 1 skipped, 16 warnings, with zero outer/runtime exit status. The selection now also includes the existing Qwen plugin test file; this is not real Qwen model or advanced E2E coverage.245942breported 673 passed, 1 failed, 1 skipped. The failure was an upstream-added test constructing a session-open message without the required output buffer. Both newly added call sites are now adapted, including the Qwen test harness sharing that same buffer; no new tests or assertions were added. The final run above includes the formerly failing case.Both WebSocket runs passed all eleven scenario checks. Counts exclude subsequent response IDs, and the other original response completed before explicit cleanup. This is one pair, not a repeatability or latency claim.
Candidate overflow passed all ten conditions:
output_backpressureerror, one failed original response ending, then closure of only the affected session. The held audio was neither journaled nor delivered. The other original response produced four chunks while held and five after the affected client's closure; a distinct replacement session produced audio and restored two active sessions. Normal and replacement sessions were explicitly closed, not observed to finish naturally.Candidate local-handle delivery received 0 old audio chunks / 0 bytes after cancellation. The other original response produced ten chunks while the slow reader was paused and three after the cancelled ending; this time it completed naturally before explicit handle cleanup. This exercises direct handles, not a complete InlineDuplexClient application.
Shared-memory cleanup warnings remain in all four runs. The deliberate server-side hold is not a physically slow network experiment, and the earlier long-input timeouts below remain unresolved. These MiniCPM observations do not validate the advanced Qwen E2E cases in #7732; that merge restriction remains subject to maintainer resolution or explicit approval.
Previous protocol-move validation at f259bb1
Rebased validation at f259bb1
The new baseline/candidate WebSocket pair, candidate overflow, candidate local-handle run, and related checks each ran once in the runtime above. All five completed with zero outer/runtime exit status; all four model probes also exited successfully.
tests/protocol/, so the count is not directly comparable to the earlier 502-case selection.Both runs passed all eleven scenario checks without errors or unexpected closure. The other original responses completed before explicit cleanup; later response IDs are excluded from these counts. This is one new pair, not a repeatability or latency claim.
Candidate WebSocket overflow passed all ten conditions:
output_backpressureerror, one failed original response ending, then closure of only the affected session. The held audio was absent from replay/send/receive records. The other original response produced four chunks while the send was held and five after the affected client's closure; a distinct replacement session produced audio and restored two active sessions. The other and replacement sessions were explicitly closed, not observed to finish naturally.Candidate local-handle delivery received 0 old audio chunks / 0 bytes after cancellation. The other session produced nine chunks during the pause and three after the cancelled ending. Both handles were explicitly closed; this is direct-handle coverage, not a complete InlineDuplexClient application.
Shared-memory cleanup warnings remain in all four model runs. Successful exit does not establish leak-free cleanup, physical playback behavior, or engine-stop latency. The deliberate server-side hold is not a physically slow network experiment. Earlier long-input timeouts below remain unresolved.
Validation before the protocol-move rebase
Review follow-up at b2bb100
d0b5699production code. Onb2bb100, the related-check command without the subsequently addedtests/protocol/selection completed with 502 passed, 1 skipped, 16 warnings, including both cases. To run only the regression:python -m pytest -q tests/engine/duplex/test_delivery.py -k queued_response_creation.realtime_events.py, so these totals are not comparable to the earlier 28-error selection. Neither run is all green, and neither working tree was changed.Original post-merge validation at d0b5699
All five original fixed runs completed with zero outer/runtime exit status: four real-model scenarios and one related-check run, each submitted once. The following results belong to
d0b5699, not the follow-up commit.d0b5699: 500 passed, 1 skipped, 16 warnings.Both runs satisfied all eleven scenario checks. Their other original responses completed before explicit cleanup; later response IDs were excluded from those counts. Client audio records matched server records by event/response identity, size and content digest. This is one pair, not a repeatability claim.
Candidate WebSocket overflow satisfied all ten conditions:
output_backpressureerror, one failed original response ending, then session closure; held audio absent from replay/send/receive records; the other original response continued across closure; a distinct replacement session produced 40,320 decoded audio bytes and restored two active sessions. Normal and replacement sessions were explicitly closed afterward, not observed to finish naturally in this scenario.Candidate local-handle observation received no old audio after cancellation; the other session continued with nine chunks during the pause and three after the cancelled ending. Both handles were explicitly closed. This exercised direct handles, not a complete InlineDuplexClient application.
Shared-memory cleanup warnings remained in all four model runs. Successful exits do not establish leak-free cleanup.
Earlier fixed-checkpoint evidence and unresolved observations
On
8ae42c6versuscf30cd2, the short-input WebSocket pair observed eight old audio chunks / 384,000 decoded bytes after cancellation submission on baseline versus zero on candidate. The held event was recorded for replay on baseline and not on candidate. Both other original responses completed naturally. This used three input repetitions and does not explain earlier failures with twelve repetitions.On
c1dff45versuscd36fe4, a separate four-run long-input cohort completed only the candidate/cancellation condition; baseline/cancellation and both no-cancellation conditions timed out while the other response continued producing audio without a recorded natural ending. All failures remain failures. These observations do not identify the cause or exclude an effect from this change.An earlier
cd36fe4overflow observer passed its ten checks, but its outer launcher exited with a shell error: that overall run was not clean. The current launch chain and observers are frozen together to avoid changes during execution. Shared-memory cleanup warnings in earlier model runs remain unresolved.These are scenario observations, not performance comparisons. ASGI send return is not physical playback; deliberate pauses do not measure engine-stop latency. Sustained memory bounds, leak-free cleanup, speech quality, and reliable long-response completion are not established.
AI assistance: Codex helped with implementation, validation, and writing.