Repository navigation
Conversation
Signed-off-by: luozijian <luozijian0924@gmail.com>
|
Contributor-side self-review for signed commit
AI assistance: ChatGPT assisted with implementation, tests and these review notes. |
|
This PR appears to belong to: docs/design/module/profiling.md, docs/design/module/entrypoints.md, docs/design/module/observability.md. Module owners: @alex-jw-brooks @linyueqian @NickCao Routing: @alex-jw-brooks via module of the changed files, module named in the PR description, semantic router, CODEOWNERS; @linyueqian via module of the changed files, module named in the PR description, semantic router, CODEOWNERS; @NickCao via module of the changed files, module named in the PR description, semantic router, CODEOWNERS @LOGO127, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
|
@LOGO127 this landed on |
|
Thanks for the pointer. I checked current I prepared a minimal test-only follow-up on current I’ll keep #7429 as the historical production fix for now and move the remaining value into that test-only follow-up rather than trying to resolve the production conflict here. |
|
Thanks for clarifying the new framework boundary. I have confirmed the production rollback is already covered by merged #7413. The independent four-case cancellation regression (attached takeover / detached reconnect, each during activation / replay) is now isolated in test-only #8200, so I am closing this older conflicting implementation rather than carrying duplicate production changes. The current #8200 attachment test file passes 25 supplemental CPU leaf tests and all applicable pre-commit hooks. That is not a full native-package or GPU result; the local installed vLLM/current-Omni API mismatch is recorded separately. #8200 remains a draft pending personal review/DCO completion. No additional production fix is claimed here. |
Purpose
Treat cancellation during duplex resume activation/replay like the existing failed-delivery rollback, then propagate CancelledError. Without it the rotated, abandoned connection generation remains current and the prior token cannot use the intended retry path.
This is a two-file fix against current main 02aaa34; it does not include PR #7413's framework changes. The same issue was discussed with its author, and a separately validated author-branch adaptation is preserved. This main version uses the current incarnation-based API.
Test Plan and Results
python -m pytest tests/entrypoints/openai/test_duplex_session_attachment.py tests/engine/duplex/test_duplex_lease.py tests/clients/test_duplex_client.py -m 'core_model and cpu' --run-level core_model -qRetained output: 63 passed, 15 warnings in 0.78s (rerun on signed head e07daf3). Python 3.12 CPU environment, actual registry/lease/client code; the local adapter disables NVML discovery only. No GPU or live model run is claimed. Earlier 66-test results against #7413 are not reused as current-main results.
AI assistance: ChatGPT assisted with implementation, tests, and this description.