Repository navigation
Conversation
Keep cancelled-fence request resources until stage cleanup completes so session close or expiry can retry a failed cleanup. Add a regression test covering failure followed by close cleanup. Signed-off-by: aayc123 <2803023760@qq.com>
|
This PR appears to belong to: docs/design/module/engine_orchestration.md, docs/design/module/observability.md. Module owners: @fake0fan @tzhouam @lishunyang12 Routing: @fake0fan via module of the changed files, semantic router, CODEOWNERS; @tzhouam via module of the changed files, semantic router, CODEOWNERS; @lishunyang12 via module of the changed files @aayc123, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
|
Self-review:
|
Signed-off-by: aayc123 <2803023760@qq.com>
Signed-off-by: aayc123 <115020989+aayc123@users.noreply.github.com>
Signed-off-by: aayc123 <2803023760@qq.com>
|
@hsliuustc0106 Thanks for the review. I removed the cancel_fence wrapper as suggested and updated the affected call sites to use the explicit prepare_cancel_fence -> cleanup -> release_fence flow. The conflict with main has also been resolved, and all CI checks now pass. Could you please take another look when you have time? |
Omni ReviewBot: no human activity for 7 days@aayc123 this pull request has had no human commit, comment or review since 2026-09-24. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
Omni ReviewBot routing recordAssigned Strict on zcode (GLM-5.3-Flash) under experiment |
Purpose
Addresses issue 8 in #7636 and the cleanup lifecycle concern in:
#7413 (comment)
Keep cancelled-fence request resources until stage cleanup completes so session close or expiry can retry a failed cleanup.
Cancellation now: prepare fence cancellation, await
stage_port.cleanup(..., abort=True), then release the cancelled fence’s resources only after cleanup succeeds. The accepted fence still advances before cleanup, so old-epoch outputs remain stale. No new retry mechanism, locking, or session state.Test Plan
Add a regression test covering failure followed by close cleanup retry. It verifies epoch advances,
runtime_signal_failedis emitted, the cancelled request ID remains registered, close retries cleanup with the same request ID andabort=True, and the resource is released after retry.pytest -q \ tests/engine/duplex/test_session_runner.py::test_stale_epoch_output_is_dropped_after_barge_in \ tests/engine/duplex/test_session_runner.py::test_cancel_cleanup_failure_retains_requests_for_close_retry \ tests/engine/duplex/test_session_runner.py::test_close_emits_session_closed_and_releases_stage_requests pytest -q tests/engine/duplex -m "core_model and cpu"vLLM Version: 0.29.0
vLLM-Omni Commit: 8fe5b42
Test Result
3 passed345 passed, 1 skippedmodel_channel.py:753,model_channel.py:797,test_session_runner.py:187,test_session_runner.py:1070.AI Assistance
Used OpenAI Codex to assist with issue analysis, implementation, regression testing. I reviewed and understand the final changes and validated them locally.