Skip to content

[Bugfix] Prevent DiffusionResultPump crash on cancelled futures (#5793) - #5983

Merged
Gaohan123 merged 3 commits into
vllm-project:mainfrom
anurag12-webster:bugfix-5793-result-pump-cancelled-future-race
Aug 15, 2026
Merged

Gaohan123 merged 3 commits into
vllm-project:mainfrom
anurag12-webster:bugfix-5793-result-pump-cancelled-future-race

Conversation

@anurag12-webster

Copy link
Copy Markdown
Contributor

What broke

MultiprocDiffusionExecutor._result_pump is the sole reader of the diffusion worker result queue, running as a singleton background thread. Every dispatch branch resolves a shared future with:

if fut is not None and not fut.done():
    fut.set_result(result)  # or fut.set_exception(...)

The not fut.done() check and the set_result()/set_exception() call are not atomic. A consumer awaiting the future via asyncio.wrap_future() under asyncio.wait_for() (an inter-output timeout, or a request abort) can cancel it in the gap between the two. When that happens, set_result()/set_exception() raises concurrent.futures.InvalidStateError, which nothing catches — killing the pump thread permanently. Since the pump is a singleton, the engine then accepts and "runs" jobs that silently never complete (/health stays 200, GPU sits idle).

Repro

concurrent.futures._base.InvalidStateError: CANCELLED: <Future at 0x... state=cancelled>
  File "vllm_omni/diffusion/executor/multiproc_executor.py", line 883, in _result_pump
    fut.set_result(msg.output)
  File ".../concurrent/futures/_base.py", line 544, in set_result
    raise InvalidStateError(...)

Reported originally in #5793 via two real triggers: a worker-side generation error aborting a request, and a client-side inter-output timeout (_ASYNC_OUTPUT_TIMEOUT) firing while the worker was still running. Both land on the same race.

Root cause

Check-then-act race on concurrent.futures.Future state across threads: the pump's done() check and its resolve call aren't atomic, so a concurrent cancel() from the consumer side can land in between.

Fix

Added _try_set_result/_try_set_exception helpers that wrap set_result/set_exception in a narrow except concurrent.futures.InvalidStateError, dropping the late resolution instead of crashing the thread. Applied at every call site in _result_pump that resolves a future pulled from the shared _rpc_futures/_output_futures dicts, plus the same pattern in shutdown()'s cleanup loop for consistency. Left untouched the two set_result calls on freshly-created local Future() objects that aren't shared with any other thread yet — no race possible there.

Testing

Added tests/diffusion/test_multiproc_result_pump.py (CPU-only, no GPU needed): a _RacyFuture subclass whose done() deterministically lies (False) while genuinely cancelled underneath, reproducing the exact race window without relying on real thread timing.

Verified on an RTX 3090 (vLLM 0.26.0, vllm-omni installed per docs):

  • With the fix: tests/diffusion/test_multiproc_result_pump.py — 3 passed.
  • With the fix reverted: all 3 fail with the exact InvalidStateError traceback from the original report, at the same fut.set_result(msg.output) call site.
  • Neighboring suite tests/diffusion/test_multiproc_engine_concurrency.py — 60 passed, no regressions.
  • ruff check, ruff format --check, and the full pre-commit hook suite pass on both changed files.

Fixes #5793

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@anurag12-webster
anurag12-webster force-pushed the bugfix-5793-result-pump-cancelled-future-race branch from 0c8c2ed to 028d28e Compare August 10, 2026 08:54
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/continuous_batching.md.

Module owners: @Isotr0py @princepride

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@anurag12-webster

Copy link
Copy Markdown
Contributor Author

@vllm-omni-review-bot

@hsliuustc0106 hsliuustc0106 added bug Something isn't working diffusion codes related to diffusion models labels Aug 11, 2026
@anurag12-webster

Copy link
Copy Markdown
Contributor Author

@Isotr0py @princepride PTAL

@anurag12-webster

Copy link
Copy Markdown
Contributor Author

hi @Isotr0py @princepride can you take a look at this

…-project#5793)

A late result for an aborted/timed-out request could race the
result-pump thread's fut.done() check and set_result()/set_exception()
call, raising InvalidStateError with nothing to catch it. Since the
pump is a singleton, this permanently kills it: the engine stays
"healthy" but never completes any job again.

Wrap each set_result/set_exception call on a shared future in a
narrow try/except InvalidStateError that drops the late resolution
instead of crashing the thread.

Rebased onto main: vllm-project#6023 rewrote _result_pump/_deliver_batch_split
and reverted to plain set_result/set_exception calls, so this race
was reproducible again. Reapplied the same guard on the new call
sites and added the regression test to tests/diffusion/test_result_pump.py
(the test file vllm-project#6023 introduced) instead of a separate file.

Signed-off-by: ANURAG KANADE <anuragkanade6@gmail.com>
@anurag12-webster
anurag12-webster force-pushed the bugfix-5793-result-pump-cancelled-future-race branch from d69bba4 to ae7b3bc Compare August 15, 2026 13:17
@anurag12-webster

Copy link
Copy Markdown
Contributor Author

rebased onto main. #6023 rewrote _result_pump/_deliver_batch_split and merged before this pr, so the InvalidStateError fix here got dropped along with the old code shape it patched. the original #5793 crash was reproducible on main again.

reapplied the same fix on top of #6023's new code, and added the regression test to tests/diffusion/test_result_pump.py (the test file #6023 introduced) instead of a separate file.

verified on a real gpu: with the fix reverted, all 3 new tests fail with the exact InvalidStateError from the original report, at the 3 call sites in the new code. with the fix applied, all 21 tests in that file pass (3 new + 18 existing, no regressions).

Comment thread vllm_omni/diffusion/executor/multiproc_executor.py Outdated
…ect#5793)

Move try_set_result/try_set_exception into vllm_omni/diffusion/utils/future_utils.py per review.

Signed-off-by: ANURAG KANADE <anuragkanade6@gmail.com>

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123 Gaohan123 added the ready label to trigger buildkite CI label Aug 15, 2026
@Gaohan123
Gaohan123 enabled auto-merge (squash) August 15, 2026 15:54
@Gaohan123
Gaohan123 merged commit 68bf8c2 into vllm-project:main Aug 15, 2026
7 of 9 checks passed
@anurag12-webster

Copy link
Copy Markdown
Contributor Author

thankyouu!! @Gaohan123

ivanusto added a commit to ivanusto/vllm-omni that referenced this pull request Aug 17, 2026
A pump thread is the sole reader of one worker's result queue. vllm-project#5983 fixed
the InvalidStateError that killed it, but three gaps around that fix remain:

1. Only the dequeue half of the loop was guarded. Anything raised while
   dispatching still escapes the thread, and the cost is not one message but
   the whole queue -- every later result goes undelivered. Extract the
   dispatch body into _dispatch_result() and wrap the call, so one bad
   message costs one message.

2. check_health() never looked at the pump. When a pump thread dies every
   worker process is still alive and every request still runs on the GPU, so
   the engine reported healthy forever while no request could return; only a
   restart recovered it. Report a dead pump as EngineDeadError.

3. A result whose waiter had already cancelled was cached in
   _completed_outputs. That cache exists for results that arrive before the
   caller asks for them; a cancelled caller never asks, so the entry is never
   popped and the unpacked output -- a whole decoded video, for a t2va
   request -- stays pinned until the process exits. Drop it instead.

The broad 'except Exception' added in (1) is deliberate: this is a top-level
thread boundary, and it logs the full traceback rather than swallowing.

Most of the diff is the dedent from extracting _dispatch_result(); the logic
change is the three points above.

Reported in vllm-project#5793 and vllm-project#5821.

Signed-off-by: ivanusto <ivanusto@gmail.com>
ivanusto added a commit to ivanusto/vllm-omni that referenced this pull request Aug 22, 2026
A pump thread is the sole reader of one worker's result queue. vllm-project#5983 fixed
the InvalidStateError that killed it, but three gaps around that fix remain:

1. Only the dequeue half of the loop was guarded. Anything raised while
   dispatching still escapes the thread, and the cost is not one message but
   the whole queue -- every later result goes undelivered. Extract the
   dispatch body into _dispatch_result() and wrap the call, so one bad
   message costs one message.

2. check_health() never looked at the pump. When a pump thread dies every
   worker process is still alive and every request still runs on the GPU, so
   the engine reported healthy forever while no request could return; only a
   restart recovered it. Report a dead pump as EngineDeadError.

3. A result whose waiter had already cancelled was cached in
   _completed_outputs. That cache exists for results that arrive before the
   caller asks for them; a cancelled caller never asks, so the entry is never
   popped and the unpacked output -- a whole decoded video, for a t2va
   request -- stays pinned until the process exits. Drop it instead.

The broad 'except Exception' added in (1) is deliberate: this is a top-level
thread boundary, and it logs the full traceback rather than swallowing.

Most of the diff is the dedent from extracting _dispatch_result(); the logic
change is the three points above.

Reported in vllm-project#5793 and vllm-project#5821.

Signed-off-by: ivanusto <ivanusto@gmail.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working diffusion codes related to diffusion models ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: DiffusionResultPump thread dies with InvalidStateError when a request is aborted — engine becomes a zombie (accepts jobs, never runs them)

4 participants