Skip to content

[Bugfix] Keep the diffusion result pump alive and make its death visible - #6253

Open
ivanusto wants to merge 1 commit into
vllm-project:mainfrom
ivanusto:fix/result-pump-survives-bad-messages
Open

ivanusto wants to merge 1 commit into
vllm-project:mainfrom
ivanusto:fix/result-pump-survives-bad-messages

Conversation

@ivanusto

@ivanusto ivanusto commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

A DiffusionResultPump thread is the sole reader of one worker's result queue. #5983 fixed the InvalidStateError that was killing it, and that fix is correct — but it left three gaps around itself, all of which were raised in #5793 / #5821 and none of which are on main today.

1. Only the dequeue half of the pump loop is guarded. _result_pump wraps dequeue in try/except; everything after it runs unguarded, and _start_result_pump passes target=self._result_pump with no wrapper. So anything raised while dispatching still escapes and kills the thread — and because the thread is the only reader of that queue, the cost is not one message but every message after it. _deliver_batch_split's batch_output.get_request_output() and the _sync_result_buffer.put() path are both live examples. This PR extracts the dispatch body into _dispatch_result() and wraps the call, so one bad message costs one message.

2. check_health() never looked at the pump. This is the part that makes the failure so unpleasant operationally: when a pump thread dies, every worker process is still alive and every request still runs to completion on the GPU — the results simply never come back. /health and /v1/models answer 200 throughout, so an operator's only signal is noticing that requests stopped completing. Now a dead pump is reported as EngineDeadError, the same as a dead worker process. Guarded on _pump_running and not _pump_stop.is_set() so a deliberate shutdown isn't reported as a failure.

3. A result whose waiter had already cancelled was cached forever. _completed_outputs exists for the race where a result arrives before the caller asks for it. A caller that cancelled — an inter-output timeout, or an abort — never asks, so the entry is never popped: _completed_outputs is only drained by wait_output_ready() and the early-arrival adopt path, and there is no TTL and no abort-side cleanup. The pinned value is the unpacked output, which for a t2va request is a whole decoded video. Dropping it is also what the elif makes explicit; the pending is None early-arrival path is untouched.

The broad except Exception added in (1) is deliberate rather than an oversight: it sits at a top-level thread boundary, and it logs the full traceback via logger.exception instead of swallowing. Per the code-quality guidance, that is the sanctioned form at such a boundary — the alternative is the status quo, where the thread dies.

Most of the diff is the dedent from extracting _dispatch_result(). The logic change is the three points above; reviewing with whitespace ignored makes that much clearer.

Reported in #5793 and #5821. #6255 covers the other still-open item from those reports — the hardcoded 30 s _ASYNC_OUTPUT_TIMEOUT that triggers the cancellation in the first place. The two are independent and can land in either order.

Test Plan

Added three test classes to tests/diffusion/test_result_pump.py (CPU-only, no GPU), and generalized the existing _feed_one_msg_to_pump helper into _feed_msgs_to_pump so a test can feed a bad message followed by a good one and assert the good one still lands.

  • TestResultPumpDispatchIsGuarded — a message that raises during dispatch (sync-buffer path, and a corrupt batch output) must not stop the next message from being delivered by the same pump thread.
  • TestResultPumpDropsOutputForAbandonedWaiter — a cancelled waiter leaves _completed_outputs empty, in both the single-request and batch-split paths, while a result with no waiter at all is still cached as before.
  • TestCheckHealthDetectsDeadPump — a dead pump thread raises EngineDeadError and sets _is_failed; a live one is healthy; a stopped one (shutdown) is not reported as dead.

vLLM Version: 0.26.1rc1.dev608+g99a10304d

vLLM-Omni Commit: baba7d1

Test Result

tests/diffusion/test_result_pump.py                   29 passed
tests/diffusion/test_multiproc_engine_concurrency.py  \
tests/diffusion/test_diffusion_engine_cleanup.py       > 81 passed
tests/diffusion/test_diffusion_engine_rpc_routing.py  /

With the multiproc_executor.py change reverted and the tests kept, the five tests that assert the new behaviour fail and the three that assert preserved behaviour still pass:

FAILED TestResultPumpDispatchIsGuarded::test_dispatch_failure_does_not_kill_the_pump
FAILED TestResultPumpDispatchIsGuarded::test_batch_split_failure_does_not_kill_the_pump
FAILED TestResultPumpDropsOutputForAbandonedWaiter::test_cancelled_waiter_output_is_not_cached
FAILED TestResultPumpDropsOutputForAbandonedWaiter::test_cancelled_waiter_batch_split_output_is_not_cached
FAILED TestCheckHealthDetectsDeadPump::test_dead_pump_thread_marks_engine_dead
5 failed, 24 passed

ruff check and ruff format --check pass on both changed files.

Copilot AI lite review requested due to automatic review settings August 17, 2026 02:44
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Hardens MultiprocDiffusionExecutor’s async result-pump so a single bad result message can’t permanently wedge a worker queue, and so pump-thread death becomes externally visible via check_health()—addressing the “zombie engine” failure mode described in #5793/#5821.

Changes:

  • Wrap per-message dispatch in the pump loop, extracting _dispatch_result() so exceptions during dispatch drop only the offending message instead of killing the pump thread.
  • Teach check_health() to treat dead result-pump threads as EngineDeadError (while not flagging intentional shutdown).
  • Stop caching outputs for already-cancelled/abandoned waiters (single-output and batch-split), avoiding long-lived memory pinning in _completed_outputs.
  • Add targeted unit tests for guarded dispatch, abandoned-waiter output dropping, and dead-pump detection.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.

File Description
vllm_omni/diffusion/executor/multiproc_executor.py Makes result dispatch exception-safe, drops outputs for cancelled waiters, and surfaces pump-thread death via health checks.
tests/diffusion/test_result_pump.py Adds/extends CPU-only tests to verify guarded dispatch, abandoned-waiter handling, and health-check behavior.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +888 to +892
elif msg.kind == AsyncOutputKind.OUTPUT_READY:
batch_id = msg.async_output_id
with self._futures_lock:
per_req_map = self._batch_split_map.pop(batch_id, None) if batch_id else None
if per_req_map is not None:
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/continuous_batching.md, docs/design/module/diffusion/index.md, docs/design/module/diffusion/diffusion_runtime.md.

Module owners: @Isotr0py @princepride @SamitHuang @wtomin @ZJY0516 @RuixiangMa @david6666666 @xuechendi @fhfuih

@ivanusto, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@ivanusto
ivanusto force-pushed the fix/result-pump-survives-bad-messages branch from 98cc8ef to 7f85b71 Compare August 17, 2026 02:52
ivanusto added a commit to ivanusto/vllm-omni that referenced this pull request Aug 17, 2026
async_output_id is what routes a result back to its waiter, so a message
without one strands its request until that request's own timeout. The branch
dropped it silently; say so at error level instead, and return early so the
lookup below no longer needs the 'if batch_id' guard.

Raised in review of vllm-project#6253.

Signed-off-by: ivanusto <ivanusto@gmail.com>
@ivanusto

Copy link
Copy Markdown
Contributor Author

Self-review.

What I checked

  • Walked the extracted _dispatch_result() against the original loop body statement by statement: every continue became return, no branch changed shape, and the only semantic edits are the three described in the PR body. Reviewing this diff with whitespace ignored is much easier than reading it raw — the dedent dominates.
  • Confirmed the failure mode is real on main rather than theoretical: _start_result_pump passes target=self._result_pump with no wrapper, and the try covers only dequeue, so anything raised below it unwinds the thread.
  • Checked that dropping the cache entry for a cancelled waiter cannot lose a live result. _completed_outputs is only read by wait_output_ready() and the early-arrival adopt path in execute_batch; both key off an id whose waiter has not been created yet (pending is None), which is the branch I left untouched.
  • Checked the check_health() guard against shutdown: shutdown() sets _pump_stop before joining and clears _pump_running after, so _pump_running and not _pump_stop.is_set() is false throughout teardown. test_stopped_pump_is_not_reported_dead pins that.
  • Verified the new tests actually bite: with multiproc_executor.py reverted and the tests kept, the five that assert new behaviour fail and the three that assert preserved behaviour still pass (output in the PR body).
  • Ran the neighbouring suites (test_multiproc_engine_concurrency.py, test_diffusion_engine_cleanup.py, test_diffusion_engine_rpc_routing.py) for regressions, plus ruff check / ruff format --check.

Where I would welcome a second opinion

  • The broad except Exception around dispatch. It is intentional — top-level thread boundary, logs the full traceback via logger.exception — but it is still a new broad catch, and I would rather a maintainer confirm that is the right call here than have it slip through.
  • Whether a dead pump should set _is_failed (making the engine permanently dead, which is what it is) or attempt a restart of the thread. I chose the former as the smaller, more predictable change; restarting a pump mid-flight would leave the in-flight futures in an ambiguous state.
  • Latest commit addresses the review note about OUTPUT_READY messages with no async_output_id: they are now logged at error level and returned early, which also removes the if batch_id guard below.

Not covered

No GPU test. Everything here is CPU-only unit-level; I have not re-run an end-to-end video generation against this branch. The original failure was reproduced on a single GB10 with MiniMax-H3 FL2VA (#5821), and the equivalent patch has been running in that deployment, but that patch predates this refactor.

@ivanusto

Copy link
Copy Markdown
Contributor Author

Status update, since this one has been sitting a while.

All checks are green on c756d03 — pre-commit, build (3.11), build (3.12) and DCO all pass — and mergeable is true, so the only thing outstanding is a review.

The one review comment on the diff has been addressed. The bot flagged that the OUTPUT_READY branch dropped messages with a missing async_output_id silently; c756d03 handles that explicitly and says why rather than just returning early:

elif msg.kind == AsyncOutputKind.OUTPUT_READY:
    batch_id = msg.async_output_id
    if not batch_id:
        # async_output_id is what routes the result back to its waiter.
        # Without it there is nobody to resolve and nothing to cache, so
        # the request hangs until its own timeout — say so rather than
        # dropping the message silently.
        logger.error("Dropping OUTPUT_READY with no async_output_id; its request cannot be resolved")
        return

That is the same principle as the PR itself: the failure was never that the pump did the wrong thing, it was that it died or discarded work without leaving a trace, so the symptom showed up much later as a request that simply never completed.

@Isotr0py @princepride @SamitHuang @wtomin @ZJY0516 @RuixiangMa @david6666666 @xuechendi @fhfuih — when one of you has a moment. Happy to rebase if main has moved under it.

pending = self._output_futures.pop(per_req_id, None)
if pending is not None and not pending.done():
try_set_result(pending, per_req_result)
elif pending is not None:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] One corrupt per-request result drops its whole batch's siblings — vllm_omni/diffusion/executor/multiproc_executor.py:954

get_request_output() raising for one request escapes _deliver_batch_split mid-loop; the split map was already popped, so every later request in the same batch loses its output and hangs until _ASYNC_OUTPUT_TIMEOUT even though only one request was corrupt. Message-level containment now holds, but inside a batch split one bad request still costs the whole batch. Smallest fix: a per-request try/except in that loop resolving the failing request with DiffusionOutput(error=...) so siblings still complete.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, and fixed in f81871b1f.

The extraction half of the loop body is now wrapped, and a request whose result cannot be read is resolved with a DiffusionOutput(error=...) of its own, so its siblings still complete:

for per_req_id, req_id in per_req_map.items():
    per_req_result: DiffusionOutput
    try:
        req_output = batch_output.get_request_output(req_id) if batch_output is not None else None
        ...
    except Exception as e:
        # The split map has already been popped, so an exception escaping
        # here would take every request later in this batch with it: they
        # would never be resolved and would hang until their own timeout.
        # One corrupt request costs one request.
        logger.exception("Failed to extract batch output for request %s", req_id)
        per_req_result = DiffusionOutput(error=f"Failed to extract batch output for request {req_id}: {e}")
    with self._futures_lock:
        ...  # resolve / drop / cache block unchanged

One deliberate choice worth flagging: the with self._futures_lock block below is left outside the try. try_set_result() already swallows InvalidStateError (#5983), and dict writes plus set_result() on a freshly constructed Future cannot raise, so widening the guard would only add a second place that has to decide what to do with a half-resolved request. The per-message guard in the pump loop stays the backstop for anything unforeseen. Happy to widen it if you would rather have belt and braces.

Tests: test_batch_split_failure_does_not_kill_the_pump previously asserted not doomed.done() with the comment "that one message is lost, as expected"; that assertion is now the opposite, since the corrupt request is failed rather than left hanging. Added test_one_corrupt_request_does_not_drop_its_batch_siblings, a three-request batch whose middle entry raises, asserting all three futures resolve and that r-2, delivered after the corrupt r-1, still gets its real result. Both fail on the previous commit and pass on this one.

# dies, every process is still alive and every request still runs on the
# GPU — the results simply never come back. Without this check the
# engine reports healthy forever and only a restart recovers it.
if self._pump_running and not self._pump_stop.is_set():

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P3] Sync async_diffusion_output.md's reliability section with the new pump semantics — vllm_omni/diffusion/executor/multiproc_executor.py:1020

Dead-pump → EngineDeadError (health/readiness flip) and the cancelled-waiter drop in _dispatch_result extend the "Reliability, Lifecycle & Timeout Behavior" contract owned by docs/design/feature/async_diffusion_output.md, but the page still describes the pump only as "resolves waiting futures or populates _completed_outputs". A short same-PR paragraph there (dispatch containment, pump-death detection, drop-on-abandoned-waiter) keeps the active design page authoritative.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed, and done in fa9b58bad.

Added item 5 to "Reliability, Lifecycle & Timeout Behavior" rather than editing items 1 to 4, so the diff stays small and #6255's rewrite of item 2 is untouched:

  1. Pump Fault Containment & Liveness:
    A pump thread is the sole reader of its worker's result queue, so anything escaping it costs every later result rather than one message. Dispatch of each dequeued message is therefore wrapped: a bad message is logged with its traceback and dropped, and the pump keeps draining. Within a batch split the containment is per request, because the split map is popped before delivery starts. Outputs whose waiter has already been cancelled or aborted are dropped instead of cached, since nobody will ever collect them and _completed_outputs has no TTL or abort cleanup (a decoded video would stay pinned until shutdown). Finally, check_health() treats a dead pump thread as EngineDeadError: the processes are all alive and the GPU work still runs, so without that check the engine reports healthy forever while no result ever comes back. Intentional shutdown is excluded, since shutdown() sets the stop event before joining.

Also extended the "Batch Split" paragraph with one sentence on the per-request containment from the other comment, since that is where a reader would look for it.

ivanusto added a commit to ivanusto/vllm-omni that referenced this pull request Aug 22, 2026
_deliver_batch_split pops the split map before it starts delivering, so an
exception raised while extracting one request's result escaped mid-loop and
took every request after it in the same batch with it: no future resolved, no
cache entry, nothing left to route a later result through. Those requests then
hung until their own timeout even though only one entry was corrupt.

Wrap the extraction so a request that cannot be read from the batch output is
resolved with a DiffusionOutput(error=...) of its own and its siblings still
complete. The resolve/cache block below is deliberately left outside the try:
try_set_result already swallows InvalidStateError, and dict writes and
set_result on a fresh Future cannot raise, so the per-message guard in the pump
loop stays the backstop for anything unforeseen.

Raised in review of vllm-project#6253.

Signed-off-by: ivanusto <ivanusto@gmail.com>
ivanusto added a commit to ivanusto/vllm-omni that referenced this pull request Aug 22, 2026
The reliability section still described the pump only as resolving waiting
futures or populating _completed_outputs, which no longer covers what it does.
Add the dispatch containment, the per-request containment inside a batch split,
the drop of outputs whose waiter has been abandoned, and the dead-pump health
check, so the design page stays authoritative for this contract.

Raised in review of vllm-project#6253.

Signed-off-by: ivanusto <ivanusto@gmail.com>
@ivanusto
ivanusto force-pushed the fix/result-pump-survives-bad-messages branch from dda442d to fa9b58b Compare August 22, 2026 16:05
ivanusto added a commit to ivanusto/vllm-omni that referenced this pull request Aug 22, 2026
async_output_id is what routes a result back to its waiter, so a message
without one strands its request until that request's own timeout. The branch
dropped it silently; say so at error level instead, and return early so the
lookup below no longer needs the 'if batch_id' guard.

Raised in review of vllm-project#6253.

Signed-off-by: ivanusto <ivanusto@gmail.com>
ivanusto added a commit to ivanusto/vllm-omni that referenced this pull request Aug 22, 2026
_deliver_batch_split pops the split map before it starts delivering, so an
exception raised while extracting one request's result escaped mid-loop and
took every request after it in the same batch with it: no future resolved, no
cache entry, nothing left to route a later result through. Those requests then
hung until their own timeout even though only one entry was corrupt.

Wrap the extraction so a request that cannot be read from the batch output is
resolved with a DiffusionOutput(error=...) of its own and its siblings still
complete. The resolve/cache block below is deliberately left outside the try:
try_set_result already swallows InvalidStateError, and dict writes and
set_result on a fresh Future cannot raise, so the per-message guard in the pump
loop stays the backstop for anything unforeseen.

Raised in review of vllm-project#6253.

Signed-off-by: ivanusto <ivanusto@gmail.com>
ivanusto added a commit to ivanusto/vllm-omni that referenced this pull request Aug 22, 2026
The reliability section still described the pump only as resolving waiting
futures or populating _completed_outputs, which no longer covers what it does.
Add the dispatch containment, the per-request containment inside a batch split,
the drop of outputs whose waiter has been abandoned, and the dead-pump health
check, so the design page stays authoritative for this contract.

Raised in review of vllm-project#6253.

Signed-off-by: ivanusto <ivanusto@gmail.com>
@ivanusto

Copy link
Copy Markdown
Contributor Author

@hsliuustc0106 thanks for the review. Both P3s are addressed, and the branch is now rebased onto current main.

P3-1, one corrupt per-request result drops its whole batch's siblings (f81871b1f)

Result extraction inside _deliver_batch_split is now wrapped per request. A request whose result cannot be read from the batch output is resolved with a DiffusionOutput(error=...) of its own, so the requests after it in the same split still complete instead of hanging until their own timeout. Details and one deliberate choice about the scope of the try are in the thread.

P3-2, sync the design page (fa9b58bad)

Added item 5, "Pump Fault Containment & Liveness", to "Reliability, Lifecycle & Timeout Behavior" in docs/design/feature/async_diffusion_output.md, covering dispatch containment, per-request containment inside a batch split, the drop for abandoned waiters, and dead-pump detection. Items 1 to 4 are untouched, so your _async_output_timeout() rewrite from #6255 stays as merged. The "Batch Split" paragraph also gains one sentence on the per-request containment. Rendered preview: https://vllm--6253.org.readthedocs.build/projects/vllm-omni/en/6253/design/feature/async_diffusion_output/

Verification

tests/diffusion/test_result_pump.py is 31 passed on the rebased tree. Reverting only multiproc_executor.py and keeping the tests turns test_batch_split_failure_does_not_kill_the_pump and the new test_one_corrupt_request_does_not_drop_its_batch_siblings red, so they do bite. Ran the neighbouring suites too (test_multiproc_engine_concurrency.py, test_diffusion_engine_cleanup.py, test_diffusion_engine_rpc_routing.py): 119 passed together, no regressions. ruff check and ruff format --check clean. DCO and Read the Docs already green on the rebased head; the GitHub Actions builds are still queued on the runner backlog rather than failing.

One ask

This PR still has no ready label, so buildkite has never run on it, unlike #6255. Could you add it, or tell me who to ask? Everything else is green and mergeable is true.

fhfuih added a commit to fhfuih/vllm-omni that referenced this pull request Sep 3, 2026
Document diffusion_kv_metadata on NewRequestData, correct multiproc
shutdown so completed async outputs drop with the executor, and tag
abort/consumer-drop async-output reclaim as pending vllm-project#6253/vllm-project#6439/vllm-project#6580.

Signed-off-by: Huang, Zeyu <11222265+fhfuih@users.noreply.github.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 16 days

@ivanusto this pull request has had no human commit, comment or review since 2026-08-22. Per repository policy it may be closed if it stays inactive.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@ivanusto

ivanusto commented Sep 8, 2026

Copy link
Copy Markdown
Contributor Author

Status update, since the bot flagged inactivity.

This one is not blocked on me. The state as of fa9b58b:

  • Both P3s from the review are addressed, and I wrote up what changed in each when I pushed them.
  • The branch is rebased on current main and shows MERGEABLE.
  • All checks are green: DCO, build (3.11), build (3.12), pre-commit, and the Read the Docs build.

There are no open review threads left on my side, so what it needs is a maintainer look rather than another push from me. @hsliuustc0106 if you have a moment, this is ready for a re-review, or for the ready label if that is what is gating it.

Happy to rebase again if it goes stale in the meantime.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 15 days

@ivanusto this pull request has had no human commit, comment or review since 2026-09-08. Please consider marking this PR as draft until work can resume. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

… outputs

The pump thread is the sole reader of a worker result queue, so an exception
escaping the dispatch costs the whole queue rather than the one message that
caused it: every later result goes undelivered while the server keeps
reporting healthy. Route each message through _dispatch_result() inside a
try, so a bad message is logged and dropped on its own.

Two deliveries could also go missing without a trace. An OUTPUT_READY with no
async_output_id has nobody to resolve and nothing to cache, and was discarded
silently; it is now logged. And inside a batch split the map has already been
popped, so a failure while extracting one request's result took every request
after it in the same batch with it; extraction is now per request, and the one
that fails resolves with its own error output.

Signed-off-by: ivanusto <ivanusto@gmail.com>
@ivanusto
ivanusto force-pushed the fix/result-pump-survives-bad-messages branch from fa9b58b to 6a4b91d Compare September 24, 2026 16:24
@ivanusto

Copy link
Copy Markdown
Contributor Author

Rebased onto current main and rewritten against what the pump looks like today, since the branch had gone stale behind #5929, #7126 and #6439. Force pushed as a single commit, 6a4b91d.

What main has already absorbed, and I dropped: the dequeue-level try/except, the multiple pump threads, the _closed re-checks, the dropped-id LRU and _finish_output. The old branch predated all of that, so I rebuilt on top of it rather than reapplying the old shape.

What is still missing on main, and is what this PR now is:

  1. Dispatch is not contained. The dequeue is wrapped, everything after it is not. This thread is the sole reader of its queue, so an exception from _deliver_batch_split or from SHM handling costs every later result on that queue while the server keeps reporting healthy. Routing moved into _dispatch_result(), called inside a try.
  2. OUTPUT_READY with no async_output_id disappears silently. _batch_split_map.pop(batch_id, ...) if batch_id else None yields None, then the single-request branch's if batch_id: is false, so nothing resolves and nothing is logged. The request hangs until its own timeout with no trace. Now logged and dropped explicitly.
  3. Batch splitting is not contained per request. The split map has already been popped, so one failing get_request_output takes every request after it in that batch. Extraction is now per request; the failing one resolves with its own DiffusionOutput(error=...) and its siblings still complete. The _closed re-check and _finish_output are untouched.

Behaviour on main is otherwise unchanged: same ordering, same locking, same shutdown semantics.

Tests. Three new cases in tests/diffusion/test_result_pump.py, added alongside the existing ones rather than replacing them, including the aborted-output cases added in #6439. They fail on main and pass here:

main this PR
test_dispatch_failure_does_not_stop_the_pump fail (later message never delivered) pass
test_output_ready_without_id_is_logged_not_silent fail (assert 0 == 1 logged errors) pass
test_one_corrupt_request_does_not_drop_its_siblings fail (siblings never resolve) pass

Whole file: 39 passed. ruff check and ruff format clean at the pinned v0.14.10.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 13 days

@ivanusto this pull request has had no human commit, comment or review since 2026-09-24. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on zcode (GLM-5.3-Flash) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 2b1ba3c3-d97a-4a4c-88f5-92a22f32f245) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 120s (try ).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working diffusion codes related to diffusion models

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants