Skip to content

[Bugfix] Make the async-output wait bound configurable and default it higher - #6255

Merged
hsliuustc0106 merged 4 commits into
vllm-project:mainfrom
ivanusto:fix/configurable-async-output-timeout
Aug 22, 2026
Merged

hsliuustc0106 merged 4 commits into
vllm-project:mainfrom
ivanusto:fix/configurable-async-output-timeout

Conversation

@ivanusto

@ivanusto ivanusto commented Aug 17, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

_ASYNC_OUTPUT_TIMEOUT is hardcoded at 30 s and used at two sites in DiffusionEngine. It bounds the wait for one step's background D2H/SHM copy — a copy that finishes in milliseconds, but that is queued behind the GPU work for that step, so the wall-clock wait tracks step time rather than copy time.

On a single GPU, large shapes legitimately run tens of seconds per step, which trips a 30 s bound and aborts the request even though the denoise completed. #5793 reported a 49-step render reaching 49/49 after 31 minutes and being discarded anyway; #5821 reported the same on 44–48 s/step. There is currently no way to raise the bound short of patching the installed package.

A tight bound also buys very little: a genuinely hung engine is surfaced elsewhere — the worker monitor and check_health() catch worker death, shutdown() settles pending futures, and the pump's dequeue loop breaks on _is_failed. What this wait catches, in practice, is a healthy render that is merely slow.

This PR:

  • reads VLLM_OMNI_ASYNC_OUTPUT_TIMEOUT, defaulting to 600 s to match _DLO_DP_WAVE_TIMEOUT_S in the same subsystem (which uses the same float(os.environ.get(...)) shape);
  • names the environment variable in the timeout log line, so an operator who hits the bound learns that the knob exists without reading the source;
  • resolves it per call rather than at import, so the value is never frozen in a module constant, a malformed value cannot break module import, and the knob is testable with monkeypatch rather than a module reload — an earlier reload-based draft of this test measurably polluted tests/diffusion/test_diffusion_engine.py when the two ran in the same session;
  • parses defensively: because the resolution now happens on the request path, a non-numeric or non-positive value is ignored with a warning_once and the default applies, rather than raising ValueError mid-generation;
  • updates docs/design/feature/async_diffusion_output.md, which documented the old 30.0s value.

Requested in #5793 (suggested fix 2) and #5821 (suggested fix 3). Note this is only the trigger, not the crash: with #5983 merged the pump now survives the cancellation this timeout causes. #6253 covers the remaining pump-robustness gaps. The three are independent.

Test Plan

New tests/diffusion/test_async_output_timeout.py (CPU-only, no GPU): default value, integer and float overrides, that the value is re-read per call rather than frozen at import, and that malformed values (non-numeric, empty, zero, negative) fall back to the default.

vLLM Version: 0.26.1rc1.dev608+g99a10304d

vLLM-Omni Commit: baba7d1

Test Result

tests/diffusion/test_async_output_timeout.py    11 passed
tests/diffusion/test_diffusion_engine.py         \
tests/diffusion/test_result_pump.py               \  133 passed, 1 failed
tests/diffusion/test_diffusion_engine_cleanup.py  /
tests/diffusion/test_multiproc_engine_concurrency.py

The single failure is test_move_tensor_tree_moves_nested_cuda_tensors_to_cpu, which needs a CUDA device (RuntimeError: Found no NVIDIA driver on your system) and fails identically on unmodified main in the same container.

ruff check and ruff format --check pass on both changed Python files.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

… higher

_ASYNC_OUTPUT_TIMEOUT was hardcoded at 30 s. It bounds the wait for one
step's background D2H/SHM copy -- a copy that takes milliseconds, but that is
queued behind the GPU work for that step, so the wall-clock wait tracks step
time. A single-GPU box legitimately runs tens of seconds per step on large
shapes, which trips the bound and aborts the request even though the denoise
completed; there was no way to raise it short of patching the installed
package.

A tight bound buys nothing here: worker death and a dead result pump are
surfaced by the worker monitor and check_health(), not by this wait. So the
only thing it catches is a healthy-but-slow render.

Read VLLM_OMNI_ASYNC_OUTPUT_TIMEOUT, defaulting to 600 s to match
_DLO_DP_WAVE_TIMEOUT_S in the same subsystem, and name the variable in the
timeout log so an operator hitting it knows the knob exists. Read per call
rather than at import so widening it does not require a restart.

Requested in vllm-project#5793 and vllm-project#5821.

Signed-off-by: ivanusto <ivanusto@gmail.com>
@ivanusto
ivanusto force-pushed the fix/configurable-async-output-timeout branch from 975e834 to 2bf8559 Compare August 17, 2026 02:52

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR makes the diffusion engine’s async-output wait timeout configurable via an environment variable and raises the default to better support long single-GPU denoise step times, preventing healthy-but-slow renders from being aborted by a hardcoded 30s bound.

Changes:

  • Replace the hardcoded async-output timeout with _async_output_timeout() reading VLLM_OMNI_ASYNC_OUTPUT_TIMEOUT (default 600s) and re-read it per call.
  • Improve timeout logging to mention the environment variable knob.
  • Add CPU-only unit tests for default/override/per-call behavior and update the async diffusion output design doc.

Reviewed changes

Copilot reviewed 3 out of 3 changed files in this pull request and generated 3 comments.

File Description
vllm_omni/diffusion/diffusion_engine.py Introduces env-var-based timeout resolution and applies it at async and sync wait sites.
tests/diffusion/test_async_output_timeout.py Adds unit tests verifying default value, overrides, and per-call re-reading.
docs/design/feature/async_diffusion_output.md Updates design documentation to reflect the new timeout mechanism and default.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread vllm_omni/diffusion/diffusion_engine.py Outdated
Comment on lines +63 to +75
def _async_output_timeout() -> float:
"""Seconds to wait for one step's background D2H/SHM copy.

The copy itself finishes in milliseconds, but it is queued behind the GPU
work for that step, so the wall-clock wait tracks step time — a single-GPU
box legitimately runs tens of seconds per step on large shapes. A tight
bound therefore does not catch a hung engine (worker death and a dead
result pump are surfaced by the worker monitor and ``check_health``); it
only aborts renders that are still making progress, throwing away the
denoise that already completed. The default matches
``_DLO_DP_WAVE_TIMEOUT_S`` in the same subsystem.
"""
return float(os.environ.get(_ASYNC_OUTPUT_TIMEOUT_ENV, _ASYNC_OUTPUT_TIMEOUT_DEFAULT))
Comment on lines 365 to 370
logger.error(
"Timed out after %.0fs waiting for async output; executor state: %s",
_ASYNC_OUTPUT_TIMEOUT,
"Timed out after %.0fs waiting for async output (raise %s to allow slower "
"steps); executor state: %s",
timeout,
_ASYNC_OUTPUT_TIMEOUT_ENV,
describe(output.async_output_id) if describe else "unavailable",
Comment on lines +32 to +35
def test_value_is_read_per_call(self, monkeypatch):
"""Read at call time rather than import time, so operators are not
forced to restart the server to widen the bound.
"""
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/index.md.

Module owners: @Isotr0py @princepride @SamitHuang @wtomin @ZJY0516 @RuixiangMa @david6666666 @xuechendi

@ivanusto, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

Three review points:

- float() on the env value ran on the request path, so an environment typo
  would start failing generations at runtime rather than at startup. Parse
  defensively: a non-numeric or non-positive value is ignored with a
  warning_once and the default applies.
- The timeout log said 'raise VLLM_OMNI_ASYNC_OUTPUT_TIMEOUT', which reads as
  raising an exception, and printed the bound with %.0f although the knob
  takes floats. Now 'set ... to a larger value', with %.1f.
- A test docstring claimed operators could widen the bound without a restart.
  Editing the environment of a running process does not generally work that
  way; the actual benefit is that the value is not frozen in a module
  constant, which is what the test now says.

Raised in review of vllm-project#6255.

Signed-off-by: ivanusto <ivanusto@gmail.com>
@ivanusto

Copy link
Copy Markdown
Contributor Author

Self-review.

What I checked

  • Both call sites of the old constant are converted (step_streaming and the sync path in add_req_and_wait_for_response); grep _ASYNC_OUTPUT_TIMEOUT leaves only the new _ENV/_DEFAULT names and the helper.
  • Confirmed the bound is not load-bearing for failure detection before raising the default, which is the part worth challenging. Worker death is caught by the worker monitor and check_health(); shutdown() settles pending futures with an exception; the pump's dequeue loop breaks on _is_failed. So a longer bound delays no real failure signal — it only stops aborting slow renders.
  • Checked the default against the neighbouring knob rather than picking a round number: _DLO_DP_WAVE_TIMEOUT_S in multiproc_executor.py is also 600 s from an env var, in the same subsystem.
  • Grepped for other references to the old 30 s value and found the design doc, which is updated in this PR.
  • ruff check / ruff format --check, plus the neighbouring diffusion suites for regressions.

Review notes addressed in the latest commit

  • Defensive parsing. The point about float() raising on the request path is a good one and it is worse than it looks: resolving per call turned what would have been a startup crash into a mid-generation failure. A non-numeric or non-positive value now falls back to the default with a warning_once, covered by parametrized tests.
  • Log wording: "raise VLLM_OMNI_ASYNC_OUTPUT_TIMEOUT" did read as raising an exception, and %.0f contradicted the float-accepting knob. Now "set ... to a larger value", with %.1f.
  • The test docstring claiming operators could widen the bound without a restart was simply wrong — editing a running process's environment does not generally work that way. Reworded to what the test actually demonstrates: the value is not frozen in a module constant.

Where I would welcome a second opinion

Not covered

No GPU test on this branch. The behaviour was originally observed on a single GB10 running MiniMax-H3 FL2VA at 44–48 s/step (#5821), where the 30 s bound aborted a completed denoise; the change here is confined to how that bound is resolved.

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

@SamitHuang PTAL

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Aug 18, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

One reference to the old constant survived the update: recipes/MiniMaxAI/MiniMax-H3-4090.md:114 still explains the dropped four-GPU Ref2VA repeats with "the engine's hardcoded 30 s async output wait (_ASYNC_OUTPUT_TIMEOUT in diffusion_engine.py)". After this PR that name no longer exists — a one-line note that the wait is now VLLM_OMNI_ASYNC_OUTPUT_TIMEOUT (default 600 s) would keep the recipe greppable for anyone hitting the same behavior.

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

On the two questions in the self-review:

  • Raising the default to 600 s: agree. The bound only aborts renders whose denoise already completed, and the failures it might be assumed to guard are surfaced earlier by the worker monitor (_start_worker_monitor, multiproc_executor.py:368) and check_health() (multiproc_executor.py:973), so a longer bound delays no real failure signal. Anchoring the default to _DLO_DP_WAVE_TIMEOUT_S (same subsystem, same 600 s) is the right call.
  • Helper shape: the per-call function is right here. The nearest precedent resolves float() at import, which turns a typo in the env var into a startup crash — acceptable there, but this knob resolves on the request path and needs the per-call read plus the defensive fallback. It's only the second consumer of the pattern in the tree, so I wouldn't generalize it into a shared helper yet.

One question on the test report: the body shows "133 passed, 1 failed" across the four neighboring suites without naming the failure. Which test failed, and does it also fail on the base commit? Given the self-review's note that an earlier draft of this file polluted test_diffusion_engine.py in-session, it's worth ruling out residual state leakage from the current version.

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

any updates?

The recipe explained the dropped four-GPU Ref2VA repeats in terms of
`_ASYNC_OUTPUT_TIMEOUT`, which this PR removes, leaving the only remaining
reference to that name in the tree dangling. Keep the historical fact -- the
wait really was a hardcoded 30 s when those numbers were taken -- and name the
env var that replaces it, so the recipe stays greppable for anyone hitting the
same behaviour.

Signed-off-by: ivanusto <ivanusto@gmail.com>
@ivanusto

Copy link
Copy Markdown
Contributor Author

Sorry for the delay, and thanks for the careful review.

Pushed 62fface, which fixes the stale reference you found. recipes/MiniMaxAI/MiniMax-H3-4090.md was the only remaining mention of _ASYNC_OUTPUT_TIMEOUT in the tree (grep -rn '_ASYNC_OUTPUT_TIMEOUT' --include='*.md' now returns nothing but the new env var). I kept the historical fact rather than rewriting it — the wait really was a hardcoded 30 s when those numbers were taken — and named the replacement so the recipe stays greppable:

That wait was a hardcoded 30 s when these numbers were taken; it is now VLLM_OMNI_ASYNC_OUTPUT_TIMEOUT (default 600 s), so a rerun on current main should not lose these repeats to the bound.

On your two points, both of which I agree with and neither of which needs a change now:

  • 600 s default. Your reasoning is the argument I wanted and could not make as precisely: the bound only aborts renders whose denoise already completed, and genuine failures surface earlier through _start_worker_monitor and check_health(), so a longer bound delays no real signal. Anchoring to _DLO_DP_WAVE_TIMEOUT_S keeps the two waits in the same subsystem consistent.
  • Per-call helper. Agreed, and it is what makes the test possible at all — a module-level constant resolved at import cannot be monkeypatched without a reload.

The two review-bot comments from 08-17 were addressed in the commits before the merge from main, so they may read as outstanding above:

  • The timeout log now reads "Timed out after %.1fs waiting for async output; set %s to a larger value to allow slower steps." — "set ... to a larger value" rather than "raise", and %.1f rather than %.0f, so a fractional value prints as given.
  • The test docstring no longer claims an operator can widen the bound on a running server. It now states the actual property: the value is resolved per call rather than captured in a module constant, so it never goes stale relative to os.environ and no module reload is needed.

CI was green on the previous head; happy to rebase or squash if you would prefer this as a single commit.

@hsliuustc0106 hsliuustc0106 added ready label to trigger buildkite CI cuda-test Used to trigger vllm-omni cuda CI separately. labels Aug 20, 2026
@ivanusto

Copy link
Copy Markdown
Contributor Author

One question on the test report: the body shows "133 passed, 1 failed" across the four neighboring suites without naming the failure. Which test failed, and does it also fail on the base commit?

Answer, with a fresh run rather than the older report.

The failure is tests/diffusion/test_diffusion_engine.py::test_move_tensor_tree_moves_nested_cuda_tensors_to_cpu, and it is a missing-GPU failure, not a regression. It is decorated @hardware_test(res={"cuda": "L4"}, num_cards=1) and its first statement is torch.arange(8, dtype=torch.float32, device="cuda") (test_diffusion_engine.py:707), so in a CPU-only container it raises:

E   RuntimeError: Found no NVIDIA driver on your system. Please check that you have an
    NVIDIA GPU and installed a driver from http://www.nvidia.com/Download/index.aspx
/usr/local/lib/python3.12/dist-packages/torch/cuda/__init__.py:529: RuntimeError

It fails identically on the base commit. Same container, same command, -p no:randomly, four suites:

commit result
62fface (this branch) 1 failed, 125 passed, 17 warnings in 6.51s
8dc3d504 (merge-base with main) 1 failed, 125 passed, 17 warnings in 6.57s

Same single failure on both sides, so the delta from this PR is zero. (The counts differ from the 133 in the description because that report was taken at baba7d1e, before the 08-19 merge from main; the tree has moved since. The failing test is the same one.)

On the state-leakage concern — that was the right thing to check, and it is clean:

  • test_async_output_timeout.py and test_diffusion_engine.py in one session: TestRequestBatchCapability::test_make_engine_runs_startup_warmup PASSED (1 failed, 44 passed — the failure is again the CUDA test above). That is the exact test the earlier importlib.reload-based draft used to break, which is why the current version resolves the value in a function and tests it with monkeypatch instead.
  • Per-file, in isolation, on this branch: test_async_output_timeout.py 11 passed · test_diffusion_engine.py 1 failed 33 passed · test_result_pump.py 21 passed · test_diffusion_engine_cleanup.py 13 passed · test_multiproc_engine_concurrency.py 58 passed. The four-suite numbers are exactly the sum of the isolated ones, so collection order changes nothing.

Two small things while you are here:

  • [Bugfix] Keep the diffusion result pump alive and make its death visible #6253 (the pump-robustness half of this pair) has no buildkite/vllm-omni check at all, because it never got the ready label that this PR has. DCO, both builds, pre-commit and readthedocs are green on it and mergeable is true, so it is only waiting on the label plus a reviewer. Could you add ready? It is 61 commits behind main now — happy to merge main into it first if you would rather see a fresh run.
  • The offer from my last comment still stands: say the word and I will squash this branch into a single commit, and/or bring it up to date with main.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

lgtm

@hsliuustc0106
hsliuustc0106 merged commit f09bdcb into vllm-project:main Aug 22, 2026
6 checks passed
leonail1 added a commit to leonail1/vllm-omni that referenced this pull request Aug 22, 2026
…project#6255

Upstream made the async-output wait bound configurable with the same
VLLM_OMNI_ASYNC_OUTPUT_TIMEOUT env var (default 600s, lazy parsing,
executor-state error message); our stage1-era constant is removed to
match upstream exactly.

Signed-off-by: zhenggang Lin <115566482+leonail1@users.noreply.github.com>
leonail1 added a commit to leonail1/vllm-omni that referenced this pull request Aug 23, 2026
…project#6255

Upstream made the async-output wait bound configurable with the same
VLLM_OMNI_ASYNC_OUTPUT_TIMEOUT env var (default 600s, lazy parsing,
executor-state error message); our stage1-era constant is removed to
match upstream exactly.

Signed-off-by: zhenggang Lin <115566482+leonail1@users.noreply.github.com>
AndyZhou952 pushed a commit to AndyZhou952/vllm-omni that referenced this pull request Aug 26, 2026
… higher (vllm-project#6255)

Signed-off-by: ivanusto <ivanusto@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Signed-off-by: AndyZhou952 <jzhoubc@connect.ust.hk>
imkingjh999 pushed a commit to imkingjh999/vllm-omni that referenced this pull request Aug 29, 2026
… higher (vllm-project#6255)

Signed-off-by: ivanusto <ivanusto@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
(cherry picked from commit f09bdcb)
JoseCarlosGarcia95 pushed a commit to valendra-tech/vllm-omni that referenced this pull request Sep 5, 2026
… higher (vllm-project#6255)

Signed-off-by: ivanusto <ivanusto@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
… higher (vllm-project#6255)

Signed-off-by: ivanusto <ivanusto@gmail.com>
Co-authored-by: Hongsheng Liu <liuhongsheng4@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working cuda-test Used to trigger vllm-omni cuda CI separately. diffusion codes related to diffusion models high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants