Skip to content

[Core][Frontend] Support request-level batching for diffusion pipelines - #4079

Merged
hsliuustc0106 merged 34 commits into
vllm-project:mainfrom
omni-nicelab:refactor/request_id
Jun 29, 2026
Merged

hsliuustc0106 merged 34 commits into
vllm-project:mainfrom
omni-nicelab:refactor/request_id

Conversation

@yJader

@yJader yJader commented Jun 2, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Implements Phase 2 of RFC #3550. This completes the diffusion request contract cleanup by making OmniDiffusionRequest represent exactly one logical request with one prompt, while runtime batching is represented by scheduler output and runner-side batch structures.

In this PR:

Request contract

  • OmniDiffusionRequest.prompts is replaced by a single prompt field; request identity remains the required single request_id. __post_init__ rejects list inputs defensively.
  • Scheduler output sends NewRequestData(request_id, req), carrying the already-initialized OmniDiffusionRequest directly so the executor/worker forward it instead of rebuilding it (which would re-trigger __post_init__ and corrupt sentinel fields like guidance_scale_provided). Fields like prompt / sampling_params / kv_sender_info are read off req rather than duplicated on the payload.

Scheduler batching

  • Scheduler batching is based on compatible independent requests via SamplingParamsKey, with conservative LoRA homogeneity constraints. num_outputs_per_prompt is added to the key so requests producing differently-shaped outputs do not get co-batched.

Request-level batch & execution path

  • DiffusionRequestBatch is introduced as the request-level batch abstraction above the step/tensor-level InputBatch (StepInputBatch alias). It wraps list[OmniDiffusionRequest] and exposes prompts / sampling_params / request_id / kv_sender_info compatibility properties so pipeline forward bodies stay close to upstream. Worker-side DiffusionRequestState (single prompt) is retained for the stepwise path only.
  • Pipeline forward is unified to forward(req: DiffusionRequestBatch) -> list[DiffusionOutput]. A class attribute supports_request_batch advertises whether the pipeline supports one batched forward.
  • The executor exposes two request-mode entries — execute_request (one execute_model RPC per scheduled request, the preserved upstream path) and execute_batch (a single execute_model_batch RPC carrying the whole DiffusionSchedulerOutput). The engine resolves request-batch capability at initialization from the configured pipeline class, including custom pipeline classes, and binds batch-capable pipelines directly to execute_batch for all request-mode cycles, including single-request cycles.
  • On the batch path the worker execute_model_batch builds the DiffusionRequestBatch and runs per-request setup (KV transfer, generator/seed), one cache refresh, and one LoRA activation per batch. Non-batch pipelines stay on the per-request execution path.
  • BatchRunnerOutput routes results per request_id, including per-request success, error, and abort outputs.
  • Diffusion IPC SHM packing now traverses both shallow RunnerOutput.result wrappers and nested BatchRunnerOutput.runner_outputs[*].result, so request-mode batch outputs with large tensors continue to use the SHM transfer path instead of falling back to pickle IPC.
  • Request-mode batching has an opt-in admission wait window through request_batch_max_wait_ms. The engine can wait briefly before the first schedule() of a wave so bursty compatible requests can accumulate, while the default 0.0 keeps the existing no-wait behavior. The scheduler exposes waiting/running queue counters for this logic, and the serve CLI / default diffusion stage config forward the option into OmniDiffusionConfig.

API boundary (breaking change)

  • The packed diffusion list-prompt submission path no longer constructs packed diffusion requests. AsyncOmni.generate(prompt=[...]) rejects early with a clear error when a diffusion stage is present; StagePool.submit_initial() / submit_update() keep a defensive list check.

Pipeline & test migration

  • Diffusion model pipelines and focused tests are migrated to the DiffusionRequestBatch forward contract; non-batch models change only signature/return type, batch-capable models add supports_request_batch = True plus per-request output splitting.
  • output_formatter.py also follows the single-request contract: it reads request.prompt, emits one OmniRequestOutput per request, and no longer keeps a request.prompts compatibility branch.

Compatibility/docs/release notes

  • This is a user-facing breaking change for diffusion AsyncOmni.generate(prompt=[...]) packed-list behavior. Users should submit multiple independent requests to use automatic scheduler batching.
  • vLLM-Omni docs now describe request-level diffusion batching, max_num_seqs, and the optional request_batch_max_wait_ms admission wait setting.

Test Plan

Core diffusion request-mode regression

Focused coverage for request contract, scheduler batching, runner behavior,
engine routing, multiprocess dispatch, and stage process integration:

python -m pytest \
  tests/diffusion/test_diffusion_request.py \
  tests/diffusion/test_diffusion_scheduler.py \
  tests/diffusion/test_diffusion_step_pipeline.py \
  tests/diffusion/test_diffusion_model_runner.py \
  tests/diffusion/test_diffusion_engine_cleanup.py \
  tests/diffusion/test_diffusion_engine_rpc_routing.py \
  tests/diffusion/test_multiproc_engine_concurrency.py \
  tests/diffusion/test_stage_diffusion_proc.py \
  -q

Touched backend and model coverage

Representative backend/model unit coverage for changed diffusion paths, excluding
hardware-dependent advanced model tests:

python -m pytest \
  tests/diffusion/diffusion_backend/test_diffusers_backend.py \
  tests/diffusion/models/dmd2/test_dmd2_request_sanitization.py \
  tests/diffusion/models/dmd2/test_dmd2_scheduler.py \
  tests/diffusion/models/flux2/test_flux2_klein_num_inference_steps.py \
  tests/diffusion/models/ovis_image/test_ovis_image.py \
  tests/diffusion/models/qwen_image/test_qwen_image_edit_plus.py \
  tests/diffusion/models/qwen_image/test_qwen_image_max_sequence_length.py \
  tests/diffusion/models/wan2_2/test_wan22_i2v_pipeline.py \
  tests/diffusion/models/wan2_2/test_wan22_pipeline_diffuse.py \
  tests/diffusion/models/wan2_2/test_wan22_vace_pipeline.py \
  -m "not advanced_model" \
  -q

Output formatter regression

Focused coverage for the single-prompt output formatting contract and the removal of the old multi-prompt compatibility path:

python -m pytest \
  tests/diffusion/test_diffusion_output_formatter.py \
  -q

IPC coverage

Large-tensor SHM packing coverage for request-mode and step-pipeline outputs:

python -m pytest \
  tests/diffusion/test_diffusion_ipc.py \
  tests/diffusion/test_diffusion_step_pipeline.py::TestIPC::test_pack_unpack_runner_output_shm \
  -q

Request-batch capability routing

Focused coverage for the supports_request_batch route from engine dispatch to
worker-side RequestBatch execution:

python -m pytest \
  tests/diffusion/test_diffusion_engine.py \
  tests/diffusion/test_diffusion_model_runner.py \
  tests/diffusion/test_multiproc_engine_concurrency.py::TestRequestModeDispatch \
  -q

Request-batch admission and CLI forwarding

Focused coverage for request-batch capability detection, admission wait,
scheduler queue counters, and the request_batch_max_wait_ms CLI/config path:

python -m pytest \
  tests/diffusion/test_diffusion_engine.py::TestRequestBatchCapability \
  tests/diffusion/test_diffusion_engine.py::TestRequestBatchAdmission \
  tests/diffusion/test_diffusion_scheduler.py \
  tests/entrypoints/test_async_omni_diffusion_config.py \
  -q

Qwen-Image performance comparison

Compare Qwen-Image performance with the same setup: 512x512, 20 denoising
steps, FLASH_ATTN, single A100, 1 warmup run, and 30 measured runs. The
baseline uses merge-base packed prompt-list; the candidate and StepScheduler
checks use two concurrent single-prompt requests with max_num_seqs: 2.

Residual old-contract check

Run a residual old-contract grep before submission:

rg "req\\.prompts|request\\.prompts|OmniDiffusionRequest\\(.*prompts|add_batch_request_async" \
  vllm_omni tests

Expected remaining add_batch_request_async hits should be limited to the abstract interface or explicit rejection stubs. Remaining req.prompts / request.prompts hits should be in DiffusionRequestBatch compatibility properties, pre-processing helpers, comments, or rejection messages — not in single-request execution paths.

Test Result

  • Core request-contract, scheduler, runner, engine-routing, stage, and multiproc coverage: 137 passed, 23 warnings in 66.91s
  • Representative touched backend/model unit coverage: 177 passed, 2 deselected, 20 warnings in 3.37s
  • IPC coverage: 10 passed, 16 warnings in 2.56s
  • Request-batch capability routing coverage: 25 passed, 16 warnings in 3.87s
  • Request-batch admission and CLI forwarding focused coverage: 72 passed, 20 warnings in 3.51s
  • KV flow regression coverage after OpenPI request-field migration: 18 passed, 17 warnings in 0.84s
  • Residual old-contract grep rerun. Remaining add_batch_request_async / req.prompts / request.prompts
    hits are in compatibility or rejection paths; focused prompts= check now only finds the
    StageDiffusionProc rejection test.

Static request-batch route checks:

  • Qwen-Image declares supports_request_batch = True.
  • Request mode keeps step_execution disabled, sets max_num_seqs: 2.
  • DiffusionEngine selects execute_batch for request-batch-capable pipelines when step_execution is disabled.

Qwen-Image request batching performance comparison:

  • Model: Qwen-Image
  • Setup: 512x512, 20 steps, FLASH_ATTN, single A100, 1 warmup run, 30 measured runs
  • Candidate uses request-mode scheduler batching with step_execution disabled. StepScheduler benchmark is intentionally not used for this PR result.
Path Commit ok/error mean ms p50 ms p90 ms amortized mean ms/image
Merge-base prompt-list b8f68174 30/0 4657.41 4658.68 4676.52 2328.70
Current request scheduler batch 7771f3c2 30/0 4664.54 4660.43 4701.68 2332.27

Current request scheduler batch delta vs merge-base prompt-list: mean +0.15%, p50 +0.04%, p90 +0.54%, amortized mean +0.15%. No significant performance regression is observed with the new request-level batching path.

@yJader
yJader force-pushed the refactor/request_id branch 4 times, most recently from 980da0e to 4f3e5ca Compare June 9, 2026 03:26
@yJader yJader changed the title [WIP] [Refactor] Move diffusion batching to request-level scheduler batches [Refactor] Move diffusion batching to request-level scheduler batches Jun 10, 2026
@yJader
yJader marked this pull request as ready for review June 10, 2026 01:55
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@yJader
yJader force-pushed the refactor/request_id branch from 7771f3c to c1859ab Compare June 10, 2026 02:03
@yJader yJader changed the title [Refactor] Move diffusion batching to request-level scheduler batches [Refactor] Migrate diffusion prompt-list batching to request-level batching Jun 10, 2026
@yJader yJader changed the title [Refactor] Migrate diffusion prompt-list batching to request-level batching [Refactor] Migrate diffusion prompt-list batching to request-level scheduler batching Jun 10, 2026
allow for future improvements.
Diffusion list-prompt input is not represented as one packed
`OmniDiffusionRequest`. Submit multiple prompts as independent requests to
use automatic request-level batching.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

how to control the batch size?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use max_num_seqs, consistent with the step-wise batching path

@SamitHuang

Copy link
Copy Markdown
Collaborator

I think it's better to compare step-level batching and request-level batching on time performance

@SamitHuang SamitHuang left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This refactor introduces request-level batching but is missing some critical validation and architectural cleanup.

  • vllm_omni/diffusion/request.py:35: If diffusion models later support multi-part single prompts (e.g., interleaved text and image dicts in a list), this check (isinstance(self.prompt, list)) will incorrectly reject them. Consider checking if the elements are valid single-prompt types rather than rejecting all lists.
  • vllm_omni/diffusion/worker/diffusion_model_runner.py:310: Hardcoding tea_cache specific logic in the model runner breaks encapsulation. The cache backend interface should define whether it requires num_inference_steps, or gracefully handle None without runner-side workarounds.
  • vllm_omni/entrypoints/async_omni.py:313: Since this is a user-facing API breaking change, ensure that this error message also points users to the updated documentation or an example script for the new independent request submission pattern.

@SamitHuang

SamitHuang commented Jun 13, 2026 •

Copy link
Copy Markdown
Collaborator

seems it does not support batching variable-length prompts in my test

@SamitHuang

Copy link
Copy Markdown
Collaborator

@knlnguyen1802 PTAL

@hsliuustc0106 hsliuustc0106 added ready label to trigger buildkite CI diffusion codes related to diffusion models merge-test label to trigger buildkite merge test CI and removed ready label to trigger buildkite CI labels Jun 29, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator
  • qwen-image
  • qwen-image-edit
  • WAN 2.1/2.2
  • SDXL
  • LTX 2.3

any updates?

… structure across various models

Signed-off-by: jader <yjader@foxmail.com>
@yJader

yJader commented Jun 29, 2026

Copy link
Copy Markdown
Contributor Author
  • qwen-image
  • qwen-image-edit
  • WAN 2.1/2.2
  • SDXL
  • LTX 2.3

any updates?

Currently request batch supports Qwen-Image, Flux, SD3.5, and LTX2.3. The implementation reuses/adapts the existing prompt-list batching path in pipelines, so we first enabled it for the models we care about most.

@SamitHuang SamitHuang added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Jun 29, 2026
@hsliuustc0106 hsliuustc0106 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI merge-test label to trigger buildkite merge test CI labels Jun 29, 2026
Signed-off-by: jader <yjader@foxmail.com>
@yJader
yJader force-pushed the refactor/request_id branch from 960df9c to dcf55c0 Compare June 29, 2026 15:29
@hsliuustc0106 hsliuustc0106 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Jun 29, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

I think this PR needs to polish for further abstraction, but let's wait for the next version

@hsliuustc0106
hsliuustc0106 merged commit c45ac74 into vllm-project:main Jun 29, 2026
9 checks passed
tzhouam added a commit that referenced this pull request Jun 30, 2026
…tch API

PR #4079 replaced per-request objects with `DiffusionRequestBatch` (whose
`sampling_params` is a read-only property) and `OmniDiffusionRequest` (which
exposes a singular `prompt`, not `prompts`). The two-stage video pipelines
were not migrated, so they crashed at worker dummy-run and the server never
became ready:

- ltx2/pipeline_ltx2.py, ltx2/pipeline_ltx2_image2video.py:
    `stage_2_req.sampling_params = req.sampling_params.clone()` hit
    "property 'sampling_params' of 'DiffusionRequestBatch' object has no
    setter". Now clone the underlying request(s) and override their
    sampling params, leaving the batch's read-only property alone.
- helios/pipeline_helios.py:
    `prepare_encode` built a single OmniDiffusionRequest then read
    `req.prompts`, hitting "'OmniDiffusionRequest' object has no attribute
    'prompts'". Now wrap the request in a DiffusionRequestBatch so the
    batch compatibility properties are available.

Fixes the "Diffusion X2V - Other Function/Accuracy Test" failures
(test_ltx2_expansion, test_video_streaming_output_similarity[helios_distilled]).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Signed-off-by: tzhouam <tzhouam@connect.ust.hk>
chickeyton added a commit to chickeyton/vllm-omni that referenced this pull request Jun 30, 2026
Integrate upstream main 0899a1a (11 commits since 1b318d1) — orchestrator
inter-stage/client output split (vllm-project#4527), diffusion request-level batching
(vllm-project#4079), speech SSE default (vllm-project#4679), and assorted model/example migrations.
Auto-merged with no conflicts.

Signed-off-by: chickeyton <ngton2014@gmail.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
chickeyton added a commit to chickeyton/vllm-omni that referenced this pull request Jun 30, 2026
…m-project#4079)

Upstream vllm-project#4079 ("request-level batching for diffusion pipelines") made
StagePool.submit_initial / submit_update reject diffusion list-prompt batch
requests. The orchestrator's AR→DiT handoff (_forward_to_next_stage) only
unwrapped a length-1 diffusion-prompt list on the streaming-update path, so the
initial submit still handed a list to submit_initial and the orchestrator thread
died on the first request reaching the DiT stage (EngineDeadError).

Unwrap a length-1 list for both the initial and the update submit so each
diffusion request is submitted independently. Multi-prompt batches (len > 1),
which previously rode the removed list-prompt path, now raise a clear error
since per-sub-request orchestration is not yet wired up.

Validated e2e on HunyuanImage TI2I (C2: 2-rep TP2 EP-off, stage-based): 36/36
requests succeeded, 0 list-prompt errors (was 1/4 then 0/16 + crash).

Signed-off-by: chickeyton <ngton2014@gmail.com>
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
zeningc added a commit to zeningc/vllm-omni that referenced this pull request Jul 1, 2026
Merging main pulled in new call sites that used the pre-rename output
queue API, which this branch had already renamed output_async_queue ->
output_sync_queue (put_nowait on the sync side). These were semantic
conflicts: git auto-merged them without a textual conflict, so they
landed referencing an attribute the constructor no longer sets.

- orchestrator.py:1251 (from vllm-project#4079, diffusion request-level batching)
  and orchestrator.py:1419 (from vllm-project#4257, Aura non-async-chunk path) still
  called `await self.output_async_queue.put(...)`, which would raise
  AttributeError on those terminal-output error/edge paths. Convert to
  `self.output_sync_queue.put_nowait(...)`.
- tests/engine/test_orchestrator_stage_input_bridge.py (new file from
  vllm-project#4257, marked core_model/cpu) constructed Orchestrator with the old
  `output_async_queue=` kwarg and failed against the new signature.
  Update to `output_sync_queue=output_q.sync_q`.

Tested on vLLM 0.24.0: engine + orchestrator unit tests (41 passed) and
a Qwen3-TTS-12Hz-0.6B streaming /v1/audio/speech smoke test.

Signed-off-by: zeningc <zening.chen@yahoo.com>
wkutak pushed a commit to wkutak/vllm-omni-cosmos-genai-nim that referenced this pull request Jul 13, 2026
…es (vllm-project#4079)

Signed-off-by: jader <yjader@foxmail.com>
Co-authored-by: Samit <285365963@qq.com>
Signed-off-by: Wojciech Kutak <wkutak@nvidia.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…es (vllm-project#4079)

Signed-off-by: jader <yjader@foxmail.com>
Co-authored-by: Samit <285365963@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion codes related to diffusion models ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants