Skip to content

[bugfix][diffusion] Enable request-level batching via online entrypoint for diffusion models - #6405

Open
wtomin wants to merge 1 commit into
vllm-project:mainfrom
wtomin:diffusion-batching
Open

wtomin wants to merge 1 commit into
vllm-project:mainfrom
wtomin:diffusion-batching

Conversation

@wtomin

@wtomin wtomin commented Aug 20, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This PR makes vllm serve --omni --max-num-seqs N correctly control the default diffusion request batch size when --diffusion-batch-size is not explicitly provided (BTW, --diffusion-batch-size is not a valid CLI argument for vllm serve).

Previously, even if users started omni serving with:

vllm serve ... --omni --max-num-seqs N

the diffusion model could still run with diffusion_batch_size=1, so the documented request batching knob did not actually take effect for diffusion workloads.

While using AsyncOmni(diffusion_batch_size=N) is working correctly, this PR wants to enable diffusion batching currently for online entrypoint.

Motivation

During diffusion stage initialization, od_config.max_num_seqs is overwritten by the resolved diffusion batch size:

od_config.max_num_seqs = batch_size

Because of this, the effective diffusion batching capacity is determined by diffusion_batch_size.

Before this change, when diffusion_batch_size was not explicitly provided, it defaulted to 1. This caused the following behavior:

  • --max-num-seqs N was accepted by the server configuration.
  • Diffusion stage startup later overwrote od_config.max_num_seqs with diffusion_batch_size.
  • Since diffusion_batch_size was still 1, diffusion request batching remained effectively limited to batch size 1.
  • Users had to pass an extra diffusion-specific batch size option even though --max-num-seqs is the documented request batching knob.

What Changed

This PR changes omni engine argument handling so that diffusion_batch_size is resolved from max_num_seqs when it is not explicitly set:

explicit_batch = kwargs.pop("diffusion_batch_size", None)
max_num_seqs = kwargs.get("max_num_seqs") or 1
diffusion_batch_size = int(
    explicit_batch if explicit_batch is not None else max_num_seqs
)

The new behavior is:

  • If diffusion_batch_size is explicitly provided, it still takes precedence.
  • If diffusion_batch_size is not provided, it defaults to max_num_seqs.
  • If neither value is provided, the fallback remains 1.

Effect After This Change

After this PR, users can configure diffusion request batching with the standard serving option:

vllm serve ... --omni --max-num-seqs 4

In this case, the diffusion model will use an effective batch size of 4, unless --diffusion-batch-size is explicitly set to another value. An example output is:

INFO 08-20 08:58:55 [diffusion_pipeline_profiler.py:32] [DiffusionPipelineProfiler] Wan22Pipeline.text_encoder.forward took 0.087888s
INFO 08-20 08:58:55 [diffusion_pipeline_profiler.py:32] [DiffusionPipelineProfiler] Wan22Pipeline.text_encoder.forward took 0.087681s
INFO 08-20 08:58:56 [diffusion_pipeline_profiler.py:32] [DiffusionPipelineProfiler] Wan22Pipeline.text_encoder.forward took 0.087087s
INFO 08-20 08:58:56 [diffusion_pipeline_profiler.py:32] [DiffusionPipelineProfiler] Wan22Pipeline.text_encoder.forward took 0.087135s
100%|███████████████████████████████████████████████████████████████████████████████████████████| 50/50 [02:33<00:00,  3.07s/it]
INFO 08-20 09:01:29 [diffusion_pipeline_profiler.py:32] [DiffusionPipelineProfiler] Wan22Pipeline.diffuse took 153.415677s
INFO 08-20 09:01:29 [diffusion_pipeline_profiler.py:32] [DiffusionPipelineProfiler] Wan22Pipeline.diffuse took 153.415595s
INFO 08-20 09:01:45 [diffusion_pipeline_profiler.py:32] [DiffusionPipelineProfiler] Wan22Pipeline.vae.decode took 16.236813s
INFO 08-20 09:01:45 [diffusion_pipeline_profiler.py:32] [DiffusionPipelineProfiler] Wan22Pipeline.forward took 169.843718s
INFO 08-20 09:01:45 [diffusion_pipeline_profiler.py:32] [DiffusionPipelineProfiler] Wan22Pipeline.vae.decode took 16.306176s
INFO 08-20 09:01:45 [diffusion_pipeline_profiler.py:32] [DiffusionPipelineProfiler] Wan22Pipeline.forward took 169.917127s
(APIServer pid=543700) INFO 08-20 09:01:45 [diffusion_engine.py:578] [RequestBatch] admission wait done waiting=4 max_batch=4 waited_ms=0.0

[RequestBatch] admission wait done waiting=4 max_batch=4 waited_ms=0.0 shows the effective batch size. Before this change, when the effective batch size is 1, the Wan22Pipeline.diffuse took around 37s.

This makes omni serving behavior match user expectations and avoids a silent mismatch where --max-num-seqs appears configured but diffusion execution still runs with batch size 1.

Backward Compatibility

This change preserves existing behavior for users who explicitly set diffusion_batch_size.

The only behavior change is for the implicit/default case: diffusion serving now follows the documented --max-num-seqs batching setting instead of silently falling back to 1.

Signed-off-by: Didan Deng <33117903+wtomin@users.noreply.github.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/entrypoints.md.

Module owners: @alex-jw-brooks @linyueqian @NickCao

@wtomin, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@wtomin

wtomin commented Aug 20, 2026 •

Copy link
Copy Markdown
Collaborator Author

I simply use max_num_seqs to overwrite diffusion_batch_size, but if --diffusion-batch-size is an acceptable CLI argument to vllm serve, I am fine with the following solution:

    1. add --diffusion-batch-size to vllm_omni/entrypoint/cli/serve.py;
    1. Enable request-level batching by vllm serve --omni --max-num-seqs 4 --diffusion-batch-size 4

cc @knlnguyen1802

@hsliuustc0106 hsliuustc0106 added bug Something isn't working diffusion codes related to diffusion models labels Aug 21, 2026
@yJader

yJader commented Aug 21, 2026

Copy link
Copy Markdown
Contributor

I hit the same issue when testing Wan 2.2. It seems #5676 allows the legacy diffusion_batch_size=1 default to overwrite max_num_seqs.

Since #1851 made max_num_seqs the scheduler concurrency setting, should we remove the internal diffusion_batch_size / batch_size handling and just use max_num_seqs everywhere? The legacy Python argument could remain as a deprecated alias.

cc @LyxWxj

@LyxWxj

LyxWxj commented Aug 21, 2026 •

Copy link
Copy Markdown
Contributor

I investigated the usages of max_num_seqs and diffusion_batch_size.

I recommend removing diffusion_batch_size from the public/runtime API. It does not represent an independent execution concept. In the current implementation, it is primarily propagated through the Stage runtime and eventually assigned to od_config.max_num_seqs:

od_config.max_num_seqs = batch_size

As a result, max_num_seqs is the parameter actually consumed by the diffusion scheduler, execution-mode selection, KV-cache profiling, NPU capacity configuration, transfer adapters, and model-specific input processors.

Current usage of diffusion_batch_size

diffusion_batch_size is propagated through:

  • omni_base.py: lines 180, 205 — extracts from kwargs (default 1), passes to AsyncOmniEngine
  • async_omni_engine.py: lines 113, 123, 300 — constructor argument, saved as member, passed to create_stage_runtime
  • stage_runtime.py: lines 90, 127, 136, 662, 673, 757, 774, 903, 966, 1034, 1098, 1120, 1137 — remote stage context field, constructor parameter, member variable, passed as batch_size when launching local Diffusion Stage, prints batch size, DistStageRuntime parameter, passed to parent class, builds and saves remote Replica context, passes to remote StageDiffusionClient, Runtime Factory parameter, passes to DistStageRuntime, passes to normal StageRuntime
  • stage_init_utils.py: line 1506 — sets od_config.max_num_seqs = batch_size
  • stage_engine_startup.py: line 1561 — sets od_config.max_num_seqs = batch_size

It is therefore an alias for the Stage-level max_num_seqs, rather than an independent configuration.

Current usage of max_num_seqs

max_num_seqs is the canonical Stage-level capacity used by:

Configuration definition and merging

  • stage_config.py: lines 310, 1011, 1024, 1035 — per-stage StageDeployConfig.max_num_seqs; CLI/runtime override into engine args; compatibility with old max_batch_size; final engine args
  • omni_config.py: lines 187, 445, 460 — scheduler override type; OmniStageSchedulerConfig; validates max_num_batched_tokens >= max_num_seqs
  • diffusion/data.py: line 894 — OmniDiffusionConfig.max_num_seqs
  • config_factory.py: line 623 — default value 1

Diffusion Engine and Scheduler

  • diffusion_engine.py: lines 118, 262, 344, 1027 — reads max sequences; checks whether model supports request-level batching; prints final value; KV profile batch size
  • base_scheduler.py: line 74 — sets max_num_running_reqs
  • platform.py: line 100 — sets NPU MC2 token capacity

Transfer and input processing

The following files also use the semantics of max_num_seqs:

Proposed change

I suggest using max_num_seqs as the single source of truth throughout the full path:

CLI / deploy config
    -> per-stage engine args
    -> OmniDiffusionConfig.max_num_seqs
    -> scheduler and model runner

Specifically:

  1. Remove the redundant diffusion_batch_size argument from OmniBase, AsyncOmniEngine, and StageRuntime.
  2. Remove the two assignments that overwrite od_config.max_num_seqs from a separate batch_size argument.
  3. Let each Stage use its resolved stage_cfg.engine_args.max_num_seqs.
  4. Update the related tests and benchmark scripts.
  5. If backward compatibility is required, keep diffusion_batch_size temporarily as a deprecated alias for max_num_seqs, emit a warning, and remove it in a later release.

This would avoid the current inconsistency where --max-num-seqs 4 is accepted by the CLI but the Stage runtime still defaults to diffusion_batch_size=1, causing the final diffusion engine configuration to become max_num_seqs=1.

If maintainers agree with this direction, I can prepare a separate follow-up PR for the cleanup. I do not think it should block this PR.

@wtomin

wtomin commented Aug 24, 2026

Copy link
Copy Markdown
Collaborator Author

I hit the same issue when testing Wan 2.2. It seems #5676 allows the legacy diffusion_batch_size=1 default to overwrite max_num_seqs.

@yJader Thanks for your comments. I believe #5676 worked fine with diffusion_batch_size>1, because in its test script run_async_validation.py, it passed diffusion_batch_size key value via:

    omni = AsyncOmni(
        model=args.model,
        diffusion_batch_size=configured_batch_size, # configured_batch_size is defined via args.batch_size
        request_batch_max_wait_ms=250.0 if concurrency > 1 else 0.0,
        enforce_eager=True,
        dtype="bfloat16",
        boundary_ratio=0.875,
        flow_shift=5.0,
        log_stats=True,
    )

The major problem this PR tries to solve is when enabling request-level batching in online serving mode, diffusion_batch_size is always its default value (1), therefore, request-level batching cannot be truly enabled.

As for whether we should remove diffusion_batch_size, I think we should ask @SamitHuang @knlnguyen1802 @zhtmike from the verl-omni side . I am wondering when request-level batching is enabled, is it possible that diffusion_batch_size and max_num_seqs are not equal? For example, max_num_seqs=8, but diffusion_batch_size maybe smaller than 8 in order to get better performance?

@SamitHuang

Copy link
Copy Markdown
Collaborator

I also noticed this bug and set

diffusion_batch_size = int(
    explicit_batch if explicit_batch is not None else max_num_seqs
)

in verl-omni in verl-project/verl-omni#408. But this PR is still worth merging for root cause fixing

@SamitHuang
SamitHuang requested a lite review from Copilot August 24, 2026 02:42
@SamitHuang SamitHuang added ready label to trigger buildkite CI cuda-test Used to trigger vllm-omni cuda CI separately. labels Aug 24, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes omni online serving so the diffusion stage’s effective request batching follows the standard --max-num-seqs knob when diffusion_batch_size is not explicitly provided, avoiding the previous silent fallback to batch size 1.

Changes:

  • Update OmniBase.__init__ to resolve diffusion_batch_size from max_num_seqs when not explicitly set.
  • Preserve explicit diffusion_batch_size precedence while keeping the fallback behavior at 1 when neither value is present.
  • Add inline rationale documenting that diffusion stage init overwrites od_config.max_num_seqs with the resolved batch size.

💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment on lines +180 to +185
# Stage init overwrites ``od_config.max_num_seqs`` with this value.
# ``vllm serve --omni --max-num-seqs N`` is the documented request-batch
# knob; use it when ``diffusion_batch_size`` is not passed explicitly.
explicit_batch = kwargs.pop("diffusion_batch_size", None)
max_num_seqs = kwargs.get("max_num_seqs") or 1
diffusion_batch_size = int(explicit_batch if explicit_batch is not None else max_num_seqs)

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at 4d003faa. Current main already fixes the original problem through the scheduler path added in #6525: --max-num-seqs reaches per-stage runtime overrides and od_config.max_num_seqs, which is the value the diffusion scheduler uses for request concurrency.

This PR only changes the legacy diffusion_batch_size value passed through AsyncOmniEngine and StageRuntime. The subprocess client logs it, and the inline client stores it only for logging. It never changes scheduler capacity. The new comment therefore contradicts the current implementation and its tests.

Please close this PR as obsolete rather than rebasing it. If there is a separate intended use for the legacy client batch value, that should be demonstrated through a current consumer and a focused test.

Validation: static review of the changed file plus the current merged tree and diffusion client/runtime call path. No PR code was executed locally.

async_chunk = kwargs.get("async_chunk")
output_modalities = kwargs.pop("output_modalities", None)
diffusion_batch_size: int = kwargs.pop("diffusion_batch_size", 1)
# Stage init overwrites ``od_config.max_num_seqs`` with this value.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[blocking] This comment is false on current main. #6525 removed the batch_size -> od_config.max_num_seqs overwrite and made --max-num-seqs flow through stage runtime overrides into the diffusion scheduler directly. The legacy value changed here is only logged by the subprocess and inline clients, so this fallback cannot increase request concurrency. The intended fix has already shipped; please close this PR as obsolete rather than rebasing an ineffective assignment.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 20 days

@wtomin this pull request has had no human commit, comment or review since 2026-09-02. Please consider marking this PR as draft until work can resume. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 30, 2026 — with ChatGPT Codex Connector
@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict on zcode (GLM-5.3-Flash) under experiment fleet-strict-cursor-grok46-zcode-glm53flash-5050-c5-z10-20261002.

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: 3e94d4ed-a015-490a-8f91-c499eebe9305) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 120s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: d90e3b0f-d209-45ff-91ce-f63541b755ac) — check zcode login and the CLI version; retrying strict/zcode/GLM-5.3-Flash in 600s (try ).

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot attempt record

Review attempt ended as failed (step 'review' (agent.review_diff): unhandled error: RuntimeError: zcode exited 1 without a result event: statusCode: undefined } Error: Turn execution failed (traceId: ee40a5c2-4d03-4c17-9d22-226aeab11648) — check zcode login and the CLI version; falling back to direct/cursor/auto).

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

Changes since the previous review

  • 0 new inline finding(s); 1 finding(s) below.

CI at 4d003faac45d (2026-10-10T03:02:10.893629+00:00): verification incomplete; required-check status is unknown. Observed Buildkite: buildkite/vllm-omni (passed).

Note: The assigned review arm strict/zcode/GLM-5.3-Flash could not complete this review, so it was produced by the fallback arm direct/cursor/auto. It is excluded from the routing experiment.

Full review analysis

PR description

This PR changes OmniBase.__init__ so an omitted diffusion_batch_size is copied from max_num_seqs, and stays 1 only when that value is also missing. An explicit diffusion_batch_size still wins, and max_num_seqs is left in the kwargs forwarded to AsyncOmniEngine. The intended effect is that vllm serve --omni --max-num-seqs N supplies diffusion batch size N instead of the previous implicit default of 1. The stage-init code that the new comment says writes this value onto od_config.max_num_seqs is not in the materialized snapshot.

Change flow

flowchart LR
  cli["[EXISTING] explicit max_num_seqs kwarg"]:::existing
  explicit["[EXISTING] explicit diffusion_batch_size"]:::existing
  resolve["[CHANGED] OmniBase.__init__ default"]:::changed
  engine["[EXISTING] AsyncOmniEngine"]:::existing
  cli --> resolve
  explicit --> resolve
  resolve --> engine
  classDef existing fill:#e5e7eb,stroke:#6b7280,color:#111827
  classDef changed fill:#fef3c7,stroke:#d97706,color:#451a03,stroke-width:2px
  classDef new fill:#dcfce7,stroke:#16a34a,color:#052e16,stroke-width:2px
  classDef removed fill:#fee2e2,stroke:#dc2626,color:#450a0a,stroke-width:2px
Loading

Findings

  • [P1] Diffusion batch default has no regression test or test plan — vllm_omni/entrypoints/omni_base.py:185
    Existing thread: #6405 (comment)
Evidence for Diffusion batch default has no regression test or test plan

OmniBase.__init__ now sets diffusion_batch_size from max_num_seqs when diffusion_batch_size is omitted (line 185) and passes that value into AsyncOmniEngine (line 210). The diff adds no test for max_num_seqs=4 forwarding batch size 4, for an explicit diffusion_batch_size still winning, or for the missing-both fallback of 1. The PR body also has no Test Plan or Test Result. On this tree, .buildkite/cuda/test-ready.yml step "Simple · Engine&Entrypoints Test" runs pytest -sv tests/entrypoints tests/engine -m 'core_model and cpu', and "Entrypoints Test" runs pytest -sv tests/entrypoints/ -m 'core_model and cuda' --run-level 'core_model'. The PR does not report either selector. Without that pin, the online path can regress to batch size 1 again with no failing test. Add a constructor unit test for those three states and record the CPU entrypoint selector in the test plan.


🤖 This review was generated by InferMatrix Copilot, an open-source repo-maintenance agent for PR review, CI debugging and issue triage. Try it on your own repo, and ⭐ star it if it helped!

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: finding feedback

[p1] Diffusion batch default has no regression test or test plan — vllm_omni/entrypoints/omni_base.py:185

See the review for details. If you are the PR author and disagree, react 👎 here; the maintainer will see your disagreement.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working cuda-test Used to trigger vllm-omni cuda CI separately. diffusion codes related to diffusion models high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants