Skip to content

[BugFix] Qwen-image performance regressed - Avoid mapping diffusion_batch_size onto scheduler max_num_seqs - #6525

Merged
Gaohan123 merged 2 commits into
vllm-project:mainfrom
NumberWan:fix/diffusion-seq-capacity-request-mode-only
Aug 25, 2026
Merged

Gaohan123 merged 2 commits into
vllm-project:mainfrom
NumberWan:fix/diffusion-seq-capacity-request-mode-only

Conversation

@NumberWan

@NumberWan NumberWan commented Aug 23, 2026 •

Copy link
Copy Markdown
Contributor

Fix Qwen-image performance Regression

Fixes

#6435

Summary

#5676 wrote od_config.max_num_seqs = batch_size on the shared diffusion stage init path. batch_size here is diffusion_batch_size (default 1, not bound to --max-num-seqs). Qwen-Image nightly serve uses --step-execution --max-num-seqs 8, so capacity became 1 and HTTP concurrency 8 queued serially.

This PR deletes those two assignments. Scheduler capacity is only --max-num-seqs / stage YAML max_num_seqs. diffusion_batch_size stays a client knob.

Wan request-level batching of width 4 must set --max-num-seqs 4 (or YAML). The default deploy file vllm_omni/deploy/wan2_2_ti2v.yaml now has max_num_seqs: 4. Replica-DP Wan YAML is unchanged (max_num_seqs: 1, one request per replica GPU).

There is no shared max(CLI, batch_size) helper.

Nightly symptom (H100, from #6435)

test_qwen_image_single_device_step_execution_high_concurrency, c=8, n=160:

Metric Current Baseline Change
throughput_qps 0.4799 0.8159 −41.2%
e2e_latency_ms 16309 9656 +68.9%

At c=2/4/8, QPS stayed ~flat (~0.47) while latency scaled nearly linearly (~4.2s → 8.3s → 16.3s). That is serial admission, not a gradual compute regression.

RCA (1× L20X, relative only)

Same case: 512×512, 20 steps, c=8, n=160, --step-execution --max-num-seqs 8.

First slowing commit: 2a315e1a — [Perf] Support request-level batching for Wan2.2 pipelines (#5676). Parent d1e230c9 is still fast. Later days stay slow; no second independent ≥15% drop.

Daily checkpoint (same case):

Day tip QPS Mean latency
8/15 (still good) ~0.70 ~11.3s
8/16 (first bad day) ~0.53 ~14.8s
8/20 (issue tip) ~0.51 ~15.4s

A/B (4 runs each):

Variant Median QPS Median mean latency vs parent
Parent d1e230c9 0.723 10.95s reference
Culprit 2a315e1a 0.502 15.59s QPS −30.6%, latency +42.4%

CV ~2–3%. Delete only od_config.max_num_seqs = batch_size on the culprit → QPS 0.712 / 11.09s (parent-level). Isolates the overwrite, not Wan pipeline files.

Knob Meaning Qwen nightly
--max-num-seqs scheduler in-flight 8
diffusion_batch_size client request-batch width (#1593) default 1

Behavior change

Call sites that relied on diffusion_batch_size=N alone to raise scheduler in-flight width will now stay at CLI/YAML max_num_seqs (often 1). Pass --max-num-seqs N as well.

Test plan

  • Unit tests: shared init batch_size=4 leaves max_num_seqs at the CLI/YAML value (1 / 8 / None), not 4.
  • Qwen-Image step-execution sweep c=1,2,4,8 (512×512 / 20 steps / n=160). Expect QPS to scale with c, not stay flat at ~0.47.
  • Wan request-level batching with --max-num-seqs 4 (YAML or CLI). diffusion_batch_size alone must not raise scheduler capacity.

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/stage_runtime.md.

Module owners: @tzhouam @fake0fan

@NumberWan, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@NumberWan
NumberWan force-pushed the fix/diffusion-seq-capacity-request-mode-only branch from 8a83219 to 294e569 Compare August 23, 2026 15:35
Shared diffusion stage init no longer overwrites max_num_seqs from the
client batch knob. Wan request batching uses --max-num-seqs or YAML.

Fixes vllm-project#6435

Signed-off-by: NumberWan <wantszkin2003@gmail.com>
@NumberWan
NumberWan force-pushed the fix/diffusion-seq-capacity-request-mode-only branch from 294e569 to b56a1ee Compare August 25, 2026 03:22
@NumberWan NumberWan changed the title [BugFix] Qwen-image performance regressed - Keep step-execution max_num_seqs free of diffusion_batch_size [BugFix] Qwen-image performance regressed - Avoid mapping diffusion_batch_size onto scheduler max_num_seqs Aug 25, 2026
@hsliuustc0106 hsliuustc0106 added the bug Something isn't working label Aug 25, 2026

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123 Gaohan123 added the ready label to trigger buildkite CI label Aug 25, 2026
@Gaohan123
Gaohan123 enabled auto-merge (squash) August 25, 2026 09:57
@Gaohan123
Gaohan123 merged commit bc8ea48 into vllm-project:main Aug 25, 2026
8 of 9 checks passed
AndyZhou952 pushed a commit to AndyZhou952/vllm-omni that referenced this pull request Aug 26, 2026
…atch_size onto scheduler max_num_seqs (vllm-project#6525)

Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Signed-off-by: AndyZhou952 <jzhoubc@connect.ust.hk>
JoseCarlosGarcia95 pushed a commit to valendra-tech/vllm-omni that referenced this pull request Sep 5, 2026
…atch_size onto scheduler max_num_seqs (vllm-project#6525)

Signed-off-by: NumberWan <wantszkin2003@gmail.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…atch_size onto scheduler max_num_seqs (vllm-project#6525)

Signed-off-by: NumberWan <wantszkin2003@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants