Repository navigation
[BugFix] Qwen-image performance regressed - Avoid mapping diffusion_batch_size onto scheduler max_num_seqs - #6525
Merged
Gaohan123 merged 2 commits intoAug 25, 2026
Conversation
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
|
This PR appears to belong to: docs/design/module/stage_runtime.md. Module owners: @tzhouam @fake0fan @NumberWan, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
NumberWan
force-pushed
the
fix/diffusion-seq-capacity-request-mode-only
branch
from
August 23, 2026 15:35
8a83219 to
294e569
Compare
2 tasks done
Shared diffusion stage init no longer overwrites max_num_seqs from the client batch knob. Wan request batching uses --max-num-seqs or YAML. Fixes vllm-project#6435 Signed-off-by: NumberWan <wantszkin2003@gmail.com>
NumberWan
force-pushed
the
fix/diffusion-seq-capacity-request-mode-only
branch
from
August 25, 2026 03:22
294e569 to
b56a1ee
Compare
NumberWan
requested review from
alex-jw-brooks and
lishunyang12
as code owners
August 25, 2026 03:22
Gaohan123
enabled auto-merge (squash)
August 25, 2026 09:57
1 task done
AndyZhou952
pushed a commit
to AndyZhou952/vllm-omni
that referenced
this pull request
Aug 26, 2026
…atch_size onto scheduler max_num_seqs (vllm-project#6525) Signed-off-by: NumberWan <wantszkin2003@gmail.com> Signed-off-by: AndyZhou952 <jzhoubc@connect.ust.hk>
JoseCarlosGarcia95
pushed a commit
to valendra-tech/vllm-omni
that referenced
this pull request
Sep 5, 2026
…atch_size onto scheduler max_num_seqs (vllm-project#6525) Signed-off-by: NumberWan <wantszkin2003@gmail.com>
khairulkabir1661
pushed a commit
to khairulkabir1661/vllm-omni
that referenced
this pull request
Sep 25, 2026
…atch_size onto scheduler max_num_seqs (vllm-project#6525) Signed-off-by: NumberWan <wantszkin2003@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fix Qwen-image performance Regression
Fixes
#6435
Summary
#5676 wrote
od_config.max_num_seqs = batch_sizeon the shared diffusion stage init path.batch_sizehere isdiffusion_batch_size(default 1, not bound to--max-num-seqs). Qwen-Image nightly serve uses--step-execution --max-num-seqs 8, so capacity became 1 and HTTP concurrency 8 queued serially.This PR deletes those two assignments. Scheduler capacity is only
--max-num-seqs/ stage YAMLmax_num_seqs.diffusion_batch_sizestays a client knob.Wan request-level batching of width 4 must set
--max-num-seqs 4(or YAML). The default deploy filevllm_omni/deploy/wan2_2_ti2v.yamlnow hasmax_num_seqs: 4. Replica-DP Wan YAML is unchanged (max_num_seqs: 1, one request per replica GPU).There is no shared
max(CLI, batch_size)helper.Nightly symptom (H100, from #6435)
test_qwen_image_single_device_step_execution_high_concurrency,c=8,n=160:throughput_qpse2e_latency_msAt
c=2/4/8, QPS stayed ~flat (~0.47) while latency scaled nearly linearly (~4.2s → 8.3s → 16.3s). That is serial admission, not a gradual compute regression.RCA (1× L20X, relative only)
Same case: 512×512, 20 steps,
c=8,n=160,--step-execution --max-num-seqs 8.First slowing commit:
2a315e1a—[Perf] Support request-level batching for Wan2.2 pipelines(#5676). Parentd1e230c9is still fast. Later days stay slow; no second independent ≥15% drop.Daily checkpoint (same case):
A/B (4 runs each):
d1e230c92a315e1aCV ~2–3%. Delete only
od_config.max_num_seqs = batch_sizeon the culprit → QPS 0.712 / 11.09s (parent-level). Isolates the overwrite, not Wan pipeline files.--max-num-seqsdiffusion_batch_sizeBehavior change
Call sites that relied on
diffusion_batch_size=Nalone to raise scheduler in-flight width will now stay at CLI/YAMLmax_num_seqs(often 1). Pass--max-num-seqs Nas well.Test plan
batch_size=4leavesmax_num_seqsat the CLI/YAML value (1 / 8 /None), not 4.c=1,2,4,8(512×512 / 20 steps /n=160). Expect QPS to scale withc, not stay flat at ~0.47.--max-num-seqs 4(YAML or CLI).diffusion_batch_sizealone must not raise scheduler capacity.