Skip to content

[Bugfix][Core] Fix CLI parallel flags losing to nested deploy parallel_config - #7786

Merged
NickCao merged 1 commit into
vllm-project:mainfrom
zwhzzz0821:fix/diffusion-cli-parallel-precedence
Sep 18, 2026
Merged

NickCao merged 1 commit into
vllm-project:mainfrom
zwhzzz0821:fix/diffusion-cli-parallel-precedence

Conversation

@zwhzzz0821

Copy link
Copy Markdown
Contributor

Purpose

Fix a precedence inversion introduced by the typed-config switch (#6849) that broke the NPU nightly HunyuanImage3 diffusion perf tests.

The legacy path StageConfig.to_omegaconf moves flat CLI parallel fields (tensor_parallel_size, ulysses_degree, ...) into the nested parallel_config dict via _apply_diffusion_parallel_runtime_overrides, so CLI flags win over a deploy YAML that nests its own parallel_config values. The new typed path _stage_engine_values skipped that step: CLI flags stayed flat, and _build_parallel_config then applied the nested deploy dict on top of the flat values — letting the deploy YAML beat the CLI.

In vllm_omni/deploy/hunyuan_image3_dit.yaml, the npu platform section sets nested parallel_config.tensor_parallel_size: 4, while the perf tests pass --tensor-parallel-size 2. The inversion turned the CLI's 2 back into the YAML's 4, so the stage demanded tp=4 × usp=2 = 8 devices on a 4-card NPU machine:

RuntimeError: Stage 0 requires 8 device(s) based on parallel_config, but 4 device(s) are available

This PR applies _apply_diffusion_parallel_runtime_overrides before reconcile_diffusion_attention_overrides in _stage_engine_values, mirroring the to_omegaconf order, and adds a regression test.

Fixes #7778

Test Plan

  • New regression test test_diffusion_cli_parallel_overrides_beat_nested_deploy_parallel_config in tests/config/test_omni_config.py: deploy config with nested parallel_config.tensor_parallel_size=4 + CLI overrides tensor_parallel_size=2, ulysses_degree=2, asserting the resolved stage has tp=2, ulysses=2, world_size=4.
  • Reproduced the failure on upstream main by resolving hunyuan_image3_dit.yaml with the perf test's CLI flags (platform forced to npu): before the fix the stage resolved to tp=4, usp=2, world_size=8; after the fix tp=2, usp=2, world_size=4.
  • Full config/engine unit suites: pytest tests/config/ tests/engine/test_stage_engine_args.py tests/engine/test_async_omni_engine_stage_init.py.
  • Real NPU e2e was not run locally (no NPU available); the NPU nightly lane should confirm.

vLLM Version: 0.29.0 (repo main)

vLLM-Omni Commit: 328d603

Test Result

  • pytest tests/config/test_omni_config.py: 221 passed, 1 skipped
  • pytest tests/config/ tests/engine/test_stage_engine_args.py tests/engine/test_async_omni_engine_stage_init.py: 884 passed, 2 skipped
  • Reproduction script (platform forced to npu, --tensor-parallel-size 2 --ulysses-degree 2 against hunyuan_image3_dit.yaml):
    • Before: tensor_parallel_size=4, ulysses_degree=2, world_size=8 (matches the nightly failure)
    • After: tensor_parallel_size=2, ulysses_degree=2, sequence_parallel_size=2, world_size=4

…y parallel_config

StageConfig.to_omegaconf moves flat CLI parallel fields (tensor_parallel_size,
ulysses_degree, ...) into the nested parallel_config dict via
_apply_diffusion_parallel_runtime_overrides, so CLI flags win over a deploy
YAML that nests its own parallel_config values. The typed-config path
(_stage_engine_values) skipped that step: the CLI flags stayed flat, and
_build_parallel_config then applied the nested deploy dict on top of the flat
values, letting the deploy YAML beat the CLI.

On the NPU nightly HunyuanImage3 perf tests this inverted precedence turned
--tensor-parallel-size 2 back into the YAML's 4, demanding 8 devices on a
4-card machine ("Stage 0 requires 8 device(s)... but 4 device(s) are available").

Apply _apply_diffusion_parallel_runtime_overrides before
reconcile_diffusion_attention_overrides in _stage_engine_values, matching the
to_omegaconf order, and add a regression test.

Fixes vllm-project#7778

Signed-off-by: zwhzzz0821 <2831474076@qq.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/vllm_omni_config.md, docs/design/module/entrypoints.md, docs/design/module/stage_runtime.md.

Module owners: @alex-jw-brooks @lishunyang12 @NickCao

Routing: @alex-jw-brooks via module of the changed files, semantic router, CODEOWNERS; @lishunyang12 via module of the changed files, semantic router, CODEOWNERS; @NickCao via module of the changed files, CODEOWNERS

@zwhzzz0821, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 328d603b6ec9 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@NickCao NickCao left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, now it indeed "Mirror StageConfig.to_omegaconf".

@NickCao NickCao added the ready label to trigger buildkite CI label Sep 18, 2026
@NickCao
NickCao enabled auto-merge (squash) September 18, 2026 12:45
@NickCao
NickCao merged commit 4b0c187 into vllm-project:main Sep 18, 2026
6 of 9 checks passed
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
…l_config (vllm-project#7786)

Signed-off-by: zwhzzz0821 <2831474076@qq.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…l_config (vllm-project#7786)

Signed-off-by: zwhzzz0821 <2831474076@qq.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ready label to trigger buildkite CI

Projects

None yet

3 participants