Repository navigation
[Bugfix][Core] Fix CLI parallel flags losing to nested deploy parallel_config - #7786
Conversation
…y parallel_config
StageConfig.to_omegaconf moves flat CLI parallel fields (tensor_parallel_size,
ulysses_degree, ...) into the nested parallel_config dict via
_apply_diffusion_parallel_runtime_overrides, so CLI flags win over a deploy
YAML that nests its own parallel_config values. The typed-config path
(_stage_engine_values) skipped that step: the CLI flags stayed flat, and
_build_parallel_config then applied the nested deploy dict on top of the flat
values, letting the deploy YAML beat the CLI.
On the NPU nightly HunyuanImage3 perf tests this inverted precedence turned
--tensor-parallel-size 2 back into the YAML's 4, demanding 8 devices on a
4-card machine ("Stage 0 requires 8 device(s)... but 4 device(s) are available").
Apply _apply_diffusion_parallel_runtime_overrides before
reconcile_diffusion_attention_overrides in _stage_engine_values, matching the
to_omegaconf order, and add a regression test.
Fixes vllm-project#7778
Signed-off-by: zwhzzz0821 <2831474076@qq.com>
|
This PR appears to belong to: docs/design/module/vllm_omni_config.md, docs/design/module/entrypoints.md, docs/design/module/stage_runtime.md. Module owners: @alex-jw-brooks @lishunyang12 @NickCao Routing: @alex-jw-brooks via module of the changed files, semantic router, CODEOWNERS; @lishunyang12 via module of the changed files, semantic router, CODEOWNERS; @NickCao via module of the changed files, CODEOWNERS @zwhzzz0821, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Omni ReviewBot triage noteAutomated triage of commit
These are automated triage suggestions only — the final decision belongs to the maintainers. |
NickCao
left a comment
There was a problem hiding this comment.
LGTM, now it indeed "Mirror StageConfig.to_omegaconf".
…l_config (vllm-project#7786) Signed-off-by: zwhzzz0821 <2831474076@qq.com> Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
…l_config (vllm-project#7786) Signed-off-by: zwhzzz0821 <2831474076@qq.com>
Purpose
Fix a precedence inversion introduced by the typed-config switch (#6849) that broke the NPU nightly HunyuanImage3 diffusion perf tests.
The legacy path
StageConfig.to_omegaconfmoves flat CLI parallel fields (tensor_parallel_size,ulysses_degree, ...) into the nestedparallel_configdict via_apply_diffusion_parallel_runtime_overrides, so CLI flags win over a deploy YAML that nests its ownparallel_configvalues. The new typed path_stage_engine_valuesskipped that step: CLI flags stayed flat, and_build_parallel_configthen applied the nested deploy dict on top of the flat values — letting the deploy YAML beat the CLI.In
vllm_omni/deploy/hunyuan_image3_dit.yaml, thenpuplatform section sets nestedparallel_config.tensor_parallel_size: 4, while the perf tests pass--tensor-parallel-size 2. The inversion turned the CLI's 2 back into the YAML's 4, so the stage demandedtp=4 × usp=2 = 8devices on a 4-card NPU machine:This PR applies
_apply_diffusion_parallel_runtime_overridesbeforereconcile_diffusion_attention_overridesin_stage_engine_values, mirroring theto_omegaconforder, and adds a regression test.Fixes #7778
Test Plan
test_diffusion_cli_parallel_overrides_beat_nested_deploy_parallel_configintests/config/test_omni_config.py: deploy config with nestedparallel_config.tensor_parallel_size=4+ CLI overridestensor_parallel_size=2, ulysses_degree=2, asserting the resolved stage hastp=2, ulysses=2, world_size=4.hunyuan_image3_dit.yamlwith the perf test's CLI flags (platform forced tonpu): before the fix the stage resolved totp=4, usp=2, world_size=8; after the fixtp=2, usp=2, world_size=4.pytest tests/config/ tests/engine/test_stage_engine_args.py tests/engine/test_async_omni_engine_stage_init.py.vLLM Version: 0.29.0 (repo main)
vLLM-Omni Commit: 328d603
Test Result
pytest tests/config/test_omni_config.py: 221 passed, 1 skippedpytest tests/config/ tests/engine/test_stage_engine_args.py tests/engine/test_async_omni_engine_stage_init.py: 884 passed, 2 skippednpu,--tensor-parallel-size 2 --ulysses-degree 2againsthunyuan_image3_dit.yaml):tensor_parallel_size=4, ulysses_degree=2, world_size=8(matches the nightly failure)tensor_parallel_size=2, ulysses_degree=2, sequence_parallel_size=2, world_size=4