Skip to content

[Core] Retain structured configs through runtime startup - #6849

Merged
linyueqian merged 3 commits into
vllm-project:mainfrom
maithilijoshi20:fix/6709-independent-runtime
Sep 17, 2026
Merged

linyueqian merged 3 commits into
vllm-project:mainfrom
maithilijoshi20:fix/6709-independent-runtime

Conversation

@maithilijoshi20

@maithilijoshi20 maithilijoshi20 commented Aug 31, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Complete the structured-config runtime-startup slice of RFC #6500.

  • Keep registered and generic-diffusion stages as BaseVllmOmniStageConfig instances through resolution, runtime planning, standard startup, and headless startup.
  • Build generic diffusion from shared normalization directly into VllmOmniConfig; the typed path no longer converts to or passes through an OmegaConf runtime stage.
  • Project typed owners at the remaining startup boundaries: connector injection, inferred KV topology, device layout, inline diffusion, local replica overrides, headless metadata, and speech-stage routing.
  • Preserve compatibility-only legacy builders for callers that have not migrated, without invoking them from the structured resolver.
  • Fix Ruff E402/format failures, stale diffusion constructor tests, registered default deploy-path propagation, CLI alias normalization, and MkDocs argparse extraction.

Test Coverage

Runtime-startup boundaries

  • Resolve generic diffusion into VllmOmniDiffusionStageConfig with typed sampling, parallel, runtime, and model owners.
  • Exercise standard startup through StageRuntime and launch_diffusion_stage_replica, stubbing only the model/process-client boundary.
  • Exercise headless startup through run_headless, launch_headless_diffusion_replicas, metadata extraction, config building, replica-group planning, and replica launch.
  • Cover CLI-over-strategy precedence, canonical diffusion output modality, typed KV connector mutation/topology, single-stage local-DP overrides, device slicing/layout, default deploy paths, DreamZero, and typed speech routing.

E2E configuration effects

These GPU tests inspect the live engine.stage_configs after startup. They verify that deploy settings have an observable runtime owner and cannot be silently ignored while the stage still starts.

  • Qwen3-Omni (2 GPU): stage type and topology; device placement; scheduler limits; cache and memory settings; trust_remote_code; default sampling parameters; connector edges; CUDA enforce_eager overrides.
  • HunyuanImage-3 (8 GPU): AR and diffusion placement; parallelism; scheduler and cache settings; sampling defaults; KV-transfer settings; connector edges; executor backend; eager mode.
  • WAN2.2 (1 GPU): generic diffusion stage type; model class; model identity; and video output modality.

Validation

  • CUDA Qwen-Image e2e suite: 3 passed , generation, log-prob output, and concurrent requests.
  • Ruff format and lint pass for the E2E configuration-effect tests.
  • Python syntax compilation and git diff --check pass.

The Qwen3-Omni, WAN2.2, and HunyuanImage-3 tests are hardware-marked. Qwen3-Omni requires 2 GPUs, WAN2.2 requires 1 GPU, and HunyuanImage-3 requires 8 H100 GPUs; CI provides execution for these hardware-specific tests.

vLLM Version: 2ae2bf5da39bc5069edc6f57270b307613c016c8

vLLM-Omni Commit: 64a03ab6b

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/stage_runtime.md.

Module owners: @tzhouam @fake0fan

@maithilijoshi20, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@maithilijoshi20 maithilijoshi20 changed the title [Refactor] Retain typed configs for runtime planning [Core] Preserve typed configuration through runtime planning Aug 31, 2026
@maithilijoshi20
maithilijoshi20 force-pushed the fix/6709-independent-runtime branch 2 times, most recently from bfb6c5e to 80fb909 Compare August 31, 2026 07:50
@maithilijoshi20
maithilijoshi20 force-pushed the fix/6709-independent-runtime branch from 80fb909 to 7b8c1d3 Compare August 31, 2026 07:59
@maithilijoshi20 maithilijoshi20 changed the title [Core] Preserve typed configuration through runtime planning [Core] Cut runtime over to structured Omni config Aug 31, 2026
@maithilijoshi20
maithilijoshi20 force-pushed the fix/6709-independent-runtime branch 2 times, most recently from 87db810 to 104bcf6 Compare August 31, 2026 08:40
@hsliuustc0106 hsliuustc0106 added core related to core module: cache, scheduler, engine, worker, modelrunner refactor refactoring for better code scalability and quality labels Aug 31, 2026
@vllm-omni-review-bot

vllm-omni-review-bot commented Aug 31, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 211d7e06f027 produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

@zwhzzz0821

Copy link
Copy Markdown
Contributor

I think we need to add sufficient end-to-end tests to demonstrate that the configuration replacement works correctly.

@maithilijoshi20
maithilijoshi20 force-pushed the fix/6709-independent-runtime branch 2 times, most recently from 2d03cde to 65a9230 Compare September 1, 2026 03:02

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed at 25f68031 after the rebase across #7413. The rebase itself is clean: vllm_omni/entrypoints/duplex/, vllm_omni/engine/duplex/, duplex_omni_engine.py and duplex_orchestrator.py are byte-identical to main, and the only omni_engine_base.py delta is the async_chunk read moving onto connector_config. The typed diffusion admission budget and the Higgs model resolution I asked about last round both check out: stage_admission.py reads cache_config.gpu_memory_utilization first and typed construction populates it, and both Higgs adapters resolve through resolve_stage_model_path.

The change I am asking for comes from the one structural move this PR makes: resolver.py stops projecting stages through stage.to_omegaconf() and hands the typed OmniStageConfig objects straight to runtime startup and to the frontend. That is the right direction, but three readers downstream still assume the OmegaConf shape and now degrade silently instead of failing. They are inline as [important]: per-stage runtime.env is dropped because _to_dict cannot convert a pydantic dataclass, nested --stage-overrides mappings now replace whole sections where the OmegaConf bridge deep-merged them, and the two video capability checks in serving_video.py still look for engine_args.model_class_name, the same reader this PR already fixed in video/generation/helpers.py. One [suggestion] on resolve_stage_model_path precedence. Once those three are covered (each is a small change plus a regression test with a typed stage) I expect to approve.

Validation: static read of the worktree against origin/main and a two-model panel; no PR code executed. GitHub Actions pre-commit, wheel build and DCO are green; the general Buildkite lane was re-fired at this head and is still running.

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Inline items for the request-changes review above (pullrequestreview-5226741717); they went out without their anchors, so here they are on the lines they refer to.

Comment thread vllm_omni/engine/stage_init_utils.py
Comment thread vllm_omni/config/resolver.py
Comment thread vllm_omni/entrypoints/openai/video/generation/helpers.py Outdated
"""Prefer typed stage model overrides, then legacy and served-model values."""
stages = getattr(engine_client, "stage_configs", ()) or ()
for stage in stages:
model_path = getattr(getattr(stage, "model_config", None), "model", None)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[suggestion] Returning the first stage that has model_config.model is fine for the shipped Higgs profiles, where no stage overrides the served model, but typed construction fills every stage's model_config.model with the served model when nothing is set, so a deployment that overrides only stage 1 would resolve stage 0's inherited value here while the legacy loop below would have returned the stage 1 override. If a later-stage override is meant to win, iterate stages in reverse or skip values equal to the served model; if not, a one-line docstring note that stage 0 wins would stop the next reader from chasing it.

Resolve the integration with current main while preserving typed config ownership through runtime startup.

Signed-off-by: maithilijoshi20 <97733343+maithilijoshi20@users.noreply.github.com>
Signed-off-by: maithilijoshi20 <97733343+maithilijoshi20@users.noreply.github.com>
@maithilijoshi20
maithilijoshi20 force-pushed the fix/6709-independent-runtime branch from 6468054 to 9a2a193 Compare September 17, 2026 02:16
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 17, 2026
Signed-off-by: maithilijoshi20 <97733343+maithilijoshi20@users.noreply.github.com>
@linyueqian linyueqian added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 17, 2026

@linyueqian linyueqian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Approving at 211d7e06, which is 9a2a1935 plus a test-only follow-up (the nested-override test now reads the Cosmos3 policy stage's diffusion_config.model_config, and the ComfyUI e2e test patches _should_serve_duplex after #7675 renamed it); the source hunks are unchanged. All three items from my request-changes round are fixed at the source, each with a test that would have caught the original: _to_dict in entrypoints/stage_utils.py now converts dataclass instances with asdict, so a typed OmniStageRuntimeConfig keeps its env through stage_runtime_env (test_stage_runtime_env_accepts_typed_runtime_config sets a variable through a typed config and reads it back inside the context); _stage_engine_values deep-merges mapping-valued CLI overrides through the same _get_recursively_merged_dict the deploy overlay uses, with omni_kv_config deliberately kept atomic, so a partial --stage-overrides no longer deletes sibling keys (test_nested_stage_override_deep_merges_structured_model_config checks that guardrails flips while policy_server_config survives, which is the Cosmos3 case I described); and the two ServingVideo capability checks now resolve the class name through one shared _stage_diffusion_model_class_name helper that video/generation/helpers.py also uses, trying diffusion_config.model_class_name, then model_config.model_arch, then the legacy engine_args, with test_typed_stage_drives_video_capability_checks feeding typed stages to both checks.

The rebase onto current main also picked up #7544 correctly: the engine's async_chunk now reads the typed connector_config first and falls back to legacy engine_args, so the any-stage rule from that PR holds for both shapes. The Higgs precedence note from last round stays a suggestion and does not affect the shipped profiles.

Validation: static comparison of the PR patch at 25f68031 against the patch at this head plus a read of the three new tests; no PR code executed. The branch is zero commits behind main with no conflicts. I re-fired the general lane at this head after the push and will merge on green.

@linyueqian
linyueqian merged commit f1a6e7c into vllm-project:main Sep 17, 2026
6 of 9 checks passed
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
…t#6849)

Signed-off-by: maithilijoshi20 <97733343+maithilijoshi20@users.noreply.github.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
@armaanamatya armaanamatya mentioned this pull request Sep 22, 2026
4 of 11 tasks
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…t#6849)

Signed-off-by: maithilijoshi20 <97733343+maithilijoshi20@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core related to core module: cache, scheduler, engine, worker, modelrunner high priority high priority issue, needs to be done asap ready label to trigger buildkite CI refactor refactoring for better code scalability and quality

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants