Skip to content

[CI][Qwen3-Omni] Compare structured-config sampling defaults against vLLM normalization - #7742

Merged
NickCao merged 1 commit into
vllm-project:mainfrom
linyueqian:fix/qwen3-omni-structured-config-sampling-test
Sep 18, 2026
Merged

NickCao merged 1 commit into
vllm-project:mainfrom
linyueqian:fix/qwen3-omni-structured-config-sampling-test

Conversation

@linyueqian

Copy link
Copy Markdown
Collaborator

Purpose

Turn the merge-pipeline step Omni · Qwen3-Omni Test green again. tests/e2e/offline_inference/test_qwen3_omni.py::test_structured_multistage_config_reaches_runtime (added in #6849) has failed on every main build that ran the step since it landed: 15444, 15507, 15510 and 15523, each at the same assertion (test_qwen3_omni.py:131, assert all(...) over the sampling defaults), with the other 12 tests in the step passing.

Closes #7732.

Root cause

The code2wav stage in vllm_omni/deploy/qwen3_omni_moe.yaml is greedy (temperature: 0.0) and spells the disabled top_k as -1. vLLM v0.29.0 SamplingParams.__post_init__ normalizes greedy sampling to top_k=0, top_p=1.0, min_p=0.0 (and quietly accepts -1 as the legacy disable sentinel), so the runtime SamplingParams held by the stage client never equals the raw YAML literal -1. The test compared the runtime object against the YAML literals directly, which fails for that one field. The thinker and talker stages match because none of their deploy values are touched by normalization.

This is a test expectation bug; the deploy values do reach the runtime.

Change

  • Keep the YAML literals as the deploy-owned expectation.
  • Assert the resolver carries them verbatim into the typed stage config (stage.model_config.default_sampling_params), where the pipeline only adds its own detokenize / stop_token_ids constraints.
  • Compare the runtime StageClient defaults against SamplingParams(**expected) built from the same values, so vLLM's own normalization applies on both sides and the check stays correct if that normalization changes again.

Test-only change; no runtime code touched.

Validation

  • ruff check and ruff format --check pass on the file.
  • The test is advanced_model and only runs in the merge pipeline (H100, 2 cards), so the PR lane cannot exercise it. The root cause is read directly from vLLM v0.29.0 vllm/sampling_params.py (greedy branch in __post_init__) and from the four failing merge logs; a maintainer with H100 access can confirm with:
pytest -s -v tests/e2e/offline_inference/test_qwen3_omni.py -k test_structured_multistage_config_reaches_runtime --run-level advanced_model

…vLLM normalization

test_structured_multistage_config_reaches_runtime has failed on every main
merge build since it landed with vllm-project#6849 (15444, 15507, 15510, 15523), always
at the sampling-defaults assertion. The code2wav stage in
vllm_omni/deploy/qwen3_omni_moe.yaml is greedy (temperature 0.0) and spells
the disabled top_k as -1; vLLM v0.29.0 SamplingParams normalizes greedy
sampling to top_k=0 (and top_p=1.0, min_p=0.0), so the runtime object never
equals the raw YAML literal.

Keep the YAML literals as the deploy-owned expectation, check that the
resolver carries them verbatim into the typed stage config, and compare the
runtime StageClient defaults against SamplingParams built from the same
values so vLLM's own normalization is applied on both sides.

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to be related to model: qwen-omni.

Model owners: @amy-why-3459 @yenuo26 @NickCao

Routing: @amy-why-3459 via semantic router; @yenuo26 via CI owner, CODEOWNERS; @NickCao via CODEOWNERS

@linyueqian, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@hsliuustc0106 hsliuustc0106 added CI/CD codes related to changes to CI/CD omni code related to omni models labels Sep 17, 2026

@NickCao NickCao left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM, comparing against the normalized configuration instead of the raw literals matches the test's intention.

@NickCao
NickCao enabled auto-merge (squash) September 17, 2026 18:18
@NickCao

NickCao commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator
pytest -s -v tests/e2e/offline_inference/test_qwen3_omni.py -k test_structured_multistage_config_reaches_runtime --run-level advanced_model

passed, 1 passed, 2 deselected, 19 warnings in 450.33s (0:07:30)

@NickCao NickCao added the ready label to trigger buildkite CI label Sep 18, 2026
@NickCao
NickCao merged commit d4ffde1 into vllm-project:main Sep 18, 2026
6 of 9 checks passed
tzhouam added a commit that referenced this pull request Sep 21, 2026
Brings in 65 main commits (825ab20..ae3880f). Motivation: the three
CI failures left on Buildkite #3060 are main drift, not rebase regressions:

- MiniCPM-o 4.5 seed-tts perf (ready + nightly): the identical request
  ("Tim Tebow ..." item, 30.2 s of silence, response never reaching
  response.done, audio_past_key_values reset at 1500) also fails on main's
  own builds #3049 and #3052, and passes on main from #3054 on, after
  [CI/Build][MiniCPM-o] Fix the Nightly tests (#7758, 6df62d5), which
  the rebase branch did not carry.
- Nemotron VoiceChat native-duplex e2e: main marked the test skip in the
  same #7758 (the pipeline declares no duplex_plugin, so /v1/realtime falls
  through to vLLM's speech-to-text handler and rejects session.update);
  main's build #3054 skipped it, ours ran it.

Conflict resolutions:

- tests/e2e/offline_inference/test_qwen3_omni.py: main's side. #7742/#7803
  compare the deploy-YAML literal (top_k -1) verbatim and normalize the
  runtime check through SamplingParams, which supersedes 3dee645's
  top_k=0 expectation.
- vllm_omni/diffusion/diffusion_kv/model_runner_backend.py: both sides
  dropped. Main (#7166) removed the scheduler_config.max_num_seqs override;
  the rebase removed the kv_caches list because f2aad6aa70's init_kv_cache
  returns the caches as a dict.
- vllm_omni/benchmarks/patch/patch.py: main's Video-MME helpers kept
  together with the rebase's get_samples(args, tokenizer, **kwargs)
  signature; both delegate paths forward **kwargs.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
…vLLM normalization (vllm-project#7742)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…vLLM normalization (vllm-project#7742)

Signed-off-by: Yueqian Lin <linyueqian@outlook.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI/CD codes related to changes to CI/CD omni code related to omni models ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[CI] Omni · Qwen3-Omni Test (advanced_model) is red on main: builds 15507 and 15510, last green 15400

4 participants