Skip to content

[CI/Build] Fix Qwen3-Omni Talker benchmark workload - #8057

Merged
Gaohan123 merged 8 commits into
vllm-project:mainfrom
psv666:fix/7915-fixed-talker-perf
Sep 24, 2026
Merged

Gaohan123 merged 8 commits into
vllm-project:mainfrom
psv666:fix/7915-fixed-talker-perf

Conversation

@psv666

@psv666 psv666 commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

The Qwen3-Omni non-async random performance test fixes Thinker output at 900 tokens
but lets Talker stop naturally. Different audio lengths change both the workload
and the fixed Thinker overhead per generated audio second, making RTF comparisons
hard to interpret. Related to #7915.

This change makes the five random concurrency points request exactly 1536 Talker
output tokens through the existing sampling_params_list API. It preserves the
Talker temperature/top-k/repetition penalty, fixes the request seed, and explicitly
sets benchmark seed and input-length range. Warmups retain the existing
max(2, concurrency) request count. Code2Wav retains server defaults. CUDA and NPU share this configuration; async/random-mm tests and
production/model/API defaults are unchanged.

Benchmark results retain ordered request_stage_metrics for formal requests,
including missing entries, without audio bodies. The runner rejects failed
requests, missing metrics and mismatched output counts for stages configured with
min_tokens == max_tokens. Tests cover exact lengths, premature termination,
missing metrics, compatibility with nonfixed workloads, warmup exclusion and
out-of-order completion.

H100 baseline status: the random row's baseline is cleared ("baseline": {}),
as agreed with the CI owner. The five historical values were measured with the
previous variable-length workload and are not comparable, so they are removed
rather than left as false regression targets. Fixed-workload H100 medians will be
added in follow-up #8100. Local L20X results are not used as H100
baselines. The random-mm rows keep their existing baselines.

This addresses benchmark comparability; it does not establish that the historical
audio-shortening behavior in #7915 has been fixed.

Test Plan

vLLM Version: 0.30.0
vLLM-Omni Commit: upstream base 75570ee3a6572a7dc7a9463ada35a9a11c367993, plus commit 491193eff401a2034b7fca119ee5f4a519d200c9.

Local environment: 2 NVIDIA L20X GPUs (~140 GiB each), driver 570.133.20,
PyTorch 2.13.0+cu129, Transformers 5.14.1, flashinfer-python 0.6.18.post1,
NumPy 2.3.5, aiohttp 3.14.1. Model:
Qwen/Qwen3-Omni-30B-A3B-Instruct, cached snapshot
26291f793822fb6be9555850f06dfe95f2d7e695. These runs used a shared host and uv
environments, not a pinned H100 CI container. The environment reported an existing
flashinfer-cubin 0.6.13 / flashinfer-python version mismatch; this same runtime was
used for both local workloads.

Validation scope: the L20X measurements below used two warmups per point,
before follow-up commit 535201079 restored max(2, concurrency). That follow-up
passed 43 CPU runner tests, a five-point runner argument check and pre-commit;
the GPU sweep has not been rerun with the restored warmup counts; only the warmup count changed and warmups are excluded from metrics, so the effect is expected to be negligible.
H100 recalibration and NPU runtime validation have not been performed because
those devices are not available locally. Follow-up validation can be arranged through CI or contributors
with access to the required hardware. The calibration procedure is tracked in #8100.

Reproduction commands for the recorded GPU measurements

Requires the repository uv environment and the cached model snapshot above.
On a shared NVIDIA host, run the server within a two-GPU reservation (gpu run).
From the checkout, start the server:

uv run --no-sync python -m vllm_omni.entrypoints.cli.main serve \
  Qwen/Qwen3-Omni-30B-A3B-Instruct --omni --no-async-chunk \
  --host 127.0.0.1 --port 8091 --stage-init-timeout 900 --init-timeout 1200

After /health is ready, run the fixed-workload sweep from another terminal:

sampling='{"return_stage_metrics":true,"sampling_params_list":[{"temperature":0.0,"max_tokens":900,"ignore_eos":true},{"temperature":0.9,"top_k":50,"repetition_penalty":1.05,"seed":0,"min_tokens":1536,"max_tokens":1536}]}'
for c in 1 4 8 16 32; do
  uv run --no-sync python -m vllm_omni.entrypoints.cli.main bench serve --omni \
    --model Qwen/Qwen3-Omni-30B-A3B-Instruct --host 127.0.0.1 --port 8091 \
    --backend openai-chat-omni --endpoint /v1/chat/completions --dataset-name random \
    --random-input-len 2500 --random-output-len 900 --random-range-ratio 0 \
    --num-prompts "$((4*c))" --max-concurrency "$c" --request-rate inf \
    --num-warmups 2 --ignore-eos --seed 0 --extra-body "$sampling" \
    --percentile-metrics ttft,tpot,itl,e2el,audio_rtf,audio_ttfp,audio_duration \
    --save-result --save-detailed --result-dir fixed-results --result-filename "c${c}.json"
done

The command above reproduces the recorded two-warmup measurements. To run the
current PR configuration, replace --num-warmups 2 with
--num-warmups "$((c < 2 ? 2 : c))".

For the original workload, use unmodified base 75570ee3a, omit --extra-body,
and use warmups max(2, c). The original seed/range defaults are already 0/0.0.
The original-main experiment exported existing formal-request stage metrics with
an external client observer after timing stopped; source and requests were unchanged.

Unit validation (CPU tests; no model weights required):

CUDA_VISIBLE_DEVICES='' uv run --no-sync python -m pytest -q -o addopts='' \
  tests/dfx/perf/tests/test_runner_metadata.py \
  tests/benchmarks/patch/test_patch.py \
  tests/entrypoints/openai_api/test_serving_chat_sampling_params.py

CI-style marker selection for the changed benchmark test:

CUDA_VISIBLE_DEVICES='' uv run --no-sync python -m pytest -q -o addopts='' \
  -m 'core_model and cpu' --run-level core_model tests/benchmarks/patch/test_patch.py

Test Result

  • Final local unit run: 158 passed, 21 warnings in 6.09s; local pre-commit gates passed (including Ruff,
    mypy, Markdown, SPDX and test markers).
  • Fixed workload: 276/276 formal requests successful, Thinker 900 and
    Talker 1536 tokens on every request, Talker finish reason length, nonempty
    audio. This covers all five concurrency points plus a concurrency-8 repeat.
  • Original main/original workload: 244/244 successful, Thinker 900 tokens
    on every request, 193 Talker natural stops and 51 stops at the default 4096-token
    limit. One fresh service covered all five points.

Fixed-workload results (two service starts, A then B; PCM sample counts are at
24 kHz and exclude warmups):

Run Concurrency Requests Mean audio RTF PCM samples: request count
A 1 4 0.150477 2943870: 4
A 4 16 0.200603 2943870: 16
A 8 32 0.216688 2943870: 31, 2947200: 1
B 8 32 0.221115 2943870: 31, 2946090: 1
B 16 64 0.289418 2943870: 60, 2947200: 4
B 32 128 0.377598 2943870: 118, 2946090: 1, 2947200: 9

Run A used GPUs 6/7; run B used GPUs 4/7 and temporary diagnostic logs. Run A stopped
after c=8 because an extra diagnostic assertion required exact waveform-length
equality; request and token checks passed. Run B recorded the length distribution
while retaining request/token checks. Temporary model instrumentation is removed
from this diff. These are functional measurements, not three-round calibration.

Original main/original-workload results (GPUs 6/7; warmups 2/4/8/16/32):

Concurrency Requests Mean audio RTF Mean Talker tokens Mean audio seconds
1 4 0.129204 2829.50 226.05453
4 16 0.157391 2684.06 214.42820
8 32 0.206043 2613.94 208.82254
16 64 0.256951 2551.55 203.83599
32 128 0.330041 2702.34 215.88926

These tables use different output lengths and warmup settings. They demonstrate
why workload control is necessary and do not measure a code speedup/regression.

Fixed-length audio ranged from 2,943,870 to 2,947,200 samples
(122.66125-122.8 seconds, approximately 0.113% spread). Diagnostic logs located the
small variation in Code2Wav batching/tail cropping: 1536 Talker steps yielded 1535
valid codec rows; a split into 1026+509 rows can return a different tail length
when a shorter segment is cropped from a padded batch. For example,
1,969,920 + 976,170 = 2,946,090 samples, matching the observed +0.0925 seconds.
This does not explain the historical large shortening, and no model fix is included.
Natural-termination and audio-quality tests remain separate.

Signed-off-by: psv666 <2693925048@qq.com>
@psv666
psv666 marked this pull request as ready for review September 23, 2026 06:48
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/benchmarking.md.

Module owners: @alex-jw-brooks @Bounty-hunter @yenuo26

Routing: @alex-jw-brooks via module of the changed files, CODEOWNERS; @Bounty-hunter via module of the changed files, CODEOWNERS; @yenuo26 via CI owner, CODEOWNERS

@psv666, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@psv666

psv666 commented Sep 23, 2026

Copy link
Copy Markdown
Contributor Author

@Gaohan123 @natureofnature @yenuo26 PTAL

@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Omni ReviewBot triage note

Automated triage of commit 3dda9d0d870b produced:

  • Priority: high. Prompt maintainer attention is suggested.

These are automated triage suggestions only — the final decision belongs to the maintainers.

Signed-off-by: psv666 <2693925048@qq.com>
@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Omni ReviewBot: superseded

The CI failure noted on 5352010795d9 refers to an earlier head; the pull request now points at beb9650fd7c7.

@yenuo26

yenuo26 commented Sep 23, 2026

Copy link
Copy Markdown
Collaborator

@vllm-omni-review-bot

@vllm-omni-review-bot

Copy link
Copy Markdown
Omni ReviewBot routing record

Assigned Strict under experiment vllm-omni-strict-5050-20260829.

# fixed stage workloads without parsing logs or storing audio payloads.
request_stage_metrics = [getattr(output, "stage_metrics", None) for output in outputs]
if any(snapshot is not None for snapshot in request_stage_metrics):
result["request_stage_metrics"] = request_stage_metrics

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If it is only added here, it will merely be written into the result file and will not appear in the benchmark's printed output. I'm not sure whether this matches your expectation.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Intentional: the runner reads the saved result JSON, so the check doesn't depend on stdout, and printing 128 per-request snapshots would flood the log. I added a one-line summary after the check passes (Fixed stage workload OK: 128 requests, num_tokens_out stage 0=900, stage 1=1536), and trimmed the persisted snapshot to num_tokens_out/finish_reason/audio_frames/audio_duration_s per stage.

for stage, expected in fixed_lengths.items():
metrics = snapshot.get(stage) if isinstance(snapshot, dict) else None
actual = metrics.get("num_tokens_out") if isinstance(metrics, dict) else None
assert actual == expected, (

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you want to add a validation here that the talker needs to reach max_token, then shouldn't the thinker also add a validation that it reaches max_tokens after ignore_eos takes effect?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. Thinker now has min_tokens == max_tokens == 900, so stage 0 enters the same check. Behavior is unchanged (Thinker has no extra stop tokens and was already 900 on all 276 local requests), but it is now enforced. Added reject cases for a missing/short Thinker.

@vllm-omni-review-bot vllm-omni-review-bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Omni ReviewBot review

Scan:

Category Result
Tests / verification 9 finding(s) below
Security no finding reported
Docs / comments no finding reported
Behavior / compatibility no finding reported
Correctness no finding reported

Validated:

  • [claim-verified] 13 omitted Code2Wav tail cloned from deploy defaults: serving_chat.py:1301-1303 appends clone_sampling_params(default_params_list[idx]).
  • [claim-verified] H100 arrays 0.169..0.4595 are still in the JSON; no L20X substitution at this head.
  • [claim-refuted] README:46 Do not merge this workload change with those old baselines vs JSON:66-74 still shipping them as baseline.H100.mean_audio_rtf
  • [validated] pipeline.py:62 Talker sampling_constraints stop_token_ids=[2150]; omni_base.py:380 merge keeps them; min_tokens==1536 is the actual pin.
  • [validated] stats.py:668-669 + definitions.py:137 snapshot keys are str(stage_id) and num_tokens_out
  • [validated] sibling test_qwen3_omni_async_chunk.json and test_qwen3_omni_multi_replicas.json have no sampling_params_list; production defaults unchanged.

The primary leftover is that the new Talker 1536 pin ships beside live historical H100 mean_audio_rtf targets the README itself says not to merge—clear or quarantine that slot (major). Two concrete follow-throughs belong in this PR: add Thinker min_tokens=900 so the new length assert actually covers the 900-token overhead the RTF story assumes (minor), and rerun the L20X five-point sweep with the restored max(2, concurrency) warmups before treating those numbers as evidence for this head (minor). The sibling multi_replicas random row is still variable-length Talker and needs the same extra_body or an explicit out-of-scope note (minor). RTD redness, README relocation, Mix {} persistence, and the tighter failed-assert are dropped as already-gated, polish, or no-issue.

Checked, no defect found:

  • tests/dfx/perf/tests/test_runner_metadata.py:708 — Accept extra_body/extra-body plus the reject snapshot table correctly exercise _assert_fixed_stage_workloads for exact 1536 and missing/wrong lengths.

Verdict: REQUEST CHANGES

Findings

  • [P2] Residual Thinker length is still unpinned: extra_body stage 0 is only max_token… — ``
    Residual Thinker length is still unpinned: extra_body stage 0 is only max_tokens 900 + ignore_eos (no min_tokens) in tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:51-55, and tests/dfx/perf/scripts/run_benchmark.py:272 only asserts stages where min_tokens==max_tokens, so stage 0 is skipped.

Evidence: tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:51-55 {"temperature": 0.0, "max_tokens": 900, "ignore_eos": true} — Thinker/stage 0 has no min_tokens. tests/dfx/perf/scripts/run_benchmark.py:272 and sampling.get("min_tokens") == sampling["max_tokens"] — only those stages enter fixed_lengths, so stage 0 is skipped. tests/dfx/perf/tests/test_runner_metadata.py:715-720 accepts {"1": {"num_tokens_out": 1536}} snapshots with stage 0 as {"max_tokens": 900} only.

  • *[P2] Resolved for order: asyncio.gather(tasks) at patch.py:3469 returns outputs in… — test_patch.py:2106
    Resolved for order: asyncio.gather(*tasks) at patch.py:3469 returns outputs in input order, and test_patch.py:2106-2107 pins that after forcing request 1 to finish before 0. Residual: the stub injects stage_metrics=snapshots[index] (test_patch.py:2055,2069) and never proves a live server num_tokens_out.

Evidence: test_patch.py:2055 snapshots = [{"1": {"num_tokens_out": 1536}}, None, {"1": {"num_tokens_out": 486}}] — injected fixture, not a server response. test_patch.py:2069 stage_metrics=snapshots[index], — monkeypatched request_func writes those values. test_patch.py:2106-2107 assert completion_order.index(1) < completion_order.index(0) / assert result["request_stage_metrics"] == snapshots — pins gather order after out-of-order completion. patch.py:3469 outputs: list[MixRequestFuncOutput] = await asyncio.gather(*tasks) — gather keeps input/task order (unchanged gather contract; this PR only consumes outputs at 3584).

  • [P2] Do not persist request_stage_metrics for empty Mix snapshots — patch.py:3585
    Do not persist request_stage_metrics for empty Mix snapshots. openai-chat-omni inits stage_metrics={} at patch.py:1862 before any SSE merge, and the new persist gate at patch.py:3585 uses any(snapshot is not None), so {} is treated as present. Every openai-chat-omni benchmark() result therefore gains request_stage_metrics as a list of empty dicts, not only the no_async_chunk extra_body path. Plain-vLLM RequestFuncOutput lacks the attribute, so getattr(..., None) omits the key.

Evidence: unchanged by this diff, present in the PR-time tree: vllm_omni/benchmarks/patch/patch.py:1862 output.stage_metrics = {} — Mix chat async_request_openai_chat_omni_completions (ASYNC_REQUEST_FUNCS['openai-chat-omni']) inits empty dict unconditionally. In this diff: vllm_omni/benchmarks/patch/patch.py:3584 request_stage_metrics = [getattr(output, "stage_metrics", None) for output in outputs]; vllm_omni/benchmarks/patch/patch.py:3585 if any(snapshot is not None for snapshot in request_stage_metrics): — {} is not None, so empty Mix inits persist on every openai-chat-omni path. Unchanged dataclass default: vllm_omni/benchmarks/patch/patch.py:883 stage_metrics: dict[str, dict] | None = None — without the chat overwrite, getattr would stay None and omit the key, matching plain-vLLM.

  • [P2] run_benchmark.py:292 not result.get("failed", 0) is inert when the key is mis… — metrics.py:292
    run_benchmark.py:292 not result.get("failed", 0) is inert when the key is missing and redundant on the MultiModals path: metrics.py increments completed only on success and sets failed=len(failed_outputs), while patch.py gathers only formal requests into outputs, so completed==num_prompt already implies failed==0. The non-MM result dict omits failed entirely. _assert_fixed_stage_workloads still returns immediately unless some stage has min_tokens==max_tokens. Require the key ("failed" in result and result["failed"]==0) or drop the conjunct.

Evidence: Changed line: run_benchmark.py:292 assert result["completed"] == num_prompt and not result.get("failed", 0), "Request failures exist" — a missing failed becomes 0, so the new conjunct never fires. Unchanged by this diff, present in the PR-time tree: metrics.py:1022 completed += 1 (only inside if outputs[i].success); metrics.py:1072 failed_outputs = [output for output in outputs if not output.success]; metrics.py:1152 failed=len(failed_outputs),; patch.py:3469 outputs: list[MixRequestFuncOutput] = await asyncio.gather(*tasks) (formal requests only; warmups discarded at patch.py:3320 _ = await asyncio.gather(*warmup_tasks)); patch.py:3521-3522 "completed": mm_metrics.completed, / "failed": mm_metrics.failed, so completed+failed==len(outputs)==num_prompt. Unchanged non-MM branch patch.py:3571-3581 sets "completed": metrics.completed and no failed key. Unchanged run_benchmark.py:274-275 if not fixed_lengths: return.

  • [P2] This PR pins Talker min_tokens==max_tokens==1536 only on tests/dfx/perf/tests/t… — tests/dfx/perf/tests/test_qwen3_omni_multi_replicas.json
    This PR pins Talker min_tokens==max_tokens==1536 only on tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:48-62. The unchanged sibling random row at tests/dfx/perf/tests/test_qwen3_omni_multi_replicas.json:28-36 still uses openai-chat-omni, random_input_len 2500, random_output_len 900, ignore_eos, and no sampling_params_list, so that replica sweep still measures variable-length Talker RTF. tests/dfx/perf/README.md documents only the no_async_chunk random sweep and never mentions multi_replicas. If #7915 incomparability is why 1536 is required, share the same extra_body or document that multi_replicas is intentionally left variable.

Evidence: unchanged by this diff, present in the PR-time tree: tests/dfx/perf/tests/test_qwen3_omni_multi_replicas.json:28-36 "dataset_name": "random", / "backend": "openai-chat-omni", / "random_input_len": 2500, / "random_output_len": 900, / "ignore_eos": true, / "percentile-metrics": "ttft,tpot,itl,e2el,audio_rtf,audio_ttfp,audio_duration", / "baseline": { — no extra_body or sampling_params_list on this random row. Changed by this diff: tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:48-62 "extra_body": { / "return_stage_metrics": true, / "sampling_params_list": [ / { / "temperature": 0.0, / "max_tokens": 900, / "ignore_eos": true / }, / { / "temperature": 0.9, / "top_k": 50, / "repetition_penalty": 1.05, / "seed": 0, / "min_tokens": 1536, / "max_tokens": 1536 — only dfx/perf JSON with min_tokens. Changed by this diff: tests/dfx/perf/README.md:3 The \random` sweep in `tests/test_qwen3_omni_no_async_chunk.json` measures` — README never mentions multi_replicas.

}
]
},
"baseline": {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] This diff pins Talker to min_tokens==max_tokens==1536 on the existing random ro…

This diff pins Talker to min_tokens==max_tokens==1536 on the existing random row (JSON:48-62) while leaving baseline.H100.mean_audio_rtf as the historical variable-length list [0.169, 0.2041, 0.2382, 0.3177, 0.4595] (JSON:66-74). The in-diff README:46 calls those five values old variable-length baselines and says "Do not merge this workload change with those old baselines," yet CUDA/NPU nightly already load this same JSON. Unchanged conftest.py:757-762 always attaches that block to result["baseline"] for each sweep step, so nightly artifacts carry new fixed-length RTF beside the old H100 targets. Clear or quarantine the H100 slot until fixed-workload H100 medians exist. NPU also inherits _assert_fixed_stage_workloads (run_benchmark.py:263) from this extra_body with no A3 snapshot in the file.

Evidence: tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:61-62 "min_tokens": 1536, / "max_tokens": 1536 — Talker fixed to 1536 on the existing random row; JSON:66-73 "baseline": { / "H100": { / "mean_audio_rtf": [ / 0.169, / 0.2041, / 0.2382, / 0.3177, / 0.4595 — historical variable-length H100 list left live. tests/dfx/perf/README.md:46 are available. **Do not merge this workload change with those old baselines.** — PR's own merge gate. Unchanged by this diff, present in the PR-time tree: tests/dfx/conftest.py:757-762 if baseline_config: / # Keep every hardware bucket; resolve list metrics to this sweep step. / result["baseline"] = resolve_baseline_for_sweep( / baseline_config, / sweep_index=sweep_index, / ) — runner always writes that JSON baseline onto the result. Unchanged by this diff, present in the PR-time tree: .buildkite/cuda/test-nightly.yml:81 - pytest -s -v tests/dfx/perf/scripts/run_benchmark.py --test-config-file tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json -m "H100 and B200 and cards_2" and .buildkite/npu/test-npu-nightly.yml:22 pytest -s -v tests/dfx/perf/scripts/run_benchmark.py -m npu --test-config-file tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json — both nightlies load this file. tests/dfx/perf/scripts/run_benchmark.py:274-280 if not fixed_lengths: / return / `` / snapshots = result.get("request_stage_metrics") / ` assert isinstance(snapshots, list) and len(snapshots) == num_prompts, (` / ` "Fixed stage workload: missing per-request stage metrics"` / ` )` — shared length assert fires for this extra_body on NPU too; JSON has no A3 baseline slot.

Suggestion: "baseline": {}

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed this must not merge with the variable-length H100 values. @yenuo26 which do you prefer: (A) run the H100 calibration described in tests/dfx/perf/qwen3_omni_fixed_workload.md and I'll put the medians into this PR, or (B) clear the H100 slot ("baseline": {}) here and recalibrate in a follow-up issue? Nightly doesn't pass --assert-baseline, but the old values would make the dashboard report a false ~15% RTF regression.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in 092a470: per @yenuo26, option B. The random row now has "baseline": {}; the random-mm rows keep their baselines. Fixed-workload H100 medians will be added in a follow-up calibration (procedure in tests/dfx/perf/qwen3_omni_fixed_workload.md).

Comment thread tests/dfx/perf/README.md Outdated
@@ -0,0 +1,83 @@
# Qwen3-Omni fixed-workload performance test

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] This diff adds tests/dfx/perf/README.md as the suite-root README, but title and…

This diff adds tests/dfx/perf/README.md as the suite-root README, but title and body only document the Qwen3-Omni no-async-chunk fixed-length sweep in tests/test_qwen3_omni_no_async_chunk.json. Unchanged sibling tests/dfx/reliability/README.md:1 indexes the whole reliability directory, and the same perf tree also holds Wan, Hunyuan, MiniCPM, Cosmos3 and other JSON configs. Move this note off the suite root (for example next to that JSON) so the perf directory index is not a single-model calibration page.

Evidence: tests/dfx/perf/README.md:1 # Qwen3-Omni fixed-workload performance test; tests/dfx/perf/README.md:3-4 The \random` sweep in `tests/test_qwen3_omni_no_async_chunk.json` measures/2500 configured input tokens, 900 Thinker output tokens, and exactly 1536 Talker— new file, entire page stays on that one JSON and H100 calibration. Unchanged by this diff, present in the PR-time tree: tests/dfx/reliability/README.md:1# L5(b) Reliability (`tests/dfx/reliability`); tests/dfx/reliability/README.md:3 This directory contains reliability fault-injection tests for key models.` Also present and unchanged: tests/dfx/perf/tests/test_wan22_i2v_vllm_omni.json, test_hunyuan_image3_it2i.json, test_minicpmo_4_5.json, test_cosmos3_vllm_omni.json and other sibling perf configs under the same suite root.

Suggestion: # Qwen3-Omni no-async-chunk fixed-workload notes

This page is not the tests/dfx/perf suite index.

"sampling_params_list": [
{
"temperature": 0.0,
"max_tokens": 900,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Thinker sampling_params_list[0] (JSON:53) sets only max_tokens=900 and ignore_e…

Thinker sampling_params_list[0] (JSON:53) sets only max_tokens=900 and ignore_eos. _assert_fixed_stage_workloads keeps a stage only when min_tokens==max_tokens (run_benchmark.py:272), so nightly never checks stage "0". The accept-path unit test uses this same Thinker-only-max_tokens shape and snapshots only stage "1" (test_runner_metadata.py:720). A short Thinker still passes. README calibration requires Thinker 900 and notes ignore_eos does not lock length. Add min_tokens=900 so stage "0" enters fixed_lengths.

Evidence: tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:51-55 "sampling_params_list": [ { "temperature": 0.0, "max_tokens": 900, "ignore_eos": true }, — Thinker has no min_tokens. tests/dfx/perf/scripts/run_benchmark.py:272 and sampling.get("min_tokens") == sampling["max_tokens"] — stage 0 is dropped because get("min_tokens") is None. tests/dfx/perf/tests/test_runner_metadata.py:720 {extra_body_key: {"sampling_params_list": [{"max_tokens": 900}, {"min_tokens": 1536, "max_tokens": 1536}]}} with snapshots only for stage "1" is accepted. tests/dfx/perf/tests/test_runner_metadata.py:756 {"sampling_params_list": [{"max_tokens": 900}]} is treated as variable and does not require snapshots. tests/dfx/perf/README.md:22 ignore_eos alone does not disable explicit stop tokens. tests/dfx/perf/README.md:59 Check every Thinker length (900),

Suggestion: {
"temperature": 0.0,
"max_tokens": 900,
"min_tokens": 900,
"ignore_eos": true
},

@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: finding feedback

[p1] This diff pins Talker to min_tokens==max_tokens==1536 on the existing random ro… — tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:66

This diff pins Talker to min_tokens==max_tokens==1536 on the existing random row (JSON:48-62) while leaving baseline.H100.mean_audio_rtf as the historical variable-length list [0.169, 0.2041, 0.2382, 0.3177, 0.4595] (JSON:66-74). The in-diff README:46 calls those five values old variable-length baselines and says "Do not merge this workload change with those old baselines," yet CUDA/NPU nightly already load this same JSON. Unchanged conftest.py:757-762 always attaches that block to result["baseline"] for each sweep step, so nightly artifacts carry new fixed-length RTF beside the old H100 targets. Clear or quarantine the H100 slot until fixed-workload H100 medians exist. NPU also inherits _assert_fixed_stage_workloads (run_benchmark.py:263) from this extra_body with no A3 snapshot in the file.

Evidence: tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json:61-62 "min_tokens": 1536, / "max_tokens": 1536 — Talker fixed to 1536 on the existing random row; JSON:66-73 "baseline": { / "H100": { / "mean_audio_rtf": [ / 0.169, / 0.2041, / 0.2382, / 0.3177, / 0.4595 — historical variable-length H100 list left live. tests/dfx/perf/README.md:46 are available. **Do not merge this workload change with those old baselines.** — PR's own merge gate. Unchanged by this diff, present in the PR-time tree: tests/dfx/conftest.py:757-762 if baseline_config: / # Keep every hardware bucket; resolve list metrics to this sweep step. / result["baseline"] = resolve_baseline_for_sweep( / baseline_config, / sweep_index=sweep_index, / ) — runner always writes that JSON baseline onto the result. Unchanged by this diff, present in the PR-time tree: .buildkite/cuda/test-nightly.yml:81 - pytest -s -v tests/dfx/perf/scripts/run_benchmark.py --test-config-file tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json -m "H100 and B200 and cards_2" and .buildkite/npu/test-npu-nightly.yml:22 pytest -s -v tests/dfx/perf/scripts/run_benchmark.py -m npu --test-config-file tests/dfx/perf/tests/test_qwen3_omni_no_async_chunk.json — both nightlies load this file. tests/dfx/perf/scripts/run_benchmark.py:274-280 if not fixed_lengths: / return / `` / snapshots = result.get("request_stage_metrics") / ` assert isinstance(snapshots, list) and len(snapshots) == num_prompts, (` / ` "Fixed stage workload: missing per-request stage metrics"` / ` )` — shared length assert fires for this extra_body on NPU too; JSON has no A3 baseline slot.

Suggestion: "baseline": {}

This finding appears in the bot's COMMENT review, but GitHub could not place it as an inline diff comment. If you are the PR author and disagree, react 👎 to this comment. The disagreement will be shown to the maintainer; it does not approve or merge the PR.

@Gaohan123 Gaohan123 added this to the v0.30.0 milestone Sep 23, 2026
@hsliuustc0106 hsliuustc0106 added the CI/CD codes related to changes to CI/CD label Sep 24, 2026
- Pin Thinker min_tokens=900 so the runner also verifies stage 0 length.
- Persist only num_tokens_out/finish_reason/audio frames/duration per stage,
  and skip empty chat-omni snapshots, instead of full per-token snapshots.
- Print a one-line summary after the fixed-length check passes.
- Drop the no-op failed-count conjunct in assert_result.
- Move the calibration notes off the perf suite root and note that
  multi_replicas is out of scope.

Signed-off-by: psv666 <2693925048@qq.com>
@psv666

psv666 commented Sep 24, 2026

Copy link
Copy Markdown
Contributor Author

Self-review / follow-up in d7f0f55:

  • Thinker length is now pinned and checked (min_tokens=900).
  • request_stage_metrics persists only 4 scalar fields per stage and skips empty chat-omni snapshots, so other openai-chat-omni runs don't gain per-token lists or {} entries.
  • Removed the inert failed conjunct; moved the notes to tests/dfx/perf/qwen3_omni_fixed_workload.md; multi_replicas is explicitly out of scope (separate baselines).
  • RTD failure was HTTP 429 fetching external inventories (psutil/pillow/typing-extensions), unrelated.
  • Still pending: H100 baseline decision (above).
  • The L20X sweep predates restoring max(2, c) warmups and was not rerun. Only the number of warmup requests changed; warmups are excluded from all metrics and the formal requests/config are identical, so the impact on the recorded numbers is expected to be negligible.
  • Checked: 163 CPU tests pass; pre-commit incl. mypy 3.10/3.12 passes.

…weep

The historical H100 mean_audio_rtf values were measured with a variable-length
Talker workload and are not comparable with the fixed 900/1536 workload.
Leave the random row without a baseline until fixed-workload H100 medians
are calibrated in a follow-up.

Signed-off-by: psv666 <2693925048@qq.com>
Move the H100 calibration procedure to vllm-project#8100 and keep a one-line rationale
in the perf config, matching other perf configs. No behavior change.

Signed-off-by: psv666 <2693925048@qq.com>
@psv666

psv666 commented Sep 24, 2026

Copy link
Copy Markdown
Contributor Author

Follow-up in beb9650: removed tests/dfx/perf/qwen3_omni_fixed_workload.md (no other perf config has a standalone notes page). The rationale is now a one-line description in the config, and the H100 calibration procedure moved to #8100. Earlier replies that reference the notes file now correspond to #8100.

@yenuo26 yenuo26 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 24, 2026
@yenuo26 yenuo26 added the ready label to trigger buildkite CI label Sep 24, 2026
@yenuo26
yenuo26 enabled auto-merge (squash) September 24, 2026 08:09
…er-perf

Signed-off-by: psv666 <2693925048@qq.com>

# Conflicts:
#	tests/dfx/perf/scripts/run_benchmark.py
auto-merge was automatically disabled September 24, 2026 09:15

Head branch was pushed to by a user without write access

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123 Gaohan123 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 24, 2026
@Gaohan123
Gaohan123 enabled auto-merge (squash) September 24, 2026 14:02
@Gaohan123
Gaohan123 disabled auto-merge September 24, 2026 15:56
@Gaohan123
Gaohan123 merged commit 2d00d86 into vllm-project:main Sep 24, 2026
7 of 9 checks passed
congw729 added a commit to congw729/vllm-omni that referenced this pull request Sep 25, 2026
Merge current main and rebuild the Qwen3-Omni baseline file from its updated configuration. Keep the random workload baseline empty after vllm-project#8057; the five archived A3 RTF values no longer describe that workload. Preserve all remaining 54 A3 values, including the five approved MiniCPM-o Aug 20-25 means.

Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI/CD codes related to changes to CI/CD ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug][Perf]: Qwen3-Omni, performance metrics regressed by more than 10% compared to the baseline in some scenarios

5 participants