Skip to content

[Benchmark] Add local OmniInteract performance cases - #6817

Merged
amy-why-3459 merged 8 commits into
vllm-project:mainfrom
natureofnature:feat/omniinteract-nightly-20260830
Sep 4, 2026
Merged

amy-why-3459 merged 8 commits into
vllm-project:mainfrom
natureofnature:feat/omniinteract-nightly-20260830

Conversation

@natureofnature

@natureofnature natureofnature commented Aug 30, 2026 •

Copy link
Copy Markdown
Collaborator

Purpose

This is the local performance-benchmark split from #5102, following the OmniInteract benchmark runner merged in #6522.

  • Define the three OmniInteract workloads in the standard tests/dfx/perf/tests JSON format.
  • Reuse one MiniCPM-o 4.5 server across 1q1a, 1q1a_math, and 1qna, with four deterministic measured videos per subset (12 total), zero benchmark warmups, and maximum concurrency two.
  • Gate successful input/response lifecycles and complete WAV, transcript, event, manifest, and result artifact publication while reporting official-evaluation eligibility separately.
  • Pin the Hub fallback revision and cover both plain org/repo and org/repo@revision filesystem paths.
  • Register MiniCPM perf JSON files with the standard local Omni launcher and keep this large, cold-download workload out of scheduled Nightly CI.

The generic performance runner gains only the two capabilities needed by these cases: an explicit num_warmups: 0 override and OmniInteract summary/artifact assertions. This PR does not change model/runtime code and does not score answer accuracy.

Test Plan

export HF_HOME=/path/to/persistent/huggingface-cache
export BENCHMARK_DIR=tests/dfx/perf/results
bash tools/nightly/run_nightly_jobs.sh \
  --test-type local \
  --model-type omni \
  --label-substr minicpmo_4_5_omniinteract

pytest -q \
  tests/dfx/perf/tests/test_runner_metadata.py \
  tests/benchmarks/test_omniinteract.py

vLLM Version: 0.28.0

Real-model E2E tested commit: 8b2b6095e1b4af2aa95c6d10243274c4aed6e6a1

Current PR head: e50be2da737c9115f41f7294050145d1120beab1

Test Result

The JSON performance workload completed on one H800, using one shared model server:

  • Perf cases: 3 passed in 2300.13s (38m 20s).
  • Real-time measured requests: 12/12 passed; every subset reported artifacts_complete=true.
  • Current focused unit suite: 67 passed.
  • JSON selector: exactly 3 tests collected.
  • Standard local launcher dry-run selected the Omni run_benchmark.py path and the target JSON.
Subset Requests Official manifest Duration (s) Output tokens Mean TTFT (ms) Mean TPOT (ms) Mean audio RTF
1q1a 4/4 4/4 313.40 913 0.244 14.796 0.841
1q1a_math 4/4 3/4 816.88 1710 0.304 28.574 1.262
1qna 4/4 3/4 769.28 2278 0.233 25.529 0.857
Total 12/12 10/12 1899.56 4901 — — —

The two successful-but-ineligible cases were excluded from official_eval_manifest.jsonl because response audio crossed the fixed video playback horizon (audio_clipped). Their transport, response lifecycle, and required-artifact checks passed; official eligibility is intentionally separate from local benchmark completion.

The real-model run used the existing local OmniInteract archive and was not rerun after the subsequent merge from main. The final PR-only follow-up removes only the redundant local-launcher subprocess test; it does not change the measured workload. On a cache miss, the checked-in configuration downloads the pinned 7.81 GiB archive into HF_HOME; subsequent local runs reuse that cache. Each subset also sends one endpoint-readiness request before its four measured cases.

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR was classified as CI work.

CI owner: @yenuo26

@natureofnature, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@natureofnature
natureofnature force-pushed the feat/omniinteract-nightly-20260830 branch from 33bb1a4 to 80591df Compare August 30, 2026 14:35
Comment thread tests/e2e/online_serving/test_minicpmo_4_5_duplex_expansion.py Outdated
@hsliuustc0106 hsliuustc0106 added the CI/CD codes related to changes to CI/CD label Aug 30, 2026

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — static review at 80591df. The gate's asserted schema keys, artifact names, WAV format, and the repo@rev fallback routing (patch.py:384-392 → HfFileSystem repo@revision) all check out against the writers, and the module-scoped server keeps the model start shared. One P3 open (inline): the reported E2E ran with a local OMNIINTERACT_ROOT, so please report one run of the env-unset fallback path the nightly will actually take — or the archive size / download+extract time — so the first scheduled run is known to fit the 120-min budget.

@natureofnature

Copy link
Copy Markdown
Collaborator Author

Author self-review at e9fd6310:

  • Reviewed the exact three-file PR diff against the current base. Scope remains Nightly wiring, its E2E assertions, and the matching documentation; there are no model, scheduler, connector, or serving-runtime changes.
  • Rechecked the test contract: one module-scoped MiniCPM-o server, 4 deterministic cases for each of 1q1a, 1q1a_math, and 1qna, no warmups, concurrency 2, and explicit input/response/artifact/manifest assertions.
  • Rechecked the dataset path: CI uses the pinned HF revision when OMNIINTERACT_ROOT is absent. The H100 preset provides a persistent node-local HF_HOME; the step now has a 300-minute cold-node budget based on the measured 7.81-GiB archive and conservative download projection recorded in the PR description.
  • Validation after the timeout-only follow-up: source YAML parses, the repository pipeline expander emits timeout_in_minutes: 300 on mithril-h100-pool, git diff --check passes, and the new commit has matching author and Signed-off-by identity.
  • The previously reported focused suite (62 passed) and 12-request E2E results remain applicable because the follow-up changes only the Buildkite timeout.

No additional blocking issue found in self-review.

@natureofnature

Copy link
Copy Markdown
Collaborator Author

@amy-why-3459 PTAL

@amy-why-3459

Copy link
Copy Markdown
Collaborator

Reviewed e9fd631. The Nightly contract itself is right: 12 real-time videos, no accuracy score, audio_clipped / cancelled stay ineligible rather than transport failures, one module-scoped MiniCPM-o, and the fallback is pinned to lucky-lance/OmniInteract@e195f75fe2666fcc5fe74f537ae49ca143a79969. patch.py routes a non-path --dataset-path to dataset_repo; HfFileSystem accepts datasets/{repo@rev}/data.tar.gz; HF reports that revision's data.tar.gz as 8,382,197,760 bytes, matching the Test Result. Function Test already ignores this file. No model/runtime diff.

Request changes on the wiring, not the assertions.

The existing MiniCPM-o Duplex Nightly step is already red on the two interrupt cases (test_duplex_soft_interrupt late playback ACK, test_duplex_server_vad_hard_interrupt response.created timeout). You reproduced both on H800 and on pristine c4192568. That is #6821 / #6716, not this PR. OmniInteract ACKs after wait_for_session_completion, which is why 12/12 can pass on a server that still fails those interrupt probes — they are different contracts and should not share one Buildkite step.

This YAML keeps all of that in one pytest, then raises the step from 50 to 300 minutes. After merge the new gate has no independent green, a warm run still pays the known interrupt failures plus ~30 min of videos, and a cold node can occupy h100_1 for most of the 300-minute budget and still end red. The same happens to any PR labeled omni-test / nightly-test.

Please split. Do not xfail the interrupt tests in this CI PR.

Also:

  • The P3 cold path is still a projection (400 MiB unauthenticated probe → ~195 min), not a completed env-unset download+extract+bench. CI has HF_TOKEN and a node-local /mnt/hf-cache, so it should be faster than that number, but test_hub_archive_uses_vllm_filesystem still only covers the plain org/repo form. After the split, report one real fallback wall-clock or add the @revision filesystem unit.
  • output.wav getnframes() > 0 is almost tautological: the writer allocates a silent 24 kHz buffer of ceil(video_duration) before painting spans, so LISTEN-only still has frames. The docs correctly allow no response audio. Assert format only, or assert audio_bytes / responses if this slice is meant to require speech.
  • Please stop importing find_vllm_cli / run_vllm_bench_subprocess from the Qwen3 accuracy helper; a local shutil.which + subprocess.run(..., check=True) is enough.

@natureofnature
natureofnature force-pushed the feat/omniinteract-nightly-20260830 branch from e9fd631 to 1d9414d Compare September 1, 2026 09:46
@natureofnature

Copy link
Copy Markdown
Collaborator Author

@amy-why-3459 Addressed at 1d9414d9 and rebased onto current main (a65e89cb).

  • Restored the existing Duplex lifecycle step to its original 50-minute command with no OmniInteract environment flag.
  • Added a separate 300-minute OmniInteract step selected by -k "omniinteract", with its own artifact upload.
  • Added a pipeline regression test that locks in this separation.
  • Extended the Hub filesystem test to cover org/repo@revision.
  • Removed the tautological non-empty-frame assertion while retaining WAV format validation.
  • Replaced the import from the Qwen3 accuracy helper with local shutil.which / subprocess.run usage.

Remote-container checks on the final files: Buildkite tests 13 passed; benchmark/CLI tests 63 passed; the dedicated selector collects exactly the three OmniInteract cases (3/6 collected, 3 deselected). The PR description now distinguishes these current-head checks from the earlier H800 12-video E2E evidence and labels the cold-download timing as a projection. PTAL.

Comment thread .buildkite/cuda/test-nightly.yml Outdated
mirror_hardwares: h100_1

- label: ":full_moon: Omni · MiniCPM-o 4.5 OmniInteract Nightly"
timeout_in_minutes: 300

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@yenuo26 PTAL

Comment thread .buildkite/cuda/test-nightly.yml Outdated
- pytest -s -v tests/e2e/online_serving/test_minicpmo_4_5_duplex_expansion.py -m "full_model and cuda and H100 and omni and cards_1" --run-level "full_model"
mirror_hardwares: h100_1

- label: ":full_moon: Omni · MiniCPM-o 4.5 OmniInteract Nightly"

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please don't include Nightly in label and please indicate the test type, such as Function or Perf.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 68b3dab. The step label is now MiniCPM-o 4.5 · OmniInteract Perf Test; Nightly is no longer part of the label.

Comment thread tests/buildkite/test_upload_pipeline.py Outdated
assert "key: upload-weekly-pipeline" in rendered


def test_minicpmo_omniinteract_nightly_isolated_from_duplex() -> None:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think this test is redundant. Maybe it's better to remove it.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Removed the dedicated Buildkite string-matching regression test. The step now uses the standard perf JSON runner and existing pipeline validation.



@hardware_test(res={"cuda": "H100"}, num_cards=1)
@pytest.mark.skipif(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please write the performance test cases in JSON format and place them under tests/dfx/perf/tests

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Addressed in 68b3dab. The three subset cases now live in tests/dfx/perf/tests/test_minicpmo_4_5_omniinteract.json and run through tests/dfx/perf/scripts/run_benchmark.py. The generic runner preserves num_warmups: 0 and asserts the OmniInteract lifecycle/artifact summary.

natureofnature and others added 4 commits September 1, 2026 22:00
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Co-authored-by: Ruirui Yang | Rein <73573651+R2-Y@users.noreply.github.com>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@natureofnature
natureofnature force-pushed the feat/omniinteract-nightly-20260830 branch from 1d9414d to 68b3dab Compare September 2, 2026 03:22
@natureofnature natureofnature changed the title [CI/Build] Add OmniInteract nightly E2E [CI/Build] Add OmniInteract performance coverage Sep 2, 2026
@natureofnature natureofnature changed the title [CI/Build] Add OmniInteract performance coverage [CI/Build] Add OmniInteract nightly E2E Sep 2, 2026

@yenuo26 yenuo26 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@natureofnature

Copy link
Copy Markdown
Collaborator Author

@ZacheryAU PTAL

@natureofnature

Copy link
Copy Markdown
Collaborator Author

@amy-why-3459 PTAL

}


def _resolve_num_warmups(params: dict[str, Any], *, default: int) -> int:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This function is kind of non-negative integer validating helper, which could be centralized to metrics/utils.py later if several modules need it.

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@natureofnature natureofnature changed the title [CI/Build] Add OmniInteract nightly E2E [Benchmark] Add local OmniInteract performance cases Sep 3, 2026
Comment thread tests/tools/test_run_nightly_jobs.py Outdated
@amy-why-3459 amy-why-3459 added the ready label to trigger buildkite CI label Sep 4, 2026
@amy-why-3459
amy-why-3459 merged commit 8d98cc0 into vllm-project:main Sep 4, 2026
8 of 9 checks passed
JoseCarlosGarcia95 pushed a commit to valendra-tech/vllm-omni that referenced this pull request Sep 5, 2026
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Co-authored-by: Ruirui Yang | Rein <73573651+R2-Y@users.noreply.github.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
ZhengWG pushed a commit to ZhengWG/vllm-omni that referenced this pull request Sep 8, 2026
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Co-authored-by: Ruirui Yang | Rein <73573651+R2-Y@users.noreply.github.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Signed-off-by: ZhengWG <zwg0606@gmail.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Co-authored-by: Ruirui Yang | Rein <73573651+R2-Y@users.noreply.github.com>
Co-authored-by: amy-why-3459 <wuhaiyan17@huawei.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI/CD codes related to changes to CI/CD ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants