Skip to content

[Benchmark] Add local OmniInteract realtime benchmark - #6522

Merged
Gaohan123 merged 28 commits into
vllm-project:mainfrom
natureofnature:feat/omniinteract-local-benchmark-20260823
Aug 25, 2026
Merged

Gaohan123 merged 28 commits into
vllm-project:mainfrom
natureofnature:feat/omniinteract-local-benchmark-20260823

Conversation

@natureofnature

@natureofnature natureofnature commented Aug 23, 2026 •

Copy link
Copy Markdown
Collaborator

This is the benchmark-runner split from #5102, after the runtime fixes in
#6360 and #6406. It intentionally leaves the 12-video Nightly workflow and
production configuration to a follow-up PR.

Purpose

  • Add OmniInteract as a normal vllm bench serve --omni dataset for the
    official 1q1a, 1q1a_math, and 1qna layouts.
  • Treat each selected video as one native-duplex WebSocket sample while using
    the standard readiness, warmup, request-rate, concurrency, metrics, and
    --save-result lifecycle.
  • Support extracted directories, local archives, and Hugging Face download,
    with deterministic selection, safe atomic extraction, bounded media tools,
    realtime pacing, and bounded concurrency.
  • Validate response identity, final-input completion, playback drain, audio,
    transcript, and .done; write evaluator-ready per-case artifacts plus the
    batch summary and official manifest.

The PR does not change scheduler, connector, model execution, prompt rollover,
or production YAML behavior. It validates transport/lifecycle/artifacts but
does not judge answer accuracy. Nightly orchestration remains a separate
follow-up.

Reviewer follow-up

  • Uses the standard dataset/module/CLI shape: the loader is under
    benchmarks/data_modules, options are registered in cli_args.py, execution
    stays in vllm bench serve, and usage is documented with other serving
    datasets rather than in a standalone command or page.
  • The standard CLI prepares selected media and reference audio before benchmark
    timing. OmniInteract defaults to three prompts when --num-prompts is
    omitted; explicit 0 still selects all. The documentation states that all
    selected media remains resident in client memory and that concurrency does
    not bound preparation memory.
  • _Playback is the single serialized timeline for response ACK and
    WAV/transcript generation. Completion requires stable response identity and
    the accepted final-input decision; invalid audio metadata and incomplete
    response pairs fail rather than being silently accepted.
  • Deferred results retain sparse clipped PCM spans and sanitized metadata, not
    collectors, base64 audio, or a dense video-length buffer. WAV materialization
    is one case at a time.
  • Media duration is rejected before decode at a configurable safety limit;
    ffmpeg duration/output and the final Python PCM buffer are independently
    bounded.
  • Per-file writes use destination-local unique temporary files and publish
    .done last. The standard CLI holds one cross-process output-root lock for
    the complete run, so case bundles and their aggregate manifest cannot be
    interleaved by another process. The public case runner only captures artifact
    context; the locked batch finalizer is the sole publication owner.
  • Clipped or cancelled successful cases remain auditable but are marked
    ineligible and omitted from official_eval_manifest.jsonl. Artifact
    publication failure revokes artifact/case success while preserving completed
    serving metrics; recovery failures are logged and reported in the batch
    summary instead of aborting metric finalization.
  • WebSocket headers follow the standard vLLM precedence: OPENAI_API_KEY, then
    explicit --header, then x-request-id. OmniInteract requires an explicit
    --endpoint /v1/realtime; an incompatible endpoint fails fast.
  • TTFT/TTFP/RTF use client receipt of response.created as their origin.
    LISTEN-only sessions and missing stage-0 timing omit unavailable samples
    instead of contributing false zeroes. Per-request missing TPOT remains aligned
    with goodput inputs, measured zero latency remains valid, and exact ITL is
    reported only when every response has complete stage-0 intervals.
  • Configure and close waits surface a stopped WebSocket reader immediately
    while still accepting a target event already received by the collector.
  • Hugging Face access uses vLLM's tagged hf_fs(). Archive extraction validates
    members, publishes a fingerprinted tree atomically, and handles concurrent
    publishers.

Usage

vllm bench serve --omni \
  --backend openai-realtime-duplex \
  --dataset-name omniinteract \
  --dataset-path /path/to/OmniInteract \
  --model openbmb/MiniCPM-o-4_5 \
  --base-url http://127.0.0.1:8000 \
  --endpoint /v1/realtime \
  --omniinteract-ref-audio /path/to/reference.wav \
  --omniinteract-output-dir ./omniinteract-artifacts \
  --num-prompts 3 \
  --max-concurrency 2 \
  --num-warmups 0 \
  --save-result

--endpoint /v1/realtime is required because the upstream serving-benchmark
default targets the completions endpoint.

Test Plan

pytest -q -o addopts= tests/benchmarks/test_omniinteract.py
pytest -q -o addopts= tests/benchmarks

pytest -q -o addopts= \
  tests/e2e/features/fullduplex/test_client.py \
  tests/e2e/online_serving/test_minicpmo_realtime_duplex_drivers.py \
  tests/examples/test_minicpmo_realtime_duplex_simple_demo.py

pytest --collect-only -q -o addopts= \
  tests/e2e/online_serving/test_minicpmo_4_5_duplex.py

E2E Benchmark Result

Current head 7713e794, MiniCPM-o 4.5, one deterministic case per subset,
--max-concurrency 1 --num-warmups 0:

Subset Requests E2EL (ms) TTFT (ms) TPOT (ms) Output tok/s Audio TTFP (ms) Audio RTF Generated audio (s) Official manifest
1q1a 1/1 156150.04 0.29 24.28 1.38 0.19 0.75 45.56 included
1q1a_math 1/1 304540.48 0.24 25.87 0.37 0.19 0.85 34.28 included
1qna 1/1 444438.23 0.26 53.72 3.49 0.18 0.85 297.60 excluded (audio_clipped=1)

All three runs completed with Successful requests: 1, Failed requests: 0,
and a per-case .done artifact. The 1qna row is reported for lifecycle and
performance evidence only; its clipped playback horizon correctly kept it out
of official_eval_manifest.jsonl. Peak output-token throughput is reported as
unavailable for these Duplex sessions because response-local timings cannot
reconstruct a valid session-global token timeline.

Regression tests on the same code: focused OmniInteract/metrics selection
52 passed; full tests/benchmarks suite 179 passed.

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/benchmarking.md.

Module owners: @alex-jw-brooks @Bounty-hunter

@natureofnature, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@natureofnature natureofnature changed the title [Misc] Add local OmniInteract realtime benchmark [Benchmark] Add local OmniInteract realtime benchmark Aug 23, 2026
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@natureofnature

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 4b956e9db6

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vllm_omni/benchmarks/data_modules/omniinteract.py Outdated
Comment thread vllm_omni/benchmarks/omniinteract.py Outdated
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@natureofnature

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. 🎉

Reviewed commit: da1ee83806

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@natureofnature

Copy link
Copy Markdown
Collaborator Author

@ZacheryAU PTAL

@natureofnature

Copy link
Copy Markdown
Collaborator Author

@amy-why-3459 @R2-Y PTAL

@amy-why-3459

Copy link
Copy Markdown
Collaborator

Thanks @natureofnature — the split from #5102 is the right shape, and
the real-model smokes (1q1a / 1q1a_math / 1qna) plus the identity
ledger look good. A few things before this lands.

  1. output.wav still overwrites overlapping responses. [Core][Benchmark]Omniinteract for Minicpm-o4.5 #5102 already
    flagged this. _build_output paints every response.audio.delta
    onto a ceil(duration) timeline; a later turn that starts before
    the previous one ends silently replaces those PCM bytes. Horizon
    clipping is documented and tested; overlap is not. The 1qna smoke
    emitted 20 responses in 300s — that path will collide. Please count
    overwritten / clipped bytes in result.json, and keep overlap-heavy
    cases out of official_eval_manifest.jsonl (or write one WAV per
    response_id). Official ASR/align will otherwise score a mixed
    timeline.

  2. --num-prompts is applied after subsets are concatenated.
    --subsets 1q1a 1q1a_math 1qna --num-prompts 3 looks like one
    video per layout; it is three videos total after a global shuffle.
    The PR body already calls this out. Either take N per subset, or
    make the CLI help / README say “total across selected subsets”.

  3. Extract lock will fail on shared filesystems. _extract_archive
    uses bare fcntl.flock. hub_prefetch.py already falls back when
    flock returns ENOLCK (FSx / Lustre / some NFS). Nightly jobs
    sharing HF_HOME are exactly that setup — please reuse that
    helper. The HF download itself is still outside the lock.

Nits: this is named vllm bench but OmniInteractCaseResult has no
TTFT / TTFP / RTF — RealtimeEventCollector.timing_summary() already
computes them; worth folding into result.json so #5069 / duplex
perf can use the same runner. The shared-client changes
(raise_if_reader_stopped, ref_audio / idle_timeout_s, fresh
session.closed) look scoped and correct.

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@natureofnature

Copy link
Copy Markdown
Collaborator Author

Correction to my previous repository-layout reply: the dedicated-subcommand interpretation is superseded.

Updated in 25c9fcba:

  • OmniInteract now uses the same entry shape as other serving datasets: vllm bench serve --omni --dataset-name omniinteract.
  • Common options are reused from bench serve (--dataset-path, --model, URL/endpoint, prompt count, concurrency, result directory, seed/shuffle); only OmniInteract-specific options live in cli_args.py.
  • The standalone subcommand was removed. Internally, only dataset_name == "omniinteract" dispatches to the existing long-lived native-duplex WebSocket runner; other datasets keep the upstream serving path.
  • The standalone OmniInteract document was removed and its dataset-specific section was folded into docs/cli/bench/serve.md.
  • Parser/config/dispatch coverage was added. Remote verification: focused suite 69 passed; full tests/benchmarks 182 passed; integrated CLI help, Ruff/format, compileall, and markdownlint passed.

b5ee886c contains only the mechanical markdownlint cleanup required after bringing the existing serving guide into this PR's touched files. The PR description and test commands now use the integrated CLI.

@natureofnature

Copy link
Copy Markdown
Collaborator Author

@codex review

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 247b2765d0

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment thread vllm_omni/benchmarks/omniinteract.py Outdated

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a9cf06df94

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +539 to +540
_clear_artifacts(directory, (*SUCCESS_ARTIFACTS, ".failed.json"))
_atomic_write_wav(directory / "output.wav", pcm, OUTPUT_SAMPLE_RATE)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep public artifact publication under the output lock

When two processes call the retained public run_omniinteract_case() for the same case/output root, or one does so while vllm bench serve holds its lease, these independent file replacements can still interleave: the public path never acquires omniinteract_output_lock, so one writer can remove another writer's files and expose a .done marker with a mixed bundle. Fresh evidence beyond the earlier artifact-lock thread is that a9cf06df removed the per-case locking from these helpers while adding the lease only around benchmarks.serve.main(); make the public mutation path participate in the same output-root lock.

Useful? React with 👍 / 👎.

@natureofnature
natureofnature force-pushed the feat/omniinteract-local-benchmark-20260823 branch 2 times, most recently from cdfba3e to e9bc9e6 Compare August 25, 2026 01:37
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@natureofnature
natureofnature force-pushed the feat/omniinteract-local-benchmark-20260823 branch from e9bc9e6 to 8475872 Compare August 25, 2026 01:44
@Gaohan123 Gaohan123 added this to the v0.28.0 milestone Aug 25, 2026
output.audio_duration = case_result.audio_bytes / (24_000 * 2)
output.audio_frames = case_result.audio_bytes // 2
session_metrics = case_result.duplex_session_metrics
output.ttft = float(session_metrics.get("mean_ttft_ms") or 0.0) / 1000.0

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Do not use response-local timing as a session-global token timeline

output.start_time marks the beginning of the entire WebSocket/video session, while output.ttft is the mean delay from each response.created event and output.itl flattens intervals from multiple responses. Downstream metrics reconstruct absolute token timestamps as output.start_time + output.ttft + cumulative_itl.

For example, if a response is created 100 seconds into a video, its tokens are incorrectly placed near the beginning of the session. Flattening multiple responses also omits the first token and inter-response gap for every response after the first. This can substantially inflate or otherwise distort max_output_tokens_per_s and the token timeline plot.

Please retain each response's response.created offset relative to stream_start and construct the timeline per response. If exact absolute timing is unavailable, the OmniInteract backend should report max_output_tokens_per_s as unavailable instead of feeding response-local timing into the session-global metric.

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@hsliuustc0106 hsliuustc0106 added benchmark/profiler/metrics/logger Codes related to benchmarks, profiler, metrics and logger system documentation Improvements or additions to documentation labels Aug 25, 2026
@Gaohan123 Gaohan123 added the ready label to trigger buildkite CI label Aug 25, 2026

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please add some e2e benchmark results in the PR description. Thanks

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
@natureofnature

Copy link
Copy Markdown
Collaborator Author

@Gaohan123 Added current-head MiniCPM-o 4.5 E2E benchmark metrics for one 1q1a, 1q1a_math, and 1qna case to the PR description. All three completed 1/1 successfully; the table also calls out that the 1qna artifact was excluded from the official manifest because audio_clipped=1.

@Gaohan123 Gaohan123 removed the ready label to trigger buildkite CI label Aug 25, 2026

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123 Gaohan123 added the ready label to trigger buildkite CI label Aug 25, 2026
@Gaohan123
Gaohan123 enabled auto-merge (squash) August 25, 2026 14:52
@Gaohan123
Gaohan123 merged commit 497c537 into vllm-project:main Aug 25, 2026
8 of 9 checks passed
AndyZhou952 pushed a commit to AndyZhou952/vllm-omni that referenced this pull request Aug 26, 2026
)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Signed-off-by: AndyZhou952 <jzhoubc@connect.ust.hk>
JoseCarlosGarcia95 pushed a commit to valendra-tech/vllm-omni that referenced this pull request Sep 5, 2026
)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
)

Signed-off-by: natureofnature <wzliu@connect.hku.hk>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

benchmark/profiler/metrics/logger Codes related to benchmarks, profiler, metrics and logger system documentation Improvements or additions to documentation ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants