Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
28 commits
Select commit Hold shift + click to select a range
62b6f7a
[Misc] Add local OmniInteract realtime benchmark
natureofnature Aug 23, 2026
4b956e9
[Benchmark] Align OmniInteract runner contracts
natureofnature Aug 23, 2026
da1ee83
[Benchmark] Harden OmniInteract cache and timeouts
natureofnature Aug 23, 2026
cf3ad26
[Benchmark] Address OmniInteract review feedback
natureofnature Aug 24, 2026
a355d8a
[Benchmark] Align OmniInteract with repository layout
natureofnature Aug 24, 2026
25c9fcb
[Benchmark] Integrate OmniInteract with bench serve
natureofnature Aug 24, 2026
b5ee886
[Docs] Format serving benchmark guide
natureofnature Aug 24, 2026
7b551a1
[Benchmark] Bound OmniInteract media and upload phases
natureofnature Aug 24, 2026
823b5c4
[Benchmark] Default OmniInteract to realtime endpoint
natureofnature Aug 24, 2026
ee803f8
[Benchmark] Preserve playback history and recover stale locks
natureofnature Aug 24, 2026
b22818d
[Benchmark] Validate final decision and session close
natureofnature Aug 24, 2026
429c1d2
[Benchmark] Address OmniInteract runner review
natureofnature Aug 24, 2026
93979b9
[Benchmark] Use tagged Hugging Face helper
natureofnature Aug 24, 2026
60c030d
[Benchmark] Correlate deferred final completion
natureofnature Aug 24, 2026
61894a9
[Benchmark] Align OmniInteract with serving lifecycle
natureofnature Aug 24, 2026
695f327
[Benchmark] Use vLLM Hugging Face filesystem
natureofnature Aug 24, 2026
202962d
[Benchmark] Require model LISTEN completion
natureofnature Aug 24, 2026
a8eb4cb
[Benchmark] Cover unmarked LISTEN events
natureofnature Aug 24, 2026
fc647a1
[Benchmark] Exclude artifact I/O from measurements
natureofnature Aug 24, 2026
e364ac0
[Benchmark] Bound deferred artifact state
natureofnature Aug 24, 2026
b71c41f
[Benchmark] Serialize artifact bundles
natureofnature Aug 24, 2026
247b276
[Benchmark] Predecode OmniInteract media
natureofnature Aug 24, 2026
8475872
[Benchmark] Simplify OmniInteract benchmark
natureofnature Aug 24, 2026
918afeb
[Benchmark] Harden OmniInteract result handling
natureofnature Aug 25, 2026
9854ebe
[Benchmark] Fix OmniInteract metric validity and safe defaults
natureofnature Aug 25, 2026
ab3aeda
Merge branch 'main' into feat/omniinteract-local-benchmark-20260823
Gaohan123 Aug 25, 2026
7713e79
[Benchmark] Avoid invalid Duplex token timeline
natureofnature Aug 25, 2026
c0127ae
Merge branch 'main' into feat/omniinteract-local-benchmark-20260823
Gaohan123 Aug 25, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
3 changes: 2 additions & 1 deletion docs/cli/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,4 +39,5 @@ vllm bench serve --omni \
--num-prompts 5
```

See [vllm bench serve](./bench/serve.md) for the full reference of all available arguments.
See [vllm bench serve](./bench/serve.md) for serving benchmark arguments and
dataset-specific examples, including OmniInteract native-duplex sessions.
44 changes: 41 additions & 3 deletions docs/cli/bench/serve.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,15 +2,15 @@
The vllm bench command launches the vLLM-Omni benchmark to evaluate the performance of multimodal models.

## Notes
vLLM-Omni registers the `openai-chat-omni`, `openai-audio-speech`, `openai-image-edits-omni`, and `daily-omni` serving benchmark backends.
vLLM-Omni registers the `openai-chat-omni`, `openai-audio-speech`, `openai-image-edits-omni`, `daily-omni`, and `openai-realtime-duplex` serving benchmark backends. It also adds the `omniinteract` dataset.

## Basic Parameter Description
You can use `vllm bench serve --omni --help=all` to get descriptions of all parameters. The commonly used parameters are described below:
- `--omni`
Enable Omni (multimodal) mode, supporting multimodal inputs and outputs such as images, videos, and audio.

- `--backend`
Specify the backend adapter. vLLM-Omni adds `openai-chat-omni`, `openai-audio-speech`, `openai-image-edits-omni`, and `daily-omni` to the upstream vLLM backend choices.
Specify the backend adapter. vLLM-Omni adds `openai-chat-omni`, `openai-audio-speech`, `openai-image-edits-omni`, `daily-omni`, and `openai-realtime-duplex` to the upstream vLLM backend choices.

- `--model`
The model identifier to load, filled according to the models supported by vLLM-Omni.
Expand All @@ -19,7 +19,7 @@ You can use `vllm bench serve --omni --help=all` to get descriptions of all para
The API endpoint exposed externally, to which clients send their requests.

- `--dataset-name`
The name of the dataset used; random-mm indicates generating random multimodal inputs (images, videos, audio).
The name of the dataset used; random-mm indicates generating random multimodal inputs (images, videos, audio), while `omniinteract` replays official OmniInteract videos.

- `--num-prompts`
The total number of requests to send, an integer.
Expand Down Expand Up @@ -269,6 +269,44 @@ We use audio generation time / audio duration to calculate RTF.

</details>

### OmniInteract Realtime Benchmark

OmniInteract runs each video as one native-duplex WebSocket sample in the standard serving-benchmark lifecycle:

```bash
vllm bench serve --omni \
--backend openai-realtime-duplex \
--dataset-name omniinteract \
--dataset-path /path/to/OmniInteract \
--model openbmb/MiniCPM-o-4_5 \
--base-url http://127.0.0.1:8000 \
--endpoint /v1/realtime \
--omniinteract-ref-audio /path/to/reference.wav \
--omniinteract-output-dir ./omniinteract-artifacts \
--num-prompts 3 \
--num-warmups 0
```

`--dataset-path` accepts an extracted directory, `data.tar[.gz]`, or a Hugging Face dataset ID; omitting it uses
`lucky-lance/OmniInteract`. `--num-prompts` is the total across subsets and defaults to 3 for OmniInteract; explicit `0`
selects all and oversize values use all available cases. Reference audio is required, and OmniInteract uses the
`/v1/realtime` endpoint.

Audio is replayed as 16 kHz PCM16 in 200 ms chunks and video at 1 FPS with real-time pacing. All selected media is decoded
before timing and remains in client memory for the run, so `--max-concurrency` does not limit media preparation memory; use
explicit `--num-prompts 0` only when the client has enough RAM for the full dataset. Media commands are bounded by
`--omniinteract-media-timeout-s`, and concurrency defaults to 1. Standard request-rate, warmup, result-saving, and summary
options apply. Use `--omniinteract-require-response` only for functional E2E cases; LISTEN is a valid benchmark result.

Each completed case writes `output.wav`, `wav_transcript.json`, `events.json`, `result.json`, and a final `.done` marker under
`--omniinteract-output-dir`. The root also contains `batch_summary.json` and `official_eval_manifest.jsonl`; failed cases write
`.failed.json`. Runs sharing one output root are serialized. Completion validates transport, response lifecycle, and artifacts,
not answer accuracy. Transcript timestamps are serialized playback-queue times. Clipped or cancelled outputs are ineligible and
omitted from the official manifest; `audio_clipped_bytes` records output beyond the rounded video horizon.

TTFT, TTFP, and RTF start at client receipt of `response.created`. TPOT/ITL use engine stage-0 timing; ITL is emitted only when
every token interval is present.

### Multi-Modal Benchmark

<details class="admonition abstract" markdown="1">
Expand Down
40 changes: 12 additions & 28 deletions examples/online_serving/minicpmo/realtime_duplex_demo.py
Original file line number Diff line number Diff line change
@@ -1,3 +1,6 @@
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM-Omni project

"""Minimal single-input MiniCPM-o 4.5 Realtime duplex demo.

Run this after starting the duplex server. Strict lifecycle, overlap, and
Expand All @@ -19,15 +22,22 @@
sys.path.insert(0, str(REPO_ROOT))

from vllm_omni.experimental.fullduplex.client import ( # noqa: E402
PCM16_BYTES_PER_SAMPLE,
PCM16_SAMPLE_RATE,
PCM16_BYTES_PER_SAMPLE, # noqa: F401 - compatibility export
PCM16_SAMPLE_RATE, # noqa: F401 - compatibility export
RealtimeDuplexClient,
RealtimeEventCollector,
build_realtime_url,
read_pcm16_wav,
wait_for,
write_pcm16_wav,
)
from vllm_omni.experimental.fullduplex.client import chunk_period_ms as _chunk_period_ms # noqa: E402
from vllm_omni.experimental.fullduplex.client import ( # noqa: E402
has_residual_model_unit as _has_residual_model_unit,
)
from vllm_omni.experimental.fullduplex.client import ( # noqa: E402
reference_audio_data_url as _ref_audio_data_url,
)


class _StreamingOutputWriter:
Expand Down Expand Up @@ -133,25 +143,6 @@ def _latest_model_decision(
return decision


def _chunk_period_ms(events: list[dict[str, object]]) -> int:
for event in reversed(events):
session = event.get("session")
if not isinstance(session, dict):
continue
capabilities = session.get("capabilities")
if not isinstance(capabilities, dict):
continue
chunk_period_ms = capabilities.get("chunk_period_ms")
if isinstance(chunk_period_ms, int) and chunk_period_ms > 0:
return chunk_period_ms
return 1000


def _has_residual_model_unit(pcm16: bytes, *, chunk_period_ms: int) -> bool:
unit_bytes = PCM16_SAMPLE_RATE * PCM16_BYTES_PER_SAMPLE * chunk_period_ms // 1000
return bool(unit_bytes > 0 and len(pcm16) % unit_bytes)


def _response_in_progress(events: list[dict[str, object]]) -> bool:
return sum(event.get("type") == "response.created" for event in events) > sum(
event.get("type") == "response.done" for event in events
Expand All @@ -168,13 +159,6 @@ def _event_count_after(
return sum(event.get("type") == event_type for event in events[index + 1 :])


def _ref_audio_data_url(path: str | None) -> str | None:
if path is None:
return None
ref_path = Path(path).expanduser()
return "data:audio/wav;base64," + base64.b64encode(ref_path.read_bytes()).decode("ascii")


async def run_demo(args: argparse.Namespace) -> dict[str, object]:
input_pcm16 = read_pcm16_wav(Path(args.input_wav))
if not input_pcm16:
Expand Down
180 changes: 179 additions & 1 deletion tests/benchmarks/metrics/test_metrics.py
Original file line number Diff line number Diff line change
@@ -1,10 +1,12 @@
# SPDX-License-Identifier: Apache-2.0
# SPDX-FileCopyrightText: Copyright contributors to the vLLM project
# SPDX-FileCopyrightText: Copyright contributors to the vLLM-Omni project

"""
Unit tests for metrics.py
"""

import math

import pytest
from vllm.benchmarks.serve import TaskType

Expand Down Expand Up @@ -66,6 +68,22 @@ def _make_output(prompt_len: int, output_tokens: int = 10) -> MixRequestFuncOutp
return output


def _calculate_test_metrics(outputs, goodput=None):
return calculate_metrics(
input_requests=[],
outputs=outputs,
dur_s=1.0,
tokenizer=None,
selected_percentiles=[50.0],
goodput_config_dict=goodput or {},
task_type=TaskType.GENERATION,
selected_percentile_metrics=[],
max_concurrency=None,
request_rate=float("inf"),
benchmark_duration=1.0,
)[0]


# ============================================================================
# total_input Tests
# ============================================================================
Expand Down Expand Up @@ -126,6 +144,166 @@ def test_audio_continuity_aggregation():
assert p99 is not None and p99 > 0.4


def test_unmeasured_duplex_latency_does_not_add_zero_samples():
measured = _make_output(100, output_tokens=2)
measured.ttft = 0.1
measured.audio_ttfp = 0.2
measured.audio_rtf = 0.5
measured.duplex_session_metrics = {
"mean_ttft_ms": 100.0,
"mean_ttfp_ms": 200.0,
"mean_rtf": 0.5,
}
listen_only = _make_output(100, output_tokens=0)
listen_only.ttft = listen_only.audio_ttfp = listen_only.audio_rtf = 0.0
listen_only.duplex_session_metrics = {
"mean_ttft_ms": None,
"mean_ttfp_ms": None,
"mean_rtf": None,
}

metrics = _calculate_test_metrics([measured, listen_only], {"ttft": 150.0})

assert metrics.mean_ttft_ms == 100.0
assert metrics.mean_audio_ttfp_ms == 200.0
assert metrics.mean_audio_rtf == 0.5
assert (metrics.num_ttft_samples, metrics.num_audio_ttfp_samples, metrics.num_audio_rtf_samples) == (1, 1, 1)
assert metrics.request_goodput == 1.0


def test_all_unmeasured_duplex_latency_is_not_reported_as_zero():
listen_only = _make_output(100, output_tokens=0)
listen_only.duplex_session_metrics = {
"mean_ttft_ms": None,
"mean_ttfp_ms": None,
"mean_rtf": None,
}

metrics = _calculate_test_metrics(
[listen_only],
{"ttft": float("inf"), "audio_ttft": float("inf")},
)

assert (metrics.num_ttft_samples, metrics.num_audio_ttfp_samples, metrics.num_audio_rtf_samples) == (0, 0, 0)
assert math.isnan(metrics.mean_ttft_ms)
assert math.isnan(metrics.mean_audio_ttfp_ms)
assert math.isnan(metrics.mean_audio_rtf)
assert metrics.request_goodput == 0.0


def test_unmeasured_duplex_tpot_does_not_add_zero_or_misalign_goodput():
missing_tpot = _make_output(100, output_tokens=5)
missing_tpot.itl = []
missing_tpot.text_latency = missing_tpot.ttft = 0.1
missing_tpot.tpot_measured = False
missing_tpot.duplex_session_metrics = {"mean_ttft_ms": 100.0}

slow_ttft = _make_output(100, output_tokens=5)
slow_ttft.ttft = 1.0
slow_ttft.text_latency = 1.4
slow_ttft.itl = [0.1] * 4
slow_ttft.duplex_session_metrics = {"mean_ttft_ms": 1000.0}

metrics = _calculate_test_metrics(
[missing_tpot, slow_ttft],
{"ttft": 500.0, "tpot": 200.0},
)

assert metrics.num_tpot_samples == 1
assert metrics.mean_tpot_ms == pytest.approx(100.0)
assert metrics.request_goodput == 0.0


def test_all_unmeasured_duplex_token_timing_is_not_reported_as_zero():
output = _make_output(100, output_tokens=5)
output.itl = []
output.text_latency = output.ttft
output.tpot_measured = False
output.duplex_session_metrics = {"mean_ttft_ms": 100.0}

metrics = _calculate_test_metrics([output], {"tpot": float("inf")})

assert (metrics.num_tpot_samples, metrics.num_itl_samples) == (0, 0)
assert math.isnan(metrics.mean_tpot_ms)
assert math.isnan(metrics.mean_itl_ms)
assert metrics.request_goodput == 0.0


def test_duplex_response_timings_do_not_build_a_session_token_timeline():
output = _make_output(100, output_tokens=5)
output.latency = 101.0
output.duplex_request_metrics = [
{"response_id": "r1", "stage0_tokens": {"itls_ms": [100.0, 100.0]}},
{"response_id": "r2", "stage0_tokens": {"itls_ms": [100.0, 100.0]}},
]
output.duplex_session_metrics = {"mean_ttft_ms": 100.0}

metrics = _calculate_test_metrics([output])

assert math.isnan(metrics.max_output_tokens_per_s)
assert metrics.max_concurrent_requests == 1
assert metrics.mean_tpot_ms == pytest.approx(100.0)
assert metrics.mean_itl_ms == pytest.approx(100.0)


def test_unmeasured_tpot_stays_missing_after_tokenizer_fallback():
output = _make_output(100, output_tokens=0)
output.generated_text = "timing metadata missing"
output.itl = []
output.text_latency = output.ttft
output.tpot_measured = False
output.duplex_session_metrics = {"mean_ttft_ms": 100.0}

def tokenizer(text, *, add_special_tokens):
assert text == output.generated_text
assert add_special_tokens is False
return type("Tokenized", (), {"input_ids": [1, 2, 3]})()

metrics, _ = calculate_metrics(
input_requests=[],
outputs=[output],
dur_s=1.0,
tokenizer=tokenizer,
selected_percentiles=[50.0],
goodput_config_dict={"tpot": float("inf")},
task_type=TaskType.GENERATION,
selected_percentile_metrics=[],
max_concurrency=None,
request_rate=float("inf"),
benchmark_duration=1.0,
)

assert metrics.total_output == 3
assert metrics.num_tpot_samples == 0
assert math.isnan(metrics.mean_tpot_ms)
assert metrics.request_goodput == 0.0


def test_measured_zero_itl_is_not_treated_as_missing():
output = _make_output(100, output_tokens=3)
output.itl = [0.0, 0.0]
output.duplex_session_metrics = {"mean_ttft_ms": 100.0}

metrics = _calculate_test_metrics([output], {"tpot": 1.0})

assert (metrics.num_tpot_samples, metrics.num_itl_samples) == (1, 2)
assert metrics.mean_tpot_ms == 0.0
assert metrics.mean_itl_ms == 0.0
assert metrics.request_goodput == 1.0


def test_duplex_goodput_does_not_pair_measurements_from_different_requests():
text_only, audio_only = _make_output(100), _make_output(100)
text_only.ttft, text_only.audio_ttfp = 0.1, 0.0
text_only.duplex_session_metrics = {"mean_ttft_ms": 100.0, "mean_ttfp_ms": None, "mean_rtf": None}
audio_only.ttft, audio_only.audio_ttfp = 0.0, 0.2
audio_only.duplex_session_metrics = {"mean_ttft_ms": None, "mean_ttfp_ms": 200.0, "mean_rtf": None}

metrics = _calculate_test_metrics([text_only, audio_only], {"ttft": 500.0, "audio_ttft": 500.0})

assert metrics.request_goodput == 0.0


# ============================================================================
# TTFT suppression for pure-audio (TTS) benchmarks
# ============================================================================
Expand Down
Loading
Loading