Skip to content

[Frontend] Add omni benchmark support for image and video endpoints - #4728

Merged
hsliuustc0106 merged 15 commits into
vllm-project:mainfrom
ZacheryAU:visual_endpoint
Sep 8, 2026
Merged

hsliuustc0106 merged 15 commits into
vllm-project:mainfrom
ZacheryAU:visual_endpoint

Conversation

@ZacheryAU

@ZacheryAU ZacheryAU commented Jun 25, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

As a further development of #3628:

Summary

This PR extends vllm bench serve --omni to benchmark OpenAI-compatible image and video generation endpoints more naturally.

  • Adds endpoint-driven backend inference for /v1/images/generations and /v1/videos
  • /v1/images/generations, /v1/images/edits, and /v1/videos can be benchmarked without passing endpoint-specific backend names.
  • Adds benchmark request handling for /v1/images/generations and async /v1/videos, including polling video jobs until completion.
  • Adds videos metrics like video duration, video frames, video throughput, video RTF, video generation latency, and also peak memory percentiles, with new metrics added in definitions.py.
  • --print-stage only prints stage benchmark results when real stage data is available.
  • Include a bugfix for Helios-Distilled for the experiment

Test Plan

vLLM Version: 0.28.0

vLLM-Omni Commit: current commit against 5d20f6b

  • Tests image-output model with vllm bench serve --omni --endpoint /v1/images/generations without --backend
  • Tests video-output model with vllm bench serve --omni --endpoint /v1/videos without --backend
  • Run python -m pytest -m 'core_model and cpu' tests/benchmarks/ tests/metrics/ tests/entrypoints/openai_api/test_video_server.py -q

Test Result

Qwen/Qwen-Image (/v1/images/generations)

Show more
============ Serving Benchmark Result ============
Successful requests:                     3         
Failed requests:                         0         
Maximum request concurrency:             1         
Benchmark duration (s):                  8.42      
Request throughput (req/s):              0.36      
Peak concurrent requests:                2.00      
-------------------Peak Memory--------------------
Mean PEAK_MEMORY_MB (MB):                58832.00  
Median PEAK_MEMORY_MB (MB):              58832.00  
P99 PEAK_MEMORY_MB (MB):                 58832.00  
----------------End-to-end Latency----------------
Mean E2EL (ms):                          2805.73   
Median E2EL (ms):                        2804.90   
P99 E2EL (ms):                           2810.49   
================== Image Result ==================
Total images generated:                  3         
Image throughput (img/s):                0.36      
Average pixels per image:                1048576.00
Mean denoise step latency (ms):          136.60    
---------------- Image Generation ----------------
Mean IMAGE_GENERATION (ms):              2731.96   
Median IMAGE_GENERATION (ms):            2731.39   
P99 IMAGE_GENERATION (ms):               2734.10   
==================================================

BestWishYsh/Helios-Distilled (/v1/videos)

Show more
tip: install termplotlib and gnuplot to plot the metrics
============ Serving Benchmark Result ============
Successful requests:                     3         
Failed requests:                         0         
Maximum request concurrency:             1         
Benchmark duration (s):                  13.86     
Request throughput (req/s):              0.22      
Peak concurrent requests:                2.00      
-------------------Peak Memory--------------------
Mean PEAK_MEMORY_MB (MB):                46833.33  
Median PEAK_MEMORY_MB (MB):              46790.00  
P99 PEAK_MEMORY_MB (MB):                 46956.60  
----------------End-to-end Latency----------------
Mean E2EL (ms):                          4620.55   
Median E2EL (ms):                        4618.87   
P99 E2EL (ms):                           4628.38   
================== Video Result ==================
Total video duration generated(s):       4.12      
Total video frames generated:            99        
Video throughput(video duration/s):      0.30      
------------------- Video RTF --------------------
Mean VIDEO_RTF:                          3.27      
Median VIDEO_RTF:                        3.28      
P99 VIDEO_RTF:                           3.28      
---------------- Video Generation ----------------
Mean VIDEO_GENERATION (ms):              4502.59   
Median VIDEO_GENERATION (ms):            4505.25   
P99 VIDEO_GENERATION (ms):               4515.02   
==================================================

pytest result

490 passed, 18 warnings in 38.86s

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits.
Credits must be used to enable repository wide code reviews.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Benchmark extension looks solid.

@hsliuustc0106 hsliuustc0106 added diffusion codes related to diffusion models enhancement New feature or request frontend code related to entrypoint omni code related to omni models labels Jul 17, 2026 — with ChatGPT Codex Connector
return dims[1]
if len(dims) == 4:
# Common layouts: [F, H, W, C] and [C, F, H, W].
if dims[0] in (1, 3, 4):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[F, H, W, C] outputs with 1, 3, or 4 frames enter this branch as if the first dimension were channels; (1, 64, 64, 3) reports 64 frames. Check the trailing channel dimension first and add short channel-last cases, since this corrupts duration, throughput, and RTF metadata.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed and moved this helper to vllm_omni/metrics/ following the idea of #5168, with unit test added in test_metrics_utils.py

stage_gen_ms = inference_time_s * 1000.0
output.video_generation_time_ms = max(output.video_generation_time_ms, stage_gen_ms)
if output.video_duration > 0 and output.latency > 0:
output.video_rtf = output.latency / output.video_duration

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This makes video RTF depend on polling cadence: output.latency includes the /v1/videos/{id} sleep and overshoot (2s by default). Use video_generation_time_ms resolved above when available, with E2E only as a fallback, and cover poll-interval independence.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

changed to video_generation_time_ms now and takes output.latency as fallback. also related tests added

Comment thread vllm_omni/benchmarks/patch/patch.py Outdated
_apply_image_metrics_from_payload(output, data)
if output.image_count <= 0:
output.image_count = int(payload.get("n") or 1)
output.success = True

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

An HTTP 200 with empty or malformed data reaches this line after fabricating image_count from requested n, so the benchmark reports a successful image that never existed. Require at least one actual image payload before marking success.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

resolve conflicts

Signed-off-by: ZacheryAU <zachery.au@gmail.com>
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
Signed-off-by: ZacheryAU <zachery.au@gmail.com>

@ZacheryAU ZacheryAU left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Self review

@@ -739,10 +739,16 @@ def build_stage_metrics(
def _infer_output_unit_type(self, request_outputs: list[Any], *, token_count: int) -> str:
final_output_type = getattr(self.stage_client, "final_output_type", None)

if self._has_image_output(request_outputs) or final_output_type in {"image", "images"}:
# Prefer declared modality over payload heuristics: video diffusion often
# stores frames in ``images`` (see serving_video / output_formatter).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Video frames may be stored in kind of images slots, and therefore the requests might be considered with "images" output, which misleads the benchmarks. Here is some related fixing.

@@ -166,7 +189,7 @@
NUM_INFERENCE_STEPS = METRIC_PREFIX + "num_inference_steps"
IMAGE_COUNT_METRIC = METRIC_PREFIX + IMAGE_COUNT
IMAGE_PIXELS_METRIC = METRIC_PREFIX + IMAGE_PIXELS
PEAK_MEMORY_MB = METRIC_PREFIX + "peak_memory_mb"
PEAK_MEMORY_MB_METRIC = METRIC_PREFIX + "peak_memory_mb"

@ZacheryAU ZacheryAU Sep 4, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Avoid conflicts of PEAK_MEMORY_MB consumption, and let Prometheus use PEAK_MEMORY_MB_METRIC instead, following the naming above.

return
endpoint = _normalize_endpoint(getattr(args, "endpoint", None))
if endpoint in _ENDPOINT_BACKEND_KEYS:
args.backend = endpoint

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Changes in this file is to get diffusion related endpoints free from assigning backends

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 4, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

do we need to append --endpoint for different output modality? how about qwen3-omni?

print_audio_metrics(selected_percentile_metrics, metrics)
if _has_image_output(metrics):
print_image_metrics(selected_percentiles or [], metrics)
if _has_video_output(metrics):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Suppress text metrics for pure image/video endpoints

The new endpoint path still calls print_text_metrics unconditionally before this video branch. Pure image/video responses have no generated text tokens, but the generic metric path still seeds a token timeline and reports text throughput; the PR sample output shows Peak output token throughput: 1.00 and Total Token throughput for image/video requests. Please gate the text section on an actual text modality/backend (and keep token counters at zero), with a pure image/video regression test.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

fixed and added tests

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

[P2] Document the new endpoint-driven benchmark interface

This PR adds three user-facing vllm bench serve --endpoint modes (/v1/images/generations, /v1/images/edits, and /v1/videos) plus new image/video/peak-memory outputs, but no documentation is updated. In particular, docs/cli/bench/serve.md still lists only the older Omni backends and percentile metrics, and gives no examples explaining that the endpoint implicitly selects the backend, that image edits need multimodal input, or that /v1/videos is asynchronous and polled. Please update the CLI section, and the image/video API docs for return_stage_metrics and response metrics, before merging.

@ZacheryAU

Copy link
Copy Markdown
Contributor Author

do we need to append --endpoint for different output modality? how about qwen3-omni?

qwen3-omni uses modalities instead of --endpoint to control output modality, and accept /v1/chat/completions instead of /v1/audio/speech

in vllm-omni/vllm_omni/entrypoints/openai/serving_speech.py:

        else:
            # Qwen omni models (Qwen3-Omni, Qwen2.5-Omni) use a "talker"
            # stage whose preprocess requires chat-templated tokens.  The
            # async-chunk orchestrator prewarms the talker via
            # compute_talker_prompt_ids_length(), which scans for Qwen
            # chat-template markers (im_start_token_id 151644).  A raw-text
            # prompt produces a 1-token placeholder that crashes the talker's
            # prefill/decode handoff.  Reject early with an actionable message.
            stage_names = {
                getattr(getattr(s, "engine_args", None), "model_stage", None) for s in self.engine_client.stage_configs
            }
            if "talker" in stage_names:
                raise ValueError(
                    "The /v1/audio/speech endpoint is only supported for "
                    "dedicated TTS models (e.g., Qwen3-TTS, Voxtral, Fish "
                    "Speech, CosyVoice3, OmniVoice, VoxCPM2). For omni "
                    "models like Qwen3-Omni, use /v1/chat/completions with "
                    '\'"modalities": ["audio"]\' instead.'
                )

models can support multiple endpoints, so --endpoint is useful for benchmarks for these models, for example minicpmo-4.5 takes /v1/chat/completions and /v1/realtime.

Signed-off-by: ZacheryAU <zachery.au@gmail.com>
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
@ZacheryAU

Copy link
Copy Markdown
Contributor Author

[P2] Document the new endpoint-driven benchmark interface

This PR adds three user-facing vllm bench serve --endpoint modes (/v1/images/generations, /v1/images/edits, and /v1/videos) plus new image/video/peak-memory outputs, but no documentation is updated. In particular, docs/cli/bench/serve.md still lists only the older Omni backends and percentile metrics, and gives no examples explaining that the endpoint implicitly selects the backend, that image edits need multimodal input, or that /v1/videos is asynchronous and polled. Please update the CLI section, and the image/video API docs for return_stage_metrics and response metrics, before merging.

added docs

form.add_field("response_format", "b64_json")
form.add_field("output_format", str(extra_body.get("output_format", "png")))
form.add_field("stream", "true")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do not unconditionally enable streaming for image edits

This request always sends stream=true, but the server rejects streaming when the model has only one stage (len(stage_configs) <= 1). As a result, /v1/images/edits benchmarks fail with HTTP 400 for every single-stage image-edit model.

Please make streaming configurable or use the non-streaming JSON response path when streaming is unsupported, and add a regression test covering a single-stage image-edit configuration.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

set stream=false as default and added related tests

Comment thread vllm_omni/benchmarks/patch/patch.py Outdated
output.latency = time.perf_counter() - st
if response.status == 200:
data = await response.json()
if not isinstance(data, Mapping):

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Measure E2E latency after consuming the response body

output.latency is recorded before await response.json() consumes and decodes the response body. Image-generation responses can contain large base64 payloads, so this excludes a potentially significant part of network transfer and JSON decoding from the reported E2E latency.

Please record the latency only after the complete body has been read, parsed, and validated, ideally in a common finalization path so successful requests are measured consistently.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

switched the timing of recording output.latency

if isinstance(value, (dict, list)):
form.add_field(key, json.dumps(value))
else:
form.add_field(key, str(value))

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Avoid forwarding a local reference through two multipart fields

image_reference is forwarded here as a generic extra_body field, while the dedicated video-reference handling later uploads the same local file as input_reference. This can produce a multipart request containing both fields, with the former still holding the raw local path, and the server may reject it during JSON parsing or conflict validation.

Please reserve the reference-related keys here and route them exclusively through the dedicated serializer/uploader. A test with a local video reference would help prevent this regression.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

reserved image_reference and input_reference

@ZacheryAU
ZacheryAU requested a review from wtomin as a code owner September 6, 2026 15:40
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
Signed-off-by: ZacheryAU <zachery.au@gmail.com>

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two remaining issues at 8de8d12a502b: fractional FPS fails during async video response construction, and structured image references can be silently dropped by the benchmark client.

Validation: static review of the pinned code and callers, syntax checks, and an independent Pydantic reproduction of the fractional-FPS validation error. PR code and GPU tests were not executed.

description="Filename of the saved output video files for this job.",
)
inference_time_s: float | None = Field(default=None, description="End-to-end inference time in seconds.")
fps: int | None = Field(default=None, description="Resolved output video frames per second, if known.")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P1] Preserve fractional FPS in video responses

VideoGenerationRequest and _parse_video_form accept fractional FPS, but create_video() passes the resolved value into this integer field through video_response_from_request(). A valid request such as fps=12.5 therefore raises a Pydantic int_from_float validation error before the queued job is created. Please use float | None here and add an async /v1/videos regression case with fractional FPS.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

changed int to float

if not reference_added:
image_reference = extra_body.get("image_reference")
if image_reference is not None:
_add_video_reference_to_form(form, image_reference)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Preserve structured image references

With --extra-body '{"image_reference":{"image_url":"https://example.com/ref.png"}}', _add_video_extra_body_to_form() excludes the reference from generic forwarding, while _add_video_reference_to_form() accepts only strings or dictionaries containing bytes. It returns False for this supported API reference object, and the return value is ignored here, so the request is sent without its conditioning image. Please serialize supported reference objects (including file_id) through the dedicated path and reject unsupported values instead of silently changing the benchmark workload; cover the object form in a regression test.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

modified

Signed-off-by: ZacheryAU <zachery.au@gmail.com>
@hsliuustc0106 hsliuustc0106 added the ready label to trigger buildkite CI label Sep 8, 2026
@hsliuustc0106
hsliuustc0106 merged commit b1ec136 into vllm-project:main Sep 8, 2026
8 of 9 checks passed
ZhengWG pushed a commit to ZhengWG/vllm-omni that referenced this pull request Sep 8, 2026
wenj-yan pushed a commit to wenj-yan/vllm-omni that referenced this pull request Sep 9, 2026
reputationly added a commit to reputationly/vllm-omni that referenced this pull request Sep 24, 2026
…ted 4)

上游 vllm-project#4728(b1ec136bd)让 generate_video_bytes 多回一个 video_metadata,返回 5 元组;
上游自己的调用方改走了 _unpack_video_generation_result,但 fork 独有的
_run_video_task_job(GPUStack facade 的 POST /v1/tasks/video/)还按 4 个解包。
9-22 合并上游(4b6c5a733)后,这条路径上每个视频任务都在生成完之后抛 ValueError:
GPU 跑满 4~5 分钟,结果丢弃,任务状态 failed,error 就是这句 unpack。

线上复现:minimax-h3-fl2va t2v,最简请求 {model, prompt, metadata.task_type} 也失败,
task_CKF7p6fSqmc2yLzun9IRB3vgOLSlJcOC 等。

改为统一走 _unpack_video_generation_result(兼容 4/5 元组);补上这条路径第一个测试,
钉住两种返回形状都能落盘并标成 COMPLETED。顺带 ruff 修了 api_server.py 原有的 import 排序。
(mypy 钩子跳过:api_server.py 里 9 个类型错误均为既有问题,与本改动无关。)

Signed-off-by: reputationly <197039020@qq.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion codes related to diffusion models enhancement New feature or request frontend code related to entrypoint high priority high priority issue, needs to be done asap omni code related to omni models ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants