Repository navigation
[Frontend] Add omni benchmark support for image and video endpoints - #4728
Conversation
|
Codex usage limits have been reached for code reviews. Please check with the admins of this repo to increase the limits by adding credits. |
hsliuustc0106
left a comment
There was a problem hiding this comment.
Benchmark extension looks solid.
| return dims[1] | ||
| if len(dims) == 4: | ||
| # Common layouts: [F, H, W, C] and [C, F, H, W]. | ||
| if dims[0] in (1, 3, 4): |
There was a problem hiding this comment.
[F, H, W, C] outputs with 1, 3, or 4 frames enter this branch as if the first dimension were channels; (1, 64, 64, 3) reports 64 frames. Check the trailing channel dimension first and add short channel-last cases, since this corrupts duration, throughput, and RTF metadata.
There was a problem hiding this comment.
fixed and moved this helper to vllm_omni/metrics/ following the idea of #5168, with unit test added in test_metrics_utils.py
| stage_gen_ms = inference_time_s * 1000.0 | ||
| output.video_generation_time_ms = max(output.video_generation_time_ms, stage_gen_ms) | ||
| if output.video_duration > 0 and output.latency > 0: | ||
| output.video_rtf = output.latency / output.video_duration |
There was a problem hiding this comment.
This makes video RTF depend on polling cadence: output.latency includes the /v1/videos/{id} sleep and overshoot (2s by default). Use video_generation_time_ms resolved above when available, with E2E only as a fallback, and cover poll-interval independence.
There was a problem hiding this comment.
changed to video_generation_time_ms now and takes output.latency as fallback. also related tests added
| _apply_image_metrics_from_payload(output, data) | ||
| if output.image_count <= 0: | ||
| output.image_count = int(payload.get("n") or 1) | ||
| output.success = True |
There was a problem hiding this comment.
An HTTP 200 with empty or malformed data reaches this line after fabricating image_count from requested n, so the benchmark reports a successful image that never existed. Require at least one actual image payload before marking success.
|
resolve conflicts |
1373d08 to
c2d0a08
Compare
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
| @@ -739,10 +739,16 @@ def build_stage_metrics( | |||
| def _infer_output_unit_type(self, request_outputs: list[Any], *, token_count: int) -> str: | |||
| final_output_type = getattr(self.stage_client, "final_output_type", None) | |||
|
|
|||
| if self._has_image_output(request_outputs) or final_output_type in {"image", "images"}: | |||
| # Prefer declared modality over payload heuristics: video diffusion often | |||
| # stores frames in ``images`` (see serving_video / output_formatter). | |||
There was a problem hiding this comment.
Video frames may be stored in kind of images slots, and therefore the requests might be considered with "images" output, which misleads the benchmarks. Here is some related fixing.
| @@ -166,7 +189,7 @@ | |||
| NUM_INFERENCE_STEPS = METRIC_PREFIX + "num_inference_steps" | |||
| IMAGE_COUNT_METRIC = METRIC_PREFIX + IMAGE_COUNT | |||
| IMAGE_PIXELS_METRIC = METRIC_PREFIX + IMAGE_PIXELS | |||
| PEAK_MEMORY_MB = METRIC_PREFIX + "peak_memory_mb" | |||
| PEAK_MEMORY_MB_METRIC = METRIC_PREFIX + "peak_memory_mb" | |||
There was a problem hiding this comment.
Avoid conflicts of PEAK_MEMORY_MB consumption, and let Prometheus use PEAK_MEMORY_MB_METRIC instead, following the naming above.
| return | ||
| endpoint = _normalize_endpoint(getattr(args, "endpoint", None)) | ||
| if endpoint in _ENDPOINT_BACKEND_KEYS: | ||
| args.backend = endpoint |
There was a problem hiding this comment.
Changes in this file is to get diffusion related endpoints free from assigning backends
|
do we need to append --endpoint for different output modality? how about qwen3-omni? |
| print_audio_metrics(selected_percentile_metrics, metrics) | ||
| if _has_image_output(metrics): | ||
| print_image_metrics(selected_percentiles or [], metrics) | ||
| if _has_video_output(metrics): |
There was a problem hiding this comment.
[P2] Suppress text metrics for pure image/video endpoints
The new endpoint path still calls print_text_metrics unconditionally before this video branch. Pure image/video responses have no generated text tokens, but the generic metric path still seeds a token timeline and reports text throughput; the PR sample output shows Peak output token throughput: 1.00 and Total Token throughput for image/video requests. Please gate the text section on an actual text modality/backend (and keep token counters at zero), with a pure image/video regression test.
There was a problem hiding this comment.
fixed and added tests
|
[P2] Document the new endpoint-driven benchmark interface This PR adds three user-facing |
qwen3-omni uses in else:
# Qwen omni models (Qwen3-Omni, Qwen2.5-Omni) use a "talker"
# stage whose preprocess requires chat-templated tokens. The
# async-chunk orchestrator prewarms the talker via
# compute_talker_prompt_ids_length(), which scans for Qwen
# chat-template markers (im_start_token_id 151644). A raw-text
# prompt produces a 1-token placeholder that crashes the talker's
# prefill/decode handoff. Reject early with an actionable message.
stage_names = {
getattr(getattr(s, "engine_args", None), "model_stage", None) for s in self.engine_client.stage_configs
}
if "talker" in stage_names:
raise ValueError(
"The /v1/audio/speech endpoint is only supported for "
"dedicated TTS models (e.g., Qwen3-TTS, Voxtral, Fish "
"Speech, CosyVoice3, OmniVoice, VoxCPM2). For omni "
"models like Qwen3-Omni, use /v1/chat/completions with "
'\'"modalities": ["audio"]\' instead.'
)models can support multiple endpoints, so --endpoint is useful for benchmarks for these models, for example minicpmo-4.5 takes |
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
added docs |
| form.add_field("response_format", "b64_json") | ||
| form.add_field("output_format", str(extra_body.get("output_format", "png"))) | ||
| form.add_field("stream", "true") | ||
|
|
There was a problem hiding this comment.
Do not unconditionally enable streaming for image edits
This request always sends stream=true, but the server rejects streaming when the model has only one stage (len(stage_configs) <= 1). As a result, /v1/images/edits benchmarks fail with HTTP 400 for every single-stage image-edit model.
Please make streaming configurable or use the non-streaming JSON response path when streaming is unsupported, and add a regression test covering a single-stage image-edit configuration.
There was a problem hiding this comment.
set stream=false as default and added related tests
| output.latency = time.perf_counter() - st | ||
| if response.status == 200: | ||
| data = await response.json() | ||
| if not isinstance(data, Mapping): |
There was a problem hiding this comment.
Measure E2E latency after consuming the response body
output.latency is recorded before await response.json() consumes and decodes the response body. Image-generation responses can contain large base64 payloads, so this excludes a potentially significant part of network transfer and JSON decoding from the reported E2E latency.
Please record the latency only after the complete body has been read, parsed, and validated, ideally in a common finalization path so successful requests are measured consistently.
There was a problem hiding this comment.
switched the timing of recording output.latency
| if isinstance(value, (dict, list)): | ||
| form.add_field(key, json.dumps(value)) | ||
| else: | ||
| form.add_field(key, str(value)) |
There was a problem hiding this comment.
Avoid forwarding a local reference through two multipart fields
image_reference is forwarded here as a generic extra_body field, while the dedicated video-reference handling later uploads the same local file as input_reference. This can produce a multipart request containing both fields, with the former still holding the raw local path, and the server may reject it during JSON parsing or conflict validation.
Please reserve the reference-related keys here and route them exclusively through the dedicated serializer/uploader. A test with a local video reference would help prevent this regression.
There was a problem hiding this comment.
reserved image_reference and input_reference
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
hsliuustc0106
left a comment
There was a problem hiding this comment.
Two remaining issues at 8de8d12a502b: fractional FPS fails during async video response construction, and structured image references can be silently dropped by the benchmark client.
Validation: static review of the pinned code and callers, syntax checks, and an independent Pydantic reproduction of the fractional-FPS validation error. PR code and GPU tests were not executed.
| description="Filename of the saved output video files for this job.", | ||
| ) | ||
| inference_time_s: float | None = Field(default=None, description="End-to-end inference time in seconds.") | ||
| fps: int | None = Field(default=None, description="Resolved output video frames per second, if known.") |
There was a problem hiding this comment.
[P1] Preserve fractional FPS in video responses
VideoGenerationRequest and _parse_video_form accept fractional FPS, but create_video() passes the resolved value into this integer field through video_response_from_request(). A valid request such as fps=12.5 therefore raises a Pydantic int_from_float validation error before the queued job is created. Please use float | None here and add an async /v1/videos regression case with fractional FPS.
There was a problem hiding this comment.
changed int to float
| if not reference_added: | ||
| image_reference = extra_body.get("image_reference") | ||
| if image_reference is not None: | ||
| _add_video_reference_to_form(form, image_reference) |
There was a problem hiding this comment.
[P2] Preserve structured image references
With --extra-body '{"image_reference":{"image_url":"https://example.com/ref.png"}}', _add_video_extra_body_to_form() excludes the reference from generic forwarding, while _add_video_reference_to_form() accepts only strings or dictionaries containing bytes. It returns False for this supported API reference object, and the return value is ignored here, so the request is sent without its conditioning image. Please serialize supported reference objects (including file_id) through the dedicated path and reject unsupported values instead of silently changing the benchmark workload; cover the object form in a regression test.
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
…llm-project#4728) Signed-off-by: ZhengWG <zwg0606@gmail.com>
…llm-project#4728) Signed-off-by: wenjie.yan <wenjyan@outlook.com>
…ted 4) 上游 vllm-project#4728(b1ec136bd)让 generate_video_bytes 多回一个 video_metadata,返回 5 元组; 上游自己的调用方改走了 _unpack_video_generation_result,但 fork 独有的 _run_video_task_job(GPUStack facade 的 POST /v1/tasks/video/)还按 4 个解包。 9-22 合并上游(4b6c5a733)后,这条路径上每个视频任务都在生成完之后抛 ValueError: GPU 跑满 4~5 分钟,结果丢弃,任务状态 failed,error 就是这句 unpack。 线上复现:minimax-h3-fl2va t2v,最简请求 {model, prompt, metadata.task_type} 也失败, task_CKF7p6fSqmc2yLzun9IRB3vgOLSlJcOC 等。 改为统一走 _unpack_video_generation_result(兼容 4/5 元组);补上这条路径第一个测试, 钉住两种返回形状都能落盘并标成 COMPLETED。顺带 ruff 修了 api_server.py 原有的 import 排序。 (mypy 钩子跳过:api_server.py 里 9 个类型错误均为既有问题,与本改动无关。) Signed-off-by: reputationly <197039020@qq.com>
Purpose
As a further development of #3628:
Summary
This PR extends
vllm bench serve --omnito benchmark OpenAI-compatible image and video generation endpoints more naturally./v1/images/generationsand/v1/videos/v1/images/generations,/v1/images/edits, and/v1/videoscan be benchmarked without passing endpoint-specific backend names./v1/images/generationsand async/v1/videos, including polling video jobs until completion.definitions.py.--print-stageonly prints stage benchmark results when real stage data is available.Test Plan
vLLM Version: 0.28.0
vLLM-Omni Commit: current commit against 5d20f6b
vllm bench serve --omni --endpoint /v1/images/generationswithout--backendvllm bench serve --omni --endpoint /v1/videoswithout--backendpython -m pytest -m 'core_model and cpu' tests/benchmarks/ tests/metrics/ tests/entrypoints/openai_api/test_video_server.py -qTest Result
Qwen/Qwen-Image (
/v1/images/generations)Show more
BestWishYsh/Helios-Distilled (
/v1/videos)Show more
pytest result
490 passed, 18 warnings in 38.86s