Repository navigation
[CI][Perf] migrate diffusion DFX benches to vllm bench serve --omni - #7737
Conversation
Signed-off-by: wangyu <410167048@qq.com>
…arity - Updated references from `run_diffusion_benchmark.py` to `run_benchmark.py` in various CI configurations and test files to standardize the script used for performance benchmarks. - Adjusted test configurations to reflect changes in dataset naming and parameters, ensuring compatibility with the new benchmark script. - Enhanced documentation to clarify the distinction between diffusion and omni benchmarks, including updates to test examples and execution guides. This refactor aims to streamline the testing process and improve maintainability across the codebase. Signed-off-by: wangyu <410167048@qq.com>
- Introduced `aggregate_stage_durations` and `print_stage_durations_metrics` functions to compute and display mean, p50, and p99 stage durations from request outputs. - Updated `MixRequestFuncOutput` to include `stage_durations` for tracking per-stage timings. - Enhanced test coverage with new tests for aggregating and printing stage durations metrics. - Refactored existing code to utilize `SimpleNamespace` for cleaner tokenization output. This update improves the observability of performance metrics during benchmarking, facilitating better analysis of stage timings. Signed-off-by: [Your Name] <your.email@example.com> Signed-off-by: wangyu <410167048@qq.com>
…uce random input length - Added "tokenizer": "gpt2" to multiple video performance test configurations for consistency. - Reduced "random_input_len" from 64 to 8 across various test cases to optimize input handling. This change enhances the uniformity of tokenizer usage and improves the efficiency of input processing in performance tests. Signed-off-by: wangyu <410167048@qq.com>
- Updated `print_stage_durations_metrics` to format output with clearer labels and units for mean, median, and p99 stage durations. - Introduced a new helper function `_stage_duration_display_name` to prettify stage names for better readability in printed metrics. - Modified test cases to reflect changes in output formatting and ensure accurate assertions. This update improves the clarity and usability of performance metrics during benchmarking, aiding in the analysis of stage timings. Signed-off-by: wangyu <410167048@qq.com>
…arity - Updated references from `run_diffusion_benchmark.py` to `run_benchmark.py` across CI configurations and test files to standardize the script used for performance benchmarks. - Adjusted artifact paths and environment variable names for consistency in performance test configurations. - Enhanced documentation to clarify the distinction between diffusion and omni benchmarks, including updates to test examples and execution guides. This refactor aims to streamline the testing process and improve maintainability across the codebase. Signed-off-by: wangyu <410167048@qq.com>
|
This PR appears to belong to: docs/design/module/benchmarking.md, docs/design/module/diffusion/index.md. Module owners: @david6666666 @Isotr0py @princepride Routing: @david6666666 via module named in the PR description; @Isotr0py via module named in the PR description; @princepride via module named in the PR description @yenuo26, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
ZacheryAU
left a comment
There was a problem hiding this comment.
Some possible deduplicating founding:
- can
print_stage_durations_metricsreuse_print_percentile_metricorprocess_one_metric? - can use
--print-stageto show stage benchmark?
And also a printing redundancy of Stage 0 Gen and stage_gen_time of stage 0 when --print-stage activated, but it depends on whether downstream prefers Stage 0 Gen.
… models - Renamed `run_benchmark.py` to `run_diffusion_benchmark.py` in CI configurations and test files to clearly differentiate between omni and diffusion benchmarks. - Adjusted artifact paths and environment variable names for consistency across performance test configurations. - Enhanced test cases to include `server_type` for diffusion models and updated benchmark parameters for clarity and uniformity. - Improved documentation to reflect changes in test execution and configuration, ensuring better guidance for contributors. These updates aim to streamline the benchmarking process and enhance the maintainability of the codebase. Signed-off-by: wangyu <410167048@qq.com>
Thanks Stage 0 Gen: removed.
|
- Updated `aggregate_stage_durations` to exclude `stage_N_gen_ms` from the metrics, ensuring clarity in the reported timings. - Modified `print_stage_durations_metrics` to utilize metrics stored in `MultiModalsBenchmarkMetrics`, improving the output format and consistency. - Adjusted test cases to validate the exclusion of `stage_N_gen_ms` and ensure accurate assertions in metrics display. These changes improve the accuracy and readability of stage duration metrics during benchmarking, facilitating better performance analysis. Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # tests/benchmarks/patch/test_patch.py # tests/dfx/perf/tests/test_hunyuanvideo15_t2v_vllm_omni.json # vllm_omni/benchmarks/patch/patch.py
- Introduced `_iter_video_reference_inputs` to yield video references from multimodal content. - Updated `_add_video_reference_to_form` to handle structured video references and integrate them into form data. - Enhanced tests to validate the extraction and addition of video references, ensuring compatibility with the existing image reference handling. These changes improve the support for video content in the benchmarking framework, aligning with the existing image reference functionality. Signed-off-by: wangyu <410167048@qq.com>
- Updated `_add_video_reference_to_form` to handle inline video data URLs more effectively, ensuring they are added as binary input references instead of JSON strings. - Adjusted tests to validate the new handling of video references, ensuring that the correct fields are populated and that no legacy fields are present. - Modified benchmark configuration to reflect changes in video dimensions. These improvements enhance the robustness of video content handling in the benchmarking framework. Signed-off-by: wangyu <410167048@qq.com>
Signed-off-by: wangyu <410167048@qq.com>
| if _add_video_reference_to_form(form, reference): | ||
| video_reference_added = True | ||
| break | ||
| if not video_reference_added: |
There was a problem hiding this comment.
[P2] Keep mixed image and inline-video references compatible with the API
When a request contains both an image reference and a data:video/... reference, the preceding image loop adds image_reference, while this loop uploads the video as input_references. _parse_video_form() explicitly rejects that combination with HTTP 400, so these mixed-reference benchmark requests fail before inference.
Please serialize the combined references using a field combination supported by the server, and add a regression test covering an image plus an inline video. This is an additional case related to the existing reference-handling discussion; the HTTP(S) video-type issue is fixed.
There was a problem hiding this comment.
A combined video-and-image mode has been added. The related selection logic is as follows:
| Scenario | Reference Form | Target Field |
|---|---|---|
| Image only | {"image_url"} / data:image / http(s) | image_reference |
| Image only | Local path / {"bytes"} | input_reference |
| Video only | {"video_url"} over http(s) | video_reference |
| Video only | data:video | input_references (upload, to avoid the ~1MB JSON limit) |
| Image + Video | Image + Video (including data:video) | image_reference + video_reference; input_references is no longer used |
1.Since input_references cannot be passed together with image_reference, the only option is to combine them as image_reference + video_reference; however, this combination easily hits the JSON 1MB limit.
2.We considered putting both the image and video into input_references, but this approach requires the model to support mixed_reference_inputs. Otherwise, all reference_inputs would be treated as video files. Currently, only MiniMax supports this.
The current implementation follows option 1.
|
fix conflicts |
- Introduced `_add_combined_video_form_references` to streamline the addition of image and video references in form data, ensuring compatibility with server requirements. - Updated `_add_video_reference_to_form` to reject unsupported `file_id` references, improving error handling. - Added new tests to validate the rejection of `file_id` references and ensure correct processing of combined image and video references. These changes enhance the robustness and clarity of video and image reference handling in the benchmarking framework. Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com> # Conflicts: # vllm_omni/benchmarks/patch/patch.py
fixed |
| if not video_refs and extra_body.get("video_reference") is not None: | ||
| video_refs = [extra_body["video_reference"]] | ||
|
|
||
| upload_inline_video = not image_refs |
There was a problem hiding this comment.
[P2] Normalize all reference forms before choosing combined form fields
The structured image/video case is fixed, but two supported input forms still produce mutually exclusive fields:
- An image URL plus a bare
data:video/...string inextra_body["video_reference"]producesimage_reference+input_references, because the string branch ignoresupload_inline_video=False. - Image upload bytes plus a structured video reference produces
input_reference+video_reference, because the image branch still uses the singular upload field.
_parse_video_form() rejects both combinations with HTTP 400, so these benchmark requests fail before inference. I reproduced both using the current serialization helpers and server validation conditions.
Please normalize the reference forms before selecting the combined encoding, ensure the upload flag also applies to bare video strings, and add regression coverage for both cases.
There was a problem hiding this comment.
I have revised the overall judgment logic as follows:
Only one category of reference
| Judgment | Input | Field |
|---|---|---|
| Image | {"image_url"}, data:image, http(s) with an image extension (.png .jpg .jpeg .webp .gif .bmp .heic .heif) |
image_reference |
| Video | {"video_url"}, http(s) with a video extension (.mp4 .mov .webm .mkv .m4v), data:video with JSON text under 1MB |
video_reference |
| Video, oversized text | Video only, and the JSON text of {"video_url":"data:video..."} is larger than 1MB |
input_references (upload the decoded bytes) |
| File | Local path, {"bytes"} |
input_reference |
| Image and video together | Both sides written as JSON: image_reference + video_reference. In this case, input_reference or input_references is no longer used. |
video+image
| Original form | Before sending |
|---|---|
Image URL, data:image, {"image_url"} |
Put into image_reference as-is |
| Image bytes, local image | Convert to data:image, put into image_reference |
Video URL, data:video, {"video_url"} |
Put into video_reference. Even if data:video exceeds 1MB, do not change it to an upload |
| Video bytes, local video | Convert to data:video, put into video_reference |
update test result
pytest -sv tests/benchmarks/patch/test_patch.py -m "core_model and cpu"
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py \
--test-config-file tests/dfx/perf/tests/test_hunyuanvideo15_i2v_vllm_omni.json \
-m "H100 and B200 and (cards_1 or cards_2)"
============ Serving Benchmark Result ============
Successful requests: 10
Failed requests: 0
Maximum request concurrency: 1
Benchmark duration (s): 80.27
Request throughput (req/s): 0.12
Peak concurrent requests: 2.00
-------------------Peak Memory--------------------
Mean PEAK_MEMORY_MB (MB): 77912.00
Median PEAK_MEMORY_MB (MB): 77912.00
P99 PEAK_MEMORY_MB (MB): 77912.00
---------------- Stage Durations -----------------
Mean Text Encoder Forward (s): 0.07
Median Text Encoder Forward (s): 0.07
P99 Text Encoder Forward (s): 0.08
Mean Vae Decode (s): 3.74
Median Vae Decode (s): 3.69
P99 Vae Decode (s): 4.35
Mean Queue Wait (ms): 0.51
Median Queue Wait (ms): 0.56
P99 Queue Wait (ms): 0.67
----------------End-to-end Latency----------------
Mean E2EL (ms): 8026.27
Median E2EL (ms): 8026.58
P99 E2EL (ms): 8029.93
================== Video Result ==================
Total video duration generated(s): 13.75
Total video frames generated: 330
Video throughput(video duration/s): 0.17
------------------- Video RTF --------------------
Mean VIDEO_RTF: 4.88
Median VIDEO_RTF: 4.84
P99 VIDEO_RTF: 5.39
---------------- Video Generation ----------------
Mean VIDEO_GENERATION (ms): 6709.08
Median VIDEO_GENERATION (ms): 6660.30
P99 VIDEO_GENERATION (ms): 7415.70
==================================================
- Refactored functions to classify and process image and video references, ensuring proper handling of bare URLs and data URIs. - Introduced new tests to validate the correct processing of video references, including oversized data videos and their handling in forms. - Improved error handling for unsupported video formats and added checks for reference types. These changes improve the robustness and clarity of the benchmarking framework's media reference handling. Signed-off-by: wangyu <410167048@qq.com>
|
The CI failure was due to a known issue(#7912) that has now been fixed. Update the code and retry. |
…llm-project#7737) Signed-off-by: wangyu <410167048@qq.com> Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
…ation Merge current main and rebuild Wan2.2 A3 baseline blocks on the new vllm bench serve --omni configuration introduced by vllm-project#7737. Rename throughput and memory metrics; convert the archived latency means from seconds to milliseconds before rounding to four decimals. Retain the selected nightly windows, including MiniCPM-o c8 Sep 18-19. Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
vllm-project#7737 migrated test_hunyuanvideo15_t2v_vllm_omni.json to the omni-bench schema (benchmark_params[].dataset_name) but missed the NPU nightly Diffusion X2V HunyuanVideo-1.5 Perf Test step, which still invoked run_diffusion_benchmark.py. That runner's is_diffusion_perf_config filter skipped both cases as omni-bench, pytest selected 0 tests and the step failed with exit 5 (issue vllm-project#8074). Switch the step to run_benchmark.py with BENCHMARK_DIR and results/*.json artifact upload, matching the already-migrated Wan22 NPU step and the CUDA HunyuanVideo-1.5 step. Add a regression test asserting every Buildkite step runs a perf JSON with the runner matching its schema. Fixes vllm-project#8074
vllm-project#7737 migrated test_hunyuanvideo15_t2v_vllm_omni.json to the omni-bench schema (benchmark_params[].dataset_name) but missed the NPU nightly Diffusion X2V HunyuanVideo-1.5 Perf Test step, which still invoked run_diffusion_benchmark.py. That runner's is_diffusion_perf_config filter skipped both cases as omni-bench, pytest selected 0 tests and the step failed with exit 5 (issue vllm-project#8074). Switch the step to run_benchmark.py with BENCHMARK_DIR and results/*.json artifact upload, matching the already-migrated Wan22 NPU step and the CUDA HunyuanVideo-1.5 step. Add a regression test asserting every Buildkite step runs a perf JSON with the runner matching its schema. Fixes vllm-project#8074 Signed-off-by: zengchuang <zengchuang3@huawei.com>
…llm-project#7737) Signed-off-by: wangyu <410167048@qq.com>
PLEASE FILL IN THE PR DESCRIPTION HERE.
Purpose
Main changes in this PR:
1. Nightly: add Lingbot video perf case
tests/dfx/perf/tests/test_lingbot_video_vllm_omni.jsonand wire it into CUDA nightly (diffusion_lingbot_perfin.buildkite/cuda/test-nightly.yml+ci_source_file_dependencies.yml).2. Migrate DFX diffusion perf cases that use OpenAI image/video APIs to
vllm bench serve --omni(run_benchmark.py)Migrated by serving API:
/v1/videostest_wan22_i2v_vllm_omni.json,test_hunyuanvideo15_t2v_vllm_omni.json,test_hunyuanvideo15_i2v_vllm_omni.json,test_lingbot_video_vllm_omni.json,test_ltx2_vllm_omni.json,test_minimax_h3_vllm_omni.json, Cosmos3 video rows intest_cosmos3_vllm_omni.json/v1/images/generationstest_cosmos3_vllm_omni.jsonBuildkite CUDA/NPU nightly perf steps for the migrated configs now call
run_benchmark.pywithBENCHMARK_DIR(instead ofrun_diffusion_benchmark.py/DIFFUSION_BENCHMARK_DIR).Kept on
run_diffusion_benchmark.py(chat / custom):/v1/chat/completionst2i/i2i: omni chat client forces SSE; diffusion image metrics/E2EL are unreliable vs the non-streaming diffusion clienttest_qwen_image_vllm_omni.json,test_qwen_image_edit_2511_vllm_omni.json,test_qwen_image_layered_vllm_omni.json,test_bagel_vllm_omni.json,test_boogu_image_vllm_omni.json,test_boogu_image_edit_vllm_omni.json,test_hunyuan_image_tp2_cfgp2.json,test_hunyuan_image_tp2_sp2.json,test_hunyuan_image_tp4.json/v1/images/edits+ custom jsonl (dataset-path-inline): omni bench has no equivalent custom loadertest_hunyuan_image3_it2i.jsonQwen-Image CUDA nightly steps remain on
run_diffusion_benchmark.py.3. Schema conversion for migrated (images/videos) configs
When converting to the omni-bench schema:
server_typedataset->dataset_name; removetask(use explicitendpoint)extra_body;enable-negative-prompt->extra_body.negative_promptrandom-mmbuckets where needednum_warmupsthroughput_qps->request_throughput,latency_mean(s) ->mean_e2el_ms,peak_memory_mb_mean->mean_peak_memory_mbtokenizer: gpt2+ shortrandom_input_lenfor Diffusers models without an HF tokenizer4. Omni bench: collect / print / save pipeline
stage_durationsWhen the server returns profiler
stage_durations(e.g. with--enable-diffusion-pipeline-profiler):stage_durationsonMixRequestFuncOutputfrom chat / images / videos payloadsstage_durations_{mean,p50,p99}and aStage Durationssection in the report JSONAlso pass
stage_durationsthroughDiffusionOutputfor HunyuanVideo 1.5 (t2v/i2v) and Cosmos3 so video runs can report Diffuse / VAE timings when the profiler is on.Docs under
docs/contributing/ci/updated for the runner split (dataset-> diffusion client;dataset_name-> omni bench).Test Plan
vLLM Version: (CI nightly image / local env)
vLLM-Omni Commit:
5bf7701baUnit / routing
Migrated perf configs (
run_benchmark.py)Test Result
Local smoke on migrated video paths (omni bench). Chat cases were not kept on omni after finding SSE chat yields empty Image/E2EL metrics (see note below).
UT
H800 vs H100 baseline
Measured on H800 (Cosmos/Wan 2026-09-19; MiniMax T2V/TI2V/DLO 2026-09-19, V2V 2026-09-20), compared with the H100 baselines embedded in the result JSON.
+means H800 is better than the H100 baseline;-means worse. Higher request throughput is better; lower mean E2E latency and peak memory are better.Cosmos3 T2I peak-memory baseline is a round 80,000 MB cap, so +75.5% is not a measured memory reduction. MiniMax V2V reference bucket was lowered to (480, 832, 209) for the 50 MiB limit (output size unchanged). HunyuanVideo15, LingBot, and LTX-2 have no H100 numeric baseline in the result or test config, so they are omitted here.
[Bug][Diffusion][LingBot] LingBotVideoPipeline ignores --enable-diffusion-pipeline-profiler (no Diffuse / stage_durations fine-grained timings) #7741
Pending: full Buildkite nightly for migrated
run_benchmark.pyjobs.BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)