Skip to content

[CI][Perf] migrate diffusion DFX benches to vllm bench serve --omni - #7737

Merged
Gaohan123 merged 23 commits into
vllm-project:mainfrom
yenuo26:perf
Sep 22, 2026
Merged

Gaohan123 merged 23 commits into
vllm-project:mainfrom
yenuo26:perf

Conversation

@yenuo26

@yenuo26 yenuo26 commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

PLEASE FILL IN THE PR DESCRIPTION HERE.

Purpose

Main changes in this PR:

1. Nightly: add Lingbot video perf case

  • Add tests/dfx/perf/tests/test_lingbot_video_vllm_omni.json and wire it into CUDA nightly (diffusion_lingbot_perf in .buildkite/cuda/test-nightly.yml + ci_source_file_dependencies.yml).

2. Migrate DFX diffusion perf cases that use OpenAI image/video APIs to vllm bench serve --omni (run_benchmark.py)

Migrated by serving API:

Endpoint Migrated configs
/v1/videos test_wan22_i2v_vllm_omni.json, test_hunyuanvideo15_t2v_vllm_omni.json, test_hunyuanvideo15_i2v_vllm_omni.json, test_lingbot_video_vllm_omni.json, test_ltx2_vllm_omni.json, test_minimax_h3_vllm_omni.json, Cosmos3 video rows in test_cosmos3_vllm_omni.json
/v1/images/generations Cosmos3 t2i row in test_cosmos3_vllm_omni.json

Buildkite CUDA/NPU nightly perf steps for the migrated configs now call run_benchmark.py with BENCHMARK_DIR (instead of run_diffusion_benchmark.py / DIFFUSION_BENCHMARK_DIR).

Kept on run_diffusion_benchmark.py (chat / custom):

Reason Configs
/v1/chat/completions t2i/i2i: omni chat client forces SSE; diffusion image metrics/E2EL are unreliable vs the non-streaming diffusion client test_qwen_image_vllm_omni.json, test_qwen_image_edit_2511_vllm_omni.json, test_qwen_image_layered_vllm_omni.json, test_bagel_vllm_omni.json, test_boogu_image_vllm_omni.json, test_boogu_image_edit_vllm_omni.json, test_hunyuan_image_tp2_cfgp2.json, test_hunyuan_image_tp2_sp2.json, test_hunyuan_image_tp4.json
/v1/images/edits + custom jsonl (dataset-path-inline): omni bench has no equivalent custom loader test_hunyuan_image3_it2i.json

Qwen-Image CUDA nightly steps remain on run_diffusion_benchmark.py.

3. Schema conversion for migrated (images/videos) configs

When converting to the omni-bench schema:

  • Case-level: drop server_type
  • Client routing: dataset -> dataset_name; remove task (use explicit endpoint)
  • Generation knobs -> extra_body; enable-negative-prompt -> extra_body.negative_prompt
  • Image/video inputs -> random-mm buckets where needed
  • Kebab-case diffusion keys -> snake_case omni keys; warmup triad -> best-effort num_warmups
  • Baseline rename: throughput_qps -> request_throughput, latency_mean (s) -> mean_e2el_ms, peak_memory_mb_mean -> mean_peak_memory_mb
  • tokenizer: gpt2 + short random_input_len for Diffusers models without an HF tokenizer

4. Omni bench: collect / print / save pipeline stage_durations

When the server returns profiler stage_durations (e.g. with --enable-diffusion-pipeline-profiler):

  • Collect: per-request stage_durations on MixRequestFuncOutput from chat / images / videos payloads
  • Aggregate / print / save: stage_durations_{mean,p50,p99} and a Stage Durations section in the report JSON

Also pass stage_durations through DiffusionOutput for HunyuanVideo 1.5 (t2v/i2v) and Cosmos3 so video runs can report Diffuse / VAE timings when the profiler is on.

Docs under docs/contributing/ci/ updated for the runner split (dataset -> diffusion client; dataset_name -> omni bench).

Test Plan

vLLM Version: (CI nightly image / local env)

vLLM-Omni Commit: 5bf7701ba

Unit / routing

pytest -q tests/benchmarks/metrics/test_metrics.py \
  tests/benchmarks/patch/test_patch.py \
  tests/dfx/perf/tests/test_runner_metadata.py

Migrated perf configs (run_benchmark.py)

export BENCHMARK_DIR=tests/dfx/perf/results

# /v1/videos
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py \
  --test-config-file tests/dfx/perf/tests/test_wan22_i2v_vllm_omni.json \
  -m "H100 and B200 and cards_1"
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py \
  --test-config-file tests/dfx/perf/tests/test_wan22_i2v_vllm_omni.json \
  -m "H100 and B200 and cards_2"
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py \
  --test-config-file tests/dfx/perf/tests/test_hunyuanvideo15_t2v_vllm_omni.json \
  -m "H100 and B200 and (cards_1 or cards_2)"
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py \
  --test-config-file tests/dfx/perf/tests/test_hunyuanvideo15_i2v_vllm_omni.json \
  -m "H100 and B200 and (cards_1 or cards_2)"
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py \
  --test-config-file tests/dfx/perf/tests/test_lingbot_video_vllm_omni.json \
  -m "H100 and B200 and cards_1"
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py \
  --test-config-file tests/dfx/perf/tests/test_ltx2_vllm_omni.json \
  -m "H100 and diffusion"
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py \
  --test-config-file tests/dfx/perf/tests/test_minimax_h3_vllm_omni.json \
  -m "H100 and B200 and diffusion"
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py \
  --test-config-file tests/dfx/perf/tests/test_cosmos3_vllm_omni.json \
  -m "H100 and B200 and cards_2"

Test Result

Local smoke on migrated video paths (omni bench). Chat cases were not kept on omni after finding SSE chat yields empty Image/E2EL metrics (see note below).

UT

image

H800 vs H100 baseline

Measured on H800 (Cosmos/Wan 2026-09-19; MiniMax T2V/TI2V/DLO 2026-09-19, V2V 2026-09-20), compared with the H100 baselines embedded in the result JSON. + means H800 is better than the H100 baseline; - means worse. Higher request throughput is better; lower mean E2E latency and peak memory are better.

Case Metric H800 H100 baseline Gap
Cosmos3 T2I 2GPU request throughput (req/s) 1.620 1.196 +35.4%
mean E2E (ms) 615 843 +27.0%
mean peak memory (MB) 19,568 80,000 +75.5%
Cosmos3 T2V 2GPU request throughput (req/s) 0.0354 0.0307 +15.3%
mean E2E (ms) 28,239 32,767 +13.8%
mean peak memory (MB) 27,768 28,762 +3.5%
Cosmos3 I2V 2GPU request throughput (req/s) 0.0354 0.0305 +16.0%
mean E2E (ms) 28,262 33,101 +14.6%
mean peak memory (MB) 27,770 29,441 +5.7%
Cosmos3 V2V 2GPU request throughput (req/s) 0.0351 0.0299 +17.3%
mean E2E (ms) 28,500 33,925 +16.0%
mean peak memory (MB) 27,768 29,443 +5.7%
Wan2.2 I2V single-device 832x480 request throughput (req/s) 0.0383 0.0361 +6.2%
mean E2E (ms) 26,090 27,621 +5.5%
mean peak memory (MB) 76,173 80,548 +5.4%
Wan2.2 I2V USP2 832x480 request throughput (req/s) 0.0525 0.0526 -0.3%
mean E2E (ms) 19,060 18,942 -0.6%
mean peak memory (MB) 47,952 50,054 +4.2%
Wan2.2 I2V USP2 1280x720 request throughput (req/s) 0.0102 0.0096 +5.9%
mean E2E (ms) 98,441 104,106 +5.4%
mean peak memory (MB) 55,623 59,718 +6.9%
MiniMax-H3 T2V 4GPU request throughput (req/s) 0.0329 0.0211 +55.8%
mean E2E (ms) 30,416 38,325 +20.6%
mean peak memory (MB) 53,462 54,220 +1.4%
MiniMax-H3 T2V DLO 4GPU request throughput (req/s) 0.0333 0.0212 +56.9%
mean E2E (ms) 30,070 38,209 +21.3%
mean peak memory (MB) 35,668 36,218 +1.5%
MiniMax-H3 TI2V 4GPU request throughput (req/s) 0.0324 0.0212 +53.0%
mean E2E (ms) 30,825 38,138 +19.2%
mean peak memory (MB) 53,838 54,623 +1.4%
MiniMax-H3 V2V 4GPU request throughput (req/s) 0.0110 0.0077 +43.0%
mean E2E (ms) 90,766 105,751 +14.2%
mean peak memory (MB) 57,486 55,181 -4.2%

Cosmos3 T2I peak-memory baseline is a round 80,000 MB cap, so +75.5% is not a measured memory reduction. MiniMax V2V reference bucket was lowered to (480, 832, 209) for the 50 MiB limit (output size unchanged). HunyuanVideo15, LingBot, and LTX-2 have no H100 numeric baseline in the result or test config, so they are omitted here.

  • Wan
============ Serving Benchmark Result ============
Successful requests:                     10
Failed requests:                         0
Maximum request concurrency:             1
Benchmark duration (s):                  260.91
Request throughput (req/s):              0.04
Peak concurrent requests:                2.00
-------------------Peak Memory--------------------
Mean PEAK_MEMORY_MB (MB):                76173.40
Median PEAK_MEMORY_MB (MB):              76088.00
P99 PEAK_MEMORY_MB (MB):                 76604.00
---------------- Stage Durations -----------------
Mean Text Encoder Forward (s):           0.07
Median Text Encoder Forward (s):         0.07
P99 Text Encoder Forward (s):            0.09
Mean Vae Encode (s):                     1.38
Median Vae Encode (s):                   1.38
P99 Vae Encode (s):                      1.38
Mean Diffuse (s):                        20.35
Median Diffuse (s):                      20.35
P99 Diffuse (s):                         20.39
Mean Vae Decode (s):                     2.40
Median Vae Decode (s):                   2.35
P99 Vae Decode (s):                      2.66
Mean Queue Wait (ms):                    0.38
Median Queue Wait (ms):                  0.36
P99 Queue Wait (ms):                     0.56
----------------End-to-end Latency----------------
Mean E2EL (ms):                          26089.74
Median E2EL (ms):                        26082.83
P99 E2EL (ms):                           26113.36
================== Video Result ==================
Total video duration generated(s):       50.62
Total video frames generated:            810
Video throughput(video duration/s):      0.19
------------------- Video RTF --------------------
Mean VIDEO_RTF:                          4.88
Median VIDEO_RTF:                        4.86
P99 VIDEO_RTF:                           5.03
---------------- Video Generation ----------------
Mean VIDEO_GENERATION (ms):              24699.17
Median VIDEO_GENERATION (ms):            24596.52
P99 VIDEO_GENERATION (ms):               25449.21
==================================================
============ Serving Benchmark Result ============
Successful requests:                     10
Failed requests:                         0
Maximum request concurrency:             1
Benchmark duration (s):                  190.61
Request throughput (req/s):              0.05
Peak concurrent requests:                2.00
-------------------Peak Memory--------------------
Mean PEAK_MEMORY_MB (MB):                47952.40
Median PEAK_MEMORY_MB (MB):              48029.00
P99 PEAK_MEMORY_MB (MB):                 48073.48
---------------- Stage Durations -----------------
Mean Text Encoder Forward (s):           0.08
Median Text Encoder Forward (s):         0.07
P99 Text Encoder Forward (s):            0.14
Mean Vae Encode (s):                     1.63
Median Vae Encode (s):                   1.47
P99 Vae Encode (s):                      2.75
Mean Diffuse (s):                        12.09
Median Diffuse (s):                      12.01
P99 Diffuse (s):                         12.45
Mean Vae Decode (s):                     2.97
Median Vae Decode (s):                   2.98
P99 Vae Decode (s):                      3.31
Mean Queue Wait (ms):                    0.48
Median Queue Wait (ms):                  0.45
P99 Queue Wait (ms):                     0.77
----------------End-to-end Latency----------------
Mean E2EL (ms):                          19060.28
Median E2EL (ms):                        19058.84
P99 E2EL (ms):                           20079.09
================== Video Result ==================
Total video duration generated(s):       50.62
Total video frames generated:            810
Video throughput(video duration/s):      0.27
------------------- Video RTF --------------------
Mean VIDEO_RTF:                          3.54
Median VIDEO_RTF:                        3.55
P99 VIDEO_RTF:                           3.83
---------------- Video Generation ----------------
Mean VIDEO_GENERATION (ms):              17922.07
Median VIDEO_GENERATION (ms):            17967.14
P99 VIDEO_GENERATION (ms):               19372.30
==================================================
  • hunyuanvideo
============ Serving Benchmark Result ============                                                                      
Successful requests:                     10                                                                             
Failed requests:                         0                                                                              
Maximum request concurrency:             1                                                                              
Benchmark duration (s):                  80.26                                                                          
Request throughput (req/s):              0.12                                                                           
Peak concurrent requests:                2.00                                                                           
-------------------Peak Memory--------------------                                                                      
Mean PEAK_MEMORY_MB (MB):                77912.00                                                                       
Median PEAK_MEMORY_MB (MB):              77912.00                                                                       
P99 PEAK_MEMORY_MB (MB):                 77912.00                                                                       
---------------- Stage Durations -----------------                                                                      
Mean Text Encoder Forward (s):           0.07                                                                           
Median Text Encoder Forward (s):         0.07                                                                           
P99 Text Encoder Forward (s):            0.08                                                                           
Mean Vae Decode (s):                     3.89                                                                           
Median Vae Decode (s):                   3.80                                                                           
P99 Vae Decode (s):                      4.58                                                                           
Mean Queue Wait (ms):                    0.51                                                                           
Median Queue Wait (ms):                  0.51                                                                           
P99 Queue Wait (ms):                     0.67                                                                           
----------------End-to-end Latency----------------                                                                      
Mean E2EL (ms):                          8025.89                                                                        
Median E2EL (ms):                        8026.63                                                                        
P99 E2EL (ms):                           8028.76                                                                        
================== Video Result ==================                                                                      
Total video duration generated(s):       13.75                                                                          
Total video frames generated:            330                                                                            
Video throughput(video duration/s):      0.17                                                                           
------------------- Video RTF --------------------                                                                      
Mean VIDEO_RTF:                          4.93                                                                           
Median VIDEO_RTF:                        4.85                                                                           
P99 VIDEO_RTF:                           5.38                                                                           
---------------- Video Generation ----------------                                                                      
Mean VIDEO_GENERATION (ms):              6773.53                                                                        
Median VIDEO_GENERATION (ms):            6672.90                                                                        
P99 VIDEO_GENERATION (ms):               7397.11                                                                        
==================================================   
============ Serving Benchmark Result ============
Successful requests:                     10
Failed requests:                         0
Maximum request concurrency:             1
Benchmark duration (s):                  60.32
Request throughput (req/s):              0.17
Peak concurrent requests:                2.00
-------------------Peak Memory--------------------
Mean PEAK_MEMORY_MB (MB):                43518.00
Median PEAK_MEMORY_MB (MB):              43518.00
P99 PEAK_MEMORY_MB (MB):                 43518.00
---------------- Stage Durations -----------------
Mean Text Encoder Forward (s):           0.07
Median Text Encoder Forward (s):         0.07
P99 Text Encoder Forward (s):            0.07
Mean Vae Encode (s):                     0.13
Median Vae Encode (s):                   0.07
P99 Vae Encode (s):                      0.33
Mean Vae Decode (s):                     1.69
Median Vae Decode (s):                   1.65
P99 Vae Decode (s):                      1.91
Mean Queue Wait (ms):                    0.37
Median Queue Wait (ms):                  0.41
P99 Queue Wait (ms):                     0.49
----------------End-to-end Latency----------------
Mean E2EL (ms):                          6030.54
Median E2EL (ms):                        6029.90
P99 E2EL (ms):                           6038.40
================== Video Result ==================
Total video duration generated(s):       13.75
Total video frames generated:            330
Video throughput(video duration/s):      0.23
------------------- Video RTF --------------------
Mean VIDEO_RTF:                          3.20
Median VIDEO_RTF:                        3.06
P99 VIDEO_RTF:                           3.71
---------------- Video Generation ----------------
Mean VIDEO_GENERATION (ms):              4401.38
Median VIDEO_GENERATION (ms):            4213.14
P99 VIDEO_GENERATION (ms):               5099.12
==================================================
============ Serving Benchmark Result ============                                                                      
Successful requests:                     3                                                                              
Failed requests:                         0                                                                              
Maximum request concurrency:             1                                                                              
Benchmark duration (s):                  6.02                                                                           
Request throughput (req/s):              0.50                                                                           
Peak concurrent requests:                2.00                                                                           
-------------------Peak Memory--------------------                                                                      
Mean PEAK_MEMORY_MB (MB):                20034.00                                                                       
Median PEAK_MEMORY_MB (MB):              20034.00                                                                       
P99 PEAK_MEMORY_MB (MB):                 20034.00                                                                       
---------------- Stage Durations -----------------                                                                      
Mean Queue Wait (ms):                    0.49                                                                           
Median Queue Wait (ms):                  0.56                                                                           
P99 Queue Wait (ms):                     0.62                                                                           
----------------End-to-end Latency----------------                                                                      
Mean E2EL (ms):                          2007.16                                                                        
Median E2EL (ms):                        2006.98                                                                        
P99 E2EL (ms):                           2007.84                                                                        
================== Video Result ==================                                                                      
Total video duration generated(s):       1.12                                                                           
Total video frames generated:            27                                                                             
Video throughput(video duration/s):      0.19                                                                           
------------------- Video RTF --------------------                                                                      
Mean VIDEO_RTF:                          0.43                                                                           
Median VIDEO_RTF:                        0.43                                                                           
P99 VIDEO_RTF:                           0.44                                                                           
---------------- Video Generation ----------------                                                                      
Mean VIDEO_GENERATION (ms):              161.08                                                                         
Median VIDEO_GENERATION (ms):            161.04                                                                         
P99 VIDEO_GENERATION (ms):               165.32                                                                         
================================================== 
  • ltx2
============ Serving Benchmark Result ============                                                                      
Successful requests:                     3                                                                              
Failed requests:                         0                                                                              
Maximum request concurrency:             1                                                                              
Benchmark duration (s):                  24.07                                                                          
Request throughput (req/s):              0.12                                                                           
Peak concurrent requests:                2.00                                                                           
-------------------Peak Memory--------------------                                                                      
Mean PEAK_MEMORY_MB (MB):                70962.00                                                                       
Median PEAK_MEMORY_MB (MB):              70962.00                                                                       
P99 PEAK_MEMORY_MB (MB):                 70962.00                                                                       
---------------- Stage Durations -----------------                                                                      
Mean Generate  Build Phase Inputs (s):   0.00                                                                           
Median Generate  Build Phase Inputs (s): 0.00                                                                           
P99 Generate  Build Phase Inputs (s):    0.00                                                                           
Mean Text Encoder Forward (s):           0.14                                                                           
Median Text Encoder Forward (s):         0.14                                                                           
P99 Text Encoder Forward (s):            0.14                                                                           
Mean Generate Run Phase (s):             6.35                                                                           
Median Generate Run Phase (s):           6.35                                                                           
P99 Generate Run Phase (s):              6.36                                                                           
Mean Vae Decode (s):                     0.08                                                                           
Median Vae Decode (s):                   0.08                                                                           
P99 Vae Decode (s):                      0.08                                                                           
Mean  Prepare Video Output For Transport (s): 0.00                                                                      
Median  Prepare Video Output For Transport (s): 0.00                                                                    
P99  Prepare Video Output For Transport (s): 0.00                                                                       
Mean Audio Vae Decode (s):               0.00                                                                           
Median Audio Vae Decode (s):             0.00                                                                           
P99 Audio Vae Decode (s):                0.00                                                                           
Mean Vocoder Forward (s):                0.01                                                                           
Median Vocoder Forward (s):              0.01                                                                           
P99 Vocoder Forward (s):                 0.01                                                                           
Mean Decode Phase (s):                   0.09                                                                           
Median Decode Phase (s):                 0.09                                                                           
P99 Decode Phase (s):                    0.10                                                                           
Mean Queue Wait (ms):                    0.33                                                                           
Median Queue Wait (ms):                  0.37                                                                           
P99 Queue Wait (ms):                     0.39                                                                           
----------------End-to-end Latency----------------                                                                      
Mean E2EL (ms):                          8022.23                                                                        
Median E2EL (ms):                        8020.89                                                                        
P99 E2EL (ms):                           8024.96                                                                        
================== Video Result ==================                                                                      
Total video duration generated(s):       5.12                                                                           
Total video frames generated:            123                                                                            
Video throughput(video duration/s):      0.21                                                                           
------------------- Video RTF --------------------                                                                      
Mean VIDEO_RTF:                          3.79                                                                           
Median VIDEO_RTF:                        3.79                                                                           
P99 VIDEO_RTF:                           3.79                                                                           
---------------- Video Generation ----------------                                                                      
Mean VIDEO_GENERATION (ms):              6472.87                                                                        
Median VIDEO_GENERATION (ms):            6472.01                                                                        
P99 VIDEO_GENERATION (ms):               6477.46                                                                        
================================================== 
  • cosmos
============ Serving Benchmark Result ============                                                                      
Successful requests:                     3                                                                              
Failed requests:                         0                                                                              
Maximum request concurrency:             1                                                                              
Benchmark duration (s):                  85.50                                                                          
Request throughput (req/s):              0.04                                                                           
Peak concurrent requests:                2.00                                                                           
-------------------Peak Memory--------------------                                                                      
Mean PEAK_MEMORY_MB (MB):                27768.00                                                                       
Median PEAK_MEMORY_MB (MB):              27768.00                                                                       
P99 PEAK_MEMORY_MB (MB):                 27768.00                                                                       
---------------- Stage Durations -----------------                                                                      
Mean Diffuse (s):                        10.87                                                                          
Median Diffuse (s):                      10.86                                                                          
P99 Diffuse (s):                         10.88                                                                          
Mean Vae Decode (s):                     11.14                                                                          
Median Vae Decode (s):                   11.14                                                                          
P99 Vae Decode (s):                      11.14                                                                          
Mean Queue Wait (ms):                    0.44                                                                           
Median Queue Wait (ms):                  0.44                                                                           
P99 Queue Wait (ms):                     0.52                                                                           
----------------End-to-end Latency----------------                                                                      
Mean E2EL (ms):                          28500.32                                                                       
Median E2EL (ms):                        28475.87                                                                       
P99 E2EL (ms):                           28553.79                                                                       
================== Video Result ==================                                                                      
Total video duration generated(s):       23.62                                                                          
Total video frames generated:            567                                                                            
Video throughput(video duration/s):      0.28                                                                           
------------------- Video RTF --------------------                                                                      
Mean VIDEO_RTF:                          3.27                                                                           
Median VIDEO_RTF:                        3.27                                                                           
P99 VIDEO_RTF:                           3.28                                                                           
---------------- Video Generation ----------------                                                                      
Mean VIDEO_GENERATION (ms):              25770.59                                                                       
Median VIDEO_GENERATION (ms):            25753.24                                                                       
P99 VIDEO_GENERATION (ms):               25826.29                                                                       
================================================== 
  • minimax
============ Serving Benchmark Result ============                                                                      
Successful requests:                     3                                                                              
Failed requests:                         0                                                                              
Maximum request concurrency:             1                                                                              
Benchmark duration (s):                  92.48                                                                          
Request throughput (req/s):              0.03                                                                           
Peak concurrent requests:                2.00                                                                           
-------------------Peak Memory--------------------                                                                      
Mean PEAK_MEMORY_MB (MB):                53838.00                                                                       
Median PEAK_MEMORY_MB (MB):              53838.00                                                                       
P99 PEAK_MEMORY_MB (MB):                 53838.00                                                                       
---------------- Stage Durations -----------------                                                                      
Mean Encode Prompt (s):                  0.29                                                                           
Median Encode Prompt (s):                0.16                                                                           
P99 Encode Prompt (s):                   0.59                                                                           
Mean  Encode Local Media (s):            0.15                                                                           
Median  Encode Local Media (s):          0.15                                                                           
P99  Encode Local Media (s):             0.16                                                                           
Mean Diffuse (s):                        26.28                                                                          
Median Diffuse (s):                      26.38                                                                          
P99 Diffuse (s):                         26.39                                                                          
Mean Video Vae Decode Latent (s):        1.94                                                                           
Median Video Vae Decode Latent (s):      1.94                                                                           
P99 Video Vae Decode Latent (s):         1.94                                                                           
Mean Audio Vae Decode Latent (s):        0.06                                                                           
Median Audio Vae Decode Latent (s):      0.06                                                                           
P99 Audio Vae Decode Latent (s):         0.06                                                                           
Mean Decode (s):                         2.01                                                                           
Median Decode (s):                       2.01                                                                           
P99 Decode (s):                          2.01                                                                           
Mean Queue Wait (ms):                    0.75                                                                           
Median Queue Wait (ms):                  0.44                                                                           
P99 Queue Wait (ms):                     1.47                                                                           
----------------End-to-end Latency----------------                                                                      
Mean E2EL (ms):                          30824.98                                                                       
Median E2EL (ms):                        30102.97                                                                       
P99 E2EL (ms):                           32229.25                                                                       
================== Video Result ==================                                                                      
Total video duration generated(s):       26.12                                                                          
Total video frames generated:            627                                                                            
Video throughput(video duration/s):      0.28                                                                           
------------------- Video RTF --------------------                                                                      
Mean VIDEO_RTF:                          3.38                                                                           
Median VIDEO_RTF:                        3.37                                                                           
P99 VIDEO_RTF:                           3.40                                                                           
---------------- Video Generation ----------------                                                                      
Mean VIDEO_GENERATION (ms):              29441.88                                                                       
Median VIDEO_GENERATION (ms):            29384.45                                                                       
P99 VIDEO_GENERATION (ms):               29571.15                                                                       
==================================================  
  • artifact_paths
7018c8fe-ac81-4302-9383-8a72550fc02b

Pending: full Buildkite nightly for migrated run_benchmark.py jobs.

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)

Signed-off-by: wangyu <410167048@qq.com>
…arity

- Updated references from `run_diffusion_benchmark.py` to `run_benchmark.py` in various CI configurations and test files to standardize the script used for performance benchmarks.
- Adjusted test configurations to reflect changes in dataset naming and parameters, ensuring compatibility with the new benchmark script.
- Enhanced documentation to clarify the distinction between diffusion and omni benchmarks, including updates to test examples and execution guides.

This refactor aims to streamline the testing process and improve maintainability across the codebase.

Signed-off-by: wangyu <410167048@qq.com>
- Introduced `aggregate_stage_durations` and `print_stage_durations_metrics` functions to compute and display mean, p50, and p99 stage durations from request outputs.
- Updated `MixRequestFuncOutput` to include `stage_durations` for tracking per-stage timings.
- Enhanced test coverage with new tests for aggregating and printing stage durations metrics.
- Refactored existing code to utilize `SimpleNamespace` for cleaner tokenization output.

This update improves the observability of performance metrics during benchmarking, facilitating better analysis of stage timings.

Signed-off-by: [Your Name] <your.email@example.com>
Signed-off-by: wangyu <410167048@qq.com>
…uce random input length

- Added "tokenizer": "gpt2" to multiple video performance test configurations for consistency.
- Reduced "random_input_len" from 64 to 8 across various test cases to optimize input handling.

This change enhances the uniformity of tokenizer usage and improves the efficiency of input processing in performance tests.

Signed-off-by: wangyu <410167048@qq.com>
- Updated `print_stage_durations_metrics` to format output with clearer labels and units for mean, median, and p99 stage durations.
- Introduced a new helper function `_stage_duration_display_name` to prettify stage names for better readability in printed metrics.
- Modified test cases to reflect changes in output formatting and ensure accurate assertions.

This update improves the clarity and usability of performance metrics during benchmarking, aiding in the analysis of stage timings.

Signed-off-by: wangyu <410167048@qq.com>
…arity

- Updated references from `run_diffusion_benchmark.py` to `run_benchmark.py` across CI configurations and test files to standardize the script used for performance benchmarks.
- Adjusted artifact paths and environment variable names for consistency in performance test configurations.
- Enhanced documentation to clarify the distinction between diffusion and omni benchmarks, including updates to test examples and execution guides.

This refactor aims to streamline the testing process and improve maintainability across the codebase.

Signed-off-by: wangyu <410167048@qq.com>
@yenuo26 yenuo26 changed the title ci(perf): migrate diffusion DFX benches to vllm bench serve --omni [CI][Perf] migrate diffusion DFX benches to vllm bench serve --omni Sep 17, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/benchmarking.md, docs/design/module/diffusion/index.md.

Module owners: @david6666666 @Isotr0py @princepride

Routing: @david6666666 via module named in the PR description; @Isotr0py via module named in the PR description; @princepride via module named in the PR description

@yenuo26, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@ZacheryAU ZacheryAU left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Some possible deduplicating founding:

  • can print_stage_durations_metrics reuse _print_percentile_metric or process_one_metric?
  • can use --print-stage to show stage benchmark?

And also a printing redundancy of Stage 0 Gen and stage_gen_time of stage 0 when --print-stage activated, but it depends on whether downstream prefers Stage 0 Gen.

… models

- Renamed `run_benchmark.py` to `run_diffusion_benchmark.py` in CI configurations and test files to clearly differentiate between omni and diffusion benchmarks.
- Adjusted artifact paths and environment variable names for consistency across performance test configurations.
- Enhanced test cases to include `server_type` for diffusion models and updated benchmark parameters for clarity and uniformity.
- Improved documentation to reflect changes in test execution and configuration, ensuring better guidance for contributors.

These updates aim to streamline the benchmarking process and enhance the maintainability of the codebase.

Signed-off-by: wangyu <410167048@qq.com>
@yenuo26
yenuo26 requested a review from wtomin as a code owner September 18, 2026 08:58
@yenuo26

yenuo26 commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator Author

Some possible deduplicating founding:

  • can print_stage_durations_metrics reuse _print_percentile_metric or process_one_metric?
  • can use --print-stage to show stage benchmark?

And also a printing redundancy of Stage 0 Gen and stage_gen_time of stage 0 when --print-stage activated, but it depends on whether downstream prefers Stage 0 Gen.

Thanks
Reuse: process_one_metric only prints fixed fields like mean_ttft_ms. _print_percentile_metric needs raw samples and opens a new section each time. Stage durations are dynamic keys (diffuse, vae.decode, …) already aggregated into maps, so they are printed like print_video_metrics / print_image_metrics, which also do not call those helpers.

Stage 0 Gen: removed. stage_N_gen_ms is no longer printed or saved under stage_durations_*. That clock stays on stage_gen_time / IMAGE_GENERATION / VIDEO_GENERATION.

--print-stage: not used for these metrics. stage_durations_mean/p50/p99 live on MultiModalsBenchmarkMetrics, same as image / video / peak memory, so they print and are saved whenever the data exists. --print-stage still only prints per-stage_id StageBenchmarkMetrics.

- Updated `aggregate_stage_durations` to exclude `stage_N_gen_ms` from the metrics, ensuring clarity in the reported timings.
- Modified `print_stage_durations_metrics` to utilize metrics stored in `MultiModalsBenchmarkMetrics`, improving the output format and consistency.
- Adjusted test cases to validate the exclusion of `stage_N_gen_ms` and ensure accurate assertions in metrics display.

These changes improve the accuracy and readability of stage duration metrics during benchmarking, facilitating better performance analysis.

Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	tests/benchmarks/patch/test_patch.py
#	tests/dfx/perf/tests/test_hunyuanvideo15_t2v_vllm_omni.json
#	vllm_omni/benchmarks/patch/patch.py
@hsliuustc0106 hsliuustc0106 added the ready label to trigger buildkite CI label Sep 19, 2026
- Introduced `_iter_video_reference_inputs` to yield video references from multimodal content.
- Updated `_add_video_reference_to_form` to handle structured video references and integrate them into form data.
- Enhanced tests to validate the extraction and addition of video references, ensuring compatibility with the existing image reference handling.

These changes improve the support for video content in the benchmarking framework, aligning with the existing image reference functionality.

Signed-off-by: wangyu <410167048@qq.com>
- Updated `_add_video_reference_to_form` to handle inline video data URLs more effectively, ensuring they are added as binary input references instead of JSON strings.
- Adjusted tests to validate the new handling of video references, ensuring that the correct fields are populated and that no legacy fields are present.
- Modified benchmark configuration to reflect changes in video dimensions.

These improvements enhance the robustness of video content handling in the benchmarking framework.

Signed-off-by: wangyu <410167048@qq.com>
Signed-off-by: wangyu <410167048@qq.com>
@Gaohan123 Gaohan123 added this to the v0.30.0 milestone Sep 20, 2026
@hsliuustc0106 hsliuustc0106 added the CI/CD codes related to changes to CI/CD label Sep 20, 2026
Comment thread vllm_omni/benchmarks/patch/patch.py Outdated
if _add_video_reference_to_form(form, reference):
video_reference_added = True
break
if not video_reference_added:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Keep mixed image and inline-video references compatible with the API

When a request contains both an image reference and a data:video/... reference, the preceding image loop adds image_reference, while this loop uploads the video as input_references. _parse_video_form() explicitly rejects that combination with HTTP 400, so these mixed-reference benchmark requests fail before inference.

Please serialize the combined references using a field combination supported by the server, and add a regression test covering an image plus an inline video. This is an additional case related to the existing reference-handling discussion; the HTTP(S) video-type issue is fixed.

@yenuo26 yenuo26 Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A combined video-and-image mode has been added. The related selection logic is as follows:

Scenario Reference Form Target Field
Image only {"image_url"} / data:image / http(s) image_reference
Image only Local path / {"bytes"} input_reference
Video only {"video_url"} over http(s) video_reference
Video only data:video input_references (upload, to avoid the ~1MB JSON limit)
Image + Video Image + Video (including data:video) image_reference + video_reference; input_references is no longer used

1.Since input_references cannot be passed together with image_reference, the only option is to combine them as image_reference + video_reference; however, this combination easily hits the JSON 1MB limit.

2.We considered putting both the image and video into input_references, but this approach requires the model to support mixed_reference_inputs. Otherwise, all reference_inputs would be treated as video files. Currently, only MiniMax supports this.
The current implementation follows option 1.

@hsliuustc0106 hsliuustc0106 added the enhancement New feature or request label Sep 20, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

fix conflicts

- Introduced `_add_combined_video_form_references` to streamline the addition of image and video references in form data, ensuring compatibility with server requirements.
- Updated `_add_video_reference_to_form` to reject unsupported `file_id` references, improving error handling.
- Added new tests to validate the rejection of `file_id` references and ensure correct processing of combined image and video references.

These changes enhance the robustness and clarity of video and image reference handling in the benchmarking framework.

Signed-off-by: wangyu <410167048@qq.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

# Conflicts:
#	vllm_omni/benchmarks/patch/patch.py
@yenuo26

yenuo26 commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator Author

fix conflicts

fixed

Comment thread vllm_omni/benchmarks/patch/patch.py Outdated
if not video_refs and extra_body.get("video_reference") is not None:
video_refs = [extra_body["video_reference"]]

upload_inline_video = not image_refs

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[P2] Normalize all reference forms before choosing combined form fields

The structured image/video case is fixed, but two supported input forms still produce mutually exclusive fields:

  • An image URL plus a bare data:video/... string in extra_body["video_reference"] produces image_reference + input_references, because the string branch ignores upload_inline_video=False.
  • Image upload bytes plus a structured video reference produces input_reference + video_reference, because the image branch still uses the singular upload field.

_parse_video_form() rejects both combinations with HTTP 400, so these benchmark requests fail before inference. I reproduced both using the current serialization helpers and server validation conditions.

Please normalize the reference forms before selecting the combined encoding, ensure the upload flag also applies to bare video strings, and add regression coverage for both cases.

@yenuo26 yenuo26 Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have revised the overall judgment logic as follows:
Only one category of reference

Judgment Input Field
Image {"image_url"}, data:image, http(s) with an image extension (.png .jpg .jpeg .webp .gif .bmp .heic .heif) image_reference
Video {"video_url"}, http(s) with a video extension (.mp4 .mov .webm .mkv .m4v), data:video with JSON text under 1MB video_reference
Video, oversized text Video only, and the JSON text of {"video_url":"data:video..."} is larger than 1MB input_references (upload the decoded bytes)
File Local path, {"bytes"} input_reference
Image and video together Both sides written as JSON: image_reference + video_reference. In this case, input_reference or input_references is no longer used.

video+image

Original form Before sending
Image URL, data:image, {"image_url"} Put into image_reference as-is
Image bytes, local image Convert to data:image, put into image_reference
Video URL, data:video, {"video_url"} Put into video_reference. Even if data:video exceeds 1MB, do not change it to an upload
Video bytes, local video Convert to data:video, put into video_reference

update test result

pytest -sv tests/benchmarks/patch/test_patch.py -m "core_model and cpu"
image
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py \
  --test-config-file tests/dfx/perf/tests/test_hunyuanvideo15_i2v_vllm_omni.json \
  -m "H100 and B200 and (cards_1 or cards_2)"
============ Serving Benchmark Result ============
Successful requests:                     10
Failed requests:                         0
Maximum request concurrency:             1
Benchmark duration (s):                  80.27
Request throughput (req/s):              0.12
Peak concurrent requests:                2.00
-------------------Peak Memory--------------------
Mean PEAK_MEMORY_MB (MB):                77912.00
Median PEAK_MEMORY_MB (MB):              77912.00
P99 PEAK_MEMORY_MB (MB):                 77912.00
---------------- Stage Durations -----------------
Mean Text Encoder Forward (s):           0.07
Median Text Encoder Forward (s):         0.07
P99 Text Encoder Forward (s):            0.08
Mean Vae Decode (s):                     3.74
Median Vae Decode (s):                   3.69
P99 Vae Decode (s):                      4.35
Mean Queue Wait (ms):                    0.51
Median Queue Wait (ms):                  0.56
P99 Queue Wait (ms):                     0.67
----------------End-to-end Latency----------------
Mean E2EL (ms):                          8026.27
Median E2EL (ms):                        8026.58
P99 E2EL (ms):                           8029.93
================== Video Result ==================
Total video duration generated(s):       13.75
Total video frames generated:            330
Video throughput(video duration/s):      0.17
------------------- Video RTF --------------------
Mean VIDEO_RTF:                          4.88
Median VIDEO_RTF:                        4.84
P99 VIDEO_RTF:                           5.39
---------------- Video Generation ----------------
Mean VIDEO_GENERATION (ms):              6709.08
Median VIDEO_GENERATION (ms):            6660.30
P99 VIDEO_GENERATION (ms):               7415.70
==================================================

- Refactored functions to classify and process image and video references, ensuring proper handling of bare URLs and data URIs.
- Introduced new tests to validate the correct processing of video references, including oversized data videos and their handling in forms.
- Improved error handling for unsupported video formats and added checks for reference types.

These changes improve the robustness and clarity of the benchmarking framework's media reference handling.

Signed-off-by: wangyu <410167048@qq.com>

@Gaohan123 Gaohan123 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. Thanks

@Gaohan123 Gaohan123 added the ready label to trigger buildkite CI label Sep 21, 2026
@Gaohan123
Gaohan123 enabled auto-merge (squash) September 21, 2026 14:47
@yenuo26

yenuo26 commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator Author

The CI failure was due to a known issue(#7912) that has now been fixed. Update the code and retry.

@yenuo26 yenuo26 removed the ready label to trigger buildkite CI label Sep 22, 2026
@yenuo26 yenuo26 added the ready label to trigger buildkite CI label Sep 22, 2026
@Gaohan123
Gaohan123 merged commit 0570b0d into vllm-project:main Sep 22, 2026
6 of 9 checks passed
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
…llm-project#7737)

Signed-off-by: wangyu <410167048@qq.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
@yenuo26
yenuo26 deleted the perf branch September 24, 2026 06:22
congw729 added a commit to congw729/vllm-omni that referenced this pull request Sep 24, 2026
…ation

Merge current main and rebuild Wan2.2 A3 baseline blocks on the new
vllm bench serve --omni configuration introduced by vllm-project#7737. Rename
throughput and memory metrics; convert the archived latency means from
seconds to milliseconds before rounding to four decimals.

Retain the selected nightly windows, including MiniCPM-o c8 Sep 18-19.

Signed-off-by: Alicia <115451386+congw729@users.noreply.github.com>
zengchuang-hw added a commit to zengchuang-hw/vllm-omni that referenced this pull request Sep 24, 2026
vllm-project#7737 migrated test_hunyuanvideo15_t2v_vllm_omni.json to the omni-bench
schema (benchmark_params[].dataset_name) but missed the NPU nightly
Diffusion X2V HunyuanVideo-1.5 Perf Test step, which still invoked
run_diffusion_benchmark.py. That runner's is_diffusion_perf_config
filter skipped both cases as omni-bench, pytest selected 0 tests and
the step failed with exit 5 (issue vllm-project#8074).

Switch the step to run_benchmark.py with BENCHMARK_DIR and results/*.json
artifact upload, matching the already-migrated Wan22 NPU step and the
CUDA HunyuanVideo-1.5 step. Add a regression test asserting every
Buildkite step runs a perf JSON with the runner matching its schema.

Fixes vllm-project#8074
zengchuang-hw added a commit to zengchuang-hw/vllm-omni that referenced this pull request Sep 24, 2026
vllm-project#7737 migrated test_hunyuanvideo15_t2v_vllm_omni.json to the omni-bench
schema (benchmark_params[].dataset_name) but missed the NPU nightly
Diffusion X2V HunyuanVideo-1.5 Perf Test step, which still invoked
run_diffusion_benchmark.py. That runner's is_diffusion_perf_config
filter skipped both cases as omni-bench, pytest selected 0 tests and
the step failed with exit 5 (issue vllm-project#8074).

Switch the step to run_benchmark.py with BENCHMARK_DIR and results/*.json
artifact upload, matching the already-migrated Wan22 NPU step and the
CUDA HunyuanVideo-1.5 step. Add a regression test asserting every
Buildkite step runs a perf JSON with the runner matching its schema.

Fixes vllm-project#8074

Signed-off-by: zengchuang <zengchuang3@huawei.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

CI/CD codes related to changes to CI/CD enhancement New feature or request high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants