Skip to content

Benchmark data statistics for each stage of omni models - #3628

Merged
Gaohan123 merged 15 commits into
vllm-project:mainfrom
ZacheryAU:fix_metric
Jun 4, 2026
Merged

Gaohan123 merged 15 commits into
vllm-project:mainfrom
ZacheryAU:fix_metric

Conversation

@ZacheryAU

@ZacheryAU ZacheryAU commented May 15, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

implement: #3967 #1361 #3593 with referring #3545

Summary

  • Collect and print the metrics within each stage of multi-stage models including:
    • stage_gen_time: Time from submitting a request to a specific stage to that stage finishing generation.
    • Serving TTFC (Time to First Chunk): Time from the HTTP request being accepted by the serving frontend to the stage producing its first non-empty output chunk. Keep the name TTFT in text stage, TTFP in audio stage for easier mapping.
    • TPOP (Time per Output Chunk): Average time from the first output chunk to stage completion, divided by the number of remaining output chunks. The TPOP abbreviation follows Qwen3.5-Omni Technical Report.
    • ICL (Inter-Chunk Latency): Time between two consecutive output chunks produced by the same stage. Keep the name ITL in text stage for easier mapping.
  • text/audio/image stages can show different groups of metrics
  • Add --print-stage to control whether or not to print stage metrics (default: false)
  • Add metrics in --percentile_metrics to adjust the printing inside stage level:
    • ttfc: enable Serving TTFC in internal stream stage (like talker in Qwen3-omni)
    • tpop: enable TPOP (include both TTFT in text stage and TTFC in internal stream stage)
    • icl: enable ICL in internal stream stage
  • Bug fixed about printing result of BAGEL and HunyuanImage-3.0
  • Add support for /v1/images/edits
  • Add Image metrics (pixels, denoise step latency, throughput, etc) in Serving Benchmark Result (see Result of HunyuanImage-3.0-Instruct)

Test Plan

  1. run vllm bench serve --omni --endpoint /v1/chat/completions --print-stage with two different types of multi-modal models (audio & image) and see if metrics within each of stages are shown
  2. run vllm bench serve --omni --endpoint /v1/imgaes/edits --print-stage for HunyuanImage-3.0-Instruct
  3. run without --print-stage as regression test to see whether only Serving Benchmark Result is shown

Test Result

Result of Qwen3-Omni-30B-A3B-Instruct

Show more
============ Serving Benchmark Result ============
Successful requests:                     3         
Failed requests:                         0         
Maximum request concurrency:             1         
Benchmark duration (s):                  50.76     
Request throughput (req/s):              0.06      
Peak concurrent requests:                2.00      
----------------End-to-end Latency----------------
Mean E2EL (ms):                          16916.12  
Median E2EL (ms):                        13670.22  
P99 E2EL (ms):                           25046.74  
================== Text Result ===================
Total input tokens:                      7536      
Total generated tokens:                  4500      
Output token throughput (tok/s):         88.66     
Peak output token throughput (tok/s):    85.00     
Peak concurrent requests:                2.00      
Total Token throughput (tok/s):          237.13    
---------------Time to First Token----------------
Mean TTFT (ms):                          693.16    
Median TTFT (ms):                        700.10    
P99 TTFT (ms):                           702.42    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          8.23      
Median TPOT (ms):                        6.18      
P99 TPOT (ms):                           12.25     
---------------Inter-token Latency----------------
Mean ITL (ms):                           12.33     
Median ITL (ms):                         12.18     
P99 ITL (ms):                            16.43     
================== Audio Result ==================
Total audio duration generated(s):       310.55    
Total audio frames generated:            7453260   
Audio throughput(audio duration/s):      6.12      
Streaming continuity OK rate:            100.00%   
---------------Time to First Packet---------------
Mean AUDIO_TTFP (ms):                    1186.69   
Median AUDIO_TTFP (ms):                  1183.57   
P99 AUDIO_TTFP (ms):                     1194.23   
-----------------Real Time Factor-----------------
Mean AUDIO_RTF:                          0.16      
Median AUDIO_RTF:                        0.16      
P99 AUDIO_RTF:                           0.17
============= Stage Benchmark Result ============= =============== Stage 0 (thinker) ================ -------------------Stage Timing------------------- Mean stage_gen_time (ms): 19579.87 Median stage_gen_time (ms): 23429.85 P99 stage_gen_time (ms): 23541.98 ================== Text Result =================== Stage generated tokens: 4500 -----------Serving Time to First Token------------ Mean Serving TTFT (ms): 689.49 Median Serving TTFT (ms): 696.86 P99 Serving TTFT (ms): 698.36 ---------------Time to First Token---------------- Mean TTFT (ms): 680.50 Median TTFT (ms): 682.27 P99 TTFT (ms): 688.57 -----Time per Output Token (excl. 1st token)------ Mean TPOT (ms): 12.34 Median TPOT (ms): 12.34 P99 TPOT (ms): 12.37 ---------------Inter-token Latency---------------- Mean ITL (ms): 12.34 Median ITL (ms): 12.26 P99 ITL (ms): 13.74 ================ Stage 1 (talker) ================ -------------------Stage Timing------------------- Mean stage_gen_time (ms): 16326.31 Median stage_gen_time (ms): 13640.64 P99 stage_gen_time (ms): 25014.86 ============= Internal Stream Result ============= -----------Serving Time to First Chunk------------ Mean Serving TTFC (ms): 876.51 Median Serving TTFC (ms): 873.99 P99 Serving TTFC (ms): 886.69 -----Time per Output Chunk (excl. 1st chunk)------ Mean TPOP (ms): 11.92 Median TPOP (ms): 12.07 P99 TPOP (ms): 12.09 ---------------Inter-chunk Latency---------------- Mean ICL (ms): 11.82 Median ICL (ms): 11.51 P99 ICL (ms): 20.44 =============== Stage 2 (code2wav) =============== -------------------Stage Timing------------------- Mean stage_gen_time (ms): 19692.02 Median stage_gen_time (ms): 20187.22 P99 stage_gen_time (ms): 25146.48 ================== Audio Result ================== Stage audio duration generated(s): 370.98 Stage audio frames generated: 8903490 -----------Serving Time to First Packet----------- Mean Serving AUDIO_TTFP (ms): 1181.69 Median Serving AUDIO_TTFP (ms): 1178.37 P99 Serving AUDIO_TTFP (ms): 1189.28 -----------------Real Time Factor----------------- Mean AUDIO_RTF: 0.16 Median AUDIO_RTF: 0.16 P99 AUDIO_RTF: 0.17 ==================================================

Result of BAGEL-7B-MoT

Show more
============ Serving Benchmark Result ============
Successful requests:                     3         
Failed requests:                         0         
Maximum request concurrency:             1         
Benchmark duration (s):                  56.43     
Request throughput (req/s):              0.05      
Peak concurrent requests:                2.00      
----------------End-to-end Latency----------------
Mean E2EL (ms):                          18808.02  
Median E2EL (ms):                        18825.57  
P99 E2EL (ms):                           18836.89  
================== Text Result ===================
Total input tokens:                      7502      
Total generated tokens:                  18        
Output token throughput (tok/s):         0.32      
Peak output token throughput (tok/s):    5.00      
Peak concurrent requests:                2.00      
Total Token throughput (tok/s):          133.27    
---------------Time to First Token----------------
Mean TTFT (ms):                          127.50    
Median TTFT (ms):                        132.21    
P99 TTFT (ms):                           137.63    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          70.52     
Median TPOT (ms):                        70.97     
P99 TPOT (ms):                           70.97     
---------------Inter-token Latency----------------
Mean ITL (ms):                           88.15     
Median ITL (ms):                         3.77      
P99 ITL (ms):                            347.22    
================== Image Result ==================
Total images generated:                  3         
Image throughput (img/s):                0.05      
Average pixels per image:                1048576.00
-----------------Image Generation-----------------
Mean IMAGE_GENERATION (ms):              18174.88  
Median IMAGE_GENERATION (ms):            18176.98  
P99 IMAGE_GENERATION (ms):               18185.23 
============= Stage Benchmark Result ============= =============== Stage 0 (thinker) ================ -------------------Stage Timing------------------- Mean stage_gen_time (ms): 875.69 Median stage_gen_time (ms): 877.06 P99 stage_gen_time (ms): 879.56 ================== Text Result =================== Stage generated tokens: 18 -----------Serving Time to First Token------------ Mean Serving TTFT (ms): 124.95 Median Serving TTFT (ms): 129.86 P99 Serving TTFT (ms): 134.73 ---------------Time to First Token---------------- Mean TTFT (ms): 121.80 Median TTFT (ms): 126.85 P99 TTFT (ms): 130.36 -----Time per Output Token (excl. 1st token)------ Mean TPOT (ms): 177.17 Median TPOT (ms): 177.40 P99 TPOT (ms): 178.17 ---------------Inter-token Latency---------------- Mean ITL (ms): 118.12 Median ITL (ms): 7.06 P99 ITL (ms): 349.32 ================= Stage 1 (dit) ================== ================== Image Result ================== Total images generated: 3 -----------------Image Generation----------------- Mean IMAGE_GENERATION (ms): 18174.88 Median IMAGE_GENERATION (ms): 18176.98 P99 IMAGE_GENERATION (ms): 18185.23 ==================================================

Result of HunyuanImage-3.0-Instruct

Show more
============ Serving Benchmark Result ============
Successful requests:                     3         
Failed requests:                         0         
Maximum request concurrency:             1         
Benchmark duration (s):                  185.15    
Request throughput (req/s):              0.02      
Peak concurrent requests:                2.00      
----------------End-to-end Latency----------------
Mean E2EL (ms):                          61712.70  
Median E2EL (ms):                        38293.20  
P99 E2EL (ms):                           116730.94 
================== Text Result ===================
Total input tokens:                      7500      
Total generated tokens:                  4718      
Output token throughput (tok/s):         25.48     
Peak output token throughput (tok/s):    21.00     
Peak concurrent requests:                2.00      
Total Token throughput (tok/s):          65.99     
---------------Time to First Token----------------
Mean TTFT (ms):                          723.66    
Median TTFT (ms):                        740.74    
P99 TTFT (ms):                           805.78    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          41.33     
Median TPOT (ms):                        49.49     
P99 TPOT (ms):                           50.03     
---------------Inter-token Latency----------------
Mean ITL (ms):                           49.10     
Median ITL (ms):                         48.59     
P99 ITL (ms):                            53.89     
================== Image Result ==================
Total images generated:                  3         
Image throughput (img/s):                0.02      
Average pixels per image:                1048576.00
Mean denoise step latency (ms):          343.58    
-----------------Image Generation-----------------
Mean IMAGE_GENERATION (ms):              17179.22  
Median IMAGE_GENERATION (ms):            17141.46  
P99 IMAGE_GENERATION (ms):               17274.08  
============= Stage Benchmark Result ============= ================== Stage 0 (AR) ================== -------------------Stage Timing------------------- Mean stage_gen_time (ms): 88400.47 Median stage_gen_time (ms): 41670.15 P99 stage_gen_time (ms): 198207.80 ================== Text Result =================== Stage generated tokens: 5344 -----------Serving Time to First Token------------ Mean Serving TTFT (ms): 696.19 Median Serving TTFT (ms): 712.40 P99 Serving TTFT (ms): 778.56 ---------------Time to First Token---------------- Mean TTFT (ms): 681.39 Median TTFT (ms): 702.00 P99 TTFT (ms): 759.29 -----Time per Output Token (excl. 1st token)------ Mean TPOT (ms): 49.64 Median TPOT (ms): 49.65 P99 TPOT (ms): 50.34 ---------------Inter-token Latency---------------- Mean ITL (ms): 49.15 Median ITL (ms): 48.63 P99 ITL (ms): 52.93 ================= Stage 1 (dit) ================== ================== Image Result ================== Total images generated: 3 -----------------Image Generation----------------- Mean IMAGE_GENERATION (ms): 17179.22 Median IMAGE_GENERATION (ms): 17141.46 P99 IMAGE_GENERATION (ms): 17274.08 ==================================================

Result of HunyuanImage-3.0-Instruct (without --print-stage)

Show more
============ Serving Benchmark Result ============
Successful requests:                     3         
Failed requests:                         0         
Maximum request concurrency:             1         
Benchmark duration (s):                  184.95    
Request throughput (req/s):              0.02      
Peak concurrent requests:                2.00      
----------------End-to-end Latency----------------
Mean E2EL (ms):                          61646.35  
Median E2EL (ms):                        38361.68  
P99 E2EL (ms):                           116797.40 
================== Text Result ===================
Total input tokens:                      7500      
Total generated tokens:                  4718      
Output token throughput (tok/s):         25.51     
Peak output token throughput (tok/s):    21.00     
Peak concurrent requests:                2.00      
Total Token throughput (tok/s):          66.06     
---------------Time to First Token----------------
Mean TTFT (ms):                          717.94    
Median TTFT (ms):                        739.31    
P99 TTFT (ms):                           788.67    
-----Time per Output Token (excl. 1st token)------
Mean TPOT (ms):                          40.89     
Median TPOT (ms):                        48.47     
P99 TPOT (ms):                           49.69     
---------------Inter-token Latency----------------
Mean ITL (ms):                           49.06     
Median ITL (ms):                         49.29     
P99 ITL (ms):                            52.41     
================== Image Result ==================
Total images generated:                  3         
Image throughput (img/s):                0.02      
Average pixels per image:                1048576.00
Mean denoise step latency (ms):          343.10    
-----------------Image Generation-----------------
Mean IMAGE_GENERATION (ms):              17155.07  
Median IMAGE_GENERATION (ms):            17135.07  
P99 IMAGE_GENERATION (ms):               17227.73  
==================================================

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 89def0fff3

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread vllm_omni/benchmarks/patch/patch.py Outdated
Comment on lines +458 to +463
if (
"stage_metrics" not in metrics
and "num_tokens_out" in metrics
and metrics.get("num_tokens_out") is not None
):
output.output_tokens = metrics.get("num_tokens_out", 0)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Keep completion token metrics when stage_metrics is attached

This condition drops num_tokens_out whenever the server includes stage_metrics, but this same commit makes finished responses always include stage_metrics (vllm_omni/metrics/stats.py sets output_to_yield.metrics = {"stage_metrics": ...} before adding token fields). As a result, text responses can leave output.output_tokens at 0 even when num_tokens_out is present, and this benchmark client does not parse usage.completion_tokens, so token-throughput/TPOT calculations become wrong in runs where tokenizer fallback is unavailable or mismatched.

Useful? React with 👍 / 👎.

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

@ZeldaHuang @amy-why-3459

@amy-why-3459

Copy link
Copy Markdown
Collaborator

@fake0fan @LJH-LBJ Please take a look as well.

Comment thread vllm_omni/benchmarks/metrics/metrics.py Outdated
to_ms=True,
)
print("{s:{c}^{n}}".format(s="Time to First Output", n=50, c="-"))
print(f"{'Mean TTFO (ms):':<40} {mean_v:<10.2f}")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does TTFO mean? Is there any documentation explaining it?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does TTFO mean? Is there any documentation explaining it?

The "Time to First Output" (TTFO) is the same concept as TTFT ("Time to First Token"), but just the "Token" is replaced by "Output" for multimodal cases, and the timestamp of TTFO is recorded inside StagePool.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does TTFO mean? Is there any documentation explaining it?

The "Time to First Output" (TTFO) is the same concept as TTFT ("Time to First Token"), but just the "Token" is replaced by "Output" for multimodal cases, and the timestamp of TTFO is recorded inside StagePool.

Can it be changed to TTFC? Talker TTFC (Time-To-First-Chunk)

Comment thread vllm_omni/benchmarks/metrics/metrics.py Outdated
to_ms=True,
)
print("{s:{c}^{n}}".format(s="Time per Output Unit (excl. 1st output)", n=50, c="-"))
print(f"{'Mean TPOU (ms):':<40} {mean_v:<10.2f}")

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Should we continue using TPOT and ITL?

@NumberWan

Copy link
Copy Markdown
Contributor

As discussed, alignment for time origins:
Proposal (client / serving SLO view)
For TTFC, measure from HTTP request accepted (benchmark: timestamp at/just before POST) until the first streamed unit the client can observe (first text/token chunk, first ar_delta, first audio frame, etc.) — not from AR/LLM stage entry inside StagePool.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label May 28, 2026
@ZacheryAU
ZacheryAU force-pushed the fix_metric branch 5 times, most recently from dc8aa6f to 85a73de Compare May 29, 2026 07:50
ZacheryAU added 3 commits June 4, 2026 04:20
…tions.py and apply it to BenchmarkMetrics).

Signed-off-by: ZacheryAU <zachery.au@gmail.com>
…PRINT_STAGE).

Signed-off-by: ZacheryAU <zachery.au@gmail.com>
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
@hsliuustc0106 hsliuustc0106 added omni-test and removed ready label to trigger buildkite CI labels Jun 4, 2026
…uested).

Signed-off-by: ZacheryAU <zachery.au@gmail.com>
…uested).

Signed-off-by: ZacheryAU <zachery.au@gmail.com>
@Gaohan123 Gaohan123 added the ready label to trigger buildkite CI label Jun 4, 2026
…ted doc).

Signed-off-by: ZacheryAU <zachery.au@gmail.com>

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

please update the docs in a follow-up PR, check https://docs.vllm.ai/en/latest/benchmarking/cli/ for a reference

@Gaohan123
Gaohan123 merged commit bfc7ced into vllm-project:main Jun 4, 2026
8 checks passed
akshatvishu pushed a commit to akshatvishu/vllm-omni that referenced this pull request Jun 13, 2026
…#3628)

Signed-off-by: ZacheryAU <70869627+ZacheryAU@users.noreply.github.com>
Signed-off-by: akshatvishu <akshatnayak197@gmail.com>
Nughm3 pushed a commit to Nughm3/vllm-omni that referenced this pull request Jun 18, 2026
…#3628)

Signed-off-by: ZacheryAU <70869627+ZacheryAU@users.noreply.github.com>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…#3628)

Signed-off-by: ZacheryAU <70869627+ZacheryAU@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

high priority high priority issue, needs to be done asap ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants