Repository navigation
Benchmark data statistics for each stage of omni models - #3628
Conversation
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 89def0fff3
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
| if ( | ||
| "stage_metrics" not in metrics | ||
| and "num_tokens_out" in metrics | ||
| and metrics.get("num_tokens_out") is not None | ||
| ): | ||
| output.output_tokens = metrics.get("num_tokens_out", 0) |
There was a problem hiding this comment.
Keep completion token metrics when stage_metrics is attached
This condition drops num_tokens_out whenever the server includes stage_metrics, but this same commit makes finished responses always include stage_metrics (vllm_omni/metrics/stats.py sets output_to_yield.metrics = {"stage_metrics": ...} before adding token fields). As a result, text responses can leave output.output_tokens at 0 even when num_tokens_out is present, and this benchmark client does not parse usage.completion_tokens, so token-throughput/TPOT calculations become wrong in runs where tokenizer fallback is unavailable or mismatched.
Useful? React with 👍 / 👎.
| to_ms=True, | ||
| ) | ||
| print("{s:{c}^{n}}".format(s="Time to First Output", n=50, c="-")) | ||
| print(f"{'Mean TTFO (ms):':<40} {mean_v:<10.2f}") |
There was a problem hiding this comment.
What does TTFO mean? Is there any documentation explaining it?
There was a problem hiding this comment.
What does TTFO mean? Is there any documentation explaining it?
The "Time to First Output" (TTFO) is the same concept as TTFT ("Time to First Token"), but just the "Token" is replaced by "Output" for multimodal cases, and the timestamp of TTFO is recorded inside StagePool.
There was a problem hiding this comment.
What does TTFO mean? Is there any documentation explaining it?
The "Time to First Output" (TTFO) is the same concept as TTFT ("Time to First Token"), but just the "Token" is replaced by "Output" for multimodal cases, and the timestamp of TTFO is recorded inside StagePool.
Can it be changed to TTFC? Talker TTFC (Time-To-First-Chunk)
| to_ms=True, | ||
| ) | ||
| print("{s:{c}^{n}}".format(s="Time per Output Unit (excl. 1st output)", n=50, c="-")) | ||
| print(f"{'Mean TPOU (ms):':<40} {mean_v:<10.2f}") |
There was a problem hiding this comment.
Should we continue using TPOT and ITL?
|
As discussed, alignment for time origins: |
dc8aa6f to
85a73de
Compare
…tions.py and apply it to BenchmarkMetrics). Signed-off-by: ZacheryAU <zachery.au@gmail.com>
…PRINT_STAGE). Signed-off-by: ZacheryAU <zachery.au@gmail.com>
Signed-off-by: ZacheryAU <zachery.au@gmail.com>
…uested). Signed-off-by: ZacheryAU <zachery.au@gmail.com>
…uested). Signed-off-by: ZacheryAU <zachery.au@gmail.com>
…ted doc). Signed-off-by: ZacheryAU <zachery.au@gmail.com>
hsliuustc0106
left a comment
There was a problem hiding this comment.
please update the docs in a follow-up PR, check https://docs.vllm.ai/en/latest/benchmarking/cli/ for a reference
…#3628) Signed-off-by: ZacheryAU <70869627+ZacheryAU@users.noreply.github.com> Signed-off-by: akshatvishu <akshatnayak197@gmail.com>
…#3628) Signed-off-by: ZacheryAU <70869627+ZacheryAU@users.noreply.github.com>
…#3628) Signed-off-by: ZacheryAU <70869627+ZacheryAU@users.noreply.github.com>
Purpose
implement: #3967 #1361 #3593 with referring #3545
Summary
--print-stageto control whether or not to print stage metrics (default: false)--percentile_metricsto adjust the printing inside stage level:/v1/images/editsServing Benchmark Result(seeResult of HunyuanImage-3.0-Instruct)Test Plan
vllm bench serve --omni --endpoint /v1/chat/completions --print-stagewith two different types of multi-modal models (audio & image) and see if metrics within each of stages are shownvllm bench serve --omni --endpoint /v1/imgaes/edits --print-stagefor HunyuanImage-3.0-Instruct--print-stageas regression test to see whether onlyServing Benchmark Resultis shownTest Result
Result of Qwen3-Omni-30B-A3B-Instruct
Show more
Result of BAGEL-7B-MoT
Show more
Result of HunyuanImage-3.0-Instruct
Show more
Result of HunyuanImage-3.0-Instruct (without --print-stage)
Show more