Repository navigation
[CI][Perf] Add Wan22 i2v perf nightly ci - #3063
Conversation
Signed-off-by: bjf-frz <frz123db@gmail.com>
Signed-off-by: bjf-frz <frz123db@gmail.com>
Signed-off-by: bjf-frz <frz123db@gmail.com>
Signed-off-by: bjf-frz <frz123db@gmail.com>
|
@yenuo26 PTAL, thx. |
hsliuustc0106
left a comment
There was a problem hiding this comment.
Good CI/Perf PR. The separation of server_type from benchmark_backend is a clean design that enables flexible benchmarking across different serving APIs.
Minor suggestions:
- PR description test results section is empty — consider adding baseline numbers or a note that this is the first run
- The new benchmark_backend field in test configs is a schema change — consider documenting this in test style docs or adding a comment in the example JSON
No blocking issues.
Signed-off-by: bjf-frz <frz123db@gmail.com>
|
@hsliuustc0106 @Gaohan123 @david6666666 @yenuo26 perf CI has passed, PTAL, thx. |
Signed-off-by: bjf-frz <frz123db@gmail.com>
|
paster your perf test results |
|
|
||
| Usage: | ||
| # Video (vllm-omni backend) | ||
| # Video (v1/videos backend) |
There was a problem hiding this comment.
v1/videos endpoint? this is not a backend
There was a problem hiding this comment.
v1/videos endpoint? this is not a backend
In the benchmark script, this is named as backend, the mapping relations look like:
backends_function_mapping = {
"2i": {
"vllm-omni": (async_request_chat_completions, "/v1/chat/completions"),
"openai": (async_request_openai_images, "/v1/images/generations"),
},
"2v": {
"v1/videos": (async_request_v1_videos, "/v1/videos"),
},
}
| i2v: | ||
| python3 benchmarks/diffusion/diffusion_benchmark_serving.py \ | ||
| --backend vllm-omni --dataset vbench --task i2v --num-prompts 10 | ||
| --backend v1/videos --dataset vbench --task i2v --num-prompts 10 |
There was a problem hiding this comment.
--task i2v corresponds to the v1/videos backend, it didn’t match before.
| If you want to use i2v, i2i dataset, you should `uv pip install gdown` first | ||
|
|
||
| Supports multiple backends: | ||
| - vllm-omni: Uses /v1/chat/completions endpoint (default) |
There was a problem hiding this comment.
chat_completion/image_generation/videos/image_edit
|
follow-up PR: change backend to endpoint |
Signed-off-by: bjf-frz <frz123db@gmail.com>
Signed-off-by: bjf-frz <frz123db@gmail.com>
Signed-off-by: bjf-frz <frz123db@gmail.com>
Owner policy: vLLM-Omni's dependency versions follow vLLM's. Upstream declares "transformers >= 5.10.4" (vllm/requirements/common.txt, unchanged between the v0.29.0 baseline 98dff2a81d74 and the pinned f2aad6aa7074). The divergence that mattered was the upper bound, not the floor, so this removes the cap and leaves "transformers >= 5.13.0". That satisfies the policy exactly. ">= 5.13.0" is a subset of ">= 5.10.4": every version it admits also satisfies upstream, and a clean install or CI resolves the same newest release vLLM resolves. The cap was the only thing that could select a different version from vLLM's, which is the behaviour worth eliminating. The cap was also inert where it mattered. docker/Dockerfile.ci installs the project and only then force-reinstalls the pinned vLLM commit wheel, which resolves vLLM's own closure. In scheduled build #3062 the first step applied the cap (transformers 5.16.1 -> 5.14.1) and the wheel step overrode it immediately (5.14.1 -> 5.17.0), so the image under test ran 5.17.0. A cap here can only downgrade an intermediate layer. It had also gone unexercised: main's #3063 shows its dependency layer CACHED with zero installs, so main's CI image carries 5.16.1 while main's file declares "< 5.15". The only reason #3062 re-resolved is that 6fc05d9 changed Dockerfile.ci and invalidated the cache. The 5.13.0 floor stays, and it is not an arbitrary divergence: it is induced by the kernels==0.16.1 pin three lines below. kernels>=0.16 needs transformers.integrations.hub_kernels LayerRepository entries to carry version=/revision=, which arrives in 5.13 (#6971). Verified that nothing else enforces the pairing: the kernels 0.16.1 wheel declares no transformers requirement at all. Dropping the floor to 5.10.4 would let pip keep an already-installed 5.10-5.12 in an existing environment, since it satisfies the weaker bound, and pair it with kernels 0.16.1. Clean installs and CI would be unaffected; existing developer environments would not. Two review iterations raised exactly this, and it is the right call. Lower the floor only together with the kernels pin. This supersedes ad6a268, which restored main's ">= 5.13.0, < 5.15" on the grounds that main declares it; matching main was the wrong target, because main's cap is itself a divergence from vLLM. Checked while here: of the packages constrained in both requirements files, einops, pydantic, safetensors and tqdm carry omni floors differing from vLLM's. None is an upper bound, so none can resolve below what vLLM asks for -- pydantic is the clearest, where omni's >= 2.1.0 is stale against vLLM's >= 2.12.0 and the stricter bound wins. Left alone: pre-existing, not rebase work. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Signed-off-by: bjf-frz <frz123db@gmail.com>
PLEASE FILL IN THE PR DESCRIPTION HERE ENSURING ALL CHECKLIST ITEMS (AT THE BOTTOM) HAVE BEEN CONSIDERED.
Purpose
This PR aims to add nightly performance coverage for Wan2.2 I2V and align diffusion benchmark backend resolution with the correct image/video serving APIs.
Test Plan
If you want to test the perf locally, run the command like:
Test Result
Essential Elements of an Effective PR Description Checklist
supported_models.mdandexamplesfor a new model. Please runmkdocs serveto sync the documentation editions to./docs.BEFORE SUBMITTING, PLEASE READ https://github.com/vllm-project/vllm-omni/blob/main/CONTRIBUTING.md (anything written below this line will be removed by GitHub Actions)