Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
23 commits
Select commit Hold shift + click to select a range
b78bca5
Add lingbot perf test
yenuo26 Sep 16, 2026
e974cc2
Refactor benchmark scripts to unify naming conventions and improve cl…
yenuo26 Sep 17, 2026
7c92ee7
Add metrics aggregation and printing for stage durations in benchmarks
yenuo26 Sep 17, 2026
5c3ca58
Update performance test configurations to use GPT-2 tokenizer and red…
yenuo26 Sep 17, 2026
8975184
Enhance stage duration metrics display in benchmarks
yenuo26 Sep 17, 2026
ddc582b
Refactor benchmark scripts to unify naming conventions and improve cl…
yenuo26 Sep 17, 2026
5bf7701
Update performance benchmark scripts and configurations for diffusion…
yenuo26 Sep 18, 2026
d080c9c
Enhance stage duration metrics handling in benchmarks
yenuo26 Sep 19, 2026
dac9015
Merge upstream/main into perf
yenuo26 Sep 19, 2026
3b325aa
Merge branch 'main' into perf
hsliuustc0106 Sep 19, 2026
6b379c1
Add video reference handling to benchmarks
yenuo26 Sep 19, 2026
3e4f1cf
Merge branch 'perf' of https://github.com/yenuo26/vllm-omni into perf
yenuo26 Sep 19, 2026
bf9e36b
Refactor video reference handling in benchmarks
yenuo26 Sep 20, 2026
8bedd93
debug for artifact_paths
yenuo26 Sep 20, 2026
1d81860
remove debug
yenuo26 Sep 20, 2026
71db77b
Enhance benchmark parameter handling and image/video reference proces…
yenuo26 Sep 20, 2026
12f9d7f
Merge branch 'main' into perf
amy-why-3459 Sep 20, 2026
d1b54f1
Refactor video reference handling and enhance test coverage
yenuo26 Sep 21, 2026
9371df2
Merge remote-tracking branch 'origin/perf' into perf
yenuo26 Sep 21, 2026
c1b1ddb
Merge remote-tracking branch 'upstream/main' into perf
yenuo26 Sep 21, 2026
fdaef00
Enhance video and image reference handling in benchmarks
yenuo26 Sep 21, 2026
0b89729
Merge branch 'main' into perf
Gaohan123 Sep 21, 2026
9f5cd26
Merge branch 'main' into perf
yenuo26 Sep 22, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
13 changes: 9 additions & 4 deletions .buildkite/common/ci_source_file_dependencies.yml
Original file line number Diff line number Diff line change
Expand Up @@ -317,7 +317,7 @@ source_file_dependencies:

diffusion_wan22_perf:
- *wan22
- tests/dfx/perf/scripts/run_diffusion_benchmark.py
- tests/dfx/perf/scripts/run_benchmark.py
- tests/dfx/perf/tests/test_wan22_i2v_vllm_omni.json

diffusion_wan22_reliability:
Expand All @@ -343,7 +343,7 @@ source_file_dependencies:

diffusion_cosmos3_perf:
- *cosmos3
- tests/dfx/perf/scripts/run_diffusion_benchmark.py
- tests/dfx/perf/scripts/run_benchmark.py
- tests/dfx/perf/tests/test_cosmos3_vllm_omni.json

diffusion_qwen_image_function:
Expand Down Expand Up @@ -416,7 +416,7 @@ source_file_dependencies:

diffusion_minimax_h3_perf:
- *minimax_h3
- tests/dfx/perf/scripts/run_diffusion_benchmark.py
- tests/dfx/perf/scripts/run_benchmark.py
- tests/dfx/perf/tests/test_minimax_h3_vllm_omni.json

diffusion_hunyuan_video_function:
Expand All @@ -430,7 +430,7 @@ source_file_dependencies:

diffusion_hunyuan_video_perf:
- *hunyuan_video
- tests/dfx/perf/scripts/run_diffusion_benchmark.py
- tests/dfx/perf/scripts/run_benchmark.py
- tests/dfx/perf/tests/test_hunyuanvideo15_t2v_vllm_omni.json
- tests/dfx/perf/tests/test_hunyuanvideo15_i2v_vllm_omni.json

Expand All @@ -457,6 +457,11 @@ source_file_dependencies:
- tests/e2e/online_serving/test_lingbot_video_moe.py
- tests/e2e/offline_inference/test_lingbot_world_v2.py

diffusion_lingbot_perf:
- *lingbot_video
- tests/dfx/perf/scripts/run_benchmark.py
- tests/dfx/perf/tests/test_lingbot_video_vllm_omni.json


# Invalid-param jobs: model business code only, plus the reliability scripts.
reliability_invalid_param_h100:
Expand Down
153 changes: 63 additions & 90 deletions .buildkite/cuda/test-nightly.yml

Large diffs are not rendered by default.

30 changes: 11 additions & 19 deletions .buildkite/cuda/test-weekly.yml
Original file line number Diff line number Diff line change
Expand Up @@ -12,14 +12,15 @@ steps:
depends_on: upload-weekly-pipeline
if: build.env("WEEKLY") == "1"
timeout_in_minutes: 60
artifact_paths:
- coverage-core-model-cpu-*.xml.gz
commands:
- |
set +e
REPORT="coverage-core-model-cpu-$${BUILDKITE_STEP_ID:-local}.xml"
pytest -sv tests/ -m 'core_model and cpu' --cov=vllm_omni --cov-report=term-missing:skip-covered --cov-report=xml:$${REPORT}
EXIT=$$?
gzip -9 -f "$${REPORT}"
buildkite-agent artifact upload "$${REPORT}.gz"
exit $$EXIT
mirror_hardwares: l4_1

Expand Down Expand Up @@ -79,41 +80,32 @@ steps:
source_file_dependencies: omni_qwen3_omni_perf
key: weekly-omni-performance-vllm-text
timeout_in_minutes: 180
artifact_paths:
- tests/dfx/perf/results/*.json
commands:
- export BENCHMARK_DIR=tests/dfx/perf/results
- |
set +e
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py --test-config-file tests/dfx/perf/tests/test_qwen3_omni_vllm_text.json
EXIT=$$?
buildkite-agent artifact upload "tests/dfx/perf/results/*.json"
exit $$EXIT
- pytest -s -v tests/dfx/perf/scripts/run_benchmark.py --test-config-file tests/dfx/perf/tests/test_qwen3_omni_vllm_text.json
mirror_hardwares: h100_2

- label: ":full_moon: Omni · Perf Test · Async Chunk · Random"
key: weekly-omni-performance-async-chunk-random
timeout_in_minutes: 180
artifact_paths:
- tests/dfx/perf/results/*.json
commands:
- export BENCHMARK_DIR=tests/dfx/perf/results
- |
set +e
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py --test-config-file tests/dfx/perf/tests/test_qwen3_omni_async_chunk.json -m "H100 and slow"
EXIT=$$?
buildkite-agent artifact upload "tests/dfx/perf/results/*.json"
exit $$EXIT
- pytest -s -v tests/dfx/perf/scripts/run_benchmark.py --test-config-file tests/dfx/perf/tests/test_qwen3_omni_async_chunk.json -m "H100 and slow"
mirror_hardwares: h100_2

- label: ":full_moon: Omni · Perf Test · Multi-Replica"
source_file_dependencies: omni_qwen3_omni_perf
key: weekly-omni-performance-multi-replicas
timeout_in_minutes: 180
artifact_paths:
- tests/dfx/perf/results/*.json
commands:
- export BENCHMARK_DIR=tests/dfx/perf/results
- |
set +e
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py --test-config-file tests/dfx/perf/tests/test_qwen3_omni_multi_replicas.json
EXIT=$$?
buildkite-agent artifact upload "tests/dfx/perf/results/*.json"
exit $$EXIT
- pytest -s -v tests/dfx/perf/scripts/run_benchmark.py --test-config-file tests/dfx/perf/tests/test_qwen3_omni_multi_replicas.json
mirror_hardwares: h100_3

- group: ":card_index_dividers: E2E Tests"
Expand Down
7 changes: 3 additions & 4 deletions .buildkite/npu/test-npu-nightly.yml
Original file line number Diff line number Diff line change
Expand Up @@ -286,11 +286,10 @@ steps:
HF_TOKEN: "${HF_TOKEN}"
DIFFUSION_ATTENTION_BACKEND: TORCH_SDPA
commands:
- export DIFFUSION_BENCHMARK_DIR=tests/dfx/perf/results
- export BENCHMARK_DIR=tests/dfx/perf/results
- |
set +e
pytest -s -v tests/dfx/perf/scripts/run_diffusion_benchmark.py -m npu --test-config-file tests/dfx/perf/tests/test_wan22_i2v_vllm_omni.json
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py -m npu --test-config-file tests/dfx/perf/tests/test_wan22_i2v_vllm_omni.json
EXIT=$$?
buildkite-agent artifact upload "tests/dfx/perf/results/diffusion_result_*.json"
buildkite-agent artifact upload "tests/dfx/perf/results/logs/*.log"
buildkite-agent artifact upload "tests/dfx/perf/results/*.json"
exit $$EXIT
52 changes: 32 additions & 20 deletions docs/contributing/ci/test_examples/l4_performance_tests.inc.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,16 +7,17 @@ When you want to add L4-level ***performance test*** cases, add entries to JSON
| Omni (nightly) | `run_benchmark.py` | `test_qwen3_omni_no_async_chunk.json`, `test_qwen3_omni_async_chunk.json` (`full_model` without `slow` in `mark`) |
| Omni (weekly) | `run_benchmark.py` | `test_qwen3_omni_async_chunk.json` (CUDA only), `test_qwen3_omni_vllm_text.json`, `test_qwen3_omni_multi_replicas.json` (`slow` in `mark`; **Perf Test** in `test-weekly.yml`) |
| TTS | `run_benchmark.py` | `test_tts.json`, `test_voxcpm2.json`, `test_higgs_audio_v3.json` |
| Diffusion | `run_diffusion_benchmark.py` | `test_qwen_image_vllm_omni.json`, `test_bagel_vllm_omni.json`, `test_wan22_i2v_vllm_omni.json`, `test_cosmos3_vllm_omni.json`, … |
| Diffusion (`/v1/chat/completions`) | `run_diffusion_benchmark.py` | `test_qwen_image_vllm_omni.json`, `test_bagel_vllm_omni.json`, … |
| Diffusion (`/v1/images/generations`, `/v1/images/edits`, `/v1/videos`) | `run_benchmark.py` | `test_wan22_i2v_vllm_omni.json`, `test_cosmos3_vllm_omni.json`, `test_lingbot_video_vllm_omni.json`, … |

#### How runners pick cases

Without **`--test-config-file`**, each runner scans all `*.json` under `tests/dfx/perf/tests/` but only keeps its own model type:

- **`run_benchmark.py`**: omni and TTS cases only (skips diffusion JSON).
- **`run_diffusion_benchmark.py`**: diffusion cases only (skips omni / TTS JSON).
- **`run_benchmark.py`**: cases whose `benchmark_params` use ``dataset_name`` (omni, TTS, and image/video OpenAI generation via `vllm bench serve --omni`).
- **`run_diffusion_benchmark.py`**: cases whose `benchmark_params` use ``dataset`` (diffusion client schema, including custom jsonl).

Diffusion cases are detected when the JSON has `server_type` (typically `"vllm-omni"`) or `"diffusion"` in the `mark` array. Omni / TTS JSON has neither.
Diffusion-script cases are detected by `is_diffusion_perf_config()`: presence of ``dataset`` (without ``dataset_name``) in `benchmark_params`. Marks / `server_type` / endpoint values are not used for this split.

#### Running perf cases

Expand Down Expand Up @@ -60,16 +61,16 @@ Pass **`--test-config-file`** to load one JSON file, or omit it for the bulk sca

##### Overview

| Field | Required | Description |
| ------------------ | -------------- | -------------------------------------------- |
| test_name | Yes | Unique identifier for the test case |
| mark | No | Pytest marks; see **`mark` field** below |
| server_params | Yes | Server-side configuration parameters |
| benchmark_params | Yes | Benchmark running parameters |
| server_type | Diffusion only | Routes case to run_diffusion_benchmark.py |
| benchmark_endpoint | Diffusion only | Benchmark API path |
| Field | Required | Description |
| ------------------ | --------- | ---------------------------------------------------------------------------------------- |
| test_name | Yes | Unique identifier for the test case |
| mark | No | Pytest marks; see **`mark` field** below |
| server_params | Yes | Server-side configuration parameters |
| benchmark_params | Yes | Benchmark running parameters |
| server_type | Diffusion | Only for diffusion-script JSON; omit on omni-bench generation cases |
| benchmark_endpoint | Optional | Legacy diffusion custom-jsonl alias; prefer `benchmark_params[].endpoint` |

Omit `mark` only for configs not meant to be filtered by `-m`. `server_type` is typically `"vllm-omni"`. `benchmark_endpoint` examples: `/v1/videos`, `/v1/images/generations`.
Omit `mark` only for configs not meant to be filtered by `-m`. Cases that call `/v1/images/edits`, `/v1/images/generations`, or `/v1/videos` use the same `benchmark_params` schema as Omni/TTS (`dataset_name`, `endpoint`, `extra_body`) and are executed by `run_benchmark.py` — do not set `server_type` or `task` on those cases. Remaining diffusion cases (usually `/v1/chat/completions`, or custom jsonl) stay on `run_diffusion_benchmark.py` and may keep `server_type`.

#### `mark` field

Expand Down Expand Up @@ -100,7 +101,7 @@ Recommended for L4 perf cases:
}
```

Multi-GPU diffusion (example: Cosmos3 with `cfg-parallel-size=2`):
Multi-GPU generation via omni bench (example: Cosmos3 with `cfg-parallel-size=2`):

```JSON
{
Expand All @@ -110,10 +111,17 @@ Multi-GPU diffusion (example: Cosmos3 with `cfg-parallel-size=2`):
"full_model",
"diffusion"
],
"server_type": "vllm-omni",
"benchmark_endpoint": "/v1/images/generations",
"server_params": { "...": "..." },
"benchmark_params": [ { "name": "1024x1024_steps4", "...": "..." } ]
"benchmark_params": [
{
"name": "1024x1024_steps4",
"dataset_name": "random",
"endpoint": "/v1/images/generations",
"num_prompts": 3,
"max_concurrency": 1,
"extra_body": { "width": 1024, "height": 1024, "num_inference_steps": 4 }
}
]
}
```

Expand All @@ -132,8 +140,8 @@ Result files use the **runtime** hardware label from `get_runtime_resource_label

Examples:

- Omni/TTS: `result_{test_name}_{optional_hw}_{dataset}_....json` under `BENCHMARK_DIR`
- Diffusion: one aggregate `diffusion_result_{config_stem}_{optional_hw}_{timestamp}.json` per source JSON file (array of all runs from that file)
- Omni/TTS and OpenAI generation endpoints (`/v1/images/*`, `/v1/videos`): `result_{test_name}_{optional_hw}_{dataset}_....json` under `BENCHMARK_DIR`
- Remaining diffusion (`run_diffusion_benchmark.py`): one aggregate `diffusion_result_{config_stem}_{optional_hw}_{timestamp}.json` per source JSON file

#### Local commands

Expand All @@ -145,6 +153,9 @@ pytest -s -v tests/dfx/perf/scripts/run_benchmark.py -m "full_model and omni and
# Single file (same selectors as the CI Perf steps)
pytest -s -v tests/dfx/perf/scripts/run_diffusion_benchmark.py \
--test-config-file tests/dfx/perf/tests/test_bagel_vllm_omni.json
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py \
--test-config-file tests/dfx/perf/tests/test_cosmos3_vllm_omni.json \
-m "H100 and B200 and cards_2"
pytest -s -v tests/dfx/perf/scripts/run_benchmark.py \
--test-config-file tests/dfx/perf/tests/test_qwen3_omni_async_chunk.json \
-m "H100 and full_model and not slow"
Expand Down Expand Up @@ -203,7 +214,8 @@ You can add any benchmark running parameters you need here. For all optional par
2. For boolean variables in the running parameters, modify them to forms such as ignore_eos: true/false and fill them into the JSON file.
3. Optionally add a `baseline` object (see **Baseline thresholds** below). If you omit `baseline` or leave it empty, the performance test still runs but does not assert metric thresholds from this field.
4. Set `"name"` on each `benchmark_params` entry for stable pytest ids and readable result keys.
5. The qps and concurrency modes are recommended to be mutually exclusive. For detailed explanations, see the table below:
5. Image/video generation cases (`/v1/images/generations`, `/v1/images/edits`, `/v1/videos`) use this same schema: set `endpoint` to the API path, put width/height/steps/frames in `extra_body`, use `dataset_name: random` for text-only inputs or `random-mm` when the request needs a synthetic image/video. Do not set `server_type` or `task`, and do not use `random-request-config` or kebab-case diffusion client fields.
6. The qps and concurrency modes are recommended to be mutually exclusive. For detailed explanations, see the table below:

| Parameter | Type | Required | Example/Values | Description |
| --------------- | ------------- | -------- | -------------------- | ------------------------------------ |
Expand Down
2 changes: 1 addition & 1 deletion docs/contributing/ci/test_execution_guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -165,7 +165,7 @@ Failed jobs: 1/2
pytest -sv tests/dfx/perf/scripts/run_benchmark.py -m "full_model and tts and H100"
pytest -sv tests/dfx/perf/scripts/run_diffusion_benchmark.py -m "full_model and diffusion and H100"
pytest -sv tests/dfx/perf/scripts/run_benchmark.py --test-config-file tests/dfx/perf/tests/test_tts.json
pytest -sv tests/dfx/perf/scripts/run_diffusion_benchmark.py --test-config-file tests/dfx/perf/tests/test_cosmos3_vllm_omni.json
pytest -sv tests/dfx/perf/scripts/run_benchmark.py --test-config-file tests/dfx/perf/tests/test_cosmos3_vllm_omni.json
```
Nightly **Perf Test** jobs in [``test-nightly.yml``](https://github.com/vllm-project/vllm-omni/blob/main/.buildkite/cuda/test-nightly.yml) use ``--test-config-file`` only (no ``-m``). Weekly **Perf Test** in [``test-weekly.yml``](https://github.com/vllm-project/vllm-omni/blob/main/.buildkite/cuda/test-weekly.yml) runs ``test_qwen3_omni_vllm_text.json`` and ``test_qwen3_omni_multi_replicas.json`` (JSON ``mark`` includes ``slow``). E2e L4 function tests use ``full_model`` + ``--run-level full_model``. Example:

Expand Down
3 changes: 2 additions & 1 deletion docs/contributing/ci/test_writing_guide.md
Original file line number Diff line number Diff line change
Expand Up @@ -159,7 +159,8 @@ When `mark` is present, it must be an **array** with exactly one ``hardware_mark
}
```

- Local bulk load: `pytest -sv tests/dfx/perf/scripts/run_diffusion_benchmark.py -m "full_model and diffusion and H100"`
- Local bulk load: `pytest -sv tests/dfx/perf/scripts/run_benchmark.py -m "full_model and H100"` (omni/TTS and `/v1/images/*` + `/v1/videos` diffusion)
- Diffusion chat-completions remaining cases: `pytest -sv tests/dfx/perf/scripts/run_diffusion_benchmark.py -m "full_model and diffusion and H100"`
- Nightly CI perf steps: `--test-config-file tests/dfx/perf/tests/test_<model>_vllm_omni.json` (file selects cases; no `-m`)
- Result filenames use **runtime** GPU detection (`get_runtime_resource_label`); `H100` is omitted on the default CI pool

Expand Down
18 changes: 9 additions & 9 deletions recipes/LTX/LTX-2.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@
## Pipelines

| `--model-class-name` | Task | Required checkpoint repositories |
|---|---|---|
| --- | --- | --- |
| `LTX2Pipeline` | LTX-2 one-stage T2V/I2V | `Lightricks/LTX-2` |
| `LTX2TwoStagePipeline` | LTX-2 ordinary two-stage T2V/I2V | `Lightricks/LTX-2` |
| `LTX2DistilledOneStagePipeline` | LTX-2 merged-distilled one-stage T2V/I2V | `rootonchair/LTX-2-19b-distilled` |
Expand Down Expand Up @@ -53,7 +53,7 @@ pipe(req, image=image, prompt=prompt)
The consolidation also removes these registry names without aliases:

| Removed name | Replacement |
|---|---|
| --- | --- |
| `LTX23Pipeline` | `LTX2Pipeline`; checkpoint metadata selects LTX-2.3 |
| `LTX2ImageToVideoPipeline` | `LTX2Pipeline` with `image=` |
| `LTX23ImageToVideoPipeline` | `LTX2Pipeline` with `image=`; checkpoint metadata selects LTX-2.3 |
Expand All @@ -67,7 +67,7 @@ offline and serving entrypoints already use named fields and are unaffected.
## One-Stage Defaults

| Parameter | LTX-2 | LTX-2.3 |
|---|---:|---:|
| --- | ---: | ---: |
| Width × height | 768 × 512 | 768 × 512 |
| Frames / frame rate | 121 / 24 | 121 / 24 |
| Denoise steps | 40 | 30 |
Expand All @@ -85,7 +85,7 @@ default to `121`.
## Two-Stage Defaults

| Parameter | Ordinary | Full-distilled |
|---|---:|---:|
| --- | ---: | ---: |
| Final width × height | 1536 × 1024 | 1536 × 1024 |
| Stage 1 width × height | 768 × 512 | 768 × 512 |
| Frames / frame rate | 121 / 24 | 121 / 24 |
Expand Down Expand Up @@ -172,7 +172,7 @@ spatio-temporal guidance (STG), cross-modality guidance, and rescaling.
Distilled stages and ordinary Stage 2 are fixed positive-only.

| Parameter | Default | Effect | Alias |
|---|---:|---|---|
| --- | ---: | --- | --- |
| `video_cfg_scale` | 3.0 | Video text CFG; `1.0` disables it | `video_cfg_guidance_scale` |
| `audio_cfg_scale` | 7.0 | Audio text CFG; `1.0` disables it | `audio_cfg_guidance_scale` |
| `video_stg_scale` | 1.0 | Video STG; `0.0` disables it | `video_stg_guidance_scale` |
Expand Down Expand Up @@ -210,7 +210,7 @@ per denoise step: `cond`, `uncond`, `ptb` (STG), and `mod`
(cross-modality). The useful balanced configurations are therefore:

| `--cfg-parallel-size` | Passes per rank | Guidance-slot utilization | Notes |
|---:|---:|---:|---|
| ---: | ---: | ---: | --- |
| `1` | 4 | 100% | Single-rank fused guidance batch |
| `2` | 2 | 100% | Recommended two-rank configuration |
| `4` | 1 | 100% | One guidance pass per rank |
Expand Down Expand Up @@ -269,7 +269,7 @@ noted below.
### Complete `forward` Surface

| Argument | Type/default | Meaning and constraints |
|---|---|---|
| --- | --- | --- |
| `req` | `DiffusionRequestBatch`, required | Only positional argument; contains prompts and per-request sampling parameters. |
| `image` | image or batch, `None` | Direct value wins over request images; no image selects T2V. I2V accepts one image per prompt, and a batch cannot mix T2V/I2V. |
| `prompt` | string or list, `None` | Positive-text fallback; request prompts win. Mutually exclusive with `prompt_embeds`. |
Expand Down Expand Up @@ -306,7 +306,7 @@ request prompt payload; LTX guidance fields live in sampling `extra_args`.
### Recipe-Specific Request Capabilities

| Override | One-stage | Ordinary two-stage | Distilled two-stage |
|---|---|---|---|
| --- | --- | --- | --- |
| Guidance | Supported | Stage 1 only; Stage 2 is positive-only | Fixed positive-only |
| Negative prompt/embeddings | Supported | Supported by Stage 1 | Rejected |
| `num_inference_steps` | Supported | Controls Stage 1; Stage 2 uses 3 | Fixed at 8 for Stage 1; Stage 2 uses 3 |
Expand Down Expand Up @@ -357,7 +357,7 @@ bundled offline CLI do not currently expose `sigmas`.
- The output audio sample rate comes from the loaded components and is not a
request parameter.
- For benchmarks, use `tests/dfx/perf/tests/test_ltx2_vllm_omni.json` with
`tests/dfx/perf/scripts/run_diffusion_benchmark.py`.
`tests/dfx/perf/scripts/run_benchmark.py`.
- Ordinary and full-distilled two-stage T2V/I2V are supported; HQ execution
remains out of scope.

Expand Down
Loading
Loading