Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
74 changes: 74 additions & 0 deletions apps/ComfyUI-vLLM-Omni/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -139,6 +139,80 @@ You can configure per-stage sampling parameters for multi-stage models.
>
> Do not use `frame` and `references` together. Task routing is automatic from which inputs you connect.

For MiniMax-H3 Ref2VA, **Video References** accepts up to 9 images (`image_1`–`image_9`),
3 videos (`video_1`–`video_3`), and 3 audio clips (`audio_1`–`audio_3`), with at most
12 connected inputs in total. Any mixture containing at least one image or video
is supported; empty and audio-only references are rejected. For example, you can
combine 6 images, 3 videos, and 3 audio clips in one request.

Within each media type, references follow slot-number order, skipping unconnected
slots. For example, connecting `image_2` and `image_9` sends `image_2` as the first
image and `image_9` as the second. Each image slot uses the first image in its batch.
Existing connections to `image_1`, `image_2`, `audio_1`, `audio_2`, `video_1`, and
`video_2` remain valid in saved workflows.

#### MiniMax-H3 Reference to Video

Load [MiniMax-H3 Reference to Video](example_workflows/vLLM-Omni%20MiniMax-H3%20Reference%20to%20Video.json)
from **Templates → ComfyUI-vLLM-Omni**, or drag the JSON onto the canvas.
It adapts the [official ComfyUI R2V workflow](https://github.com/Comfy-Org/workflow_templates/blob/main/templates/video_minimax_h3_r2v.json)
to remote vLLM-Omni execution. Model loading, sampling, and VAE decoding run on
the server; ComfyUI loads references and saves the returned video.

Use an extension version containing the audio-preserving output fix from
[#7456](https://github.com/vllm-project/vllm-omni/pull/7456). Older versions can
drop the server's audio while decoding the response.

1. Start a Ref2VA-capable service using the [MiniMax-H3 recipe](../../recipes/MiniMaxAI/MiniMax-H3.md).
Set **Generate Video** to its `/v1` URL and served model name.
2. Select your own image in **Load Image**. The template's filenames are placeholders
relative to ComfyUI's input directory; no sample assets or model weights are bundled.
3. Connect **Load Video** and **Load Audio** when needed. Duplicate the loaders to
fill more reference slots, or disconnect the image when using video-only input.
Keep the references connected to **Generate Video** and leave `frame` disconnected.
4. Match prompt tags such as `<Picture 1>`, `<Video 1>`, and `<Audio 1>` to the
connected references, counting each media type separately and skipping empty slots.
5. Run the workflow. **Save Video** writes an MP4 under `output/video/` and preserves
the generated audio when the audio-output prerequisite is installed.

The default is **1344×768, 24 FPS, 124 frames** (about 5.17 seconds), 50 sampling
points, video flow shift 12, audio flow shift 3, and seed 42. The duration widget
is set to 5.167 seconds, which converts to 124 frames at 24 FPS. Other H3 canvas presets
are 1024×768, 768×768, 768×1024, and 768×1344. Keep frame counts at `17k+5` within
the 4–15 second output range; examples are 107, 124, 209, and 345 frames.
Reference video/audio clips must each be 2–15 seconds, with at most 15 seconds of
reference video and at most 15 seconds of standalone reference audio per request.

For **Ref2VA Turbo**, use the same template. Configure a Ref2VA-only server with
`--lora-backend peft --lora-path "$TURBO_LORA" --task-type ref2va`, following the
recipe's hardware setup. For `minimax_h3_ref2v_turbo_4step_v0.1_bf16.safetensors`,
set **LoRA** `local_path` to that file's path on the server and connect it to
**Generate Video**. Use `num_inference_steps=5`, `flow_shift=12`,
`audio_flow_shift=3`, and LoRA `scale=1`. The API counts sigma points, so five
points produce four denoiser evaluations. Use the Diffusers artifact, not its
`_comfyui_` export. To return to base mode, disconnect LoRA and restore 50 steps.
The 8-step v1.0 Ref2VA adapter instead requires nine sampling points and video
flow shift 6. FastH3 is a separate startup-fused T2VA deployment.

Validate the template's wiring and defaults locally with:

```bash
python -m pytest tests/e2e/features/comfyui/test_h3_reference_workflow.py -q
```

For a real-model check, run the workflow against H3 and inspect the saved file:

```bash
ffprobe -v error -show_entries stream=codec_type,width,height,r_frame_rate,sample_rate,channels \
-show_entries format=duration -of json "$OUTPUT_MP4"
```

Check for 24 FPS video, stereo audio, the requested canvas, and aligned duration;
also play the result to assess reference conditioning and audio/video synchronization.
Record the server commit, model/adapter, input assets, prompt, settings, and output
alongside the result. Local schema or mocked-server checks do not establish H3
generation quality or replace this real-model validation.

#### H3 video upscale (WF-07)

The **vLLM-Omni MiniMax H3 Video Upscale** template generates video remotely, upscales it with SeedVR2, and saves the original and upscaled videos with the generated audio and FPS. See [workflow setup](docs/wf07-h3-upscale.md).
Expand Down
36 changes: 24 additions & 12 deletions apps/ComfyUI-vLLM-Omni/comfyui_vllm_omni/nodes.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,9 @@
from .utils.logger import get_logger
from .utils.models import lookup_model_spec
from .utils.types import (
MAX_REFERENCE_AUDIOS,
MAX_REFERENCE_IMAGES,
MAX_REFERENCE_VIDEOS,
AudioFormat,
AutoregressionSamplingParams,
DiffusionSamplingParams,
Expand Down Expand Up @@ -918,6 +921,10 @@ def INPUT_TYPES(cls):
"audio_2": ("AUDIO",),
"video_1": ("VIDEO",),
"video_2": ("VIDEO",),
# Append ports to preserve connections in saved workflows.
**{f"image_{i}": ("IMAGE",) for i in range(3, MAX_REFERENCE_IMAGES + 1)},
**{f"audio_{i}": ("AUDIO",) for i in range(3, MAX_REFERENCE_AUDIOS + 1)},
**{f"video_{i}": ("VIDEO",) for i in range(3, MAX_REFERENCE_VIDEOS + 1)},
},
}

Expand All @@ -934,21 +941,26 @@ def get_references(
audio_2: AudioInput | None = None,
video_1: VideoInput | None = None,
video_2: VideoInput | None = None,
image_3: torch.Tensor | None = None,
image_4: torch.Tensor | None = None,
image_5: torch.Tensor | None = None,
image_6: torch.Tensor | None = None,
image_7: torch.Tensor | None = None,
image_8: torch.Tensor | None = None,
image_9: torch.Tensor | None = None,
audio_3: AudioInput | None = None,
video_3: VideoInput | None = None,
**kwargs,
):
if kwargs:
logger.info("Uncaught kwargs: %s", kwargs)
refs = VideoReferences()
if image_1 is not None:
refs["image_1"] = image_1
if image_2 is not None:
refs["image_2"] = image_2
if audio_1 is not None:
refs["audio_1"] = audio_1
if audio_2 is not None:
refs["audio_2"] = audio_2
if video_1 is not None:
refs["video_1"] = video_1
if video_2 is not None:
refs["video_2"] = video_2
for kind, values in (
("image", (image_1, image_2, image_3, image_4, image_5, image_6, image_7, image_8, image_9)),
("video", (video_1, video_2, video_3)),
("audio", (audio_1, audio_2, audio_3)),
):
for index, value in enumerate(values, start=1):
if value is not None:
refs[f"{kind}_{index}"] = value
return (refs,)
78 changes: 38 additions & 40 deletions apps/ComfyUI-vLLM-Omni/comfyui_vllm_omni/utils/api_client.py
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@

from .format import (
audio_to_base64,
audio_to_bytes,
base64_to_audio,
base64_to_image_tensor,
bytes_to_audio,
Expand All @@ -31,7 +32,13 @@
)
from .logger import get_logger, pretty_printer
from .models import lookup_model_spec
from .types import AudioFormat
from .types import (
MAX_REFERENCE_AUDIOS,
MAX_REFERENCE_IMAGES,
MAX_REFERENCE_VIDEOS,
MAX_TOTAL_REFERENCES,
AudioFormat,
)

logger = get_logger(__name__)

Expand Down Expand Up @@ -304,38 +311,44 @@ async def generate_video(

# === multimodal inputs (first-last-frames, references, etc.) ===
input_reference_image: torch.Tensor | None = None
audio_reference: AudioInput | None = None
reference_videos: list[VideoInput] = []
video_task: str | None = None

if frame is not None:
input_reference_image = frame
video_task = "fl2va"
elif references is not None:
images = [references[k] for k in ("image_1", "image_2") if k in references and references[k] is not None]
audios = [references[k] for k in ("audio_1", "audio_2") if k in references and references[k] is not None]
videos = [references[k] for k in ("video_1", "video_2") if k in references and references[k] is not None]

if not images and not audios and not videos:
raise ValueError("references is empty; connect at least one image/audio/video.")

if videos:
if images or audios:
raise ValueError(
"Video references cannot be combined with image or audio references. "
"Connect only video_1/video_2 for multi-video Ref2VA."
)
reference_videos = videos
video_task = "ref2va"
elif len(images) == 1 and len(audios) == 1:
input_reference_image = images[0]
audio_reference = audios[0]
video_task = "ref2va"
else:
reference_formats = (
("image", MAX_REFERENCE_IMAGES, "png", "image/png", image_tensor_to_png_bytes),
("video", MAX_REFERENCE_VIDEOS, "mp4", "video/mp4", video_to_bytes),
("audio", MAX_REFERENCE_AUDIOS, "mp3", "audio/mpeg", audio_to_bytes),
)
supported_inputs = {f"{kind}_{i}" for kind, limit, *_ in reference_formats for i in range(1, limit + 1)}
connected = {name: value for name, value in references.items() if value is not None}
unsupported = connected.keys() - supported_inputs
if unsupported:
raise ValueError(f"Unsupported reference input(s): {', '.join(sorted(unsupported))}.")
if not any(name.startswith(("image_", "video_")) for name in connected):
raise ValueError(
"Invalid references combination. Supported modes: "
"(1) one or more videos only, or (2) exactly one image and one audio."
"references requires at least one image or video; audio-only inputs are not supported."
)
if len(connected) > MAX_TOTAL_REFERENCES:
raise ValueError(
f"references supports at most {MAX_TOTAL_REFERENCES} inputs in total "
f"(up to {MAX_REFERENCE_IMAGES} images, {MAX_REFERENCE_VIDEOS} videos, "
f"and {MAX_REFERENCE_AUDIOS} audios)."
)
for kind, limit, extension, content_type, encode in reference_formats:
for index in range(1, limit + 1):
name = f"{kind}_{index}"
if name in connected:
filename = f"{name}.{extension}"
form.add_field(
"input_references",
encode(connected[name], filename),
filename=filename,
content_type=content_type,
)
video_task = "ref2va"
else:
video_task = "t2va"

Expand All @@ -348,21 +361,6 @@ async def generate_video(
content_type="image/png",
)

if audio_reference is not None:
form.add_field(
"audio_reference",
json.dumps({"audio_url": audio_to_base64(audio_reference)}, ensure_ascii=False),
)

for idx, video in enumerate(reference_videos, start=1):
video_filename = f"reference_{idx}.mp4"
form.add_field(
"input_references",
video_to_bytes(video, video_filename),
filename=video_filename,
content_type="video/mp4",
)

# === model specific params. Either use a specialized builder, or add flattened fields as-is ===
if model_params is not None:
model_params = dict(model_params)
Expand Down
6 changes: 6 additions & 0 deletions apps/ComfyUI-vLLM-Omni/comfyui_vllm_omni/utils/types.py
Original file line number Diff line number Diff line change
Expand Up @@ -30,6 +30,12 @@ class MiniMaxH3ModelSpecificParams(dict):
pass


MAX_REFERENCE_IMAGES = 9
MAX_REFERENCE_VIDEOS = 3
MAX_REFERENCE_AUDIOS = 3
MAX_TOTAL_REFERENCES = 12


class VideoReferences(dict):
pass

Expand Down
Loading
Loading