Skip to content

[Frontend] Add MiniMax-H3 latent editing workflows (WF-05) - #7898

Merged
princepride merged 14 commits into
vllm-project:mainfrom
FayeSpica:feat/comfyui-wf-05
Sep 21, 2026
Merged

princepride merged 14 commits into
vllm-project:mainfrom
FayeSpica:feat/comfyui-wf-05

Conversation

@FayeSpica

@FayeSpica FayeSpica commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Implement WF-05 from #7380 using the latent-edit interfaces from #7465 and #7575.

  1. Add one ComfyUI workflow covering object removal, inpainting, continuation, and extension.
  2. Build independent spatial masks with native ComfyUI nodes and calculate temporal masks with MiniMax-H3 Temporal Mask.
  3. Preview each mask as a red overlay on source video using native compositing nodes, with Preview above Result for each case.
  4. Upload video and audio masks as JSON files to match the server API, including small masks and audio scalars.
  5. Add workflow usage guides for the four editing cases.

Workflow Screenshot

WF-05

Test Plan

Run MiniMax-H3 FL2VA:

vllm serve /root/.cache/huggingface/hub/models--MiniMaxAI--MiniMax-H3/snapshots/6818f6c32d12b210915e44ad56a4228c2608f160/FL2VA \
  --omni \
  --host 0.0.0.0 --port 8000 \
  --trust-remote-code \
  --served-model-name MiniMaxAI/MiniMax-H3 \
  --task-type fl2va \
  --num-gpus 8 --usp 8 --ring 1 \
  --text-encoder-tp-size 8 \
  --enable-distributed-layerwise-offload \
  --vae-parallel-mode tile \
  --vae-use-tiling \
  --vae-patch-parallel-size 8 \
  --enable-diffusion-pipeline-profiler \
  --diffusion-attention-backend FLASH_ATTN \
  --init-timeout 1800

Use the same source clip for all four cases. Object removal and inpainting have independent generated rectangular masks, currently 180 × 168 at x=1074, y=493 on a 1344 × 768 canvas. Adapt their position and size to the input. Continuation and extension use automatically calculated temporal masks. The mask image and static preview below are historical input illustrations, not required workflow inputs.

Case Expected edit Duration Audio mask
Object removal Reconstruct surrounding wall in the bag region 5 s 0
Inpainting Curled orange cat gently swaying its tail 5 s 0
Continuation Overcast sky, torrential rain and distant lightning; preserve_fraction=0.2 5 s 0.8
Extension Summer starry night; camera follows a rising firework upward until it bursts 10 s 0.9
Input Media
Input video
school_generated3.mp4
Original spatial mask (reference only) bag_mask3
Static mask preview (input illustration)
WF05-mask-preview.mp4

Run each case's Mask Preview and Result outputs separately against a compatible H3 service. Inspect spatial coverage, temporal preservation, and the generated video/audio. Running the entire workflow submits all four generation branches.

vLLM Version: v0.29.0

vLLM-Omni Commit: On the EDIT-02 stack merged with main at 23f412644.

Test Result

Object removal

Tips: Include handles, shadows, and some surrounding background in the mask. In this clip, expanding each side by 32 pixels improved removal: (x, y, width, height) changed from (1106, 525, 116, 104) to (1074, 493, 180, 168). This is a clip-specific margin. For symmetric expansion by d, use (x-d, y-d, width+2d, height+2d) within the canvas bounds.

Output Media
Mask Preview
WF05-object-removal-mask-preview.mp4
Result
WF05-object-removal.mp4

Inpainting

Output Media
Mask Preview
WF05-inpainting-mask-preview.mp4
Result
WF05-inpainting.mp4

Continuation

Output Media
Mask Preview
WF05-continuation-mask-preview.mp4
Result
WF05-continuation.mp4

Extension

Output Media
Mask Preview
WF05-extension-mask-preview.mp4
Result
WF05-extension.mp4

The attachments above are manual preview and generation examples, not placeholders. Prompt adherence can vary across inputs.

BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)

avicii-forever and others added 11 commits September 15, 2026 16:45
Implements EDIT-02 from vllm-project#7380: a Latent Mask Editing node that uploads source media and serializes video/audio noise masks to /v1/videos.

Signed-off-by: chen hongwei <1792043268@qq.com>
Use a supported H3 duration (4.458s -> 107 frames, on the 17n+5 lattice) in the latent-mask editing example workflow; 1.0s maps to 24 frames and is rejected.

Move the latent-mask tests under tests/e2e/features/comfyui and reuse its CPU fixtures (drop the CUDA skip and ComfyUI checkout requirement).

Signed-off-by: chen hongwei <1792043268@qq.com>
…ent-mask

# Conflicts:
#	apps/ComfyUI-vLLM-Omni/comfyui_vllm_omni/utils/api_client.py
The merge resolution dropped the `from .latent_mask import` line and the
`latent_edit` parameter from `generate_video`, leaving the latent-mask block
referencing undefined names (ruff F821). Restore both.

Signed-off-by: chen hongwei <1792043268@qq.com>
Add two example workflows demonstrating the two video_mask modes:
- Image Mask: a 2D mask (LoadImageMask) broadcast to every frame
- Temporal Mask: a 3D temporal mask (SolidMask + BatchMasksNode) for
  per-frame control (first half preserved, second half regenerated)

Signed-off-by: chen hongwei <1792043268@qq.com>
Add trailing newline and normalize to LF for pre-commit.

Signed-off-by: chen hongwei <1792043268@qq.com>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
@hsliuustc0106 hsliuustc0106 added the frontend code related to entrypoint label Sep 20, 2026
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
@FayeSpica
FayeSpica marked this pull request as ready for review September 21, 2026 06:57
@FayeSpica

Copy link
Copy Markdown
Contributor Author

@princepride Could you review this PR when you have time? Thanks!

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to be related to model: MinimaxH3.

Model owners: @david6666666 @alex-jw-brooks @fhfuih

Routing: @david6666666 via model owner; @alex-jw-brooks via CODEOWNERS; @fhfuih via CODEOWNERS

@FayeSpica, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@princepride

Copy link
Copy Markdown
Collaborator

Great Job!!! Thank you so much, I succeed reproduce it!

@princepride princepride left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@princepride princepride added the ready label to trigger buildkite CI label Sep 21, 2026
@princepride
princepride enabled auto-merge (squash) September 21, 2026 12:45
@princepride
princepride merged commit fa1d03a into vllm-project:main Sep 21, 2026
6 of 9 checks passed
avicii-forever pushed a commit to avicii-forever/vllm-omni that referenced this pull request Sep 21, 2026
Keep audio_mask as a scalar FLOAT (the interface WF-05 vllm-project#7898 builds on), and
add an optional audio_temporal_mask MASK input for time-varying audio masks.
The temporal mask takes precedence over the scalar when both are connected.

This keeps the frontend backward-compatible with the WF-05 workflow while
still supporting temporal audio masks.
avicii-forever added a commit to avicii-forever/vllm-omni that referenced this pull request Sep 21, 2026
avicii-forever added a commit to avicii-forever/vllm-omni that referenced this pull request Sep 21, 2026
y-null pushed a commit to y-null/vllm-omni that referenced this pull request Sep 22, 2026
…ect#7898)

Signed-off-by: chen hongwei <1792043268@qq.com>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: chen hongwei <1792043268@qq.com>
Co-authored-by: chen hongwei <54873389+avicii-forever@users.noreply.github.com>
Signed-off-by: y-null <y-null@users.noreply.github.com>
y-null pushed a commit to y-null/vllm-omni that referenced this pull request Sep 22, 2026
…ect#7898)

Co-authored-by: chen hongwei <1792043268@qq.com>
Co-authored-by: chen hongwei <54873389+avicii-forever@users.noreply.github.com>

Signed-off-by: Weiming Liao <liaowm5@gmail.com>
mlaneuville pushed a commit to mlaneuville/vllm-omni that referenced this pull request Sep 22, 2026
…ect#7898)

Signed-off-by: chen hongwei <1792043268@qq.com>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: chen hongwei <1792043268@qq.com>
Co-authored-by: chen hongwei <54873389+avicii-forever@users.noreply.github.com>
Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661 pushed a commit to khairulkabir1661/vllm-omni that referenced this pull request Sep 25, 2026
…ect#7898)

Signed-off-by: chen hongwei <1792043268@qq.com>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Co-authored-by: chen hongwei <1792043268@qq.com>
Co-authored-by: chen hongwei <54873389+avicii-forever@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

frontend code related to entrypoint ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants