Repository navigation
[Frontend] Add MiniMax-H3 latent editing workflows (WF-05) - #7898
Merged
Merged
Conversation
Implements EDIT-02 from vllm-project#7380: a Latent Mask Editing node that uploads source media and serializes video/audio noise masks to /v1/videos. Signed-off-by: chen hongwei <1792043268@qq.com>
Use a supported H3 duration (4.458s -> 107 frames, on the 17n+5 lattice) in the latent-mask editing example workflow; 1.0s maps to 24 frames and is rejected. Move the latent-mask tests under tests/e2e/features/comfyui and reuse its CPU fixtures (drop the CUDA skip and ComfyUI checkout requirement). Signed-off-by: chen hongwei <1792043268@qq.com>
…ent-mask # Conflicts: # apps/ComfyUI-vLLM-Omni/comfyui_vllm_omni/utils/api_client.py
The merge resolution dropped the `from .latent_mask import` line and the `latent_edit` parameter from `generate_video`, leaving the latent-mask block referencing undefined names (ruff F821). Restore both. Signed-off-by: chen hongwei <1792043268@qq.com>
Add two example workflows demonstrating the two video_mask modes: - Image Mask: a 2D mask (LoadImageMask) broadcast to every frame - Temporal Mask: a 3D temporal mask (SolidMask + BatchMasksNode) for per-frame control (first half preserved, second half regenerated) Signed-off-by: chen hongwei <1792043268@qq.com>
Add trailing newline and normalize to LF for pre-commit. Signed-off-by: chen hongwei <1792043268@qq.com>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
Signed-off-by: Weiming Liao <liaowm5@gmail.com>
FayeSpica
marked this pull request as ready for review
September 21, 2026 06:57
FayeSpica
requested review from
NickCao,
alex-jw-brooks and
yenuo26
as code owners
September 21, 2026 06:57
Contributor
Author
|
@princepride Could you review this PR when you have time? Thanks! |
|
This PR appears to be related to model: MinimaxH3. Model owners: @david6666666 @alex-jw-brooks @fhfuih Routing: @david6666666 via model owner; @alex-jw-brooks via CODEOWNERS; @fhfuih via CODEOWNERS @FayeSpica, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
Collaborator
|
Great Job!!! Thank you so much, I succeed reproduce it! |
princepride
enabled auto-merge (squash)
September 21, 2026 12:45
avicii-forever
pushed a commit
to avicii-forever/vllm-omni
that referenced
this pull request
Sep 21, 2026
Keep audio_mask as a scalar FLOAT (the interface WF-05 vllm-project#7898 builds on), and add an optional audio_temporal_mask MASK input for time-varying audio masks. The temporal mask takes precedence over the scalar when both are connected. This keeps the frontend backward-compatible with the WF-05 workflow while still supporting temporal audio masks.
avicii-forever
added a commit
to avicii-forever/vllm-omni
that referenced
this pull request
Sep 21, 2026
avicii-forever
added a commit
to avicii-forever/vllm-omni
that referenced
this pull request
Sep 21, 2026
y-null
pushed a commit
to y-null/vllm-omni
that referenced
this pull request
Sep 22, 2026
…ect#7898) Signed-off-by: chen hongwei <1792043268@qq.com> Signed-off-by: Weiming Liao <liaowm5@gmail.com> Co-authored-by: chen hongwei <1792043268@qq.com> Co-authored-by: chen hongwei <54873389+avicii-forever@users.noreply.github.com> Signed-off-by: y-null <y-null@users.noreply.github.com>
y-null
pushed a commit
to y-null/vllm-omni
that referenced
this pull request
Sep 22, 2026
…ect#7898) Co-authored-by: chen hongwei <1792043268@qq.com> Co-authored-by: chen hongwei <54873389+avicii-forever@users.noreply.github.com> Signed-off-by: Weiming Liao <liaowm5@gmail.com>
mlaneuville
pushed a commit
to mlaneuville/vllm-omni
that referenced
this pull request
Sep 22, 2026
…ect#7898) Signed-off-by: chen hongwei <1792043268@qq.com> Signed-off-by: Weiming Liao <liaowm5@gmail.com> Co-authored-by: chen hongwei <1792043268@qq.com> Co-authored-by: chen hongwei <54873389+avicii-forever@users.noreply.github.com> Signed-off-by: Matthieu Laneuville <matthieu.laneuville@surf.nl>
khairulkabir1661
pushed a commit
to khairulkabir1661/vllm-omni
that referenced
this pull request
Sep 25, 2026
…ect#7898) Signed-off-by: chen hongwei <1792043268@qq.com> Signed-off-by: Weiming Liao <liaowm5@gmail.com> Co-authored-by: chen hongwei <1792043268@qq.com> Co-authored-by: chen hongwei <54873389+avicii-forever@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
Implement WF-05 from #7380 using the latent-edit interfaces from #7465 and #7575.
Workflow Screenshot
Test Plan
Run MiniMax-H3 FL2VA:
Use the same source clip for all four cases. Object removal and inpainting have independent generated rectangular masks, currently 180 × 168 at x=1074, y=493 on a 1344 × 768 canvas. Adapt their position and size to the input. Continuation and extension use automatically calculated temporal masks. The mask image and static preview below are historical input illustrations, not required workflow inputs.
school_generated3.mp4
WF05-mask-preview.mp4
Run each case's Mask Preview and Result outputs separately against a compatible H3 service. Inspect spatial coverage, temporal preservation, and the generated video/audio. Running the entire workflow submits all four generation branches.
vLLM Version: v0.29.0
vLLM-Omni Commit: On the EDIT-02 stack merged with main at
23f412644.Test Result
Object removal
Tips: Include handles, shadows, and some surrounding background in the mask. In this clip, expanding each side by 32 pixels improved removal:
(x, y, width, height)changed from(1106, 525, 116, 104)to(1074, 493, 180, 168). This is a clip-specific margin. For symmetric expansion byd, use(x-d, y-d, width+2d, height+2d)within the canvas bounds.WF05-object-removal-mask-preview.mp4
WF05-object-removal.mp4
Inpainting
WF05-inpainting-mask-preview.mp4
WF05-inpainting.mp4
Continuation
WF05-continuation-mask-preview.mp4
WF05-continuation.mp4
Extension
WF05-extension-mask-preview.mp4
WF05-extension.mp4
The attachments above are manual preview and generation examples, not placeholders. Prompt adherence can vary across inputs.
BEFORE SUBMITTING: read CONTRIBUTING.md and run the precheck-pr skill with the code agent for a self-check against project conventions.
(anything written below this line will be removed by GitHub Actions)