Skip to content

[Frontend] MiniMax-H3: Add latent-mask editing to the ComfyUI extension - #7575

Merged
princepride merged 5 commits into
vllm-project:mainfrom
avicii-forever:avicii/edit-02-h3-latent-mask
Oct 6, 2026
Merged

princepride merged 5 commits into
vllm-project:mainfrom
avicii-forever:avicii/edit-02-h3-latent-mask

Conversation

@avicii-forever

@avicii-forever avicii-forever commented Sep 15, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Related: #7380

Add MiniMax H3 latent-mask editing to the ComfyUI extension. A new Latent Mask Editing node uploads the source media and serializes the video/audio noise masks the /v1/videos API accepts, and Generate Video accepts a new latent_edit input that forwards it to the client.

  • source_video / source_audio upload the media to edit.
  • video_mask is a ComfyUI mask image (0 preserves, 1 regenerates, fractional blends). A 2D mask [H, W] is applied to every frame; a 3D mask [T, H, W] is a temporal mask for continuation or extension.
  • audio_mask is a scalar in [0, 1].

The mask is resized to the H3 latent grid using the server's own shape rules replicated client-side: the frame count snaps up to the 17n+5 lattice and width/height floor to a multiple of 32, so the serialized [Tv, H/16, W/16] grid matches what the server computes. The client sends source_video (mp4), source_audio (mp3), video_noise_mask (JSON grid), and audio_noise_mask (JSON scalar) multipart fields. A non-trivial mask requires its matching source, and a source without a mask is rejected.

Test Plan

  • Unit: python -m pytest tests/e2e/features/comfyui/test_latent_mask.py — raw-mask JSON serialization (2D/3D video masks, scalar/1D/2D audio masks).
  • E2E: python -m pytest tests/e2e/features/comfyui/test_latent_mask_editing.py — asserts the exact multipart fields the client produces (masks uploaded as JSON file parts).
  • Real API server: python -m pytest tests/e2e/features/comfyui/test_comfyui_integration.py -k latent_mask — boots the real /v1/videos server (mocked AsyncOmni) and exercises the latent-mask path.

Test Result

Environment: NVIDIA H20, Python 3.12.13, pytest 9.1.1, torch 2.11.0+cu126.

Unit — tests/e2e/features/comfyui/test_latent_mask.py:

collected 10 items

test_scalar_mask_to_json PASSED
test_video_mask_grid_shape PASSED
test_video_mask_grid_shape_non_aligned_frames PASSED
test_video_mask_grid_shape_short_clip PASSED
test_video_mask_grid_shape_unaligned_width PASSED
test_video_mask_3d_input PASSED
test_video_mask_temporal PASSED
test_video_mask_temporal_resize PASSED
test_video_mask_uniform_preserved PASSED
test_video_mask_values_in_range PASSED

10 passed

E2E — tests/e2e/features/comfyui/test_latent_mask_editing.py (against a mock /v1/videos server):

collected 1 item

test_latent_mask_editing.py::test_latent_edit_serialization PASSED

1 passed

test_latent_edit_serialization asserts the exact multipart fields the client produced: source_video file part (video/mp4), audio_noise_mask scalar "0.5", and video_noise_mask resized to the aligned [Tv, H/16, W/16] grid — 7x6x10 for a 160x120 / 22-frame request (height 120 floors to 96).

@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to be related to model: MinimaxH3.

Model owners: @david6666666 @alex-jw-brooks @fhfuih

Routing: @david6666666 via semantic router, model owner; @alex-jw-brooks via CODEOWNERS; @fhfuih via CODEOWNERS

@avicii-forever, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@avicii-forever
avicii-forever force-pushed the avicii/edit-02-h3-latent-mask branch 2 times, most recently from 179d56f to 8143c18 Compare September 15, 2026 08:45
160,
120,
24,
1.0

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use a supported H3 duration; one second is rejected by the server.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the review! I'll update the example workflow to use a supported H3 duration.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed — the example workflow now uses duration: 4.458 (107 frames at 24 fps, which lies on H3's 17n+5 frame lattice); the previous 1.0 mapped to 24 frames and was rejected.

except ImportError:
pytest.skip("ComfyUI / extension import unavailable", allow_module_level=True)

pytestmark = pytest.mark.skipif(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Move these tests under tests/e2e/features/comfyui and reuse its CPU fixtures.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the review! I'll move these tests under tests/e2e/features/comfyui and reuse its CPU fixtures.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed — the tests now live under tests/e2e/features/comfyui/ and reuse that directory's CPU fixtures (the conftest.py mocks for comfy_api / comfy_extras). The CUDA skip and the ComfyUI-checkout requirement are removed.

@avicii-forever

Copy link
Copy Markdown
Contributor Author

Merged latest main (fa639b88) into this branch to clear the merge conflict.

Why

This branch was based on e0c4f97, and main has since advanced 55 commits —
including #7483 ("Add MiniMax-H3 references and Ref2VA workflow"), which
rewrote the reference-input handling in api_client.py, nodes.py,
types.py, and README.md — the same files this PR touches. #7483 removed the
old audio_reference / reference_videos parameters and replaced them with a
generic reference_formats loop, which conflicted with the latent-mask editing block this PR inserts right after the old reference block.

What changed

  • api_client.py (the only manual conflict resolution): kept main's
    reference_formats handling, dropped the now-obsolete audio_reference /
    reference_videos blocks, and kept this PR's latent-mask editing block
    untouched. Net diff vs main is now just +45 lines (the latent block).
  • nodes.py / types.py / README.md: auto-merged — this PR's
    VLLMOmniLatentMaskEditing / LatentMaskEditing coexist with [Frontend] Add MiniMax-H3 references and Ref2VA workflow #7483's
    VLLMOmniVideoReferences / MAX_REFERENCE_*.
  • This PR's own content is byte-identical (no changes):
    latent_mask.py, both test files, mock_videos_server.py, the example
    workflow JSON, __init__.py, web/main.js.

Verification

@hsliuustc0106 hsliuustc0106 added frontend code related to entrypoint enhancement New feature or request labels Sep 16, 2026
@avicii-forever

avicii-forever commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor Author

End-to-end verification on real MiniMax-H3

Verified EDIT-01 (#7465, server) and EDIT-02 (this PR, frontend) together on real hardware.

Environment

Results

Test Result
Plain T2VA (107 frames / 4.458 s) valid H.264+AAC video, ~12 s
Latent-mask editing, 2D image mask partial mask regenerated the masked region
Latent-mask editing, 3D temporal mask first half preserved source, second half regenerated

Operational note
The Generate Video node auto-fills aspect_ratio only when lookup_model_spec matches MiniMax-H3 (it matches the last path segment). When serving from a local path, start the server with --served-model-name MiniMaxAI/MiniMax-H3 and keep model = MiniMaxAI/MiniMax-H3.

Example workflows (added to example_workflows/)

  1. vLLM-Omni Latent Mask Editing - Image Mask.json — 2D image mask (LoadImageMask), one mask broadcast to every frame
  2. vLLM-Omni Latent Mask Editing - Temporal Mask.json — 3D temporal mask (SolidMask + BatchMasksNode), per-frame control

2D image mask workflow:
![2D image mask workflow](
图片掩码工作流

3D temporal mask workflow:
![3D temporal mask workflow]
视频掩码工作流

@avicii-forever

Copy link
Copy Markdown
Contributor Author

Update: re-verified against the latest #7465

#7465 has been rebased onto a newer main (2ab5d170), which resolves the
conflict list in the previous comment (the author handled it during the rebase).

Re-ran the smoke test against the updated combined branch
(2ab5d170 + #7465 + this PR):

Test Result Time
Plain T2VA completed 8.4 s
Latent-mask editing, 2D image mask completed 7.0 s
Latent-mask editing, 3D temporal mask completed 6.7 s

No regression from the rebase.

@princepride

Copy link
Copy Markdown
Collaborator

Can we use another example?😂 Maybe we can refer comfyui's H3 mask edit example.

@princepride princepride left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks — node registration, types and colour family follow the existing VideoReferences pattern, and area-resizing the mask onto the latent grid is reasonable.

Requesting changes:

  1. Depends on an unsettled server contract. This client hard-codes #7465's current form fields and its 1 MiB text limit, while #7465 is unmerged and has open change requests on exactly that API surface. Please land this after #7465 and follow its final contract.
  2. The mock server can't catch contract drift (inline). The existing test_comfyui_integration.py fixture runs the real API server with a mocked AsyncOmni; please use it instead of mock_videos_server.py.
  3. Temporal mask mapping doesn't follow H3's VAE frame grouping (inline).
  4. Client re-implements server rules — both the shape lattice in latent_mask.py and the mask/source validation in api_client.py (inline). Longer term it may be simpler for the server to accept any mask resolution and resize itself, so the client doesn't have to mirror the lattice.

Example workflows: the three templates are ~500 lines each and largely duplicate each other; please keep one (or two: image mask + temporal mask). They use a 160x120 test canvas (floored to 160x96 server-side) rather than a realistic H3 resolution such as 1344x768. More importantly, the default vLLM-Omni Latent Mask Editing.json feeds SolidMask 1.0, i.e. an all-generate mask, which the server treats as a no-op — so the template that says "Partially restyle the clip" doesn't edit the video at all, only the audio. Please default to a mask that actually preserves something.

)
if video_mask is not None:
mask_json = video_mask_to_grid_json(video_mask, width=width, height=height, num_frames=num_frames)
if len(mask_json) > 1024 * 1024:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This threshold mirrors the server's form-text limit (Starlette's 1 MiB max_part_size for non-file fields). Since the server accepts the mask as a JSON file part, please always send it that way and drop the magic number, so the client isn't coupled to a server-side constant.

if video_mask is None and audio_mask is None:
raise ValueError("Latent-mask editing requires at least one mask.")

video_mask_trivial = video_mask is None or bool((video_mask == 1.0).all().item())

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The trivial-mask / required-source checks re-implement the server's validation (with a .item() on the mask), and the server already returns a clear 400 for these cases. Keeping only the "at least one mask" check would be enough on the client; scalar_mask_to_json's range check is likewise already covered by the node widget limits.

return str(value)


def _align_frame_count(frame_count: int) -> int:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This mirrors the server's 17n+5 lattice, _video_latent_t and the //32 canvas floor, so any server change will silently produce wrong-shaped masks. Some dead code here too: the frame_count <= 0 branch is unreachable for a real request, and in _video_latent_t the <= 5 branch returns the same 2 the formula already gives for an aligned count of 5. The while loop can be replaced by arithmetic (n + (5 - n) % 17).

if grid.shape[0] == 1:
grid = grid.expand(tv, gh, gw)
elif grid.shape[0] != tv:
grid = F.interpolate(grid.unsqueeze(0).unsqueeze(0), size=(tv, gh, gw), mode="nearest").squeeze(0).squeeze(0)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nearest interpolation spreads mask slices uniformly over latent frames, but H3's VAE groups frames non-uniformly (first 5 frames -> 2 latents, then every 17 frames -> 5 latents). For a per-frame mask (T == num_frames, the natural ComfyUI shape), each latent picks a single frame's value, so a preserve->regenerate boundary inside a latent's frame group can end up preserved. Please map by the VAE grouping and take the max per group (regenerate wins) when T == num_frames. The relationship between T and num_frames is also not validated at the moment.

from fastapi.responses import Response
from starlette.datastructures import UploadFile

app = FastAPI()

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This mock records fields without validating them, so it can't detect client/server contract drift. For example, the 160x120 / 22-frame (~0.9 s) request in test_latent_edit_serialization would be rejected by the real server (resolve_minimax_h3_shape requires 4-15 s) but passes here. Please reuse the real-API-server fixture from test_comfyui_integration.py and drop this file.

"source_video": ("VIDEO",),
"source_audio": ("AUDIO",),
"video_mask": ("MASK",),
"audio_mask": ("FLOAT", {"default": -1.0, "min": -1.0, "max": 1.0, "step": 0.01}),

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Using -1.0 as an "unset" sentinel means any value in (-1, 0) is silently treated as unset. Making audio_mask an optional input socket (forceInput) would give None when unconnected and remove the sentinel.

Comment on lines +399 to +402
else:
form.add_field("video_noise_mask", mask_json)
if audio_mask is not None:
form.add_field("audio_noise_mask", scalar_mask_to_json(audio_mask))

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
else:
form.add_field("video_noise_mask", mask_json)
if audio_mask is not None:
form.add_field("audio_noise_mask", scalar_mask_to_json(audio_mask))

Could we upload both masks as JSON files with a filename and content_type="application/json", regardless of payload size? Since #7465, the server declares both video_noise_mask and audio_noise_mask as UploadFile. The small-video-mask branch and the audio-mask field here send plain strings, which are rejected before inference.

check:

video_noise_mask: UploadFile | None = File(default=None),
audio_noise_mask: UploadFile | None = File(default=None),

@avicii-forever
avicii-forever force-pushed the avicii/edit-02-h3-latent-mask branch from 661c49e to 435542f Compare September 21, 2026 15:33
@avicii-forever

Copy link
Copy Markdown
Contributor Author

Rebased onto latest main and updated for review. Depends on #7947 (server-side raw-mask acceptance), which should land first.

Changes since the last review

Verification

  • Client unit tests: 32 passed (CPU).
  • Real-API-server integration suite (test_comfyui_integration.py): 37 passed.
  • ComfyUI end-to-end (real MiniMax-H3, small canvas): workflow completed, produced a valid H.264+AAC video.

@avicii-forever

Copy link
Copy Markdown
Contributor Author

Rebased onto latest main and updated for review. Depends on #7947 (server-side raw-mask acceptance), which should land first.

What changed

Server side (separate PR #7947)

  • _canonical_video_edit_mask now accepts raw masks — a 2D spatial [H, W] (broadcast over time) and a 3D frame-space [T, H, W] (one slice per source frame) — and resolves them to the latent grid server-side.
  • _canonical_audio_edit_mask now accepts a raw temporal [T] and channel-major [C, T] mask.
  • The temporal max-pool uses the VAE's exact structure, determined by probing the MiniMax-H3 video VAE: pad to a multiple of 17, stride-4 max-pool each 17-frame clip into 5 tokens, drop the last 3 (vae_clip_length=17, vae_ratio_t=4, vae_token_drop=3). The count ceil(T/17)*5 - 3 matches the VAE for T ∈ {5, 17, 22, 34, 39, 51, 107}.

Client side (this PR)

  • Masks are uploaded as JSON file parts (video_noise_mask / audio_noise_mask are UploadFile), matching the merged [Model][Frontend] MiniMax-H3: Add latent-mask editing #7465 contract; the 1 MiB string-field branch is gone.
  • The client no longer mirrors the H3 shape lattice — it sends raw masks, and the server ([Model] MiniMax-H3: accept raw video masks and resolve the latent grid server-side #7947) does the resize + temporal pooling.
  • audio_mask stays a scalar FLOAT (uniform, backward compatible); a new optional audio_temporal_mask (MASK) input supports time-varying audio masks and takes precedence when connected.
  • Dropped the mock_videos_server.py; the e2e serialization test asserts multipart fields in-process, and test_comfyui_integration.py gained a real-API-server latent-mask test.
  • Example workflows: removed the no-op all-generate template, switched to a realistic 1344×768 canvas, and normalized the two remaining templates to the same schema.

What I deliberately did NOT change (and why)

Support for the WF-05 workflow PR (#7898)

#7898 (FayeSpica's WF-05) is branched directly from this PR, so I kept its interfaces intact and verified the combined client + WF-05 nodes on real hardware:

  1. Scalar audio_mask is unchanged — [Frontend] Add MiniMax-H3 latent editing workflows (WF-05) #7898's four cases (object removal / inpainting / continuation / extension) use scalar audio masks, which keep working with no change to the workflow.
  2. Raw-mask serialization is compatible with the WF-05 temporal-mask node. VLLMOmniMiniMaxH3TemporalMask emits a [latent_t, 1, 1] mask; the server's raw path interprets it correctly (1×1 spatial broadcast, latent_t temporal passthrough), so the node's output needs no rework.
  3. Both node families register together. I verified VLLMOmniLatentMaskEditing, VLLMOmniMiniMaxH3TemporalMask, and VLLMOmniGenerateVideo all load and run against this client in one ComfyUI instance.

When #7898 rebases onto this PR, the only thing its author needs to do is drop the inherited grid-serialization helpers (video_mask_to_grid_json) in favor of the raw-mask path here; the temporal-mask node and the WF-05 workflow itself should carry over as-is.

Verification

  • Server mask canonicalization tests: 8 passed (CPU).
  • Real-API-server integration suite (test_comfyui_integration.py): 37 passed.
  • ComfyUI end-to-end (real MiniMax-H3): workflow completed and produced a valid H.264 + AAC video. Note: run at a 256×144 canvas because this 2×97 GB instance cannot hold the 135 GB FL2VA model plus a 1344×768 generation without offload; the full-resolution run is deferred to a larger instance.

@avicii-forever
avicii-forever force-pushed the avicii/edit-02-h3-latent-mask branch 3 times, most recently from 6ef3458 to 4fd43a8 Compare September 22, 2026 01:09
@avicii-forever

Copy link
Copy Markdown
Contributor Author

E2E verification against #7898's WF-05 workflow

Ran this PR's raw-mask client + VLLMOmniMiniMaxH3TemporalMask node against the raw-mask server (#7947) end-to-end on real hardware.

Environment: 2× RTX PRO 6000 Blackwell 96 GB (192 GB total), TP2, no offload; vllm 0.29.0 / torch 2.13.0+cu130; served model MiniMaxAI/MiniMax-H3 (FL2VA).

Test: loaded vLLM-Omni MiniMax-H3 Latent Mask Editing.json (WF-05, 4 scenarios) into ComfyUI and ran the whole graph. Produced valid H.264+AAC videos:

  • ✅ continuation — 5.17 s / 124 frames
  • ✅ extension
  • ✅ inpainting
  • ⏳ object-removal (in progress)

Workflow portability issues (upstream in #7898's WF-05 JSON, not this PR's code) — worked around locally for the test:

  1. The WF-05 workflow hard-codes http://127.0.0.1:8001/v1 while the frontend default is :8000; patched to :8000 locally.
  2. It depends on the MarkdownNote custom node (not installed); skipped (documentation only).
  3. It ships with an empty LoadVideo.file; filled with test_source.mp4 for the run.

…ask node (supports WF-05 workflows)

Signed-off-by: chen hongwei <1792043268@qq.com>
Signed-off-by: chen hongwei <1792043268@qq.com>
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 7 days

@avicii-forever this pull request has had no human commit, comment or review since 2026-09-22. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

princepride and others added 3 commits October 5, 2026 17:05
… uploads

- Temporal Mask node now emits one slice per output frame instead of one per
  latent, so the server's frame-to-latent pooling preserves the intended
  prefix (WF-05 extension example: 32 latents instead of 9).
- Area-downsample video masks by the VAE spatial stride and round values
  before upload so they stay within the server's 8 MiB mask limit.
- Drop the audio_temporal_mask input: no node produces a [T]/[C, T] tensor and
  ComfyUI MASK inputs are [H, W]/[B, H, W].
- Restore the WF-05 README entry and remove the unused mock_videos_server.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: Wang Zhipeng <wangzhipeng628@gmail.com>
…nvas

The server now pads a short [T, H, W] video mask with its last slice instead
of stretching it, so the Temporal Mask template's two-slice SolidMask batch
([0, 1]) preserved only the first latent and regenerated almost the whole
clip. Build the mask with the MiniMax-H3 Temporal Mask node (one slice per
output frame, continuation at preserve_fraction 0.5) as WF-05 does.

Also drop the all-generate default template (SolidMask 1.0 is a no-op for the
video; the Image Mask template covers the same flow) and move both remaining
templates from the 160x120 test canvas to 1344x768. Add CPU tests that check
the template wiring, node interfaces, canvas, and the per-frame mask source.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: princepride <wangzhipeng628@gmail.com>
The Image Mask and Temporal Mask templates duplicate WF-05: its object
removal and inpainting cases send the same 2D spatial mask, and its
continuation case uses the identical Load Video -> Get Video Components ->
Temporal Mask wiring. Remove the two small templates and point the workflow
tests at WF-05, checking every case's mask source, and that each temporal
case's mode and duration match its Generate Video node.

Default WF-05's Generate Video URL to :8000 like the other templates, and
note in its usage guide that Load Image (as Mask) can replace a case's
rectangle for an arbitrary-shape mask.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Signed-off-by: princepride <wangzhipeng628@gmail.com>

@princepride princepride left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@princepride princepride added the ready label to trigger buildkite CI label Oct 6, 2026
@princepride
princepride enabled auto-merge (squash) October 6, 2026 16:01
@princepride
princepride merged commit bad88bc into vllm-project:main Oct 6, 2026
7 of 9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request frontend code related to entrypoint ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants