Skip to content

[Model] Accelerate MiniMax-H3 end-to-end inference - #7519

Draft
lishunyang12 wants to merge 5 commits into
vllm-project:mainfrom
lishunyang12:perf/minimax-h3-ultra-e2e
Draft

lishunyang12 wants to merge 5 commits into
vllm-project:mainfrom
lishunyang12:perf/minimax-h3-ultra-e2e

Conversation

@lishunyang12

@lishunyang12 lishunyang12 commented Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator

#7415
Spliting into

ID Scope Depends on
P1 Move H3 VSA policy into the model directory; keep shared interfaces generic #7535 —
P2 Exact AdaLN cache, builder, and adapter/schedule identity checks P1
P3 DiT MXFP8 projections and SwiGLU fusion P2
P4 Generic Sage provider and H3 VSA integration P1
P5 Chunked video D2H, bounded CPU staging, and MP4 output pipeline —
P6 VAE MXFP8, paired/mixed batching, gather and audio overlap P3, P5
P7 Optional generic FlashInfer Ulysses/RDMA transport P1
P8 Gate projection overlap, QK/V splitting, early Q preparation, coarse/fine overlap, and O producer lookahead P3, P4, P7
P9 Final preset, reproducible recipe, and cumulative optimization results P6, P8

Module layout and validation

The H3 directory now has 39 Python modules instead of 45. Component quantization policy, VAE batching, VAE parallel output and VSA producer lifetime are consolidated; the model-specific VSA entry is attention/vsa.py. Shared platform selection, transport and operators retain their owners. No VDN implementation or new dependency is added.

At 7ba27060, the related regression group reports 337 passed, 20 failed. All 20 failures also occur in the pre-change dda16f4b checkout after removing its unsupported OffloadPlan.resident_offload_submodules keyword solely to allow imports (that comparison reports 327 passed, 21 failed). Nine implementation-adapter tests are brought over from P1, and its dispatch typing fix resolves the remaining one-test difference. The outstanding failures concern media output/capability contracts and fixtures missing acceleration state; this draft is not merge-qualified.

Pre-commit passes. All 61 moved function/class ASTs match after explicit import and symbol normalization. The consolidated modules import with optional kernel-provider imports blocked. No full-model E2E or VDN qualification was run.

Signed-off-by: lishunyang12 <lishunyang12@users.noreply.github.com>
@lishunyang12 lishunyang12 changed the title [Perf] Add MiniMax-H3 Ultra end-to-end acceleration [Perf] Add MiniMax-H3 real-time end-to-end acceleration Sep 14, 2026
@lishunyang12 lishunyang12 changed the title [Perf] Add MiniMax-H3 real-time end-to-end acceleration [Perf] Add MiniMax-H3 real-time gen pipeline on SM120 Sep 14, 2026
Signed-off-by: lishunyang12 <lishunyang12@users.noreply.github.com>
@lishunyang12 lishunyang12 changed the title [Perf] Add MiniMax-H3 real-time gen pipeline on SM120 [Model] Accelerate MiniMax-H3 end-to-end inference Sep 14, 2026
Signed-off-by: lishunyang12 <lishunyang12@users.noreply.github.com>
Signed-off-by: lishunyang12 <lishunyang12@users.noreply.github.com>
Signed-off-by: lishunyang12 <lishunyang12@users.noreply.github.com>
@hsliuustc0106 hsliuustc0106 added diffusion codes related to diffusion models Kernel optimization Codes related to optimize kernel execution to improve hardware utilization enhancement New feature or request labels Sep 15, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

This PR touches vllm_omni/diffusion/, tests/diffusion/, docs/user_guide/, recipes/MiniMaxAI/, tools/minimax_h3/ (51 files). Based on CODEOWNERS coverage of the changed files, the most-related reviewers appear to be:

@wtomin @Bounty-hunter @fhfuih

Could one of you take a look when you get a chance? Thanks!

@hsliuustc0106

Copy link
Copy Markdown
Collaborator

Ready for full review when draft status is removed and the merge conflicts are resolved.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 30, 2026 — with ChatGPT Codex Connector
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion codes related to diffusion models enhancement New feature or request high priority high priority issue, needs to be done asap Kernel optimization Codes related to optimize kernel execution to improve hardware utilization

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants