Skip to content

feat(magi2): add fused BF16 routed MoE path - #7206

Merged
hsliuustc0106 merged 1 commit into
vllm-project:mainfrom
yeahdongcn:xd/magi2-bf16-moe-kernel
Sep 28, 2026
Merged

hsliuustc0106 merged 1 commit into
vllm-project:mainfrom
yeahdongcn:xd/magi2-bf16-moe-kernel

Conversation

@yeahdongcn

@yeahdongcn yeahdongcn commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Add a fused BF16 routed-MoE path for MAGI-2 Preview and use it by default for
supported eager GPU inference. CUDA falls back to the native path below 4096
tokens, where launch overhead dominates; CPU, non-BF16, deterministic, and
torch.compile execution retain the native implementation.

The fast path keeps FP32 accumulation through SwiGLU7 and the down projection.
Gate/up weights are packed once after checkpoint loading and cached by
parameter/storage version. Route metadata uses reusable upper-bound buffers;
the live padded count stays device-side, avoiding per-forward allocations and
CPU/GPU synchronization.

No environment variable is required.

Validation

python -m pytest --noconftest -o addopts='' -q \
  tests/diffusion/models/magi2/test_bf16_moe_wiring.py \
  tests/diffusion/models/magi2/test_moe_kernel_contract.py \
  tests/benchmarks/test_magi2_bf16_moe_benchmark.py \
  -m 'core_model and cpu'

Result on commit aaeede2d:

  • CPU: 42 passed
  • NVIDIA H20 / CUDA / Triton 3.7.1: 16 passed
  • MTT S5000 / MUSA / Triton 3.2.0: 16 passed, 2 skipped
    (BF16 atomic_add is unavailable before Triton-MUSA 3.6; the fused path
    does not use it)
  • targeted pre-commit: passed

Production-local shape (3 EP4-local heads, 256 experts/head, top-k 6,
hidden size 256, intermediate size 1280):

# H20 native atomic reference
python -m benchmarks.kernels.benchmark_magi2_bf16_moe \
  --tokens 4096 --heads 3 --experts 256 --top-k 6 \
  --warmup 10 --iters 30

# S5000 portable native reference
python -m benchmarks.kernels.benchmark_magi2_bf16_moe \
  --tokens 4096 --heads 3 --experts 256 --top-k 6 \
  --warmup 10 --iters 30 --native-deterministic
Device Native p50 Default fused p50 Speedup
MTT S5000 60-SM, Triton 3.2.0 12.357 ms 4.957 ms 2.493x
NVIDIA H20, Triton 3.7.1 3.320 ms 2.377 ms 1.397x

Both runs passed numerical parity. Maximum absolute difference was
0.001953125 on S5000 and 0 on H20. The benchmark uses alternating
native/default calls with a 512 MiB cache flush outside timed regions; it is an
eager operator diagnostic rather than a full-video throughput claim.

Scope

This PR contains the BF16 routed-MoE kernels, MAGI-2 wiring, focused tests, and
the reproducible microbenchmark. It does not include EP framework changes,
attention, mHC, SwiGLU7 activation, sampler changes, or MUSA-only APIs. #7156
remains the integration reference.

@yeahdongcn
yeahdongcn force-pushed the xd/magi2-bf16-moe-kernel branch from d21f835 to cbd6d64 Compare September 7, 2026 10:45
@yeahdongcn yeahdongcn changed the title refactor(magi2): isolate BF16 fused MoE kernels feat(magi2): add opt-in BF16 routed MoE kernels Sep 8, 2026
@yeahdongcn
yeahdongcn force-pushed the xd/magi2-bf16-moe-kernel branch from cbd6d64 to 5455f6d Compare September 8, 2026 04:02
@hsliuustc0106 hsliuustc0106 added Kernel optimization Codes related to optimize kernel execution to improve hardware utilization diffusion codes related to diffusion models labels Sep 8, 2026
@hsliuustc0106

Copy link
Copy Markdown
Collaborator

This PR touches tests/diffusion/, vllm_omni/diffusion/ (4 files). Based on CODEOWNERS coverage of the changed files, the most-related reviewers appear to be:

@Bounty-hunter @fhfuih @wtomin

Could one of you take a look when you get a chance? Thanks!

@yeahdongcn
yeahdongcn requested a review from ywang96 as a code owner September 14, 2026 00:44
@vllm-omni-review-bot

Copy link
Copy Markdown

This PR appears to belong to: docs/design/module/diffusion/offloader.md, docs/design/module/diffusion/diffusion_model_integration.md, docs/design/module/diffusion/index.md.

Module owners: @Bounty-hunter @fhfuih @wtomin

Routing: @Bounty-hunter via module of the changed files, semantic router, CODEOWNERS; @fhfuih via module of the changed files, semantic router, CODEOWNERS; @wtomin via module of the changed files, semantic router, CODEOWNERS

@yeahdongcn, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer.

Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment.

@hsliuustc0106 hsliuustc0106 added the high priority high priority issue, needs to be done asap label Sep 14, 2026
Comment thread vllm_omni/diffusion/models/magi2/mh_moe.py Outdated
Comment thread vllm_omni/diffusion/models/magi2/mh_moe.py Outdated
Comment thread vllm_omni/diffusion/models/magi2/fused_moe_kernels.py Outdated
@hsliuustc0106 hsliuustc0106 added the enhancement New feature or request label Sep 14, 2026

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Two P2s block this at the current head:

  1. MAGI2_USE_BF16_MOE_KERNEL is unregistered in environment_variable_inventory.py — tests/config/test_environment_variables.py::test_statically_resolvable_environment_accesses_are_classified fails on a clean checkout of 977f59a, so any cpu/core_model lane will go red.

  2. The per-call packed_w13 repack makes the opt-in path a net regression at production expert counts. Dispatch-mode A/B on L20X (12 heads × 256 experts, top-k 6, env toggled in one process): 512 tok 3374→11963 µs (0.28×), 4096 tok 6443→13137 µs (0.49×), 10000 tok 12832→17054 µs (0.75×). The stack alone measures 8.85 ms p50 (4.03 GB) per call at those dims. The reported 7.9–11.2× reproduce only at the benchmark's synthetic default --experts 8.

The rest verifies clean on my side: 44 CPU + 16 CUDA tests pass, parity is exact (rel-L2 0.0 at every shape including real dims), routing weights are applied exactly once in FP32, and the deterministic/env gating is correct. Packing once per weight version plus a real-dims row in the perf table should get this approve-able.

Comment thread vllm_omni/diffusion/models/magi2/mh_moe.py Outdated
@yeahdongcn
yeahdongcn force-pushed the xd/magi2-bf16-moe-kernel branch from 977f59a to aaeede2 Compare September 16, 2026 13:47
@yeahdongcn yeahdongcn changed the title feat(magi2): add opt-in BF16 routed MoE kernels feat(magi2): add fused BF16 routed MoE path Sep 16, 2026
@vllm-omni-review-bot

Copy link
Copy Markdown

Omni ReviewBot: no human activity for 7 days

@yeahdongcn this pull request has had no human commit, comment or review since 2026-09-16. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state.

To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline.

@hsliuustc0106 hsliuustc0106 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both blockers are resolved at this head, and the env-var one the better way: instead of registering MAGI2_USE_BF16_MOE_KERNEL, the flag is gone entirely — the fused path now self-selects (MUSA always; CUDA at ≥4096 tokens; still respects the registered MAGI2_DETERMINISTIC and skips under torch.compile), so the static-classification test passes by construction. The per-forward packed_w13 re-stack is replaced by _get_bf16_packed_w13() with an id/data_ptr/_version invalidation key that correctly handles in-place reloads and mmap swaps, and the stale Triton constexpr comment is fixed. Mergeable, GitHub checks green; the kernel itself is unchanged from the A/B-verified revision.

@hsliuustc0106 hsliuustc0106 added ready label to trigger buildkite CI cuda-test Used to trigger vllm-omni cuda CI separately. labels Sep 25, 2026
@vllm-omni-review-bot

vllm-omni-review-bot commented Sep 25, 2026 •

Copy link
Copy Markdown

Omni ReviewBot: superseded

The CI failure noted on 19c927d2b90c refers to an earlier head; the pull request now points at b82e7fc1078f.

Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
@yeahdongcn
yeahdongcn force-pushed the xd/magi2-bf16-moe-kernel branch from 19c927d to b82e7fc Compare September 25, 2026 11:42
@hsliuustc0106 hsliuustc0106 added ready label to trigger buildkite CI and removed ready label to trigger buildkite CI labels Sep 28, 2026
@hsliuustc0106
hsliuustc0106 merged commit 29b818a into vllm-project:main Sep 28, 2026
6 checks passed
IDEA-V added a commit to IDEA-V/vllm-omni that referenced this pull request Sep 28, 2026
main moved the route layout and expert kernels out of mh_moe.py into
fused_moe_kernels.py (vllm-project#7206).  Keep that split: the fused top-k routing
stays in mh_moe.py, and the rewritten global_sort_routes, its reference
and _route_gather_kernel replace the old global_sort_routes where main
put it.  The routing test and benchmark import the layout functions from
their new module.

Signed-off-by: Weitian Wang <wangweitian@hotmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cuda-test Used to trigger vllm-omni cuda CI separately. diffusion codes related to diffusion models enhancement New feature or request high priority high priority issue, needs to be done asap Kernel optimization Codes related to optimize kernel execution to improve hardware utilization ready label to trigger buildkite CI

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants