Repository navigation
feat(magi2): add fused BF16 routed MoE path - #7206
Conversation
d21f835 to
cbd6d64
Compare
cbd6d64 to
5455f6d
Compare
|
This PR touches tests/diffusion/, vllm_omni/diffusion/ (4 files). Based on CODEOWNERS coverage of the changed files, the most-related reviewers appear to be: Could one of you take a look when you get a chance? Thanks! |
e69066a to
cca79ef
Compare
92c8144 to
977f59a
Compare
|
This PR appears to belong to: docs/design/module/diffusion/offloader.md, docs/design/module/diffusion/diffusion_model_integration.md, docs/design/module/diffusion/index.md. Module owners: @Bounty-hunter @fhfuih @wtomin Routing: @Bounty-hunter via module of the changed files, semantic router, CODEOWNERS; @fhfuih via module of the changed files, semantic router, CODEOWNERS; @wtomin via module of the changed files, semantic router, CODEOWNERS @yeahdongcn, please review your own changes and leave a short self-review comment describing what you checked. PRs without author self-review may not be assigned a reviewer. Please take a look when you have a chance. If you would like an automated review, mention @vllm-omni-review-bot in a comment. |
hsliuustc0106
left a comment
There was a problem hiding this comment.
Two P2s block this at the current head:
-
MAGI2_USE_BF16_MOE_KERNELis unregistered inenvironment_variable_inventory.py—tests/config/test_environment_variables.py::test_statically_resolvable_environment_accesses_are_classifiedfails on a clean checkout of 977f59a, so any cpu/core_model lane will go red. -
The per-call
packed_w13repack makes the opt-in path a net regression at production expert counts. Dispatch-mode A/B on L20X (12 heads × 256 experts, top-k 6, env toggled in one process): 512 tok 3374→11963 µs (0.28×), 4096 tok 6443→13137 µs (0.49×), 10000 tok 12832→17054 µs (0.75×). The stack alone measures 8.85 ms p50 (4.03 GB) per call at those dims. The reported 7.9–11.2× reproduce only at the benchmark's synthetic default--experts 8.
The rest verifies clean on my side: 44 CPU + 16 CUDA tests pass, parity is exact (rel-L2 0.0 at every shape including real dims), routing weights are applied exactly once in FP32, and the deterministic/env gating is correct. Packing once per weight version plus a real-dims row in the perf table should get this approve-able.
977f59a to
aaeede2
Compare
Omni ReviewBot: no human activity for 7 days@yeahdongcn this pull request has had no human commit, comment or review since 2026-09-16. Please confirm the current plan and next step. The author or a maintainer decides whether to change the PR state. To keep it moving, any one of these is enough: push an update, reply to the open blocker, or post the current plan and timeline. |
hsliuustc0106
left a comment
There was a problem hiding this comment.
Both blockers are resolved at this head, and the env-var one the better way: instead of registering MAGI2_USE_BF16_MOE_KERNEL, the flag is gone entirely — the fused path now self-selects (MUSA always; CUDA at ≥4096 tokens; still respects the registered MAGI2_DETERMINISTIC and skips under torch.compile), so the static-classification test passes by construction. The per-forward packed_w13 re-stack is replaced by _get_bf16_packed_w13() with an id/data_ptr/_version invalidation key that correctly handles in-place reloads and mmap swaps, and the stale Triton constexpr comment is fixed. Mergeable, GitHub checks green; the kernel itself is unchanged from the A/B-verified revision.
Omni ReviewBot: supersededThe CI failure noted on |
Signed-off-by: Xiaodong Ye <xiaodong.ye@mthreads.com>
19c927d to
b82e7fc
Compare
main moved the route layout and expert kernels out of mh_moe.py into fused_moe_kernels.py (vllm-project#7206). Keep that split: the fused top-k routing stays in mh_moe.py, and the rewritten global_sort_routes, its reference and _route_gather_kernel replace the old global_sort_routes where main put it. The routing test and benchmark import the layout functions from their new module. Signed-off-by: Weitian Wang <wangweitian@hotmail.com>
Summary
Add a fused BF16 routed-MoE path for MAGI-2 Preview and use it by default for
supported eager GPU inference. CUDA falls back to the native path below 4096
tokens, where launch overhead dominates; CPU, non-BF16, deterministic, and
torch.compileexecution retain the native implementation.The fast path keeps FP32 accumulation through SwiGLU7 and the down projection.
Gate/up weights are packed once after checkpoint loading and cached by
parameter/storage version. Route metadata uses reusable upper-bound buffers;
the live padded count stays device-side, avoiding per-forward allocations and
CPU/GPU synchronization.
No environment variable is required.
Validation
Result on commit
aaeede2d:(
BF16 atomic_addis unavailable before Triton-MUSA 3.6; the fused pathdoes not use it)
Production-local shape (
3EP4-local heads,256experts/head, top-k6,hidden size
256, intermediate size1280):Both runs passed numerical parity. Maximum absolute difference was
0.001953125on S5000 and0on H20. The benchmark uses alternatingnative/default calls with a 512 MiB cache flush outside timed regions; it is an
eager operator diagnostic rather than a full-video throughput claim.
Scope
This PR contains the BF16 routed-MoE kernels, MAGI-2 wiring, focused tests, and
the reproducible microbenchmark. It does not include EP framework changes,
attention, mHC, SwiGLU7 activation, sampler changes, or MUSA-only APIs. #7156
remains the integration reference.