Repository navigation
[NPU] Enable MXFP8 low-latency DeepEP dispatch for Fp8MoEMethod FP4 experts - #40519
Open
iridiumine wants to merge 3 commits into
Open
iridiumine wants to merge 3 commits into
iridiumine wants to merge 3 commits into
Conversation
iridiumine
requested review from
Alisehen,
AniZpZ,
BBuf,
Edwardf0t1,
FlamingoPg,
HaiShaw,
OrangeRedeng,
b8zhong,
ch-wan and
mmangkad
as code owners
September 21, 2026 03:22
iridiumine
force-pushed
the
pr-npu-fp4-mxfp8-dispatch
branch
from
September 22, 2026 09:11
1df6be8 to
6466ae9
Compare
iridiumine
commented
Sep 22, 2026
Fp8MoEMethod with is_fp4_expert on NPU delegates execution to NPUW4A8MXFP4MoEMethod kernels but never configured the DeepEP dispatcher quant dtype, leaving dispatch in BF16 and forcing a separate DynamicMxQuant before each GMM. Declare the required dispatch wire format as a DISPATCHER_QUANT_CONFIG class attribute on NPUW4A8MXFP4MoEMethod (low_latency=mxfp8, normal stays BF16 since A5 MXFP8 normal dispatch is intranode-only), have Fp8MoEMethod pick it up from the attached w13_kernel instead of platform if/else patches, and reuse the same declaration in NPUW4A8MXFP4FusedMoEMethod, replacing its hand-written copy.
iridiumine
force-pushed
the
pr-npu-fp4-mxfp8-dispatch
branch
from
September 23, 2026 08:41
3d5920d to
12b643b
Compare
sglang-npu-bot
approved these changes
Sep 23, 2026
Collaborator
|
/tag-and-rerun-ci |
Contributor
Author
|
/rerun-failed-ci |
3 similar comments
Contributor
Author
|
/rerun-failed-ci |
Contributor
Author
|
/rerun-failed-ci |
Contributor
Author
|
/rerun-failed-ci |
Open
3 of 5 tasks
Contributor
Author
|
/rerun-failed-ci |
OrangeRedeng
reviewed
Oct 1, 2026
Comment on lines
172
to
175
| if hasattr(layer, "dispatcher"): | ||
| layer.dispatcher.set_quant_config( | ||
| { | ||
| "normal_dispatcher_output_dtype": "bf16", | ||
| "low_latency_dispatcher_output_dtype": "mxfp8", | ||
| } | ||
| dict(self.w13_kernel.DISPATCHER_QUANT_CONFIG) | ||
| ) |
Collaborator
There was a problem hiding this comment.
Hi! My PR include #26408 an alternative solution that makes it possible to automatically determine the type of output dispatcher (fixing a long‑broken strategy), and it is implemented for both normal and low‑latency modes.
Collaborator
There was a problem hiding this comment.
I’d also prefer a single interface for interacting with the dispatcher type, for example the change introduced in PR #39589 has now broken automatic detection of the dispatcher output dtype in the TP (tensor parallelism) case.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
On NPU, FP4-expert models routed through Fp8MoEMethod ( is_fp4_expert=True , e.g. MiMo-V2.5-Pro with modelopt_fp4) delegate execution to NPUW4A8MXFP4MoEMethod kernels via the ASCEND MoE runner. However, process_weights_after_loading only runs Fp8MoEMethod 's own body, so the DeepEP dispatcher quant config was never set — dispatch fell back to BF16 and a separate DynamicMxQuant kernel ran before every grouped matmul (~15ms/layer at 16K-token prefill).
The native FP4 path ( NPUW4A8MXFP4FusedMoEMethod , DSV4) already sets the same dispatcher config; the Fp8MoEMethod path was simply missing it.
Modifications
With this, A5 low-latency dispatch quantizes activations to MXFP8 in-kernel (with E8M0 block scales) and the GMMs consume them directly; the three previously duplicated config dicts collapse into a single kernel-owned declaration.
Accuracy Tests
Speed Tests and Profiling
MiMo-V2.5-Pro-FP4 (DFlash), single-node A5, DeepEP low-latency, prefill with 16K-token random inputs.
Per-layer kernel profile (avg ms, before → after):
End-to-end (
bench_serving, 128 requests, rate=1.0, saturated queueing):Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): ❌ Run #37717189731
Latest PR Test (Extra): ❌ Run #37717189408
Latest PR Test (AMD ROCm 10): ❌ Run #37717189609