Fix MXFP4 scale placeholder initialization - #33500
Merged
Fridge003 merged 1 commit intoAug 6, 2026
Merged
Conversation
Contributor
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
weireweire
marked this pull request as ready for review
August 4, 2026 06:54
Contributor
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
weireweire
requested review from
Alisehen,
AniZpZ,
BBuf,
Edwardf0t1,
FlamingoPg,
HaiShaw,
OrangeRedeng,
b8zhong and
mmangkad
as code owners
August 4, 2026 06:54
weireweire
force-pushed
the
fix/mxfp4-valid-ue8m0-placeholders
branch
from
August 4, 2026 06:55
9f28f8d to
4026381
Compare
Contributor
Author
|
/tag-and-rerun-ci |
mmangkad
approved these changes
Aug 4, 2026
Contributor
Author
Collaborator
|
@weireweire could you merge main? Although the ci finished, it fast-failed and many tests never actually ran |
Root cause: serialized MXFP4 scale parameters were zero-filled even though they store raw UE8M0 bytes. Presharded post-load transforms can consume these placeholders before cached tensors are restored, and dummy initialization skips integer tensors entirely. Fix: initialize both MoE scale tensors to the neutral UE8M0 encoding for 1.0 while keeping packed weights zero-filled. Real checkpoint loads still overwrite the placeholders. Validation: full pre-commit passes.
weireweire
force-pushed
the
fix/mxfp4-valid-ue8m0-placeholders
branch
from
August 5, 2026 05:32
4026381 to
06ebf8b
Compare
Contributor
Author
|
sure, rebased, lets wait another round. |
Contributor
Author
|
/rerun-failed-ci |
mmangkad
approved these changes
Aug 6, 2026
Fridge003
approved these changes
Aug 6, 2026
sagearc
pushed a commit
to sagearc/sglang
that referenced
this pull request
Aug 13, 2026
…zation (sgl-project#33500) (sgl-project#34032) Co-authored-by: weireweire <weiliangl@nvidia.com> Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com> Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>
saturn-acc
pushed a commit
to saturn-acc/sglang
that referenced
this pull request
Aug 16, 2026
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Atituiset
pushed a commit
to Atituiset/sglang
that referenced
this pull request
Sep 10, 2026
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Failure
A real-weight Kimi K3 presharded reload with the MegaMoE/DeepGEMM path failed while rebuilding the post-processed parameter layout. Reload invokes process_weights_after_loading on a fresh model skeleton before copying the cached tensors, and the fresh MXFP4 scale placeholders reached DeepGEMM first.
The workers aborted in deep_gemm/include/deep_gemm/impls/smxx_layout.cuh:131, surfaced as CUDA error 719 from transform_sf_into_required_layout. This happened before the cached scales could be restored.
Root cause
Serialized MXFP4 scales are stored as raw uint8 UE8M0 bytes. Fresh modules previously zero-filled these tensors. UE8M0 byte 0 represents 2^-127; converting it to FP32 produces a subnormal value instead of the normalized power-of-two representation required by the DeepGEMM layout transform.
Dummy initialization only randomizes floating-point tensors, so the uint8 scale tensors also retained their zero placeholders.
Fix
Initialize the two serialized MXFP4 MoE scale tensors with byte 127, the UE8M0 representation of 1.0. This provides a neutral normalized placeholder for post-load processing and dummy execution while preserving normal checkpoint-loading behavior.
Validation
The behavior was validated with actual Kimi K3 weights by seeding the same scale placeholders to byte 127 immediately before post-processing, which is equivalent to the initialization performed by this change. A TP8 disaggregated deployment with one prefill worker and eight decode workers completed all 144 target/draft rank reloads without the DeepGEMM assertion. An 8K-input/1K-output benchmark then completed 1280/1280 requests.
The full pre-commit suite also passed.
CI States
Latest PR Test (Base): 🚫 Run #30978432947
Latest PR Test (Extra): ❌ Run #30978432790