Skip to content

Fix MXFP4 scale placeholder initialization - #33500

Merged
Fridge003 merged 1 commit into
sgl-project:mainfrom
weireweire:fix/mxfp4-valid-ue8m0-placeholders
Aug 6, 2026
Merged

Fridge003 merged 1 commit into
sgl-project:mainfrom
weireweire:fix/mxfp4-valid-ue8m0-placeholders

Conversation

@weireweire

@weireweire weireweire commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Initialize serialized MXFP4 MoE scale placeholders with the raw UE8M0 encoding of 1.0 instead of zero.
  • Keep packed weight placeholders zero-filled; real checkpoint loading still overwrites the scale placeholders.

Failure

A real-weight Kimi K3 presharded reload with the MegaMoE/DeepGEMM path failed while rebuilding the post-processed parameter layout. Reload invokes process_weights_after_loading on a fresh model skeleton before copying the cached tensors, and the fresh MXFP4 scale placeholders reached DeepGEMM first.

The workers aborted in deep_gemm/include/deep_gemm/impls/smxx_layout.cuh:131, surfaced as CUDA error 719 from transform_sf_into_required_layout. This happened before the cached scales could be restored.

Root cause

Serialized MXFP4 scales are stored as raw uint8 UE8M0 bytes. Fresh modules previously zero-filled these tensors. UE8M0 byte 0 represents 2^-127; converting it to FP32 produces a subnormal value instead of the normalized power-of-two representation required by the DeepGEMM layout transform.

Dummy initialization only randomizes floating-point tensors, so the uint8 scale tensors also retained their zero placeholders.

Fix

Initialize the two serialized MXFP4 MoE scale tensors with byte 127, the UE8M0 representation of 1.0. This provides a neutral normalized placeholder for post-load processing and dummy execution while preserving normal checkpoint-loading behavior.

Validation

The behavior was validated with actual Kimi K3 weights by seeding the same scale placeholders to byte 127 immediately before post-processing, which is equivalent to the initialization performed by this change. A TP8 disaggregated deployment with one prefill worker and eight decode workers completed all 144 target/draft rank reloads without the DeepGEMM assertion. An 8K-input/1K-output benchmark then completed 1280/1280 requests.

The full pre-commit suite also passed.


CI States

Latest PR Test (Base): 🚫 Run #30978432947
Latest PR Test (Extra): ❌ Run #30978432790

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@weireweire
weireweire marked this pull request as ready for review August 4, 2026 06:54
@weireweire
weireweire requested a review from ch-wan as a code owner August 4, 2026 06:54
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@weireweire

Copy link
Copy Markdown
Contributor Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Aug 4, 2026
@weireweire

Copy link
Copy Markdown
Contributor Author

@mmangkad thanks for review, the ci finished and the failure is not related(fixed in #33509). could we merge?

@mmangkad

mmangkad commented Aug 5, 2026

Copy link
Copy Markdown
Collaborator

@weireweire could you merge main? Although the ci finished, it fast-failed and many tests never actually ran

Root cause: serialized MXFP4 scale parameters were zero-filled even though they store raw UE8M0 bytes. Presharded post-load transforms can consume these placeholders before cached tensors are restored, and dummy initialization skips integer tensors entirely.

Fix: initialize both MoE scale tensors to the neutral UE8M0 encoding for 1.0 while keeping packed weights zero-filled. Real checkpoint loads still overwrite the placeholders.

Validation: full pre-commit passes.
@weireweire
weireweire force-pushed the fix/mxfp4-valid-ue8m0-placeholders branch from 4026381 to 06ebf8b Compare August 5, 2026 05:32
@weireweire

Copy link
Copy Markdown
Contributor Author

sure, rebased, lets wait another round.

@weireweire

Copy link
Copy Markdown
Contributor Author

/rerun-failed-ci

@Fridge003
Fridge003 merged commit fe55d78 into sgl-project:main Aug 6, 2026
307 of 361 checks passed
Fridge003 pushed a commit that referenced this pull request Aug 7, 2026
…zation (#33500) (#34032)

Co-authored-by: weireweire <weiliangl@nvidia.com>
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
sagearc pushed a commit to sagearc/sglang that referenced this pull request Aug 13, 2026
…zation (sgl-project#33500) (sgl-project#34032)

Co-authored-by: weireweire <weiliangl@nvidia.com>
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Signed-off-by: Sage Ahrac <sagiahrak@gmail.com>
saturn-acc pushed a commit to saturn-acc/sglang that referenced this pull request Aug 16, 2026
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Atituiset pushed a commit to Atituiset/sglang that referenced this pull request Sep 10, 2026
Co-authored-by: weireweire <20922698+weireweire@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants