Repository navigation
Conversation
…1051) Revive the first PR2 tranche of the weight loader v2 migration plan (sgl-project#31051), originally drafted in sgl-project#31264 and stalled since July. The changes are rebased from the original commit f64a164 onto current main; the intervening refactor wave (PP layer construction, boundary renames) did not touch any load_weights region, so the three-way merge resolved cleanly. - auto_loader: add load_with_stacked_dispatch helper used by module-level MLP/Attention loaders - opt-in v2 dense loaders (behind SGLANG_ENABLE_WEIGHT_LOADER_V2, default off) for qwen3, gemma2, glm4, granite, internlm2, olmo and olmo2, plus the qwen3_classification mandatory wrapper; legacy paths preserved verbatim as _legacy_load_weights - tests: CPU stacked-dispatch unit test, Qwen3-0.6B v1/v2 state-dict equivalence test, e2e generation-match extension Original implementation from the stalled draft sgl-project#31264 (commit f64a164) by JD-ETH; revived, rebased and verified against the current protocol contract in sgl-project#31051. Co-authored-by: JD-ETH <jaedon.guo@gmail.com> Signed-off-by: joeqth <joe_qth110@163.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
The weight loader v2 migration plan (#31051) stalled after PR1 (#28671) was
merged: the PR2–PR6 drafts (#31264–#31268) have been closed with no activity
since July. This PR revives the first PR2 tranche so the migration can
continue.
Full credit for the implementation goes to @JD-ETH — this PR rebases their
draft #31264 (commit f64a164) onto current main.
Modifications
model_loader/auto_loader.py: addload_with_stacked_dispatchhelper usedby module-level MLP/Attention loaders
SGLANG_ENABLE_WEIGHT_LOADER_V2, defaultoff) for
qwen3,gemma2,glm4,granite,internlm2,olmo,olmo2and the
qwen3_classificationmandatory wrapper — same env-gated dual-pathshape as llama/qwen2 from PR1
_legacy_load_weightsequivalence, e2e generation-match extension
Rebase note: the intervening refactor wave (#41816
make_pp_layers, #41750boundary renames) did not touch any
load_weightsregion; the three-waymerge resolved cleanly. Every order-sensitive legacy branch (tie-embedding
copies, prefix fixes, PP layer filtering, kv-scale remap) is preserved in
the v2 paths — e.g. qwen3's tie copy keeps the
is_last_rankguard and thestacked-parameter ordering is unchanged.
Accuracy Tests
so model outputs are unaffected.
test_qwen3_v1_v2_state_dict_identical(Qwen3-0.6B)compares the full state dict between V1 and V2 with
rtol=0, atol=0.test_weight_loader_v2_e2e.pyextension compares generation outputsV1 vs V2.
test/registered/unit/model_loader/test_stacked_params_dispatch.py.Speed Tests and Profiling
Not applicable — weight loading only, no forward-path changes.
Checklist
CI States
Latest PR Test (Base): ❌ Run #36873446858
Latest PR Test (Extra): ❌ Run #36873446450
Latest PR Test (AMD ROCm 10): ❌ Run #36873446674