Skip to content

feat(model_loader): revive weight loader v2 dense tranche 1 (PR2 of #31051) - #42069

Open
joeqth wants to merge 1 commit into
sgl-project:mainfrom
joeqth:feat/weight-loader-v2-pr2-dense-revival
Open

joeqth wants to merge 1 commit into
sgl-project:mainfrom
joeqth:feat/weight-loader-v2-pr2-dense-revival

Conversation

@joeqth

@joeqth joeqth commented Oct 1, 2026 •

Copy link
Copy Markdown

Motivation

The weight loader v2 migration plan (#31051) stalled after PR1 (#28671) was
merged: the PR2–PR6 drafts (#31264–#31268) have been closed with no activity
since July. This PR revives the first PR2 tranche so the migration can
continue.

Full credit for the implementation goes to @JD-ETH — this PR rebases their
draft #31264 (commit f64a164) onto current main.

Modifications

  • model_loader/auto_loader.py: add load_with_stacked_dispatch helper used
    by module-level MLP/Attention loaders
  • Opt-in v2 dense loaders (behind SGLANG_ENABLE_WEIGHT_LOADER_V2, default
    off) for qwen3, gemma2, glm4, granite, internlm2, olmo, olmo2
    and the qwen3_classification mandatory wrapper — same env-gated dual-path
    shape as llama/qwen2 from PR1
  • Legacy paths preserved verbatim as _legacy_load_weights
  • Tests: CPU stacked-dispatch unit test, Qwen3-0.6B v1/v2 state-dict
    equivalence, e2e generation-match extension

Rebase note: the intervening refactor wave (#41816 make_pp_layers, #41750
boundary renames) did not touch any load_weights region; the three-way
merge resolved cleanly. Every order-sensitive legacy branch (tie-embedding
copies, prefix fixes, PP layer filtering, kv-scale remap) is preserved in
the v2 paths — e.g. qwen3's tie copy keeps the is_last_rank guard and the
stacked-parameter ordering is unchanged.

Accuracy Tests

  • With the flag off (default) the legacy paths are byte-identical renames,
    so model outputs are unaffected.
  • Env-on equivalence: test_qwen3_v1_v2_state_dict_identical (Qwen3-0.6B)
    compares the full state dict between V1 and V2 with rtol=0, atol=0.
  • E2E: test_weight_loader_v2_e2e.py extension compares generation outputs
    V1 vs V2.
  • CPU mapping unit test: test/registered/unit/model_loader/test_stacked_params_dispatch.py.
  • I could not run GPU tests locally; CI will exercise the equiv/e2e tests.

Speed Tests and Profiling

Not applicable — weight loading only, no forward-path changes.

Checklist


CI States

Latest PR Test (Base): ❌ Run #36873446858
Latest PR Test (Extra): ❌ Run #36873446450
Latest PR Test (AMD ROCm 10): ❌ Run #36873446674

…1051)

Revive the first PR2 tranche of the weight loader v2 migration plan
(sgl-project#31051), originally drafted in sgl-project#31264 and stalled since July. The
changes are rebased from the original commit f64a164 onto current
main; the intervening refactor wave (PP layer construction, boundary
renames) did not touch any load_weights region, so the three-way
merge resolved cleanly.

- auto_loader: add load_with_stacked_dispatch helper used by
  module-level MLP/Attention loaders
- opt-in v2 dense loaders (behind SGLANG_ENABLE_WEIGHT_LOADER_V2,
  default off) for qwen3, gemma2, glm4, granite, internlm2, olmo and
  olmo2, plus the qwen3_classification mandatory wrapper; legacy
  paths preserved verbatim as _legacy_load_weights
- tests: CPU stacked-dispatch unit test, Qwen3-0.6B v1/v2 state-dict
  equivalence test, e2e generation-match extension

Original implementation from the stalled draft sgl-project#31264 (commit
f64a164) by JD-ETH; revived, rebased and verified against the
current protocol contract in sgl-project#31051.

Co-authored-by: JD-ETH <jaedon.guo@gmail.com>
Signed-off-by: joeqth <joe_qth110@163.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant