Skip to content

[diffusion] loader: filter duplicate precision variants across custom loaders - #37616

Merged
mickqian merged 4 commits into
sgl-project:mainfrom
dujifeng:fix/sd3-text-encoder-weight-load
Sep 3, 2026
Merged

mickqian merged 4 commits into
sgl-project:mainfrom
dujifeng:fix/sd3-text-encoder-weight-load

Conversation

@dujifeng

@dujifeng dujifeng commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

Some Diffusers-style checkpoints contain both canonical safetensors files and precision-specific copies, for example model.safetensors and model.fp16.safetensors. Loading both can produce duplicate tensor names or make the selected value depend on file order.

The original fix covered transformers and text encoders, but the same checkpoint discovery pattern exists in other customized component loaders.

Modifications

  • Define precision-variant selection in the shared weight-source layer and apply it by default during local safetensors discovery.
  • Prefer a canonical checkpoint family across sharded and unsharded layouts while preserving precision-only checkpoints.
  • Route transformer, text/image encoder, bridge, and upsampler discovery through the shared path; VAE retains one explicit raw-candidate path because its model-specific precision selector must run first.
  • Cover adapter, vocoder, sound tokenizer, and diffusion decoder through the shared state-dict loader.
  • Cover explicit transformer weight sources, including local directories and Hugging Face inventories.
  • Apply the same selection to post-training weight updates and checkpoint-size accounting.
  • Keep safetensors indexes authoritative and run model-specific VAE selection before generic deduplication.
  • Add regression tests for canonical, explicit-variant, indexed, precision-only-index, shared state-dict, checkpoint-size, and text encoder selection.

Accuracy Tests

Previously run by the author:

  • python -m pytest -q sglang/multimodal_gen/test/unit/test_transformer_quant.py: 68 passed, 18 subtests passed.
  • python -m pytest -q sglang/multimodal_gen/test/unit/test_text_encoder_loader.py: 36 passed, 3 subtests passed.

Latest update:

  • All repository pre-commit hooks for changed files pass.
  • Full unit coverage is running in CI.

Speed Tests and Profiling

Not applicable. File selection happens only during model loading and does not modify the inference hot path.


CI States

Latest PR Test (Base): ✅ Run #33747564434
Latest PR Test (Extra): ❌ Run #33747563650
Latest PR Test (AMD ROCm 7.2): ⏳ Run #33747564498

@github-actions github-actions Bot added quant LLM Quantization diffusion SGLang Diffusion labels Sep 2, 2026
@mickqian mickqian changed the title [diffusion] loader: filter duplicate text encoder precision variants [diffusion] loader: filter duplicate precision variants across custom loaders Sep 3, 2026
@mickqian

mickqian commented Sep 3, 2026

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 3, 2026
@mickqian

mickqian commented Sep 3, 2026

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@mickqian

mickqian commented Sep 3, 2026

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@mickqian
mickqian merged commit 9cb38a3 into sgl-project:main Sep 3, 2026
129 of 143 checks passed
StevenChenSE pushed a commit to StevenChenSE/sglang that referenced this pull request Sep 6, 2026
…oaders (sgl-project#37616)

Co-authored-by: mickqian <mickqian@users.noreply.github.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
sglang-miles-h3 loaded both model.safetensors and model.fp16.safetensors
for SD3.5's CLIP text encoders, which left pooled_projections wrong
(cosine 0.33 vs diffusers). sglang-miles carries the loader fix
(sgl-project/sglang#37616), so the pooled embedding now matches
diffusers (cosine 0.99996) and every downstream metric moves.

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35798499525).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
Attribution: sgl-project/sglang#37616 (9cb38a3d57, "filter duplicate
precision variants across custom loaders") is the only commit between
sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f)
that moves this test. A CI-runner bisect (record-e2e-standards on
<commit> + the four sglang-miles-h3 picks) shows the series equal to the
old standard up to 9cb38a3d57^ and equal to this standard from
9cb38a3d57 on.

Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors
for its CLIP text encoders. The old loader globbed both, which left
pooled_projections wrong: cosine 0.33 vs diffusers' encode_prompt, vs
0.99996 after #37616. encoder_hidden_states were identical on both sides.
Every rollout and train metric moves downstream of the pooled embedding.

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35798499525).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
Attribution: sgl-project/sglang#37616 (9cb38a3d57, "filter duplicate
precision variants across custom loaders") is the only commit between
sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f)
that moves this test. A CI-runner bisect (record-e2e-standards on
<commit> + the four sglang-miles-h3 picks) shows the series equal to the
old standard up to 9cb38a3d57^ and equal to this standard from
9cb38a3d57 on.

Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors
(bit-identical) for its CLIP text encoders. Before #37616, sgl-d's own
CLIP loader saw both and raised "Duplicate tensor names detected across
safetensors files" (517 names for CLIP-G). The pipeline then silently fell
back to transformers' CLIPTextModel, which has no projection head and
drops text_projection.weight ("UNEXPECTED"). So the CLIP-G half of
pooled_projections was the unprojected pooler output: cos -0.016 vs the
projected embedding. CLIP-L's projection is near-identity (cos 1.0). The
concatenated 2048-d vector gives cos 0.333 vs diffusers' encode_prompt,
with norms 50.43 vs 44.68, exactly as measured. After #37616 the sgl-d
CLIPTextModelWithProjection loads, all 149 + 389 params match the HF
files, and cos is 0.99996. encoder_hidden_states (penultimate layer, no
projection) were identical on both sides.

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35798499525).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
Attribution: sgl-project/sglang#37616 (9cb38a3d57) only, the same cause
as the SD3 NFT standard. Probes on the CI runner (<commit> + the four
sglang-miles-h3 picks) give: 5375babbac == 9cb38a3d57^ == old standard;
9cb38a3d57 == 5a8da8cc3f == this standard.

Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors
(bit-identical) for its CLIP text encoders. Before #37616, sgl-d's own
CLIP loader saw both and raised "Duplicate tensor names detected across
safetensors files" (517 names for CLIP-G). The pipeline then silently fell
back to transformers' CLIPTextModel, which has no projection head and
drops text_projection.weight ("UNEXPECTED"). So the CLIP-G half of
pooled_projections was the unprojected pooler output: cos -0.016 vs the
projected embedding. CLIP-L's projection is near-identity (cos 1.0). The
concatenated 2048-d vector gives cos 0.333 vs diffusers' encode_prompt,
with norms 50.43 vs 44.68, exactly as measured. After #37616 the sgl-d
CLIPTextModelWithProjection loads, all 149 + 389 params match the HF
files, and cos is 0.99996. encoder_hidden_states (penultimate layer, no
projection) were identical on both sides.

The step-0 mean reward drops 0.416 -> 0.375 here, but that is sampling
noise from the recipe's 8 prompts, not a regression. In a paired run of
128 test prompts x 4 seeds with this recipe's sampling (512 samples per
side), the old base scores 0.4366 and sglang-miles 0.4686: +0.032, 95% CI
[+0.0006, +0.064]. Resampling 8 of those prompts gives a new-old spread of
[-0.087, +0.155], and a <= -0.040 draw has probability ~0.12.

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35927893055).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 24, 2026
Attribution: sgl-project/sglang#37616 (9cb38a3d57) only, the same cause
as the SD3 NFT standard. Probes on the CI runner (<commit> + the four
sglang-miles-h3 picks) give: 5375babbac == 9cb38a3d57^ == old standard;
9cb38a3d57 == 5a8da8cc3f == this standard.

Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors
(bit-identical) for its CLIP text encoders. Before #37616, sgl-d's own
CLIP loader saw both and raised "Duplicate tensor names detected across
safetensors files" (517 names for CLIP-G). The pipeline then silently fell
back to transformers' CLIPTextModel, which has no projection head and
drops text_projection.weight ("UNEXPECTED"). So the CLIP-G half of
pooled_projections was the unprojected pooler output: cos -0.016 vs the
projected embedding. CLIP-L's projection is near-identity (cos 1.0). The
concatenated 2048-d vector gives cos 0.333 vs diffusers' encode_prompt,
with norms 50.43 vs 44.68, exactly as measured. After #37616 the sgl-d
CLIPTextModelWithProjection loads, all 149 + 389 params match the HF
files, and cos is 0.99996. encoder_hidden_states (penultimate layer, no
projection) were identical on both sides.

The step-0 mean reward drops 0.416 -> 0.375 here, but that is sampling
noise from the recipe's 8 prompts, not a regression. In a paired run of
128 test prompts x 4 seeds with this recipe's sampling (512 samples per
side), the old base scores 0.4366 and sglang-miles 0.4686: +0.032, 95% CI
[+0.0006, +0.064]. Resampling 8 of those prompts gives a new-old spread of
[-0.087, +0.155], and a <= -0.040 draw has probability ~0.12.

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35927893055).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 25, 2026
The previous curve came from a sglang build whose SD3.5 CLIP text
encoders fell back to transformers' CLIPTextModel and dropped
text_projection (fixed by sgl-project/sglang#37616). The new curve is the
same recipe run on sgl-project/sglang#40786.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion quant LLM Quantization run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants