Skip to content

[Diffusion][sglang-miles] Bring sglang-miles-h3 H3 rollout commits onto sglang-miles + LoRA-wrap fixes - #40786

Merged
Zhichenzzz merged 6 commits into
sgl-project:sglang-milesfrom
Rockdu:sglang-miles-h3-to-sglang-miles
Sep 25, 2026
Merged

Zhichenzzz merged 6 commits into
sgl-project:sglang-milesfrom
Rockdu:sglang-miles-h3-to-sglang-miles

Conversation

@Rockdu

@Rockdu Rockdu commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Review guide

TL;DR: the sglang-miles-h3 PR + two LoRA fixes.

  • 4 H3 cherry-picks: already reviewed on sglang-miles-h3, applied cleanly. Skim.
  • 2 LoRA fixes: 447815eb7f (Qwen-Image, +1) and 9895674013 (MiniMax-H3, +3). Review these.

Brings the four commits that exist only on sglang-miles-h3 onto the current sglang-miles. It also fixes two sglang-miles fast paths that break once LoRA wraps a layer, which blocked miles_diffusion RL.

Cherry-picks (applied cleanly)

Fixes on top

Both bugs are also on main; a separate PR will port them.

Validation

miles_diffusion e2e is green on this PR (radixark/miles_diffusion#257, 13/13). Four standards were re-recorded for upstream sglang-miles changes, not these commits (#37616 for SD3, #36680 for Qwen-Image, #35796 for H3); attribution is in #257.

🤖 Generated with Claude Code

@github-actions github-actions Bot added lora diffusion SGLang Diffusion labels Sep 22, 2026
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
sglang-miles-h3 loaded both model.safetensors and model.fp16.safetensors
for SD3.5's CLIP text encoders, which left pooled_projections wrong
(cosine 0.33 vs diffusers). sglang-miles carries the loader fix
(sgl-project/sglang#37616), so the pooled embedding now matches
diffusers (cosine 0.99996) and every downstream metric moves.

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35798499525).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
With the LoRA forward overrides dropped from the qwen_image parity patch,
rollout LoRA now runs sgl-d's native path. The train<->rollout residual
grows slightly (log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the
reward series barely moves.

Recorded on the h200/5gpu runner against sgl-project/sglang#40786
(run 35803865496).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
Attribution: sgl-project/sglang#37616 (9cb38a3d57, "filter duplicate
precision variants across custom loaders") is the only commit between
sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f)
that moves this test. A CI-runner bisect (record-e2e-standards on
<commit> + the four sglang-miles-h3 picks) shows the series equal to the
old standard up to 9cb38a3d57^ and equal to this standard from
9cb38a3d57 on.

Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors
for its CLIP text encoders. The old loader globbed both, which left
pooled_projections wrong: cosine 0.33 vs diffusers' encode_prompt, vs
0.99996 after #37616. encoder_hidden_states were identical on both sides.
Every rollout and train metric moves downstream of the pooled embedding.

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35798499525).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
Attribution, split by source:
- Step 0 (first rollout and the first train step): only this PR's
  7d83ef7, which drops the LoRA forward overrides from the qwen_image
  parity patch. The old sglang base (5375babbac + sglang-miles-h3 picks)
  with this branch already gives this standard's step-0 values bit for
  bit.
- Everything after the first weight update: sgl-project/sglang#36680
  (71cee04ebe, "Optimize Qwen-Image TP collectives and attention"), the
  only contributing commit in a CI-runner bisect from 5375babbac to
  5a8da8cc3f with this branch held fixed. It packs add_{q,k,v}_proj into
  one to_added_qkv, so once LoRA is live the added-text Q/K/V run as
  one packed GEMM plus a stacked LoRA delta instead of three GEMMs.

Effect: the train<->rollout residual grows slightly
(log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series
barely moves.

Recorded on the h200/5gpu runner against sgl-project/sglang#40786
(run 35803865496).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
Attribution: sgl-project/sglang#35796 (0447ade326, "fall back to a
component's default attention backend") is the only commit between
sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f)
that moves this test. This comes from a CI-runner bisect of <commit> +
the four sglang-miles-h3 picks (+ the MiniMax-H3 LoRA-wrap fix where
#37903's code exists); 5375babbac reproduces the old standard bit for
bit.

Why: #35796 pins the H3 video VAE's attention to
default_attention_backend=TORCH_SDPA ({FA, TORCH_SDPA} supported), where
it used to inherit the global FA backend. The decoded frames move by
rounding, so PickScore rewards shift in the 5th decimal (step 0 mean
0.74778 -> 0.74774), and the GRPO advantages and train metrics follow.
The DiT path is unchanged and the train<->rollout residual keeps its
magnitude (log_prob_mean_abs_diff 4.13e-5 vs 4.08e-5).

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35812567370).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Rockdu Rockdu changed the title [Diffusion][sglang-miles] Bring sglang-miles-h3 MiniMax H3 rollout commits onto sglang-miles [Diffusion][sglang-miles] Bring sglang-miles-h3 H3 rollout commits onto sglang-miles + LoRA-wrap fixes Sep 23, 2026
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
Attribution: sgl-project/sglang#37616 (9cb38a3d57, "filter duplicate
precision variants across custom loaders") is the only commit between
sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f)
that moves this test. A CI-runner bisect (record-e2e-standards on
<commit> + the four sglang-miles-h3 picks) shows the series equal to the
old standard up to 9cb38a3d57^ and equal to this standard from
9cb38a3d57 on.

Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors
(bit-identical) for its CLIP text encoders. Before #37616, sgl-d's own
CLIP loader saw both and raised "Duplicate tensor names detected across
safetensors files" (517 names for CLIP-G). The pipeline then silently fell
back to transformers' CLIPTextModel, which has no projection head and
drops text_projection.weight ("UNEXPECTED"). So the CLIP-G half of
pooled_projections was the unprojected pooler output: cos -0.016 vs the
projected embedding. CLIP-L's projection is near-identity (cos 1.0). The
concatenated 2048-d vector gives cos 0.333 vs diffusers' encode_prompt,
with norms 50.43 vs 44.68, exactly as measured. After #37616 the sgl-d
CLIPTextModelWithProjection loads, all 149 + 389 params match the HF
files, and cos is 0.99996. encoder_hidden_states (penultimate layer, no
projection) were identical on both sides.

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35798499525).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
Attribution, split by source:
- Step 0 (first rollout and the first train step): only this PR's
  7d83ef7, which drops the LoRA forward overrides from the qwen_image
  parity patch. The old sglang base (5375babbac + sglang-miles-h3 picks)
  with this branch already gives this standard's step-0 values bit for
  bit.
- Everything after the first weight update: sgl-project/sglang#36680
  (71cee04ebe, "Optimize Qwen-Image TP collectives and attention"), the
  only contributing commit in a CI-runner bisect from 5375babbac to
  5a8da8cc3f with this branch held fixed. It packs add_{q,k,v}_proj into
  one to_added_qkv, so once LoRA is live the added-text Q/K/V run as
  one packed GEMM plus a stacked LoRA delta instead of three GEMMs.

Effect: the train<->rollout residual grows slightly
(log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series
barely moves.

Recorded on the h200/5gpu runner against sgl-project/sglang#40786
(run 35803865496).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
Attribution: sgl-project/sglang#35796 (0447ade326, "fall back to a
component's default attention backend") is the only commit between
sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f)
that moves this test. This comes from a CI-runner bisect of <commit> +
the four sglang-miles-h3 picks (+ the MiniMax-H3 LoRA-wrap fix where
#37903's code exists); 5375babbac reproduces the old standard bit for
bit.

Why: #35796 pins the H3 video VAE's attention to
default_attention_backend=TORCH_SDPA ({FA, TORCH_SDPA} supported), where
it used to inherit the global FA backend. The decoded frames move by
rounding, so PickScore rewards shift in the 5th decimal (step 0 mean
0.74778 -> 0.74774), and the GRPO advantages and train metrics follow.
The DiT path is unchanged and the train<->rollout residual keeps its
magnitude (log_prob_mean_abs_diff 4.13e-5 vs 4.08e-5).

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35812567370).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 23, 2026
Attribution: sgl-project/sglang#37616 (9cb38a3d57) only, the same cause
as the SD3 NFT standard. Probes on the CI runner (<commit> + the four
sglang-miles-h3 picks) give: 5375babbac == 9cb38a3d57^ == old standard;
9cb38a3d57 == 5a8da8cc3f == this standard.

Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors
(bit-identical) for its CLIP text encoders. Before #37616, sgl-d's own
CLIP loader saw both and raised "Duplicate tensor names detected across
safetensors files" (517 names for CLIP-G). The pipeline then silently fell
back to transformers' CLIPTextModel, which has no projection head and
drops text_projection.weight ("UNEXPECTED"). So the CLIP-G half of
pooled_projections was the unprojected pooler output: cos -0.016 vs the
projected embedding. CLIP-L's projection is near-identity (cos 1.0). The
concatenated 2048-d vector gives cos 0.333 vs diffusers' encode_prompt,
with norms 50.43 vs 44.68, exactly as measured. After #37616 the sgl-d
CLIPTextModelWithProjection loads, all 149 + 389 params match the HF
files, and cos is 0.99996. encoder_hidden_states (penultimate layer, no
projection) were identical on both sides.

The step-0 mean reward drops 0.416 -> 0.375 here, but that is sampling
noise from the recipe's 8 prompts, not a regression. In a paired run of
128 test prompts x 4 seeds with this recipe's sampling (512 samples per
side), the old base scores 0.4366 and sglang-miles 0.4686: +0.032, 95% CI
[+0.0006, +0.064]. Resampling 8 of those prompts gives a new-old spread of
[-0.087, +0.155], and a <= -0.040 draw has probability ~0.12.

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35927893055).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-wrapped layers

sgl-project#36680 added _split_unquantized_merged_linear, which reads
output_partition_sizes and weight straight off to_added_qkv. Once LoRA wraps
the layer (MergedColumnParallelLinearWithLoRA) the attribute is missing, and
splitting the base weight would drop the LoRA delta anyway. That broke
miles_diffusion test_qwenimage_pickscore_grpo_5xGPU on sglang-miles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Rockdu
Rockdu force-pushed the sglang-miles-h3-to-sglang-miles branch from 0d57314 to 1dab2e0 Compare September 24, 2026 23:45
… path

sgl-project#37903 added _accepts_mxfp8_input, which reads linear.quant_method. A
LoRA wrapper (RowParallelLinearWithLoRA) has no such attribute and must
run its own forward to apply the delta, so it now reports False. This
broke miles_diffusion test_h3_t2va_grpo_2xGPU on sglang-miles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Rockdu
Rockdu force-pushed the sglang-miles-h3-to-sglang-miles branch from 1dab2e0 to 9895674 Compare September 24, 2026 23:46
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 24, 2026
Attribution, split by source:
- Step 0 (first rollout and the first train step): only this PR's
  7d83ef7, which drops the LoRA forward overrides from the qwen_image
  parity patch. The old sglang base (5375babbac + sglang-miles-h3 picks)
  with this branch already gives this standard's step-0 values bit for
  bit.
- Everything after the first weight update: sgl-project/sglang#36680
  (71cee04ebe, "Optimize Qwen-Image TP collectives and attention"), the
  only contributing commit in a CI-runner bisect from 5375babbac to
  5a8da8cc3f with this branch held fixed. It packs add_{q,k,v}_proj into
  one to_added_qkv, so once LoRA is live the added-text Q/K/V run as
  one packed GEMM plus a stacked LoRA delta instead of three GEMMs.

Effect: the train<->rollout residual grows slightly
(log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series
barely moves.

Recorded on the h200/5gpu runner against sgl-project/sglang#40786
(run 35803865496).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 24, 2026
Attribution: sgl-project/sglang#35796 (0447ade326, "fall back to a
component's default attention backend") is the only commit between
sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f)
that moves this test. This comes from a CI-runner bisect of <commit> +
the four sglang-miles-h3 picks (+ the MiniMax-H3 LoRA-wrap fix where
#37903's code exists); 5375babbac reproduces the old standard bit for
bit.

Why: #35796 pins the H3 video VAE's attention to
default_attention_backend=TORCH_SDPA ({FA, TORCH_SDPA} supported), where
it used to inherit the global FA backend. The decoded frames move by
rounding, so PickScore rewards shift in the 5th decimal (step 0 mean
0.74778 -> 0.74774), and the GRPO advantages and train metrics follow.
The DiT path is unchanged and the train<->rollout residual keeps its
magnitude (log_prob_mean_abs_diff 4.13e-5 vs 4.08e-5).

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35812567370).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 24, 2026
Attribution: sgl-project/sglang#37616 (9cb38a3d57) only, the same cause
as the SD3 NFT standard. Probes on the CI runner (<commit> + the four
sglang-miles-h3 picks) give: 5375babbac == 9cb38a3d57^ == old standard;
9cb38a3d57 == 5a8da8cc3f == this standard.

Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors
(bit-identical) for its CLIP text encoders. Before #37616, sgl-d's own
CLIP loader saw both and raised "Duplicate tensor names detected across
safetensors files" (517 names for CLIP-G). The pipeline then silently fell
back to transformers' CLIPTextModel, which has no projection head and
drops text_projection.weight ("UNEXPECTED"). So the CLIP-G half of
pooled_projections was the unprojected pooler output: cos -0.016 vs the
projected embedding. CLIP-L's projection is near-identity (cos 1.0). The
concatenated 2048-d vector gives cos 0.333 vs diffusers' encode_prompt,
with norms 50.43 vs 44.68, exactly as measured. After #37616 the sgl-d
CLIPTextModelWithProjection loads, all 149 + 389 params match the HF
files, and cos is 0.99996. encoder_hidden_states (penultimate layer, no
projection) were identical on both sides.

The step-0 mean reward drops 0.416 -> 0.375 here, but that is sampling
noise from the recipe's 8 prompts, not a regression. In a paired run of
128 test prompts x 4 seeds with this recipe's sampling (512 samples per
side), the old base scores 0.4366 and sglang-miles 0.4686: +0.032, 95% CI
[+0.0006, +0.064]. Resampling 8 of those prompts gives a new-old spread of
[-0.087, +0.155], and a <= -0.040 draw has probability ~0.12.

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35927893055).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Rockdu
Rockdu marked this pull request as ready for review September 24, 2026 23:47
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 25, 2026
The previous curve came from a sglang build whose SD3.5 CLIP text
encoders fell back to transformers' CLIPTextModel and dropped
text_projection (fixed by sgl-project/sglang#37616). The new curve is the
same recipe run on sgl-project/sglang#40786.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

@Zhichenzzz Zhichenzzz left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM if just changing multi-modal

@Zhichenzzz
Zhichenzzz merged commit 106ef6d into sgl-project:sglang-miles Sep 25, 2026
78 of 87 checks passed
Rockdu added a commit to radixark/miles_diffusion that referenced this pull request Sep 25, 2026
sgl-project/sglang#40786 landed the sglang-miles-h3 commits on sglang-miles,
so CI, the Docker image and the docs now follow sglang-miles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ziang-and pushed a commit to zianglih/sglang that referenced this pull request Oct 2, 2026
…PC LoRA updates (sgl-project#40786)

Keep H3 rollout and streamed-weight behavior on the release implementation. The release already contains the LoRA wrapper type guards.

Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com>
Co-authored-by: niehen6174 <niehen6174@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion lora

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants