Repository navigation
[Diffusion][sglang-miles] Bring sglang-miles-h3 H3 rollout commits onto sglang-miles + LoRA-wrap fixes - #40786
Merged
Zhichenzzz merged 6 commits intoSep 25, 2026
Conversation
…sgl-project#34365 (sgl-project#35598) Co-authored-by: niehen6174 <niehen6174@gmail.com>
…ing headers, opt-in uint8 video (sgl-project#36754)
This was referenced Sep 22, 2026
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 23, 2026
sglang-miles-h3 loaded both model.safetensors and model.fp16.safetensors for SD3.5's CLIP text encoders, which left pooled_projections wrong (cosine 0.33 vs diffusers). sglang-miles carries the loader fix (sgl-project/sglang#37616), so the pooled embedding now matches diffusers (cosine 0.99996) and every downstream metric moves. Recorded on the h200/3gpu runner against sgl-project/sglang#40786 (run 35798499525). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 23, 2026
With the LoRA forward overrides dropped from the qwen_image parity patch, rollout LoRA now runs sgl-d's native path. The train<->rollout residual grows slightly (log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series barely moves. Recorded on the h200/5gpu runner against sgl-project/sglang#40786 (run 35803865496). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 23, 2026
Attribution: sgl-project/sglang#37616 (9cb38a3d57, "filter duplicate precision variants across custom loaders") is the only commit between sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f) that moves this test. A CI-runner bisect (record-e2e-standards on <commit> + the four sglang-miles-h3 picks) shows the series equal to the old standard up to 9cb38a3d57^ and equal to this standard from 9cb38a3d57 on. Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors for its CLIP text encoders. The old loader globbed both, which left pooled_projections wrong: cosine 0.33 vs diffusers' encode_prompt, vs 0.99996 after #37616. encoder_hidden_states were identical on both sides. Every rollout and train metric moves downstream of the pooled embedding. Recorded on the h200/3gpu runner against sgl-project/sglang#40786 (run 35798499525). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 23, 2026
Attribution, split by source: - Step 0 (first rollout and the first train step): only this PR's 7d83ef7, which drops the LoRA forward overrides from the qwen_image parity patch. The old sglang base (5375babbac + sglang-miles-h3 picks) with this branch already gives this standard's step-0 values bit for bit. - Everything after the first weight update: sgl-project/sglang#36680 (71cee04ebe, "Optimize Qwen-Image TP collectives and attention"), the only contributing commit in a CI-runner bisect from 5375babbac to 5a8da8cc3f with this branch held fixed. It packs add_{q,k,v}_proj into one to_added_qkv, so once LoRA is live the added-text Q/K/V run as one packed GEMM plus a stacked LoRA delta instead of three GEMMs. Effect: the train<->rollout residual grows slightly (log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series barely moves. Recorded on the h200/5gpu runner against sgl-project/sglang#40786 (run 35803865496). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 23, 2026
Attribution: sgl-project/sglang#35796 (0447ade326, "fall back to a component's default attention backend") is the only commit between sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f) that moves this test. This comes from a CI-runner bisect of <commit> + the four sglang-miles-h3 picks (+ the MiniMax-H3 LoRA-wrap fix where #37903's code exists); 5375babbac reproduces the old standard bit for bit. Why: #35796 pins the H3 video VAE's attention to default_attention_backend=TORCH_SDPA ({FA, TORCH_SDPA} supported), where it used to inherit the global FA backend. The decoded frames move by rounding, so PickScore rewards shift in the 5th decimal (step 0 mean 0.74778 -> 0.74774), and the GRPO advantages and train metrics follow. The DiT path is unchanged and the train<->rollout residual keeps its magnitude (log_prob_mean_abs_diff 4.13e-5 vs 4.08e-5). Recorded on the h200/3gpu runner against sgl-project/sglang#40786 (run 35812567370). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 23, 2026
Attribution: sgl-project/sglang#37616 (9cb38a3d57, "filter duplicate precision variants across custom loaders") is the only commit between sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f) that moves this test. A CI-runner bisect (record-e2e-standards on <commit> + the four sglang-miles-h3 picks) shows the series equal to the old standard up to 9cb38a3d57^ and equal to this standard from 9cb38a3d57 on. Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors (bit-identical) for its CLIP text encoders. Before #37616, sgl-d's own CLIP loader saw both and raised "Duplicate tensor names detected across safetensors files" (517 names for CLIP-G). The pipeline then silently fell back to transformers' CLIPTextModel, which has no projection head and drops text_projection.weight ("UNEXPECTED"). So the CLIP-G half of pooled_projections was the unprojected pooler output: cos -0.016 vs the projected embedding. CLIP-L's projection is near-identity (cos 1.0). The concatenated 2048-d vector gives cos 0.333 vs diffusers' encode_prompt, with norms 50.43 vs 44.68, exactly as measured. After #37616 the sgl-d CLIPTextModelWithProjection loads, all 149 + 389 params match the HF files, and cos is 0.99996. encoder_hidden_states (penultimate layer, no projection) were identical on both sides. Recorded on the h200/3gpu runner against sgl-project/sglang#40786 (run 35798499525). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 23, 2026
Attribution, split by source: - Step 0 (first rollout and the first train step): only this PR's 7d83ef7, which drops the LoRA forward overrides from the qwen_image parity patch. The old sglang base (5375babbac + sglang-miles-h3 picks) with this branch already gives this standard's step-0 values bit for bit. - Everything after the first weight update: sgl-project/sglang#36680 (71cee04ebe, "Optimize Qwen-Image TP collectives and attention"), the only contributing commit in a CI-runner bisect from 5375babbac to 5a8da8cc3f with this branch held fixed. It packs add_{q,k,v}_proj into one to_added_qkv, so once LoRA is live the added-text Q/K/V run as one packed GEMM plus a stacked LoRA delta instead of three GEMMs. Effect: the train<->rollout residual grows slightly (log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series barely moves. Recorded on the h200/5gpu runner against sgl-project/sglang#40786 (run 35803865496). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 23, 2026
Attribution: sgl-project/sglang#35796 (0447ade326, "fall back to a component's default attention backend") is the only commit between sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f) that moves this test. This comes from a CI-runner bisect of <commit> + the four sglang-miles-h3 picks (+ the MiniMax-H3 LoRA-wrap fix where #37903's code exists); 5375babbac reproduces the old standard bit for bit. Why: #35796 pins the H3 video VAE's attention to default_attention_backend=TORCH_SDPA ({FA, TORCH_SDPA} supported), where it used to inherit the global FA backend. The decoded frames move by rounding, so PickScore rewards shift in the 5th decimal (step 0 mean 0.74778 -> 0.74774), and the GRPO advantages and train metrics follow. The DiT path is unchanged and the train<->rollout residual keeps its magnitude (log_prob_mean_abs_diff 4.13e-5 vs 4.08e-5). Recorded on the h200/3gpu runner against sgl-project/sglang#40786 (run 35812567370). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 23, 2026
Attribution: sgl-project/sglang#37616 (9cb38a3d57) only, the same cause as the SD3 NFT standard. Probes on the CI runner (<commit> + the four sglang-miles-h3 picks) give: 5375babbac == 9cb38a3d57^ == old standard; 9cb38a3d57 == 5a8da8cc3f == this standard. Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors (bit-identical) for its CLIP text encoders. Before #37616, sgl-d's own CLIP loader saw both and raised "Duplicate tensor names detected across safetensors files" (517 names for CLIP-G). The pipeline then silently fell back to transformers' CLIPTextModel, which has no projection head and drops text_projection.weight ("UNEXPECTED"). So the CLIP-G half of pooled_projections was the unprojected pooler output: cos -0.016 vs the projected embedding. CLIP-L's projection is near-identity (cos 1.0). The concatenated 2048-d vector gives cos 0.333 vs diffusers' encode_prompt, with norms 50.43 vs 44.68, exactly as measured. After #37616 the sgl-d CLIPTextModelWithProjection loads, all 149 + 389 params match the HF files, and cos is 0.99996. encoder_hidden_states (penultimate layer, no projection) were identical on both sides. The step-0 mean reward drops 0.416 -> 0.375 here, but that is sampling noise from the recipe's 8 prompts, not a regression. In a paired run of 128 test prompts x 4 seeds with this recipe's sampling (512 samples per side), the old base scores 0.4366 and sglang-miles 0.4686: +0.032, 95% CI [+0.0006, +0.064]. Resampling 8 of those prompts gives a new-old spread of [-0.087, +0.155], and a <= -0.040 draw has probability ~0.12. Recorded on the h200/3gpu runner against sgl-project/sglang#40786 (run 35927893055). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-wrapped layers sgl-project#36680 added _split_unquantized_merged_linear, which reads output_partition_sizes and weight straight off to_added_qkv. Once LoRA wraps the layer (MergedColumnParallelLinearWithLoRA) the attribute is missing, and splitting the base weight would drop the LoRA delta anyway. That broke miles_diffusion test_qwenimage_pickscore_grpo_5xGPU on sglang-miles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu
force-pushed
the
sglang-miles-h3-to-sglang-miles
branch
from
September 24, 2026 23:45
0d57314 to
1dab2e0
Compare
… path sgl-project#37903 added _accepts_mxfp8_input, which reads linear.quant_method. A LoRA wrapper (RowParallelLinearWithLoRA) has no such attribute and must run its own forward to apply the delta, so it now reports False. This broke miles_diffusion test_h3_t2va_grpo_2xGPU on sglang-miles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu
force-pushed
the
sglang-miles-h3-to-sglang-miles
branch
from
September 24, 2026 23:46
1dab2e0 to
9895674
Compare
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 24, 2026
Attribution, split by source: - Step 0 (first rollout and the first train step): only this PR's 7d83ef7, which drops the LoRA forward overrides from the qwen_image parity patch. The old sglang base (5375babbac + sglang-miles-h3 picks) with this branch already gives this standard's step-0 values bit for bit. - Everything after the first weight update: sgl-project/sglang#36680 (71cee04ebe, "Optimize Qwen-Image TP collectives and attention"), the only contributing commit in a CI-runner bisect from 5375babbac to 5a8da8cc3f with this branch held fixed. It packs add_{q,k,v}_proj into one to_added_qkv, so once LoRA is live the added-text Q/K/V run as one packed GEMM plus a stacked LoRA delta instead of three GEMMs. Effect: the train<->rollout residual grows slightly (log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series barely moves. Recorded on the h200/5gpu runner against sgl-project/sglang#40786 (run 35803865496). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 24, 2026
Attribution: sgl-project/sglang#35796 (0447ade326, "fall back to a component's default attention backend") is the only commit between sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f) that moves this test. This comes from a CI-runner bisect of <commit> + the four sglang-miles-h3 picks (+ the MiniMax-H3 LoRA-wrap fix where #37903's code exists); 5375babbac reproduces the old standard bit for bit. Why: #35796 pins the H3 video VAE's attention to default_attention_backend=TORCH_SDPA ({FA, TORCH_SDPA} supported), where it used to inherit the global FA backend. The decoded frames move by rounding, so PickScore rewards shift in the 5th decimal (step 0 mean 0.74778 -> 0.74774), and the GRPO advantages and train metrics follow. The DiT path is unchanged and the train<->rollout residual keeps its magnitude (log_prob_mean_abs_diff 4.13e-5 vs 4.08e-5). Recorded on the h200/3gpu runner against sgl-project/sglang#40786 (run 35812567370). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 24, 2026
Attribution: sgl-project/sglang#37616 (9cb38a3d57) only, the same cause as the SD3 NFT standard. Probes on the CI runner (<commit> + the four sglang-miles-h3 picks) give: 5375babbac == 9cb38a3d57^ == old standard; 9cb38a3d57 == 5a8da8cc3f == this standard. Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors (bit-identical) for its CLIP text encoders. Before #37616, sgl-d's own CLIP loader saw both and raised "Duplicate tensor names detected across safetensors files" (517 names for CLIP-G). The pipeline then silently fell back to transformers' CLIPTextModel, which has no projection head and drops text_projection.weight ("UNEXPECTED"). So the CLIP-G half of pooled_projections was the unprojected pooler output: cos -0.016 vs the projected embedding. CLIP-L's projection is near-identity (cos 1.0). The concatenated 2048-d vector gives cos 0.333 vs diffusers' encode_prompt, with norms 50.43 vs 44.68, exactly as measured. After #37616 the sgl-d CLIPTextModelWithProjection loads, all 149 + 389 params match the HF files, and cos is 0.99996. encoder_hidden_states (penultimate layer, no projection) were identical on both sides. The step-0 mean reward drops 0.416 -> 0.375 here, but that is sampling noise from the recipe's 8 prompts, not a regression. In a paired run of 128 test prompts x 4 seeds with this recipe's sampling (512 samples per side), the old base scores 0.4366 and sglang-miles 0.4686: +0.032, 95% CI [+0.0006, +0.064]. Resampling 8 of those prompts gives a new-old spread of [-0.087, +0.155], and a <= -0.040 draw has probability ~0.12. Recorded on the h200/3gpu runner against sgl-project/sglang#40786 (run 35927893055). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu
marked this pull request as ready for review
September 24, 2026 23:47
Rockdu
requested review from
HaiShaw,
mickqian,
ping1jing2 and
yichiche
as code owners
September 24, 2026 23:47
Rockdu
requested review from
AgainstEntropy,
BBuf and
kevin-mii
as code owners
September 24, 2026 23:47
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 25, 2026
The previous curve came from a sglang build whose SD3.5 CLIP text encoders fell back to transformers' CLIPTextModel and dropped text_projection (fixed by sgl-project/sglang#37616). The new curve is the same recipe run on sgl-project/sglang#40786. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Zhichenzzz
approved these changes
Sep 25, 2026
Zhichenzzz
left a comment
Collaborator
There was a problem hiding this comment.
LGTM if just changing multi-modal
Rockdu
added a commit
to radixark/miles_diffusion
that referenced
this pull request
Sep 25, 2026
sgl-project/sglang#40786 landed the sglang-miles-h3 commits on sglang-miles, so CI, the Docker image and the docs now follow sglang-miles. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ziang-and
pushed a commit
to zianglih/sglang
that referenced
this pull request
Oct 2, 2026
…PC LoRA updates (sgl-project#40786) Keep H3 rollout and streamed-weight behavior on the release implementation. The release already contains the LoRA wrapper type guards. Co-authored-by: WenhaoZhang <42087078+niehen6174@users.noreply.github.com> Co-authored-by: niehen6174 <niehen6174@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Review guide
TL;DR: the
sglang-miles-h3PR + two LoRA fixes.sglang-miles-h3, applied cleanly. Skim.447815eb7f(Qwen-Image, +1) and9895674013(MiniMax-H3, +3). Review these.Brings the four commits that exist only on
sglang-miles-h3onto the currentsglang-miles. It also fixes two sglang-miles fast paths that break once LoRA wraps a layer, which blocked miles_diffusion RL.Cherry-picks (applied cleanly)
latent_step_indices; without it Wan22 GRPO fails on plainsglang-miles.Fixes on top
447815eb7f):_get_added_qkv_projectionsonly takes [Diffusion] Optimize Qwen-Image TP collectives and attention #36680's packed split path for a bareMergedColumnParallelLinear. On a LoRA wrapper it raisedAttributeError: ... no attribute 'output_partition_sizes', and it would also have dropped the LoRA delta.9895674013):_accepts_mxfp8_input([diffusion] model: support VDN-H3 (hybrid window softmax + Video Delta linear attention MiniMax-H3, 8-NFE distill) with a hybrid_window_attn_h3 backend #37903) returns False for anything that is not aLinearBase, so LoRA wrappers stay off the MXFP8 input path. Before this it raisedAttributeError: 'RowParallelLinearWithLoRA' object has no attribute 'quant_method'.Both bugs are also on
main; a separate PR will port them.Validation
miles_diffusion e2e is green on this PR (radixark/miles_diffusion#257, 13/13). Four standards were re-recorded for upstream
sglang-mileschanges, not these commits (#37616 for SD3, #36680 for Qwen-Image, #35796 for H3); attribution is in #257.🤖 Generated with Claude Code