Skip to content

bump SGLang-diffusion to sglang-miles branch - #257

Merged
Rockdu merged 9 commits into
mainfrom
ci/sglang-miles-h3-to-sglang-miles
Sep 26, 2026
Merged

Rockdu merged 9 commits into
mainfrom
ci/sglang-miles-h3-to-sglang-miles

Conversation

@Rockdu

@Rockdu Rockdu commented Sep 22, 2026 •

Copy link
Copy Markdown
Collaborator

Moves miles_diffusion from sglang-miles-h3 to sglang-miles, now that sgl-project/sglang#40786 (the four H3 picks plus two LoRA-wrap fixes) has landed there.

Changes

  • ci: build the rollout engine from sglang-miles (DEFAULT_SGLANG_REF in .github/workflows/_run-ci.yml, SGLANG_DIFFUSION_BRANCH in docker/Dockerfile, and the docs).
  • qwen_image rollout patch: drop the LoRA forward overrides, since they break on sglang's packed added-QKV layer (#36680). Also follow the new apply_qk_norm_with_optional_rope signature and keep the fused QK-norm/RoPE epilogue off.
  • Re-record four e2e standards for upstream sglang-miles changes (details in each commit message):
    • SD3 NFT, SD3.5 OCR: #37616 (CLIP text encoders now keep text_projection)
    • Qwen-Image: the patch change above + #36680
    • MiniMax-H3: #35796 (H3 VAE attention defaults to SDPA)
  • Docs: replace the SD3 DiffusionNFT PickScore curve and eval numbers with a run on #40786.

Validation

e2e 13/13 green against #40786. 400-rollout Qwen-Image recipe run: https://wandb.ai/radixarkai/miles-diffusion-CIs/runs/npcsyfm8

🤖 Generated with Claude Code

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Rockdu Rockdu added run-ci-e2e Run e2e metric-regression tests on this PR and removed run-ci-e2e Run e2e metric-regression tests on this PR labels Sep 22, 2026
@Rockdu Rockdu added run-ci-e2e Run e2e metric-regression tests on this PR and removed run-ci-e2e Run e2e metric-regression tests on this PR labels Sep 22, 2026
@Rockdu
Rockdu force-pushed the ci/sglang-miles-h3-to-sglang-miles branch 2 times, most recently from 5b2b56a to 650be62 Compare September 23, 2026 15:34
@Rockdu Rockdu changed the title [DO NOT MERGE] ci: e2e against sglang-miles + H3 rollout commits (sglang#40786) fix: run on sglang-miles (Qwen-Image parity patch + attributed e2e standards) Sep 23, 2026
Rockdu and others added 2 commits September 23, 2026 15:50
Attribution: sgl-project/sglang#37616 (9cb38a3d57, "filter duplicate
precision variants across custom loaders") is the only commit between
sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f)
that moves this test. A CI-runner bisect (record-e2e-standards on
<commit> + the four sglang-miles-h3 picks) shows the series equal to the
old standard up to 9cb38a3d57^ and equal to this standard from
9cb38a3d57 on.

Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors
(bit-identical) for its CLIP text encoders. Before #37616, sgl-d's own
CLIP loader saw both and raised "Duplicate tensor names detected across
safetensors files" (517 names for CLIP-G). The pipeline then silently fell
back to transformers' CLIPTextModel, which has no projection head and
drops text_projection.weight ("UNEXPECTED"). So the CLIP-G half of
pooled_projections was the unprojected pooler output: cos -0.016 vs the
projected embedding. CLIP-L's projection is near-identity (cos 1.0). The
concatenated 2048-d vector gives cos 0.333 vs diffusers' encode_prompt,
with norms 50.43 vs 44.68, exactly as measured. After #37616 the sgl-d
CLIPTextModelWithProjection loads, all 149 + 389 params match the HF
files, and cos is 0.99996. encoder_hidden_states (penultimate layer, no
projection) were identical on both sides.

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35798499525).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ity patch

sgl-project/sglang#36680 packs Qwen-Image's add_{q,k,v}_proj into one
MergedColumnParallelLinear (to_added_qkv). The parity patch forced every
LoRA layer through a 2D PEFT-ordered delta, which crashes on the merged
layer's 3D stacked LoRA (size 8 vs 3072). Let sgl-d's native LoRA
forwards run; the norm/RoPE/split_seqs parity patches stay.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Rockdu
Rockdu force-pushed the ci/sglang-miles-h3-to-sglang-miles branch from 650be62 to 9b3ba42 Compare September 23, 2026 22:50
Rockdu and others added 4 commits September 24, 2026 16:46
…used epilogue off

sgl-project/sglang#33555 passes freqs_complex to
apply_qk_norm_with_optional_rope, which the patched replacement rejected.
The new fused QK-norm/RoPE epilogue (sm90, unquantized, on by default)
also bypasses the patched norms and RoPE, so disable it when the patch
group is applied.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Attribution, split by source:
- Step 0 (first rollout and the first train step): only this PR's
  7d83ef7, which drops the LoRA forward overrides from the qwen_image
  parity patch. The old sglang base (5375babbac + sglang-miles-h3 picks)
  with this branch already gives this standard's step-0 values bit for
  bit.
- Everything after the first weight update: sgl-project/sglang#36680
  (71cee04ebe, "Optimize Qwen-Image TP collectives and attention"), the
  only contributing commit in a CI-runner bisect from 5375babbac to
  5a8da8cc3f with this branch held fixed. It packs add_{q,k,v}_proj into
  one to_added_qkv, so once LoRA is live the added-text Q/K/V run as
  one packed GEMM plus a stacked LoRA delta instead of three GEMMs.

Effect: the train<->rollout residual grows slightly
(log_prob_mean_abs_diff ~3.1e-5 -> 3.2-4.7e-5) and the reward series
barely moves.

Recorded on the h200/5gpu runner against sgl-project/sglang#40786
(run 35803865496).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Attribution: sgl-project/sglang#35796 (0447ade326, "fall back to a
component's default attention backend") is the only commit between
sglang-miles-h3's merge base (5375babbac) and sglang-miles (5a8da8cc3f)
that moves this test. This comes from a CI-runner bisect of <commit> +
the four sglang-miles-h3 picks (+ the MiniMax-H3 LoRA-wrap fix where
#37903's code exists); 5375babbac reproduces the old standard bit for
bit.

Why: #35796 pins the H3 video VAE's attention to
default_attention_backend=TORCH_SDPA ({FA, TORCH_SDPA} supported), where
it used to inherit the global FA backend. The decoded frames move by
rounding, so PickScore rewards shift in the 5th decimal (step 0 mean
0.74778 -> 0.74774), and the GRPO advantages and train metrics follow.
The DiT path is unchanged and the train<->rollout residual keeps its
magnitude (log_prob_mean_abs_diff 4.13e-5 vs 4.08e-5).

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35812567370).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Attribution: sgl-project/sglang#37616 (9cb38a3d57) only, the same cause
as the SD3 NFT standard. Probes on the CI runner (<commit> + the four
sglang-miles-h3 picks) give: 5375babbac == 9cb38a3d57^ == old standard;
9cb38a3d57 == 5a8da8cc3f == this standard.

Why: SD3.5-medium ships both model.safetensors and model.fp16.safetensors
(bit-identical) for its CLIP text encoders. Before #37616, sgl-d's own
CLIP loader saw both and raised "Duplicate tensor names detected across
safetensors files" (517 names for CLIP-G). The pipeline then silently fell
back to transformers' CLIPTextModel, which has no projection head and
drops text_projection.weight ("UNEXPECTED"). So the CLIP-G half of
pooled_projections was the unprojected pooler output: cos -0.016 vs the
projected embedding. CLIP-L's projection is near-identity (cos 1.0). The
concatenated 2048-d vector gives cos 0.333 vs diffusers' encode_prompt,
with norms 50.43 vs 44.68, exactly as measured. After #37616 the sgl-d
CLIPTextModelWithProjection loads, all 149 + 389 params match the HF
files, and cos is 0.99996. encoder_hidden_states (penultimate layer, no
projection) were identical on both sides.

The step-0 mean reward drops 0.416 -> 0.375 here, but that is sampling
noise from the recipe's 8 prompts, not a regression. In a paired run of
128 test prompts x 4 seeds with this recipe's sampling (512 samples per
side), the old base scores 0.4366 and sglang-miles 0.4686: +0.032, 95% CI
[+0.0006, +0.064]. Resampling 8 of those prompts gives a new-old spread of
[-0.087, +0.155], and a <= -0.040 draw has probability ~0.12.

Recorded on the h200/3gpu runner against sgl-project/sglang#40786
(run 35927893055).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Rockdu
Rockdu force-pushed the ci/sglang-miles-h3-to-sglang-miles branch from 9b3ba42 to 944201c Compare September 24, 2026 23:47
The previous curve came from a sglang build whose SD3.5 CLIP text
encoders fell back to transformers' CLIPTextModel and dropped
text_projection (fixed by sgl-project/sglang#37616). The new curve is the
same recipe run on sgl-project/sglang#40786.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Rockdu
Rockdu marked this pull request as ready for review September 25, 2026 22:47
sgl-project/sglang#40786 landed the sglang-miles-h3 commits on sglang-miles,
so CI, the Docker image and the docs now follow sglang-miles.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Rockdu Rockdu added run-ci-e2e Run e2e metric-regression tests on this PR and removed run-ci-e2e Run e2e metric-regression tests on this PR labels Sep 25, 2026
@Rockdu Rockdu changed the title fix: run on sglang-miles (Qwen-Image parity patch + attributed e2e standards) bump SGLang-diffusion to sglang-miles branch Sep 25, 2026
Rockdu added a commit that referenced this pull request Sep 25, 2026
One run of scripts/run_diffusion_grpo_pickscore_5gpu_flowgrpo_aligned.py
(wandb radixarkai/miles-diffusion-CIs/runs/npcsyfm8) on the code of #257:
eval/pickscore_test rises from 0.859 at rollout 29 to 0.900 at rollout 389.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Rockdu added a commit that referenced this pull request Sep 25, 2026
One run of scripts/run_diffusion_grpo_pickscore_5gpu_flowgrpo_aligned.py
(wandb radixarkai/miles-diffusion-CIs/runs/npcsyfm8) on the code of #257:
eval/pickscore_test rises from 0.859 at rollout 29 to 0.900 at rollout 389.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@Rockdu
Rockdu merged commit e2cfb16 into main Sep 26, 2026
25 of 30 checks passed
OnePunchMonk pushed a commit to OnePunchMonk/miles_diffusion that referenced this pull request Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci-e2e Run e2e metric-regression tests on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant