Skip to content

[diffusion] Allow Cache-DiT with DiT layerwise offload - #35858

Merged
mickqian merged 6 commits into
sgl-project:mainfrom
niehen6174:feat/layerwise-cache-dit-coexist
Aug 30, 2026
Merged

mickqian merged 6 commits into
sgl-project:mainfrom
niehen6174:feat/layerwise-cache-dit-coexist

Conversation

@niehen6174

@niehen6174 niehen6174 commented Aug 21, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

DiT layerwise offload and Cache-DiT can run together. The old startup ValueError was an implementation accident, not a real mutex.

On one 24GB 4090 D, MiniMax-H3 fl2va 1344×768 · 107 frames · 50 NFE (kitchen_int8 + FA, DiT+TE layerwise). Layerwise-only vs the same commit without this patch is within noise (+0.1%). Prefetch including last-layer wrap is unchanged.

Config timed e2e Denoise vs FA PSNR vs FA
Layerwise only 721.2 s 700.9 s — —
+ Cache-DiT R=0.24 W=2 MC=16 173.2 s 153.9 s 4.16× 15.90dB
+ sage_attn + same 106.4 s 87.1 s 6.78× 15.89dB
+ Cache-DiT F1B2 W=8 R=0.08 MC=2 392.3 s 372.7 s 1.84× 30.10dB

Recommended quality setting: Fn=1 Bn=2 W=8 R=0.08 MC=2 (30.1 dB, 1.84× vs FA). Spectrum PR #35684 11/5/1.0 was 285.3 s / 27.4 dB — faster, slightly lower PSNR.

Why they conflicted

Layerwise: most weights stay on CPU. The running layer is copied to GPU; the next layer is prefetched in the background.

Cache-DiT: a step may skip middle blocks. DBCache always runs the first Fn blocks, then on a hit skips Mn and (optionally) runs the last Bn blocks. Skipped blocks are never forwarded.These axes are orthogonal. Skipping a block means less H2D, not more.

Layerwise assumed every step walks 0..N-1 in order:

  1. After layer i, prefetch i+1.
  2. Near the tail, % N wraps to the next step's layer 0.

A Cache-DiT hit is 0 ──skip 1–5──→ 6 → 7 (8-layer, Fn=1,Bn=2):

Layerwise assumed:  0 → 1 → 2 → 3 → 4 → 5 → 6 → 7
Cache-DiT actually: 0 ──────── skip 1–5 ──────── → 6 → 7

Layer 1 is prefetched during layer 0 and never posts a release hook, or a wrap/release leaves empty((1,)) on the next compute layer → shape mismatch. The old fix banned the combination at startup.

Default Bn=0: a hit ends after layer 0. Layer 1 sits on GPU until the next step unless prepare drops it.

What changed

Trust only layers that actually ran. A full-stack step still prefetches as before (including last-layer wrap).

  • Jump: release the unused gap; sync-load the destination if needed.
  • Next step: prepare drops leftover prefetch (Bn=0).
  • Startup allows Cache-DiT + DiT layerwise. FSDP stays incompatible.

No Fn/Bn plan is published. The first layer after a skip may sync-load one layer of PCIe. Default Bn=0 has no Bn to load.

Before After
Full stack (layerwise only) Sequential prefetch + wrap Same
Cache-DiT skip Prefetch i+1 / wrap → empty / crash Drop the gap, sync-load dest
Combined Startup ValueError Allowed

Tests

MiniMax-H3 on 4090 D

Protocol unless noted: 1344×768, 107 frames, kitchen_int8 + FA, --performance-mode memory, --layerwise-offload-components dit,text_encoder.

Cache-DiT without DiT layerwise OOM on 24GB (22.94 GiB used, +932 MiB). Same on baseline. Not a regression.

Scheme sweep (same process, forced skip so DBCache actually jumps). 8 requested steps → 7 NFE; SCM rerun used 9 steps → 8 NFE because cache_dit.steps_mask only allows 4 or 6 when total_steps < 8. That limit is upstream Cache-DiT, not this patch. No shape-mismatch / empty weight crash on any scheme.

Scheme Steps Denoise E2E Peak Result
Layerwise only 8 97.6 s 134.4 s 21.1 GB ok
DBCache F1B0 8 29.3 s 47.9 s 18.4 GB ok
DBCache F1B2 8 32.0 s 50.7 s 18.1 GB ok
DBCache F2B0 8 30.6 s 49.4 s 17.5 GB ok
F1B0 + TaylorSeer O1 8 29.3 s 47.8 s 17.6 GB ok
F1B0 + SCM fast 9 71.2 s 115.3 s 22.9 GB ok (5 compute / 3 cache)
F1B0 + SCM medium 9 83.5 s 102.4 s 18.2 GB ok (6 compute / 2 cache)

Do not stack Spectrum / TeaCache with Cache-DiT. Hooks cover whole-block forward() only.

Example

# Do not pass --quality (H3 generic Cache-DiT is disabled if quality is explicit).
sglang generate \
  --model-path MiniMaxAI/MiniMax-H3 \
  --model-variant fl2va \
  --quantization kitchen_int8 \
  --attention-backend fa \
  --performance-mode memory \
  --layerwise-offload-components dit,text_encoder \
  --dit-offload-prefetch-size 1 \
  --dit-layerwise-resident-layers 0 \
  --enable-cache-dit true \
  --cache-dit-params '{"Fn_compute_blocks":1,"Bn_compute_blocks":2,"max_warmup_steps":8,"residual_diff_threshold":0.08,"max_continuous_cached_steps":2}'

Layerwise only video:

cache-dit.mp4

Layerwise +cache-dit video:

original.mp4

Need comfy-kitchen for kitchen_int8.


CI States

Latest PR Test (Base): ✅ Run #33241840779
Latest PR Test (Extra): ❌ Run #33241840609
Latest PR Test (AMD ROCm 7.2): ❌ Run #33241840692

Trust only layers that actually ran so a Cache-DiT hit does not leave
prefetched weights or empty shells on skipped blocks.
Drop the startup mutex now that skipped blocks are released instead of
reused. FSDP remains incompatible.
Update the H3 cookbook, Cache-DiT docs, CLI help, and performance skill.
FSDP stays incompatible; quality=high is unchanged.
@github-actions github-actions Bot added documentation Improvements or additions to documentation diffusion SGLang Diffusion labels Aug 21, 2026
Keep both the Cache-DiT skip-aware layerwise tests and main's non-layer parking tests.
@mickqian

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Aug 29, 2026
@mickqian
mickqian merged commit e9a7157 into sgl-project:main Aug 30, 2026
134 of 144 checks passed
StevenChenSE pushed a commit to StevenChenSE/sglang that referenced this pull request Sep 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion documentation Improvements or additions to documentation run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants