Repository navigation
[diffusion] Allow Cache-DiT with DiT layerwise offload - #35858
Merged
mickqian merged 6 commits intoAug 30, 2026
Merged
Conversation
Trust only layers that actually ran so a Cache-DiT hit does not leave prefetched weights or empty shells on skipped blocks.
Drop the startup mutex now that skipped blocks are released instead of reused. FSDP remains incompatible.
Update the H3 cookbook, Cache-DiT docs, CLI help, and performance skill. FSDP stays incompatible; quality=high is unchanged.
niehen6174
requested review from
AgainstEntropy,
BBuf,
HaiShaw,
JustinTong0323,
mickqian,
ping1jing2,
sogalin,
wisclmy0611,
yichiche and
zijiexia
as code owners
August 21, 2026 11:02
This was referenced Aug 22, 2026
Keep both the Cache-DiT skip-aware layerwise tests and main's non-layer parking tests.
mickqian
approved these changes
Aug 29, 2026
Collaborator
|
/tag-and-rerun-ci |
StevenChenSE
pushed a commit
to StevenChenSE/sglang
that referenced
this pull request
Sep 6, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
DiT layerwise offload and Cache-DiT can run together. The old startup
ValueErrorwas an implementation accident, not a real mutex.On one 24GB 4090 D, MiniMax-H3
fl2va1344×768 · 107 frames · 50 NFE (kitchen_int8+ FA, DiT+TE layerwise). Layerwise-only vs the same commit without this patch is within noise (+0.1%). Prefetch including last-layer wrap is unchanged.R=0.24W=2MC=16sage_attn+ sameF1B2 W=8 R=0.08 MC=2Recommended quality setting:
Fn=1 Bn=2 W=8 R=0.08 MC=2(30.1 dB, 1.84× vs FA). Spectrum PR #3568411/5/1.0was 285.3 s / 27.4 dB — faster, slightly lower PSNR.Why they conflicted
Layerwise: most weights stay on CPU. The running layer is copied to GPU; the next layer is prefetched in the background.
Cache-DiT: a step may skip middle blocks. DBCache always runs the first
Fnblocks, then on a hit skipsMnand (optionally) runs the lastBnblocks. Skipped blocks are neverforwarded.These axes are orthogonal. Skipping a block means less H2D, not more.Layerwise assumed every step walks
0..N-1in order:i, prefetchi+1.% Nwraps to the next step's layer 0.A Cache-DiT hit is
0 ──skip 1–5──→ 6 → 7(8-layer,Fn=1,Bn=2):Layer 1 is prefetched during layer 0 and never posts a release hook, or a wrap/release leaves
empty((1,))on the next compute layer → shape mismatch. The old fix banned the combination at startup.Default
Bn=0: a hit ends after layer 0. Layer 1 sits on GPU until the next step unless prepare drops it.What changed
Trust only layers that actually ran. A full-stack step still prefetches as before (including last-layer wrap).
preparedrops leftover prefetch (Bn=0).No
Fn/Bnplan is published. The first layer after a skip may sync-load one layer of PCIe. DefaultBn=0has no Bn to load.i+1/ wrap → empty / crashValueErrorTests
MiniMax-H3 on 4090 D
Protocol unless noted: 1344×768, 107 frames,
kitchen_int8+ FA,--performance-mode memory,--layerwise-offload-components dit,text_encoder.Cache-DiT without DiT layerwise OOM on 24GB (22.94 GiB used, +932 MiB). Same on baseline. Not a regression.
Scheme sweep (same process, forced skip so DBCache actually jumps). 8 requested steps → 7 NFE; SCM rerun used 9 steps → 8 NFE because
cache_dit.steps_maskonly allows 4 or 6 whentotal_steps < 8. That limit is upstream Cache-DiT, not this patch. No shape-mismatch / empty weight crash on any scheme.F1B0F1B2F2B0F1B0+ TaylorSeer O1F1B0+ SCMfastF1B0+ SCMmediumDo not stack Spectrum / TeaCache with Cache-DiT. Hooks cover whole-block
forward()only.Example
Layerwise only video:
cache-dit.mp4
Layerwise +cache-dit video:
original.mp4
Need
comfy-kitchenforkitchen_int8.CI States
Latest PR Test (Base): ✅ Run #33241840779
Latest PR Test (Extra): ❌ Run #33241840609
Latest PR Test (AMD ROCm 7.2): ❌ Run #33241840692