Skip to content

[Diffusion] Separate a use-scoped layerwise release from release_all - #40590

Merged
mickqian merged 1 commit into
sgl-project:mainfrom
mickqian:mick/layerwise-end-use-api
Sep 22, 2026
Merged

mickqian merged 1 commit into
sgl-project:mainfrom
mickqian:mick/layerwise-end-use-api

Conversation

@mickqian

@mickqian mickqian commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator

First of three, splitting a naming problem from the behaviour change it hides.

The two questions that were one call

LayerwiseOffloadStrategy.finish_use called manager.release_all(), and
release_all did what its name says — dropped every layer, the resident set
included. But its two callers want different things:

caller what it means
enable_offload a full reset: sync to CPU, leave nothing on the device
finish_use this component's use ended

A use ending is the narrower question. The streamed window is certainly dead;
whether the resident set is depends on whether anything will want it before
another stage needs the room. There was no way to say that, so finish_use said
"drop everything".

Why it matters

That is why --layerwise-resident-layers does nothing for any component whose
use is a single forward pass rather than a denoise loop. The set is prefetched at
prepare_for_use and dropped at finish_use, every request. For the DiT one use
spans all the denoise steps, so the set survives the loop and earns its memory
back — which is the case the feature was built for, under its original name
--dit-layerwise-resident-layers. The per-component override generalised the
flag to other components and inherited a DiT-shaped scope with it.

Measured on Qwen-Image-2.1 / one RTX 5090, five requests per arm:

p50 s/step steady VRAM the manager reports
baseline 14.136 s 0.3173 17047 MiB resident=0/66
--layerwise-resident-layers text_encoder=0.8 14.168 s 0.3173 17047 MiB resident=53/66

The flag is accepted, resident=53/66 is logged, and nothing moves. Its own help
text meanwhile promises the opposite, naming this exact case:

Resident layers are transferred once at startup rather than streamed, so they
cut the transfer of every pass including the first — an auxiliary component
that runs once per request still benefits
, it just recovers the VRAM once per
request instead of once per denoising step.

What this PR does

Only names the two operations apart:

release_after_use(keep_resident=False)   the use ended
release_all()                            a full reset

finish_use now says release_after_use(). With the default argument that is
byte-identical to what it did before, and test_release_after_use_defaults_to_the_old_release_all
pins it. enable_offload keeps release_all(), and
test_release_all_still_drops_everything pins that it did not inherit the
resident-set exemption.

No behaviour change. keep_resident=True has no caller in this PR; giving
that decision a home is #2 of the stack, and it needs a memory picture this layer
does not have — a naive local threshold takes the same pipeline from 14.136 s to
13.737 s on a 31.4 GiB card and stops a 23.5 GiB one from starting at all.

Naming

release_after_use rather than end_use: ComponentManager.end_use(use, module=…)
already exists with an unrelated signature, and manager.end_use(...) would read
ambiguously across the two. release_after_use also sits in this file's existing
vocabulary — release_layer, release_all, _release_unneeded_streamed_layers.

Tests

Four, in test_layerwise_offload.py: the default equals the old call; release_all
still drops everything; keep_resident=True keeps exactly _retained_set and
skips the first-pass fault-in those layers do not need; keep_resident=True with
no residents configured is still a full release.

Whole-file run: 81 failed / 51 passed against clean main's 81 failed / 47 passed —
the same pre-existing failures on a machine without CUDA, plus these four.

🤖 Generated with Claude Code


CI States

Latest PR Test (Base): ✅ Run #35611211982
Latest PR Test (Extra): ✅ Run #35611211877
Latest PR Test (AMD ROCm 10): ❌ Run #35611211509

…yerwise offload

`LayerwiseOffloadStrategy.finish_use` called `manager.release_all()`, and
`release_all` did exactly what its name says: dropped every layer, the resident
set included. Those are two different questions. When a component's use ends the
streamed window is certainly dead, but whether the resident set is depends on
whether anything will want it before another stage needs the room — and
`release_all` is also the right call for the full reset in `enable_offload`,
which must leave nothing behind.

Conflating them is why `--layerwise-resident-layers` does nothing for any
component whose use is a single forward pass rather than a denoise loop: the set
is prefetched at the start of the use and dropped at the end of it, every
request. For the DiT one use spans all denoise steps, so the set survives and
earns its memory back; for a text encoder it never survives anything.

This change only names the two operations apart:

  release_after_use(keep_resident=False)  the use ended
  release_all()                           a full reset

`finish_use` now says `release_after_use()`, and the default argument makes that
byte-identical to what it did before — a test pins that. `enable_offload` keeps
`release_all()`. `keep_resident=True` has no caller yet; giving the decision a
home is a separate change.

`release_after_use` rather than `end_use` because `ComponentManager.end_use`
already exists with an unrelated signature, and `manager.end_use(...)` would
read ambiguously across the two.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@mickqian

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci extra

@mickqian

Copy link
Copy Markdown
Collaborator Author

Stack: this PR → #40592 (declare the retain decision on the use) → #40593 (make the flags describe what actually happens, independent). Only #40592 carries this one's diff; #40593 is based on main.

@github-actions github-actions Bot added run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) labels Sep 21, 2026
@mickqian mickqian changed the title [diffusion] Separate "this use ended" from "release everything" in layerwise offload (1/3) [Diffusion] Separate a use-scoped layerwise release from release_all Sep 21, 2026
@mickqian

Copy link
Copy Markdown
Collaborator Author

CI status note (from the author's triage, so reviewers need not re-derive it).

Every red on this PR falls in one of three buckets, none of them from the diff:

  1. NVIDIA multimodal-gen-test-2-gpu (0) — Load Latency (excluding warmup) perf validations only. All 15 validation failures in the latest run are that one metric; no correctness or e2e/denoise metric fails. Measured/expected load on runner h100-novita9-gpu-23: wan2_2_t2v_a14b_2gpu 1.48× on the first attempt and 1.26× on the in-job retry, wan2_1_i2v_14b_480P_2gpu 1.58× → 1.09×, wan2_2_i2v_a14b 1.24×, while the small checkpoints pass (flux2_modelopt_fp8_tp2 0.97×, qwen_image_t2i_2_gpus 1.12×). Second read of the same weights on the same machine is faster and the biggest checkpoints are worst: that is a cold page cache on the runner, not code. The control: [Diffusion] Correct the resident-layer help text to match its scope #40593, whose diff is help text and a doc, fails the same job with the same metric on the same runner pool.
  2. AMD extra-a-*, multimodal-gen-test-*-amd — extra-a-test-1-gpu-small-amd fails identically on [Diffusion] Separate a use-scoped layerwise release from release_all #40590, [diffusion] feat: allow a component use retain its layerwise resident set #40592, [Diffusion] Correct the resident-layer help text to match its scope #40593 and [Diffusion] Add a permanent lifetime for layerwise resident layers #40599 (job logs are not uploaded by the runner); multimodal-gen-test-1-gpu-amd (0) is the whole unit-test suite timing out at the 30-minute budget with test_layerwise_offload.py fully PASSED before the cut. The sibling extra-a-test-1-gpu-large-amd passes on every PR.
  3. NPU multimodal-gen-test-4-npu-a3 (0) — wan2_2_t2v_14b_w8a8_2npu: [consistency] GT image not found (404 on the pinned ci-data revision), with e2e 191197 ms vs baseline 191226 ms.

diffusion-coverage-check and the *-finish jobs are aggregators of the above (coverage saw 9 of 28 2-GPU cases because the cancelled partitions uploaded no report).
For this PR specifically, one earlier attempt of multimodal-gen-test-2-gpu (1) also failed Average Denoise Step on ltx_2_3_two_stage_ti2v_2gpus / ltx_2_5_diffusion_decoder_2gpus. Neither case reaches the changed code: the first runs the DiT resident with cfg_parallel and the second uses component-offload; the layerwise components in both have resident=0, where release_after_use() releases exactly what release_all() did (pinned by test_release_after_use_defaults_to_the_old_release_all). The step timings in that run decay from 606 ms to 348 ms across the loop against a flat ~300 ms baseline — a GPU recovering from contention, not a fixed per-step cost. The same job passes on #40592, which carries this PR's diff unchanged.

@mickqian
mickqian merged commit 5f9c6b9 into sgl-project:main Sep 22, 2026
315 of 343 checks passed
mickqian added a commit to mickqian/sglang that referenced this pull request Sep 22, 2026
Resolves the textual overlap with sgl-project#40593 (help strings, CLI reference,
readiness log): the merged text keeps sgl-project#40593's user vocabulary and adds
the lifetime axis on top, and the readiness log names the lifetime
(`(forward)` / `(permanent)`) where sgl-project#40593 wrote `(per request)`.
Brings in sgl-project#40590 and sgl-project#40592, which this PR composes with unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant