Skip to content

[Diffusion] Correct the resident-layer help text to match its scope - #40593

Merged
mickqian merged 3 commits into
sgl-project:mainfrom
mickqian:mick/layerwise-resident-help-truth
Sep 22, 2026
Merged

mickqian merged 3 commits into
sgl-project:mainfrom
mickqian:mick/layerwise-resident-help-truth

Conversation

@mickqian

@mickqian mickqian commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator

Third of three, independent of the other two — this one needs no behaviour change
at all, because the behaviour is fine and the description is not.

Stack

  1. [Diffusion] Separate a use-scoped layerwise release from release_all #40590 — separate release_after_use from release_all
  2. [diffusion] feat: allow a component use retain its layerwise resident set #40592 — declare the retain decision on the use
  3. this PR — make the flags describe what actually happens (independent, based on main)

The claims

--dit-layerwise-resident-layers:

keep this many DiT layers permanently resident on GPU … resident layers are
transferred once (not re-streamed every step)

--layerwise-resident-layers:

Resident layers are transferred once at startup rather than streamed, so they
cut the transfer of every pass including the first — an auxiliary component
that runs once per request still benefits
, it just recovers the VRAM once per
request instead of once per denoising step.

What happens

The resident set is released when a component's use ends —
LayerwiseOffloadStrategy.finish_use → release_all(), whose own docstring says
it "ends the denoise stage that the resident set is scoped to".

For a DiT that costs nothing: one use spans every denoise step, so "transferred
once per use" and "transferred once per request" coincide, which is the case both
flags were written against. For a component whose use is a single forward pass
they do not coincide at all — the set is prefetched at the start of the use and
dropped at the end, and the next request transfers all of it again.

Measured on Qwen-Image-2.1, one RTX 5090, five requests per arm:

p50 s/step steady VRAM logged
baseline 14.136 s 0.3173 17047 MiB resident=0/66
--layerwise-resident-layers text_encoder=0.8 14.168 s 0.3173 17047 MiB resident=53/66

Identical memory, identical latency, identical per-step time. The setting is
accepted and reported and has no effect — and the help text for it names exactly
this case as one that benefits.

What this PR changes

Both help strings, the CLI reference and the readiness log, in the user's own
terms. The layers are held while the component does its work for a request and
released when it finishes, so they are transferred once per request rather than
once per step or pass -- and not kept for the life of the server. The readiness
log says resident=53/66 (per request) where the bare number read as a standing
state.

"Per request" rather than the runtime's "per use" on purpose: a use is a
ComponentUse, the runtime's unit of a component's work, and not a word a user
setting a CLI flag has any reason to know. Verified against the runtime before
wording it that way: the set is pinned by the layer-0 pre-hook on the component's
first forward, released by finish_use, and finish_request forces
preferred = False for any module holding residents, so nothing keeps it past
the request. In every ordinary pipeline the two scopes coincide.

docs/docs/sglang-diffusion/api/cli.mdx carried the same claim in its own words
-- "Resident layers are transferred once at startup, so they are removed from
every pass" -- and its per-component example set the flag on text_encoder,
which is the one component where it demonstrably does nothing. Both corrected;
the example now uses transformer=4. (Not dit=4, which I tried first: the
per-component map does no aliasing -- _parse_component_value_map is literal
and _pick matches the component's real name -- so dit=4 is silently
ignored, a worse example than the one it replaced.) The help's own example is
now video_vae=36, the one measured per-component recipe in the repo.

One correction to my own first wording, from reading the cookbook: the help said
the flag buys nothing for "an auxiliary component", and that is too broad in the
direction that costs a user performance. deployment_cookbook.mdx measures
--layerwise-resident-layers video_vae=36 at 13 s of decode against 150 s
streamed, because a video VAE's single use makes many passes over its layers --
the same shape as a DiT across denoise steps, not the text encoder's single
forward pass. The condition is how many times the component runs its layers per request, not
whether it is auxiliary, and the help now says so.

Nothing else. The diagnostic actually worth having — telling an operator at
startup that their setting cannot pay — has to know how many times the component
runs per request. estimate_layerwise_layer_uses in auto_residency.py computes
that, from warmup records this code has no access to, so the warning belongs with
the residency planner. I have not approximated it here.

Tests

test_resident_layer_help_describes_the_actual_scope asserts the retired phrases
are gone and the scope is stated, because a help string is exactly the kind of
claim that drifts back when someone edits nearby.

test_server_args.py: 4 failed / 195 passed against clean main's 4 failed / 194
passed — the same pre-existing failures, plus this one.

CI then caught what that run did not cover: test_layerwise_offload.py's
test_configure_logs_component_start_and_completion pins the readiness log line
verbatim, so annotating it broke that assertion. Fixed in 0e18e49, with a
comment on the assertion saying why the words are there.

test_resident_layer_help_describes_the_actual_scope now anchors on the
user-facing claims -- once per request, life of the server, has no effect
-- rather than on runtime vocabulary (6db268e).

🤖 Generated with Claude Code


CI States

Latest PR Test (Base): ✅ Run #35619823972
Latest PR Test (Extra): ✅ Run #35619823478
Latest PR Test (AMD ROCm 10): ❌ Run #35619824063

Both resident-layer flags describe behaviour the runtime does not have.

--dit-layerwise-resident-layers calls the layers "permanently resident on GPU"
and says they are "transferred once". --layerwise-resident-layers goes further:
"transferred once at startup rather than streamed", cutting "the transfer of
every pass including the first", and names the case explicitly -- "an auxiliary
component that runs once per request still benefits".

The resident set is released when a component's use ends. For a DiT that costs
nothing, because one use spans every denoise step. For a component whose use is
a single forward pass, the set is prefetched at the start of the use and dropped
at the end of it, and the next request transfers all of it again. Measured on
Qwen-Image-2.1, `--layerwise-resident-layers text_encoder=0.8` reports
`resident=53/66` and changes neither steady VRAM (17047 MiB either way) nor
latency (14.168s against 14.136s).

Corrects both help texts and adds "(per use)" to the readiness log, where the
bare `resident=53/66` reads as a standing state.

No behaviour change; this is the claim catching up with the code. The warning
worth having -- telling an operator their setting cannot pay -- needs to know
how many times a component runs per request, which `estimate_layerwise_layer_uses`
computes from warmup records in auto_residency.py rather than anywhere this code
can see, so it belongs with the residency planner and is not approximated here.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@mickqian

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci extra

@github-actions github-actions Bot added the diffusion SGLang Diffusion label Sep 21, 2026
@github-actions github-actions Bot added run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) labels Sep 21, 2026
@mickqian mickqian changed the title [diffusion] Say what the layerwise resident set actually does (3/3) [Diffusion] Correct the resident-layer help text to match its scope Sep 21, 2026
Three follow-ups to the same mistake, found by CI.

`test_configure_logs_component_start_and_completion` pins the startup log
line verbatim, and adding "(per use)" to it broke that assertion. Updated,
with a comment saying why the words are there so the next edit does not
quietly drop them.

`docs/.../api/cli.mdx` repeats the claim the flags were making -- "Resident
layers are transferred once at startup, so they are removed from every
pass" -- and its own per-component example sets the flag on `text_encoder`,
which is exactly the component where it does nothing. Corrected both, and
moved the example to `dit`.

The per-component help said the flag buys nothing for "an auxiliary
component". That is too broad in the direction that costs a user
performance: `deployment_cookbook.mdx` measures `video_vae=36` at 13 s of
decode against 150 s streamed, because a video VAE's single use makes many
passes over the layers, the same shape as a DiT across denoise steps. The
condition is the number of passes per use, not whether a component is
auxiliary, so the help now says that.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 21, 2026
"Use" is a `ComponentUse` -- the runtime's unit of a component's work, and
not a word a user setting a CLI flag has any reason to know. Both help
strings, the CLI reference and the readiness log now say what happens in
the terms the user already has: the layers are held while the component
does its work for a request and released when it finishes, so they are
transferred once per request rather than once per step or pass. The flag
pays for a component that runs its layers many times per request (a DiT
across denoise steps, a video VAE across latent chunks) and does nothing
for one that runs them once (a text encoder).

Verified against the runtime before wording it: the set is pinned by the
layer-0 pre-hook and released by `finish_use`; `finish_request` forces
`preferred = False` for any module holding residents, so nothing keeps it
past the request. "Per request" is the user-facing truth of that.

Also corrects the example key I had just introduced. `dit=4` is silently
ignored by `--layerwise-resident-layers`: `_parse_component_value_map`
does no aliasing and the DiT's component name is `transformer`. The doc
example now uses `transformer=4`; the help uses `video_vae=36`, the one
measured per-component recipe in the repo.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@mickqian

Copy link
Copy Markdown
Collaborator Author

CI status note (from the author's triage, so reviewers need not re-derive it).

Every red on this PR falls in one of three buckets, none of them from the diff:

  1. NVIDIA multimodal-gen-test-2-gpu (0) — Load Latency (excluding warmup) perf validations only. All 15 validation failures in the latest run are that one metric; no correctness or e2e/denoise metric fails. Measured/expected load on runner h100-novita9-gpu-23: wan2_2_t2v_a14b_2gpu 1.48× on the first attempt and 1.26× on the in-job retry, wan2_1_i2v_14b_480P_2gpu 1.58× → 1.09×, wan2_2_i2v_a14b 1.24×, while the small checkpoints pass (flux2_modelopt_fp8_tp2 0.97×, qwen_image_t2i_2_gpus 1.12×). Second read of the same weights on the same machine is faster and the biggest checkpoints are worst: that is a cold page cache on the runner, not code. The control: [Diffusion] Correct the resident-layer help text to match its scope #40593, whose diff is help text and a doc, fails the same job with the same metric on the same runner pool.
  2. AMD extra-a-*, multimodal-gen-test-*-amd — extra-a-test-1-gpu-small-amd fails identically on [Diffusion] Separate a use-scoped layerwise release from release_all #40590, [diffusion] feat: allow a component use retain its layerwise resident set #40592, [Diffusion] Correct the resident-layer help text to match its scope #40593 and [Diffusion] Add a permanent lifetime for layerwise resident layers #40599 (job logs are not uploaded by the runner); multimodal-gen-test-1-gpu-amd (0) is the whole unit-test suite timing out at the 30-minute budget with test_layerwise_offload.py fully PASSED before the cut. The sibling extra-a-test-1-gpu-large-amd passes on every PR.
  3. NPU multimodal-gen-test-4-npu-a3 (0) — wan2_2_t2v_14b_w8a8_2npu: [consistency] GT image not found (404 on the pinned ci-data revision), with e2e 191197 ms vs baseline 191226 ms.

diffusion-coverage-check and the *-finish jobs are aggregators of the above (coverage saw 9 of 28 2-GPU cases because the cancelled partitions uploaded no report).

@mickqian
mickqian merged commit 4a1b69a into sgl-project:main Sep 22, 2026
172 of 184 checks passed
mickqian added a commit to mickqian/sglang that referenced this pull request Sep 22, 2026
Resolves the textual overlap with sgl-project#40593 (help strings, CLI reference,
readiness log): the merged text keeps sgl-project#40593's user vocabulary and adds
the lifetime axis on top, and the readiness log names the lifetime
(`(forward)` / `(permanent)`) where sgl-project#40593 wrote `(per request)`.
Brings in sgl-project#40590 and sgl-project#40592, which this PR composes with unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion documentation Improvements or additions to documentation run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant