Skip to content

[diffusion] Auto residency: per-pipeline containment fixes and CI baselines (4/4) - #35335

Open
mickqian wants to merge 132 commits into
feat/residency-apply-validatefrom
mick/diffusion-auto-residency
Open

mickqian wants to merge 132 commits into
feat/residency-apply-validatefrom
mick/diffusion-auto-residency

Conversation

@mickqian

@mickqian mickqian commented Aug 18, 2026 •

Copy link
Copy Markdown
Collaborator

Fourth and last layer of the split of this PR's original 12.8k-line diff. The first three layers are now their own PRs, stacked:

  1. [diffusion] Measure warmup memory and layer usage per phase for residency calibration (1/4) #37916 — measure warmup memory and layer usage per phase (record types, estimators, probe request, calibration gate)
  2. [diffusion] Plan component residency from calibrated warmup records (2/4) #37917 — plan component residency from calibrated records (pure planner library, placement_budget)
  3. [diffusion] Apply, validate and roll back calibrated residency after warmup (3/4) #37918 — apply / validate / roll back calibrated residency after warmup (worker, scheduler, client, layerwise layouts, ServerArgs modes)
  4. this PR — the per-pipeline containment fixes and CI baselines that calibrated placement surfaced

Every layer's tree is a prefix of this one; the union is byte-identical to the original branch merged with main.

What stays here

  • LTX-2: preserve the legacy two-stage placement and the original device mode's offloaded DiT; drop the H200-name special case for the resident auto gate.
  • Hunyuan3D paint, LoRA pipeline, LoRA linear: runtime LoRA weights stay inference-only (no grad, versioned adapters) so promotions do not clone them.
  • SANA-WM: keep the text encoders and VAE resident when ≥ 70 GiB is available (the two-stage path is not numerically invariant under layerwise offload), with the two test_server_args tests that assert those deployment hints.
  • Timestep preparation: log shapes instead of tensors (the tensor log forced a CUDA sync).
  • CI baselines: refreshed VRAM/perf baselines and consistency thresholds (H100, 5090) for the calibrated placements; config tests for LTX-2.5 and Hunyuan3D native textures.

CI States

Latest PR Test (Base): 🚫 Run #34829072659
Latest PR Test (Extra): ❌ Run #34829072485
Latest PR Test (AMD ROCm 10): ❌ Run #34829072688

@mickqian

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Aug 18, 2026
@mickqian

Copy link
Copy Markdown
Collaborator Author

Ran this on a 12 GiB RTX 3060 (on top of #35538, which frees ~4.8 GiB there). It never promoted anything, for three separate reasons, now fixed in 8053c90:

  • A warmup probe that did not fit aborted the whole warmup. The synthetic pass sends two probes; 832x480x17f ran out of memory, the loop raised, the second probe never ran, and the failed record disabled the estimate. Probes are now retried about half as large (frames first, then area), capped at three attempts, and only for memory errors.
  • A failed probe blocked the estimate outright. That is correct when the probe is at or below the target — the card cannot hold the default workload as configured, and promoting weights only adds to it — but a probe that failed above the target says nothing about the target. The block is kept for the first case with a skip reason that names the failing size, and the record is dropped in the second. test_failed_warmup_disables_estimation passed target_units=None, so it was exercising the unknown-target branch rather than the failed-record rule; it is now split into three cases that actually cover it.
  • The 4 GiB reserve floor only binds below ~40 GiB. Above that the 10% fraction dominates, so the floor only ever applies to the cards it was not sized for — on 12 GiB it fences off a third of the device. Capped at a fifth of the budget; datacenter behaviour is unchanged (the existing 20 GiB test still expects 4 GiB).
  • The gate was also invisible: --performance-mode auto with any other warmup mode is skipped at debug level, so nothing is printed at all. Now logged at info when auto was requested. Worth noting in the docs too: only sglang serve reaches this hook, sglang generate never does.

62 unit tests pass.

Separate issue, not addressed here: on that 3060, --performance-mode auto keeps ~9.9 GiB live during warmup and the failing allocation stays 738 MiB no matter how small the probe gets (it still OOMs at 48x16), so degradation cannot rescue it. umt5-xxl is 21.16 GB fp32 and ~10.6 GB in bf16, which matches. --performance-mode memory on the same box is fine. That looks like an auto-mode offload-policy problem rather than anything in this PR.

@mickqian

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

mickqian and others added 12 commits August 26, 2026 13:54
Under --performance-mode auto with server warmup, measure per-rank peak
GPU memory during the synthetic warmup, extrapolate it to the model's
default workload (scaling only the activation part above the pre-forward
allocated baseline), and promote implicitly offloaded components
(component offload -> resident, layerwise offload -> fully loaded) when
the estimate plus promoted weights fits under max(10% VRAM, 4 GiB)
reserve. The decision is computed from all-gathered rank reports so every
rank promotes identically; any rank failure rolls all ranks back, and the
residency freezes before /health turns ready.

Orchestration: warmup (measure) -> AutoResidencyReq(apply) fanned out
through the existing control-req path -> re-run the synthetic warmup
under the final residency (physically moves promoted components with
per-use dtype semantics, rebuilds compile caches, and proves the layout
fits) -> ready. A failed re-warm rolls back and re-warms the original
strategy; only a failed rollback aborts startup.

Excluded paths: explicit placement, FSDP, diffusers backend, BCG,
cache-dit, dynamic batching, dp>1, realtime/disagg, and quantized
checkpoints (residency shifts measurably moved fp8 DiT outputs before).
Kill switch: SGLANG_DIFFUSION_DISABLE_AUTO_RESIDENCY.

Also fixes ridden-along bugs:
- server warmup frame caps now re-apply the model frame contract
  (LongLive2's capped 17 frames -> 5 latent frames broke its 8-frame
  causal block, so every server warmup failed silently under fail-open;
  now re-aligned to 29)
- get_can_stay_resident_components sized components from the loader's
  GPU-load delta, which is ~0 for exactly the offloaded components it
  reports on; it now sizes live modules through layerwise CPU buffers
- enable_offload() re-registered forward hooks on managers that were
  never disabled, double-firing every layer hook

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ation

Single-point extrapolation cannot separate constant peak costs (streamed
layer weights, attention workspace, tiled VAE decode) from workload-linear
activations when offload leaves the pre-forward baseline nearly empty:
measured on Wan2.1-14B/H100, a ~30 GiB default-workload peak was estimated
as 182.9 GiB and promotion never fired. Frame-capped video server warmup
now adds one smaller calibration size (9 frames, auto mode only) so the
estimator can fit peak = constant + slope * units and extrapolate only the
measured linear part; a measurement at or above the target bounds the peak
directly, and a single usable size keeps the conservative baseline split.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…args

After warmup the log now states that residency adjustment is starting,
what changed (mode -> resident per component), the equivalent
--component-residency flags users can pin to freeze the placement, and
the SGLANG_DIFFUSION_DISABLE_AUTO_RESIDENCY kill switch. Rollbacks log
that the startup-configured residency is in effect again.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
get_can_stay_resident_components now reuses collect_promotion_candidates
and the shared H2D-savings ranking, so the post-request hint names the
same components in the same order auto mode would promote (still raw
capacity: no reserve/margin, and explicitly offloaded components stay
listed). Drops the static OFFLOAD_DISABLE_RECOMMENDATION_ORDER and
points the hint at --component-residency instead of legacy flags.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measurement correctness:
- empty_cache + reset before each measured warmup forward, so a later
  shape reusing the previous request's allocator pool cannot flatten the
  calibration slope to zero (under-estimate -> post-ready OOM)
- run the calibration forward AFTER the main warmup request: one-time
  allocations land in the larger point and can only steepen the slope
- only is_warmup reqs are measured; torch-compile prewarm reqs run a
  different offload layout and poisoned the fit in either direction
- record failed warmup forwards even on the propagate re-raise path
- skip promotion entirely under --enable-torch-compile: compile warmup
  runs a stripped memory layout (layerwise DiT, aux components evicted)
  whose peaks are far below a real request's
- unknown default workload now skips instead of trusting the capped
  warmup peak with no margin; supported_resolutions picks the largest
- warmup and target frame counts share one contract helper that also
  applies the real request's num_gpus latent alignment (extracted from
  SamplingParams._adjust_visual_fields)

Distributed and rollback robustness:
- fence everything before the first all-gather into a skip report
  (a raise there parked peer ranks in the collective until timeout);
  worker-side dp>1 guard; pipeline-None guard
- error gathers filter on `is not None` and describe_error() keeps
  empty-string exceptions (str(AssertionError()) == "") visible
- rollback failures inside apply raise AutoResidencyRollbackError and
  reach ROLLBACK_FAILED (abort) instead of masquerading as rolled_back;
  rollback_promotions undoes every component and aggregates errors
- disable_offload re-arms hooks when load_all_layers fails, so a failed
  promotion cannot leave a hook-less enabled manager serving (1,)
  placeholders
- mechanism/mode mismatches (offload_during_compile window) are
  excluded from candidates and guarded at apply/rollback
- _required_resident_components records an owner feature; auto-residency
  rollback can no longer release a loader's hard requirement

Startup contract:
- apply-RPC failures honor the warmup fail-open contract instead of
  SIGTERMing startup; the post-rollback restore warmup keeps the
  fail-closed contract of explicit --warmup-resolutions
- calibration request is gated on auto_residency_skip_reason, so the
  kill switch / quantized / manual configs no longer pay an extra
  warmup forward they never consume
- warmup records are consumed exactly once and the handler is one-shot;
  re-warm passes no longer double-advance the warmup progress bar
- debug residency hint is fenced (never fails a completed request),
  resolves the default workload once, and sizes layerwise components
  from metadata instead of materializing per-weight views

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The cross-process control requests (LoRA ops, shutdown, realtime
release, disagg stats, auto residency) are IPC contracts between the
HTTP process and the scheduler workers, not utilities; the grab-bag
entrypoints/utils.py hid that. Pure move: class bodies are unchanged,
the new module stays import-light because both processes load it, and
openai/utils.py keeps re-exporting the LoRA protocol names.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Opportunistic migration per the no-dataclasses rule while the classes
moved files: same fields, same defaults, positional construction still
supported. No dataclasses.asdict/fields/replace call sites exist for
these types, and msgspec Structs pickle across the existing
broadcast_pyobj/ZMQ transport (AutoResidencyReq already proved the
path).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Running this on a 12 GiB RTX 3060 it never promoted anything, for three
separate reasons.

A warmup probe that did not fit aborted the whole warmup. The synthetic pass
sends two probes; when 832x480x17f ran out of memory the loop raised, later
probes never ran, and the failed record disabled the estimate. Retry the probe
about half as large instead -- frames first, since they drive video activation
size and cutting them leaves the spatial kernels at their serving shape -- and
give up after three attempts so a failure that is not about probe size does not
walk the workload down to nothing. Only memory errors are retried; anything
else fails the same way at every size and should surface.

A failed probe then blocked the estimate outright. That is right when the probe
is at or below the target -- the card cannot hold the default workload as it is
already configured, and making weights resident only adds to it -- but a probe
that failed above the target says nothing about the target. Keep the block for
the first case, with a skip reason that names the size that failed instead of
"no usable warmup measurement", and drop the record in the second.

The 4 GiB reserve floor only binds below ~40 GiB. Above that the 10% fraction
dominates (10% of 80 GiB is 8 GiB), so the floor only ever applies to the cards
it was not sized for: on a 12 GiB device a flat 4 GiB fences off a third of the
card. Cap the floor at a fifth of the budget, which leaves datacenter behaviour
unchanged.

Finally, the gate itself was invisible: `--performance-mode auto` with any
other warmup mode is skipped, and only at debug level, so the user who asked
for auto saw nothing at all. Log it at info when auto was requested.
Mick Qian added 3 commits September 8, 2026 23:01
…-residency

# Conflicts:
#	python/sglang/multimodal_gen/runtime/pipelines_core/lora/pipeline.py
…text

The snapshot-offload tests build the pipeline as a bare namespace; the
transition wrapper must not require the attribute.
@Jiminator
Jiminator deleted the branch feat/residency-apply-validate September 14, 2026 04:44
@Jiminator Jiminator closed this Sep 14, 2026
@Jiminator
Jiminator deleted the mick/diffusion-auto-residency branch September 14, 2026 04:44
@alexnails
alexnails restored the mick/diffusion-auto-residency branch September 14, 2026 05:48
@hnyls2002 hnyls2002 removed the run-ci CI: run the baseline test suite on this PR label Sep 14, 2026
@hnyls2002 hnyls2002 reopened this Sep 14, 2026
@mickqian

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

This branch was successfully deployed

1 active deployment
staging - docs — 6f473764 Deployed Sep 12, 2026 by mintlify[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

diffusion SGLang Diffusion documentation Improvements or additions to documentation lora Multi-modal multi-modal language model quant LLM Quantization run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants