Repository navigation
Conversation
mickqian
requested review from
AgainstEntropy,
BBuf,
HaiShaw,
ping1jing2 and
yichiche
as code owners
August 18, 2026 14:11
Collaborator
Author
|
/tag-and-rerun-ci |
Collaborator
Author
|
Ran this on a 12 GiB RTX 3060 (on top of #35538, which frees ~4.8 GiB there). It never promoted anything, for three separate reasons, now fixed in 8053c90:
62 unit tests pass. Separate issue, not addressed here: on that 3060, |
Collaborator
Author
|
/tag-and-rerun-ci |
4 of 5 tasks
mickqian
requested review from
Fridge003,
Kangyan-Zhou,
bingxche,
ispobock and
merrymercy
as code owners
August 20, 2026 13:44
Under --performance-mode auto with server warmup, measure per-rank peak GPU memory during the synthetic warmup, extrapolate it to the model's default workload (scaling only the activation part above the pre-forward allocated baseline), and promote implicitly offloaded components (component offload -> resident, layerwise offload -> fully loaded) when the estimate plus promoted weights fits under max(10% VRAM, 4 GiB) reserve. The decision is computed from all-gathered rank reports so every rank promotes identically; any rank failure rolls all ranks back, and the residency freezes before /health turns ready. Orchestration: warmup (measure) -> AutoResidencyReq(apply) fanned out through the existing control-req path -> re-run the synthetic warmup under the final residency (physically moves promoted components with per-use dtype semantics, rebuilds compile caches, and proves the layout fits) -> ready. A failed re-warm rolls back and re-warms the original strategy; only a failed rollback aborts startup. Excluded paths: explicit placement, FSDP, diffusers backend, BCG, cache-dit, dynamic batching, dp>1, realtime/disagg, and quantized checkpoints (residency shifts measurably moved fp8 DiT outputs before). Kill switch: SGLANG_DIFFUSION_DISABLE_AUTO_RESIDENCY. Also fixes ridden-along bugs: - server warmup frame caps now re-apply the model frame contract (LongLive2's capped 17 frames -> 5 latent frames broke its 8-frame causal block, so every server warmup failed silently under fail-open; now re-aligned to 29) - get_can_stay_resident_components sized components from the loader's GPU-load delta, which is ~0 for exactly the offloaded components it reports on; it now sizes live modules through layerwise CPU buffers - enable_offload() re-registered forward hooks on managers that were never disabled, double-firing every layer hook Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ation Single-point extrapolation cannot separate constant peak costs (streamed layer weights, attention workspace, tiled VAE decode) from workload-linear activations when offload leaves the pre-forward baseline nearly empty: measured on Wan2.1-14B/H100, a ~30 GiB default-workload peak was estimated as 182.9 GiB and promotion never fired. Frame-capped video server warmup now adds one smaller calibration size (9 frames, auto mode only) so the estimator can fit peak = constant + slope * units and extrapolate only the measured linear part; a measurement at or above the target bounds the peak directly, and a single usable size keeps the conservative baseline split. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…args After warmup the log now states that residency adjustment is starting, what changed (mode -> resident per component), the equivalent --component-residency flags users can pin to freeze the placement, and the SGLANG_DIFFUSION_DISABLE_AUTO_RESIDENCY kill switch. Rollbacks log that the startup-configured residency is in effect again. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
get_can_stay_resident_components now reuses collect_promotion_candidates and the shared H2D-savings ranking, so the post-request hint names the same components in the same order auto mode would promote (still raw capacity: no reserve/margin, and explicitly offloaded components stay listed). Drops the static OFFLOAD_DISABLE_RECOMMENDATION_ORDER and points the hint at --component-residency instead of legacy flags. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Measurement correctness: - empty_cache + reset before each measured warmup forward, so a later shape reusing the previous request's allocator pool cannot flatten the calibration slope to zero (under-estimate -> post-ready OOM) - run the calibration forward AFTER the main warmup request: one-time allocations land in the larger point and can only steepen the slope - only is_warmup reqs are measured; torch-compile prewarm reqs run a different offload layout and poisoned the fit in either direction - record failed warmup forwards even on the propagate re-raise path - skip promotion entirely under --enable-torch-compile: compile warmup runs a stripped memory layout (layerwise DiT, aux components evicted) whose peaks are far below a real request's - unknown default workload now skips instead of trusting the capped warmup peak with no margin; supported_resolutions picks the largest - warmup and target frame counts share one contract helper that also applies the real request's num_gpus latent alignment (extracted from SamplingParams._adjust_visual_fields) Distributed and rollback robustness: - fence everything before the first all-gather into a skip report (a raise there parked peer ranks in the collective until timeout); worker-side dp>1 guard; pipeline-None guard - error gathers filter on `is not None` and describe_error() keeps empty-string exceptions (str(AssertionError()) == "") visible - rollback failures inside apply raise AutoResidencyRollbackError and reach ROLLBACK_FAILED (abort) instead of masquerading as rolled_back; rollback_promotions undoes every component and aggregates errors - disable_offload re-arms hooks when load_all_layers fails, so a failed promotion cannot leave a hook-less enabled manager serving (1,) placeholders - mechanism/mode mismatches (offload_during_compile window) are excluded from candidates and guarded at apply/rollback - _required_resident_components records an owner feature; auto-residency rollback can no longer release a loader's hard requirement Startup contract: - apply-RPC failures honor the warmup fail-open contract instead of SIGTERMing startup; the post-rollback restore warmup keeps the fail-closed contract of explicit --warmup-resolutions - calibration request is gated on auto_residency_skip_reason, so the kill switch / quantized / manual configs no longer pay an extra warmup forward they never consume - warmup records are consumed exactly once and the handler is one-shot; re-warm passes no longer double-advance the warmup progress bar - debug residency hint is fenced (never fails a completed request), resolves the default workload once, and sizes layerwise components from metadata instead of materializing per-weight views Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The cross-process control requests (LoRA ops, shutdown, realtime release, disagg stats, auto residency) are IPC contracts between the HTTP process and the scheduler workers, not utilities; the grab-bag entrypoints/utils.py hid that. Pure move: class bodies are unchanged, the new module stays import-light because both processes load it, and openai/utils.py keeps re-exporting the LoRA protocol names. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Opportunistic migration per the no-dataclasses rule while the classes moved files: same fields, same defaults, positional construction still supported. No dataclasses.asdict/fields/replace call sites exist for these types, and msgspec Structs pickle across the existing broadcast_pyobj/ZMQ transport (AutoResidencyReq already proved the path). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Running this on a 12 GiB RTX 3060 it never promoted anything, for three separate reasons. A warmup probe that did not fit aborted the whole warmup. The synthetic pass sends two probes; when 832x480x17f ran out of memory the loop raised, later probes never ran, and the failed record disabled the estimate. Retry the probe about half as large instead -- frames first, since they drive video activation size and cutting them leaves the spatial kernels at their serving shape -- and give up after three attempts so a failure that is not about probe size does not walk the workload down to nothing. Only memory errors are retried; anything else fails the same way at every size and should surface. A failed probe then blocked the estimate outright. That is right when the probe is at or below the target -- the card cannot hold the default workload as it is already configured, and making weights resident only adds to it -- but a probe that failed above the target says nothing about the target. Keep the block for the first case, with a skip reason that names the size that failed instead of "no usable warmup measurement", and drop the record in the second. The 4 GiB reserve floor only binds below ~40 GiB. Above that the 10% fraction dominates (10% of 80 GiB is 8 GiB), so the floor only ever applies to the cards it was not sized for: on a 12 GiB device a flat 4 GiB fences off a third of the card. Cap the floor at a fifth of the budget, which leaves datacenter behaviour unchanged. Finally, the gate itself was invisible: `--performance-mode auto` with any other warmup mode is skipped, and only at debug level, so the user who asked for auto saw nothing at all. Log it at info when auto was requested.
added 5 commits
September 7, 2026 00:00
# Conflicts: # python/sglang/multimodal_gen/test/server/perf_baselines/5090.json # python/sglang/multimodal_gen/test/server/perf_baselines/h100.json
added 3 commits
September 8, 2026 23:01
…nto l4-auto-residency
…-residency # Conflicts: # python/sglang/multimodal_gen/runtime/pipelines_core/lora/pipeline.py
…text The snapshot-offload tests build the pipeline as a bare namespace; the transition wrapper must not require the attribute.
Collaborator
Author
|
/tag-and-rerun-ci |
This was referenced Sep 16, 2026
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fourth and last layer of the split of this PR's original 12.8k-line diff. The first three layers are now their own PRs, stacked:
placement_budget)Every layer's tree is a prefix of this one; the union is byte-identical to the original branch merged with main.
What stays here
originaldevice mode's offloaded DiT; drop the H200-name special case for the resident auto gate.test_server_argstests that assert those deployment hints.CI States
Latest PR Test (Base): 🚫 Run #34829072659
Latest PR Test (Extra): ❌ Run #34829072485
Latest PR Test (AMD ROCm 10): ❌ Run #34829072688