Repository navigation
Pin the LoRA rollout adapter under SGLang DP attention - #3747
Closed
yushengsu-thu wants to merge 2 commits into
Closed
yushengsu-thu wants to merge 2 commits into
yushengsu-thu wants to merge 2 commits into
Conversation
SGLang's DP-attention LoRA (sgl-project/sglang#36389) routes MLP/MoE tokens by adapter slots all-gathered across DP ranks, so it serves only pinned adapters: every rank must keep an adapter in the same GPU slot. A pinned adapter also needs a slot besides the one the anti-starvation check keeps for the base model. - Register the trainer-pushed adapter with pinned=True under DP attention. - Pin the adapter the engine loads from disk (rollout-only / skip-sync runs). - Give the single-adapter engine max_loras_per_batch=2 under DP attention. Without DP attention, nothing changes: one slot and an evictable adapter.
yushengsu-thu
requested review from
Shi-Dong,
Zhichenzzz,
fzyzcjy,
maocheng23 and
yueming-yuan
as code owners
September 28, 2026 22:03
There was a problem hiding this comment.
Claude Code Review
This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.
Tip: disable this comment in your organization's Code Review settings.
Callers and tests that build partial args without sglang_enable_dp_attention crashed on the direct attribute read. Default it the way the neighbouring LoRA helpers read their optional flags.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
With sgl-project/sglang#36389 (ported to
sglang-milesin sgl-project/sglang#41595), SGLang serves LoRA under DP attention only from pinned adapters. miles registers its adapter unpinned and gives the engine a single LoRA slot, but pinning needs at least two, so DP-attention LoRA RL would fail at adapter registration.Under
--sglang-enable-dp-attention, this PR:pinned=True,max_loras_per_batchfrom 1 to 2.Without DP attention nothing changes. Merge this before sgl-project/sglang#41595.
Testing: new unit tests for the slot count, the pinned startup path and pinned registration; CI passes.