Skip to content

Pin the LoRA rollout adapter under SGLang DP attention - #3747

Closed
yushengsu-thu wants to merge 2 commits into
mainfrom
yusheng/pin-lora-under-dp-attention
Closed

yushengsu-thu wants to merge 2 commits into
mainfrom
yusheng/pin-lora-under-dp-attention

Conversation

@yushengsu-thu

@yushengsu-thu yushengsu-thu commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

With sgl-project/sglang#36389 (ported to sglang-miles in sgl-project/sglang#41595), SGLang serves LoRA under DP attention only from pinned adapters. miles registers its adapter unpinned and gives the engine a single LoRA slot, but pinning needs at least two, so DP-attention LoRA RL would fail at adapter registration.

Under --sglang-enable-dp-attention, this PR:

  • registers the adapter with pinned=True,
  • pins the adapter the engine loads from disk at startup,
  • raises max_loras_per_batch from 1 to 2.

Without DP attention nothing changes. Merge this before sgl-project/sglang#41595.

Testing: new unit tests for the slot count, the pinned startup path and pinned registration; CI passes.

SGLang's DP-attention LoRA (sgl-project/sglang#36389) routes MLP/MoE tokens by
adapter slots all-gathered across DP ranks, so it serves only pinned adapters:
every rank must keep an adapter in the same GPU slot. A pinned adapter also
needs a slot besides the one the anti-starvation check keeps for the base model.

- Register the trainer-pushed adapter with pinned=True under DP attention.
- Pin the adapter the engine loads from disk (rollout-only / skip-sync runs).
- Give the single-adapter engine max_loras_per_batch=2 under DP attention.

Without DP attention, nothing changes: one slot and an evictable adapter.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This repository is configured for manual code reviews. Comment @claude review for a one-time review, or @claude review always to subscribe this PR to a review on every future push.

Tip: disable this comment in your organization's Code Review settings.

Callers and tests that build partial args without sglang_enable_dp_attention
crashed on the direct attribute read. Default it the way the neighbouring LoRA
helpers read their optional flags.
@yushengsu-thu yushengsu-thu added run-ci-lora run-ci-lora-native Run native (raw-mode) LoRA plugin e2e tests run-ci-multi-lora Run Tinker cookbook and multi-tenant gateway GPU e2e tests labels Sep 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci-lora run-ci-lora-native Run native (raw-mode) LoRA plugin e2e tests run-ci-multi-lora Run Tinker cookbook and multi-tenant gateway GPU e2e tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant