Skip to content

[cherrypick from #36389] [sglang][lora] Support DP attention in LoRA backends - #41595

Open
yushengsu-thu wants to merge 70 commits into
sglang-milesfrom
yusheng/cherrypick-36389-sglang-miles
Open

yushengsu-thu wants to merge 70 commits into
sglang-milesfrom
yusheng/cherrypick-36389-sglang-miles

Conversation

@yushengsu-thu

@yushengsu-thu yushengsu-thu commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Port #36389 to the sglang-miles branch.

LoRA under --enable-dp-attention now routes each model section with the token layout it runs on. Attention and other DP-local sections use each rank's own batch info. The MLP/MoE after the DP gather, and a DP-gathered LM head, use a TP-global view whose per-token adapter ids are all-gathered across DP ranks. Every adapter in a multi-adapter batch is applied to the right tokens, including the tokens that other ranks' experts and shared-expert shards process.

On sglang-miles today, DP-attention LoRA whose targets include shared or routed experts produces wrong logprobs (see Testing). The existing path knows only the local tokens; at best it attributes other ranks' tokens to the single loaded adapter.

As in #36389, LoRA with DP attention requires:

  • the Triton LoRA backend,
  • --dp-size equal to --tp-size,
  • pinned adapters, at most --max-loras-per-batch - 1 of them,
  • no --enable-lora-overlap-loading.

LoRA prefill runs eagerly under DP attention; decode CUDA graphs are supported. radixark/miles#3747 pins the miles rollout adapter to match.

Port notes

  • Replaces sglang-miles' own DP-attention MoE-LoRA path from 216f478 ([11/54]): the single-adapter gathered-tail stamping, get_gathered_moe_num_tokens, the MoE graph buffers scaled by DP size inside init_cuda_graph_moe_buffers, and the idle-rank MoE prepare_lora_batch call in ForwardBatch.init_new. [sglang][lora] Support DP attention in LoRA backends #36389 prepares the LoRA batch after the DP sync on every rank; the idle-rank call would run its collectives on idle ranks only. test_moe_lora_tail_stamp.py, which pinned the stamping, is removed.
  • Keeps sglang-miles' quant-agnostic device lookup for the MoE graph buffers.
  • Uses dp_attention.get_attention_dp_rank(), the slot dp_gather_replicate writes this rank's rows to on sglang-miles, instead of main-only dp_slot_in.
  • Publishes the LoRA layout in Qwen35FlashInferLayerCommunicator, whose fused paths return without reaching LayerCommunicator.prepare_attn / prepare_mlp.
  • Applies the incomplete-unload bookkeeping to register_lora_adapter (streamed, upsert, defer_publish) and to staged adapters whose discard fails.
  • Drops main-only context that [sglang][lora] Support DP attention in LoRA backends #36389 touched (defer_moe_finalize, UnreducedOutput).
  • Tests: the new DP-attention unit test stubs get_attention_dp_rank, the cleanup test exercises register_lora_adapter (sglang-miles has no load_lora_adapter_from_tensors), and sglang-miles' own LoRA tests set the new fields on their __new__-built doubles.
  • Contains [cherrypick from #36389] [LoRA] Let DP-attention LoRA run idle ranks and eager prefill (partial) #41507's two hunks unchanged, so the two apply cleanly in either order.

Testing

All runs are on an 8×H200 devbox.

  • Unit tests (test/registered/unit/{lora,managers,model_executor,layers,models,batch_overlap,spec} and the runtime-context tests): 2689 passed / 46 failed, against 2668 / 47 on sglang-miles. There are no new failures; the 21 extra passes are [sglang][lora] Support DP attention in LoRA backends #36389's tests plus one test that fails on sglang-miles.
  • pre-commit run on the changed files passes. Ruff F821/F811/F401/F841 reports only the two findings sglang-miles already has.
  • GPU: DeepSeek-V2-Lite with --tp 8 --dp 8 --enable-dp-attention --ep 8 --moe-dense-tp-size 1 --enable-dp-lm-head, triton MoE and LoRA backends, --lora-use-virtual-experts, and two pinned random adapters on o_proj, the dense MLP, the shared experts and all routed experts. Served LoRA is compared with the same adapter merged into the base weights, in the same layout, over 164 prompt tokens. "Delta corr" is the correlation between (LoRA − base) and (merged − base).
mean abs dlogprob LoRA effect delta corr greedy outputs identical
sglang-miles + #41507 0.55 0.61 0.40 2/16
this PR 0.042 0.63 0.9985 14/16
  • Base, a1 and a2 requests mixed across DP ranks match the same requests sent one at a time within batch noise (mean abs dlogprob 0.019 / 0.028 / 0.035).
  • Without DP attention (--tp 8 --ep 8), prompt logprobs are bit-identical to sglang-miles. Greedy decode differs only within the run-to-run variation of the same build, from split-K atomics in the MoE LoRA kernels.
  • No scheduler exceptions in any run, including idle DP ranks and decode CUDA graphs.

CI States

Latest PR Test (Base): ❌ Run #36490009615
Latest PR Test (Extra): ❌ Run #36490009135
Latest PR Test (AMD ROCm 10): ❌ Run #36490009492

ishandhanani and others added 30 commits September 16, 2026 15:31
Co-authored-by: William Arnold <7565007+Aphoh@users.noreply.github.com>
…s in exact-token preprocessing (#30368) (#40009)

Signed-off-by: Zhuangcheng(Jesse) Gu <zcgu@connect.hku.hk>
Co-authored-by: Zhuangcheng(Jesse) Gu <40918450+Chokoyo@users.noreply.github.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Mick <mickjagger19@icloud.com>
…-gather case (#39899) (#40012)

Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
… dropping swiglu_limit clamped SwiGLU activation (#39920) (#40035)

Co-authored-by: Shijin Zhang <75300765+Dovis01@users.noreply.github.com>
…39915) (#40137)

Signed-off-by: William Arnold <7565007+Aphoh@users.noreply.github.com>
Co-authored-by: William Arnold <7565007+Aphoh@users.noreply.github.com>
Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
…ense) (#18639)

(cherry picked from commit a6cc76d)
(cherry picked from commit 29d5285)
…ort (#18642)

(cherry picked from commit f78421f)
(cherry picked from commit 72e9b93)
(cherry picked from commit cd31ae6)
(cherry picked from commit 77b31fd)
(cherry picked from commit f9747a5)
(cherry picked from commit 6e7483c)
(cherry picked from commit c4b6f74)
(cherry picked from commit 049efd7)
(cherry picked from commit d669fb3)
(cherry picked from commit e18e48c)
(cherry picked from commit ab3c05f)
(cherry picked from commit 17c3be8)
(cherry picked from commit 16b8d2e)
(cherry picked from commit 4661ad8)
(cherry picked from commit 3a4659e)
(cherry picked from commit 547283d)
…h/RL fixes (#25141, #29874, #31251)

(cherry picked from commit 7fea23f)
(cherry picked from commit 00639c4)
…27728)

(cherry picked from commit 3e32648)
(cherry picked from commit b63c007)
…psert (#27268, #31759, #30913)

(cherry picked from commit f8a11ee)
(cherry picked from commit 8528623)
…capture + meta_info threading

(cherry picked from commit 2aee68b)
(cherry picked from commit 9e6e62c)
… for spec draft worker(s) (#27749, #28575, #18565, #22663, #28001, #29675, #27750)

(cherry picked from commit 4290cf4)
(cherry picked from commit 3c0c432)
… from RL weight check (#29339)

(cherry picked from commit 5d36347)
(cherry picked from commit 3e92eeb)
…ht updates (#30421)

(cherry picked from commit b09c3c7)
(cherry picked from commit 02715e0)
(cherry picked from commit 5944cf6)
(cherry picked from commit f2ba83f)
… into a host-local checkpoint (#30366, #28524)

(cherry picked from commit f09356f)
(cherry picked from commit 00a2ca2)
… multi-lora needs (#30912)

(cherry picked from commit 80aa4ec)
(cherry picked from commit f0cbd48)
… in retract (#31962)

(cherry picked from commit e9ed3cb)
(cherry picked from commit 9fa67e5)
…h the merged per-role payload

(cherry picked from commit 8fea606)
(cherry picked from commit a264f30)
…t to v0.5.16's parallel state

(cherry picked from commit 19207ba)
(cherry picked from commit 174ddab)
(cherry picked from commit 5296432)
(cherry picked from commit 4c974df)

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

deepseek documentation Improvements or additions to documentation jit-kernel lora

Projects

None yet

Development

Successfully merging this pull request may close these issues.