Repository navigation
[cherrypick from #36389] [sglang][lora] Support DP attention in LoRA backends - #41595
Open
yushengsu-thu wants to merge 70 commits into
Open
yushengsu-thu wants to merge 70 commits into
yushengsu-thu wants to merge 70 commits into
Conversation
…s in exact-token preprocessing (#30368) (#40009) Signed-off-by: Zhuangcheng(Jesse) Gu <zcgu@connect.hku.hk> Co-authored-by: Zhuangcheng(Jesse) Gu <40918450+Chokoyo@users.noreply.github.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> Co-authored-by: Mick <mickjagger19@icloud.com>
…39915) (#40137) Signed-off-by: William Arnold <7565007+Aphoh@users.noreply.github.com> Co-authored-by: William Arnold <7565007+Aphoh@users.noreply.github.com> Co-authored-by: ishandhanani <82981111+ishandhanani@users.noreply.github.com>
guapisolo
force-pushed
the
sglang-miles
branch
from
October 2, 2026 06:45
14a1fa7 to
6c94e57
Compare
guapisolo
requested review from
1am9trash,
AgainstEntropy,
Alisehen,
AniZpZ,
ByronHsu,
CatherineSue,
Duyi-Wang,
FlamingoPg,
OrangeRedeng,
Qiaolin-Yu,
ShangmingCai,
YAMY1234,
b8zhong,
hebiao064,
hubertlu-tw,
kevin-mii,
key4ng,
kkHuang-amd,
kpham-sgl,
mickqian,
mmangkad,
niehen6174,
ping1jing2,
rainj-me,
slin1237,
yctseng0211,
yhyang201,
yichiche and
yizhang2077
as code owners
October 2, 2026 06:45
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Port #36389 to the
sglang-milesbranch.LoRA under
--enable-dp-attentionnow routes each model section with the token layout it runs on. Attention and other DP-local sections use each rank's own batch info. The MLP/MoE after the DP gather, and a DP-gathered LM head, use a TP-global view whose per-token adapter ids are all-gathered across DP ranks. Every adapter in a multi-adapter batch is applied to the right tokens, including the tokens that other ranks' experts and shared-expert shards process.On
sglang-milestoday, DP-attention LoRA whose targets include shared or routed experts produces wrong logprobs (see Testing). The existing path knows only the local tokens; at best it attributes other ranks' tokens to the single loaded adapter.As in #36389, LoRA with DP attention requires:
--dp-sizeequal to--tp-size,--max-loras-per-batch - 1of them,--enable-lora-overlap-loading.LoRA prefill runs eagerly under DP attention; decode CUDA graphs are supported. radixark/miles#3747 pins the miles rollout adapter to match.
Port notes
get_gathered_moe_num_tokens, the MoE graph buffers scaled by DP size insideinit_cuda_graph_moe_buffers, and the idle-rank MoEprepare_lora_batchcall inForwardBatch.init_new. [sglang][lora] Support DP attention in LoRA backends #36389 prepares the LoRA batch after the DP sync on every rank; the idle-rank call would run its collectives on idle ranks only.test_moe_lora_tail_stamp.py, which pinned the stamping, is removed.dp_attention.get_attention_dp_rank(), the slotdp_gather_replicatewrites this rank's rows to on sglang-miles, instead of main-onlydp_slot_in.Qwen35FlashInferLayerCommunicator, whose fused paths return without reachingLayerCommunicator.prepare_attn/prepare_mlp.register_lora_adapter(streamed, upsert,defer_publish) and to staged adapters whose discard fails.defer_moe_finalize,UnreducedOutput).get_attention_dp_rank, the cleanup test exercisesregister_lora_adapter(sglang-miles has noload_lora_adapter_from_tensors), and sglang-miles' own LoRA tests set the new fields on their__new__-built doubles.Testing
All runs are on an 8×H200 devbox.
test/registered/unit/{lora,managers,model_executor,layers,models,batch_overlap,spec}and the runtime-context tests): 2689 passed / 46 failed, against 2668 / 47 onsglang-miles. There are no new failures; the 21 extra passes are [sglang][lora] Support DP attention in LoRA backends #36389's tests plus one test that fails onsglang-miles.pre-commit runon the changed files passes. Ruff F821/F811/F401/F841 reports only the two findingssglang-milesalready has.--tp 8 --dp 8 --enable-dp-attention --ep 8 --moe-dense-tp-size 1 --enable-dp-lm-head, triton MoE and LoRA backends,--lora-use-virtual-experts, and two pinned random adapters ono_proj, the dense MLP, the shared experts and all routed experts. Served LoRA is compared with the same adapter merged into the base weights, in the same layout, over 164 prompt tokens. "Delta corr" is the correlation between (LoRA − base) and (merged − base).sglang-miles+ #41507--tp 8 --ep 8), prompt logprobs are bit-identical tosglang-miles. Greedy decode differs only within the run-to-run variation of the same build, from split-K atomics in the MoE LoRA kernels.CI States
Latest PR Test (Base): ❌ Run #36490009615
Latest PR Test (Extra): ❌ Run #36490009135
Latest PR Test (AMD ROCm 10): ❌ Run #36490009492