Repository navigation
[DP Attn] Fix crash for no token all-gather case - #39899
Merged
ByronHsu merged 3 commits intoSep 17, 2026
Merged
Conversation
ByronHsu
requested review from
BBuf,
Edwardf0t1,
Fridge003,
HaiShaw,
Ying1123,
ch-wan,
ispobock and
merrymercy
as code owners
September 17, 2026 04:46
Collaborator
Author
|
/tag-and-rerun-ci |
Collaborator
Author
|
/rerun-test test/registered/dp_attn/test_dp_attention.py test/registered/dp_attn/test_dp_attention_bcg_kl.py test/registered/cuda_graph/breakable/test_breakable_cuda_graph.py test/registered/sampling/test_original_logprobs.py |
Contributor
|
Results for 🚀 🚀 🚀 |
fungaren
pushed a commit
to fungaren/sglang
that referenced
this pull request
Sep 20, 2026
…-gather case (sgl-project#39899) (sgl-project#40012) Co-authored-by: Byron Hsu <byronhsu1230@gmail.com> Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why?
With DP attention, DeepEP, DP LM head,
--moe-dense-tp-size 1, and breakable prefill graphs, one request can crash an idle DP rank:What?
IDLEmode in logits metadata so idle ranks skip last-token selection. Attention and MLP execution keep the padded batch mode.test_mlp_sync_pad_unpad.py.Verification?
Reproduced on upstream
main(2733afe54e) with two H200s andQwen/Qwen3-30B-A3B-FP8. Both runs used the same source and environment; the after run applies only the logits fix. Runtime: PyTorch2.13.0+cu130, Transformers5.12.1, andsglang-kernel==0.4.7.Server command, from the source checkout with
MODEL_PATHpointing to the local model:PYTHONPATH=python CUDA_VISIBLE_DEVICES=0,1 OMP_NUM_THREADS=8 \ SGLANG_DEEPEP_NUM_MAX_DISPATCH_TOKENS_PER_RANK=256 \ python -m sglang.launch_server \ --model-path "$MODEL_PATH" \ --host 127.0.0.1 --port 31899 \ --tp-size 2 --dp-size 2 --ep-size 2 \ --enable-dp-attention --enable-dp-lm-head \ --moe-dense-tp-size 1 \ --moe-a2a-backend deepep --deepep-mode auto \ --moe-runner-backend deep_gemm \ --attention-backend fa3 \ --mem-fraction-static 0.5 --context-length 8192 \ --max-running-requests 32 --chunked-prefill-size 4096 \ --cuda-graph-max-bs-decode 8 \ --cuda-graph-backend-prefill breakable \ --cuda-graph-max-bs-prefill 128 \ --skip-server-warmup --watchdog-timeout 120 --random-seed 636278555After the server is ready, send exactly one request without a generation health check first:
Observed on upstream
main:The patched response starts with
Paris. Which of the following is theand reportscompletion_tokens: 8,dp_rank: 0. No warmup or generation health request was sent before the test request.Test?
Regression command:
PYTHONPATH=python python test/registered/unit/model_executor/test_mlp_sync_pad_unpad.py TestMlpSyncPadUnpad.test_idle_rank_does_not_index_dummy_last_tokenObserved against upstream main (
2733afe54e) and this patch:Full
test_mlp_sync_pad_unpad.py: all 7 tests pass.Ruff, isort, and
git diff --check: pass.CI on
a0d36515d1: all 10 CPU partitions pass, including the regression above.Focused CUDA CI: 47 DP-attention tests and 4 DP-attention graph tests pass, 15 breakable-graph tests pass, and original-logprob tests pass.
Remaining broad CI jobs: pending.
CI States
Latest PR Test (Base): ⏳ Run #35245918478
Latest PR Test (Extra): ❌ Run #35245913894
Latest PR Test (AMD ROCm 10): ⏳ Run #35245913706