Skip to content

[DP Attn] Fix crash for no token all-gather case - #39899

Merged
ByronHsu merged 3 commits into
sgl-project:mainfrom
ByronHsu:fix-idle-prefill-logits-upstream
Sep 17, 2026
Merged

ByronHsu merged 3 commits into
sgl-project:mainfrom
ByronHsu:fix-idle-prefill-logits-upstream

Conversation

@ByronHsu

@ByronHsu ByronHsu commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Why?

With DP attention, DeepEP, DP LM head, --moe-dense-tp-size 1, and breakable prefill graphs, one request can crash an idle DP rank:

One request routed to rank 0
  |
  +-> Rank 0: processes the prompt
  |
  +-> Rank 1: IDLE, 0 local tokens
        |
        v
      Prefill graph padding creates a dummy EXTEND request of length 0
        |
        v
      Empty batch falls back to eager execution
        |
        v
      Logits selection: cumsum([0]) - 1 = [-1]
      hidden_states[-1] -> IndexError on the empty tensor
        |
        v
      Worker fails; client connection closes without a response

What?

  • Preserve the original IDLE mode in logits metadata so idle ranks skip last-token selection. Attention and MLP execution keep the padded batch mode.
  • Add one CPU regression test to the existing test_mlp_sync_pad_unpad.py.

Verification?

Reproduced on upstream main (2733afe54e) with two H200s and Qwen/Qwen3-30B-A3B-FP8. Both runs used the same source and environment; the after run applies only the logits fix. Runtime: PyTorch 2.13.0+cu130, Transformers 5.12.1, and sglang-kernel==0.4.7.

Server command, from the source checkout with MODEL_PATH pointing to the local model:

PYTHONPATH=python CUDA_VISIBLE_DEVICES=0,1 OMP_NUM_THREADS=8 \
SGLANG_DEEPEP_NUM_MAX_DISPATCH_TOKENS_PER_RANK=256 \
python -m sglang.launch_server \
  --model-path "$MODEL_PATH" \
  --host 127.0.0.1 --port 31899 \
  --tp-size 2 --dp-size 2 --ep-size 2 \
  --enable-dp-attention --enable-dp-lm-head \
  --moe-dense-tp-size 1 \
  --moe-a2a-backend deepep --deepep-mode auto \
  --moe-runner-backend deep_gemm \
  --attention-backend fa3 \
  --mem-fraction-static 0.5 --context-length 8192 \
  --max-running-requests 32 --chunked-prefill-size 4096 \
  --cuda-graph-max-bs-decode 8 \
  --cuda-graph-backend-prefill breakable \
  --cuda-graph-max-bs-prefill 128 \
  --skip-server-warmup --watchdog-timeout 120 --random-seed 636278555

After the server is ready, send exactly one request without a generation health check first:

curl http://127.0.0.1:31899/generate \
  -H 'Content-Type: application/json' \
  -d '{
    "text": "The capital city of France is",
    "routed_dp_rank": 0,
    "sampling_params": {
      "temperature": 0,
      "max_new_tokens": 8,
      "ignore_eos": true
    },
    "stream": false
  }'

Observed on upstream main:

Before: rank 1 raises IndexError at hidden_states[last_index];
        server exits and the client connection closes without a response.
After:  HTTP 200; 8 generated tokens; 0.137 s client request time.

The patched response starts with Paris. Which of the following is the and reports completion_tokens: 8, dp_rank: 0. No warmup or generation health request was sent before the test request.

Test?

  • Regression command: PYTHONPATH=python python test/registered/unit/model_executor/test_mlp_sync_pad_unpad.py TestMlpSyncPadUnpad.test_idle_rank_does_not_index_dummy_last_token

    Observed against upstream main (2733afe54e) and this patch:

    Before: FAIL - IndexError on the empty hidden-state tensor
    After:  PASS
    
  • Full test_mlp_sync_pad_unpad.py: all 7 tests pass.

  • Ruff, isort, and git diff --check: pass.

  • CI on a0d36515d1: all 10 CPU partitions pass, including the regression above.

  • Focused CUDA CI: 47 DP-attention tests and 4 DP-attention graph tests pass, 15 breakable-graph tests pass, and original-logprob tests pass.

  • Remaining broad CI jobs: pending.


CI States

Latest PR Test (Base): ⏳ Run #35245918478
Latest PR Test (Extra): ❌ Run #35245913894
Latest PR Test (AMD ROCm 10): ⏳ Run #35245913706

@ByronHsu ByronHsu changed the title Fix logits selection on idle DP prefill ranks [DP Attn] Fix crash for no token all-gather case Sep 17, 2026
@ByronHsu

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 17, 2026
@ByronHsu

Copy link
Copy Markdown
Collaborator Author

/rerun-test test/registered/dp_attn/test_dp_attention.py test/registered/dp_attn/test_dp_attention_bcg_kl.py test/registered/cuda_graph/breakable/test_breakable_cuda_graph.py test/registered/sampling/test_original_logprobs.py

@github-actions

github-actions Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/dp_attn/test_dp_attention.py test/registered/dp_attn/test_dp_attention_bcg_kl.py test/registered/cuda_graph/breakable/test_breakable_cuda_graph.py test/registered/sampling/test_original_logprobs.py:

🚀 2-gpu-h100 (2 tests): ✅ View workflow run

cd test/ && python3 registered/dp_attn/test_dp_attention.py
cd test/ && python3 registered/dp_attn/test_dp_attention_bcg_kl.py

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/cuda_graph/breakable/test_breakable_cuda_graph.py

🚀 1-gpu-5090 (1 test): ✅ View workflow run

cd test/ && python3 registered/sampling/test_original_logprobs.py

@ByronHsu
ByronHsu merged commit a98d921 into sgl-project:main Sep 17, 2026
68 of 95 checks passed
Qiaolin-Yu pushed a commit that referenced this pull request Sep 17, 2026
…-gather case (#39899) (#40012)

Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
fungaren pushed a commit to fungaren/sglang that referenced this pull request Sep 20, 2026
…-gather case (sgl-project#39899) (sgl-project#40012)

Co-authored-by: Byron Hsu <byronhsu1230@gmail.com>
Co-authored-by: Byron Hsu <byron+per@periodiclabs.ai>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant