Skip to content

[PD] Fix KV cache corruption on abort by notifying ongoing prefill - #27372

Merged
ShangmingCai merged 10 commits into
mainfrom
feat/pd-abort-notification
Jun 5, 2026
Merged

ShangmingCai merged 10 commits into
mainfrom
feat/pd-abort-notification

Conversation

@ShangmingCai

@ShangmingCai ShangmingCai commented Jun 5, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

A lightweight alternative to #24580. This should address almost all race conditions caused by long ITL aborts.

Note: this PR cannot address the case when an ongoing prefill has already called a transfer sync that would take a very long time (maximum value -> mooncake timeout: 30s), but this case has an extremely low probability of happening (extremely hard to reproduce). We could introduce deferred release of decode-side KVCache to fix this as well, but this could harm the performance of the pre-allocation process of decode (considering this aborted req is a large ITL req that will occupy many many KVCache slots).

Modifications

Accuracy Tests

Speed Tests and Profiling

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #27023869436
Latest PR Test (Extra): ❌ Run #27023869006

…ing slot release

Signed-off-by: Shangming Cai <csmthu@gmail.com>
@ShangmingCai

Copy link
Copy Markdown
Collaborator Author

/rerun-group disaggregation

@github-actions

github-actions Bot commented Jun 5, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-group disaggregation:

🚀 2-gpu-h100 (3 tests): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_basic.py
cd test/ && python3 registered/disaggregation/test_disaggregation_decode_offload.py
cd test/ && python3 registered/disaggregation/test_disaggregation_optimistic_prefill.py

🚀 8-gpu-h20 (4 tests): ❌ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_decode_radix_cache.py
cd test/ && python3 registered/disaggregation/test_disaggregation_different_tp.py
cd test/ && python3 registered/disaggregation/test_disaggregation_dp_attention.py
cd test/ && python3 registered/disaggregation/test_disaggregation_pp.py

🚀 8-gpu-h200 (1 test): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_dsv4.py

🚀 8-gpu-h200 (1 test): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_hybrid_attention.py

🚀 1-gpu-5090 (2 tests): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_xpu.py
cd test/ && python3 registered/disaggregation/test_specv2_kvcache_offloading.py

🚀 4-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_epd_disaggregation.py

⛔ registered/disaggregation/test_disaggregation_aarch64.py: Unknown runner_config 4-gpu-gb200 in test/registered/disaggregation/test_disaggregation_aarch64.py — not in scripts/ci/runner_configs.yml.

Known runner_configs: 1-gpu-large, 1-gpu-small, 2-gpu-large, 4-gpu-b200, 4-gpu-h100, 8-gpu-b200, 8-gpu-h20, 8-gpu-h200, deepep-4-gpu-b200, deepep-4-gpu-h100, deepep-8-gpu-h200

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a deferred KV cache release mechanism for decode-initiated aborts to prevent RDMA corruption from in-flight prefill writes. It implements an abort notification flow using ABORT and ABORT_ACK messages, along with a configurable grace period. The review feedback highlights a critical bug where an incorrect attribute name (decode_initiated_abort instead of abort_initiated) disables the deferred release mechanism. Additionally, the reviewer noted potential background thread crashes due to unvalidated message parsing and a potential memory leak from late-arriving ABORT_ACK messages.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread python/sglang/srt/disaggregation/decode.py Outdated
Comment thread python/sglang/srt/disaggregation/mooncake/conn.py
Comment thread python/sglang/srt/disaggregation/mooncake/conn.py
Comment thread python/sglang/srt/disaggregation/decode.py Outdated
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Signed-off-by: Shangming Cai <csmthu@gmail.com>
Signed-off-by: Shangming Cai <csmthu@gmail.com>
@ShangmingCai

Copy link
Copy Markdown
Collaborator Author

/rerun-group disaggregation

@github-actions

github-actions Bot commented Jun 5, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-group disaggregation:

🚀 2-gpu-h100 (3 tests): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_basic.py
cd test/ && python3 registered/disaggregation/test_disaggregation_decode_offload.py
cd test/ && python3 registered/disaggregation/test_disaggregation_optimistic_prefill.py

🚀 8-gpu-h20 (4 tests): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_decode_radix_cache.py
cd test/ && python3 registered/disaggregation/test_disaggregation_different_tp.py
cd test/ && python3 registered/disaggregation/test_disaggregation_dp_attention.py
cd test/ && python3 registered/disaggregation/test_disaggregation_pp.py

🚀 8-gpu-h200 (1 test): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_dsv4.py

🚀 8-gpu-h200 (1 test): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_hybrid_attention.py

🚀 1-gpu-5090 (2 tests): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_xpu.py
cd test/ && python3 registered/disaggregation/test_specv2_kvcache_offloading.py

🚀 4-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_epd_disaggregation.py

⛔ registered/disaggregation/test_disaggregation_aarch64.py: Unknown runner_config 4-gpu-gb200 in test/registered/disaggregation/test_disaggregation_aarch64.py — not in scripts/ci/runner_configs.yml.

Known runner_configs: 1-gpu-large, 1-gpu-small, 2-gpu-large, 4-gpu-b200, 4-gpu-h100, 8-gpu-b200, 8-gpu-h20, 8-gpu-h200, deepep-4-gpu-b200, deepep-4-gpu-h100, deepep-8-gpu-h200

Signed-off-by: Shangming Cai <csmthu@gmail.com>
@ShangmingCai ShangmingCai changed the title [PD] Fix KV cache corruption on abort by notifying prefill and deferring slot release [PD] Fix KV cache corruption on abort by notifying ongoing prefill Jun 5, 2026
Chronostasys pushed a commit to MindLab-Research/sglang that referenced this pull request Aug 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant