Repository navigation
[PD] Preserve abort ACKs until in-flight KV transfers drain - #40645
Conversation
|
/tag-and-rerun-ci |
| if ( | ||
| room_active | ||
| or self._staging_outstanding.get(room_to_be_aborted, 0) > 0 | ||
| ): |
There was a problem hiding this comment.
nit: with this new condition, the elif at L2438 is only reached when the room is inactive and outstanding <= 0, so its == 0 re-check can never usefully be false. The one state where it differs is a hypothetically negative counter, where the elif silently drops the abort (no ack, no registration → decode waits out the release timeout), while a plain else would correctly ack — a negative counter still means nothing is in flight.
- elif self._staging_outstanding.get(room_to_be_aborted, 0) == 0:
- # Concluded/unknown AND quiescent: ack now. A cleared
- # room is not automatically quiescent -- clear() can
- # drop a room whose chunk is still transferring.
+ else:
+ # Concluded/unknown AND quiescent (the branch above
+ # already took every case with writes outstanding):
+ # ack now so decode releases without the timeout.Purely cosmetic/robustness — fine to take as a follow-up; the current code is behaviorally identical in every reachable state.
|
/rerun-group disaggregation |
|
Results for 🚀 🚀 🚀 🚀 🚀 🚀 🚀 ⛔ |
Pick up sgl-project#40645, sgl-project#40711, the shared prefill->decode status plumbing (sgl-project#36612) and runtime role switching (sgl-project#28403). Conflict resolution: - common: keep upstream's update_status structure and still reset the deferred-ACK state when a room lifecycle starts. CommonKVSender.clear() follows sgl-project#40645: ACK now when nothing is in flight, else keep the target. - mooncake: the PR's unconditional register-then-try-ACK already covers sgl-project#40645 and sgl-project#40711. - nixl: keep upstream's settle-then-conclude exception path and still poison the ACK target, since that chunk stays counted. - mori: re-apply the drain-aware ACK path on upstream's conclude_transfer / conclude_failure API and list-returning _submit_kv_transfer; register the drain threads with teardown() and stop them with a sentinel. - tests: update stubs for the new NIXL exception path and for the deferred-ACK fields CommonKVManager now always owns. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Motivation
The prefill scheduler can clear a failed sender while its KV transfer is still writing. Clearing the pending abort-ACK destination loses the drain notification, leaving decode to release device KV on timeout. An abort arriving after room cleanup can also miss registration while writes remain outstanding.
Modifications
Preserve the ACK destination until outstanding writes drain, acknowledge immediately when already drained, and register late Mooncake aborts for rooms with outstanding transfers. Add a regression test covering cleanup before and after drain and exactly-once acknowledgment.
Validation
83406b6f45.test/registered/kernel/quantization/test_mxfp8_kv_reserved_slot.pypath; only that hook was skipped for the commit.CI States
Latest PR Test (Base): 🚫 Run #35667160829
Latest PR Test (Extra): ❌ Run #35667160700
Latest PR Test (AMD ROCm 10): ⏳ Run #35667160749