Skip to content

[Kimi-K3] Deferred decode-side KV release (#35049 + #35360) on the zmq-mitigation base - #36610

Draft
hanming-lu wants to merge 3 commits into
kimi-k3-zmq-max-sockets-mitigationfrom
kimi-k3-deferred-release-35049-35360
Draft

hanming-lu wants to merge 3 commits into
kimi-k3-zmq-max-sockets-mitigationfrom
kimi-k3-deferred-release-35049-35360

Conversation

@hanming-lu

@hanming-lu hanming-lu commented Aug 27, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

This is the exact Kimi-K3 serving tree used to A/B-test the deferred
decode-side KV release fix on a GB300 PD deployment. It is
kimi-k3-zmq-max-sockets-mitigation (6dfdebf) with the deferred-release work
cherry-picked on top, so the branch can be built into a serving image and
compared directly against the unpatched build.

Cherry-picks, in order:

upstream commit why
#30545 05c7ebf64c [Disagg][StagingBuffer][2/2] Support radix cache — prerequisite for the two below
#35049 97dedd1ce9 [PD] Deferred decode-side KV release for aborts mid-transfer
#35360 adca19c497 [PD] Deferred decode-side KV release for the NIXL backend

Three conflicts came up, all purely additive (two from typing import ...
lines, plus new blocks in environ.py and mooncake/conn.py); each was
resolved by taking the incoming side.

The resulting tree hash is a7845862cd43eecbd923fff557df1f180fcd8502, which
matches the tree that was actually exercised on hardware — that equality is the
point of this branch, so a reviewer can confirm the image and the measurements
came from the same bytes.

What the measurements showed

Reproduced on a 2-prefill/1-decode Kimi-K3 deployment (mooncake, dp16 prefill ×2

  • dp32 decode) with a deterministic barrier that parks one prefill's
    send_kvcache until released, so the abort/free/realloc ordering is enforced
    rather than raced. Greedy decoding, 512 output tokens.

Without the fix, request B was allocated exactly the 32 KV pages and the
mamba slot that request A had just freed. A's stale write then landed in them
mid-generation: a checksum over 64 head-of-prompt KV rows (rows decode never
rewrites) changed 4 ms after prefill logged WRITE_END ret=0, and B's output
diverged from its own baseline at token 269, with 241 of 512 tokens differing.
B returned HTTP 200 and finish_reason=length throughout, so nothing surfaced
to the client. Notably A's prefill had already returned
finish_reason={'type': 'abort'} seconds earlier — the scheduler considered the
request finished while its transfer worker went on to write successfully.

With the fix, when the stalled write completes inside the hold window,
decode logs DEFER_HOLD instead of FREE, B is not admitted while the pages
are held, and B's output is bit-identical to baseline.

One caveat worth reviewer attention

In every control run the hold ended via the timeout, not an ack:

Deferred KV release for room ... timed out after 30.0s without a full drain ack
from prefill; releasing anyway.

The drain ack is sent by the prefill transfer worker, which is precisely the
thread that is stuck, so it cannot arrive while the write is stalled and
SGLANG_DISAGGREGATION_DEFERRED_DECODE_KV_RELEASE_TIMEOUT always decides. When
the write outlives that timeout the original window reopens — in that
configuration B diverged at token 314 (197/512 tokens) with the fix enabled.

So on this workload the feature bounds the exposure to the timeout rather than
closing it. Whether that is acceptable depends on how long a real stalled RDMA
transfer can run; it may be worth either blocking the release until the ack
regardless of elapsed time, or having the sender drop the write when it observes
the room already failed at the point of send rather than only at dequeue.

Draft: opened to share the tested tree and these results, not proposing new code.

Test plan

  • git rev-parse HEAD^{tree} = a7845862cd43eecbd923fff557df1f180fcd8502,
    matching the tree that ran on hardware (independently re-derived by redoing
    the cherry-picks from 6dfdebf0fb in a clean worktree).
  • End-to-end PD serving on 16 GB300 nodes: both arms brought up, health-checked,
    and driven through the three scenarios in the table above.
  • Greedy determinism established first — three runs of the same prompt through
    two different prefill engines returned byte-identical output ids.
  • Ordinary mooncake transfers validated against the live stack before any change.
  • Not run: no unit or CI suite was executed for this branch; the cherry-picked
    PRs ship their own tests (test_deferred_decode_kv_release.py,
    test_nixl_deferred_kv_release.py) but they were not run here. Each arm was
    measured once, so there is no variance estimate.

CI States

Latest PR Test (Base): ❌ Run #34829317759
Latest PR Test (Extra): ❌ Run #34829317288
Latest PR Test (AMD ROCm 10): ➖ No AMD PR run found for this commit.

@Jiminator Jiminator closed this Sep 14, 2026
@Jiminator
Jiminator deleted the kimi-k3-deferred-release-35049-35360 branch September 14, 2026 04:42
@alexnails
alexnails restored the kimi-k3-deferred-release-35049-35360 branch September 14, 2026 05:43
@hnyls2002 hnyls2002 reopened this Sep 14, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants