[Kimi-K3][LMCache] Fuse the state load leg, and fence the KV save against the forward - #2132
zejunchen-zejun wants to merge 2 commits into
Conversation
🏷️ CI GuideRuns automatically on every eligible PR before approval:
Heavy model tests:
|
…arration Responding to review of #2132. **F1 is rejected on the facts, and its remedy adopted anyway.** The claim was that `stats()` has zero callers repo-wide, so the invariant audit never runs. It does run: `PagedStateCheckpointCoordinator.checkpoint_fates` reaches it as `getattr(self.offload, "stats", None)` (page_unit_checkpoint.py:1025), which `BlockManager.state_checkpoint_fates` folds into `checkpoint_funnel`. Verified by construction, not by reading: corrupting the accounting makes `state_offload_invariant_violations` go to 1 in the funnel dict. But a competent reviewer grepped and concluded the opposite, which is itself the finding -- a chain whose middle link is a `getattr` cannot be seen, so it now has a test that asserts a violation reaches the funnel, and the docstring names the prefixed key rather than the bare one. **F2 was asked as a question and the answer is that the path is covered.** A state load armed and then refused after the arm settles through `Scheduler._drop_state_load` -> `abandon_load` before `build_connector_meta` clears its bookkeeping. Now tested: the dispatch is settled, the invariant holds, and -- the part worth pinning -- an abandon does NOT retract the hash, because nothing was attempted and the bytes are still there. **F3 accepted, narrowly.** Two blocks in `state_offload.py` argued against the previous design rather than describing this one, which after a squash-merge a reader cannot check. Both had one load-bearing sentence buried in the history: why `stores_refused` is separate from `stores_failed`, and why capability is derived from config rather than the connector's name. Kept those, dropped the narration. 4259 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
…arration Responding to review of #2132. **F1 is rejected on the facts, and its remedy adopted anyway.** The claim was that `stats()` has zero callers repo-wide, so the invariant audit never runs. It does run: `PagedStateCheckpointCoordinator.checkpoint_fates` reaches it as `getattr(self.offload, "stats", None)` (page_unit_checkpoint.py:1025), which `BlockManager.state_checkpoint_fates` folds into `checkpoint_funnel`. Verified by construction, not by reading: corrupting the accounting makes `state_offload_invariant_violations` go to 1 in the funnel dict. But a competent reviewer grepped and concluded the opposite, which is itself the finding -- a chain whose middle link is a `getattr` cannot be seen, so it now has a test that asserts a violation reaches the funnel, and the docstring names the prefixed key rather than the bare one. **F2 was asked as a question and the answer is that the path is covered.** A state load armed and then refused after the arm settles through `Scheduler._drop_state_load` -> `abandon_load` before `build_connector_meta` clears its bookkeeping. Now tested: the dispatch is settled, the invariant holds, and -- the part worth pinning -- an abandon does NOT retract the hash, because nothing was attempted and the bytes are still there. **F3 accepted, narrowly.** Two blocks in `state_offload.py` argued against the previous design rather than describing this one, which after a squash-merge a reader cannot check. Both had one load-bearing sentence buried in the history: why `stores_refused` is separate from `stores_failed`, and why capability is derived from config rather than the connector's name. Kept those, dropped the narration. 4259 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
0b54083 to
94caab3
Compare
There was a problem hiding this comment.
🟡 Changes recommended
Unresolved critical routing and completion issues, plus a state-index invariant issue, remain.
Get a fresh assessment by requesting another Copilot review.
Pull request overview
This PR fences dense KV saves against forward writes and fuses Kimi-K3 recurrent-state loads into the KV load lifecycle.
Changes:
- Adds producer-event synchronization for dense KV gathers.
- Centralizes state-load ownership, slot reclamation, and completion accounting.
- Expands connector, scheduler, lifecycle, and regression-test coverage.
File summaries
| File | Reviewed changes and final findings |
|---|---|
tests/test_state_offload_index.py |
Tests lifecycle invariants and reclamation. |
tests/test_state_checkpoint.py |
Updates state checkpoint coverage. |
tests/test_page_unit_checkpoint.py |
Tests page-unit checkpoint behavior. |
tests/test_multi_connector.py |
Tests composite connector routing. |
tests/test_lmcache_offload_connector.py.names |
Updates test-name metadata. |
tests/test_lmcache_offload_connector.py |
Tests fused load and state-tier behavior. |
tests/test_kv_drain_liveness.py |
Tests idle transfer draining and completion liveness. |
tests/test_block_manager.py |
Tests state-slot ownership changes. |
atom/model_engine/state_offload.py |
Centralizes state-load lifecycle accounting. Moderate (2 votes): the invariant audit compares cardinalities rather than key sets. |
atom/model_engine/scheduler.py |
Integrates fused-load completion and reclamation handling. |
atom/model_engine/pp_engine_core.py |
Propagates connector completion events. |
atom/model_engine/engine_core.py |
Updates idle KV-work draining. |
atom/model_engine/block_manager.py |
Transfers orphaned state-slot ownership. |
atom/kv_transfer/offload/metadata.py |
Adds state-load and producer-event metadata. |
atom/kv_transfer/offload/hybrid/kimi_k3/state_tier.py |
Implements synchronous state loads and reports. |
atom/kv_transfer/offload/hybrid/kimi_k3/state_object.py |
Updates state transfer integration. |
atom/kv_transfer/offload/hybrid/kimi_k3/connector.py |
Fuses K3 KV/state loads. Critical (2 votes): hash-only verdict keys can be tombstoned after the first quorum, dropping later repeated-load verdicts. Critical (2 votes): load verdicts returning True can be misclassified as terminal saves. |
atom/kv_transfer/offload/dense/connector.py |
Adds KV producer fencing and load refactoring seams. |
atom/kv_transfer/offload/connector.py |
Updates scheduler-shell forwarding. |
atom/kv_transfer/offload/_offload_common.py |
Revises shared completion handling. |
atom/kv_transfer/disaggregation/types.py |
Adjusts connector metadata types. |
atom/kv_transfer/disaggregation/multi/multi_connector.py |
Updates composite connector routing. Critical (1 vote): the dense/K3 routing can resume the forward before K3’s state-only task completes. |
Review details
- Files reviewed: 22/22 changed files
- Comments generated: 4
- Review effort level: Lite
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Review — Kimi-K3 state-offload refactorReviewed at The direction of this PR is right, and worth saying before the findings: it is a consolidation, not a bolt-on. The state-load lifecycle used to be spread across Blockers1. The state-load verdict is emitted only on ranks whose KV leg succeeded, so the quorum for that key is never reached — permanently [verified]# hybrid/kimi_k3/connector.py:310
ok = self._load_kv_bytes(req)
if ok and req.state_load_spec is not None:
ok = self._load_state_bytes(req) # the only thing that writes _hash_verdicts
Then: # disaggregation/aggregator.py:79-83
for key, reports in list(self._reports.items()):
if len(reports) < self._world_size:
continue # not deleted, no TTL, no eviction
This needs no unusual configuration; it happens in a fully homogeneous TP build the first time one rank's LRU differs from another's, which is the normal state of independent per-rank caches. A rank that does not run the state leg must still emit a neutral report on the key (or the key must be self-clearing). Related hardening, latent rather than reachable today: the tier- 2.
|
…arration Responding to review of #2132. **F1 is rejected on the facts, and its remedy adopted anyway.** The claim was that `stats()` has zero callers repo-wide, so the invariant audit never runs. It does run: `PagedStateCheckpointCoordinator.checkpoint_fates` reaches it as `getattr(self.offload, "stats", None)` (page_unit_checkpoint.py:1025), which `BlockManager.state_checkpoint_fates` folds into `checkpoint_funnel`. Verified by construction, not by reading: corrupting the accounting makes `state_offload_invariant_violations` go to 1 in the funnel dict. But a competent reviewer grepped and concluded the opposite, which is itself the finding -- a chain whose middle link is a `getattr` cannot be seen, so it now has a test that asserts a violation reaches the funnel, and the docstring names the prefixed key rather than the bare one. **F2 was asked as a question and the answer is that the path is covered.** A state load armed and then refused after the arm settles through `Scheduler._drop_state_load` -> `abandon_load` before `build_connector_meta` clears its bookkeeping. Now tested: the dispatch is settled, the invariant holds, and -- the part worth pinning -- an abandon does NOT retract the hash, because nothing was attempted and the bytes are still there. **F3 accepted, narrowly.** Two blocks in `state_offload.py` argued against the previous design rather than describing this one, which after a squash-merge a reader cannot check. Both had one load-bearing sentence buried in the history: why `stores_refused` is separate from `stores_failed`, and why capability is derived from config rather than the connector's name. Kept those, dropped the narration. 4259 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
44ca569 to
62ae93c
Compare
Reply to the review on PR #2132Thanks — this is the most useful review this branch has had. Every finding I Four commits on top of the reviewed
BlockersFinding 1 — verdict quorum unreachableConfirmed exactly as described. Fix ( On the latent asymmetry you flagged alongside it — the tier- Finding 2 — reclaim window measured from the wrong instantConfirmed. Fix ( PerformanceFinding 6 —
|
| item | verdict | action |
|---|---|---|
dense/connector.py:195 fence unguarded |
reachable | fixed, cfe9ff20c |
scheduler.py:3151 missed-hash skew |
reachable, reproduced | fixed, 44ca56923 |
connector.py:520 save-only spec |
not reachable | hardened anyway, 44ca56923 |
connector.py:536 dropped state load |
not reachable as described | see below — it caught something else |
connector.py:531 negative slot |
reachable | fixed, 09af0e092 |
| doc drift (4 items) | n/a | fixed, cfe9ff20c |
block_manager.py:2417 re-admit slot leak |
accepted | not fixed — see Deferred |
state_tier.py:114 shared load executor |
accepted | not fixed — see Deferred |
multi_connector.py:567 state-only dead under multi |
accepted | not fixed — see Deferred |
dense/connector.py:195 had three problems in one hunk. A raise escaped —
ModelRunner.process_kvconnector_output has no handler, so a failure building
the event dropped every KV load and save dispatched that step, a far larger
blast radius than the corruption the fence prevents. It now falls back to
torch.cuda.synchronize(): strictly stronger ordering, so degrading to it
cannot produce the torn gather. One event per request bought no ordering over
one per step, so it is hoisted, as both other in-tree fences do. And the event
defaulted to blocking=False while its consumer synchronize()s it for the
whole fenced forward — a spun core for that window, per save; now blocking=True.
Two regression tests, the second verified to fail by the raise escaping.
scheduler.py:3151 is real and I reproduced it end-to-end. The worker writes
the miss verdict in load_state, then _finish_load runs _lookup_unpin — a
real LMCache call, under a different lock — before recording the failure, while
get_finished drains the two in the opposite order. A step landing in that
window publishes the verdict on step N and the failure on step N+1; step N
discarded the unmatched miss, step N+1 saw an empty set. verdict_step <= failure_step always, so it never self-corrects. Same end state as finding 1.
Fix: a miss is evidence about the hash and forget needs no request id, so
the retraction no longer goes through the per-request report at all. missing=
stays on fail_load, which uses it to decide slot reuse — that one is genuinely
per-request.
connector.py:520 is not reachable: the shape needs the sid popped from
_reqs_need_recv during the build, whose only in-build pop is the
_decide_load_after_alloc-False branch — which already ran one call earlier in
should_park_for_load_after_alloc, whose False routes to _drop_state_load,
which clears load_hash before the build. Hardened anyway, because relying on
that made the loop's correctness depend on a field a different owner writes; it
now tests req.load_spec is not None itself.
Something I have to report against myself
09af0e092 — the same commit that fixed findings 1, 2, 6, 7 and 8 — also
restored a guard from the merge base that I had deleted earlier on this branch
(ba5df6304). I restored it verbatim without checking that it still composes.
It does not, in two independent ways, and 6288d3853 backs it out.
It called BlockManager.cancel_state_load — a method this branch deleted.
hasattr(BlockManager, "cancel_state_load") is False, and the only remaining
reference in the tree was the restored call itself. So it raised AttributeError
instead of guarding, and that escapes _process_engine_step_inner,
_process_engine_step (only a finally) and busy_loop (no except): a dead
engine core.
And its predicate, load_hash != -1 and not boundary_tokens, is exactly the
fused state-only shape. At the merge base that shape returned
(False, "per_req_cache_state_boundary"), so needs_remote_load was False and
the guard was never reached for it. The fused path returns
(True, "state_only_load") — so the guard matches every state-only load, and
repairing the AttributeError would have silently dropped the capability this
branch exists to add.
Why it is gone rather than repaired: what the guard was for — two transfers
reporting once — cannot arise while one connector owns both legs. A state-only
load moves no KV (_load_kv_bytes treats lmc <= hbm as the no-op success it
is) and reports once. It arises under kv_connector: multi, where a PD sub wins
the KV leg and _update_waiting_for_remote_kv unparks on that sub's report
alone. That unpark takes no state-leg predicate at the merge base either, so it
is a pre-existing scheduler-wide gap, not something to patch K3-shaped here. The
reasoning is left at the call site so the next reader does not restore it again.
Nothing dynamic could have caught this: test_scheduler.py does not import in
this environment. So 6288d3853 adds
tests/test_scheduler_block_manager_contract.py, a static check that every
BlockManager method the scheduler calls exists — verified to fail against the
deleted-method call.
Deferred, with reasons
Findings 3, 4, 5. I agree with all three, including the part I expected to
argue with: making pp_size > 1 fatal in the worker is redundant against an
engine-side gate that has already declined gracefully, and raising out of
register_kv_caches wedges the server between "load model runner success" and
"ready", so the carefully written message is exactly the one the operator never
sees. Not defending it.
Not doing it here. Moving state_offload.py:418-612 into kv_transfer,
collapsing _build_state_tier's six refusals onto the predicate that already
exists, and deleting the duplicate owner is a structural change with its own
blast radius, and this PR's whole argument is that it is narrower than the one
it replaces. Happy to open it immediately if you would rather review them
together.
The multi unpark gap. _update_waiting_for_remote_kv unparks on one
connector's report with no state-leg predicate, and MultiConnector passes
finished_recving through ungated. I diffed both against the merge base:
identical. Your multi_connector.py:567 note is the same root cause by a
different branch (no sub wins, the state-only load is disowned into a full
recompute — silently dead rather than racing); the other branch is a PD sub
winning the KV leg and unparking while the state H2D is still writing. One fix
in the scheduler, not two K3-shaped patches.
I would fold one more thing into that change: _offload_subconfig compares
kv_connector as a raw string while sub-connectors are built through
KVConnectorFactory.canonical_name, which strips, casefolds and resolves
aliases. So the "two offload sub-connectors are unrepresentable" guarantee that
multi_connector.py leans on in three comments does not hold:
['lmcache_offload', 'lmcache_offload'] -> raises
['lmcache_offload', 'LMCacheConnectorV1'] -> passes # alias
['lmcache_mp', 'lmcache_offload'] -> passes # not in _STATE_TIER_BACKENDS
state_tier.py:114. Accepted as stated, and worth naming plainly: fusing
the completion did not require fusing the execution, and putting the state
leg on the shared max_workers=1 load pool turns a joint load's TTFT from
max(kv, state) into kv + state under a burst. _LOAD_WAIT_WARN_MS went in
the same hunk, so nothing reports it. I have not measured it and I am not going
to claim it is small. It wants its own change with a number attached.
block_manager.py:2417. Accepted. Restoring per-admission identity means
the index holds a list per req_id again rather than a single pending — a data
structure change, not a patch.
Verification
blackclean;ruffclean on every changed file.- Full suite 143 failed / 5167 passed, against 143 / 5157 at the merge
base09cabf1d2— identical failure set, +10 new tests. That also settles the
question you left open: the 9 failures intest_deepseek_v4_wo_a_dequant.py
andtest_heavy_ci_gate.pyare present at the merge base, so they are
pre-existing. - Every regression test added here was run against the previous behaviour and
confirmed to fail first.
Heavy accuracy CI is still skipped (draft, unapproved, no ci:* label). Given
that the accuracy fix is the headline, I would rather run ci:atom before you
approve than after — say the word and I will add the label.
…arration Responding to review of #2132. **F1 is rejected on the facts, and its remedy adopted anyway.** The claim was that `stats()` has zero callers repo-wide, so the invariant audit never runs. It does run: `PagedStateCheckpointCoordinator.checkpoint_fates` reaches it as `getattr(self.offload, "stats", None)` (page_unit_checkpoint.py:1025), which `BlockManager.state_checkpoint_fates` folds into `checkpoint_funnel`. Verified by construction, not by reading: corrupting the accounting makes `state_offload_invariant_violations` go to 1 in the funnel dict. But a competent reviewer grepped and concluded the opposite, which is itself the finding -- a chain whose middle link is a `getattr` cannot be seen, so it now has a test that asserts a violation reaches the funnel, and the docstring names the prefixed key rather than the bare one. **F2 was asked as a question and the answer is that the path is covered.** A state load armed and then refused after the arm settles through `Scheduler._drop_state_load` -> `abandon_load` before `build_connector_meta` clears its bookkeeping. Now tested: the dispatch is settled, the invariant holds, and -- the part worth pinning -- an abandon does NOT retract the hash, because nothing was attempted and the bytes are still there. **F3 accepted, narrowly.** Two blocks in `state_offload.py` argued against the previous design rather than describing this one, which after a squash-merge a reader cannot check. Both had one load-bearing sentence buried in the history: why `stores_refused` is separate from `stores_failed`, and why capability is derived from config rather than the connector's name. Kept those, dropped the narration. 4259 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
44b86db to
bbf5531
Compare
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
It changes correctness-critical offload synchronization and multi-component load lifecycle semantics, which warrants careful human review plus targeted GPU validation.
Review effort: Lite
Findings: 1
Open (2)
…arration Responding to review of #2132. **F1 is rejected on the facts, and its remedy adopted anyway.** The claim was that `stats()` has zero callers repo-wide, so the invariant audit never runs. It does run: `PagedStateCheckpointCoordinator.checkpoint_fates` reaches it as `getattr(self.offload, "stats", None)` (page_unit_checkpoint.py:1025), which `BlockManager.state_checkpoint_fates` folds into `checkpoint_funnel`. Verified by construction, not by reading: corrupting the accounting makes `state_offload_invariant_violations` go to 1 in the funnel dict. But a competent reviewer grepped and concluded the opposite, which is itself the finding -- a chain whose middle link is a `getattr` cannot be seen, so it now has a test that asserts a violation reaches the funnel, and the docstring names the prefixed key rather than the bare one. **F2 was asked as a question and the answer is that the path is covered.** A state load armed and then refused after the arm settles through `Scheduler._drop_state_load` -> `abandon_load` before `build_connector_meta` clears its bookkeeping. Now tested: the dispatch is settled, the invariant holds, and -- the part worth pinning -- an abandon does NOT retract the hash, because nothing was attempted and the bytes are still there. **F3 accepted, narrowly.** Two blocks in `state_offload.py` argued against the previous design rather than describing this one, which after a squash-merge a reader cannot check. Both had one load-bearing sentence buried in the history: why `stores_refused` is separate from `stores_failed`, and why capability is derived from config rather than the connector's name. Kept those, dropped the narration. 4259 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
bbf5531 to
a82902b
Compare
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
It makes correctness-critical changes across connector/scheduler/engine boundaries (including GPU stream ordering), and should receive final human review despite strong test coverage.
Review effort: Lite
Findings: 1
…arration Responding to review of #2132. **F1 is rejected on the facts, and its remedy adopted anyway.** The claim was that `stats()` has zero callers repo-wide, so the invariant audit never runs. It does run: `PagedStateCheckpointCoordinator.checkpoint_fates` reaches it as `getattr(self.offload, "stats", None)` (page_unit_checkpoint.py:1025), which `BlockManager.state_checkpoint_fates` folds into `checkpoint_funnel`. Verified by construction, not by reading: corrupting the accounting makes `state_offload_invariant_violations` go to 1 in the funnel dict. But a competent reviewer grepped and concluded the opposite, which is itself the finding -- a chain whose middle link is a `getattr` cannot be seen, so it now has a test that asserts a violation reaches the funnel, and the docstring names the prefixed key rather than the bare one. **F2 was asked as a question and the answer is that the path is covered.** A state load armed and then refused after the arm settles through `Scheduler._drop_state_load` -> `abandon_load` before `build_connector_meta` clears its bookkeeping. Now tested: the dispatch is settled, the invariant holds, and -- the part worth pinning -- an abandon does NOT retract the hash, because nothing was attempted and the bytes are still there. **F3 accepted, narrowly.** Two blocks in `state_offload.py` argued against the previous design rather than describing this one, which after a squash-merge a reader cannot check. Both had one load-bearing sentence buried in the history: why `stores_refused` is separate from `stores_failed`, and why capability is derived from config rather than the connector's name. Kept those, dropped the narration. 4259 passed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
a82902b to
f70842a
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
A fused load where KV succeeds but the state leg fails can currently fail to force a recompute in vLLM because no load-error blocks are recorded in the state-only/no-op KV shape.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 2
Open (4)
Resolved since last review (1)
There was a problem hiding this comment.
Copilot review overview
🔵 Needs a closer look
It changes correctness-critical ordering and cross-thread/cross-process offload load/save lifecycles across many components, so it warrants final human review plus targeted GPU validation.
Review effort: Lite
Findings: 2
ea3f48d to
08dc08e
Compare
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Critical state-load generation, state-only lookup, and orphan-lifecycle issues remain unresolved.
Review effort: Lite
Findings: 4
Open (7)
Preserve KV-tier coverage for degenerate state-only loads · New Include load generation in connector completion identity · New Make orphan ownership generation-specific · New Reject late completions from reclaimed request generations · New Reset stale offload_load_cancelled before new arbitration · New Count reclaimed orphaned loads as abandoned This audit only compares cardinalities, so it reports a healthy index if one hash is removed and a…
Resolved since last review (2)
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Critical unresolved state-load abandonment, verdict-generation, and vLLM miss-handling issues remain.
Review effort: Lite
Findings: 3
Open (3)
Resolved since last review (7)
Reject late completions from reclaimed request generations Make orphan ownership generation-specific Include load generation in connector completion identity Preserve KV-tier coverage for degenerate state-only loads Reset stale offload_load_cancelled before new arbitration Count reclaimed orphaned loads as abandoned This audit only compares cardinalities, so it reports a healthy index if one hash is removed and a…
| # Hashes whose state `get` missed on some rank. Drained by the engine, | ||
| # which is the only owner of the index that advertises them. | ||
| self._state_load_missed: set[int] = set() |
| # Anything left never reached the metadata (its KV leg was refused after | ||
| # the arm), so drop it rather than attach it to a later step's request. | ||
| self._state_load_seqs.clear() |
| def fail_load(self, req_id, *, missing: bool = False) -> None: | ||
| """No usable load came back. | ||
|
|
||
| `missing` is what separates the two failures the fused load can report. | ||
| The verdict on `failed_loading` covers BOTH legs, so it may mean the KV |
Review — fuse the Kimi-K3 state load legReviewed Fusing the two legs into one task and one completion is the right move, and the deletions are most of why: one report instead of two, one quorum, one place that decides a load is done. The verdict key carrying the load generation ( Most of what follows collapses into two roots rather than fifteen independent defects, so I have organised it that way.
Separate from both, one ordering hazard sits at the heart of the new verdict protocol, so it goes first. Provenance: [verified] = I read the deciding lines at this head. [reported] = the shape matches the code but I did not trace it end to end. 1. The neutral verdict is written before the KV leg runs, and can tombstone the key the real verdict needs [verified]
self._note_state_leg_unrun(req) # :332 -> verdict True for key K
ok = self._load_kv_bytes(req) # :333 -> a multi-MB LMCache retrieve
if ok and req.state_load_spec is not None:
ok = self._load_state_bytes(req) # :335 -> the real verdict for key KBoth reports use the same
So a The "reportable exactly once per key" property is not incidental here; it is what makes the second report unreportable. The window is a full LMCache retrieve against a ~10 ms engine step. Whether the tick actually interleaves is yours to rule out — but the safe shape is to write the neutral verdict only on the paths that are about to skip the leg, not unconditionally ahead of it. 2. Root 1: four places where the connector declined and the index was not toldA An armed leg whose request already parked is dropped with nothing left to report it [reported].
3. Root 2: the sentinel KV spec, and the two costs it is already imposingA failed state-only load records load-error blocks over a fully-resident, possibly shared KV range [reported]. A state-only load ships the full prompt and the full block table for a leg defined to move zero bytes [reported]. With 4. A default-argument change that was not swept to its callers [verified]
On the ATOM path retraction survives through The reasoning in the new docstring is right — a joint 5. PP>1 went from graceful degradation to a hard raise inside the worker [verified]Main logged "the state tier is unsupported under pipeline parallelism (pipeline_parallel_size=%d); paged KV is unaffected" and returned, leaving the dense paged-KV leg working. This head raises This repo's recorded failure mode for that shape is a worker dying between "load model runner success" and the ready signal: the RPC deadlocks and The raise was also moved above the soft Refusing loudly is the right instinct; the right depth is the engine's config-validation path, which already computes 6. Slot lifetime: two ways a destination slot is released while it may still be written [reported]
And the reclaimer itself may never run. 7. A rank without a tier votes on nothing, and the key never dies [reported]
If tier presence is asymmetric across TP ranks it is worse: the tier-bearing ranks report the key, the tier-less rank does not, and 8. Deleting the tier's own executor serialises the legs and unbounds the staging buffers [reported]Two costs from dropping the dedicated 1-worker load executor and its Latency. TTFT for a joint load goes from Memory. A cheaper fusion keeps the state leg on its own executor and fuses only the report: submit inside 9.
|
A state load's lifecycle held three facts in three owners across two
processes: "this request owes a report on hash H" (StateOffloadIndex,
engine), "its state slot stays off the free list"
(BlockManager._orphan_load_slots, engine) and "have both legs landed"
(_JointPark, worker). No object could state
dispatched == settled + outstanding
so no test could assert it.
The state leg now rides the request as `LMCacheReqMeta.state_load_spec`
and runs inside the KV leg's own task, so one dispatch emits exactly one
completion on every path, including a raise. Dense's `_do_load_req` is
split into `_load_kv_bytes` + `_finish_load` to give it that seam. A
state-only load travels the ordinary load path on a no-op KV spec
(hbm == lmc). `StateOffloadIndex` becomes the sole engine-side owner: it
takes the destination slot, absorbs orphan parking, and audits its own
invariant (surfaced as `state_offload_invariant_violations`).
Deleted: `_JointPark`, `metadata.state_loads`, the state-load disposition
channel, the state tier's load executor and staging-lane semaphore, the
engine-side load transport on BlockManager, `_publish_state_loads` /
`_settle_state_load` / `_abandon_state_load`, and the state-only park
branch. PP > 1 now refuses at startup instead of warning.
Rebased onto main (ba51495). The dense save producer fence this branch
carried is dropped in favour of main's #2339, which orders the pack stream
on the device rather than host-synchronizing; `producer_event` on
LMCacheReqMeta and its tests go with it, and staging.py's fence docs now
point at #2339's path. Conflicts with #2238 and #2305 resolved keeping both
sides.
Squashed from 20 commits; the pre-rebase history is at ea3f48d.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…pin check Review round on 08dc08e. A state-only load could be parked and then never reported. Main's #2305 re-takes the lookup pin at build time and drops a load whose tier hit is below `lmcache_cached_tokens`. The state-only no-op spec aims that at the HBM length, and HBM can hold more of the prefix than the KV tier does (1024 resident, 768 in the tier, boundary 768), so the load was dropped after the request had parked. `_ensure_lookup_pin` now skips a spec that reads nothing from the tier, and the state-only branch clears a stale `transfer_end_tokens` so the spec really reads nothing. Under `kv_connector: multi` every state-only resume was declined. Fusing the legs moved the park decision to the KV load owner, and a state-only load has none: every sub's KV answer is 0. That is the production agentic shape ([mooncake producer, lmcache_offload]); main parked it unconditionally. The composite now asks the tier sub when the engine secured a state leg (`offload_joint.load_hash`) and no sub owns KV. `seq.offload_load_cancelled` was never cleared, so a request whose tier sub lost one arbitration kept its state leg suppressed after winning a later one: the KV load went out alone and the forward would resume over a slot nothing filled. Each arbitration now starts clean. State verdicts are keyed by the load generation (`LoadOperationId`), not the request id: a re-admitted request loading the same hash again reused its first load's tombstoned key and its miss was dropped. `orphan()` claims a slot only if it is the one the load writes into. The request id names the request, not the admission, so a later admission's teardown was told the index had its slot, and leaked it. `reclaim` settles as an abandon, so the outcome counters add up to `settled` again. Each new test fails against the previous code. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
3ea7ab1 to
3a61b3d
Compare
|
@valarLip All of it is in Where I verified a [reported] item differently from how it was written, I say so below. §1 — neutral verdict written before the KV leg — agreed, fixed (together with §8)Confirmed. The neutral verdict is now written only on the paths where the leg does not run: the leg was refused (coverage, Tests: §2 / Root 1 — the connector declined and nothing told the indexParked load dropped at build time — agreed, fixed generally. This was not only a K3 problem. #2305's
Re-admission refused while an orphan is outstanding — not reachable on the native path today. A torn-down admission's slot is still protected (see §6). In the native scheduler, a request with a state load in flight is never deallocated and then re-admitted:
§3 / Root 2 — the sentinel KV spec — agreed, done the way you suggestedA state-only load is now a request of its own:
This deletes the Tests: A correction to something I said earlier: in the previous round I told Copilot that the §4 —
|



This PR restructures the Kimi-K3 LMCache offload tier that #2053 landed: the recurrent-state load leg now rides the KV load task, so one dispatch emits exactly one completion and the invariant
dispatched == settled + outstandingis something one object can state and one test can assert.The defect
A state load's lifecycle held three facts in three owners across two processes:
StateOffloadIndexBlockManager._orphan_load_slots_JointParkNo object could state
dispatched == settled + outstanding, so no test could assert it. That is the shape behind the load-path findings this subsystem kept producing._JointParkgrew to 213 lines / 12 methods, every one of them added by a bug fix.The change
The state leg rides the request as
LMCacheReqMeta.state_load_spec, the same shape DSV4'sslot_load_specuses, and runs inside the KV leg's own task:Dense's
_do_load_reqis split into_load_kv_bytes+_finish_loadto give it that seam (behaviour-preserving for dense). A state-only load (KV resident, state in the tier) travels the ordinary load path on a no-op KV spec (hbm == lmc).StateOffloadIndexbecomes the sole engine-side owner: it takes the destination slot, absorbs orphan parking, and audits its own invariant (state_offload_invariant_violationsincheckpoint_funnel).Deleted (each greps to zero):
_JointPark·metadata.state_loads·STATE_LOAD_DISPOSITION_CHANNEL·take_state_load_survived·_publish_state_loads·_settle_state_load·_abandon_state_load·_orphan_load_slots·reconcile_orphan_load_slots·take_state_loads·_tier_can_serve·staging_lanes. The state tier loses its own load executor.Rules the fusion creates, each tested:
ok=Falsemay mean the KV leg, so it does not retract the state hash; only a real stategetmiss does. Every rank reports a verdict (a neutral one if it never ran the leg), or the TP quorum is never reached.reclaimfrees only orphans.Unchanged:
_joint_kv_boundary,_chain_to,_gated_hit,_commit_joint_boundary,_no_joint,allocate,_state_leg_secured,disown_claimed_prefix. Joint resume at any checkpoint rung, and claiming HBM-resident KV so only the delta transfers, both stand. The store leg is out of scope.Also fixed
KimiK3OffloadScheduler.connector_completionended in a barereturn False, so with early block release on,dense.page.source_safe/dense.page.storewere logged as unhandled and their deferred blocks were only freed by the stall timeout. Pre-existing on main; it now delegates to the base._offload_subconfigcompared raw strings while sub-connectors resolve through the registry (strip, casefold, aliases), andlmcache_mpwas missing. Both sets are now compared against the canonical name.settle_state_store(ok=False)) had no test anywhere in the repo; now they do.From review of this PR (
3ea7ab115):kv_connector: multi. Fusing moved the park decision to the KV load owner, and a state-only load has none. That is the production agentic shape ([mooncake producer, lmcache_offload]), which main parked unconditionally. The composite now asks the tier sub when the engine secured a state leg and no sub owns KV.offload_load_cancelledcould suppress a state leg on re-arbitration; each arbitration now starts clean.orphan()matches the slot, not just the request id;reclaimcounts as an abandon.Behaviour changes to sign off on
reclaimgives up._load_kv_bytes/_finish_load). The only observable reordering is[OFFLOAD-LOAD-PROF]logging before the unpin.Known limitations
LocalDiskBackendnever scans its directory).multi, if a P/D sub wins the KV leg while a state leg is also pending, the unpark follows that sub's report alone. This gap is also on main; it is noted at the call site inscheduler.py.Testing
CPU suite (
--ignore=tests/plugin, two modules with missing deps excluded): 6780+ passed; the failure set is identical toorigin/main(28 pre-existing,test_pd_pp/test_mooncake_rail_address).tests/plugin: failure set identical to the parent commit. Every new regression test was checked to fail against the code before its fix.Needs GPU validation before merge (owned by the tester): Kimi-K3 TP8/DCP8 on the AtoMesh
multiconfig, comparing against main:delta_accinside the ±0.024 temp=0 band, withstate_offload_invariant_violations == 0.🤖 Generated with Claude Code