Skip to content

[DSV4] Add PAGE-backed state checkpoints - #1894

Merged
valarLip merged 14 commits into
mainfrom
feat/dsv4-paged-state-checkpoints
Aug 15, 2026
Merged

valarLip merged 14 commits into
mainfrom
feat/dsv4-paged-state-checkpoints

Conversation

@yhl-amd

@yhl-amd yhl-amd commented Aug 14, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Keep each resident request on the existing contiguous, fixed-address Active Slot.
  • Store immutable DSV4 state checkpoints in an ordered set of arbitrary PAGE-sized units, so checkpoint allocation no longer requires a contiguous 13-20-PAGE extent.
  • Make PAGE-backed checkpointing the only path for StateTransfer.copy(); remove ATOM_V4_PAGED_STATE_CHECKPOINTS, the old slot-backed copy checkpoint fallback, and STATE_CKPT_EXTRA_ENTRIES Active Slot headroom.
  • Carry runtime geometry as a frozen PagedStateCheckpointSpec (page-unit bytes, slot bytes, layout id) instead of mutating Config; units-per-checkpoint is derived and never transmitted.
  • Make StateTransfer the single backend checkpoint capability: copy(layout_id), fork(tokens), and none() are invariant-checked. StateRuntime combines that capability with the optional PagedStateCheckpointSpec, validates the complete combination/layout once, and crosses runner RPC as one nested typed payload.
  • Name Active Slot relocation and PAGE checkpoint scatter/gather separately as relocate_state_slots and execute_paged_state_copies; PD/RDMA get_kv_transfer_tensors remains independent.
  • Add a CheckpointRecord control plane with COPYING/READY/EVICTING states, generation checks, reader pins, LRU, and atomic whole-record eviction.
  • Carry one typed StateMaintenanceOps bundle through ScheduledBatch; its relocations, checkpoint stores, and checkpoint restores are drained once for a real batch and executed before the forward. DSV4 Main/Indexer scatter/gather uses one descriptor-driven Triton byte-copy launch.

Allocation and correctness invariants

  • One PAGE remains one CompositePageUnit; one state image owns K arbitrary unit ids in canonical order.
  • A prefix hash points to one READY checkpoint record, never to individual fragments.
  • COPYING records are not visible to prefix lookup.
  • PAGE allocation can reclaim only an entire unpinned checkpoint; no fragment is independently reused.
  • A checkpoint hit first allocates a complete Active Slot and pins its source units before allocating fresh PAGE blocks. PAGE fragments are never adopted as a runtime slot.
  • Main KV and FP8/FP4 Indexer regions share the same block-id ownership transition.
  • Copy-transfer backends publish stable PAGE/state layout and geometry through the explicit runner -> engine -> BlockManager runtime spec. Missing or drifted geometry fails at startup; there is no slot-backed fallback.
  • GDN StateTransfer.fork() checkpointing remains group-backed and is not part of the removed DSV4 copy path.

Implementation notes

  • BlockPool.reserve_units/release_units records checkpoint identity, generation, and piece index for each reserved block id.
  • ModelRunner derives the checkpoint spec from PoolPlan plus transfer.paged_layout_id and constructs one validated StateRuntime; EngineCore rebuilds that runtime from its nested wire payload and passes the same object explicitly through Scheduler to BlockManager.
  • PageUnitCheckpointStore owns hash lookup, ordered units, LRU, pins, two-phase publication, and deferred cleanup.
  • The copy planner intersects the Active Slot and Composite PAGE logical byte streams into physical spans, expands them into 4 KiB tiles, and launches one raw-byte Triton copy for all store/restore ops in the batch.
  • Pipeline parallel and RapidServe configurations fail fast until they have a cross-stage/cross-scheduler ACK and rollback protocol.

Validation

Passed in the current workspace:

  • python3 -m compileall -q atom ...
  • git diff --check
  • Black 26.5.1 and Ruff on the changed paged-state control-plane files
  • 19 pure control-plane tests for StateTransfer, StateRuntime, StateMaintenanceOps, PagedStateCheckpointSpec, and PageUnitCheckpointStore
  • control-plane smoke covering mandatory PAGE construction, store -> READY, restore into a distinct contiguous Active Slot, and removal of the slot-backed fallback
  • explicit allocation from non-contiguous free PAGE ids

Added unit coverage for:

  • arbitrary-unit reservation and atomic release
  • COPYING visibility and READY publication
  • multi-reader pins and whole-record LRU eviction
  • unindex/clear during in-flight copies
  • protected-hit admission accounting
  • StateRuntime COPY/FORK/NONE combinations, illegal transfer/spec/layout pairs, nested wire revalidation, derived unit count, and explicit Scheduler/BlockManager propagation
  • empty batches retaining queued state maintenance, and real batches carrying relocation/store/restore exactly once
  • segmented-stream boundary and partial-tail planning
  • GPU random-byte scatter/gather round trip when a CUDA/ROCm device is available

GitHub Pre Checkin, including Black, Ruff, schema validation, and the full non-GPU unit suite, passes for the validated StateRuntime revision. Offline inference also passes.

DSV4 GPU BF16/FP8/FP4 bitwise and performance validation still requires a ROCm/Triton machine.

@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every eligible PR before approval:

  • ✅ Pre Checkin: Black, Ruff, catalog schema validation, non-GPU unit tests

Heavy model tests:

  • ✅ Run after the PR is approved and Pre Checkin passes
  • ✅ Run immediately when an approval review is submitted
  • ✅ Can be requested before approval with labels
Label Tests
ci:full Run all heavy PR model tests: native ATOM, vLLM, and SGLang
ci:atom Run native ATOM model accuracy tests
ci:vllm Run ATOM vLLM OOT model accuracy tests
ci:sglang Run ATOM SGLang model accuracy tests

Heavy jobs are skipped when the PR is not approved and no matching ci:* label is present.
Add labels via the sidebar or gh pr edit 1894 --add-label <label>

@yhl-amd
yhl-amd force-pushed the feat/dsv4-paged-state-checkpoints branch from 700539a to e73e82f Compare August 14, 2026 03:20
@valarLip

Copy link
Copy Markdown
Collaborator

Reviewed at 891952b2b. The design direction is right — PAGE-backing is what finally decouples checkpoint capacity from max_num_seqs, and it sidesteps the layer-major/entry-major ordering problem that made an elastic KV↔arena boundary intractable. Two correctness holes should block the merge though, plus one hardening item. Both blockers have a red repro at the bottom of this comment.

Both blockers come from the same asymmetry:

where it begins when it is completed
store begin_store, inside take_checkpoint_ops — only ever for a batch that really runs complete_inflight
restore begin_restore, from BlockManager._attach_state_group during allocate() — at admission complete_inflight, unconditionally

The reader pin and the queued gather are taken together but completed apart.


Blocker 1 — an empty batch unpins a gather that has not run yet

Scheduler.schedule calls complete_previous_state_batch() unconditionally at the top, but skips take_state_maintenance_ops() when the batch is empty (deliberately — test_empty_batch_does_not_drain_state_maintenance pins that). So a restore queued on a tick that produces no batch has its source unpinned on the next tick while its op is still in _restore_ops.

The record then becomes READY with pin_count == 0, which is exactly what ensure_free_units looks for. Its units are recycled into ordinary KV blocks by _fresh_block, and whenever the queued gather finally rides a batch it copies those recycled bytes into the resuming request's Active Slot — stale compressor ring and stale SWA window, the #1417 shape the two-phase publication exists to prevent.

Reachable path (scheduler.py):

1261  self.block_manager.allocate(seq, num_cached_blocks)
        └─ _attach_state_group → begin_restore → pin += 1, op appended to _restore_ops
1272  needs_remote_load = self._confirm_remote_load_after_alloc(seq, needs_remote_load)
1277  self._park_for_remote_load(seq, skipped_waiting_requests)
1278  continue

_park_for_remote_load only appends to skipped_waiting_requests and sets WAITING_FOR_REMOTE_KVS — it does not deallocate. The seq never reaches _schedule_prefill_seq, so it never enters scheduled_seqs. With nothing else running, the batch is empty.

This is not gated away: model_runner.py:1674-1685 rejects pipeline_parallel_size > 1 and enable_rapidserve, but not a KV transfer config, and self.kv_connector = get_kvconnector("scheduler", config) (scheduler.py:708) is independent of RapidServe. So the PD/connector path and PAGE checkpoints coexist.

Worth noting the scope is narrow: a non-empty batch is fine — the op is drained into it and the gather executes before the parked seq ever runs. Only the "allocated but the batch came out empty" tick bites.

Verified end to end on the fixture below by asserting the buggy behaviour instead of the invariant: after complete_previous_batch() the pin is 0, ensure_free_units evicts the record, pool.num_free rises by its whole unit count, and the still-queued op goes on naming those units.

Blocker 2 — deallocate does not cancel a gather aimed at the released slot

BlockManager.deallocate calls paged_state_checkpoints.forget_pending(seq), whose whole body is self._pending.pop(id(seq), None) — store intents only. A CheckpointRestoreOp already queued in _restore_ops with dst_slot == seq.per_req_cache_group survives, and the very next line returns that slot to the free list via self.state.release(...).

The next request to pop the slot gets the dead request's checkpoint image scattered over its Active Slot before its own forward — wrong compressor state and wrong sliding window for a request that never asked to resume. If two restores name the same recycled slot in one batch, the two gathers race inside a single Triton launch.

Triggers are all ordinary: a request is admitted on a checkpoint hit (op queued), then aborted, or its remote KV load fails, or it is preempted before the op drains.

Repro output:

AssertionError: a gather into a slot that has gone back on the free list must not
survive deallocate; got (CheckpointRestoreOp(dst_slot=4, unit_ids=(0, 1, 2),
total_bytes=25, layout_id='layout-v1'),)

Suggested direction for both

Same fix shape: tie the pin and the cancellation to the moment the op actually leaves the queue, not to the tick. Concretely, have complete_inflight's restore half only settle restores that were handed out by a previous take_checkpoint_ops, and have forget_pending also drop _restore_ops entries whose dst_slot matches the sequence's group.


Hardening — block_manager.py:443

assert restored, "gated PAGE checkpoint disappeared before attach"

python -O erases this, and the request then proceeds on a freshly popped slot still holding the previous occupant's compressor ring and SWA window — silently. The same file already makes this call the other way at __init__ lines 141-143 ("rather than asserting, which python -O would drop"), where the interval snap became a warning. Either raise unconditionally, or fall back to "no resume" and drop the hit.


Non-blocking notes

  • Checkpoint stores now draw from the live KV pool with no watermark. begin_store → reserve_units → pool.pop() + allocate() takes from the same pool that serves may_append, and the resulting COPYING record is refused by both has_available_units and ensure_free_units (both require READY) for a full tick. Removing STATE_CKPT_EXTRA_ENTRIES at the same time leaves no headroom knob at all. The failure mode changes character here: today a checkpoint that cannot find a free slot is simply dropped and live requests are untouched; after this PR a burst of simultaneous rung crossings (uniform-length requests under a benchmark all cross the ladder together) can lock N × units_per_checkpoint blocks that neither can_append nor _fresh_block can reclaim. On a 256-seq V4 config that measures 14 units per checkpoint, 64 concurrent stores is ~1.4 GB unreclaimable for a tick. A reserve fraction on checkpoint units, or making COPYING cancellable, would keep the upside without the new starvation mode.

  • launch_copy_spans builds the descriptor in Python on the pre-forward critical path. Spans are expanded to one entry per 4 KiB tile, so a ~21 MB Active Slot is ~5,200 tiles × 3 arrays (src_ptrs int64, dst_ptrs int64, valid_bytes) ≈ 15k Python ints marshalled through a synchronizing pageable torch.tensor(..., device=cuda) per op, inside CommonAttentionBuilder.build(). Since the Active Slot is contiguous and PAGE units are equal-sized, the mapping is affine — the only per-op input is the K unit ids. Precomputing the logical span table once and passing just the unit-id table (letting the kernel do unit = offset // unit_bytes) turns this into a block table, the same shape the KV path already uses.

  • BlockPool.reserve_units is documented as atomic but re-enters through the eviction callback. reserve_units → pop → allocate → _unindex → on_evict → _record_evicted → unindex → _release_record → release_units → free mutates _free, _vacant, _cached and _raw_unit_owner in the middle of the reservation loop. Benign today (the re-entrant path only adds free blocks, and a pending hash always names a block held by a live sequence so it cannot be the popped victim), but the "Atomically reserve" docstring is not accurate. Minor: the # Validate ownership before releasing any unit. comment sits below the loop it describes.

Checked, no issue found

Wire format (StateRuntime.to_wire/from_wire → EngineCore → Scheduler → BlockManager) is consistent; h if num_cached_blocks > 0 else -1 at block_manager.py:405 short-circuits fine on a cold start; the byte-stream ordering is symmetric between _active_slot_segments and _page_unit_segments for store and restore; block_bytes/slot_bytes are linear in row width so the sizing cross-check in allocate_per_req_cache holds; slot rows and block envelopes provably do not overlap; no stale references remain to state_copy_pairs, copy_state_entries, release_state_pins, state_transfer_kind, state_fork_tokens, pending_checkpoint, or STATE_CKPT_EXTRA_ENTRIES; TBO's build_ubatch_* paths derive from the single build() result so maintenance ops still execute exactly once per batch; _checkpoint_room's block_manager.state.transfer.forks correctly evaluates False under PAGE checkpointing, so DSV4 spec-decode boundaries stay checkpointable.

Failing PP and RapidServe fast rather than half-supporting them is the right call — PP plus stateful attention has never actually started on our box, and the deferred-hash path there would have stored the wrong bytes silently.

On validation

The PR notes that DSV4 GPU bitwise and performance validation still needs a ROCm/Triton machine. Worth flagging that this PR replaces the path #1417 was filed against, and the slot-backed scheme it replaces was signed off with 1319-question GSM8K at n≥3 per arm plus the #1417 coherence probe — that model needs n=3 because same-code run-to-range reaches 1.6pp. Happy to run gate A (V4-Flash-DSpark tp2, 1319 × 3) and the #1417 prefix-hit coherence probe here once the two blockers are fixed; running them before that would just produce dirty numbers.


Repro — tests/test_page_unit_checkpoint_restore_lifetime.py (both tests fail at 891952b2b)
# SPDX-License-Identifier: MIT

"""Repro for the two restore-lifetime holes in PAGE-backed state checkpoints.

Both come from the same asymmetry. A *store* is begun inside
`take_checkpoint_ops`, so it only ever exists for a batch that really runs. A
*restore* is begun from `BlockManager._attach_state_group` during `allocate()`,
so its reader pin and its queued gather are taken at admission — one for the
tick, the other for whenever a batch next carries it. Nothing keeps the two
together, and both of these tests show them coming apart.
"""

import pytest

from atom.model_engine.block_pool import BlockPool
from atom.model_engine.page_unit_checkpoint import (
    READY,
    PagedStateCheckpointCoordinator,
    PagedStateCheckpointSpec,
)


class _Seq:
    """The only thing `forget_pending` reads off a sequence is its identity."""

    def __init__(self, group):
        self.per_req_cache_group = group
        self.has_per_req_cache = True


def make_coordinator(num_units=20, unit_bytes=10, slot_bytes=25):
    pool = BlockPool(num_units)
    spec = PagedStateCheckpointSpec(
        page_unit_bytes=unit_bytes,
        slot_bytes=slot_bytes,
        layout_id="layout-v1",
    )
    return pool, PagedStateCheckpointCoordinator(pool, spec, enabled=True)


def publish(coordinator, prefix_hash, src_slot=0):
    """A checkpoint whose store has already ridden a batch."""
    assert coordinator.store.begin_store(prefix_hash, src_slot=src_slot) is not None
    coordinator.complete_previous_batch()
    checkpoint_id = next(
        cid
        for cid, record in coordinator.store.records.items()
        if record.prefix_hash == prefix_hash
    )
    assert coordinator.store.records[checkpoint_id].state == READY
    return checkpoint_id


def test_an_empty_batch_unpins_a_gather_that_has_not_run_yet():
    """The pin has to outlive the gather, not the tick.

    `Scheduler.schedule` calls `complete_previous_state_batch` unconditionally
    but skips `take_state_maintenance_ops` when the batch is empty. A request
    admitted on a tick that produces no batch — the PD consumer path parks it
    in `_park_for_remote_load` after `allocate` — therefore has its source
    unpinned while its gather is still queued.

    What that costs, measured on this fixture by asserting the buggy behaviour
    instead of the invariant: once the pin is gone the record is READY and
    unpinned, `ensure_free_units` evicts it, `pool.num_free` rises by its whole
    unit count, and the still-queued gather goes on naming units that
    `_fresh_block` is now free to hand to anyone. That is the #1417 shape.
    """
    _, coordinator = make_coordinator()
    checkpoint_id = publish(coordinator, prefix_hash=7)

    # Admission takes the pin and queues the gather.
    assert coordinator.begin_restore(7, dst_slot=1)
    assert coordinator.store.records[checkpoint_id].pin_count == 1

    # The tick produced an empty batch, so the bundle is NOT drained. The next
    # tick completes the previous batch all the same.
    coordinator.complete_previous_batch()

    assert coordinator.store.records[checkpoint_id].pin_count == 1, (
        "a gather that has not ridden a batch must keep its source pinned; "
        "the pin belongs to the op, not to the tick"
    )


def test_deallocate_does_not_cancel_a_gather_aimed_at_the_released_slot():
    """`forget_pending` drops store intents only; restores keep their target.

    `BlockManager.deallocate` calls `forget_pending(seq)` and then returns
    `seq.per_req_cache_group` to the state pool's free list. A gather already
    queued against that slot survives both, so the next request to pop the slot
    gets the dead request's checkpoint image written over its Active Slot
    before its own forward.
    """
    _, coordinator = make_coordinator()
    publish(coordinator, prefix_hash=7)

    seq = _Seq(group=4)
    assert coordinator.begin_restore(7, dst_slot=seq.per_req_cache_group)

    # The request is aborted / its remote load fails / it is preempted.
    coordinator.forget_pending(seq)

    _, restores = coordinator.take_checkpoint_ops()
    assert restores == (), (
        "a gather into a slot that has gone back on the free list must not "
        f"survive deallocate; got {restores}"
    )

@yhl-amd

yhl-amd commented Aug 15, 2026

Copy link
Copy Markdown
Contributor Author

Addressed the maintainer review blockers in 895827b:

  • Restore reservations now remain queued until take_checkpoint_ops hands them to a real batch; complete_previous_batch settles inflight restores only.
  • deallocate cancels any queued gather targeting the released Active Slot and drops its reader pin before the slot returns to the freelist.
  • A disappeared gated checkpoint now returns the newly allocated slot and raises RuntimeError unconditionally.
  • Corrected the reserve_units atomic wording and ownership-validation comment placement.

Added regressions for the empty-batch lifetime, queued cancellation including EVICTING release, recycled-slot safety, and the fail-fast path. Local validation: 291 related BlockManager/Scheduler/state tests passed; Black, changed-file Ruff, compileall, and diff check passed.

The watermark and copy-descriptor observations remain non-blocking follow-up work pending stress/performance data.

@yhl-amd

yhl-amd commented Aug 15, 2026

Copy link
Copy Markdown
Contributor Author

E2E validation on 895827b8 completed with 8x MI355X using image rocm/atom-dev@sha256:bd915ac97e249e1cf43282bc9232b79e18be76bece0b8788caa463e3501a3ec6.

  • Full GSM8K, 64-shot, 1,319 requests per round; every prompt exceeded the 8,192-token checkpoint interval (min 9,098, max 12,021, mean 10,463.28).
  • Round 1: strict 0.9613343442, flexible 0.9605761941, 548 s. The 1,319 token-ID records were fsynced; SHA256 6192978780a1fe3602ea904cc05218656733808c168c856818e2e117b35cb155.
  • Round 2 replay on the same server: 1,319/1,319 HTTP 200 on attempt 1, 381.42 s. Re-evaluated with the same lm-eval filters: strict/flexible 0.9620924943 (1,269/1,319), so no accuracy regression.
  • Round-2 cache delta: 8,249,600 cached / 13,801,066 full tokens = 59.7751% realized hit; 74.7982% compressed hit; 15.0231% lost to the 8K checkpoint boundary. Checkpoints kept +1,320, dropped 0, evicted 0.
  • Startup reported PAGE-backed checkpoints with 1,521,920-byte units, 25,395,200-byte slots, and 17 units/checkpoint. Server logs show 8,192-token restores (cached: [8192]); no traceback/assert/runtime/OOM/GPU-fault patterns, and the container was not OOM-killed or restarted.

@yhl-amd

yhl-amd commented Aug 15, 2026

Copy link
Copy Markdown
Contributor Author

Follow-up MTP/DSpark coverage on the same 895827b8 image and 64-shot GSM8K workload:

  • Re-ran both rounds with real MTP-1 enabled: --method mtp --num-speculative-tokens 1; startup confirmed SpeculativeConfig(method=mtp, num_spec_tokens=1) and loaded the mtp.0.* draft layer.
  • MTP changes the state geometry from ring 128 / 25,395,200-byte slot / 17 PAGE units to ring 129 / 26,782,720-byte slot / 18 PAGE units.
  • Round 1: strict 0.9545109932 (1,259/1,319), flexible 0.9537528431; 534 s; acceptance 95.9064% (60,398/62,976), 1.9591 tok/fwd.
  • Round 2: 1,319/1,319 first-attempt success; strict 0.9583017437 (1,264/1,319), flexible 0.9575435936; 366.57 s; acceptance 95.8537% (60,292/62,900). Accuracy improved by 5 questions vs the MTP cold round, so the 18-unit checkpoint restore path did not regress accuracy.
  • Round-2 cache delta: 7,365,120 cached / 13,801,066 full = 53.3663% realized hit; 66.5734% compressed hit; 13.2071% lost to the 8K boundary. Checkpoints kept +1,321, dropped 0, evicted 0. No fatal/OOM/GPU-fault patterns.
  • Compared with non-MTP, MTP was ~2.6% faster on round 1 and ~3.9% faster on round 2. Scores varied by ~0.38-0.68 pp across the distinct batching/spec paths; the within-MTP checkpoint-hit round did not decline.

DSpark is not validly testable with this local checkpoint: it has num_nextn_predict_layers=1 and mtp.0.*, but no dspark_block_size, DSpark target-layer metadata, or separate DSpark draft checkpoint. DSpark is a different block drafter, not a runtime toggle that can be safely forced onto serial-MTP weights. A V4-Pro-DSpark checkpoint is required for that follow-up.

@valarLip
valarLip merged commit c2d40e2 into main Aug 15, 2026
31 of 32 checks passed
@valarLip
valarLip deleted the feat/dsv4-paged-state-checkpoints branch August 15, 2026 15:31
ganyi1996ppo added a commit that referenced this pull request Aug 19, 2026
… a -1 interval

Rebased onto main's PAGE-backed checkpoint work (#1874, #1894, #1943, #1880),
which rewrote this subsystem underneath the branch. Squashed to one commit
because the three original commits each re-conflicted against the new base and
against each other's resolutions; the reasoning from all three is kept below.

--- allocate state slots per need, not by group

A "state cache group" was `1 + num_spec` slots wide and was the unit of
everything: allocation, admission, sizing, and the checkpoint index. But a
checkpoint has no speculation to roll back -- it holds a committed state -- so
filing one cost a full group and wasted `num_spec/(1 + num_spec)` of its bytes.
At two speculative tokens that is two thirds.

The slot is now the unit. `--state-checkpoint-slots 64` buys 64 checkpoints for
64 slots instead of 192, and the slots it no longer takes stay in the paged KV
pool. This is only possible because a request's slots need not be adjacent,
which the kernels never required: the ssm kernel gathers each index out of the
indices tensor and the conv path is handed column 0 alone. Contiguity was
manufactured by `prepare_state_indices` writing `arange(base, base + width)`;
it now writes the seq's own slot list straight in.

`StateGroupPool` -> `StateSlotPool`, `Sequence.per_req_cache_group` ->
`state_slots`, whose element 0 is the committed state. The setter re-points [0]
and preserves [1:], because speculation scratch persists across forwards.
`--state-checkpoint-groups` still parses, as an alias.

--- -1 turns off the interval ladder without turning off checkpointing

The interval is a guess about where reuse will resume; a demand rung is a
position a request was actually refused at. On the SemiAnalysis cc-traces the
8192 ladder placed ~30x the writes of the demand rung alone and caught reuse the
demand already reaches -- 0.0% of resumes landed on a ladder rung -- while every
rung costs the prompt that keeps it an extra prefill chunk.

  >0  a rung every N tokens (unchanged, still the default)
   0  state checkpointing off entirely (unchanged)
  -1  no interval rungs; the demand rung and prompt-end anchor still place them

-1 rather than reusing 0 because 0 is the documented contract and is reachable
by accident: the grid snap rounds an off-grid interval down and can land on 0,
so a --block-size typo currently fails safe. Three of the four sites are not the
arithmetic you would guess -- `pos % interval` under -1 admits *every* position
rather than none, and `pos - last < -1` is true for every pos.

--- make the demand rung switchable

A demand rung is 47% of checkpoint writes on the cc-traces and reads back 2.8%
of the time, against 85.2% for a prompt-end anchor. Gated independently of
--state-checkpoint-interval-tokens, because the demand is not part of the
interval grid. Default unchanged. The refusal is still measured when the
placement is off -- switching off a rung must not blind the diagnostic that
justifies it.

--- reconciliation with main

Dropped: the copy/pending-checkpoint path (`_commit_pending`, `record_copy`,
`take_copies`, `Sequence.pending_checkpoint`). DeepSeek-V4 checkpointing moved
to `PagedStateCheckpointCoordinator`, and `StateSlotPool` now rejects
`transfer.copies` outright. The fork path (GDN), where the measured wins are, is
kept in full.

`readable_midstep` is carried on main's `StateTransfer` in `state_runtime.py`
(wire format included) rather than on the branch's copy.

Two bugs this rebase exposed, both fixed here:

- `PagedStateCheckpointCoordinator` did not implement the midstep half of the
  `StateCache` protocol, so `checkpoint_cut` raised on every V4 batch. A PAGE
  image is not readable midstep, so the three methods are the no-ops the
  protocol documents.
- `_record_checkpoint_end` could place the anchor past the last matchable block.
  `can_allocate` stops one block short of the prompt, so a checkpoint filed
  under the final block's hash is one no scan looks up -- and being stored, it
  evicted the ladder rung that would have served the resume, taking an identical
  re-request from 8 hit blocks to 0. Capped at `(n_hash_blocks - 1) * hbs`.

Tests: `tests/test_state_checkpoint.py` 170 passed. Full suite 45 failed /
2300 passed, a strict subset of origin/main's own 114 pre-existing failures --
zero regressions, verified by set difference against a clean origin/main
worktree. black clean; ruff no new findings.

Co-Authored-By: Claude <noreply@anthropic.com>
ganyi1996ppo added a commit that referenced this pull request Aug 19, 2026
… a -1 interval

Rebased onto main's PAGE-backed checkpoint work (#1874, #1894, #1943, #1880),
which rewrote this subsystem underneath the branch. Squashed to one commit
because the three original commits each re-conflicted against the new base and
against each other's resolutions; the reasoning from all three is kept below.

--- allocate state slots per need, not by group

A "state cache group" was `1 + num_spec` slots wide and was the unit of
everything: allocation, admission, sizing, and the checkpoint index. But a
checkpoint has no speculation to roll back -- it holds a committed state -- so
filing one cost a full group and wasted `num_spec/(1 + num_spec)` of its bytes.
At two speculative tokens that is two thirds.

The slot is now the unit. `--state-checkpoint-slots 64` buys 64 checkpoints for
64 slots instead of 192, and the slots it no longer takes stay in the paged KV
pool. This is only possible because a request's slots need not be adjacent,
which the kernels never required: the ssm kernel gathers each index out of the
indices tensor and the conv path is handed column 0 alone. Contiguity was
manufactured by `prepare_state_indices` writing `arange(base, base + width)`;
it now writes the seq's own slot list straight in.

`StateGroupPool` -> `StateSlotPool`, `Sequence.per_req_cache_group` ->
`state_slots`, whose element 0 is the committed state. The setter re-points [0]
and preserves [1:], because speculation scratch persists across forwards.
`--state-checkpoint-groups` still parses, as an alias.

--- -1 turns off the interval ladder without turning off checkpointing

The interval is a guess about where reuse will resume; a demand rung is a
position a request was actually refused at. On the SemiAnalysis cc-traces the
8192 ladder placed ~30x the writes of the demand rung alone and caught reuse the
demand already reaches -- 0.0% of resumes landed on a ladder rung -- while every
rung costs the prompt that keeps it an extra prefill chunk.

  >0  a rung every N tokens (unchanged, still the default)
   0  state checkpointing off entirely (unchanged)
  -1  no interval rungs; the demand rung and prompt-end anchor still place them

-1 rather than reusing 0 because 0 is the documented contract and is reachable
by accident: the grid snap rounds an off-grid interval down and can land on 0,
so a --block-size typo currently fails safe. Three of the four sites are not the
arithmetic you would guess -- `pos % interval` under -1 admits *every* position
rather than none, and `pos - last < -1` is true for every pos.

--- make the demand rung switchable

A demand rung is 47% of checkpoint writes on the cc-traces and reads back 2.8%
of the time, against 85.2% for a prompt-end anchor. Gated independently of
--state-checkpoint-interval-tokens, because the demand is not part of the
interval grid. Default unchanged. The refusal is still measured when the
placement is off -- switching off a rung must not blind the diagnostic that
justifies it.

--- DeepSeek-V4

Unaffected by the slot-vs-group change, and that claim is now checked rather
than asserted: DSV4 declares `entries_per_req=1` unconditionally (the MTP/DSpark
lookahead widens the slot via `win_with_spec`, it never multiplies the count),
so `state_slots_per_req == 1`, `pop_many(1)` pops the same index `pop()` did,
and slot == group exactly as before. `v4_pool_geometry.py` and `sub_pool_spec.py`
are untouched, so the `_physical_slots` reversal and DSV4's pool size are both
byte-identical to main. DSV4 passes no `extra_entries`, so `--state-checkpoint-
slots` is inert for it.

The *anchor*, though, did reach DSV4 -- and cost it. `PagedStateCheckpointCoord-
inator.applies()` is true for any V4 seq, so `_record_checkpoint_end` reserved a
prompt-end anchor and `checkpoint_cut` shortened a prefill chunk onto it. But the
coordinator files one pending checkpoint per seq (`_pending[id(seq)]`, and it
`del`s `boundary_blocks`), so the prompt-end checkpoint landing a chunk later
overwrote the anchor before either was stored. Measured: one extra prefill chunk
per prompt for a hit rate that did not move (identical re-send 0 blocks either
way, continuation 11 either way).

So the anchor is now gated on `keeps_interior_boundaries`, which `StateSlotPool`
answers True (each boundary is its own slot in the index, so both survive) and
the PAGE coordinator answers False. Asked as a capability rather than by naming
the backend, so a future multi-boundary copy class opts in by answering yes.
DSV4 is back to main's one-cut prefill; GDN keeps the anchor. Pinned by
`test_a_last_boundary_only_class_is_not_anchored_for`.

--- reconciliation with main

Dropped: the copy/pending-checkpoint path (`_commit_pending`, `record_copy`,
`take_copies`, `Sequence.pending_checkpoint`). DeepSeek-V4 checkpointing moved
to `PagedStateCheckpointCoordinator`, and `StateSlotPool` now rejects
`transfer.copies` outright. The fork path (GDN), where the measured wins are, is
kept in full.

`readable_midstep` is carried on main's `StateTransfer` in `state_runtime.py`
(wire format included) rather than on the branch's copy.

Two bugs this rebase exposed, both fixed here:

- `PagedStateCheckpointCoordinator` did not implement the midstep half of the
  `StateCache` protocol, so `checkpoint_cut` raised on every V4 batch. This was
  introduced *by this branch*, not latent in main: `readable_midstep` does not
  exist on main at all, and main's `checkpoint_cut` never consults it. A PAGE
  image is not readable midstep, so the three methods are the no-ops the
  protocol documents.
- `_record_checkpoint_end` could place the anchor past the last matchable block.
  `can_allocate` stops one block short of the prompt, so a checkpoint filed
  under the final block's hash is one no scan looks up -- and being stored, it
  evicted the ladder rung that would have served the resume, taking an identical
  re-request from 8 hit blocks to 0. Capped at `(n_hash_blocks - 1) * hbs`.

Tests: `tests/test_state_checkpoint.py` 171 passed; the state/checkpoint and
DSV4/LMCache suites together 632 passed. Full suite 45 failed /
2300 passed, a strict subset of origin/main's own 114 pre-existing failures --
zero regressions, verified by set difference against a clean origin/main
worktree. black clean; ruff no new findings.

Co-Authored-By: Claude <noreply@anthropic.com>
zejunchen-zejun pushed a commit that referenced this pull request Aug 20, 2026
Rebuilds the K3 work on main's restructured offload package instead of
merging the old branch: `connector.py` is now a dispatch shell selecting a
layout from config alone, so K3 becomes a third layout (`kimi_k3`) beside
`dense` and `hybrid`, sharing dense's KV path and adding only the KDA
per-request state tier.

  atom/kv_transfer/offload/hybrid/kimi_k3/
    connector.py    KimiK3OffloadConnector (worker) / KimiK3OffloadScheduler
    state_tier.py   the tier and its joint park
    state_object.py the entry codec
    staging.py      single-entry staged transfer

The scheduler half is six small overrides on `DenseOffloadScheduler` rather
than a fork of it:

  * `_decide_load_after_alloc` clamps a hybrid's KV leg to the boundary the
    state leg is aimed at, and refuses the load outright when there is none.
    A hybrid's state is the compressed history of exactly `[0, hbm)`, so
    raising the KV-loaded length past that has the forward skip tokens the
    recurrence never saw -- wrong output, no exception.
  * `should_defer_free` releases the blocks of a save that was never handed
    out once the backend stops draining. Holding them is what turned a
    stopped backend into a stopped engine.
  * `_may_emit_save` bounds how many requests may have a save outstanding,
    since each one pins its blocks. Dense gains the seam (unbounded there,
    it does not pin); `max_pending_saves` moves to `_offload_common` so both
    it and DSV4 read the same `OFFLOAD_MAX_PENDING_SAVES`.
  * `build_connector_meta` drains the state-load queue onto the metadata.
  * `connector_completion` routes the tier's two channels.

Reports ride main's generic `ConnectorCompletion` channels rather than extra
`KVConnectorOutput` fields, so TP quorum is the aggregator's existing job and
`multi_connector` needs only the tier re-exposure and `enqueue_state_loads`.

Env vars: four are gone. `OFFLOAD_STATE` and `OFFLOAD_KV_FOR_HYBRID` because
a resumable prefix needs both legs -- state is offloaded exactly when KV is,
and the connector being configured is the only switch. `OFFLOAD_SAVE_MAX_INFLIGHT`
and `OFFLOAD_SAVE_STALL_S` fold into `OFFLOAD_MAX_PENDING_SAVES` and a
constant. `OFFLOAD_STATE_MIN_LOAD_TOKENS` defaulted to 0 and a token floor is
the wrong unit for a flat-cost entry. The staging ring is gated on the hosting
connector (`state_pool(..., offload_hosted=)`) so a deployment without one
does not pay for a ring nothing can drain. `docs/environment_variables.md`
gains a full offload section covering main's existing vars as well.

`STATE_CKPT_EXTRA_ENTRIES` is deliberately NOT reintroduced: the old branch
predates #1894, which deleted it, and it is a state-pool sizing knob with no
bearing on offload. `state_pool` keeps `extra_entries` as the backend-declared
parameter main left it as.

Also removed as dead: `STATE_OFFLOAD_LOADS_WIRED` (permanently True),
`clamp_state_boundary` (no caller), `StateOffloadIndex.kv_offload_enabled`
(always True once the index is only built for a hosting connector), a
duplicate `_submit_state_spills` override, and ~150 lines of `staging.py`
that re-implemented `atom_lmcache_staging`. Its `_env_flag` fix -- strip, and
read empty as off, so `VAR=` does not turn a flag ON -- moves to the shared
copy.

Measured on c10 (Kimi-K3, TP8, 3600s): P90 interactivity 31.19 -> 41.03,
per-GPU throughput 4313 -> 4784 tok/s, combined hit rate 91.62% -> 93.21%.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
zejunchen-zejun pushed a commit that referenced this pull request Aug 20, 2026
Rebuilds the K3 work on main's restructured offload package instead of
merging the old branch: `connector.py` is now a dispatch shell selecting a
layout from config alone, so K3 becomes a third layout (`kimi_k3`) beside
`dense` and `hybrid`, sharing dense's KV path and adding only the KDA
per-request state tier.

  atom/kv_transfer/offload/hybrid/kimi_k3/
    connector.py    KimiK3OffloadConnector (worker) / KimiK3OffloadScheduler
    state_tier.py   the tier and its joint park
    state_object.py the entry codec
    staging.py      single-entry staged transfer

The scheduler half is six small overrides on `DenseOffloadScheduler` rather
than a fork of it:

  * `_decide_load_after_alloc` clamps a hybrid's KV leg to the boundary the
    state leg is aimed at, and refuses the load outright when there is none.
    A hybrid's state is the compressed history of exactly `[0, hbm)`, so
    raising the KV-loaded length past that has the forward skip tokens the
    recurrence never saw -- wrong output, no exception.
  * `should_defer_free` releases the blocks of a save that was never handed
    out once the backend stops draining. Holding them is what turned a
    stopped backend into a stopped engine.
  * `_may_emit_save` bounds how many requests may have a save outstanding,
    since each one pins its blocks. Dense gains the seam (unbounded there,
    it does not pin); `max_pending_saves` moves to `_offload_common` so both
    it and DSV4 read the same `OFFLOAD_MAX_PENDING_SAVES`.
  * `build_connector_meta` drains the state-load queue onto the metadata.
  * `connector_completion` routes the tier's two channels.

Reports ride main's generic `ConnectorCompletion` channels rather than extra
`KVConnectorOutput` fields, so TP quorum is the aggregator's existing job and
`multi_connector` needs only the tier re-exposure and `enqueue_state_loads`.

Env vars: four are gone. `OFFLOAD_STATE` and `OFFLOAD_KV_FOR_HYBRID` because
a resumable prefix needs both legs -- state is offloaded exactly when KV is,
and the connector being configured is the only switch. `OFFLOAD_SAVE_MAX_INFLIGHT`
and `OFFLOAD_SAVE_STALL_S` fold into `OFFLOAD_MAX_PENDING_SAVES` and a
constant. `OFFLOAD_STATE_MIN_LOAD_TOKENS` defaulted to 0 and a token floor is
the wrong unit for a flat-cost entry. The staging ring is gated on the hosting
connector (`state_pool(..., offload_hosted=)`) so a deployment without one
does not pay for a ring nothing can drain. `docs/environment_variables.md`
gains a full offload section covering main's existing vars as well.

`STATE_CKPT_EXTRA_ENTRIES` is deliberately NOT reintroduced: the old branch
predates #1894, which deleted it, and it is a state-pool sizing knob with no
bearing on offload. `state_pool` keeps `extra_entries` as the backend-declared
parameter main left it as.

Also removed as dead: `STATE_OFFLOAD_LOADS_WIRED` (permanently True),
`clamp_state_boundary` (no caller), `StateOffloadIndex.kv_offload_enabled`
(always True once the index is only built for a hosting connector), a
duplicate `_submit_state_spills` override, and ~150 lines of `staging.py`
that re-implemented `atom_lmcache_staging`. Its `_env_flag` fix -- strip, and
read empty as off, so `VAR=` does not turn a flag ON -- moves to the shared
copy.

Measured on c10 (Kimi-K3, TP8, 3600s): P90 interactivity 31.19 -> 41.03,
per-GPU throughput 4313 -> 4784 tok/s, combined hit rate 91.62% -> 93.21%.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
zejunchen-zejun pushed a commit that referenced this pull request Aug 26, 2026
Rebuilds the K3 work on main's restructured offload package instead of
merging the old branch: `connector.py` is now a dispatch shell selecting a
layout from config alone, so K3 becomes a third layout (`kimi_k3`) beside
`dense` and `hybrid`, sharing dense's KV path and adding only the KDA
per-request state tier.

  atom/kv_transfer/offload/hybrid/kimi_k3/
    connector.py    KimiK3OffloadConnector (worker) / KimiK3OffloadScheduler
    state_tier.py   the tier and its joint park
    state_object.py the entry codec
    staging.py      single-entry staged transfer

The scheduler half is six small overrides on `DenseOffloadScheduler` rather
than a fork of it:

  * `_decide_load_after_alloc` clamps a hybrid's KV leg to the boundary the
    state leg is aimed at, and refuses the load outright when there is none.
    A hybrid's state is the compressed history of exactly `[0, hbm)`, so
    raising the KV-loaded length past that has the forward skip tokens the
    recurrence never saw -- wrong output, no exception.
  * `should_defer_free` releases the blocks of a save that was never handed
    out once the backend stops draining. Holding them is what turned a
    stopped backend into a stopped engine.
  * `_may_emit_save` bounds how many requests may have a save outstanding,
    since each one pins its blocks. Dense gains the seam (unbounded there,
    it does not pin); `max_pending_saves` moves to `_offload_common` so both
    it and DSV4 read the same `OFFLOAD_MAX_PENDING_SAVES`.
  * `build_connector_meta` drains the state-load queue onto the metadata.
  * `connector_completion` routes the tier's two channels.

Reports ride main's generic `ConnectorCompletion` channels rather than extra
`KVConnectorOutput` fields, so TP quorum is the aggregator's existing job and
`multi_connector` needs only the tier re-exposure and `enqueue_state_loads`.

Env vars: four are gone. `OFFLOAD_STATE` and `OFFLOAD_KV_FOR_HYBRID` because
a resumable prefix needs both legs -- state is offloaded exactly when KV is,
and the connector being configured is the only switch. `OFFLOAD_SAVE_MAX_INFLIGHT`
and `OFFLOAD_SAVE_STALL_S` fold into `OFFLOAD_MAX_PENDING_SAVES` and a
constant. `OFFLOAD_STATE_MIN_LOAD_TOKENS` defaulted to 0 and a token floor is
the wrong unit for a flat-cost entry. The staging ring is gated on the hosting
connector (`state_pool(..., offload_hosted=)`) so a deployment without one
does not pay for a ring nothing can drain. `docs/environment_variables.md`
gains a full offload section covering main's existing vars as well.

`STATE_CKPT_EXTRA_ENTRIES` is deliberately NOT reintroduced: the old branch
predates #1894, which deleted it, and it is a state-pool sizing knob with no
bearing on offload. `state_pool` keeps `extra_entries` as the backend-declared
parameter main left it as.

Also removed as dead: `STATE_OFFLOAD_LOADS_WIRED` (permanently True),
`clamp_state_boundary` (no caller), `StateOffloadIndex.kv_offload_enabled`
(always True once the index is only built for a hosting connector), a
duplicate `_submit_state_spills` override, and ~150 lines of `staging.py`
that re-implemented `atom_lmcache_staging`. Its `_env_flag` fix -- strip, and
read empty as off, so `VAR=` does not turn a flag ON -- moves to the shared
copy.

Measured on c10 (Kimi-K3, TP8, 3600s): P90 interactivity 31.19 -> 41.03,
per-GPU throughput 4313 -> 4784 tok/s, combined hit rate 91.62% -> 93.21%.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
zejunchen-zejun pushed a commit that referenced this pull request Aug 27, 2026
Rebuilds the K3 work on main's restructured offload package instead of
merging the old branch: `connector.py` is now a dispatch shell selecting a
layout from config alone, so K3 becomes a third layout (`kimi_k3`) beside
`dense` and `hybrid`, sharing dense's KV path and adding only the KDA
per-request state tier.

  atom/kv_transfer/offload/hybrid/kimi_k3/
    connector.py    KimiK3OffloadConnector (worker) / KimiK3OffloadScheduler
    state_tier.py   the tier and its joint park
    state_object.py the entry codec
    staging.py      single-entry staged transfer

The scheduler half is six small overrides on `DenseOffloadScheduler` rather
than a fork of it:

  * `_decide_load_after_alloc` clamps a hybrid's KV leg to the boundary the
    state leg is aimed at, and refuses the load outright when there is none.
    A hybrid's state is the compressed history of exactly `[0, hbm)`, so
    raising the KV-loaded length past that has the forward skip tokens the
    recurrence never saw -- wrong output, no exception.
  * `should_defer_free` releases the blocks of a save that was never handed
    out once the backend stops draining. Holding them is what turned a
    stopped backend into a stopped engine.
  * `_may_emit_save` bounds how many requests may have a save outstanding,
    since each one pins its blocks. Dense gains the seam (unbounded there,
    it does not pin); `max_pending_saves` moves to `_offload_common` so both
    it and DSV4 read the same `OFFLOAD_MAX_PENDING_SAVES`.
  * `build_connector_meta` drains the state-load queue onto the metadata.
  * `connector_completion` routes the tier's two channels.

Reports ride main's generic `ConnectorCompletion` channels rather than extra
`KVConnectorOutput` fields, so TP quorum is the aggregator's existing job and
`multi_connector` needs only the tier re-exposure and `enqueue_state_loads`.

Env vars: four are gone. `OFFLOAD_STATE` and `OFFLOAD_KV_FOR_HYBRID` because
a resumable prefix needs both legs -- state is offloaded exactly when KV is,
and the connector being configured is the only switch. `OFFLOAD_SAVE_MAX_INFLIGHT`
and `OFFLOAD_SAVE_STALL_S` fold into `OFFLOAD_MAX_PENDING_SAVES` and a
constant. `OFFLOAD_STATE_MIN_LOAD_TOKENS` defaulted to 0 and a token floor is
the wrong unit for a flat-cost entry. The staging ring is gated on the hosting
connector (`state_pool(..., offload_hosted=)`) so a deployment without one
does not pay for a ring nothing can drain. `docs/environment_variables.md`
gains a full offload section covering main's existing vars as well.

`STATE_CKPT_EXTRA_ENTRIES` is deliberately NOT reintroduced: the old branch
predates #1894, which deleted it, and it is a state-pool sizing knob with no
bearing on offload. `state_pool` keeps `extra_entries` as the backend-declared
parameter main left it as.

Also removed as dead: `STATE_OFFLOAD_LOADS_WIRED` (permanently True),
`clamp_state_boundary` (no caller), `StateOffloadIndex.kv_offload_enabled`
(always True once the index is only built for a hosting connector), a
duplicate `_submit_state_spills` override, and ~150 lines of `staging.py`
that re-implemented `atom_lmcache_staging`. Its `_env_flag` fix -- strip, and
read empty as off, so `VAR=` does not turn a flag ON -- moves to the shared
copy.

Measured on c10 (Kimi-K3, TP8, 3600s): P90 interactivity 31.19 -> 41.03,
per-GPU throughput 4313 -> 4784 tok/s, combined hit rate 91.62% -> 93.21%.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
ganyi1996ppo added a commit that referenced this pull request Aug 27, 2026
… a -1 interval

Rebased onto main's PAGE-backed checkpoint work (#1874, #1894, #1943, #1880),
which rewrote this subsystem underneath the branch. Squashed to one commit
because the three original commits each re-conflicted against the new base and
against each other's resolutions; the reasoning from all three is kept below.

--- allocate state slots per need, not by group

A "state cache group" was `1 + num_spec` slots wide and was the unit of
everything: allocation, admission, sizing, and the checkpoint index. But a
checkpoint has no speculation to roll back -- it holds a committed state -- so
filing one cost a full group and wasted `num_spec/(1 + num_spec)` of its bytes.
At two speculative tokens that is two thirds.

The slot is now the unit. `--state-checkpoint-slots 64` buys 64 checkpoints for
64 slots instead of 192, and the slots it no longer takes stay in the paged KV
pool. This is only possible because a request's slots need not be adjacent,
which the kernels never required: the ssm kernel gathers each index out of the
indices tensor and the conv path is handed column 0 alone. Contiguity was
manufactured by `prepare_state_indices` writing `arange(base, base + width)`;
it now writes the seq's own slot list straight in.

`StateGroupPool` -> `StateSlotPool`, `Sequence.per_req_cache_group` ->
`state_slots`, whose element 0 is the committed state. The setter re-points [0]
and preserves [1:], because speculation scratch persists across forwards.
`--state-checkpoint-groups` still parses, as an alias.

--- -1 turns off the interval ladder without turning off checkpointing

The interval is a guess about where reuse will resume; a demand rung is a
position a request was actually refused at. On the SemiAnalysis cc-traces the
8192 ladder placed ~30x the writes of the demand rung alone and caught reuse the
demand already reaches -- 0.0% of resumes landed on a ladder rung -- while every
rung costs the prompt that keeps it an extra prefill chunk.

  >0  a rung every N tokens (unchanged, still the default)
   0  state checkpointing off entirely (unchanged)
  -1  no interval rungs; the demand rung and prompt-end anchor still place them

-1 rather than reusing 0 because 0 is the documented contract and is reachable
by accident: the grid snap rounds an off-grid interval down and can land on 0,
so a --block-size typo currently fails safe. Three of the four sites are not the
arithmetic you would guess -- `pos % interval` under -1 admits *every* position
rather than none, and `pos - last < -1` is true for every pos.

--- make the demand rung switchable

A demand rung is 47% of checkpoint writes on the cc-traces and reads back 2.8%
of the time, against 85.2% for a prompt-end anchor. Gated independently of
--state-checkpoint-interval-tokens, because the demand is not part of the
interval grid. Default unchanged. The refusal is still measured when the
placement is off -- switching off a rung must not blind the diagnostic that
justifies it.

--- DeepSeek-V4

Unaffected by the slot-vs-group change, and that claim is now checked rather
than asserted: DSV4 declares `entries_per_req=1` unconditionally (the MTP/DSpark
lookahead widens the slot via `win_with_spec`, it never multiplies the count),
so `state_slots_per_req == 1`, `pop_many(1)` pops the same index `pop()` did,
and slot == group exactly as before. `v4_pool_geometry.py` and `sub_pool_spec.py`
are untouched, so the `_physical_slots` reversal and DSV4's pool size are both
byte-identical to main. DSV4 passes no `extra_entries`, so `--state-checkpoint-
slots` is inert for it.

The *anchor*, though, did reach DSV4 -- and cost it. `PagedStateCheckpointCoord-
inator.applies()` is true for any V4 seq, so `_record_checkpoint_end` reserved a
prompt-end anchor and `checkpoint_cut` shortened a prefill chunk onto it. But the
coordinator files one pending checkpoint per seq (`_pending[id(seq)]`, and it
`del`s `boundary_blocks`), so the prompt-end checkpoint landing a chunk later
overwrote the anchor before either was stored. Measured: one extra prefill chunk
per prompt for a hit rate that did not move (identical re-send 0 blocks either
way, continuation 11 either way).

So the anchor is now gated on `keeps_interior_boundaries`, which `StateSlotPool`
answers True (each boundary is its own slot in the index, so both survive) and
the PAGE coordinator answers False. Asked as a capability rather than by naming
the backend, so a future multi-boundary copy class opts in by answering yes.
DSV4 is back to main's one-cut prefill; GDN keeps the anchor. Pinned by
`test_a_last_boundary_only_class_is_not_anchored_for`.

--- reconciliation with main

Dropped: the copy/pending-checkpoint path (`_commit_pending`, `record_copy`,
`take_copies`, `Sequence.pending_checkpoint`). DeepSeek-V4 checkpointing moved
to `PagedStateCheckpointCoordinator`, and `StateSlotPool` now rejects
`transfer.copies` outright. The fork path (GDN), where the measured wins are, is
kept in full.

`readable_midstep` is carried on main's `StateTransfer` in `state_runtime.py`
(wire format included) rather than on the branch's copy.

Two bugs this rebase exposed, both fixed here:

- `PagedStateCheckpointCoordinator` did not implement the midstep half of the
  `StateCache` protocol, so `checkpoint_cut` raised on every V4 batch. This was
  introduced *by this branch*, not latent in main: `readable_midstep` does not
  exist on main at all, and main's `checkpoint_cut` never consults it. A PAGE
  image is not readable midstep, so the three methods are the no-ops the
  protocol documents.
- `_record_checkpoint_end` could place the anchor past the last matchable block.
  `can_allocate` stops one block short of the prompt, so a checkpoint filed
  under the final block's hash is one no scan looks up -- and being stored, it
  evicted the ladder rung that would have served the resume, taking an identical
  re-request from 8 hit blocks to 0. Capped at `(n_hash_blocks - 1) * hbs`.

Tests: `tests/test_state_checkpoint.py` 171 passed; the state/checkpoint and
DSV4/LMCache suites together 632 passed. Full suite 45 failed /
2300 passed, a strict subset of origin/main's own 114 pre-existing failures --
zero regressions, verified by set difference against a clean origin/main
worktree. black clean; ruff no new findings.

Co-Authored-By: Claude <noreply@anthropic.com>
@ganyi1996ppo ganyi1996ppo mentioned this pull request Aug 27, 2026
ganyi1996ppo added a commit that referenced this pull request Aug 27, 2026
… a -1 interval

Rebased onto main's PAGE-backed checkpoint work (#1874, #1894, #1943, #1880),
which rewrote this subsystem underneath the branch. Squashed to one commit
because the three original commits each re-conflicted against the new base and
against each other's resolutions; the reasoning from all three is kept below.

--- allocate state slots per need, not by group

A "state cache group" was `1 + num_spec` slots wide and was the unit of
everything: allocation, admission, sizing, and the checkpoint index. But a
checkpoint has no speculation to roll back -- it holds a committed state -- so
filing one cost a full group and wasted `num_spec/(1 + num_spec)` of its bytes.
At two speculative tokens that is two thirds.

The slot is now the unit. `--state-checkpoint-slots 64` buys 64 checkpoints for
64 slots instead of 192, and the slots it no longer takes stay in the paged KV
pool. This is only possible because a request's slots need not be adjacent,
which the kernels never required: the ssm kernel gathers each index out of the
indices tensor and the conv path is handed column 0 alone. Contiguity was
manufactured by `prepare_state_indices` writing `arange(base, base + width)`;
it now writes the seq's own slot list straight in.

`StateGroupPool` -> `StateSlotPool`, `Sequence.per_req_cache_group` ->
`state_slots`, whose element 0 is the committed state. The setter re-points [0]
and preserves [1:], because speculation scratch persists across forwards.
`--state-checkpoint-groups` still parses, as an alias.

--- -1 turns off the interval ladder without turning off checkpointing

The interval is a guess about where reuse will resume; a demand rung is a
position a request was actually refused at. On the SemiAnalysis cc-traces the
8192 ladder placed ~30x the writes of the demand rung alone and caught reuse the
demand already reaches -- 0.0% of resumes landed on a ladder rung -- while every
rung costs the prompt that keeps it an extra prefill chunk.

  >0  a rung every N tokens (unchanged, still the default)
   0  state checkpointing off entirely (unchanged)
  -1  no interval rungs; the demand rung and prompt-end anchor still place them

-1 rather than reusing 0 because 0 is the documented contract and is reachable
by accident: the grid snap rounds an off-grid interval down and can land on 0,
so a --block-size typo currently fails safe. Three of the four sites are not the
arithmetic you would guess -- `pos % interval` under -1 admits *every* position
rather than none, and `pos - last < -1` is true for every pos.

--- make the demand rung switchable

A demand rung is 47% of checkpoint writes on the cc-traces and reads back 2.8%
of the time, against 85.2% for a prompt-end anchor. Gated independently of
--state-checkpoint-interval-tokens, because the demand is not part of the
interval grid. Default unchanged. The refusal is still measured when the
placement is off -- switching off a rung must not blind the diagnostic that
justifies it.

--- DeepSeek-V4

Unaffected by the slot-vs-group change, and that claim is now checked rather
than asserted: DSV4 declares `entries_per_req=1` unconditionally (the MTP/DSpark
lookahead widens the slot via `win_with_spec`, it never multiplies the count),
so `state_slots_per_req == 1`, `pop_many(1)` pops the same index `pop()` did,
and slot == group exactly as before. `v4_pool_geometry.py` and `sub_pool_spec.py`
are untouched, so the `_physical_slots` reversal and DSV4's pool size are both
byte-identical to main. DSV4 passes no `extra_entries`, so `--state-checkpoint-
slots` is inert for it.

The *anchor*, though, did reach DSV4 -- and cost it. `PagedStateCheckpointCoord-
inator.applies()` is true for any V4 seq, so `_record_checkpoint_end` reserved a
prompt-end anchor and `checkpoint_cut` shortened a prefill chunk onto it. But the
coordinator files one pending checkpoint per seq (`_pending[id(seq)]`, and it
`del`s `boundary_blocks`), so the prompt-end checkpoint landing a chunk later
overwrote the anchor before either was stored. Measured: one extra prefill chunk
per prompt for a hit rate that did not move (identical re-send 0 blocks either
way, continuation 11 either way).

So the anchor is now gated on `keeps_interior_boundaries`, which `StateSlotPool`
answers True (each boundary is its own slot in the index, so both survive) and
the PAGE coordinator answers False. Asked as a capability rather than by naming
the backend, so a future multi-boundary copy class opts in by answering yes.
DSV4 is back to main's one-cut prefill; GDN keeps the anchor. Pinned by
`test_a_last_boundary_only_class_is_not_anchored_for`.

--- reconciliation with main

Dropped: the copy/pending-checkpoint path (`_commit_pending`, `record_copy`,
`take_copies`, `Sequence.pending_checkpoint`). DeepSeek-V4 checkpointing moved
to `PagedStateCheckpointCoordinator`, and `StateSlotPool` now rejects
`transfer.copies` outright. The fork path (GDN), where the measured wins are, is
kept in full.

`readable_midstep` is carried on main's `StateTransfer` in `state_runtime.py`
(wire format included) rather than on the branch's copy.

Two bugs this rebase exposed, both fixed here:

- `PagedStateCheckpointCoordinator` did not implement the midstep half of the
  `StateCache` protocol, so `checkpoint_cut` raised on every V4 batch. This was
  introduced *by this branch*, not latent in main: `readable_midstep` does not
  exist on main at all, and main's `checkpoint_cut` never consults it. A PAGE
  image is not readable midstep, so the three methods are the no-ops the
  protocol documents.
- `_record_checkpoint_end` could place the anchor past the last matchable block.
  `can_allocate` stops one block short of the prompt, so a checkpoint filed
  under the final block's hash is one no scan looks up -- and being stored, it
  evicted the ladder rung that would have served the resume, taking an identical
  re-request from 8 hit blocks to 0. Capped at `(n_hash_blocks - 1) * hbs`.

Tests: `tests/test_state_checkpoint.py` 171 passed; the state/checkpoint and
DSV4/LMCache suites together 632 passed. Full suite 45 failed /
2300 passed, a strict subset of origin/main's own 114 pre-existing failures --
zero regressions, verified by set difference against a clean origin/main
worktree. black clean; ruff no new findings.

Co-Authored-By: Claude <noreply@anthropic.com>
valarLip pushed a commit that referenced this pull request Aug 28, 2026
* feat(state-cache): per-slot allocation, a switchable demand rung, and a -1 interval

Rebased onto main's PAGE-backed checkpoint work (#1874, #1894, #1943, #1880),
which rewrote this subsystem underneath the branch. Squashed to one commit
because the three original commits each re-conflicted against the new base and
against each other's resolutions; the reasoning from all three is kept below.

--- allocate state slots per need, not by group

A "state cache group" was `1 + num_spec` slots wide and was the unit of
everything: allocation, admission, sizing, and the checkpoint index. But a
checkpoint has no speculation to roll back -- it holds a committed state -- so
filing one cost a full group and wasted `num_spec/(1 + num_spec)` of its bytes.
At two speculative tokens that is two thirds.

The slot is now the unit. `--state-checkpoint-slots 64` buys 64 checkpoints for
64 slots instead of 192, and the slots it no longer takes stay in the paged KV
pool. This is only possible because a request's slots need not be adjacent,
which the kernels never required: the ssm kernel gathers each index out of the
indices tensor and the conv path is handed column 0 alone. Contiguity was
manufactured by `prepare_state_indices` writing `arange(base, base + width)`;
it now writes the seq's own slot list straight in.

`StateGroupPool` -> `StateSlotPool`, `Sequence.per_req_cache_group` ->
`state_slots`, whose element 0 is the committed state. The setter re-points [0]
and preserves [1:], because speculation scratch persists across forwards.
`--state-checkpoint-groups` still parses, as an alias.

--- -1 turns off the interval ladder without turning off checkpointing

The interval is a guess about where reuse will resume; a demand rung is a
position a request was actually refused at. On the SemiAnalysis cc-traces the
8192 ladder placed ~30x the writes of the demand rung alone and caught reuse the
demand already reaches -- 0.0% of resumes landed on a ladder rung -- while every
rung costs the prompt that keeps it an extra prefill chunk.

  >0  a rung every N tokens (unchanged, still the default)
   0  state checkpointing off entirely (unchanged)
  -1  no interval rungs; the demand rung and prompt-end anchor still place them

-1 rather than reusing 0 because 0 is the documented contract and is reachable
by accident: the grid snap rounds an off-grid interval down and can land on 0,
so a --block-size typo currently fails safe. Three of the four sites are not the
arithmetic you would guess -- `pos % interval` under -1 admits *every* position
rather than none, and `pos - last < -1` is true for every pos.

--- make the demand rung switchable

A demand rung is 47% of checkpoint writes on the cc-traces and reads back 2.8%
of the time, against 85.2% for a prompt-end anchor. Gated independently of
--state-checkpoint-interval-tokens, because the demand is not part of the
interval grid. Default unchanged. The refusal is still measured when the
placement is off -- switching off a rung must not blind the diagnostic that
justifies it.

--- DeepSeek-V4

Unaffected by the slot-vs-group change, and that claim is now checked rather
than asserted: DSV4 declares `entries_per_req=1` unconditionally (the MTP/DSpark
lookahead widens the slot via `win_with_spec`, it never multiplies the count),
so `state_slots_per_req == 1`, `pop_many(1)` pops the same index `pop()` did,
and slot == group exactly as before. `v4_pool_geometry.py` and `sub_pool_spec.py`
are untouched, so the `_physical_slots` reversal and DSV4's pool size are both
byte-identical to main. DSV4 passes no `extra_entries`, so `--state-checkpoint-
slots` is inert for it.

The *anchor*, though, did reach DSV4 -- and cost it. `PagedStateCheckpointCoord-
inator.applies()` is true for any V4 seq, so `_record_checkpoint_end` reserved a
prompt-end anchor and `checkpoint_cut` shortened a prefill chunk onto it. But the
coordinator files one pending checkpoint per seq (`_pending[id(seq)]`, and it
`del`s `boundary_blocks`), so the prompt-end checkpoint landing a chunk later
overwrote the anchor before either was stored. Measured: one extra prefill chunk
per prompt for a hit rate that did not move (identical re-send 0 blocks either
way, continuation 11 either way).

So the anchor is now gated on `keeps_interior_boundaries`, which `StateSlotPool`
answers True (each boundary is its own slot in the index, so both survive) and
the PAGE coordinator answers False. Asked as a capability rather than by naming
the backend, so a future multi-boundary copy class opts in by answering yes.
DSV4 is back to main's one-cut prefill; GDN keeps the anchor. Pinned by
`test_a_last_boundary_only_class_is_not_anchored_for`.

--- reconciliation with main

Dropped: the copy/pending-checkpoint path (`_commit_pending`, `record_copy`,
`take_copies`, `Sequence.pending_checkpoint`). DeepSeek-V4 checkpointing moved
to `PagedStateCheckpointCoordinator`, and `StateSlotPool` now rejects
`transfer.copies` outright. The fork path (GDN), where the measured wins are, is
kept in full.

`readable_midstep` is carried on main's `StateTransfer` in `state_runtime.py`
(wire format included) rather than on the branch's copy.

Two bugs this rebase exposed, both fixed here:

- `PagedStateCheckpointCoordinator` did not implement the midstep half of the
  `StateCache` protocol, so `checkpoint_cut` raised on every V4 batch. This was
  introduced *by this branch*, not latent in main: `readable_midstep` does not
  exist on main at all, and main's `checkpoint_cut` never consults it. A PAGE
  image is not readable midstep, so the three methods are the no-ops the
  protocol documents.
- `_record_checkpoint_end` could place the anchor past the last matchable block.
  `can_allocate` stops one block short of the prompt, so a checkpoint filed
  under the final block's hash is one no scan looks up -- and being stored, it
  evicted the ladder rung that would have served the resume, taking an identical
  re-request from 8 hit blocks to 0. Capped at `(n_hash_blocks - 1) * hbs`.

Tests: `tests/test_state_checkpoint.py` 171 passed; the state/checkpoint and
DSV4/LMCache suites together 632 passed. Full suite 45 failed /
2300 passed, a strict subset of origin/main's own 114 pre-existing failures --
zero regressions, verified by set difference against a clean origin/main
worktree. black clean; ruff no new findings.

Co-Authored-By: Claude <noreply@anthropic.com>

* remove gpu unit test

Signed-off-by: ganyi <ygan@amd.com>

* remove the triton import part

Signed-off-by: ganyi <ygan@amd.com>

* match main's array('i') token_ids contract in a relocation test

`#1990` added an assertion that `Block.token_ids` is an `array('i')`, not a
list -- a list never compares equal to what the production publish paths
store, so every hit on the block would read as a hash collision. This test
was written before that landed and still passed a bare list.

The file already has `toks()` for exactly this; the test just did not use it.

* feat(state-cache): keep Kimi-K3's KDA checkpoints as PAGE images

A KDA Active Slot is 53.6 MiB. Held as a checkpoint it competed with live
requests for the pool that admits them, so retaining one cost the workload the
concurrency it was retained for. Held as PAGE units it is 127 ordinary KV
blocks -- 0.112% of the paged pool -- drawn from the same free list as
everything else and evicted by the same LRU.

This is the mechanism `main` already ships and DeepSeek-V4 already uses
(`PagedStateCheckpointCoordinator`). Nothing about the coordinator changes;
what is added is the source side of the copy for a state that is two strided
tensors rather than one contiguous slab.

`plan_segmented_copy` intersects two ordered byte streams and needs neither
block alignment nor equal segments, so the state tensors keep their layout: a
slot is 138 ranges (69 conv + 69 ssm) and the planner cuts them against 127
units. `_checkpoint_layer_ranges` is the sole owner of that order -- both the
sizes and the addresses read it, because a plan cut against one order and
addressed through another lands whole layers in the wrong unit.

Two things the port had to get right, both now asserted rather than assumed:

- A PAGE unit is a *logical* block, but `kv_cache` is shaped in physical ones
  and K3's `block_ratio` is 128. `_page_unit_regions` derives its stride from
  `runner.block_size` and checks `num_rows * region == page_unit_bytes`, so a
  granularity mix-up is a startup error instead of 127 blocks of scrambled
  state. Unit ids are range-checked against the logical count for the same
  reason.
- K3's slots are strided by `num_slots`, so an off-by-one in
  `(layer * num_slots + slot)` lands inside a neighbouring request's live state
  rather than off the end of the tensor. V4 cannot fail this way and its tests
  do not look for it; `test_no_bystander_slot_is_touched` does.

`state_spec` now asks for no spare checkpoint slots under PAGE.
`--state-checkpoint-slots` buys Active Slots for checkpoints to sit in, which a
copy does not need -- 1.7 GiB reserved for nothing, and it is the same memory
the paged pool wants in order to absorb the images.

Both fall back to `fork` under pipeline parallelism and RapidServe, where
`get_num_blocks` raises on a copying transfer: answering `copy` there would
turn "K3 keeps no state cache" into "K3 does not start".

The dtype objection in the old `state_transfer` docstring is retired, not
ignored. It was that `_state_dtypes` gives kimi_linear an fp32 v side while the
chunked states are bf16, so a checkpoint cut from the kernel's `h` would hand
cached requests a rounded state. A PAGE image is copied out of the slot and
back into a slot -- both fp32, no kernel output in between, no conversion
anywhere. Both dtypes are named in the layout id, so a build that changed
either cannot read another's images.

Not yet flipped on in anger: `execute_paged_state_copies` is reachable only
from `build()`, and the GPU verification (probe at conc 1 and 8 against the
known-good 0/1 and 0/8, then GSM8K, then a matched-N hit-rate A/B) is the next
step.

Known follow-up, measured before it is fixed: the coordinator keeps one
checkpoint per sequence (`_pending` is last-writer-wins), where the fork path
indexed every boundary. A 24k prompt files at 8192/16384/24576 today and would
keep only the last. If the A/B shows the drop, the lever is to make `_pending`
hold a list -- deliberately not bundled here, because a mechanism swap plus a
policy change is a regression nobody can attribute.

* feat(state-cache): keep every boundary a PAGE seq reaches, not just its last

`_pending` was keyed by sequence, so a prompt's second checkpoint overwrote its
first before either was stored. That made the prompt-end anchor worthless --
it sits under a block from the prompt's end, lands in the same or the adjacent
prefill chunk, and was reliably the loser. `_record_checkpoint_end` reads
`keeps_interior_boundaries` and duly declined to reserve one.

Keyed by `(sequence, prefix hash)` both survive. Reaching the *same* hash twice
still collapses, which is what the hash in the key is for: that is one boundary
reached again, not two boundaries.

This matters because the anchor is the placement that pays. The measurement is
already in `_record_checkpoint_end`'s docstring: of 4,808 cc-trace resumes with
a nonzero KV hit, 93.5% land on a previous prompt end and 0.0% on the 8192
ladder. The ladder was cutting a prefill chunk every 8192 tokens to store
something nothing ever resumed from -- and on this workload a prompt averages
117k tokens, so that is ~14 rungs per request, each one a shortened forward and
an image in the paged pool.

What makes keeping both affordable is the price a PAGE image pays: 127 blocks,
0.112% of the paged pool, against a whole 53.6 MiB Active Slot under `fork`.
The measured run that preceded this kept 1,508 checkpoints with
`checkpoints_evicted: 0` -- capacity was never the binding constraint.

Run with `--state-checkpoint-interval-tokens -1` to drop the ladder entirely
and leave the anchor and the demand as the only two placements.

Three tests changed rather than deleted, because each pinned the old behaviour
deliberately and each now pins its replacement:

- `test_latest_pending_checkpoint_replaces_the_previous_intent` becomes
  `test_two_boundaries_of_one_seq_are_both_stored`, plus a new sibling for the
  same-hash-twice case.
- `test_a_last_boundary_only_class_is_not_anchored_for` becomes
  `test_both_classes_are_anchored_for`.
- Two demand tests rested a tightened pool on "exactly one image"; a prompt now
  stores two, so they spend down to the deepest -- which is both the resume
  target and what `_next_victim` would keep longest.

Not yet measured. The preceding PAGE run at conc 8 reached 91.79% at N=791
against a 96.9% trace ceiling; this is the change aimed at that gap, and the
A/B is the next step.

* refactor(state-cache): drop the parts of the PR nothing reads

Three removals, none of which change behaviour. Verified against the same
4610-passed baseline, and `ruff` on the touched files goes 16 -> 14 findings.

`cache_pressure.py` had no importer anywhere in the tree, and the log field
its regex parses (`Cached/Total:`) was renamed to `Cached/Reusable:` by this
same PR -- so it could not have matched a line this branch produces.

`keeps_interior_boundaries` was a capability hook with one reader and no
implementor that answered `False`: the `getattr` default was `True`, both
classes set `True`, and the case it existed for -- the PAGE coordinator
overwriting its own anchor -- was fixed earlier in this branch by re-keying
`_pending` on `(seq, hash)`. The measurement that justified it (of 4,808
cc-trace resumes with a nonzero KV hit, 93.5% land on a previous prompt end,
0.0% on the 8192 ladder) moves onto `checkpoint`, which is where the key it
argues for lives.

`_log_frequency` and its four `reqs_*` counters cost four `__slots__` entries
and four per-request branches to render one log line, and are read by nothing
else -- not `metrics.py`, not either aggregation tuple in `llm_engine.py`.
`_log_pools` stays: its three rates are pure ratios of totals already kept,
and the paged/state split is this PR's central claim. `_log_pressure` stays
because `checkpoints_*` and `demands_recorded` do reach Prometheus.

Left alone deliberately: the `record_relocation` / `take_relocations` /
`relocate_state_slots` chain is equally unreachable, but it is that way on
`main` too. Deleting main's debt from this branch would widen the diff it is
meant to narrow.

* docs(state-cache): tighten the comments this PR added

No code changes; 375 tests pass and `ruff` on the touched files stays at 14
findings against main's 16.

The bf16/fp32 accuracy argument was written out in full three times --
`GDNStateMixin.state_transfer`, `pop_last_intermediate_states`, and inverted
again in `_KimiMLAGDNCommon.state_transfer` -- each time as a rebuttal to an
objection nobody raised, and two of the three cited
`tests/test_gdn_state_checkpoint_gpu.py`, deleted in a04ce7f. It now lives
once, in the present tense, where the dtypes are chosen; the other two point
at it. That alone is ~30 lines and both dead citations.

Two measurements had spread to four and five sites. The prompt-end anchor's
read-back rate stays in `_record_checkpoint_end`, which exists because of it;
the demand rung's stays in `mark_speculative`, the only place it decides
behaviour, and in the `--state-checkpoint-demand` help text, where a CLI user
cannot follow a code reference. `config.py`, `envs.py`, `sequence.py`,
`page_unit_checkpoint.py` and `checkpointers_at` now reference rather than
restate, so there is one copy to update when the number moves.

The rest is history that git already holds: what an earlier Python-loop
version got wrong, what the upstream branch does with `state_cache_base`,
what this pool "used to allocate", which objection "kept this on fork". Each
is restated as the invariant it was arguing for. Also two stragglers of the
group->slot rename in `attention_gdn.py`, and a call-site comment in
`gdn_attn.py` that restated `_checkpoint_targets`' own docstring.

Left long on purpose: `_page_unit_regions`' logical-vs-physical block-id trap
(K3's block_ratio is 128, and getting it wrong scrambles 127 blocks silently),
`_assert_checkpoint_geometry_still_holds`, the conv-window claim in
`state_transfer`, and `CacheStats`' argument for `reusable` over `full` as the
denominator -- that last reads like a rebuttal but the objection is one a
reader will actually raise.

* docs: describe the two model-agnostic features and the instrumentation

The description covered only the K3 PAGE port, which is 1,221 of the 4,694
added lines. Three things it shipped were undocumented:

Per-slot allocation. `StateGroupPool` -> `StateSlotPool`, and a request's state
goes from one fixed-width group of `1 + num_spec` adjacent slots to a list of
ids that need not be adjacent. The point is that a checkpoint takes one slot
rather than a whole group, since a resumed prefix has no speculation to roll
back. Documents the one consumer that reads past element 0 -- the spec-decode
path, which stopped deriving the set from `base = group * slots_per_group` --
and states why DeepSeek-V4 is a rename rather than a behaviour change.

Midstep checkpoints. A mamba-like backend can now take every boundary a forward
covers out of the chunk kernel's own `h`, instead of having its prefill cut so
the forward *ends* on each one. Covers the reserve/publish/cancel split (the
bytes do not exist when the destination must be chosen), the `is_end` targets
that read the runtime slot because `h` does not hold the final state, and the
paired gate in `checkpoint_cut`/`checkpointers_at` -- suppressing one alone
keeps zero checkpoints with no error. Names Qwen3-Next and Qwen3.5 as the
models on this path and K3 as the one that cannot be, and adds the latter to
the follow-ups.

Hit-rate instrumentation. Every measurement in this PR was read off these
lines. The `[Cache Stats]` denominator was `full`, which includes the trailing
block `can_allocate` never matches -- so it charged both pools for a block
neither was offered and reported an unreachable ceiling; it is now `reusable`.
`[Cache Pools]` splits the series into `paged * state = combined`, which is
what showed the paged index matching 99.4% while the state gate discarded it.
`[Checkpoint Fates]` separates four fates that argue for different fixes, and
`kept: 1508, dropped: 0, evicted: 0` is the evidence behind the "capacity
stopped being the binding constraint" claim.

Also refreshes the numbers the rebase and the two cleanup commits invalidated:
33 files / +4694, the current commit hashes, the per-file table, and the test
baseline (4610 passed / 50 pre-existing failures, 40 of them sglang files that
score identically on origin/main).

* docs(state-cache): KDA's interior h exists; aiter just does not return it

The follow-up said K3 cannot be `readable_midstep` because
`chunk_kimi_delta_attn` "exposes only `output_final_state`". True of the API,
misleading about the cause: in aiter's
`_triton_kernels/chunk_delta_attn/chunk_fwd.py` the per-chunk `h` is computed
at line 170 -- by `chunk_gated_delta_rule_fwd_h`, the same function the GDN
path uses -- consumed by `chunk_gla_fwd_o`, then set to None at line 202 and
left out of the returned tuple.

So the two backends differ in plumbing, not in what their kernels produce.
ATOM vendors GDN's chunk entry under `model_ops/fla_ops/`, which is why
`keep_intermediate_states` could be added there; KDA goes out to aiter, which
has no equivalent. Whoever picks this up is adding a return value, not an
algorithm -- worth stating, because the old wording invites the conclusion
that the kernel would have to be rewritten.

Behaviour is unchanged: K3 stays `readable_midstep = False` and keeps cutting
a chunk per placement. Under the shipped anchor-only policy that is 0% of
prompts cut at 1.00 checkpoints per request, since the anchor lands where the
last prefill chunk was going to end anyway.

* test(gdn): pin the claim readable_midstep rests on, on real hardware

`readable_midstep` asserts that `h[:, j]` is the recurrent state after
`j * 64` tokens, and `BlockManager` acts on it by suppressing `checkpoint_cut`
outright -- the prefill runs full length and the boundaries are harvested from
`h` afterwards. If that assertion is false, every checkpoint the readable path
stores is subtly wrong: a resuming request inherits a state its prefix never
produced, silently.

`TestMidstepCheckpoints` pins everything *around* the claim (which positions
are chosen, reserve/publish/cancel, that the cut is suppressed) but stubs the
kernel, so it cannot see the claim itself fail. This asks the kernel.

Measured on MI355, 8 chunks of 64: all 7 interior boundaries are **bit-exact**
against a forward stopped at that position -- `torch.equal`, not a tolerance,
which is the right bar because both arms round the same fp32 value into the
same dtype (`h` is `k.new_empty`; `_state_dtypes` returns `config.torch_dtype`).

Two smaller guards alongside it: popping consumes the reference, so a later
layer cannot read the previous one's `h` and file it under its own slot; and a
forward that was not asked to keep retains nothing, so the plugins that never
pop do not pin a large tensor past their last forward.

Needs one GPU and a few hundred MB -- no server, no TP, no weights -- and
skips at module level otherwise, following `test_compress_chunk_equivalence`.

* docs: record the midstep hardware result, and narrow what is still unmeasured

"Qwen3.5 is not measured" was true when written and is now too blunt. The claim
`readable_midstep` rests on -- that `h[:, j]` equals the state a forward
stopped at `j * 64` would leave -- has been asked of the kernel directly: 7 of 7
interior boundaries bit-exact on MI355. That belongs in Verification, because it
is the one part of the midstep path a CPU test cannot reach and a failure there
would be silent.

What remains unmeasured is narrower and worth saying precisely: no server has
been stood up on a readable backend, so there is no hit rate, accuracy, or TTFT
for it. Named the three things a single sequence through one kernel cannot show
-- the per-sequence `chunk_offsets[row]` base with two prefills in a batch, the
ordering an `is_end` target depends on, and a resume landing on a stored midstep
boundary -- so the gap is actionable rather than a blanket disclaimer.

* fix(state-cache): address review findings 1, 7, 10 and the instrumentation

Findings from @valarLip on #2045, each re-verified against the code rather
than taken on report -- two of the sixteen did not survive that check
(`_rehome_checkpoint` does not exist; `chunk_gated_delta_rule` carries
`@torch.compiler.disable`, so the CUDA-graph half of #12 cannot happen).

**#1, a regression this branch introduced.** `eb058321d` re-keyed `_pending`
to `(seq, hash)` so two boundaries of one prompt could coexist, and did not
touch the drain, which still resolves a single `seq.state_slot` for all of
them. Both images are then copied out of whatever the last forward left there,
filing the earlier hash over the later state -- a request resuming on it
continues from ahead of its own prefix, and `_validate_paged_state_op` passes
because layout, size and unit count are all still correct.

`_supersede` keeps one pending boundary per sequence. Ordinarily a drain
follows every forward and both boundaries are stored correctly from their own
slots; the exception is a pass that schedules nothing, where
`state_maintenance_ops=None` carries `_pending` into the next drain. The newer
boundary wins because it is the one the slot holds, and the older is counted
`dropped` -- it is reuse the placement asked for and did not get. The two tests
that pinned the old behaviour asserted coexistence without asserting each was
stored from its own slot, which the drain cannot do; they now pin the fix and a
sibling covers the ordinary drain-between-forwards case.

**#7** was the same invariant read from the other end: the descriptor buffer is
sized `2 * max_num_seqs` on "one store per sequence", which the re-key removed
and `_supersede` restores. No resize -- the docstrings here and in
`deepseek_v4_attn.py` now name what holds the bound instead of asserting it.

**#2/#3/#4/#5/#8 are all on the midstep write path, so `readable_midstep` goes
back to False.** The write path declines on six conditions `commit_midstep`
cannot see and publishes the hash regardless; `_checkpoint_targets` indexes
three differently scoped sequence lists with one `i`; the SSM read floors to a
64 grid `midstep_positions` does not enforce (`hash_block_size` defaults to
16); the conv window is `conv_kernel-1+num_spec` in the kernel and
`conv_kernel-1` in the guard. Each stores a findable image holding the wrong
state. None of it has run under a server -- K3 takes the PAGE path and cannot
reach it -- so the honest state is off. The machinery and its bit-exactness
test stay; `test_midstep_is_off_in_production.py` pins the decision, and is
deliberately not behind `importorskip` so the non-GPU runner actually runs it.

**#10** `pool_pressure` read `self.state`, which under PAGE is a different
object built with `StateTransfer.none()` that never sees a `checkpoint()`. It
printed four zeros for the life of the server while `checkpoint_funnel`, the
next method, reported the real numbers from the coordinator.

Smaller: `chunks_cut_for_end` and `checkpoints_orphaned` reach both aggregation
whitelists (a cut counter without its sibling is unreadable; `orphaned` argues
for a bigger paged pool where `evicted` argues for a bigger state pool);
`paged_hit` is dropped as a second name for `compressed_hit`;
`_warn_if_unschedulable` compares against `state_slots_per_req` again, so a
pool too narrow for one request warns instead of waiting forever in silence.

`clear_index` still moves no counter, now stated as a decision: each fate
argues for a different fix and an operator emptying the cache argues for none.

* fix(state-cache): address review findings 6, 9, 11, 13 and 14

**#9 inverts the policy it implements, on the majority of prompts.**
`mark_speculative` exists so a guessed resume point is spent before a known
one -- anchors are read back 85.2% of the time against a demand rung's 2.8%.
Both call sites gated on `if anchor and pos != anchor`, so a seq whose
`checkpoint_end_pos` is 0 demoted *nothing* and filed its guesses at the LRU
tail beside real anchors. `_record_checkpoint_end` leaves it at 0 on four
paths, one being every prompt too short for a keepable end -- the common shape
of an agentic first turn. `_anchor_of` answers None rather than 0 there, which
compares unequal to every position, so those seqs demote everything. Shared by
both sites because two spellings of one rule is how they drift apart.

`publish_midstep(seq=None)` still demotes nothing, now as the stated other end
of the rule: a caller with no sequence cannot tell a guess from knowledge, and
over-keeping costs one eviction where over-demoting spends an anchor.

**#6 could take the engine down over a log line.** The two asserts run for
every prefill seq, over counters with four independent writers -- the
CPU-offload wake sets `num_cached_tokens` without touching the hit-block
counters the rest derive from, so an LMCache resume that loads more prefix than
the GPU index held produces `cached > wanted` legitimately. Now a warning and a
clamp, which also removes a behaviour difference between `-O` and not.

**#11 had two consumers disagreeing about what -1 means.** `BlockManager`
clamps to `max(-1, ...)` and reads -1 as "grid off, anchor and demand still
placing"; the DSV4 offload policy clamped to `max(0, ...)`, folding it into 0,
which for that consumer means no sidecar checkpoints at all -- so the engine
kept checkpointing while offload resume silently degraded to zero reuse. With
no grid to align to the sidecar now takes `resume_alignment` alone.

**#13b/#13c.** `_extend_hash_chain` sat one line above the `_has_page_units`
refusal, so a 128k prompt queued behind a full pool paid ~2000 xxhash rounds
per waiting request per pass for a list that was then discarded. Moved below
it; verified nothing between consumes it, and that its one reader is
`midstep_positions`. The comment claiming it "reads `checkpoint_end_pos`" was
false and is replaced with what actually orders the call.

**#14.** `chunk_gated_delta_rule_fwd` returned `h` unconditionally, so the
caller's frame pinned ~33 MB for the rest of that layer's forward even with
`keep_intermediate_states=False` -- every GDN prefill with checkpointing off,
which is the default, and every vLLM/SGLang/rtpllm caller. The flag now reaches
the producer, so the reference dies with the fwd frame. No compute changes; the
kernel computed it either way. `test_gdn_midstep_state_gpu.py` covers both
values and still passes on hardware.

Suite: 4622 passed against 4618 before, with the same 50 pre-existing failures.

* docs: carry the slot rename into the guides, and document the new knobs

The rename landed in code and left five guides describing a `group` model that
no longer exists. One of them was actively dangerous: the `deallocate` snippet
in the scheduling guide released `seq.per_req_cache_group` — a single slot —
where the real function calls `release_many(seq.state_slots)`. Copied as
written it leaks `num_spec` slots per request, and admission cannot see the
loss because it gates on the free list this never returns them to.

Corrected across `scheduling_kv_cache_guide.md` (the pool construction snippet,
the allocation and deallocation prose, the pool-field list, the Sequence table,
and the fork-checkpoint capacity paragraph), plus the Sequence rows in
`architecture_guide.md` and the GDN state paragraph in
`model_support_guide.md`. `state_slots` is documented as a list with `[0]`
committed and `[1:]` rollback, explicitly not adjacent, with `state_slot` as
the property over element 0 — which is the contract a backend has to know
before it indexes anything.

Newly documented rather than merely renamed:

  * `--state-checkpoint-interval-tokens -1`. The guides described `0` as the
    only off switch, so the ladder-off-but-anchor-on regime this PR added was
    reachable and undocumented.
  * `--state-checkpoint-slots` (and its `--state-checkpoint-groups` alias),
    with the note that a PAGE backend zeroes it out.
  * `--state-checkpoint-demand` / `--no-state-checkpoint-demand`.
  * `ATOM_STATE_CHECKPOINT_DEMAND`, under a new "State checkpoints" section in
    `environment_variables.md` — it had no entry at all.

Every symbol the guides now name was checked to exist in `atom/`. The two
`*_plan.md` files still say `group`; they are dated design notes rather than
reference docs, and rewriting them would misrepresent what was planned.

* revert: drop two changes that belong to other PRs

Neither touches per-request state, checkpoints, or the pools. They rode along
on this branch and widen its review surface for no reason.

`triton_merge_attn_states.py` moves `prefill_tokens_with_context` off
`tl.constexpr`. That is a real fix — a per-batch token count as a constexpr
mints a fresh kernel per distinct batch size, 184 of them in one 8-minute
agentic run — but it is an attention-kernel compile-time bug, not a state-cache
one, and belongs in a PR that says so.

`tests/plugin/test_vllm_kimi_k3.py` moved its registry check out-of-process to
survive `sys.modules` damage other plugin tests do. Also genuine, also
unrelated; verified it does not pollute the session on its own.

`test_rtpllm_forward_context_semantics.py` is NOT reverted, though it looked
like the same category. Its change makes the stubs it installs restore what
they displaced, and without it `atom.model_ops.attention_gdn` and
`atom.utils.forward_context` stay shadowed for the rest of the session:
reverting it turned 3 collection errors into 6, taking
`test_cudagraph_capture_bounds.py` (9 passed alone) and three sglang plugin
modules down with `cannot import name ... (unknown location)`. That is
load-bearing for whether this branch's own suite can be run at all.

* remove --state-checkpoint-slots, which never took effect

The flag sized a flat cushion of spare Active Slots for checkpoints to sit in.
It defaults to 0, DeepSeek-V4 never declared it, and Kimi-K3 overrode it back
to 0 — so on every shipped path it added nothing, and the only configuration
where it did anything was GDN's `fork` with someone passing a value by hand,
which no measurement in this PR or before it covers.

What it was for is real: a checkpoint held as a slot competes with live
requests, so how many can be retained is set by concurrency rather than by how
much reuse the traffic has. The PAGE path solves that properly, by keeping the
image in KV blocks instead of a slot. A cushion would buy the same decoupling
for `fork` at the price of a knob nobody can size without measuring first.

Removing it collapses two things it had propped up. `_KimiMLAGDNCommon.state_spec`
existed only to zero the field and is deleted — with the flag gone the base
spec is already right, and `super()` needs no correction. And
`TestTheSpareSlotsGoBackToTheKvPool` went with it: it monkeypatched
`GDNStateMixin.state_spec` to a lambda returning a literal `extra_entries=32`,
so it asserted against its own stub and would have passed unchanged if
production stopped reading the field altogether. That is the shape @valarLip
flagged, and deleting the feature removes the test's subject rather than its
substitute.

`SubPoolSpec.extra_entries` stays. It is the sizing layer's general capability,
no backend passes a nonzero value today, and the test that pins its arithmetic
now says so — a future cushion should get a flat one, not `width x` what it
asked for.

4620 passed against 4622 before, the difference being the two deleted tests;
same 50 pre-existing failures. `ruff` on the touched files matches origin/main
exactly.

* test: keep the midstep watch on the CPU-only side of the aiter line

`test_gdn_does_not_declare_itself_midstep_readable` reached
`GDNStateMixin.state_transfer`, which means importing `gdn_attn`, which
imports aiter at module level. The non-GPU CI runner installs CPU torch and
neither aiter nor triton, so that is a collection-time
ModuleNotFoundError, not a skip -- it failed the job.

Split the flag's two halves by what a CPU runner can actually see. The PAGE
coordinator's `readable_midstep`, the three-call midstep protocol on
`StateCache`, and `StateTransfer`'s field are all pure Python and stay
watched here. GDN's declaration is the half that needs aiter; it belongs
with the kernel tests, and the docstring now says so rather than leaving the
next reader to rediscover it by breaking CI.

Coverage lost is one assertion, not the mechanism: `TestMidstepCheckpoints`
builds its own `StateTransfer(readable_midstep=True)` and exercises the
write path regardless of what production declares.

---------

Signed-off-by: ganyi <ygan@amd.com>
Co-authored-by: Claude <noreply@anthropic.com>
Co-authored-by: Guanbao Yu <Guanbao.Yu@amd.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants