perf(v4): a smaller PAGE checkpoint image, a 60x cheaper copy, and no reserve floor - #1943
Merged
Merged
Conversation
A PAGE-backed checkpoint copied the whole Active Slot. Most of it is dead at the boundary the checkpoint sits on. DeepSeek-V4's HCA compressor pools `ratio` tokens with no overlap, so the first pool at or after a boundary P covers `[P, P + 128)` — every row of it written by the very forward that reads it — and a checkpoint is aligned to `hash_block_size`, a multiple of 128. Its two fields are 51% of a slot and a resumer never reads a byte of them. The sliding windows sharing the slot are a sliding window and stay whole; the padding between the two halves belongs to neither. `StateField.in_checkpoint` lets a field say it is not carried, `checkpoint_ranges_for` turns the declaration into byte ranges, and the DSV4 builder composes those with the window rows. `PagedStateCheckpointSpec.image_bytes` prices the result, so on V4-Flash-DSpark a checkpoint costs 7 PAGE units instead of 14 and displaces half the KV history it used to. The rule holds only while every compress ratio divides the block size a checkpoint is aligned to. `_assert_ratios_divide_block` refuses a build where it does not, because that failure is silent — a resumer reading stale KV for its first pool, which costs a little accuracy and nothing else. Two workers disagreeing about the rule would read one image at two layouts, so `layout_id` names what is dropped and takes a new version. The offset walk that `entry_bytes_for`, the arena's field offsets and the new ranges all need is now written once, in `field_extents`; the three have to agree and previously each carried its own copy. Gates on V4-Flash-DSpark tp2, bf16 KV, DSpark-5, interval 256. Arms were alternated one fresh server each: this box has a per-process spread that sequential arms read as a regression, twice. units_per_checkpoint 14 -> 7, image 10,288,128 B = 48.6% of a slot #1417 coherence 0/3 collapsed throughput n=3 39,000 vs 37,538 tok/s acceptance n=3 42.93% vs 41.97%, overlapping GSM8K flexible n=3 0.9510 vs 0.9530, overlapping, both inside the 0.9522 +/- 0.0059 band Copy cost is descriptor-bound, not bandwidth-bound: halving the bytes left the measured 2.7 ms per op (0.87 plan + 1.85 launch, 135 spans) unchanged, which is what caps the gain at 3.9%.
A PAGE-backed store called `ensure_free_units(units_per_checkpoint)` and took whatever it got. Its units are then unreclaimable until the next `complete_inflight` publishes the record: `ensure_free_units` only evicts READY ones, and an in-flight store is COPYING. A burst of requests crossing a rung together therefore drains the free list, and the raise lands not on the checkpoint — which is best-effort and would have been happy to be dropped — but on whichever live request calls `_fresh_block` next, where it is an AssertionError with no path back. A store now has to leave `reserve_units` behind and returns None when it cannot, which routes it into `checkpoints_dropped`, the same answer the pool already gives when nothing can be evicted. The floor is one batch's worth of new blocks — a chunk of prefill plus at most one append per running sequence — because one batch is exactly how long a store's units stay unreclaimable. On V4-Flash-DSpark that derives to 320 of 103,809 units, 0.31% of the pool, and it is logged at startup: the number comes from the batch shape rather than a flag, so a zero would quietly mean there is no floor at all. Not a quota on how much checkpoints may hold in total. They can already be reclaimed on demand — `_ensure_page_units` evicts READY records for live KV — so the steady state is self-correcting and only the in-flight window needed protecting. Verified on V4-Flash-DSpark tp2 bf16 DSpark-5, interval 256: the floor changes nothing at this scale, `checkpoints_kept` 81 / `dropped` 0 / `evicted` 0 / hit 83.2%, all identical to the run without it.
Copying a checkpoint cost 0.92 ms per op against 0.018 ms of kernel. The other 98% was describing the copy: 22 throwaway tensor views per PAGE unit to learn addresses, then every span expanded to 4 KiB tiles on the host and shipped as three Python lists through three pageable transfers — 2,632 tiles where there were 135 spans. Both go. `_page_unit_regions` works out `(base, num_bytes)` per region once. Blocks sit back to back in every pool, so a block's address is `base + id * bytes` and slicing the tensors to find it was buying one multiplication with a view. `tensor_segment`'s contiguity check comes along, asked once of the layout instead of every time of a slice. `launch_copy_spans` uploads one descriptor row per span and lets the grid's second axis cut the tiles. That axis has to be as tall as the widest span, so where spans differ most programs find nothing to do — 94% of them here, whose spans run from 8 KiB to 1.4 MB. Measured before being believed: an empty program is cheap enough that the trade is not close, and the kernel is unchanged at 0.018 ms. per op, DSV4 narrowed image before after region addresses 0.260 0.030 plan_segmented_copy 0.061 0.061 launch_copy_spans 0.600 0.040 total 0.92 0.13 7.1x End to end, arms alternated three fresh servers each: TTFT 202 -> 172 ms, 3/3 rounds lower, ranges disjoint. Throughput +2.0% and acceptance -0.4 pp both overlap and are not claimed — this workload runs OSL 64, so its total is prefill-bound and a decode-side instance spread of 7% sits on top. Correctness is pinned by an oracle rather than by inspection: the addresses are asserted equal to what the tensor slicing produced, and the new kernel's bytes equal to what the tile kernel wrote. Both were checked to fail — swapping the layer and block strides, shortening a region, filling the descriptor's destination column from its source, and dropping a byte per span each turn a test red.
…reads
Two changes to the PAGE-backed checkpoint path: describing a copy no longer
costs per span, and an image no longer carries the entry's interleave
padding. The first is what makes the second free.
Describing a copy cost about 0.53 us a span -- one Python loop to intersect
the slot's ranges with the image's PAGE regions, another to read three
fields off each span into the descriptor. On the DeepSeek-V4 image that is
0.107 ms an op against 0.018 ms of actually copying, and it is what decided
whether a finer image was affordable: every byte saved cost host time to
describe.
None of it has to be paid per op. Which source segment meets which
destination segment, at what offset into each and for how many bytes,
follows from the two streams' *sizes*; addresses enter only when a copy is
issued. Both streams are geometry -- the slot's ranges come from the
layout, and every image is `units_per_checkpoint` units of identical region
sizes, which `_validate_paged_state_op` already insisted on. So the
intersection is walked once for the life of the pool and an op becomes two
gathers and two adds over precomputed offset arrays.
`plan_segmented_copy` now takes sizes and returns a `SegmentedCopyPlan`:
five parallel int64 arrays, no addresses. `write_descriptor` fills a
`(spans, 3)` block from one base per segment of each stream, which is where
the caller's geometry enters. A store and a restore are the same
intersection read opposite ways, so `forward=False` reuses the plan rather
than cutting a second one, and every op of a batch shares one descriptor
and one launch. `ByteSegment`, `CopySpan`, `tensor_segment` and
`launch_copy_spans` go with it; V4's three segment builders collapse into
`_checkpoint_slot_bases` (a `[group, segment]` matrix built once, so the
per-op source side is a row lookup) and `_page_unit_bases` (one outer
product).
With a span costing nothing to describe, the image can drop the rows the
row space only holds so that one index formula can serve every layer of a
compress class. The interleave runs by ring *position*, not by layer: rows
`[c*run_rows, (c+1)*run_rows)` hold every layer's positions for run `c`, so
the rows the construction skips are reachable by no `(layer, position)`
pair at all -- nothing writes or reads them, and a checkpoint image, which
is only ever gathered back into a slot and never read by an attention
kernel, does not owe them. `ClassLayout.entry_row_runs` enumerates what is
reachable; on the DSpark configuration the rest is 17.3% of the entry.
This keeps every ring position, so it needs no phase input and the range
list is still computed once. `layout_id` goes to v3 because an image is no
longer a subsequence of the slot's rows -- a v2 reader would gather every
window row shifted.
image 10,288,128 -> 9,060,352 B (48.6% -> 42.8% of a slot)
PAGE units 7 -> 6
per op 0.107 -> 0.044 ms
Gates on DeepSeek-V4-Flash-DSpark bf16 tp2, `--num-speculative-tokens 5`,
`--state-checkpoint-interval-tokens 256`:
- 1483 unit tests, including `entry_row_runs` checked against
`ring_offset_for` over the whole (layers, stride, ring_slots) product
and against the real Flash geometry's `ring_row`.
- The descriptor is asserted equal row for row to the walk it replaces,
on four image shapes, before any timing.
- #1417 coherence probe 0/27 collapsed, hits at 256 and 512 -- the
resumer reading window rows only the gather can have given it. A
control arm with packing disabled and nothing else changed: also 0/27.
- Acceptance rate, the probe this change would show up in first, on a
workload with 83% prefix hits: 42.28 packed against 42.28 unpacked
over four alternating instances per arm. On GSM8K with a
counterbalanced order: 64.460 against 64.465, and the same 0.96125
score.
- GSM8K 1319-question, three runs: 0.9484 / 0.9553 / 0.9484, inside the
0.9522 +/- 0.0059 band this configuration has held all along.
Throughput came out 3.2% lower in the packed arm (p = 0.29, arm ranges
overlapping, and the two position-matched pairs disagree in sign). Reading
it as an effect would need many more samples than it is worth: there is no
mechanism -- the image is smaller, takes fewer PAGE units, and costs
0.044 ms an op against 0.042 -- and the metric that would see a real
regression first shows nothing.
The floor added for live KV was handed to `ensure_free_units` as part of the count it must reach. That reads like a request and behaves like a demand: `ensure_free_units` gives up only after it has evicted every READY checkpoint it can, so whenever live KV held the rest of the pool a single `begin_store` emptied the entire cache and still returned `None` -- and did it again on the next batch, and the next. Reproduced on the real classes: a 100-unit pool with 50 READY checkpoints and a drained free list loses all 50 to one dropped store. The units it freed did go to live KV, but `_fresh_block` already takes those on demand, one at a time, at the moment they are actually needed. Nothing was gained for the cost of the cache. Ask whether the floor is reachable before asking to reach it: `reclaimable_units` is the ceiling eviction could raise the free list to, so a store that cannot clear the floor now refuses without evicting anything. Recycling is unaffected -- a store still takes the oldest checkpoint's units when the floor allows it. The victim predicate `ensure_free_units` walked inline becomes `_evictable`, so the two cannot come to disagree about what is evictable. The rationale for the floor was also wrong, which is what made its size look indefensible. It said the hazard was a burst of stores holding `COPYING` units that live KV cannot reclaim. That window does not exist on the normal path: `schedule` calls `complete_previous_state_batch` before it allocates anything, so an allocating `_fresh_block` always sees the previous batch's stores as READY. The reachable hazard is the other unevictable state. `BlockManager.allocate` pins a restore and then asks for fresh blocks *in the same pass*, and the pin holds until the next `complete_previous_state_batch` -- so one pass of prefix hits can pin every checkpoint it resumes from and then find nothing left to evict. A floor of one pass's worth of new blocks (`ceil(max_num_batched_tokens / block_size)` + `max_num_seqs`) is exactly what keeps that pass from ever having to evict, which makes how much of the cache it pinned stop mattering. It is sized against the pass, not against the unevictable set. `test_the_floor_survives_a_pass_that_pins_the_whole_cache` pins that invariant, which had no coverage at all; it fails with the floor set to zero. Also: the derived-reserve log no longer fires when the coordinator ends up disabled, where it reported a floor nothing would ever consult; and the four address caches on the V4 builder carry the constraint that makes them safe, since a reallocating pool would turn a stale one into a copy to the wrong address rather than a crash.
…fuses A demand is an instruction to cut a prefill chunk onto a rung, and that cut costs the request a forward -- the same forward the interval grid exists to amortize, and the one this guide measures at 17.5% of throughput when spent unconditionally. `begin_store` then drops the checkpoint whenever taking it would leave the pool under the floor, so under pool pressure the ladder was buying forwards for stores that were already going to be refused. The ladder now asks the question the store will ask, through the same expression: `has_room_for_store` is what `begin_store` refuses on, so the two cannot come to disagree about what is affordable. The gate goes on `_record_checkpoint_demand` rather than on either reader, because `checkpoint_cut` and `checkpointers_at` have to agree position for position and both read the one field -- gating the field keeps that free. What is deliberately NOT suppressed is the attribution. `num_wanted_hit_blocks`, and hence `Lost-to-checkpoint`, still say the reuse was declined for want of a checkpoint, because it was; only the instruction to act on it is withheld. `demands_declined_no_room` joins the funnel so the difference is visible, which keeps "the ladder is quiet because there is no demand" apart from "the ladder is quiet because the pool is tight" -- the whole reason that funnel is assembled stage by stage. The fork path is untouched: a fork checkpoint costs the paged pool nothing, so `_checkpoint_has_room` is trivially true where there is no PAGE-backed coordinator at all. Also corrects the comment on `reserve_units` itself, which still described the floor as protecting in-flight `COPYING` units. That window does not exist on the normal path -- `schedule` publishes the previous batch's stores before it allocates anything. The reachable one is a pass that pins every checkpoint it resumes from, which is what the floor is sized against. `test_a_demand_the_floor_would_refuse_is_not_recorded` fails with the gate short-circuited, and asserts the attribution survives it.
…against it Two halves, both of them subtraction. ## The copy path `launch_copy_descriptor` opened a rectangular grid, `(spans, ceil(widest / TILE))`, which gives every span as many programs as the *widest* one needs. On a DeepSeek-V4 image, whose spans run 8 KiB to 1.4 MB, that is 46,364 programs to do 2,631 tiles of work. One op could afford the waste; `execute_paged_state _copies` batches every op of a step into one launch, and a batch cannot. The grid is now one program per tile that exists, from a tiling that is a pure function of the plan's geometry -- computed once when the plan is cut, resident on the device from first use, shared by every op. That also retires `widest`, a parameter whose only failure mode was silent: pass one too small and every longer span was truncated, byte-correct on its prefix and stale on its tail. The descriptor is then built for the whole batch in one pass instead of one copy at a time. `write_descriptor` takes `(copies, segments)` base arrays and `_page_unit_bases` grew an image axis to match; store and restore are batched apart because they read the same intersection in opposite directions. At these sizes a numpy call is nearly all call overhead -- a span table is a few hundred entries -- so paying it per copy was what made describing a batch a quarter of the path once the kernel stopped being the bottleneck. Finally the kernel runs on four warps rather than eight. Its speed turns out to be set by the width of one lane's access, `TILE / (num_warps * 64)`: three unrelated (TILE, warps) pairs that land on sixteen bytes measured within 0.5% of each other, while eight warps halves the width and costs 12%. This is recorded in a comment because it reads like a knob to turn up, and turning it up makes it slower. At 256 ops the path measures 1.29 ms against 15 ms, and the kernel is 88% of what is left. Every step was checked against the previous implementation as an oracle, byte for byte, before it was timed. ## The reserve `reserve_units` is deleted rather than resized. It had two stated purposes and neither survived being measured. The first was to keep `_fresh_block` from raising. It cannot be reached: a READY unpinned checkpoint is *already* available to live KV, since `has_available_units` counts it and `ensure_free_units` will spend it, so the size of the cache is not the variable. What competes is the unevictable set -- `COPYING`, or held by a restore pin -- and that set is confined to one pass. `schedule` publishes the previous batch's stores and releases its pins before it allocates anything, and this batch's stores are taken at batch construction, after every allocation. The one overlap is `allocate`, which pins a restore and then asks for fresh blocks in the same pass -- and its own `can_allocate` counted that pin. `may_append` never overlaps at all, because the decode loop runs only in a pass that scheduled no prefill. Driven through the real `BlockManager` across a grid of pool shapes, thirty-eight gated runs never reached the raise while bypassing the gate reached it in fifteen of nineteen. Under contention the reachable outcome is a refused admission, which the next pass retries. The second was to make that gate never refuse -- "sized so the set is never consulted". It cannot do that either: the floor is a chunk's worth of blocks, `ceil(max_num_batched_tokens / block_size)`, while `allocate` takes a whole prompt's block table, up to `max_model_len` of them. A single 128K admission asks for more than the entire floor. No reserved quantity can promise live KV a block, because live KV's demand is unbounded and legitimate, and that is now said in the code so the next reader stops looking for one. What the floor did do was evict. `begin_store` asked `ensure_free_units` for `needed + reserve`, so one accepted store spent up to fifty-five checkpoints building a cushion for a hazard that is not there. It now asks for `needed`: free units first, and at most one image's worth of eviction for the shortfall. Two silent failure modes go with it -- a reserve larger than the pool would have disabled checkpoints permanently behind a healthy-looking startup line, and the floor's prefill term used `block_size` where DCP wants `hash_block_size`. Eviction eligibility and eviction policy are separated on the way past. `_is_evictable` says whether a checkpoint may be spent and is the single rule `has_available_units` and `ensure_free_units` share; `_next_victim` says which one to spend first and is the only place the policy lives, least recently used today. `has_available_units` stops at the shortfall rather than totalling the cache, which is both the faster answer and the reason a future policy cannot move the gate: the eligible set decides it, not the order it is walked in. That walk is per-sequence per-pass, and a warm pool holds `num_kvcache_blocks / units_per_checkpoint` checkpoints -- ten thousand of them here, previously summed in full for an answer one or two settle. ## Also `_assert_ratios_divide_block` is now `_assert_ratios_divide_the_alignment`, because it stopped asking about `block_size` when it started asking about `kv_cache_block_size * decode_context_parallel_size`. `_invalidate_pool_caches` gives the five address caches on the V4 builder one place to be dropped from, which whoever wires an elastic pool has to call -- a stale one is a copy to the wrong slot, not a crash. The vLLM bridge's HCA fields carry `in_checkpoint =False` like the native list, since `layout_id` is derived from the native list alone and cannot fence a disagreement between the two. And `checkpoint_bytes_for` goes, along with the duplicated flattening that had left it with no production caller. ## Gates Unit tests 1497 (from 1443 on main); the nine mutations the new assertions exist to catch were each confirmed to fail them. The #1417 prefix-hit gate is 0/6 collapsed with a hit at position 512, which the resumer never wrote and can only have gathered from its image. Startup geometry is unchanged to the byte -- `image_bytes=9060352`, `units_per_checkpoint=6`, layout v3 -- and the reserve line is gone. `demands_declined_no_room` and `checkpoints_dropped` are both zero. Six counterbalanced server instances put GSM8K, throughput, TTFT and MTP acceptance all in overlapping intervals, with each arm's own spread several times the difference between them.
…leaves The gate asked whether an image fits, and then the very admission that asked took its block table. A pool with room for an image but not for the request *and* an image answered yes, `begin_store` refused many forwards later, and the prefill chunk the gate exists to withhold had already been bought -- with `demands_declined_no_room` at zero, so the funnel showed nothing. It now asks `num_new_blocks + units_per_checkpoint`, with the same `protected_hash` `can_allocate` passes to `_has_page_units` on the next line, so the two gates of one pass agree on what eviction could reclaim. It is asked afresh on every attempt, because a demand affordable when it was recorded is not still affordable once it is not; the sequence carries a marker per counter so only the counting stays once per admission. It remains a sample even so, and the comment says which loss it removes and which `checkpoints_dropped` still owns. The reachability refusal moves from `begin_store` into `ensure_free_units`. The bare loop gives up only after evicting everything it can, so an unreachable count destroyed the cache on the way to saying no -- and only `_fresh_block` asking for a single unit kept the other caller from needing it. Staging. `launch_copy_descriptor` now takes a descriptor already resident. A pageable `torch.from_numpy(x).to(dev)` issued from `build()` synchronizes the current stream, so the host waits out the whole enqueued forward rather than the 800 KB: measured 2.9 ms behind 4 ms of work against 0.1 ms through a pinned `CpuGpuBuffer`, a cost the transfer's own size says nothing about. The kernel's tile and span counts stop being `tl.constexpr` -- specialising on them keyed the compiled kernel to a pool geometry, so every image shape missed the on-disk cache to save a divide the copy does not notice. New `AttentionMetadataBuilder.warmup_per_req_cache`, called once after the pools are installed. Everything the copy path builds lazily -- the plan, the slot views, the slot base table, the tiling's upload, the pinned buffer, the Triton JIT -- otherwise lands inside the batch of whichever request first crosses a rung. Guards that could not fire, or fired on the wrong thing: - `_assert_ratios_divide_block` compared `CSA_RATIO`/`HCA_RATIO` against an alignment `config.py` pins to 256 for every `DeepseekV4*`, so it could only restate the config. Renamed `_assert_ratios_divide_the_alignment` and aimed at `hf_config.compress_ratios`, which is what a variant is free to change. A non-positive alignment gets its own refusal, since every ratio divides zero and the ratio check would otherwise accuse an empty list. - `merge_abutting` silently merged unordered runs: `[(0,400),(256,64), (320,64)]` returned `[(0,400),(256,128)]`, double-counting 128 bytes into every consumer of `checkpoint_image_bytes`, its own cross-check included. - `write_descriptor` let numpy broadcast a short `dst_bases`, aiming every copy of a batch at the first image's addresses, and accepted a flat one that failed two lines later about an array the caller never passed. - `_page_unit_regions` keys on the addresses it was built from. Half of them come from pools `_invalidate_pool_caches` does not own, so that hook could never have been the invariant. - `checkpoint_ranges_for` no longer emits zero-length ranges, which `plan_segmented_copy` refuses on the first copy -- after sizing, the cross-check and startup had all passed. `has_room_for_store` and `reclaimable_units` are gone; neither had a production caller left. The first survives in the tests as `an_image_fits_on_its_own`, which is the contrast the new gate is read against. 1522 passed / 52 skipped. Every new assertion was checked against a mutation restoring the behaviour it describes.
Contributor
🏷️ CI GuideRuns automatically on every eligible PR before approval:
Heavy model tests:
|
The non-GPU CI runner collected it as an error: every class in the file reads
unbound methods off `DeepseekV4AttentionMetadataBuilder`, and that module does
`from aiter import dtypes` at load. One error in 1321 collected items, and the
only test module in the repo importing that chain without a guard.
Two things the obvious guard gets wrong, both found by rebuilding the runner's
shape locally (an empty `aiter` namespace package shadowing the real one,
which reproduces the message verbatim):
`importorskip("aiter")`, which is what the neighbouring V4 kernel tests use,
is not the right question here. The failure reads "cannot import name
'dtypes' from 'aiter' (unknown location)" rather than "no module named", so
`aiter` is resolving as a namespace package and a guard on it can succeed and
leave the real import to fail anyway. Asked of the module actually needed.
`exc_type=ImportError` is required, not tidiness. The module *is* found, so a
bare `importorskip` treats an ImportError out of it as the caller's mistake:
a deprecation warning on pytest 9.0, which is what this box has, and an error
from 9.1, which is what CI runs -- so the unqualified form would have left CI
red. Naming the type also keeps the skip narrow: anything that is not an
ImportError still fails.
Verified both ways, because a guard that skips everywhere is not a fix: under
the runner's shape with 9.1 semantics the module skips where it used to error,
and against the real aiter on this box its 26 tests still run.
Those 26 assertions are now CI-skipped, as every V4-kernel-adjacent module
already is. Splitting the file would not recover any of them -- there is no
class in it that does not go through the builder.
ganyi1996ppo
added a commit
that referenced
this pull request
Aug 19, 2026
… a -1 interval Rebased onto main's PAGE-backed checkpoint work (#1874, #1894, #1943, #1880), which rewrote this subsystem underneath the branch. Squashed to one commit because the three original commits each re-conflicted against the new base and against each other's resolutions; the reasoning from all three is kept below. --- allocate state slots per need, not by group A "state cache group" was `1 + num_spec` slots wide and was the unit of everything: allocation, admission, sizing, and the checkpoint index. But a checkpoint has no speculation to roll back -- it holds a committed state -- so filing one cost a full group and wasted `num_spec/(1 + num_spec)` of its bytes. At two speculative tokens that is two thirds. The slot is now the unit. `--state-checkpoint-slots 64` buys 64 checkpoints for 64 slots instead of 192, and the slots it no longer takes stay in the paged KV pool. This is only possible because a request's slots need not be adjacent, which the kernels never required: the ssm kernel gathers each index out of the indices tensor and the conv path is handed column 0 alone. Contiguity was manufactured by `prepare_state_indices` writing `arange(base, base + width)`; it now writes the seq's own slot list straight in. `StateGroupPool` -> `StateSlotPool`, `Sequence.per_req_cache_group` -> `state_slots`, whose element 0 is the committed state. The setter re-points [0] and preserves [1:], because speculation scratch persists across forwards. `--state-checkpoint-groups` still parses, as an alias. --- -1 turns off the interval ladder without turning off checkpointing The interval is a guess about where reuse will resume; a demand rung is a position a request was actually refused at. On the SemiAnalysis cc-traces the 8192 ladder placed ~30x the writes of the demand rung alone and caught reuse the demand already reaches -- 0.0% of resumes landed on a ladder rung -- while every rung costs the prompt that keeps it an extra prefill chunk. >0 a rung every N tokens (unchanged, still the default) 0 state checkpointing off entirely (unchanged) -1 no interval rungs; the demand rung and prompt-end anchor still place them -1 rather than reusing 0 because 0 is the documented contract and is reachable by accident: the grid snap rounds an off-grid interval down and can land on 0, so a --block-size typo currently fails safe. Three of the four sites are not the arithmetic you would guess -- `pos % interval` under -1 admits *every* position rather than none, and `pos - last < -1` is true for every pos. --- make the demand rung switchable A demand rung is 47% of checkpoint writes on the cc-traces and reads back 2.8% of the time, against 85.2% for a prompt-end anchor. Gated independently of --state-checkpoint-interval-tokens, because the demand is not part of the interval grid. Default unchanged. The refusal is still measured when the placement is off -- switching off a rung must not blind the diagnostic that justifies it. --- reconciliation with main Dropped: the copy/pending-checkpoint path (`_commit_pending`, `record_copy`, `take_copies`, `Sequence.pending_checkpoint`). DeepSeek-V4 checkpointing moved to `PagedStateCheckpointCoordinator`, and `StateSlotPool` now rejects `transfer.copies` outright. The fork path (GDN), where the measured wins are, is kept in full. `readable_midstep` is carried on main's `StateTransfer` in `state_runtime.py` (wire format included) rather than on the branch's copy. Two bugs this rebase exposed, both fixed here: - `PagedStateCheckpointCoordinator` did not implement the midstep half of the `StateCache` protocol, so `checkpoint_cut` raised on every V4 batch. A PAGE image is not readable midstep, so the three methods are the no-ops the protocol documents. - `_record_checkpoint_end` could place the anchor past the last matchable block. `can_allocate` stops one block short of the prompt, so a checkpoint filed under the final block's hash is one no scan looks up -- and being stored, it evicted the ladder rung that would have served the resume, taking an identical re-request from 8 hit blocks to 0. Capped at `(n_hash_blocks - 1) * hbs`. Tests: `tests/test_state_checkpoint.py` 170 passed. Full suite 45 failed / 2300 passed, a strict subset of origin/main's own 114 pre-existing failures -- zero regressions, verified by set difference against a clean origin/main worktree. black clean; ruff no new findings. Co-Authored-By: Claude <noreply@anthropic.com>
ganyi1996ppo
added a commit
that referenced
this pull request
Aug 19, 2026
… a -1 interval Rebased onto main's PAGE-backed checkpoint work (#1874, #1894, #1943, #1880), which rewrote this subsystem underneath the branch. Squashed to one commit because the three original commits each re-conflicted against the new base and against each other's resolutions; the reasoning from all three is kept below. --- allocate state slots per need, not by group A "state cache group" was `1 + num_spec` slots wide and was the unit of everything: allocation, admission, sizing, and the checkpoint index. But a checkpoint has no speculation to roll back -- it holds a committed state -- so filing one cost a full group and wasted `num_spec/(1 + num_spec)` of its bytes. At two speculative tokens that is two thirds. The slot is now the unit. `--state-checkpoint-slots 64` buys 64 checkpoints for 64 slots instead of 192, and the slots it no longer takes stay in the paged KV pool. This is only possible because a request's slots need not be adjacent, which the kernels never required: the ssm kernel gathers each index out of the indices tensor and the conv path is handed column 0 alone. Contiguity was manufactured by `prepare_state_indices` writing `arange(base, base + width)`; it now writes the seq's own slot list straight in. `StateGroupPool` -> `StateSlotPool`, `Sequence.per_req_cache_group` -> `state_slots`, whose element 0 is the committed state. The setter re-points [0] and preserves [1:], because speculation scratch persists across forwards. `--state-checkpoint-groups` still parses, as an alias. --- -1 turns off the interval ladder without turning off checkpointing The interval is a guess about where reuse will resume; a demand rung is a position a request was actually refused at. On the SemiAnalysis cc-traces the 8192 ladder placed ~30x the writes of the demand rung alone and caught reuse the demand already reaches -- 0.0% of resumes landed on a ladder rung -- while every rung costs the prompt that keeps it an extra prefill chunk. >0 a rung every N tokens (unchanged, still the default) 0 state checkpointing off entirely (unchanged) -1 no interval rungs; the demand rung and prompt-end anchor still place them -1 rather than reusing 0 because 0 is the documented contract and is reachable by accident: the grid snap rounds an off-grid interval down and can land on 0, so a --block-size typo currently fails safe. Three of the four sites are not the arithmetic you would guess -- `pos % interval` under -1 admits *every* position rather than none, and `pos - last < -1` is true for every pos. --- make the demand rung switchable A demand rung is 47% of checkpoint writes on the cc-traces and reads back 2.8% of the time, against 85.2% for a prompt-end anchor. Gated independently of --state-checkpoint-interval-tokens, because the demand is not part of the interval grid. Default unchanged. The refusal is still measured when the placement is off -- switching off a rung must not blind the diagnostic that justifies it. --- DeepSeek-V4 Unaffected by the slot-vs-group change, and that claim is now checked rather than asserted: DSV4 declares `entries_per_req=1` unconditionally (the MTP/DSpark lookahead widens the slot via `win_with_spec`, it never multiplies the count), so `state_slots_per_req == 1`, `pop_many(1)` pops the same index `pop()` did, and slot == group exactly as before. `v4_pool_geometry.py` and `sub_pool_spec.py` are untouched, so the `_physical_slots` reversal and DSV4's pool size are both byte-identical to main. DSV4 passes no `extra_entries`, so `--state-checkpoint- slots` is inert for it. The *anchor*, though, did reach DSV4 -- and cost it. `PagedStateCheckpointCoord- inator.applies()` is true for any V4 seq, so `_record_checkpoint_end` reserved a prompt-end anchor and `checkpoint_cut` shortened a prefill chunk onto it. But the coordinator files one pending checkpoint per seq (`_pending[id(seq)]`, and it `del`s `boundary_blocks`), so the prompt-end checkpoint landing a chunk later overwrote the anchor before either was stored. Measured: one extra prefill chunk per prompt for a hit rate that did not move (identical re-send 0 blocks either way, continuation 11 either way). So the anchor is now gated on `keeps_interior_boundaries`, which `StateSlotPool` answers True (each boundary is its own slot in the index, so both survive) and the PAGE coordinator answers False. Asked as a capability rather than by naming the backend, so a future multi-boundary copy class opts in by answering yes. DSV4 is back to main's one-cut prefill; GDN keeps the anchor. Pinned by `test_a_last_boundary_only_class_is_not_anchored_for`. --- reconciliation with main Dropped: the copy/pending-checkpoint path (`_commit_pending`, `record_copy`, `take_copies`, `Sequence.pending_checkpoint`). DeepSeek-V4 checkpointing moved to `PagedStateCheckpointCoordinator`, and `StateSlotPool` now rejects `transfer.copies` outright. The fork path (GDN), where the measured wins are, is kept in full. `readable_midstep` is carried on main's `StateTransfer` in `state_runtime.py` (wire format included) rather than on the branch's copy. Two bugs this rebase exposed, both fixed here: - `PagedStateCheckpointCoordinator` did not implement the midstep half of the `StateCache` protocol, so `checkpoint_cut` raised on every V4 batch. This was introduced *by this branch*, not latent in main: `readable_midstep` does not exist on main at all, and main's `checkpoint_cut` never consults it. A PAGE image is not readable midstep, so the three methods are the no-ops the protocol documents. - `_record_checkpoint_end` could place the anchor past the last matchable block. `can_allocate` stops one block short of the prompt, so a checkpoint filed under the final block's hash is one no scan looks up -- and being stored, it evicted the ladder rung that would have served the resume, taking an identical re-request from 8 hit blocks to 0. Capped at `(n_hash_blocks - 1) * hbs`. Tests: `tests/test_state_checkpoint.py` 171 passed; the state/checkpoint and DSV4/LMCache suites together 632 passed. Full suite 45 failed / 2300 passed, a strict subset of origin/main's own 114 pre-existing failures -- zero regressions, verified by set difference against a clean origin/main worktree. black clean; ruff no new findings. Co-Authored-By: Claude <noreply@anthropic.com>
ganyi1996ppo
added a commit
that referenced
this pull request
Aug 27, 2026
… a -1 interval Rebased onto main's PAGE-backed checkpoint work (#1874, #1894, #1943, #1880), which rewrote this subsystem underneath the branch. Squashed to one commit because the three original commits each re-conflicted against the new base and against each other's resolutions; the reasoning from all three is kept below. --- allocate state slots per need, not by group A "state cache group" was `1 + num_spec` slots wide and was the unit of everything: allocation, admission, sizing, and the checkpoint index. But a checkpoint has no speculation to roll back -- it holds a committed state -- so filing one cost a full group and wasted `num_spec/(1 + num_spec)` of its bytes. At two speculative tokens that is two thirds. The slot is now the unit. `--state-checkpoint-slots 64` buys 64 checkpoints for 64 slots instead of 192, and the slots it no longer takes stay in the paged KV pool. This is only possible because a request's slots need not be adjacent, which the kernels never required: the ssm kernel gathers each index out of the indices tensor and the conv path is handed column 0 alone. Contiguity was manufactured by `prepare_state_indices` writing `arange(base, base + width)`; it now writes the seq's own slot list straight in. `StateGroupPool` -> `StateSlotPool`, `Sequence.per_req_cache_group` -> `state_slots`, whose element 0 is the committed state. The setter re-points [0] and preserves [1:], because speculation scratch persists across forwards. `--state-checkpoint-groups` still parses, as an alias. --- -1 turns off the interval ladder without turning off checkpointing The interval is a guess about where reuse will resume; a demand rung is a position a request was actually refused at. On the SemiAnalysis cc-traces the 8192 ladder placed ~30x the writes of the demand rung alone and caught reuse the demand already reaches -- 0.0% of resumes landed on a ladder rung -- while every rung costs the prompt that keeps it an extra prefill chunk. >0 a rung every N tokens (unchanged, still the default) 0 state checkpointing off entirely (unchanged) -1 no interval rungs; the demand rung and prompt-end anchor still place them -1 rather than reusing 0 because 0 is the documented contract and is reachable by accident: the grid snap rounds an off-grid interval down and can land on 0, so a --block-size typo currently fails safe. Three of the four sites are not the arithmetic you would guess -- `pos % interval` under -1 admits *every* position rather than none, and `pos - last < -1` is true for every pos. --- make the demand rung switchable A demand rung is 47% of checkpoint writes on the cc-traces and reads back 2.8% of the time, against 85.2% for a prompt-end anchor. Gated independently of --state-checkpoint-interval-tokens, because the demand is not part of the interval grid. Default unchanged. The refusal is still measured when the placement is off -- switching off a rung must not blind the diagnostic that justifies it. --- DeepSeek-V4 Unaffected by the slot-vs-group change, and that claim is now checked rather than asserted: DSV4 declares `entries_per_req=1` unconditionally (the MTP/DSpark lookahead widens the slot via `win_with_spec`, it never multiplies the count), so `state_slots_per_req == 1`, `pop_many(1)` pops the same index `pop()` did, and slot == group exactly as before. `v4_pool_geometry.py` and `sub_pool_spec.py` are untouched, so the `_physical_slots` reversal and DSV4's pool size are both byte-identical to main. DSV4 passes no `extra_entries`, so `--state-checkpoint- slots` is inert for it. The *anchor*, though, did reach DSV4 -- and cost it. `PagedStateCheckpointCoord- inator.applies()` is true for any V4 seq, so `_record_checkpoint_end` reserved a prompt-end anchor and `checkpoint_cut` shortened a prefill chunk onto it. But the coordinator files one pending checkpoint per seq (`_pending[id(seq)]`, and it `del`s `boundary_blocks`), so the prompt-end checkpoint landing a chunk later overwrote the anchor before either was stored. Measured: one extra prefill chunk per prompt for a hit rate that did not move (identical re-send 0 blocks either way, continuation 11 either way). So the anchor is now gated on `keeps_interior_boundaries`, which `StateSlotPool` answers True (each boundary is its own slot in the index, so both survive) and the PAGE coordinator answers False. Asked as a capability rather than by naming the backend, so a future multi-boundary copy class opts in by answering yes. DSV4 is back to main's one-cut prefill; GDN keeps the anchor. Pinned by `test_a_last_boundary_only_class_is_not_anchored_for`. --- reconciliation with main Dropped: the copy/pending-checkpoint path (`_commit_pending`, `record_copy`, `take_copies`, `Sequence.pending_checkpoint`). DeepSeek-V4 checkpointing moved to `PagedStateCheckpointCoordinator`, and `StateSlotPool` now rejects `transfer.copies` outright. The fork path (GDN), where the measured wins are, is kept in full. `readable_midstep` is carried on main's `StateTransfer` in `state_runtime.py` (wire format included) rather than on the branch's copy. Two bugs this rebase exposed, both fixed here: - `PagedStateCheckpointCoordinator` did not implement the midstep half of the `StateCache` protocol, so `checkpoint_cut` raised on every V4 batch. This was introduced *by this branch*, not latent in main: `readable_midstep` does not exist on main at all, and main's `checkpoint_cut` never consults it. A PAGE image is not readable midstep, so the three methods are the no-ops the protocol documents. - `_record_checkpoint_end` could place the anchor past the last matchable block. `can_allocate` stops one block short of the prompt, so a checkpoint filed under the final block's hash is one no scan looks up -- and being stored, it evicted the ladder rung that would have served the resume, taking an identical re-request from 8 hit blocks to 0. Capped at `(n_hash_blocks - 1) * hbs`. Tests: `tests/test_state_checkpoint.py` 171 passed; the state/checkpoint and DSV4/LMCache suites together 632 passed. Full suite 45 failed / 2300 passed, a strict subset of origin/main's own 114 pre-existing failures -- zero regressions, verified by set difference against a clean origin/main worktree. black clean; ruff no new findings. Co-Authored-By: Claude <noreply@anthropic.com>
Merged
ganyi1996ppo
added a commit
that referenced
this pull request
Aug 27, 2026
… a -1 interval Rebased onto main's PAGE-backed checkpoint work (#1874, #1894, #1943, #1880), which rewrote this subsystem underneath the branch. Squashed to one commit because the three original commits each re-conflicted against the new base and against each other's resolutions; the reasoning from all three is kept below. --- allocate state slots per need, not by group A "state cache group" was `1 + num_spec` slots wide and was the unit of everything: allocation, admission, sizing, and the checkpoint index. But a checkpoint has no speculation to roll back -- it holds a committed state -- so filing one cost a full group and wasted `num_spec/(1 + num_spec)` of its bytes. At two speculative tokens that is two thirds. The slot is now the unit. `--state-checkpoint-slots 64` buys 64 checkpoints for 64 slots instead of 192, and the slots it no longer takes stay in the paged KV pool. This is only possible because a request's slots need not be adjacent, which the kernels never required: the ssm kernel gathers each index out of the indices tensor and the conv path is handed column 0 alone. Contiguity was manufactured by `prepare_state_indices` writing `arange(base, base + width)`; it now writes the seq's own slot list straight in. `StateGroupPool` -> `StateSlotPool`, `Sequence.per_req_cache_group` -> `state_slots`, whose element 0 is the committed state. The setter re-points [0] and preserves [1:], because speculation scratch persists across forwards. `--state-checkpoint-groups` still parses, as an alias. --- -1 turns off the interval ladder without turning off checkpointing The interval is a guess about where reuse will resume; a demand rung is a position a request was actually refused at. On the SemiAnalysis cc-traces the 8192 ladder placed ~30x the writes of the demand rung alone and caught reuse the demand already reaches -- 0.0% of resumes landed on a ladder rung -- while every rung costs the prompt that keeps it an extra prefill chunk. >0 a rung every N tokens (unchanged, still the default) 0 state checkpointing off entirely (unchanged) -1 no interval rungs; the demand rung and prompt-end anchor still place them -1 rather than reusing 0 because 0 is the documented contract and is reachable by accident: the grid snap rounds an off-grid interval down and can land on 0, so a --block-size typo currently fails safe. Three of the four sites are not the arithmetic you would guess -- `pos % interval` under -1 admits *every* position rather than none, and `pos - last < -1` is true for every pos. --- make the demand rung switchable A demand rung is 47% of checkpoint writes on the cc-traces and reads back 2.8% of the time, against 85.2% for a prompt-end anchor. Gated independently of --state-checkpoint-interval-tokens, because the demand is not part of the interval grid. Default unchanged. The refusal is still measured when the placement is off -- switching off a rung must not blind the diagnostic that justifies it. --- DeepSeek-V4 Unaffected by the slot-vs-group change, and that claim is now checked rather than asserted: DSV4 declares `entries_per_req=1` unconditionally (the MTP/DSpark lookahead widens the slot via `win_with_spec`, it never multiplies the count), so `state_slots_per_req == 1`, `pop_many(1)` pops the same index `pop()` did, and slot == group exactly as before. `v4_pool_geometry.py` and `sub_pool_spec.py` are untouched, so the `_physical_slots` reversal and DSV4's pool size are both byte-identical to main. DSV4 passes no `extra_entries`, so `--state-checkpoint- slots` is inert for it. The *anchor*, though, did reach DSV4 -- and cost it. `PagedStateCheckpointCoord- inator.applies()` is true for any V4 seq, so `_record_checkpoint_end` reserved a prompt-end anchor and `checkpoint_cut` shortened a prefill chunk onto it. But the coordinator files one pending checkpoint per seq (`_pending[id(seq)]`, and it `del`s `boundary_blocks`), so the prompt-end checkpoint landing a chunk later overwrote the anchor before either was stored. Measured: one extra prefill chunk per prompt for a hit rate that did not move (identical re-send 0 blocks either way, continuation 11 either way). So the anchor is now gated on `keeps_interior_boundaries`, which `StateSlotPool` answers True (each boundary is its own slot in the index, so both survive) and the PAGE coordinator answers False. Asked as a capability rather than by naming the backend, so a future multi-boundary copy class opts in by answering yes. DSV4 is back to main's one-cut prefill; GDN keeps the anchor. Pinned by `test_a_last_boundary_only_class_is_not_anchored_for`. --- reconciliation with main Dropped: the copy/pending-checkpoint path (`_commit_pending`, `record_copy`, `take_copies`, `Sequence.pending_checkpoint`). DeepSeek-V4 checkpointing moved to `PagedStateCheckpointCoordinator`, and `StateSlotPool` now rejects `transfer.copies` outright. The fork path (GDN), where the measured wins are, is kept in full. `readable_midstep` is carried on main's `StateTransfer` in `state_runtime.py` (wire format included) rather than on the branch's copy. Two bugs this rebase exposed, both fixed here: - `PagedStateCheckpointCoordinator` did not implement the midstep half of the `StateCache` protocol, so `checkpoint_cut` raised on every V4 batch. This was introduced *by this branch*, not latent in main: `readable_midstep` does not exist on main at all, and main's `checkpoint_cut` never consults it. A PAGE image is not readable midstep, so the three methods are the no-ops the protocol documents. - `_record_checkpoint_end` could place the anchor past the last matchable block. `can_allocate` stops one block short of the prompt, so a checkpoint filed under the final block's hash is one no scan looks up -- and being stored, it evicted the ladder rung that would have served the resume, taking an identical re-request from 8 hit blocks to 0. Capped at `(n_hash_blocks - 1) * hbs`. Tests: `tests/test_state_checkpoint.py` 171 passed; the state/checkpoint and DSV4/LMCache suites together 632 passed. Full suite 45 failed / 2300 passed, a strict subset of origin/main's own 114 pre-existing failures -- zero regressions, verified by set difference against a clean origin/main worktree. black clean; ruff no new findings. Co-Authored-By: Claude <noreply@anthropic.com>
valarLip
pushed a commit
that referenced
this pull request
Aug 28, 2026
* feat(state-cache): per-slot allocation, a switchable demand rung, and a -1 interval Rebased onto main's PAGE-backed checkpoint work (#1874, #1894, #1943, #1880), which rewrote this subsystem underneath the branch. Squashed to one commit because the three original commits each re-conflicted against the new base and against each other's resolutions; the reasoning from all three is kept below. --- allocate state slots per need, not by group A "state cache group" was `1 + num_spec` slots wide and was the unit of everything: allocation, admission, sizing, and the checkpoint index. But a checkpoint has no speculation to roll back -- it holds a committed state -- so filing one cost a full group and wasted `num_spec/(1 + num_spec)` of its bytes. At two speculative tokens that is two thirds. The slot is now the unit. `--state-checkpoint-slots 64` buys 64 checkpoints for 64 slots instead of 192, and the slots it no longer takes stay in the paged KV pool. This is only possible because a request's slots need not be adjacent, which the kernels never required: the ssm kernel gathers each index out of the indices tensor and the conv path is handed column 0 alone. Contiguity was manufactured by `prepare_state_indices` writing `arange(base, base + width)`; it now writes the seq's own slot list straight in. `StateGroupPool` -> `StateSlotPool`, `Sequence.per_req_cache_group` -> `state_slots`, whose element 0 is the committed state. The setter re-points [0] and preserves [1:], because speculation scratch persists across forwards. `--state-checkpoint-groups` still parses, as an alias. --- -1 turns off the interval ladder without turning off checkpointing The interval is a guess about where reuse will resume; a demand rung is a position a request was actually refused at. On the SemiAnalysis cc-traces the 8192 ladder placed ~30x the writes of the demand rung alone and caught reuse the demand already reaches -- 0.0% of resumes landed on a ladder rung -- while every rung costs the prompt that keeps it an extra prefill chunk. >0 a rung every N tokens (unchanged, still the default) 0 state checkpointing off entirely (unchanged) -1 no interval rungs; the demand rung and prompt-end anchor still place them -1 rather than reusing 0 because 0 is the documented contract and is reachable by accident: the grid snap rounds an off-grid interval down and can land on 0, so a --block-size typo currently fails safe. Three of the four sites are not the arithmetic you would guess -- `pos % interval` under -1 admits *every* position rather than none, and `pos - last < -1` is true for every pos. --- make the demand rung switchable A demand rung is 47% of checkpoint writes on the cc-traces and reads back 2.8% of the time, against 85.2% for a prompt-end anchor. Gated independently of --state-checkpoint-interval-tokens, because the demand is not part of the interval grid. Default unchanged. The refusal is still measured when the placement is off -- switching off a rung must not blind the diagnostic that justifies it. --- DeepSeek-V4 Unaffected by the slot-vs-group change, and that claim is now checked rather than asserted: DSV4 declares `entries_per_req=1` unconditionally (the MTP/DSpark lookahead widens the slot via `win_with_spec`, it never multiplies the count), so `state_slots_per_req == 1`, `pop_many(1)` pops the same index `pop()` did, and slot == group exactly as before. `v4_pool_geometry.py` and `sub_pool_spec.py` are untouched, so the `_physical_slots` reversal and DSV4's pool size are both byte-identical to main. DSV4 passes no `extra_entries`, so `--state-checkpoint- slots` is inert for it. The *anchor*, though, did reach DSV4 -- and cost it. `PagedStateCheckpointCoord- inator.applies()` is true for any V4 seq, so `_record_checkpoint_end` reserved a prompt-end anchor and `checkpoint_cut` shortened a prefill chunk onto it. But the coordinator files one pending checkpoint per seq (`_pending[id(seq)]`, and it `del`s `boundary_blocks`), so the prompt-end checkpoint landing a chunk later overwrote the anchor before either was stored. Measured: one extra prefill chunk per prompt for a hit rate that did not move (identical re-send 0 blocks either way, continuation 11 either way). So the anchor is now gated on `keeps_interior_boundaries`, which `StateSlotPool` answers True (each boundary is its own slot in the index, so both survive) and the PAGE coordinator answers False. Asked as a capability rather than by naming the backend, so a future multi-boundary copy class opts in by answering yes. DSV4 is back to main's one-cut prefill; GDN keeps the anchor. Pinned by `test_a_last_boundary_only_class_is_not_anchored_for`. --- reconciliation with main Dropped: the copy/pending-checkpoint path (`_commit_pending`, `record_copy`, `take_copies`, `Sequence.pending_checkpoint`). DeepSeek-V4 checkpointing moved to `PagedStateCheckpointCoordinator`, and `StateSlotPool` now rejects `transfer.copies` outright. The fork path (GDN), where the measured wins are, is kept in full. `readable_midstep` is carried on main's `StateTransfer` in `state_runtime.py` (wire format included) rather than on the branch's copy. Two bugs this rebase exposed, both fixed here: - `PagedStateCheckpointCoordinator` did not implement the midstep half of the `StateCache` protocol, so `checkpoint_cut` raised on every V4 batch. This was introduced *by this branch*, not latent in main: `readable_midstep` does not exist on main at all, and main's `checkpoint_cut` never consults it. A PAGE image is not readable midstep, so the three methods are the no-ops the protocol documents. - `_record_checkpoint_end` could place the anchor past the last matchable block. `can_allocate` stops one block short of the prompt, so a checkpoint filed under the final block's hash is one no scan looks up -- and being stored, it evicted the ladder rung that would have served the resume, taking an identical re-request from 8 hit blocks to 0. Capped at `(n_hash_blocks - 1) * hbs`. Tests: `tests/test_state_checkpoint.py` 171 passed; the state/checkpoint and DSV4/LMCache suites together 632 passed. Full suite 45 failed / 2300 passed, a strict subset of origin/main's own 114 pre-existing failures -- zero regressions, verified by set difference against a clean origin/main worktree. black clean; ruff no new findings. Co-Authored-By: Claude <noreply@anthropic.com> * remove gpu unit test Signed-off-by: ganyi <ygan@amd.com> * remove the triton import part Signed-off-by: ganyi <ygan@amd.com> * match main's array('i') token_ids contract in a relocation test `#1990` added an assertion that `Block.token_ids` is an `array('i')`, not a list -- a list never compares equal to what the production publish paths store, so every hit on the block would read as a hash collision. This test was written before that landed and still passed a bare list. The file already has `toks()` for exactly this; the test just did not use it. * feat(state-cache): keep Kimi-K3's KDA checkpoints as PAGE images A KDA Active Slot is 53.6 MiB. Held as a checkpoint it competed with live requests for the pool that admits them, so retaining one cost the workload the concurrency it was retained for. Held as PAGE units it is 127 ordinary KV blocks -- 0.112% of the paged pool -- drawn from the same free list as everything else and evicted by the same LRU. This is the mechanism `main` already ships and DeepSeek-V4 already uses (`PagedStateCheckpointCoordinator`). Nothing about the coordinator changes; what is added is the source side of the copy for a state that is two strided tensors rather than one contiguous slab. `plan_segmented_copy` intersects two ordered byte streams and needs neither block alignment nor equal segments, so the state tensors keep their layout: a slot is 138 ranges (69 conv + 69 ssm) and the planner cuts them against 127 units. `_checkpoint_layer_ranges` is the sole owner of that order -- both the sizes and the addresses read it, because a plan cut against one order and addressed through another lands whole layers in the wrong unit. Two things the port had to get right, both now asserted rather than assumed: - A PAGE unit is a *logical* block, but `kv_cache` is shaped in physical ones and K3's `block_ratio` is 128. `_page_unit_regions` derives its stride from `runner.block_size` and checks `num_rows * region == page_unit_bytes`, so a granularity mix-up is a startup error instead of 127 blocks of scrambled state. Unit ids are range-checked against the logical count for the same reason. - K3's slots are strided by `num_slots`, so an off-by-one in `(layer * num_slots + slot)` lands inside a neighbouring request's live state rather than off the end of the tensor. V4 cannot fail this way and its tests do not look for it; `test_no_bystander_slot_is_touched` does. `state_spec` now asks for no spare checkpoint slots under PAGE. `--state-checkpoint-slots` buys Active Slots for checkpoints to sit in, which a copy does not need -- 1.7 GiB reserved for nothing, and it is the same memory the paged pool wants in order to absorb the images. Both fall back to `fork` under pipeline parallelism and RapidServe, where `get_num_blocks` raises on a copying transfer: answering `copy` there would turn "K3 keeps no state cache" into "K3 does not start". The dtype objection in the old `state_transfer` docstring is retired, not ignored. It was that `_state_dtypes` gives kimi_linear an fp32 v side while the chunked states are bf16, so a checkpoint cut from the kernel's `h` would hand cached requests a rounded state. A PAGE image is copied out of the slot and back into a slot -- both fp32, no kernel output in between, no conversion anywhere. Both dtypes are named in the layout id, so a build that changed either cannot read another's images. Not yet flipped on in anger: `execute_paged_state_copies` is reachable only from `build()`, and the GPU verification (probe at conc 1 and 8 against the known-good 0/1 and 0/8, then GSM8K, then a matched-N hit-rate A/B) is the next step. Known follow-up, measured before it is fixed: the coordinator keeps one checkpoint per sequence (`_pending` is last-writer-wins), where the fork path indexed every boundary. A 24k prompt files at 8192/16384/24576 today and would keep only the last. If the A/B shows the drop, the lever is to make `_pending` hold a list -- deliberately not bundled here, because a mechanism swap plus a policy change is a regression nobody can attribute. * feat(state-cache): keep every boundary a PAGE seq reaches, not just its last `_pending` was keyed by sequence, so a prompt's second checkpoint overwrote its first before either was stored. That made the prompt-end anchor worthless -- it sits under a block from the prompt's end, lands in the same or the adjacent prefill chunk, and was reliably the loser. `_record_checkpoint_end` reads `keeps_interior_boundaries` and duly declined to reserve one. Keyed by `(sequence, prefix hash)` both survive. Reaching the *same* hash twice still collapses, which is what the hash in the key is for: that is one boundary reached again, not two boundaries. This matters because the anchor is the placement that pays. The measurement is already in `_record_checkpoint_end`'s docstring: of 4,808 cc-trace resumes with a nonzero KV hit, 93.5% land on a previous prompt end and 0.0% on the 8192 ladder. The ladder was cutting a prefill chunk every 8192 tokens to store something nothing ever resumed from -- and on this workload a prompt averages 117k tokens, so that is ~14 rungs per request, each one a shortened forward and an image in the paged pool. What makes keeping both affordable is the price a PAGE image pays: 127 blocks, 0.112% of the paged pool, against a whole 53.6 MiB Active Slot under `fork`. The measured run that preceded this kept 1,508 checkpoints with `checkpoints_evicted: 0` -- capacity was never the binding constraint. Run with `--state-checkpoint-interval-tokens -1` to drop the ladder entirely and leave the anchor and the demand as the only two placements. Three tests changed rather than deleted, because each pinned the old behaviour deliberately and each now pins its replacement: - `test_latest_pending_checkpoint_replaces_the_previous_intent` becomes `test_two_boundaries_of_one_seq_are_both_stored`, plus a new sibling for the same-hash-twice case. - `test_a_last_boundary_only_class_is_not_anchored_for` becomes `test_both_classes_are_anchored_for`. - Two demand tests rested a tightened pool on "exactly one image"; a prompt now stores two, so they spend down to the deepest -- which is both the resume target and what `_next_victim` would keep longest. Not yet measured. The preceding PAGE run at conc 8 reached 91.79% at N=791 against a 96.9% trace ceiling; this is the change aimed at that gap, and the A/B is the next step. * refactor(state-cache): drop the parts of the PR nothing reads Three removals, none of which change behaviour. Verified against the same 4610-passed baseline, and `ruff` on the touched files goes 16 -> 14 findings. `cache_pressure.py` had no importer anywhere in the tree, and the log field its regex parses (`Cached/Total:`) was renamed to `Cached/Reusable:` by this same PR -- so it could not have matched a line this branch produces. `keeps_interior_boundaries` was a capability hook with one reader and no implementor that answered `False`: the `getattr` default was `True`, both classes set `True`, and the case it existed for -- the PAGE coordinator overwriting its own anchor -- was fixed earlier in this branch by re-keying `_pending` on `(seq, hash)`. The measurement that justified it (of 4,808 cc-trace resumes with a nonzero KV hit, 93.5% land on a previous prompt end, 0.0% on the 8192 ladder) moves onto `checkpoint`, which is where the key it argues for lives. `_log_frequency` and its four `reqs_*` counters cost four `__slots__` entries and four per-request branches to render one log line, and are read by nothing else -- not `metrics.py`, not either aggregation tuple in `llm_engine.py`. `_log_pools` stays: its three rates are pure ratios of totals already kept, and the paged/state split is this PR's central claim. `_log_pressure` stays because `checkpoints_*` and `demands_recorded` do reach Prometheus. Left alone deliberately: the `record_relocation` / `take_relocations` / `relocate_state_slots` chain is equally unreachable, but it is that way on `main` too. Deleting main's debt from this branch would widen the diff it is meant to narrow. * docs(state-cache): tighten the comments this PR added No code changes; 375 tests pass and `ruff` on the touched files stays at 14 findings against main's 16. The bf16/fp32 accuracy argument was written out in full three times -- `GDNStateMixin.state_transfer`, `pop_last_intermediate_states`, and inverted again in `_KimiMLAGDNCommon.state_transfer` -- each time as a rebuttal to an objection nobody raised, and two of the three cited `tests/test_gdn_state_checkpoint_gpu.py`, deleted in a04ce7f. It now lives once, in the present tense, where the dtypes are chosen; the other two point at it. That alone is ~30 lines and both dead citations. Two measurements had spread to four and five sites. The prompt-end anchor's read-back rate stays in `_record_checkpoint_end`, which exists because of it; the demand rung's stays in `mark_speculative`, the only place it decides behaviour, and in the `--state-checkpoint-demand` help text, where a CLI user cannot follow a code reference. `config.py`, `envs.py`, `sequence.py`, `page_unit_checkpoint.py` and `checkpointers_at` now reference rather than restate, so there is one copy to update when the number moves. The rest is history that git already holds: what an earlier Python-loop version got wrong, what the upstream branch does with `state_cache_base`, what this pool "used to allocate", which objection "kept this on fork". Each is restated as the invariant it was arguing for. Also two stragglers of the group->slot rename in `attention_gdn.py`, and a call-site comment in `gdn_attn.py` that restated `_checkpoint_targets`' own docstring. Left long on purpose: `_page_unit_regions`' logical-vs-physical block-id trap (K3's block_ratio is 128, and getting it wrong scrambles 127 blocks silently), `_assert_checkpoint_geometry_still_holds`, the conv-window claim in `state_transfer`, and `CacheStats`' argument for `reusable` over `full` as the denominator -- that last reads like a rebuttal but the objection is one a reader will actually raise. * docs: describe the two model-agnostic features and the instrumentation The description covered only the K3 PAGE port, which is 1,221 of the 4,694 added lines. Three things it shipped were undocumented: Per-slot allocation. `StateGroupPool` -> `StateSlotPool`, and a request's state goes from one fixed-width group of `1 + num_spec` adjacent slots to a list of ids that need not be adjacent. The point is that a checkpoint takes one slot rather than a whole group, since a resumed prefix has no speculation to roll back. Documents the one consumer that reads past element 0 -- the spec-decode path, which stopped deriving the set from `base = group * slots_per_group` -- and states why DeepSeek-V4 is a rename rather than a behaviour change. Midstep checkpoints. A mamba-like backend can now take every boundary a forward covers out of the chunk kernel's own `h`, instead of having its prefill cut so the forward *ends* on each one. Covers the reserve/publish/cancel split (the bytes do not exist when the destination must be chosen), the `is_end` targets that read the runtime slot because `h` does not hold the final state, and the paired gate in `checkpoint_cut`/`checkpointers_at` -- suppressing one alone keeps zero checkpoints with no error. Names Qwen3-Next and Qwen3.5 as the models on this path and K3 as the one that cannot be, and adds the latter to the follow-ups. Hit-rate instrumentation. Every measurement in this PR was read off these lines. The `[Cache Stats]` denominator was `full`, which includes the trailing block `can_allocate` never matches -- so it charged both pools for a block neither was offered and reported an unreachable ceiling; it is now `reusable`. `[Cache Pools]` splits the series into `paged * state = combined`, which is what showed the paged index matching 99.4% while the state gate discarded it. `[Checkpoint Fates]` separates four fates that argue for different fixes, and `kept: 1508, dropped: 0, evicted: 0` is the evidence behind the "capacity stopped being the binding constraint" claim. Also refreshes the numbers the rebase and the two cleanup commits invalidated: 33 files / +4694, the current commit hashes, the per-file table, and the test baseline (4610 passed / 50 pre-existing failures, 40 of them sglang files that score identically on origin/main). * docs(state-cache): KDA's interior h exists; aiter just does not return it The follow-up said K3 cannot be `readable_midstep` because `chunk_kimi_delta_attn` "exposes only `output_final_state`". True of the API, misleading about the cause: in aiter's `_triton_kernels/chunk_delta_attn/chunk_fwd.py` the per-chunk `h` is computed at line 170 -- by `chunk_gated_delta_rule_fwd_h`, the same function the GDN path uses -- consumed by `chunk_gla_fwd_o`, then set to None at line 202 and left out of the returned tuple. So the two backends differ in plumbing, not in what their kernels produce. ATOM vendors GDN's chunk entry under `model_ops/fla_ops/`, which is why `keep_intermediate_states` could be added there; KDA goes out to aiter, which has no equivalent. Whoever picks this up is adding a return value, not an algorithm -- worth stating, because the old wording invites the conclusion that the kernel would have to be rewritten. Behaviour is unchanged: K3 stays `readable_midstep = False` and keeps cutting a chunk per placement. Under the shipped anchor-only policy that is 0% of prompts cut at 1.00 checkpoints per request, since the anchor lands where the last prefill chunk was going to end anyway. * test(gdn): pin the claim readable_midstep rests on, on real hardware `readable_midstep` asserts that `h[:, j]` is the recurrent state after `j * 64` tokens, and `BlockManager` acts on it by suppressing `checkpoint_cut` outright -- the prefill runs full length and the boundaries are harvested from `h` afterwards. If that assertion is false, every checkpoint the readable path stores is subtly wrong: a resuming request inherits a state its prefix never produced, silently. `TestMidstepCheckpoints` pins everything *around* the claim (which positions are chosen, reserve/publish/cancel, that the cut is suppressed) but stubs the kernel, so it cannot see the claim itself fail. This asks the kernel. Measured on MI355, 8 chunks of 64: all 7 interior boundaries are **bit-exact** against a forward stopped at that position -- `torch.equal`, not a tolerance, which is the right bar because both arms round the same fp32 value into the same dtype (`h` is `k.new_empty`; `_state_dtypes` returns `config.torch_dtype`). Two smaller guards alongside it: popping consumes the reference, so a later layer cannot read the previous one's `h` and file it under its own slot; and a forward that was not asked to keep retains nothing, so the plugins that never pop do not pin a large tensor past their last forward. Needs one GPU and a few hundred MB -- no server, no TP, no weights -- and skips at module level otherwise, following `test_compress_chunk_equivalence`. * docs: record the midstep hardware result, and narrow what is still unmeasured "Qwen3.5 is not measured" was true when written and is now too blunt. The claim `readable_midstep` rests on -- that `h[:, j]` equals the state a forward stopped at `j * 64` would leave -- has been asked of the kernel directly: 7 of 7 interior boundaries bit-exact on MI355. That belongs in Verification, because it is the one part of the midstep path a CPU test cannot reach and a failure there would be silent. What remains unmeasured is narrower and worth saying precisely: no server has been stood up on a readable backend, so there is no hit rate, accuracy, or TTFT for it. Named the three things a single sequence through one kernel cannot show -- the per-sequence `chunk_offsets[row]` base with two prefills in a batch, the ordering an `is_end` target depends on, and a resume landing on a stored midstep boundary -- so the gap is actionable rather than a blanket disclaimer. * fix(state-cache): address review findings 1, 7, 10 and the instrumentation Findings from @valarLip on #2045, each re-verified against the code rather than taken on report -- two of the sixteen did not survive that check (`_rehome_checkpoint` does not exist; `chunk_gated_delta_rule` carries `@torch.compiler.disable`, so the CUDA-graph half of #12 cannot happen). **#1, a regression this branch introduced.** `eb058321d` re-keyed `_pending` to `(seq, hash)` so two boundaries of one prompt could coexist, and did not touch the drain, which still resolves a single `seq.state_slot` for all of them. Both images are then copied out of whatever the last forward left there, filing the earlier hash over the later state -- a request resuming on it continues from ahead of its own prefix, and `_validate_paged_state_op` passes because layout, size and unit count are all still correct. `_supersede` keeps one pending boundary per sequence. Ordinarily a drain follows every forward and both boundaries are stored correctly from their own slots; the exception is a pass that schedules nothing, where `state_maintenance_ops=None` carries `_pending` into the next drain. The newer boundary wins because it is the one the slot holds, and the older is counted `dropped` -- it is reuse the placement asked for and did not get. The two tests that pinned the old behaviour asserted coexistence without asserting each was stored from its own slot, which the drain cannot do; they now pin the fix and a sibling covers the ordinary drain-between-forwards case. **#7** was the same invariant read from the other end: the descriptor buffer is sized `2 * max_num_seqs` on "one store per sequence", which the re-key removed and `_supersede` restores. No resize -- the docstrings here and in `deepseek_v4_attn.py` now name what holds the bound instead of asserting it. **#2/#3/#4/#5/#8 are all on the midstep write path, so `readable_midstep` goes back to False.** The write path declines on six conditions `commit_midstep` cannot see and publishes the hash regardless; `_checkpoint_targets` indexes three differently scoped sequence lists with one `i`; the SSM read floors to a 64 grid `midstep_positions` does not enforce (`hash_block_size` defaults to 16); the conv window is `conv_kernel-1+num_spec` in the kernel and `conv_kernel-1` in the guard. Each stores a findable image holding the wrong state. None of it has run under a server -- K3 takes the PAGE path and cannot reach it -- so the honest state is off. The machinery and its bit-exactness test stay; `test_midstep_is_off_in_production.py` pins the decision, and is deliberately not behind `importorskip` so the non-GPU runner actually runs it. **#10** `pool_pressure` read `self.state`, which under PAGE is a different object built with `StateTransfer.none()` that never sees a `checkpoint()`. It printed four zeros for the life of the server while `checkpoint_funnel`, the next method, reported the real numbers from the coordinator. Smaller: `chunks_cut_for_end` and `checkpoints_orphaned` reach both aggregation whitelists (a cut counter without its sibling is unreadable; `orphaned` argues for a bigger paged pool where `evicted` argues for a bigger state pool); `paged_hit` is dropped as a second name for `compressed_hit`; `_warn_if_unschedulable` compares against `state_slots_per_req` again, so a pool too narrow for one request warns instead of waiting forever in silence. `clear_index` still moves no counter, now stated as a decision: each fate argues for a different fix and an operator emptying the cache argues for none. * fix(state-cache): address review findings 6, 9, 11, 13 and 14 **#9 inverts the policy it implements, on the majority of prompts.** `mark_speculative` exists so a guessed resume point is spent before a known one -- anchors are read back 85.2% of the time against a demand rung's 2.8%. Both call sites gated on `if anchor and pos != anchor`, so a seq whose `checkpoint_end_pos` is 0 demoted *nothing* and filed its guesses at the LRU tail beside real anchors. `_record_checkpoint_end` leaves it at 0 on four paths, one being every prompt too short for a keepable end -- the common shape of an agentic first turn. `_anchor_of` answers None rather than 0 there, which compares unequal to every position, so those seqs demote everything. Shared by both sites because two spellings of one rule is how they drift apart. `publish_midstep(seq=None)` still demotes nothing, now as the stated other end of the rule: a caller with no sequence cannot tell a guess from knowledge, and over-keeping costs one eviction where over-demoting spends an anchor. **#6 could take the engine down over a log line.** The two asserts run for every prefill seq, over counters with four independent writers -- the CPU-offload wake sets `num_cached_tokens` without touching the hit-block counters the rest derive from, so an LMCache resume that loads more prefix than the GPU index held produces `cached > wanted` legitimately. Now a warning and a clamp, which also removes a behaviour difference between `-O` and not. **#11 had two consumers disagreeing about what -1 means.** `BlockManager` clamps to `max(-1, ...)` and reads -1 as "grid off, anchor and demand still placing"; the DSV4 offload policy clamped to `max(0, ...)`, folding it into 0, which for that consumer means no sidecar checkpoints at all -- so the engine kept checkpointing while offload resume silently degraded to zero reuse. With no grid to align to the sidecar now takes `resume_alignment` alone. **#13b/#13c.** `_extend_hash_chain` sat one line above the `_has_page_units` refusal, so a 128k prompt queued behind a full pool paid ~2000 xxhash rounds per waiting request per pass for a list that was then discarded. Moved below it; verified nothing between consumes it, and that its one reader is `midstep_positions`. The comment claiming it "reads `checkpoint_end_pos`" was false and is replaced with what actually orders the call. **#14.** `chunk_gated_delta_rule_fwd` returned `h` unconditionally, so the caller's frame pinned ~33 MB for the rest of that layer's forward even with `keep_intermediate_states=False` -- every GDN prefill with checkpointing off, which is the default, and every vLLM/SGLang/rtpllm caller. The flag now reaches the producer, so the reference dies with the fwd frame. No compute changes; the kernel computed it either way. `test_gdn_midstep_state_gpu.py` covers both values and still passes on hardware. Suite: 4622 passed against 4618 before, with the same 50 pre-existing failures. * docs: carry the slot rename into the guides, and document the new knobs The rename landed in code and left five guides describing a `group` model that no longer exists. One of them was actively dangerous: the `deallocate` snippet in the scheduling guide released `seq.per_req_cache_group` — a single slot — where the real function calls `release_many(seq.state_slots)`. Copied as written it leaks `num_spec` slots per request, and admission cannot see the loss because it gates on the free list this never returns them to. Corrected across `scheduling_kv_cache_guide.md` (the pool construction snippet, the allocation and deallocation prose, the pool-field list, the Sequence table, and the fork-checkpoint capacity paragraph), plus the Sequence rows in `architecture_guide.md` and the GDN state paragraph in `model_support_guide.md`. `state_slots` is documented as a list with `[0]` committed and `[1:]` rollback, explicitly not adjacent, with `state_slot` as the property over element 0 — which is the contract a backend has to know before it indexes anything. Newly documented rather than merely renamed: * `--state-checkpoint-interval-tokens -1`. The guides described `0` as the only off switch, so the ladder-off-but-anchor-on regime this PR added was reachable and undocumented. * `--state-checkpoint-slots` (and its `--state-checkpoint-groups` alias), with the note that a PAGE backend zeroes it out. * `--state-checkpoint-demand` / `--no-state-checkpoint-demand`. * `ATOM_STATE_CHECKPOINT_DEMAND`, under a new "State checkpoints" section in `environment_variables.md` — it had no entry at all. Every symbol the guides now name was checked to exist in `atom/`. The two `*_plan.md` files still say `group`; they are dated design notes rather than reference docs, and rewriting them would misrepresent what was planned. * revert: drop two changes that belong to other PRs Neither touches per-request state, checkpoints, or the pools. They rode along on this branch and widen its review surface for no reason. `triton_merge_attn_states.py` moves `prefill_tokens_with_context` off `tl.constexpr`. That is a real fix — a per-batch token count as a constexpr mints a fresh kernel per distinct batch size, 184 of them in one 8-minute agentic run — but it is an attention-kernel compile-time bug, not a state-cache one, and belongs in a PR that says so. `tests/plugin/test_vllm_kimi_k3.py` moved its registry check out-of-process to survive `sys.modules` damage other plugin tests do. Also genuine, also unrelated; verified it does not pollute the session on its own. `test_rtpllm_forward_context_semantics.py` is NOT reverted, though it looked like the same category. Its change makes the stubs it installs restore what they displaced, and without it `atom.model_ops.attention_gdn` and `atom.utils.forward_context` stay shadowed for the rest of the session: reverting it turned 3 collection errors into 6, taking `test_cudagraph_capture_bounds.py` (9 passed alone) and three sglang plugin modules down with `cannot import name ... (unknown location)`. That is load-bearing for whether this branch's own suite can be run at all. * remove --state-checkpoint-slots, which never took effect The flag sized a flat cushion of spare Active Slots for checkpoints to sit in. It defaults to 0, DeepSeek-V4 never declared it, and Kimi-K3 overrode it back to 0 — so on every shipped path it added nothing, and the only configuration where it did anything was GDN's `fork` with someone passing a value by hand, which no measurement in this PR or before it covers. What it was for is real: a checkpoint held as a slot competes with live requests, so how many can be retained is set by concurrency rather than by how much reuse the traffic has. The PAGE path solves that properly, by keeping the image in KV blocks instead of a slot. A cushion would buy the same decoupling for `fork` at the price of a knob nobody can size without measuring first. Removing it collapses two things it had propped up. `_KimiMLAGDNCommon.state_spec` existed only to zero the field and is deleted — with the flag gone the base spec is already right, and `super()` needs no correction. And `TestTheSpareSlotsGoBackToTheKvPool` went with it: it monkeypatched `GDNStateMixin.state_spec` to a lambda returning a literal `extra_entries=32`, so it asserted against its own stub and would have passed unchanged if production stopped reading the field altogether. That is the shape @valarLip flagged, and deleting the feature removes the test's subject rather than its substitute. `SubPoolSpec.extra_entries` stays. It is the sizing layer's general capability, no backend passes a nonzero value today, and the test that pins its arithmetic now says so — a future cushion should get a flat one, not `width x` what it asked for. 4620 passed against 4622 before, the difference being the two deleted tests; same 50 pre-existing failures. `ruff` on the touched files matches origin/main exactly. * test: keep the midstep watch on the CPU-only side of the aiter line `test_gdn_does_not_declare_itself_midstep_readable` reached `GDNStateMixin.state_transfer`, which means importing `gdn_attn`, which imports aiter at module level. The non-GPU CI runner installs CPU torch and neither aiter nor triton, so that is a collection-time ModuleNotFoundError, not a skip -- it failed the job. Split the flag's two halves by what a CPU runner can actually see. The PAGE coordinator's `readable_midstep`, the three-call midstep protocol on `StateCache`, and `StateTransfer`'s field are all pure Python and stay watched here. GDN's declaration is the half that needs aiter; it belongs with the kernel tests, and the docstring now says so rather than leaving the next reader to rediscover it by breaking CI. Coverage lost is one assertion, not the mechanism: `TestMidstepCheckpoints` builds its own `StateTransfer(readable_midstep=True)` and exercises the write path regardless of what production declares. --------- Signed-off-by: ganyi <ygan@amd.com> Co-authored-by: Claude <noreply@anthropic.com> Co-authored-by: Guanbao Yu <Guanbao.Yu@amd.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this is
Eight commits on the DeepSeek-V4 PAGE-backed state checkpoint path. Three strands, each independently measured:
1. The image
units_per_checkpointimage_bytesTwo cuts.
StateField.in_checkpoint=Falsedrops HCA state, which is 52% of a slot and dead by construction: a compressor that poolsratiotokens with no overlap begins its first pool exactly on the boundary a resumer starts at, so every row it reads is one it writes. ThenUnifiedPoolGeometry.entry_row_runsdrops the rows between interleaved per-layer windows — reachable by no(layer, position)pair, so nothing writes or reads them — worth a further 17.3% of the entry on the DSpark configuration.Both rules are named in
layout_idand fenced by its version, because two workers disagreeing about either would read one image at two layouts.2. The copy
At 256 ops the whole path went from 15 ms to 1.29 ms, with the kernel now 88% of it rather than 2%.
What changed: the plan is intersected once and reused (
SegmentedCopyPlan— which source segment meets which destination segment follows from the two streams' sizes; addresses enter only when a copy is issued); the grid is one program per tile that exists rather than(spans, ceil(widest / TILE)), which on an image whose spans run 8 KiB to 1.4 MB had 94% of its programs finding nothing to do; the descriptor is staged through a pinnedCpuGpuBufferrather than uploaded pageably frombuild(), which synchronizes the current stream and made the host wait out the whole enqueued forward (2.9 ms behind 4 ms of work, against 0.1 ms staged — a cost the transfer's own 800 KB says nothing about); and the kernel's tile/span counts stopped beingtl.constexpr, which had keyed the compiled kernel to a pool geometry.num_warpsis pinned at 4 with the reason written down: this kernel is nothing but load and store, and its speed is set by bytes per lane — sixteen (dwordx4) is the fast point, and eight warps halves the width and costs 12%.New
AttentionMetadataBuilder.warmup_per_req_cache, so the plan, slot views, slot base table, tiling upload, pinned buffer and Triton JIT are not paid inside the batch of whichever request first crosses a rung.3. The floor and the gates
reserve_unitsis deleted rather than resized. Both of its stated purposes were tested and neither held: a COPYING window it was said to protect does not exist on the normal path, and it could not make the admission gate always answer yes either (a floor of 320 blocks against the 512 one 128K prompt allocates). Measured: 38 gated runs, zero raises; the bypass control raised 15 of 19.What the floor did do was evict — it was handed to
ensure_free_unitsas part of the count, so a store spent tens of checkpoints building a cushion that bought nothing.Eligibility and policy are now separate:
_is_evictableis the one rule every caller shares,_next_victimis the only place the LRU lives, andhas_available_unitsstops at the shortfall rather than totalling a cache that holds ten thousand checkpoints here.Two gate corrections:
num_new_blocks + units_per_checkpoint, with the sameprotected_hashcan_allocatepasses to_has_page_unitson the next line. Asked afresh per attempt; counted once, via a marker per counter rather than a position the gate overwrites.complete_inflight, so a store now leaves one batch's worth of live-KV demand behind (320 of 103,809 units on V4-Flash-DSpark, 0.31%, logged at startup).Guards that could not fire, or fired on the wrong thing
_assert_ratios_divide_blockcomparedCSA_RATIO/HCA_RATIOagainst an alignmentconfig.pypins to 256 for everyDeepseekV4*— it could only restate the config. Now_assert_ratios_divide_the_alignment, aimed athf_config.compress_ratios.merge_abuttingsilently merged unordered runs:[(0,400),(256,64),(320,64)]returned[(0,400),(256,128)], double-counting 128 bytes into every consumer ofcheckpoint_image_bytes— its own cross-check included.write_descriptorlet numpy broadcast a shortdst_bases, aiming every copy of a batch at the first image's addresses._page_unit_regionsnow keys on the addresses it was built from; half of them come from pools_invalidate_pool_cachesdoes not own.checkpoint_ranges_forno longer emits zero-length ranges, whichplan_segmented_copyrefuses on the first copy — after sizing, the cross-check and startup had all passed.has_room_for_storeandreclaimable_unitsare gone; neither had a production caller left.Test plan
.github/scripts/run_unit_tests.sh.blackclean;ruffcount identical tomain.A B B A A B, six server instances, benchmark ×2 and GSM8K each. All five metrics overlapping. A second six-instance run for MTP acceptance rate was also overlapping (arm difference 0.37–0.46 pp against A's own 1.44–1.94 pp spread). Strict alternation was not used: it confounds arm with position parity, and sequential arms read as a regression twice on this shape.Not covered
_state_fieldsand the two plugin bridges). Left as a follow-up refactor.pipeline_parallel_size > 1the pinned descriptor is neither in theforward_varsring nor covered by_stage_h2d_done, which isNonethere. Atpp_size == 1— every configuration that runs today —_gate_staging_reusecovers it correctly.