Skip to content

perf: unify metadata publication and reuse page tables and Triton kernels - #2399

Merged
valarLip merged 5 commits into
mainfrom
perf/unified-h2d-pr
Sep 25, 2026
Merged

valarLip merged 5 commits into
mainfrom
perf/unified-h2d-pr

Conversation

@valarLip

@valarLip valarLip commented Sep 25, 2026 •

Copy link
Copy Markdown
Collaborator

Forward preparation uploads many small metadata buffers and repacks page tables even when the page mapping is unchanged. Changing batch sizes and prefill lengths also creates new Triton specializations for values that only control runtime bounds.

This PR publishes metadata at its first consumer under a shared ownership contract, reuses page-table state across attention backends, and reuses Triton kernels across dynamic lengths. ATOM_H2D_BACKEND=direct remains the default; packed combines eligible metadata into one H2D followed by scatter.

Changes

  • Add checked publication with source lifetime, stream/device identity, semantic padding, and failure handling owned by the existing runner slots. A shared producer decorator checks write permission before host data can be overwritten. Packed arena ownership follows actual packed transfers, including mixed direct/packed publication.

  • Share versioned page-table preparation across common attention, MHA/MLA, V4, and V4.1. Unchanged mappings avoid page-ID scans and uploads; appends and request reordering update the affected rows. CPU preparation and successful GPU publication have separate revisions. PP/TBO buffers retain independent state, and V4.1 uses the common entry point.

  • Integrate runner, attention, draft/GDN, and Engram consumers at their first-use boundaries. Keep uniform sampling filters on the CPU, reuse independent TBO buffers, and preserve graph padding and fixed destination storage.

  • Remove exact-length specialization from SWA writes and 13 additional Triton kernels. Preserve model geometry, tiles, and useful bounded alignment variants. Cached KV quantization keeps vectorized accesses and supports inputs exceeding 2^31 elements; only the existing small-input reduction configuration uses two-way loop unrolling.

  • Add ATOM:: runner stage annotations when profiling is active, document the publication contract, and organize H2D regressions by owner/consumer responsibility.

  • Bind and retain the model-runner TCPStore before spawning workers. All managed ranks connect as clients, removing the gap between probing a free port and binding it later. Single-node runs allocate the port with port=0; multi-node runs share an explicit fixed port. RapidServe retains separate prefill/decode stores, and standalone ModelRunner callers keep their existing rendezvous path.

Validation

Validation used an AMD Instinct MI355X with PyTorch 2.9.1+ROCm7.1.1 and Triton 3.8.0.

  • H2D/shared-page-table commits: 892 passed, 6 skipped; skips require additional GPUs or an external V4.1 reference. The core, profiler, and SWA commits were also checked independently after history cleanup.
  • Final dynamic-Triton changes: 176 passed, covering numerical output, tail maxima and sentinels, request boundaries, graph replay, and JIT reuse. The 13 targeted kernels reuse one variant across six changing values within a fixed geometry/alignment class.
  • TCPStore startup and related regressions: 123 passed, including real ModelRunner initialization and collectives for TP, DP, TP+DP, PP, PCP, DCP and eight-rank DP; standalone TP2 compatibility; port ownership/release; fixed-port conflicts; and remote-node client selection. Multi-node configuration was tested over loopback, not across physical hosts.
  • A real 1,048,576-token × 12-head KV input passed full output-byte validation. The previous implementation failed to compile this shape because its K/V constexpr element counts inferred incompatible integer types.
  • Formatting and diff checks passed. Ruff passes for the dynamic-Triton changes and introduces no new findings in the startup fix. Existing diagnostics include three Optional annotation warnings and four exception-handling warnings in the manager.

The suites overlap and their counts should not be added. Performance checks use 1M context and operator/metadata microbenchmarks, not end-to-end model throughput. In a controlled shared-buffer comparison, 4096-token KV quantization measured 11.95 → 11.02 μs after the reduction adjustment (11.14 μs before the Triton changes). Some small dynamic shapes retain overhead, including about 0.30 μs for 257-token KV; the accepted shared-page-table path also retains about 2.4 μs in its B256 single-request subsequent-chunk case. Full-model serving, quality, and throughput validation was not rerun.

Publish persistent metadata at its first consumer boundary using checked
direct copies or one packed H2D followed by scatter. Keep source lifetime,
padding, stream identity, PP slots and explicit republication under one owner.
Reject producer reentry before host writes and track actual packed-arena use.

Share versioned page-table preparation across attention backends, including
V4.1. Reuse unchanged maps, copy appended or reordered rows selectively, and
record GPU revisions only after successful publication. Preserve worker append
lineage without retaining page-id tuples or changing scheduler bookkeeping.

Keep uniform sampling filters on the CPU, reuse independent TBO buffers, and
combine Engram tentative cursors with snapshots. Group H2D regressions by
contract and consumer; retain CPU page-map state tests separately.
Label staging reuse, preparation, model execution and postprocessing while
the runner's torch profiler is active. Include batch token counts on forward
markers and avoid record_function scopes when profiling is disabled.
Read the per-batch write width from the launch grid instead of specializing
on each prefill length. Keep head_dim constexpr for the fixed layout and
retain source bounds, ring indexing and padding behavior.

Cover changing widths and odd head dimensions with GPU ring-write regressions.
Move runtime bounds and strides out of exact-value specialization across
Gemma RMSNorm, MRoPE, cached KV quantization, MLA conversion, Qwen4 metadata,
M3 context partitioning, and V4.1 packed gather. Read redundant counts from
the launch grid while retaining model geometry and useful bounded variants.

Preserve vectorized KV accesses with fixed row widths and int64 lengths.
Unroll only the existing small-input reduction configuration and avoid
allocating unused single-pass scratch. Keep all bounds and padding guards.

Add one focused suite covering output correctness, tail boundaries, large
integer signatures, and compiled-kernel reuse. Relevant GPU regressions:
176 passed. Black, Ruff, and diff checks passed for this change.
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every eligible PR before approval:

  • ✅ Pre Checkin: Black, Ruff, catalog schema validation, non-GPU unit tests

Heavy model tests:

  • ✅ Run after the PR is approved and Pre Checkin passes
  • ✅ Run immediately when an approval review is submitted
  • ✅ Can be requested before approval with labels
Label Tests
ci:full Run all heavy PR model tests: native ATOM, vLLM, and SGLang
ci:atom Run native ATOM model accuracy tests
ci:vllm Run ATOM vLLM OOT model accuracy tests
ci:sglang Run ATOM SGLang model accuracy tests

Heavy jobs are skipped when the PR is not approved and no matching ci:* label is present.
Add labels via the sidebar or gh pr edit 2399 --add-label <label>

@valarLip
valarLip merged commit 1be9a86 into main Sep 25, 2026
31 of 34 checks passed
@valarLip
valarLip deleted the perf/unified-h2d-pr branch September 25, 2026 06:29
yhl-amd pushed a commit that referenced this pull request Sep 28, 2026
…ether

Every backend kept `block_regions` and `block_tensor_views` as two lists
paired by hand (DSV4 even collected a third list of sources to zip later).
`KVTransferTensors.add_block_region(tensor, semantic_role=...)` now appends
both from one tensor: the region's addresses and a zero-copy
`uint8 [num_units, 1, unit_bytes]` alias of exactly those bytes, cut from a
larger allocation when `total_bytes` says so.

DSV4, MHA, the MHA draft and MLA all publish through it. MHA and the draft
already published this shape. MLA's views change from `[n, block_size, width]`
to the same byte form; LMCache stores the same bytes in the same order either
way, and neither its object key nor ATOM's model namespace depends on view
shape, so existing cache entries stay valid. The MLA builder test's module
stubs had drifted since #2399 (`atom.utils.block_tables`); they are fixed so
the test runs again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
valarLip pushed a commit that referenced this pull request Sep 28, 2026
* feat(offload): support native DSv4 checkpoints with LMCache MP

Save and restore DSv4 PAGE KV together with the matching recurrent STATE
checkpoint through a standalone LMCache multiprocess server. The lmcache_mp
connector selects its implementation from backend capabilities: backends
that publish a PagedStateCheckpointSpec and execute_paged_state_copies use
the native PAGE/STATE path, others keep the PAGE-only transport. Saves
lease immutable READY checkpoint PAGE units instead of snapshotting the
live SLOT; restores reserve fresh units and adopt them as a local READY
checkpoint.

Also:
- allow single-host DP and DP-attention, scoping MP sessions per replica;
- release source PAGE blocks per chunk and reacquire a finished request's
  still-resident prefix at save admission;
- allow partial release of recurrent-state requests only when the
  connector guarantees an independent state lease;
- read offload env knobs through atom.utils.envs;
- refuse PAGE-copied state backends that lack the native contract.

Requires LMCache dev@05fc77a (LMCache/LMCache#5132), pinned on main by
#2395. Start the server with --null-block-id -1 --separate-object-groups.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(offload): pass --eviction-policy to the LMCache MP server example

The pinned LMCache (05fc77a) requires --eviction-policy on `lmcache server`;
without it the documented command exits 2 during argument parsing.

Reported-by: kvnloo
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): skip native saves shorter than OFFLOAD_MIN_SAVE_TOKENS

Native MP stored every READY checkpoint boundary, including whole short
prompts, although a prefix below OFFLOAD_MIN_LOAD_TOKENS (8192 by default,
equal to OFFLOAD_MIN_SAVE_TOKENS) can never be loaded back. On DSv4-Pro TP8
with a 1K/1K workload at concurrency 64, offload stored 257 such prefixes and
cost 12.4% output throughput; skipping them stores none and leaves 1.6-2.2%,
within run-to-run noise. 16K prompts still save as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(offload): keep early release safe without a BlockManager; one save threshold

Review follow-ups (#2250 §1, §4).

- Unbound schedulers (the vLLM plugin, where vLLM owns the blocks) cannot
  reacquire a finished request's prefix, so teardown again freezes the block
  table and leases the unemitted suffix, and the final save reads exactly
  those leased blocks. The late-save reacquire path now runs only when a
  BlockManager is bound, instead of keying off an empty block table that the
  plugin never clears.
- The shared base no longer reads OFFLOAD_MIN_SAVE_TOKENS, so dense and hybrid
  behave as before this PR. A late save with nothing savable resident now
  retires the request instead of emitting an empty save forever.
- Native MP applies OFFLOAD_MIN_SAVE_TOKENS to the absolute boundary for
  normal and late saves alike, so a long request's short tail is still
  stored, and the floor can never reach boundary 0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): bound unprovable transfers; fence native restores

Review follow-ups (#2250 §3, §7).

- A submission that raises, a restore that raises, or a restore event whose
  query raises may still be touching engine memory, so its lease is kept -- but
  no longer forever. _UncertainSubmission now fails the transfer after
  lmcache.mp.uncertain_transfer_timeout_s (default: twice lmcache.mp.mq_timeout)
  with a warning, so the save or load settles and its lease, budget, pending
  slot and descriptor slot are released. Before, one ZMQ hiccup could stop all
  native saves and loads and keep EngineCore busy-looping.
- PAGE-only MP treats a raising save submission the same way instead of as a
  definite failure. The failure reported no quiescence, so under #2339 the save
  was never retried and its lease was never released. The now-dead
  _immediate_save_failures set is removed.
- Native restores are fenced both ways: the restore stream waits for work
  already on the compute stream (write-after-read on the SLOT's previous
  occupant), and every step's compute stream waits on in-flight restore
  events (read-after-write, including SLOT relocations in build()). The fence
  is on the GPU, so the host does not block.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): end sessions after the last save; unwind failed admissions

Review follow-ups (#2250 §7).

- The MP scheduler ended a request's session in request_finished, but with
  early release its final save is emitted and submitted under that session
  afterwards. Session ends are now deferred until no save of the request is
  tracked or in flight. They are swept on retirement and at every metadata
  build.
- Native save admission returns the checkpoint lease and budget charge if
  anything after the acquire raises; the pin is never timeout-reclaimable, so
  it used to leak for the process lifetime. Load admission computes the
  boundary hash before reserving units, and releases the units, the budget
  and any suspended local restore if the rest raises.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): keep draft PAGE regions out of the native image

Review follow-ups (#2250 §5, §7).

- DSv4 now publishes paged_state_region_count, the number of its own PAGE
  regions. A draft with a pool of its own appends regions after them, and the
  native layout validates checkpoint coverage and builds STATE aliases from
  the leading regions only. Draft rows are registered as ordinary PAGE KV and
  never folded into the checkpoint image. Before, DSv4 with such a draft
  failed registration with "PAGE regions do not cover the native PAGE unit".
- Auto rank collapse is off when a DSpark draft's backend owns a KV pool:
  that draft publishes per-rank PAGE regions, so the complete PAGE object is
  not replicated and registration would otherwise reject the requested
  collapse.
- The native layout registers its PAGE group as uint8 views, like the
  PAGE-only path and like its own STATE aliases. LMCache's ROCm raw-pointer
  fallback cannot express FP8 through the CUDA array interface.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(attention): take the descriptor slot in DSv4.1 paged-state copies

Review follow-up (#2250 §5).

descriptor_slot reached the base class, DSv4 and GDN but not DSv4.1, whose
execute_paged_state_copies(stores, restores) would raise TypeError inside the
native restore's exception handler and be reported as a failed restore.
StateCopies now keeps an independent pinned staging buffer and upload fence
per descriptor slot, like the other backends. A new AST sweep test fails if
any attention backend's implementation lacks the parameter.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(offload): seed native and late-save hash chains from cache_seed

Review follow-up (#2250 §2).

The native boundary-hash fallback and BlockManager.acquire_offload_prefix both
started their chains at -1, while every BlockManager chain starts at
seq.cache_seed, which multimodal requests set. For those requests every
native save was skipped, every restored checkpoint was published under a key
no loader looks up, and late saves found nothing resident. The native
fallback now extends its chain through the new BlockManager.prefix_hash_chain
(the manager's own _chain_to), so both _boundary_hash branches share one
seed, algorithm and token slicing. acquire_offload_prefix seeds from
seq.cache_seed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(offload): fail closed on disabled paged-state checkpoints and composite state release

Review follow-ups (#2250 §6).

- The MP scheduler shell picked the native-state scheduler whenever the
  checkpoint coordinator existed, but the coordinator exists without prefix
  caching and then can never hold a READY image, so offload silently did
  nothing. A PAGE-only fallback would restore KV under stale recurrent state.
  Binding now fails with a message naming --enable-prefix-caching.
- MultiConnectorScheduler.can_partially_deallocate_state now requires every
  still-deferring sub to guarantee state safety. Before, one declaring sub was
  enough, and the state slot could be recycled while a non-declaring sub (e.g.
  a dense PAGE sub) still needed the request alive. The test now separates
  all-of from any-of.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* perf(lmcache_mp): cache the native save frontier; reserve restore slots early

Review follow-ups (#2250 §5, §8).

- Native _save_frontier walked the prompt checkpoint by checkpoint for every
  tracked request on every scheduler step. PageUnitCheckpointStore now has a
  generation that is bumped whenever the READY set can change, and the answer
  is cached per request on (frontier, floor, generation).
- Restore descriptor slots >= 1 were first allocated as pinned memory on the
  connector thread mid-serving. The native worker now reserves them at
  registration through the new builder hook reserve_checkpoint_descriptors
  (base builders allocate their descriptor buffers, DSv4.1 its per-slot
  staging).
- Document why single-host DP replicas share one (model_name, worker_id,
  world_size) identity. It is the content-addressed storage namespace, while
  the server registers GPU memory per instance_id, refcounts layout
  descriptors per (model_name, world_size), and _mp_session_id scopes
  sessions per replica.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(offload): document every offload env var; guard it with a test

Review follow-ups (#2250 §8, §9).

- docs/environment_variables.md listed only OFFLOAD_MAX_PENDING_SAVES and the
  LMCache pin timeout, and still said they bypass atom.utils.envs. All 16
  ATOM-owned offload knobs are now documented with type, default, precedence
  and invalid-value behavior. A new test fails if an OFFLOAD_*/LMCACHE_* var
  registered in envs.py is missing from the reference.
- README and envs.py describe OFFLOAD_MIN_SAVE_TOKENS as the native absolute
  boundary for normal and late saves, and the README documents the
  uncertain-transfer bound.
- Test that a restore raising after it took a descriptor slot fails the load
  and returns the slot once the uncertainty bound expires.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(lmcache_mp): make the raising-restore test independent of CUDA

On a CPU-only torch, torch.cuda.Event is a dummy class that raises when it is
built, which happens before the restore takes its descriptor slot, so the test's
"slot is held until the bound" assertion failed in CI. Stub the Event so the
restore fails deterministically at the stream fence, after the slot is taken,
on every runner.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): release transfer memory only on a report, fail stop at a deadline

The MP path answered "what settles a transfer whose outcome is unprovable"
three ways: never (native checkpoint pins, a zeroed abandon timeout), after
600 s (the uncertainty bound, which then freed load destinations and PAGE
sources a server might still be using), and never-unbounded (a future whose
poll keeps raising). One rule now: memory under an MP transfer is released
only on a terminal report, and a transfer still not terminal after
`lmcache.mp.transfer_deadline_s` (default 1200 s) raises
`LMCacheTransferUnprovable`, stopping the engine.

- Worker: every pending PAGE-only and native load/save carries its start
  time and is checked against the deadline, which bounds a raising
  submission, a future whose poll keeps raising, and a restore whose event
  cannot be queried. A descriptor that cannot be built is provably unsent
  and still fails immediately; only the transport call is unprovable.
- Scheduler: a watchdog over dispatched saves, loads and native checkpoint
  sources catches a completion that never arrives (lost report, silent TP
  rank), with a margin so the worker fails first.
- `save_abandon_timeout_s` is inherited again, so the engine-wide stalled
  save, state-pin and orphan-load-slot reclaimers stay on; MP opts out
  through its own `abandon_save` / `reclaim_stale_leases` instead.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(offload): lease the final save source when prefix caching is off

The late final save reacquired a finished request's prefix through the hash
index whenever a BlockManager was bound. With prefix caching off nothing is
indexed, so the lookup found nothing and the save was dropped without a
warning -- the combination the offload README ships for `lmcache_offload`.
Hash reacquire now requires a bound manager with prefix caching; otherwise
teardown leases the unemitted suffix of the frozen table, as on main. A
late save that finds part of its prefix evicted is counted in
`truncated_late_saves`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(offload): reacquire a final save source only after a partial release

`protected_block_ids` marked the request finished before the scheduler
decided whether it could release it partially. When per-request state made
that unsafe, the request was deferred whole and kept its table, but the next
metadata build still took the late-save path and claimed every block a
second time through the hash index (dropping the save if a hash was
missing). `activate_block_leases`, which only the partial-release branch
calls, now marks the release, and only that mark selects the reacquire path.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): keep the native STATE image pinned until the STORE terminal

The worker released a save's checkpoint image as soon as any PAGE chunk
milestone reached the boundary. Chunk milestones report PAGE token ranges;
nothing in LMCache promises that a chunk completes only after every engine
group registered at it, STATE groups included, has been read. The pinned
LMCache emits no chunk events, so this path never ran, but it would have
been unsound on the first server that did. The STATE source-safe channel is
removed: the image pin is released at the STORE terminal, and chunk
milestones release PAGE leases only.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(envs): read an empty offload knob as its default everywhere

`OFFLOAD_PUBLICATION_TIMEOUT_S=` made `float('')` raise out of connector
init, and `OFFLOAD_COPY_WORKERS` / `OFFLOAD_LOAD_WORKERS` died with a bare
`invalid literal for int()`. The offload section now states one policy:
unset or empty is the default; a set but unusable value warns and falls back
for knobs that only tune reuse, and is rejected at startup, naming the
variable, for widths, sizes and timeouts. The transfer-mode comment no
longer lists `engine_driven` as valid.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): keep failed adoptions releasable and surface broken invariants

- `adopt_transfer_units` popped the transfer record before adopting; if
  adoption raised, the units were reserved but unreachable by any release
  path. The record is now removed only after adoption returns.
- A native restore that loses the publish race to an identical image was
  dropped silently; it is now logged as a deduplication.
- The worker's own pending-save refusal is unreachable while the scheduler
  enforces the same bound; if it ever fires it now logs an error instead of
  passing as an ordinary save failure.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(lmcache_mp): run a native restore end to end and fix drifted copy doubles

The restore-copy doubles were positional-only, so the production call with
`descriptor_slot=` would have raised inside `_begin_restore` and been
reported as a failed restore; the success tests replaced `_begin_restore`
outright, so no test ran the happy path. The doubles now take the production
signature, and a new test drives a terminal retrieve through `get_finished`
into the real `_begin_restore`, then asserts the copy's descriptor slot, that
the load finishes only once the restore event is done, and that the slot is
returned. The CUDA stand-ins are shared with the fence test.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* perf(lmcache_mp): invalidate a request's save frontier only on its own checkpoints

The per-request `_save_frontier` memo was keyed on the store-wide
generation, so any request's publish or eviction forced every tracked
request to rescan its prompt on the next step. The store now keeps a
bounded log of which prefix hash each generation bump touched; a request
rescans only when a change hits one of the boundaries it scanned (its answer
or anything above it), or when the log cannot tell.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(attentions): one fenced descriptor staging for every checkpoint copier

DSv4.1 kept its own `_DescriptorStaging`, a third copy of the per-slot
pinned descriptor pool the base builder shares with DSv4 and GDN, and the
only one that fenced reuse: a non-blocking H2D reads its pinned rows when
the stream reaches it, so refilling a buffer whose last upload is still
queued rewrites the descriptor under an earlier copy. `DescriptorStaging`
in `pool_layout/paged_state_copy.py` now owns the per-slot buffers and the
fence, and the base builder (DSv4, GDN) and DSv4.1's `StateCopies` both use
it, so the base path gains the fence too.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(lmcache_mp): validate PAGE views once for both registrations

`build_native_state_mp_layout` and `_build_cache_views` each validated the
backend-published PAGE views, and had already drifted: PAGE-only required
full contiguity and compared total bytes against the view, native checked
inner contiguity plus the block stride against the region. Both now call
`validate_page_views` (new `mp/page_views.py`), which applies the stricter
union -- shape, tight block-major stride, unit and total bytes, aliasing,
forward indexing, one device -- and, for native, that a block's physical
slots divide the block size. Each builder keeps only what is its own:
PAGE-only rejects stateful fields; native checks the image coverage.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(lmcache_mp): split the native worker's completion poll

`get_finished` was 91 lines at six levels of nesting, with three
near-identical recovery blocks. It is now a loop over `_poll_native_save`
and `_poll_native_load`; the load's retrieve-then-restore progression lives
in `_advance_native_load`, and both unprovable-restore paths share
`_hold_unprovable_restore`. No behaviour change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(lmcache_mp): split mp/backend.py by responsibility

`mp/backend.py` had grown to 1332 lines holding configuration, PAGE view
validation, transfer bookkeeping, lookups and both connector halves. It is
split without behaviour change into:

- `deployment.py`: config validation, TP/DP topology and rank collapse,
  model namespace, server adapters
- `transfer.py`: operation identity, terminal detection, the transfer
  deadline and `LMCacheTransferUnprovable`
- `lookup.py`: the scheduler's lookup client and read-lock bookkeeping
- `page_views.py`: PAGE-only `_build_cache_views`, beside the shared view
  validation it uses
- `worker.py` / `scheduler.py`: the PAGE-only connector halves

The native modules, the public shell and the tests import from the new
homes; tests patch each name where it is looked up.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(offload): keep clock reclaim off MP requests; warn on truncated late saves

- With the abandon window inherited again, a slow but legitimate MP save
  deferred past it was "abandoned" (a no-op for MP) and, a minute later,
  logged as a wedged P/D send -- both long before MP's own transfer
  deadline. A connector can now answer `waits_for_transfer_report(seq)`;
  the stalled-save reclaim skips such requests. LMCache MP answers True,
  the in-process shell forwards (default False), and the composite answers
  True only if every sub still deferring the request does, so P/D sends and
  in-process saves keep their abandon path.
- A late save that finds part of its prefix evicted now logs a warning with
  the running count, not a debug line: early release returns the tail's
  blocks at teardown, and this is the cost of that trade.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(kv_transfer): publish each PAGE region and its byte view together

Every backend kept `block_regions` and `block_tensor_views` as two lists
paired by hand (DSV4 even collected a third list of sources to zip later).
`KVTransferTensors.add_block_region(tensor, semantic_role=...)` now appends
both from one tensor: the region's addresses and a zero-copy
`uint8 [num_units, 1, unit_bytes]` alias of exactly those bytes, cut from a
larger allocation when `total_bytes` says so.

DSV4, MHA, the MHA draft and MLA all publish through it. MHA and the draft
already published this shape. MLA's views change from `[n, block_size, width]`
to the same byte form; LMCache stores the same bytes in the same order either
way, and neither its object key nor ATOM's model namespace depends on view
shape, so existing cache entries stay valid. The MLA builder test's module
stubs had drifted since #2399 (`atom.utils.block_tables`); they are fixed so
the test runs again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(kv_transfer): make a PAGE unit one value, its region and view together

`add_block_region` paired a region and its view at one call site, but the
pair was only a convention: both lists stayed public and mutable, the
constructor still took them separately, DSV4 and MLA used a throwaway
`KVTransferTensors` as a scratch builder, the draft merge extended the two
lists by hand, and tensor code lived in the torch-free `types` contract.

- `PageRegion(region, view)` is a frozen value; `KVTransferTensors.pages`
  holds them. `block_regions` and `block_tensor_views` are read-only tuples
  derived from it, so no code can add to one without the other.
- `page_region(tensor, ...)` in the new `disaggregation/page_region.py`
  builds one from the owning tensor (zero-copy byte view, contiguity and
  size checks); `types.py` no longer touches torch.
- Builders pass `pages=[...]` straight to the final object: no scratch
  instance. The address-only producer (Qwen4 exp) publishes
  `PageRegion(region)` with no view, which LMCache MP refuses as before.
- `merge_pages(other)` replaces the draft merge in `ModelRunner`, carrying
  the gcd replication rule with it.

`validate_page_views` stays on the LMCache MP side: it is the consumer
checking what it is handed, not a second copy of construction.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Honglie Yi <hyi@crsuse2-m2m-v2-020.us-east2-a.compute.internal>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
jiayyu added a commit that referenced this pull request Sep 29, 2026
Squash of the 33 commits of fpz/mixed_mla_dispatch_v4 (history preserved in
fpz/mixed_mla_dispatch_v4_backup_20260929), rebased onto main e59d422.

Ported onto main's checked H2D publication (#2399):
- The mixed prefill bank is now a bound publication slot with its own owner,
  groups and event (ModelRunner.build_h2d_slot), one per forward_vars ring
  slot. Swapping forward_vars alone made the prefill producers publish the
  runner's (stale) sources through h2d_groups.
- cu_seqlens_q has one writer per buffer per epoch: publish_cu_seqlens_q
  publishes the decode segment's spans on a mixed step, and
  mixed_prefill_bank_active the prefill segment's into the bank. prepare_mixed
  no longer re-uploads it.
- Dense decode block tables go through block_table_state; V4 merges positions
  on the device instead of re-staging a published buffer.
- Mixed deferred input ids use fill_deferred_decode_ids over the decode region.
- The TBO prefill-segment rebuild swaps owner and groups with the bank.
- _flash_attn_prefill takes `out` (FlyDSL FP8 copies into it); the chunk
  helpers take the decode-first prefill budget; _settle_prefill_chunks ends
  a chunk at num_tokens, per main.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant