Skip to content

ci(lmcache): ship LMCache wheels as Docker Hub images, not GitHub Releases - #2395

Merged
valarLip merged 6 commits into
mainfrom
ci/lmcache-wheel-dockerhub-main
Sep 24, 2026
Merged

valarLip merged 6 commits into
mainfrom
ci/lmcache-wheel-dockerhub-main

Conversation

@yhl-amd

@yhl-amd yhl-amd commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Why

lmcache-rocm-wheel.yaml (#2388) published wheels as GitHub Releases targeting GITHUB_SHA. The lmcache-v0.5.6.dev98-g05fc77a0-rocm-torch210 tag it created on main was picked up by setuptools-scm's git describe and could not be parsed, so pip install -e . failed in every Pre Checkin run (e.g. https://github.com/ROCm/ATOM/actions/runs/35996699729). That tag and release have been deleted. Until this merges, dispatching the workflow from main with publish=true would do the same thing again. After it, the workflow creates no git tags or releases; the only tag creators left are ATOM's own vX.Y.Z release workflows, so no git describe filter is added.

What

1. lmcache-rocm-wheel.yaml: publish to Docker Hub, not Releases

  • publish pushes a FROM scratch image holding only the wheel, as rocm/atom-dev:lmcache-v<version>-g<sha8>-rocm-torch210, and outputs its digest. It uses the repo's Docker Hub credentials (scoped to the push step) via docker-auth + docker_push_retry.sh. gh release is gone, and the job has contents: read only.
  • Overwrite guard: only no such manifest counts as the tag being free. Rate limits, 5xx and network errors fail the job.
  • check-inputs job: requires a full 40-hex lmcache_commit on every path, requires source_run_id to be numeric, and rejects source_run_id without publish.
  • source_run_id: publishes the wheel an earlier run built. That run must be this workflow, on the default or current branch, with a successful build job in the latest unexpired attempt. Without it, publish uses the build's own attempt, so "Re-run failed jobs" works.
  • bump-pr (from ci: open a pin-bump PR after publishing an LMCache wheel #2391, via feat(offload): support native DSv4 checkpoints with LMCache MP #2250): after publishing, opens a draft PR that moves the Dockerfile pin with bump_lmcache_wheel_pin.py and lists Requires-Dist changes.

2. Dockerfiles: install LMCache from the digest-pinned wheel image

docker/Dockerfile and docker/atom_release.dockerfile move from LMCache's 0.5.5rc3 release wheel to 0.5.6.dev98 (dev@05fc77a, contains LMCache/LMCache#5132):

ARG LMCACHE_WHEEL_IMAGE="rocm/atom-dev:lmcache-v0.5.6.dev98-g05fc77a0-rocm-torch210@sha256:d3cfe74f…"
FROM ${LMCACHE_WHEEL_IMAGE} AS lmcache_wheel
...
ARG LMCACHE_WHEEL_SHA256=a5fe8f3f…
COPY --from=lmcache_wheel / /tmp/lmcache-wheel/

The sha256 check, --no-deps install and import validation are kept, and lmcache.__version__ is asserted against the wheel name. grpcio>=1.78.0 and protobuf>=6.31.1,<7 are added for the newer wheel. This supersedes #2394, which made the same change on #2250's branch.

Validation

  • Pre Checkin (non-GPU unit tests, Black, Ruff), workflow lint, and the offline inference smoke test pass on this PR. The smoke test's first attempt died on a Docker Hub token expiring mid-pull (unauthorized); the rerun passed.
  • The wheel image was published by https://github.com/ROCm/ATOM/actions/runs/36002407218 from run 35984043135's artifact. An anonymous pull contains exactly that wheel, with sha256 a5fe8f3f….
  • https://github.com/ROCm/ATOM/actions/runs/36009420936 ran the fixed publish path end to end: it resolved the artifact, checked the sha256, then refused to overwrite the existing tag.
  • LMCache 0.5.5rc3 → 0.5.6.dev98 on main's code: in rocm/atom-dev:latest (2026-09-23 nightly, MI355X), I ran all 72 test files that touch lmcache/kv_transfer/offload with this branch's source, once on the image's 0.5.5rc3 and once after installing the new wheel from the image above. Both runs gave 2375 passed. Both had the same 10 failures (vLLM plugin layout/SeqView and scheduler metrics), which are already present on 0.5.5rc3 because the nightly lags main.
  • The Dockerfile's LMCache RUN step was run verbatim in that image: OK: lmcache 0.5.6.dev98+g05fc77a0… HIP cuda_ops.
  • actionlint 1.7.7 clean.
  • Not run: a full image build, and serving or accuracy with LMCache offload enabled on the new LMCache.

🤖 Generated with Claude Code

A GitHub Release needs a git tag. The lmcache-v...-rocm-torch210 tag
created on main was picked up by setuptools-scm's git describe and broke
`pip install -e .` in every Pre Checkin run. Push the wheel as a
FROM scratch image, rocm/atom-dev:lmcache-v<version>-g<sha8>-rocm-torch210,
instead; nothing appears under Releases or Tags, and the publish job no
longer has contents: write.

source_run_id publishes the wheel an earlier run built instead of
rebuilding it.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every eligible PR before approval:

  • ✅ Pre Checkin: Black, Ruff, catalog schema validation, non-GPU unit tests

Heavy model tests:

  • ✅ Run after the PR is approved and Pre Checkin passes
  • ✅ Run immediately when an approval review is submitted
  • ✅ Can be requested before approval with labels
Label Tests
ci:full Run all heavy PR model tests: native ATOM, vLLM, and SGLang
ci:atom Run native ATOM model accuracy tests
ci:vllm Run ATOM vLLM OOT model accuracy tests
ci:sglang Run ATOM SGLang model accuracy tests

Heavy jobs are skipped when the PR is not approved and no matching ci:* label is present.
Add labels via the sidebar or gh pr edit 2395 --add-label <label>

The lmcache-v0.5.6.dev98-g05fc77a0-rocm-torch210 release tag is now the
nearest tag reachable from main, and setuptools_scm cannot parse it as a
PEP 440 version, so `pip install -e .` fails in the Pre Checkin CI.

Restrict git describe to tags matching v[0-9]* so non-ATOM release tags
are ignored while the ATOM version is still derived from vX.Y.Z tags.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@yhl-amd yhl-amd changed the title ci(lmcache): publish LMCache wheels as Docker Hub images instead of GitHub Releases fix(build): keep non-ATOM tags out of the version; publish LMCache wheels to Docker Hub Sep 24, 2026
- atom-release.yaml: git describe only matches v[0-9]* tags, like
  setuptools-scm now does, so a stray tag cannot become the release
  version.
- check-inputs job validates lmcache_commit (also on the source_run_id
  path), source_run_id's format, and rejects source_run_id without
  publish instead of finishing green having done nothing.
- publish resolves exactly one artifact: the build's own attempt (so
  "Re-run failed jobs" finds it), or for source_run_id the latest
  unexpired attempt of a run of this workflow on the default or current
  branch whose build job succeeded.
- Only "no such manifest" counts as the tag being free; rate limits,
  5xx and network errors now fail instead of allowing an overwrite.
- Docker Hub credentials are scoped to the push step.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
yhl-amd pushed a commit that referenced this pull request Sep 24, 2026
Same fixes as #2395: check-inputs job (full commit SHA on every path,
numeric source_run_id, source_run_id requires publish), exactly one
artifact resolved per publish (build's own attempt, or the latest
unexpired attempt of a successful build of this workflow on the default
or current branch), only "no such manifest" allows a push, and Docker
Hub credentials scoped to the push step.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Move main's images from LMCache's 0.5.5rc3 release wheel to the
0.5.6.dev98 (dev@05fc77a) wheel published by lmcache-rocm-wheel.yaml,
pulled from rocm/atom-dev by digest via a FROM scratch stage and
checked by sha256. The workflow's bump-pr job and
bump_lmcache_wheel_pin.py move this pin from now on. Adds grpcio and
protobuf, which the newer wheel requires.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@yhl-amd yhl-amd changed the title fix(build): keep non-ATOM tags out of the version; publish LMCache wheels to Docker Hub fix(build): keep non-ATOM tags out of the version; ship LMCache wheels as Docker Hub images Sep 24, 2026
The LMCache wheel workflow no longer creates tags, and the only
remaining tag creators are ATOM's own vX.Y.Z releases, so drop the
--match 'v[0-9]*' filters from pyproject.toml and atom-release.yaml.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@yhl-amd yhl-amd changed the title fix(build): keep non-ATOM tags out of the version; ship LMCache wheels as Docker Hub images ci(lmcache): ship LMCache wheels as Docker Hub images, not GitHub Releases Sep 24, 2026
- bump-pr gets actions: read; download-artifact with run-id reads the
  artifact through the Actions API and failed without it.
- bump-pr can be re-run: it force-pushes its own per-wheel branch and
  reuses an open PR instead of failing on the existing branch.
- source_run_id: the source run's build job name carries the full
  commit it was dispatched with; require it to equal lmcache_commit
  (the wheel name only has 8 hex digits).
- Read the pushed digest from the registry (buildx imagetools) rather
  than RepoDigests, which differs under the containerd image store.
- The summary only names the build image for a build in this run.
- Dockerfiles: bind-mount the wheel stage in the install RUN instead of
  COPY + rm, so no layer keeps the wheel.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@valarLip
valarLip merged commit a5ad0a5 into main Sep 24, 2026
24 of 25 checks passed
@valarLip
valarLip deleted the ci/lmcache-wheel-dockerhub-main branch September 24, 2026 15:13
yhl-amd pushed a commit that referenced this pull request Sep 24, 2026
Take main's LMCache wheel pipeline from #2395 for the bump script, the
wheel workflow and both Dockerfiles. The GitHub Release this branch pinned
was deleted by #2395; main pins the same dev@05fc77a wheel (sha256
a5fe8f3f...) as a digest-pinned Docker Hub image, so this branch no longer
changes those files.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
yhl-amd pushed a commit that referenced this pull request Sep 25, 2026
Save and restore DSv4 PAGE KV together with the matching recurrent STATE
checkpoint through a standalone LMCache multiprocess server. The lmcache_mp
connector selects its implementation from backend capabilities: backends
that publish a PagedStateCheckpointSpec and execute_paged_state_copies use
the native PAGE/STATE path, others keep the PAGE-only transport. Saves
lease immutable READY checkpoint PAGE units instead of snapshotting the
live SLOT; restores reserve fresh units and adopt them as a local READY
checkpoint.

Also:
- allow single-host DP and DP-attention, scoping MP sessions per replica;
- release source PAGE blocks per chunk and reacquire a finished request's
  still-resident prefix at save admission;
- allow partial release of recurrent-state requests only when the
  connector guarantees an independent state lease;
- read offload env knobs through atom.utils.envs;
- refuse PAGE-copied state backends that lack the native contract.

Requires LMCache dev@05fc77a (LMCache/LMCache#5132), pinned on main by
#2395. Start the server with --null-block-id -1 --separate-object-groups.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
valarLip pushed a commit that referenced this pull request Sep 28, 2026
* feat(offload): support native DSv4 checkpoints with LMCache MP

Save and restore DSv4 PAGE KV together with the matching recurrent STATE
checkpoint through a standalone LMCache multiprocess server. The lmcache_mp
connector selects its implementation from backend capabilities: backends
that publish a PagedStateCheckpointSpec and execute_paged_state_copies use
the native PAGE/STATE path, others keep the PAGE-only transport. Saves
lease immutable READY checkpoint PAGE units instead of snapshotting the
live SLOT; restores reserve fresh units and adopt them as a local READY
checkpoint.

Also:
- allow single-host DP and DP-attention, scoping MP sessions per replica;
- release source PAGE blocks per chunk and reacquire a finished request's
  still-resident prefix at save admission;
- allow partial release of recurrent-state requests only when the
  connector guarantees an independent state lease;
- read offload env knobs through atom.utils.envs;
- refuse PAGE-copied state backends that lack the native contract.

Requires LMCache dev@05fc77a (LMCache/LMCache#5132), pinned on main by
#2395. Start the server with --null-block-id -1 --separate-object-groups.

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(offload): pass --eviction-policy to the LMCache MP server example

The pinned LMCache (05fc77a) requires --eviction-policy on `lmcache server`;
without it the documented command exits 2 during argument parsing.

Reported-by: kvnloo
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): skip native saves shorter than OFFLOAD_MIN_SAVE_TOKENS

Native MP stored every READY checkpoint boundary, including whole short
prompts, although a prefix below OFFLOAD_MIN_LOAD_TOKENS (8192 by default,
equal to OFFLOAD_MIN_SAVE_TOKENS) can never be loaded back. On DSv4-Pro TP8
with a 1K/1K workload at concurrency 64, offload stored 257 such prefixes and
cost 12.4% output throughput; skipping them stores none and leaves 1.6-2.2%,
within run-to-run noise. 16K prompts still save as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(offload): keep early release safe without a BlockManager; one save threshold

Review follow-ups (#2250 §1, §4).

- Unbound schedulers (the vLLM plugin, where vLLM owns the blocks) cannot
  reacquire a finished request's prefix, so teardown again freezes the block
  table and leases the unemitted suffix, and the final save reads exactly
  those leased blocks. The late-save reacquire path now runs only when a
  BlockManager is bound, instead of keying off an empty block table that the
  plugin never clears.
- The shared base no longer reads OFFLOAD_MIN_SAVE_TOKENS, so dense and hybrid
  behave as before this PR. A late save with nothing savable resident now
  retires the request instead of emitting an empty save forever.
- Native MP applies OFFLOAD_MIN_SAVE_TOKENS to the absolute boundary for
  normal and late saves alike, so a long request's short tail is still
  stored, and the floor can never reach boundary 0.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): bound unprovable transfers; fence native restores

Review follow-ups (#2250 §3, §7).

- A submission that raises, a restore that raises, or a restore event whose
  query raises may still be touching engine memory, so its lease is kept -- but
  no longer forever. _UncertainSubmission now fails the transfer after
  lmcache.mp.uncertain_transfer_timeout_s (default: twice lmcache.mp.mq_timeout)
  with a warning, so the save or load settles and its lease, budget, pending
  slot and descriptor slot are released. Before, one ZMQ hiccup could stop all
  native saves and loads and keep EngineCore busy-looping.
- PAGE-only MP treats a raising save submission the same way instead of as a
  definite failure. The failure reported no quiescence, so under #2339 the save
  was never retried and its lease was never released. The now-dead
  _immediate_save_failures set is removed.
- Native restores are fenced both ways: the restore stream waits for work
  already on the compute stream (write-after-read on the SLOT's previous
  occupant), and every step's compute stream waits on in-flight restore
  events (read-after-write, including SLOT relocations in build()). The fence
  is on the GPU, so the host does not block.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): end sessions after the last save; unwind failed admissions

Review follow-ups (#2250 §7).

- The MP scheduler ended a request's session in request_finished, but with
  early release its final save is emitted and submitted under that session
  afterwards. Session ends are now deferred until no save of the request is
  tracked or in flight. They are swept on retirement and at every metadata
  build.
- Native save admission returns the checkpoint lease and budget charge if
  anything after the acquire raises; the pin is never timeout-reclaimable, so
  it used to leak for the process lifetime. Load admission computes the
  boundary hash before reserving units, and releases the units, the budget
  and any suspended local restore if the rest raises.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): keep draft PAGE regions out of the native image

Review follow-ups (#2250 §5, §7).

- DSv4 now publishes paged_state_region_count, the number of its own PAGE
  regions. A draft with a pool of its own appends regions after them, and the
  native layout validates checkpoint coverage and builds STATE aliases from
  the leading regions only. Draft rows are registered as ordinary PAGE KV and
  never folded into the checkpoint image. Before, DSv4 with such a draft
  failed registration with "PAGE regions do not cover the native PAGE unit".
- Auto rank collapse is off when a DSpark draft's backend owns a KV pool:
  that draft publishes per-rank PAGE regions, so the complete PAGE object is
  not replicated and registration would otherwise reject the requested
  collapse.
- The native layout registers its PAGE group as uint8 views, like the
  PAGE-only path and like its own STATE aliases. LMCache's ROCm raw-pointer
  fallback cannot express FP8 through the CUDA array interface.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(attention): take the descriptor slot in DSv4.1 paged-state copies

Review follow-up (#2250 §5).

descriptor_slot reached the base class, DSv4 and GDN but not DSv4.1, whose
execute_paged_state_copies(stores, restores) would raise TypeError inside the
native restore's exception handler and be reported as a failed restore.
StateCopies now keeps an independent pinned staging buffer and upload fence
per descriptor slot, like the other backends. A new AST sweep test fails if
any attention backend's implementation lacks the parameter.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(offload): seed native and late-save hash chains from cache_seed

Review follow-up (#2250 §2).

The native boundary-hash fallback and BlockManager.acquire_offload_prefix both
started their chains at -1, while every BlockManager chain starts at
seq.cache_seed, which multimodal requests set. For those requests every
native save was skipped, every restored checkpoint was published under a key
no loader looks up, and late saves found nothing resident. The native
fallback now extends its chain through the new BlockManager.prefix_hash_chain
(the manager's own _chain_to), so both _boundary_hash branches share one
seed, algorithm and token slicing. acquire_offload_prefix seeds from
seq.cache_seed.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(offload): fail closed on disabled paged-state checkpoints and composite state release

Review follow-ups (#2250 §6).

- The MP scheduler shell picked the native-state scheduler whenever the
  checkpoint coordinator existed, but the coordinator exists without prefix
  caching and then can never hold a READY image, so offload silently did
  nothing. A PAGE-only fallback would restore KV under stale recurrent state.
  Binding now fails with a message naming --enable-prefix-caching.
- MultiConnectorScheduler.can_partially_deallocate_state now requires every
  still-deferring sub to guarantee state safety. Before, one declaring sub was
  enough, and the state slot could be recycled while a non-declaring sub (e.g.
  a dense PAGE sub) still needed the request alive. The test now separates
  all-of from any-of.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* perf(lmcache_mp): cache the native save frontier; reserve restore slots early

Review follow-ups (#2250 §5, §8).

- Native _save_frontier walked the prompt checkpoint by checkpoint for every
  tracked request on every scheduler step. PageUnitCheckpointStore now has a
  generation that is bumped whenever the READY set can change, and the answer
  is cached per request on (frontier, floor, generation).
- Restore descriptor slots >= 1 were first allocated as pinned memory on the
  connector thread mid-serving. The native worker now reserves them at
  registration through the new builder hook reserve_checkpoint_descriptors
  (base builders allocate their descriptor buffers, DSv4.1 its per-slot
  staging).
- Document why single-host DP replicas share one (model_name, worker_id,
  world_size) identity. It is the content-addressed storage namespace, while
  the server registers GPU memory per instance_id, refcounts layout
  descriptors per (model_name, world_size), and _mp_session_id scopes
  sessions per replica.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* docs(offload): document every offload env var; guard it with a test

Review follow-ups (#2250 §8, §9).

- docs/environment_variables.md listed only OFFLOAD_MAX_PENDING_SAVES and the
  LMCache pin timeout, and still said they bypass atom.utils.envs. All 16
  ATOM-owned offload knobs are now documented with type, default, precedence
  and invalid-value behavior. A new test fails if an OFFLOAD_*/LMCACHE_* var
  registered in envs.py is missing from the reference.
- README and envs.py describe OFFLOAD_MIN_SAVE_TOKENS as the native absolute
  boundary for normal and late saves, and the README documents the
  uncertain-transfer bound.
- Test that a restore raising after it took a descriptor slot fails the load
  and returns the slot once the uncertainty bound expires.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(lmcache_mp): make the raising-restore test independent of CUDA

On a CPU-only torch, torch.cuda.Event is a dummy class that raises when it is
built, which happens before the restore takes its descriptor slot, so the test's
"slot is held until the bound" assertion failed in CI. Stub the Event so the
restore fails deterministically at the stream fence, after the slot is taken,
on every runner.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): release transfer memory only on a report, fail stop at a deadline

The MP path answered "what settles a transfer whose outcome is unprovable"
three ways: never (native checkpoint pins, a zeroed abandon timeout), after
600 s (the uncertainty bound, which then freed load destinations and PAGE
sources a server might still be using), and never-unbounded (a future whose
poll keeps raising). One rule now: memory under an MP transfer is released
only on a terminal report, and a transfer still not terminal after
`lmcache.mp.transfer_deadline_s` (default 1200 s) raises
`LMCacheTransferUnprovable`, stopping the engine.

- Worker: every pending PAGE-only and native load/save carries its start
  time and is checked against the deadline, which bounds a raising
  submission, a future whose poll keeps raising, and a restore whose event
  cannot be queried. A descriptor that cannot be built is provably unsent
  and still fails immediately; only the transport call is unprovable.
- Scheduler: a watchdog over dispatched saves, loads and native checkpoint
  sources catches a completion that never arrives (lost report, silent TP
  rank), with a margin so the worker fails first.
- `save_abandon_timeout_s` is inherited again, so the engine-wide stalled
  save, state-pin and orphan-load-slot reclaimers stay on; MP opts out
  through its own `abandon_save` / `reclaim_stale_leases` instead.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(offload): lease the final save source when prefix caching is off

The late final save reacquired a finished request's prefix through the hash
index whenever a BlockManager was bound. With prefix caching off nothing is
indexed, so the lookup found nothing and the save was dropped without a
warning -- the combination the offload README ships for `lmcache_offload`.
Hash reacquire now requires a bound manager with prefix caching; otherwise
teardown leases the unemitted suffix of the frozen table, as on main. A
late save that finds part of its prefix evicted is counted in
`truncated_late_saves`.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(offload): reacquire a final save source only after a partial release

`protected_block_ids` marked the request finished before the scheduler
decided whether it could release it partially. When per-request state made
that unsafe, the request was deferred whole and kept its table, but the next
metadata build still took the late-save path and claimed every block a
second time through the hash index (dropping the save if a hash was
missing). `activate_block_leases`, which only the partial-release branch
calls, now marks the release, and only that mark selects the reacquire path.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): keep the native STATE image pinned until the STORE terminal

The worker released a save's checkpoint image as soon as any PAGE chunk
milestone reached the boundary. Chunk milestones report PAGE token ranges;
nothing in LMCache promises that a chunk completes only after every engine
group registered at it, STATE groups included, has been read. The pinned
LMCache emits no chunk events, so this path never ran, but it would have
been unsound on the first server that did. The STATE source-safe channel is
removed: the image pin is released at the STORE terminal, and chunk
milestones release PAGE leases only.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(envs): read an empty offload knob as its default everywhere

`OFFLOAD_PUBLICATION_TIMEOUT_S=` made `float('')` raise out of connector
init, and `OFFLOAD_COPY_WORKERS` / `OFFLOAD_LOAD_WORKERS` died with a bare
`invalid literal for int()`. The offload section now states one policy:
unset or empty is the default; a set but unusable value warns and falls back
for knobs that only tune reuse, and is rejected at startup, naming the
variable, for widths, sizes and timeouts. The transfer-mode comment no
longer lists `engine_driven` as valid.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(lmcache_mp): keep failed adoptions releasable and surface broken invariants

- `adopt_transfer_units` popped the transfer record before adopting; if
  adoption raised, the units were reserved but unreachable by any release
  path. The record is now removed only after adoption returns.
- A native restore that loses the publish race to an identical image was
  dropped silently; it is now logged as a deduplication.
- The worker's own pending-save refusal is unreachable while the scheduler
  enforces the same bound; if it ever fires it now logs an error instead of
  passing as an ordinary save failure.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(lmcache_mp): run a native restore end to end and fix drifted copy doubles

The restore-copy doubles were positional-only, so the production call with
`descriptor_slot=` would have raised inside `_begin_restore` and been
reported as a failed restore; the success tests replaced `_begin_restore`
outright, so no test ran the happy path. The doubles now take the production
signature, and a new test drives a terminal retrieve through `get_finished`
into the real `_begin_restore`, then asserts the copy's descriptor slot, that
the load finishes only once the restore event is done, and that the slot is
returned. The CUDA stand-ins are shared with the fence test.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* perf(lmcache_mp): invalidate a request's save frontier only on its own checkpoints

The per-request `_save_frontier` memo was keyed on the store-wide
generation, so any request's publish or eviction forced every tracked
request to rescan its prompt on the next step. The store now keeps a
bounded log of which prefix hash each generation bump touched; a request
rescans only when a change hits one of the boundaries it scanned (its answer
or anything above it), or when the log cannot tell.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(attentions): one fenced descriptor staging for every checkpoint copier

DSv4.1 kept its own `_DescriptorStaging`, a third copy of the per-slot
pinned descriptor pool the base builder shares with DSv4 and GDN, and the
only one that fenced reuse: a non-blocking H2D reads its pinned rows when
the stream reaches it, so refilling a buffer whose last upload is still
queued rewrites the descriptor under an earlier copy. `DescriptorStaging`
in `pool_layout/paged_state_copy.py` now owns the per-slot buffers and the
fence, and the base builder (DSv4, GDN) and DSv4.1's `StateCopies` both use
it, so the base path gains the fence too.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(lmcache_mp): validate PAGE views once for both registrations

`build_native_state_mp_layout` and `_build_cache_views` each validated the
backend-published PAGE views, and had already drifted: PAGE-only required
full contiguity and compared total bytes against the view, native checked
inner contiguity plus the block stride against the region. Both now call
`validate_page_views` (new `mp/page_views.py`), which applies the stricter
union -- shape, tight block-major stride, unit and total bytes, aliasing,
forward indexing, one device -- and, for native, that a block's physical
slots divide the block size. Each builder keeps only what is its own:
PAGE-only rejects stateful fields; native checks the image coverage.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(lmcache_mp): split the native worker's completion poll

`get_finished` was 91 lines at six levels of nesting, with three
near-identical recovery blocks. It is now a loop over `_poll_native_save`
and `_poll_native_load`; the load's retrieve-then-restore progression lives
in `_advance_native_load`, and both unprovable-restore paths share
`_hold_unprovable_restore`. No behaviour change.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(lmcache_mp): split mp/backend.py by responsibility

`mp/backend.py` had grown to 1332 lines holding configuration, PAGE view
validation, transfer bookkeeping, lookups and both connector halves. It is
split without behaviour change into:

- `deployment.py`: config validation, TP/DP topology and rank collapse,
  model namespace, server adapters
- `transfer.py`: operation identity, terminal detection, the transfer
  deadline and `LMCacheTransferUnprovable`
- `lookup.py`: the scheduler's lookup client and read-lock bookkeeping
- `page_views.py`: PAGE-only `_build_cache_views`, beside the shared view
  validation it uses
- `worker.py` / `scheduler.py`: the PAGE-only connector halves

The native modules, the public shell and the tests import from the new
homes; tests patch each name where it is looked up.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(offload): keep clock reclaim off MP requests; warn on truncated late saves

- With the abandon window inherited again, a slow but legitimate MP save
  deferred past it was "abandoned" (a no-op for MP) and, a minute later,
  logged as a wedged P/D send -- both long before MP's own transfer
  deadline. A connector can now answer `waits_for_transfer_report(seq)`;
  the stalled-save reclaim skips such requests. LMCache MP answers True,
  the in-process shell forwards (default False), and the composite answers
  True only if every sub still deferring the request does, so P/D sends and
  in-process saves keep their abandon path.
- A late save that finds part of its prefix evicted now logs a warning with
  the running count, not a debug line: early release returns the tail's
  blocks at teardown, and this is the cost of that trade.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(kv_transfer): publish each PAGE region and its byte view together

Every backend kept `block_regions` and `block_tensor_views` as two lists
paired by hand (DSV4 even collected a third list of sources to zip later).
`KVTransferTensors.add_block_region(tensor, semantic_role=...)` now appends
both from one tensor: the region's addresses and a zero-copy
`uint8 [num_units, 1, unit_bytes]` alias of exactly those bytes, cut from a
larger allocation when `total_bytes` says so.

DSV4, MHA, the MHA draft and MLA all publish through it. MHA and the draft
already published this shape. MLA's views change from `[n, block_size, width]`
to the same byte form; LMCache stores the same bytes in the same order either
way, and neither its object key nor ATOM's model namespace depends on view
shape, so existing cache entries stay valid. The MLA builder test's module
stubs had drifted since #2399 (`atom.utils.block_tables`); they are fixed so
the test runs again.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(kv_transfer): make a PAGE unit one value, its region and view together

`add_block_region` paired a region and its view at one call site, but the
pair was only a convention: both lists stayed public and mutable, the
constructor still took them separately, DSV4 and MLA used a throwaway
`KVTransferTensors` as a scratch builder, the draft merge extended the two
lists by hand, and tensor code lived in the torch-free `types` contract.

- `PageRegion(region, view)` is a frozen value; `KVTransferTensors.pages`
  holds them. `block_regions` and `block_tensor_views` are read-only tuples
  derived from it, so no code can add to one without the other.
- `page_region(tensor, ...)` in the new `disaggregation/page_region.py`
  builds one from the owning tensor (zero-copy byte view, contiguity and
  size checks); `types.py` no longer touches torch.
- Builders pass `pages=[...]` straight to the final object: no scratch
  instance. The address-only producer (Qwen4 exp) publishes
  `PageRegion(region)` with no view, which LMCache MP refuses as before.
- `merge_pages(other)` replaces the draft merge in `ModelRunner`, carrying
  the gcd replication rule with it.

`validate_page_views` stays on the LMCache MP side: it is the consumer
checking what it is handed, not a second copy of construction.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Honglie Yi <hyi@crsuse2-m2m-v2-020.us-east2-a.compute.internal>
Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants