Repository navigation
Conversation
Port vLLM's KV offloading connector to Apple Silicon (MLX): - MetalOffloadingConnector + MetalKVOffloadHandler: synchronous bidirectional block mover over the MLX paged KV cache (unified memory: a transfer is an on-package copy). Registration-time guards: CACHE_ATTRS completeness introspection, spec-vs-handler block-byte cross-check, block-id range validation. - MetalSharedOffloadRegion: anonymous in-process RAM region for the tiering host pool (macOS has no tmpfs; a file-backed mmap would write KV churn back to SSD through APFS). Tiering enforces the single-process executor. - MetalFileSystemTierManager: macOS disk tier — F_NOCACHE, 0600 blocks under a Spotlight-excluded .noindex directory, Time Machine exclusion off the startup path, thread QoS (loads USER_INITIATED / stores UTILITY), layout-keyed store paths (dtype/quant disjoint), size-validating lookups + failed-load negative cache (no promotion livelock on corrupt files). - Platform config: performs the --kv-offloading-size translation, routes connector/spec lookups to the Metal classes, fail-fast guards for unsupported connectors/specs/tiers/roles/parallelism, pool-undersized and PYTHONHASHSEED warnings. - Scheduler side (eviction policies, block tracking, tiering cascade, fs tier orchestration) is inherited from upstream vLLM unmodified. Verified end-to-end on 1B-70B (fp16/bf16/TurboQuant KV, MoE, hybrid-attention guard rejection): byte-exact restores vs a no-offload resume control; disk restore 6-11x faster than recompute; cross-restart persistence; bench serve shows 2-4% connector overhead under saturation. Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
- Handler round-trips (fp16/bf16/TurboQuant), block_size_factor sub-slot striding incl. unaligned first block, error paths (unsupported dtype, OOB ids, short CPU-block lists, uncovered per-block arrays), spec-vs-handler byte cross-check. - CUDA parity: host-pool byte placement pinned to upstream's compute_sub_block_ptrs (the CUDA DMA descriptor math); scheduler surface asserted to be upstream code by function identity. - Shared-region tests: two-sided scheduler/worker carve identity, TurboQuant through the region's byte-level memoryview contract. - fs tier: store/load semantics, layout signature disjointness, directory permission behavior, end-to-end manager lifecycle incl. the failed-load livelock guard, QoS, Time Machine exclusion (slow). - Platform config guards: translation, connector/spec/tier/backend rejection matrix, kv_role normalization. Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
Architecture and tier rationale for unified memory (wired MLX pool vs pageable host pool vs NVMe), MLX read/write ordering contract, resume-numerics analysis (restored-prefix tail recompute is a different kernel batch shape; byte-exact restore does not imply bit-identical greedy continuations on near-tie prompts), verified results 1B-70B, macOS fs-tier integration notes, and future work (MLX-streams transfer overlap, hybrid-model multi-group support, object-store tier, KV compression). Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
Fraction-derived KV budgets are wired (non-pageable) and know nothing about host state: fraction 0.8 on a 128GB machine planned a ~90GB pool for a 1.5B model and drove free memory to 11% under load (observed live; the same pattern has hardware-reset development machines). compute_safe_kv_budget clamps the paged-attention plan against live host memory: - hard limit (always on): the plan must fit in currently-available memory — an oversized plan is cut rather than thrashing from the first request; - soft floor (VLLM_METAL_MIN_FREE_FRACTION, default 15% of total RAM, min 4GB): the budget is reduced so wiring can never push free memory below the floor, with a prominent warning explaining the clamp and how to size intentionally; - VLLM_METAL_DISABLE_MEMORY_GUARD=1 opts out of the floor (never the cannot-fit cut) for hosts managed externally. Validated by replaying the incident config: budget clamped 76.1GB -> 56.6GB at startup, serve healthy, free memory held above the floor. Small-RAM recipes (fraction 0.8 on 16GB CI Macs) are trimmed gently, not gutted — covered by unit tests. Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
The upstream defect the fs-tier lookup hardening works around is now filed with a full trace and a runnable no-GPU reproduction: vllm-project/vllm#49176 (failed secondary-tier load livelocks the request; the async lookup cache is never invalidated on load failure). The fix here is self-contained in the Metal tier and does not depend on the upstream resolution. |
CI's lint job runs mypy without the full vLLM install, so upstream types degrade to Any and two sites need explicit typing: annotate the lookup-manager replacement in the fs tier, and assert-narrow cache_config before writing kv_offloading_size (a non-None size implies it exists). Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
The soft floor's 4GB absolute minimum and the 4GB planning slack both assume 16GB-class machines; on GitHub's ~7GB macos-15 runners they exceed what the host can give and clamp every KV budget to zero (observed in PR CI). Cap the floor minimum at 25% of RAM and the slack at RAM/16 — behavior at 16GB and above is unchanged (invariance test added), and a regression test replays the CI runner's numbers. Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
|
CI failures diagnosed and fixed in the two commits just pushed:
Full local gates re-run green (1,545 non-slow tests, shellcheck/ruff/format/mypy). Workflows will need re-approval to run. |
fcntl is POSIX and present on Linux; only the F_NOCACHE constant is Darwin-specific and stays getattr-guarded. The conditional import made mypy --platform linux (the ubuntu lint jobs) see an undefined name at the use site. Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
The memory-safety clamp samples psutil available memory at plan time, after the model weights are loaded and resident — so subtracting model_memory from the available-derived limits counted it twice. Noise on large hosts; on GitHub's 6GB macos-15 runner (2.6GB available) it zeroed every KV budget even after the constants were scaled. Only allocations still to come (overhead, hybrid reservation) belong in planned_other_bytes. Regression test updated to the runner's observed numbers (6GB total / 2.6GB available). Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
|
Round two of CI fixes pushed (4bdf7e1, ee97a45):
Local gates green: 1,545 non-slow tests, shellcheck/ruff/format, mypy on both platforms. Workflows will need re-approval once more — thanks for the quick turnarounds. |
|
I think I understand the idea: save evicted KV blocks so repeated prefixes can be restored instead of recomputed. My concern is that the win seems narrow on Mac, since normal prefix caching already handles the common case. What workload needs this beyond prefix caching? |
|
Fair concern — and you're right that for a single session re-hitting a resident prefix, in-GPU prefix caching already handles it. But I'd actually argue the win is widest on Mac, not narrowest, for three reasons specific to the platform: 1. Prefixes evict sooner on Mac. Wired (non-pageable) KV memory is scarce on a unified-memory machine — it's shared with the OS, apps, and the model weights, so you can't give the KV cache the headroom a dedicated 80GB GPU has. Prefix caching's reuse window is therefore shorter here: the working set overflows the wired pool sooner, and once a block is LRU-dropped, prefix caching has nothing left. Offload extends that window into the pageable host pool and a disk tier instead of dropping to zero. 2. Recompute costs more on Mac. Prefill is the scarce resource on Apple Silicon, so the work a restore avoids is exactly the expensive part — and the penalty grows with model size: ~6× at 1.5B up to 11× at 70B (78.8s cold prefill → 7.1s disk restore). The larger the model you run on a Mac, the more a lost prefix hurts. 3. The spill is nearly free here. wired→pageable is a page-table change, not a PCIe copy — the host tier costs on CUDA (pinned buffers, bus transfer) mostly don't exist on unified memory. So the break-even for "restore instead of recompute" sits much lower on Mac. Add restarts on top — dev loops, redeploys, sleep/wake all discard the in-memory prefix cache while the disk tier survives (a fresh process restores a served prefix on its first request: 0.37s vs 1.42s cold at 1.5B, 7.3s vs 78.8s at 70B) — and it's also the building block for KV-aware routing (llm-d). So I'd flip the framing: prefix caching's window is narrowest exactly on the platform where losing a prefix costs the most to rebuild. Happy to add a memory-pressure or restart benchmark if a concrete case would land it better. |
|
Quick follow-up on the fs-tier lookup hardening: the upstream fix for that livelock is now open as a PR, not just a filed issue — vllm-project/vllm#49328. It's a three-layer fix in core vLLM's tiering (invalidate the stale lookup verdict on a failed promotion, size-validate fs lookups instead of checking bare existence, and quarantine provably-bad block files instead of deleting them), with a no-GPU regression test that reproduces the busy-loop. Once that lands in core, some of the hardening in this PR becomes belt-and-suspenders — the plugin inherits the fix. Until it's released, the guard here is what keeps the Metal fs tier safe on current vLLM, so I'd keep it either way and revisit once #49328 merges. |
LxYuan0420
left a comment
There was a problem hiding this comment.
Since the fs-tier work is still being settled upstream in vllm#49328, let’s close this for now and revisit after it lands.
|
Superseded by #681, a fresh PR on main at vLLM 0.28.0, which carries vllm-project/vllm#49328. Same scope, host pool plus fs disk tier, with the benchmarks this one lacked. |
Use core vLLM's KV offloading on Apple Silicon: the same --kv-offloading-size and --kv-transfer-config flags as CUDA, the same OffloadingConnector, and upstream's TieringOffloadingSpec for the host pool plus fs disk tiers. Scheduler-side orchestration is inherited. Only the platform-bound pieces are replaced. - The worker. Upstream's copy engine is CUDA-only, and a torch alias of the KV cache goes stale because the MLX cache rebinds each layer's array on every write. MetalKVOffloadWorker moves blocks between the live MLX arrays and the host pool, gathering through the current list entry so reads are ordered after pending writes. - The shared region. Upstream's is /dev/shm plus madvise. macOS has neither, so MetalSharedOffloadRegion is anonymous RAM shared within one process. That is why offloading requires the single-process executor. - The fs tier. fcntl(F_NOCACHE) where upstream would use O_DIRECT, 0o600 block files under a 0o700 root, and a blocks.noindex directory so Spotlight ignores the churn. The on-disk format is upstream's, byte for byte. The platform hook translates --kv-offloading-size itself, because vLLM's own translation runs after the hook and would select a connector this platform cannot serve. Unsupported configurations fail at config time with a specific message: the obj tier (NIXL has no macOS build), pipeline parallelism, multi-process executors, hybrid and sliding-window models, and draft-model speculative decoding. The model runner keeps the connector step open across the execute_model / sample_tokens split and closes it when the deferred decode is submitted, so stores are queued before the next step's forward can overwrite their blocks. When KV cache events are enabled globally, enable_kv_events is set on each tier; the reverse combination is refused. A KV-aware router then sees BlockStored events without the two switches being lined up by hand. 76 tests: byte-exact round trips (fp16, bf16, TurboQuant), block placement checked against upstream's pointer arithmetic on the real region, the store fence across the deferred decode, the fs tier end to end including upstream's failed-load negative cache, and the config guards. Nothing changes when offloading is off. Every new call site is behind has_kv_transfer_group(). Supersedes vllm-project#530. Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
Use core vLLM's KV offloading on Apple Silicon: the same --kv-offloading-size and --kv-transfer-config flags as CUDA, the same OffloadingConnector, and upstream's TieringOffloadingSpec for the host pool plus fs disk tiers. Scheduler-side orchestration is inherited. Only the platform-bound pieces are replaced. - The worker. Upstream's copy engine is CUDA-only, and a torch alias of the KV cache goes stale because the MLX cache rebinds each layer's array on every write. MetalKVOffloadWorker moves blocks between the live MLX arrays and the host pool, gathering through the current list entry so reads are ordered after pending writes. - The shared region. Upstream's is /dev/shm plus madvise. macOS has neither, so MetalSharedOffloadRegion is anonymous RAM shared within one process. That is why offloading requires the single-process executor. - The fs tier. fcntl(F_NOCACHE) where upstream would use O_DIRECT, 0o600 block files under a 0o700 root, and a blocks.noindex directory so Spotlight ignores the churn. The on-disk format is upstream's, byte for byte. The platform hook translates --kv-offloading-size itself, because vLLM's own translation runs after the hook and would select a connector this platform cannot serve. Unsupported configurations fail at config time with a specific message: the obj tier (NIXL has no macOS build), pipeline parallelism, multi-process executors, hybrid and sliding-window models, and draft-model speculative decoding. The model runner keeps the connector step open across the execute_model / sample_tokens split and closes it when the deferred decode is submitted, so stores are queued before the next step's forward can overwrite their blocks. When KV cache events are enabled globally, enable_kv_events is set on each tier; the reverse combination is refused. A KV-aware router then sees BlockStored events without the two switches being lined up by hand. 76 tests: byte-exact round trips (fp16, bf16, TurboQuant), block placement checked against upstream's pointer arithmetic on the real region, the store fence across the deferred decode, the fs tier end to end including upstream's failed-load negative cache, and the config guards. Nothing changes when offloading is off. Every new call site is behind has_kv_transfer_group(). Supersedes vllm-project#530. Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
Use core vLLM's KV offloading on Apple Silicon: the same --kv-offloading-size and --kv-transfer-config flags as CUDA, the same OffloadingConnector, and upstream's TieringOffloadingSpec for the host pool plus fs disk tiers. Scheduler-side orchestration is inherited. Only the platform-bound pieces are replaced. - The worker. Upstream's copy engine is CUDA-only, and a torch alias of the KV cache goes stale because the MLX cache rebinds each layer's array on every write. MetalKVOffloadWorker moves blocks between the live MLX arrays and the host pool, gathering through the current list entry so reads are ordered after pending writes. - The shared region. Upstream's is /dev/shm plus madvise. macOS has neither, so MetalSharedOffloadRegion is anonymous RAM shared within one process. That is why offloading requires the single-process executor. - The fs tier. fcntl(F_NOCACHE) where upstream would use O_DIRECT, 0o600 block files under a 0o700 root, and a blocks.noindex directory so Spotlight ignores the churn. The on-disk format is upstream's, byte for byte. The platform hook translates --kv-offloading-size itself, because vLLM's own translation runs after the hook and would select a connector this platform cannot serve. Unsupported configurations fail at config time with a specific message: the obj tier (NIXL has no macOS build), pipeline parallelism, multi-process executors, hybrid and sliding-window models, and draft-model speculative decoding. The model runner keeps the connector step open across the execute_model / sample_tokens split and closes it when the deferred decode is submitted, so stores are queued before the next step's forward can overwrite their blocks. When KV cache events are enabled globally, enable_kv_events is set on each tier; the reverse combination is refused. A KV-aware router then sees BlockStored events without the two switches being lined up by hand. 76 tests: byte-exact round trips (fp16, bf16, TurboQuant), block placement checked against upstream's pointer arithmetic on the real region, the store fence across the deferred decode, the fs tier end to end including upstream's failed-load negative cache, and the config guards. Nothing changes when offloading is off. Every new call site is behind has_kv_transfer_group(). Supersedes vllm-project#530. Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
Why
The cost of long-context inference lives in prefill. System prompts, RAG
contexts, and multi-turn sessions re-process the same tokens over and over,
and on Apple Silicon — where prefill compute is the scarce resource — that
tax is paid in full every time a prefix falls out of the GPU cache. vLLM
already solved this class of problem with the
OffloadingConnector(KVblocks spill to host RAM and secondary tiers, and are restored instead of
recomputed), but the implementation is CUDA-only. Macs get nothing.
A note on what "offloading" means on unified memory, since there is no
separate device memory to offload from: the tier boundary on Metal is
wired vs pageable, not device vs host. The MLX paged KV cache lives in
wired (non-pageable) memory, hard-capped by the memory-fraction budget —
that budget is the scarce resource. The offload pool is plain pageable
memory outside that budget, free for macOS to reclaim under pressure, and
the disk tier sits below it. Offloading extends effective KV capacity beyond
the wired cap; what a restore saves is prefill compute, not memory-bus
distance.
That framing is also why this is backwards for the platform, twice over:
restore crosses PCIe from pinned host pools; on Apple Silicon it's an
on-package copy at 5–9 GB/s with no pinning, no staging, no DMA
descriptors. The feature costs less on Metal than where it was invented.
recompute advantage scales with model size: 6× at 1.5B, 8× at 32B, 11×
at 70B (78.8s cold prefill → 7.1s disk restore for Llama-3.3-70B). The
bigger the model, the more an evicted prefix costs to recompute — and the
more this pays back.
With the disk tier and
PYTHONHASHSEEDpinned, reuse also survives serverrestarts: a fresh process restores a previously-served prefix from disk on
its first request (0.37s vs 1.42s cold at 1.5B; 7.3s vs 78.8s at 70B). A Mac
that has prefilled a context once never pays full price for it again.
What this PR does
Wires vLLM's native offloading (
--kv-offloading-size, upstream'sOffloadingConnector) into vllm-metal, with a pageable host-pool tier, anfsdisk tier, and cross-restart persistence:function identity, and byte placement in the host pool is pinned to
upstream's own
compute_sub_block_ptrsmath (the CUDA DMA descriptorlayout, runnable on any platform) across block-size factors, alignment
cases, and KV dtypes.
MetalKVOffloadWorker(implements 0.25.1's
OffloadingWorker) moves blocks between the wiredMLX paged KV cache and a host pool, following the cache kernels'
write-then-rebind idiom so MLX graph provenance orders reads after writes.
fsdisk tier is upstream's tier with macOS behavior layered on:0o600files /0o700dirs (KV blocks are conversation-derived data withpresence-testable hash names),
blocks.noindexso Spotlight never indexeschurn, async Time Machine exclusion,
F_NOCACHE(upstream'sO_DIRECTsilently degrades to buffered on macOS), QoS-tagged I/O threads, and an
8-byte CRC32 footer per block file verified on every load — torn
writes and bit rot become clean misses instead of poisoned KV.
models (e.g. gemma-4), NIXL
objtiers, heterogeneous KV dtypes, andunsupported connector configs are rejected at startup with actionable
messages. A memory-safety guard clamps the paged-attention plan against
live host memory (hard cannot-fit cut + configurable free-RAM floor), so
oversized fractions degrade gracefully instead of wiring the machine into
the ground.
testing: the tiering async-lookup cache is never invalidated when a load
fails, which livelocks a request into re-promoting a deleted block forever.
The Metal tier validates file size at lookup and negative-caches failed
loads (upstream issue to follow — the fix here is self-contained).
Deliberate deltas vs the CUDA path
Studied and intentionally not ported, with reasons: pinned host pools and
cudaHostRegister(no-ops on unified memory), page prefaulting (activelyharmful — wires cold pages), stream/descriptor pooling and GDS (no
equivalent I/O path). Transfers currently run synchronously in the submit
hooks; the worker API is shaped for a later move to
mx.async_eval+thread-local streams (MLX ≥0.32), which is the top item in the design doc's
future work.
On "warm output differs from cold" — read before testing
If you drive this with a repetitive prompt and diff greedy outputs against
cold prefill, you will eventually see a divergence and suspect corruption.
It is not. The restore path is byte-exact (env-gated per-row CRC
instrumentation verifies stored == loaded == post-scatter GPU bytes), and
the divergence reproduces with no offloading configured — plain GPU
prefix-cache resume. Mechanism: restored-prefix + short-tail recompute is a
different kernel batch shape than chunked cold prefill; fp16 logits differ
by ~1e-3 nats, and on near-tie tokens (we measured a 0.045-nat top-2 margin
at a flip point) greedy argmax flips deterministically. This is inherent to
resume-then-continue on any backend, CUDA included. The correct oracle —
used by our tests — compares restored output against a no-offload resume
of the same prefix, which is byte-identical across all tested models.
Evidence
E2E matrix (local harness, resume oracle above; byte-identical output in
every cell):
"guarded" = enabling offloading for this model is rejected at startup with
a
NotImplementedErrornaming the reason (hybrid attention: mixedsliding/full layers need per-group KV handling the single-group port doesn't
have yet); plain serving without offloading is unaffected. Rejection is the
designed behavior, not a failure: accepting these configs would offload only
part of each block's state and restore corrupted KV silently. The guard path
itself is under test — these two models are e2e cells asserting the
rejection fires cleanly.
cascaded and restored.
vllm bench serve, sonnet dataset, 100prompts, per the contributing guide): connector overhead 1.9–4.6% —
median TTFT +1.9%, TPOT +2.3%, E2EL +3.5%, P99s within 3.1%, throughput
unchanged within noise.
scripts/lint.shandscripts/test.shgreen. The stack was additionallyvalidated end-to-end from a scratch install on a second machine (M-series
Studio), same byte-exactness and scaling behavior.
OffloadingWorker/LookupResultinterfaces);rebased on current
main.Known limitations (guarded, documented, future work in the design doc)
sliding-window models are cleanly rejected at config time (three gemma
variants tested for the rejection path). gemma-2 works where
sliding_window == max_model_len; beyond-window is untested.yet).
eviction are listed in the design doc's future work, alongside a
purgeable host pool,
MADV_FREEeviction hooks, and streams overlap.Try it
Serve a long prompt, evict it (fill the cache with other traffic or restart
the server), send it again, and watch
vllm:prompt_tokens_by_source_total{source="external_kv_transfer"}gononzero while the response returns in a fraction of the cold time.
Design doc:
docs/offload-design.md(architecture, correctness analysis,resume-numerics writeup, future work). Diagram:
docs/offload-architecture.mmd.Generated with Claude Code