Skip to content

feat: KV offloading for the Metal backend (pageable host pool, disk tier, cross-restart reuse) - #530

Closed
RobbieJ wants to merge 8 commits into
vllm-project:mainfrom
RobbieJ:kv-offload-metal-pr
Closed

RobbieJ wants to merge 8 commits into
vllm-project:mainfrom
RobbieJ:kv-offload-metal-pr

Conversation

@RobbieJ

@RobbieJ RobbieJ commented Jul 20, 2026

Copy link
Copy Markdown
Contributor

Why

The cost of long-context inference lives in prefill. System prompts, RAG
contexts, and multi-turn sessions re-process the same tokens over and over,
and on Apple Silicon — where prefill compute is the scarce resource — that
tax is paid in full every time a prefix falls out of the GPU cache. vLLM
already solved this class of problem with the OffloadingConnector (KV
blocks spill to host RAM and secondary tiers, and are restored instead of
recomputed), but the implementation is CUDA-only. Macs get nothing.

A note on what "offloading" means on unified memory, since there is no
separate device memory to offload from: the tier boundary on Metal is
wired vs pageable, not device vs host. The MLX paged KV cache lives in
wired (non-pageable) memory, hard-capped by the memory-fraction budget —
that budget is the scarce resource. The offload pool is plain pageable
memory outside that budget, free for macOS to reclaim under pressure, and
the disk tier sits below it. Offloading extends effective KV capacity beyond
the wired cap; what a restore saves is prefill compute, not memory-bus
distance.

That framing is also why this is backwards for the platform, twice over:

  • Unified memory makes offloading unusually cheap here. On CUDA, a
    restore crosses PCIe from pinned host pools; on Apple Silicon it's an
    on-package copy at 5–9 GB/s with no pinning, no staging, no DMA
    descriptors. The feature costs less on Metal than where it was invented.
  • The win grows exactly where Macs hurt most. Measured restore-vs-
    recompute advantage scales with model size: 6× at 1.5B, 8× at 32B, 11×
    at 70B
    (78.8s cold prefill → 7.1s disk restore for Llama-3.3-70B). The
    bigger the model, the more an evicted prefix costs to recompute — and the
    more this pays back.

With the disk tier and PYTHONHASHSEED pinned, reuse also survives server
restarts: a fresh process restores a previously-served prefix from disk on
its first request (0.37s vs 1.42s cold at 1.5B; 7.3s vs 78.8s at 70B). A Mac
that has prefilled a context once never pays full price for it again.

What this PR does

Wires vLLM's native offloading (--kv-offloading-size, upstream's
OffloadingConnector) into vllm-metal, with a pageable host-pool tier, an
fs disk tier, and cross-restart persistence:

  • The scheduler side is upstream code, unmodified. Tests assert this by
    function identity, and byte placement in the host pool is pinned to
    upstream's own compute_sub_block_ptrs math (the CUDA DMA descriptor
    layout, runnable on any platform) across block-size factors, alignment
    cases, and KV dtypes.
  • Only the worker-side data plane is Metal: MetalKVOffloadWorker
    (implements 0.25.1's OffloadingWorker) moves blocks between the wired
    MLX paged KV cache and a host pool, following the cache kernels'
    write-then-rebind idiom so MLX graph provenance orders reads after writes.
  • The fs disk tier is upstream's tier with macOS behavior layered on:
    0o600 files / 0o700 dirs (KV blocks are conversation-derived data with
    presence-testable hash names), blocks.noindex so Spotlight never indexes
    churn, async Time Machine exclusion, F_NOCACHE (upstream's O_DIRECT
    silently degrades to buffered on macOS), QoS-tagged I/O threads, and an
    8-byte CRC32 footer per block file verified on every load — torn
    writes and bit rot become clean misses instead of poisoned KV.
  • Config-time guards, not runtime surprises: hybrid/sliding-window
    models (e.g. gemma-4), NIXL obj tiers, heterogeneous KV dtypes, and
    unsupported connector configs are rejected at startup with actionable
    messages. A memory-safety guard clamps the paged-attention plan against
    live host memory (hard cannot-fit cut + configurable free-RAM floor), so
    oversized fractions degrade gracefully instead of wiring the machine into
    the ground.
  • The tier's lookup also works around a live upstream defect we found while
    testing: the tiering async-lookup cache is never invalidated when a load
    fails, which livelocks a request into re-promoting a deleted block forever.
    The Metal tier validates file size at lookup and negative-caches failed
    loads (upstream issue to follow — the fix here is self-contained).

Deliberate deltas vs the CUDA path

Studied and intentionally not ported, with reasons: pinned host pools and
cudaHostRegister (no-ops on unified memory), page prefaulting (actively
harmful — wires cold pages), stream/descriptor pooling and GDS (no
equivalent I/O path). Transfers currently run synchronously in the submit
hooks; the worker API is shaped for a later move to mx.async_eval +
thread-local streams (MLX ≥0.32), which is the top item in the design doc's
future work.

On "warm output differs from cold" — read before testing

If you drive this with a repetitive prompt and diff greedy outputs against
cold prefill, you will eventually see a divergence and suspect corruption.
It is not. The restore path is byte-exact (env-gated per-row CRC
instrumentation verifies stored == loaded == post-scatter GPU bytes), and
the divergence reproduces with no offloading configured — plain GPU
prefix-cache resume. Mechanism: restored-prefix + short-tail recompute is a
different kernel batch shape than chunked cold prefill; fp16 logits differ
by ~1e-3 nats, and on near-tie tokens (we measured a 0.045-nat top-2 margin
at a flip point) greedy argmax flips deterministically. This is inherent to
resume-then-continue on any backend, CUDA included. The correct oracle —
used by our tests — compares restored output against a no-offload resume
of the same prefix, which is byte-identical across all tested models.

Evidence

E2E matrix (local harness, resume oracle above; byte-identical output in
every cell):

Model serve host pool disk cascade cross-restart cold → disk restore
Qwen2.5-1.5B-4bit pass pass pass pass 2.0s → 0.33s
Qwen2.5-1.5B + TurboQuant KV pass pass pass pass —
Qwen2.5-7B-4bit pass pass pass pass 7.8s → 0.8s
Llama-3.2-3B-4bit pass pass pass pass 3.3s → 0.45s
Llama-3.2-1B (bf16 KV) pass pass pass pass 1.0s → 0.34s
gemma-2-2b-4bit pass pass pass pass 1.5s → 0.37s
Qwen2.5-32B-4bit pass pass pass pass 33.6s → 4.1s (8.3×)
Qwen3-30B-A3B-4bit (MoE) pass pass pass pass 3.6s → 0.64s
gemma-4-e4b / gemma-4-31b (hybrid) pass guarded guarded guarded clean config-time rejection
Llama-3.3-70B-4bit pass — pass pass 78.8s → 7.1s (11.1×)

"guarded" = enabling offloading for this model is rejected at startup with
a NotImplementedError naming the reason (hybrid attention: mixed
sliding/full layers need per-group KV handling the single-group port doesn't
have yet); plain serving without offloading is unaffected. Rejection is the
designed behavior, not a failure: accepting these configs would offload only
part of each block's state and restore corrupted KV silently. The guard path
itself is under test — these two models are e2e cells asserting the
rejection fires cleanly.

  • Transfer bandwidth 5–9 GB/s across sizes; disk stores up to 18.9 GB
    cascaded and restored.
  • No serving regression (vllm bench serve, sonnet dataset, 100
    prompts, per the contributing guide): connector overhead 1.9–4.6% —
    median TTFT +1.9%, TPOT +2.3%, E2EL +3.5%, P99s within 3.1%, throughput
    unchanged within noise.
  • 56 offload unit/parity tests; full non-slow suite (1,541) green;
    scripts/lint.sh and scripts/test.sh green. The stack was additionally
    validated end-to-end from a scratch install on a second machine (M-series
    Studio), same byte-exactness and scaling behavior.
  • Targets vLLM 0.25.1 (OffloadingWorker / LookupResult interfaces);
    rebased on current main.

Known limitations (guarded, documented, future work in the design doc)

  • Single KV cache group / uniform full attention — hybrid and
    sliding-window models are cleanly rejected at config time (three gemma
    variants tested for the rejection path). gemma-2 works where
    sliding_window == max_model_len; beyond-window is untested.
  • Single worker rank; synchronous transfers (no compute/transfer overlap
    yet).
  • The disk store has no GC or size cap (upstream behavior); retention and
    eviction are listed in the design doc's future work, alongside a
    purgeable host pool, MADV_FREE eviction hooks, and streams overlap.

Try it

PYTHONHASHSEED=0 VLLM_METAL_USE_PAGED_ATTENTION=1 \
vllm serve mlx-community/Qwen2.5-1.5B-Instruct-4bit \
  --kv-offloading-backend native --kv-offloading-size 8 \
  --kv-transfer-config '{"kv_connector_extra_config":
    {"secondary_tiers": [{"type": "fs", "root_dir": "/tmp/kv-store"}]}}'

Serve a long prompt, evict it (fill the cache with other traffic or restart
the server), send it again, and watch
vllm:prompt_tokens_by_source_total{source="external_kv_transfer"} go
nonzero while the response returns in a fraction of the cold time.

Design doc: docs/offload-design.md (architecture, correctness analysis,
resume-numerics writeup, future work). Diagram:
docs/offload-architecture.mmd.


Generated with Claude Code

RobbieJ added 4 commits July 19, 2026 22:46
Port vLLM's KV offloading connector to Apple Silicon (MLX):

- MetalOffloadingConnector + MetalKVOffloadHandler: synchronous
  bidirectional block mover over the MLX paged KV cache (unified
  memory: a transfer is an on-package copy). Registration-time
  guards: CACHE_ATTRS completeness introspection, spec-vs-handler
  block-byte cross-check, block-id range validation.
- MetalSharedOffloadRegion: anonymous in-process RAM region for the
  tiering host pool (macOS has no tmpfs; a file-backed mmap would
  write KV churn back to SSD through APFS). Tiering enforces the
  single-process executor.
- MetalFileSystemTierManager: macOS disk tier — F_NOCACHE, 0600
  blocks under a Spotlight-excluded .noindex directory, Time Machine
  exclusion off the startup path, thread QoS (loads USER_INITIATED /
  stores UTILITY), layout-keyed store paths (dtype/quant disjoint),
  size-validating lookups + failed-load negative cache (no
  promotion livelock on corrupt files).
- Platform config: performs the --kv-offloading-size translation,
  routes connector/spec lookups to the Metal classes, fail-fast
  guards for unsupported connectors/specs/tiers/roles/parallelism,
  pool-undersized and PYTHONHASHSEED warnings.
- Scheduler side (eviction policies, block tracking, tiering
  cascade, fs tier orchestration) is inherited from upstream vLLM
  unmodified.

Verified end-to-end on 1B-70B (fp16/bf16/TurboQuant KV, MoE,
hybrid-attention guard rejection): byte-exact restores vs a
no-offload resume control; disk restore 6-11x faster than
recompute; cross-restart persistence; bench serve shows 2-4%
connector overhead under saturation.

Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
- Handler round-trips (fp16/bf16/TurboQuant), block_size_factor
  sub-slot striding incl. unaligned first block, error paths
  (unsupported dtype, OOB ids, short CPU-block lists, uncovered
  per-block arrays), spec-vs-handler byte cross-check.
- CUDA parity: host-pool byte placement pinned to upstream's
  compute_sub_block_ptrs (the CUDA DMA descriptor math); scheduler
  surface asserted to be upstream code by function identity.
- Shared-region tests: two-sided scheduler/worker carve identity,
  TurboQuant through the region's byte-level memoryview contract.
- fs tier: store/load semantics, layout signature disjointness,
  directory permission behavior, end-to-end manager lifecycle incl.
  the failed-load livelock guard, QoS, Time Machine exclusion (slow).
- Platform config guards: translation, connector/spec/tier/backend
  rejection matrix, kv_role normalization.

Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
Architecture and tier rationale for unified memory (wired MLX pool
vs pageable host pool vs NVMe), MLX read/write ordering contract,
resume-numerics analysis (restored-prefix tail recompute is a
different kernel batch shape; byte-exact restore does not imply
bit-identical greedy continuations on near-tie prompts), verified
results 1B-70B, macOS fs-tier integration notes, and future work
(MLX-streams transfer overlap, hybrid-model multi-group support,
object-store tier, KV compression).

Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
Fraction-derived KV budgets are wired (non-pageable) and know nothing
about host state: fraction 0.8 on a 128GB machine planned a ~90GB pool
for a 1.5B model and drove free memory to 11% under load (observed
live; the same pattern has hardware-reset development machines).

compute_safe_kv_budget clamps the paged-attention plan against live
host memory:
- hard limit (always on): the plan must fit in currently-available
  memory — an oversized plan is cut rather than thrashing from the
  first request;
- soft floor (VLLM_METAL_MIN_FREE_FRACTION, default 15% of total RAM,
  min 4GB): the budget is reduced so wiring can never push free
  memory below the floor, with a prominent warning explaining the
  clamp and how to size intentionally;
- VLLM_METAL_DISABLE_MEMORY_GUARD=1 opts out of the floor (never the
  cannot-fit cut) for hosts managed externally.

Validated by replaying the incident config: budget clamped
76.1GB -> 56.6GB at startup, serve healthy, free memory held above
the floor. Small-RAM recipes (fraction 0.8 on 16GB CI Macs) are
trimmed gently, not gutted — covered by unit tests.

Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, you can upgrade your account or add credits to your account and enable them for code reviews in your settings.

@RobbieJ

RobbieJ commented Jul 20, 2026

Copy link
Copy Markdown
Contributor Author

The upstream defect the fs-tier lookup hardening works around is now filed with a full trace and a runnable no-GPU reproduction: vllm-project/vllm#49176 (failed secondary-tier load livelocks the request; the async lookup cache is never invalidated on load failure). The fix here is self-contained in the Metal tier and does not depend on the upstream resolution.

RobbieJ added 2 commits July 20, 2026 13:07
CI's lint job runs mypy without the full vLLM install, so upstream types
degrade to Any and two sites need explicit typing: annotate the
lookup-manager replacement in the fs tier, and assert-narrow cache_config
before writing kv_offloading_size (a non-None size implies it exists).

Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
The soft floor's 4GB absolute minimum and the 4GB planning slack both
assume 16GB-class machines; on GitHub's ~7GB macos-15 runners they exceed
what the host can give and clamp every KV budget to zero (observed in PR
CI). Cap the floor minimum at 25% of RAM and the slack at RAM/16 —
behavior at 16GB and above is unchanged (invariance test added), and a
regression test replays the CI runner's numbers.

Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
@RobbieJ

RobbieJ commented Jul 20, 2026

Copy link
Copy Markdown
Contributor Author

CI failures diagnosed and fixed in the two commits just pushed:

  • lint: the lint job runs mypy without the full vLLM install, so upstream types resolve to Any and two sites needed explicit typing (annotation on the fs-tier lookup-manager replacement; assert-narrowing before a cache_config write). No behavior change.
  • test: the memory-safety guard's floor minimum (4GB) and planning slack (4GB) were sized for 16GB-class machines; on the ~7GB macos-15 runner they exceed what the host can give and clamped the KV budget to zero. Both now scale to the machine (floor minimum capped at 25% of RAM, slack at RAM/16); behavior at 16GB+ is unchanged, with an invariance test, and a regression test replays the runner's exact numbers.

Full local gates re-run green (1,545 non-slow tests, shellcheck/ruff/format/mypy). Workflows will need re-approval to run.

RobbieJ added 2 commits July 20, 2026 14:46
fcntl is POSIX and present on Linux; only the F_NOCACHE constant is
Darwin-specific and stays getattr-guarded. The conditional import made
mypy --platform linux (the ubuntu lint jobs) see an undefined name at
the use site.

Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
The memory-safety clamp samples psutil available memory at plan time,
after the model weights are loaded and resident — so subtracting
model_memory from the available-derived limits counted it twice. Noise
on large hosts; on GitHub's 6GB macos-15 runner (2.6GB available) it
zeroed every KV budget even after the constants were scaled. Only
allocations still to come (overhead, hybrid reservation) belong in
planned_other_bytes. Regression test updated to the runner's observed
numbers (6GB total / 2.6GB available).

Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
@RobbieJ

RobbieJ commented Jul 20, 2026

Copy link
Copy Markdown
Contributor Author

Round two of CI fixes pushed (4bdf7e1, ee97a45):

  • lint (ubuntu): fcntl was imported under a darwin platform guard, so Linux mypy saw an undefined name (previously hidden by fail-fast cancellation). Now imported unconditionally — it is POSIX; only the F_NOCACHE constant is Darwin-specific and remains getattr-guarded. Both mypy platforms now run in our local gate (mypy --platform linux mirrors the ubuntu jobs).
  • test (macos-15): root cause found via the runner's own log (6GB total, 2.6GB available): the memory guard sampled available memory after the model weights were resident but still subtracted model_memory from the available-derived limits — a double-count that is noise on large hosts and zeroed the budget on the small runner. Only still-to-come allocations (overhead, hybrid reservation) count now; the regression test replays the runner's observed numbers.

Local gates green: 1,545 non-slow tests, shellcheck/ruff/format, mypy on both platforms. Workflows will need re-approval once more — thanks for the quick turnarounds.

@LxYuan0420

Copy link
Copy Markdown
Collaborator

I think I understand the idea: save evicted KV blocks so repeated prefixes can be restored instead of recomputed. My concern is that the win seems narrow on Mac, since normal prefix caching already handles the common case. What workload needs this beyond prefix caching?

@RobbieJ

RobbieJ commented Jul 21, 2026

Copy link
Copy Markdown
Contributor Author

Fair concern — and you're right that for a single session re-hitting a resident prefix, in-GPU prefix caching already handles it. But I'd actually argue the win is widest on Mac, not narrowest, for three reasons specific to the platform:

1. Prefixes evict sooner on Mac. Wired (non-pageable) KV memory is scarce on a unified-memory machine — it's shared with the OS, apps, and the model weights, so you can't give the KV cache the headroom a dedicated 80GB GPU has. Prefix caching's reuse window is therefore shorter here: the working set overflows the wired pool sooner, and once a block is LRU-dropped, prefix caching has nothing left. Offload extends that window into the pageable host pool and a disk tier instead of dropping to zero.

2. Recompute costs more on Mac. Prefill is the scarce resource on Apple Silicon, so the work a restore avoids is exactly the expensive part — and the penalty grows with model size: ~6× at 1.5B up to 11× at 70B (78.8s cold prefill → 7.1s disk restore). The larger the model you run on a Mac, the more a lost prefix hurts.

3. The spill is nearly free here. wired→pageable is a page-table change, not a PCIe copy — the host tier costs on CUDA (pinned buffers, bus transfer) mostly don't exist on unified memory. So the break-even for "restore instead of recompute" sits much lower on Mac.

Add restarts on top — dev loops, redeploys, sleep/wake all discard the in-memory prefix cache while the disk tier survives (a fresh process restores a served prefix on its first request: 0.37s vs 1.42s cold at 1.5B, 7.3s vs 78.8s at 70B) — and it's also the building block for KV-aware routing (llm-d).

So I'd flip the framing: prefix caching's window is narrowest exactly on the platform where losing a prefix costs the most to rebuild. Happy to add a memory-pressure or restart benchmark if a concrete case would land it better.

@RobbieJ

RobbieJ commented Jul 21, 2026

Copy link
Copy Markdown
Contributor Author

Quick follow-up on the fs-tier lookup hardening: the upstream fix for that livelock is now open as a PR, not just a filed issue — vllm-project/vllm#49328. It's a three-layer fix in core vLLM's tiering (invalidate the stale lookup verdict on a failed promotion, size-validate fs lookups instead of checking bare existence, and quarantine provably-bad block files instead of deleting them), with a no-GPU regression test that reproduces the busy-loop.

Once that lands in core, some of the hardening in this PR becomes belt-and-suspenders — the plugin inherits the fix. Until it's released, the guard here is what keeps the Metal fs tier safe on current vLLM, so I'd keep it either way and revisit once #49328 merges.

@LxYuan0420 LxYuan0420 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Since the fs-tier work is still being settled upstream in vllm#49328, let’s close this for now and revisit after it lands.

@RobbieJ

RobbieJ commented Sep 3, 2026

Copy link
Copy Markdown
Contributor Author

Superseded by #681, a fresh PR on main at vLLM 0.28.0, which carries vllm-project/vllm#49328. Same scope, host pool plus fs disk tier, with the benchmarks this one lacked.

RobbieJ added a commit to RobbieJ/vllm-metal that referenced this pull request Sep 5, 2026
Use core vLLM's KV offloading on Apple Silicon: the same
--kv-offloading-size and --kv-transfer-config flags as CUDA, the same
OffloadingConnector, and upstream's TieringOffloadingSpec for the host
pool plus fs disk tiers. Scheduler-side orchestration is inherited.

Only the platform-bound pieces are replaced.

- The worker. Upstream's copy engine is CUDA-only, and a torch alias of
  the KV cache goes stale because the MLX cache rebinds each layer's
  array on every write. MetalKVOffloadWorker moves blocks between the
  live MLX arrays and the host pool, gathering through the current list
  entry so reads are ordered after pending writes.
- The shared region. Upstream's is /dev/shm plus madvise. macOS has
  neither, so MetalSharedOffloadRegion is anonymous RAM shared within
  one process. That is why offloading requires the single-process
  executor.
- The fs tier. fcntl(F_NOCACHE) where upstream would use O_DIRECT,
  0o600 block files under a 0o700 root, and a blocks.noindex directory
  so Spotlight ignores the churn. The on-disk format is upstream's, byte
  for byte.

The platform hook translates --kv-offloading-size itself, because vLLM's
own translation runs after the hook and would select a connector this
platform cannot serve. Unsupported configurations fail at config time
with a specific message: the obj tier (NIXL has no macOS build),
pipeline parallelism, multi-process executors, hybrid and sliding-window
models, and draft-model speculative decoding.

The model runner keeps the connector step open across the
execute_model / sample_tokens split and closes it when the deferred
decode is submitted, so stores are queued before the next step's
forward can overwrite their blocks.

When KV cache events are enabled globally, enable_kv_events is set on
each tier; the reverse combination is refused. A KV-aware router then
sees BlockStored events without the two switches being lined up by hand.

76 tests: byte-exact round trips (fp16, bf16, TurboQuant), block
placement checked against upstream's pointer arithmetic on the real
region, the store fence across the deferred decode, the fs tier end to
end including upstream's failed-load negative cache, and the config
guards.

Nothing changes when offloading is off. Every new call site is behind
has_kv_transfer_group().

Supersedes vllm-project#530.

Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
RobbieJ added a commit to RobbieJ/vllm-metal that referenced this pull request Sep 7, 2026
Use core vLLM's KV offloading on Apple Silicon: the same
--kv-offloading-size and --kv-transfer-config flags as CUDA, the same
OffloadingConnector, and upstream's TieringOffloadingSpec for the host
pool plus fs disk tiers. Scheduler-side orchestration is inherited.

Only the platform-bound pieces are replaced.

- The worker. Upstream's copy engine is CUDA-only, and a torch alias of
  the KV cache goes stale because the MLX cache rebinds each layer's
  array on every write. MetalKVOffloadWorker moves blocks between the
  live MLX arrays and the host pool, gathering through the current list
  entry so reads are ordered after pending writes.
- The shared region. Upstream's is /dev/shm plus madvise. macOS has
  neither, so MetalSharedOffloadRegion is anonymous RAM shared within
  one process. That is why offloading requires the single-process
  executor.
- The fs tier. fcntl(F_NOCACHE) where upstream would use O_DIRECT,
  0o600 block files under a 0o700 root, and a blocks.noindex directory
  so Spotlight ignores the churn. The on-disk format is upstream's, byte
  for byte.

The platform hook translates --kv-offloading-size itself, because vLLM's
own translation runs after the hook and would select a connector this
platform cannot serve. Unsupported configurations fail at config time
with a specific message: the obj tier (NIXL has no macOS build),
pipeline parallelism, multi-process executors, hybrid and sliding-window
models, and draft-model speculative decoding.

The model runner keeps the connector step open across the
execute_model / sample_tokens split and closes it when the deferred
decode is submitted, so stores are queued before the next step's
forward can overwrite their blocks.

When KV cache events are enabled globally, enable_kv_events is set on
each tier; the reverse combination is refused. A KV-aware router then
sees BlockStored events without the two switches being lined up by hand.

76 tests: byte-exact round trips (fp16, bf16, TurboQuant), block
placement checked against upstream's pointer arithmetic on the real
region, the store fence across the deferred decode, the fs tier end to
end including upstream's failed-load negative cache, and the config
guards.

Nothing changes when offloading is off. Every new call site is behind
has_kv_transfer_group().

Supersedes vllm-project#530.

Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
RobbieJ added a commit to RobbieJ/vllm-metal that referenced this pull request Sep 9, 2026
Use core vLLM's KV offloading on Apple Silicon: the same
--kv-offloading-size and --kv-transfer-config flags as CUDA, the same
OffloadingConnector, and upstream's TieringOffloadingSpec for the host
pool plus fs disk tiers. Scheduler-side orchestration is inherited.

Only the platform-bound pieces are replaced.

- The worker. Upstream's copy engine is CUDA-only, and a torch alias of
  the KV cache goes stale because the MLX cache rebinds each layer's
  array on every write. MetalKVOffloadWorker moves blocks between the
  live MLX arrays and the host pool, gathering through the current list
  entry so reads are ordered after pending writes.
- The shared region. Upstream's is /dev/shm plus madvise. macOS has
  neither, so MetalSharedOffloadRegion is anonymous RAM shared within
  one process. That is why offloading requires the single-process
  executor.
- The fs tier. fcntl(F_NOCACHE) where upstream would use O_DIRECT,
  0o600 block files under a 0o700 root, and a blocks.noindex directory
  so Spotlight ignores the churn. The on-disk format is upstream's, byte
  for byte.

The platform hook translates --kv-offloading-size itself, because vLLM's
own translation runs after the hook and would select a connector this
platform cannot serve. Unsupported configurations fail at config time
with a specific message: the obj tier (NIXL has no macOS build),
pipeline parallelism, multi-process executors, hybrid and sliding-window
models, and draft-model speculative decoding.

The model runner keeps the connector step open across the
execute_model / sample_tokens split and closes it when the deferred
decode is submitted, so stores are queued before the next step's
forward can overwrite their blocks.

When KV cache events are enabled globally, enable_kv_events is set on
each tier; the reverse combination is refused. A KV-aware router then
sees BlockStored events without the two switches being lined up by hand.

76 tests: byte-exact round trips (fp16, bf16, TurboQuant), block
placement checked against upstream's pointer arithmetic on the real
region, the store fence across the deferred decode, the fs tier end to
end including upstream's failed-load negative cache, and the config
guards.

Nothing changes when offloading is off. Every new call site is behind
has_kv_transfer_group().

Supersedes vllm-project#530.

Signed-off-by: Robbie J <RobbieJ@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants