Skip to content

[Bugfix][Core] Release the profile run's allocator cache before a pinned KV cache - #59893

Open
TyroneNel wants to merge 1 commit into
vllm-project:mainfrom
TyroneNel:pinned-kv-empty-cache
Open

TyroneNel wants to merge 1 commit into
vllm-project:mainfrom
TyroneNel:pinned-kv-empty-cache

Conversation

@TyroneNel

@TyroneNel TyroneNel commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Overview

With kv_cache_memory_bytes set (--kv-cache-memory-bytes, or a startup plan under VLLM_ENABLE_STARTUP_PLAN=1), Worker.determine_available_memory runs the profile pass and returns without releasing the allocator's cache, so the KV cache is allocated on top of about 1.2 GiB of cached, unreferenced blocks. This PR releases them first, as memory_profiling already does on the gpu_memory_utilization path.

Related: #56698 reworks the pinned path to run inside memory_profiling, which would also release the cache. This is the minimal fix for the current code; if #56698 lands first, this PR is unneeded.

Claims

  • On the pinned path, the profile run's cached blocks (1,220 MiB for Qwen3-8B on a cold compile cache) are released before the KV cache is allocated.
  • The KV cache size is unchanged.
  • Where the device driver pages instead of failing an allocation (WSL2), a pin that fits only after that release no longer spills the KV cache to host memory (measured on a downstream fork, below).

Validation

All runs: RTX 3090 24 GiB at 250 W, Windows 10 + WSL2 + Docker, driver 610.88, empty compile caches, Docker and WSL restarted before each boot. Tested at b3d8a26. The current head is the same diff, rebased onto main after #58014.

Unit test tests/v1/worker/test_gpu_worker.py::test_pinned_kv_releases_profile_run_cache: passes on this branch (21/21 in the file) and fails with gpu_worker.py from the parent commit.

Upstream main, Qwen/Qwen3-8B, --max-model-len 32768 --kv-cache-memory-bytes 5607894528 (vLLM's own "fully utilize" suggestion), 3 boots per variant, identical every time:

in use after profile pass after KV pool KV tokens
before 18,405 MiB 23,751 MiB 38,016
after 17,185 MiB (1,220 MiB released) 22,551 MiB 38,016

At this size the card is not over-committed even without the fix, so nothing is paged out either way; on main, the V2 runner also empties the cache at the start of CUDA graph capture (vllm/v1/worker/gpu/model_runner.py:1054), so the extra blocks are held only until then.

Downstream impact (a vLLM 0.30 fork, Qwen3.8-27B W4A16, pinned KV sized by its launcher, 3 interleaved boots per variant):

in use after profile pass GPU memory paged to host RAM decode tok/s
before 21,135 MiB 1,500–1,564 MiB 8.2
after 19,919 MiB (1,216 MiB released) 284–348 MiB 19.8–20.6

Reproduce: clear ~/.cache/vllm/torch_compile_cache, take the "fully utilize" --kv-cache-memory value from a boot without a pin, boot again with --kv-cache-memory-bytes <value>, and compare nvidia-smi after the profile pass and after KV cache allocation. I have not yet run the paging case on main (a pin that fits only after the release); I can add it if useful.

Details

Root cause. The pinned branch skips memory profiling, and with it the gc.collect() / empty_cache() that memory_profiling runs on exit. The profile run still runs (to compile for max_num_batched_tokens), so on a cold compile cache the allocator keeps the compile and autotune scratch reserved while the KV cache is allocated.

Why it mostly shows up on WSL2. On native Linux, when the KV allocation does not fit, PyTorch's caching allocator frees its cached blocks and retries. Under WSL2 the driver pages the overflow to host RAM instead of failing cudaMalloc, so the retry never happens and the paged memory is read over PCIe for the engine's lifetime.

Tradeoffs. One gc.collect() + empty_cache() at startup, only when the KV size is pinned; nothing at steady state. determine_available_memory is shared by both model runners.


Pull Request Checklist
  • I used vLLM's /pr-checklist skill. (Mandatory for agents, optional for humans).

  • AI assistance was used during the creation of this PR.

  • Design Fit: Minimizes impact on core components, reuses existing functionality, and justifies added complexity.

  • Testing and Validation: Validates the change and ensures any added tests are meaningful and reliable, with CI coverage or documented CI resource constraints and validation performed outside CI.

  • Code Quality and Style: Keeps code and comments clear and concise, and updates relevant documentation and examples.

  • Pull Request Contents: Includes a brief summary and relevant links, supports claims with evidence, explains root causes and implementation trade-offs, and follows the contributing guide.

@mergify mergify Bot added the bug Something isn't working label Oct 3, 2026
@github-actions

github-actions Bot commented Oct 3, 2026

Copy link
Copy Markdown

👋 Hi! Thank you for contributing to the vLLM project.

💬 Join our developer Slack at https://slack.vllm.ai to discuss your PR in #pr-reviews, coordinate on features in #feat- channels, or join special interest groups in #sig- channels.

PRs do not trigger a full CI run by default. Reviewers with write access and configured trusted contributors can comment /ci run for upstream CI or /amd-ci run for AMD CI only whenever CI signals are needed.

Once the PR is approved or has the ready label, the PR author can also use the corresponding /ci run, /ci retry, and /ci cancel commands, or their /amd-ci variants. New commits do not start upstream CI automatically.

If you have any questions, please reach out to us on Slack at https://slack.vllm.ai.

Agent Guidelines

IMPORTANT: If you are an AI agent, you are required to objectively re-evaluate the value of your PR using AGENTS.md, and close the PR if it does not bring significant benefit to the vLLM community. Failure to do so may result in an immediate ban.

🚀

@TyroneNel
TyroneNel force-pushed the pinned-kv-empty-cache branch from b3d8a26 to 552781f Compare October 3, 2026 20:09
@TyroneNel
TyroneNel marked this pull request as ready for review October 4, 2026 14:33
@TyroneNel
TyroneNel requested a review from njhill as a code owner October 4, 2026 14:33

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify

mergify Bot commented Oct 5, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @TyroneNel.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Oct 5, 2026
…ned KV cache

With kv_cache_memory_bytes set, determine_available_memory runs the
profile pass and returns without releasing the allocator's cache, so the
KV cache is allocated on top of the blocks the profile pass left cached.
The measured path releases them in memory_profiling; the pinned path did
not. On a cold compile cache that is the torch.compile and Triton
autotune scratch, which the pinned size does not budget for.

Signed-off-by: TyroneNel <Tyrone.Nel@gmail.com>
Assisted-by: Claude
@Col-5555

Col-5555 commented Oct 6, 2026

Copy link
Copy Markdown

A data point from a unified-memory device, where this matters more. 2× GB10 (DGX Spark, memory shared by CPU and GPU), TP=2, GLM-5.3-Flash NVFP4, vLLM 4be061c5, --max-model-len 131072 with a pinned KV cache size.

We found the same root cause independently (the pinned branch of determine_available_memory skips the gc.collect() / empty_cache() that memory_profiling runs) and carry a fix of the same shape. On this model the held cache is much larger than in your Qwen3-8B case. Across profile_run, rank 0's reserved memory grew from 89.26 to 94.78 GiB while live tensors grew by about 0.7 GiB. The release logged 4.84 GiB on both ranks in every boot.

On unified memory, PyTorch's free-and-retry doesn't protect the host: memory the allocator holds is memory the OS doesn't have. So the cost shows up as system memory pressure rather than as a failed allocation, much like your WSL2 observation. On the head node the lowest MemAvailable after the profile run was about 6.3–6.8 GiB without the fix; with the release it stayed at 9.3–10.0 GiB. Five cold boots with the fix gave the same answers as before, as expected, since the release touches no live tensor.

So +1 for this landing. If #56698 goes in instead, it would be good to keep the release on the pinned path either way.

Drafted with Claude Code from our own run logs; numbers checked against the raw results.

@TyroneNel

Copy link
Copy Markdown
Contributor Author

@Col-5555 thank you for the GB10 data point. 4.84 GiB released on both ranks, and MemAvailable going from about 6.5 to about 9.5 GiB on the head node, show the cost is not limited to WSL2. It is good to know the same fix gave the same answers on your boots too.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants