Repository navigation
Conversation
Setting `--kv-cache-memory-bytes` skipped memory profiling entirely, so the pinned size was never checked against what the model needs for itself. The rank starts, serves short requests, then OOMs mid-request once traffic reaches the activation peak. Under `gpu_memory_utilization` this stays invisible: the KV size is derived from a budget that is only a fraction of the device, and whatever the profile missed lands in the slice left unallocated. A pinned size has no such slice. Two things were missing from the pinned path: - The profile did not run at all. It now does, and the pinned size is checked against its result. If it does not fit, startup fails and names the largest size that does, rather than deferring the failure to a request. - `profile_run` forwards with `skip_attn=True`, so attention never runs and neither its activations nor the device memory its first launch takes are in the profile. Only when a size is pinned, build the minimal KV cache the CUDA graph profiling already uses and run a ladder of batch widths through it, and account for what the caching allocator holds beyond the live peak. On one rank of a pipeline-parallel Qwen3.8 serve, profiled non-KV usage goes from 57.52 GiB to 59.11 GiB against a measured 59.24 GiB, and a pinned 4.1 GiB that used to OOM under load is now refused at startup. The `gpu_memory_utilization` path is unchanged: every addition sits inside the pinned branch. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Signed-off-by: lesj0610 <lesj0610@godoiksan.org>
…mory-bytes-profiling Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
…mory-bytes-profiling Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
…mory-bytes-profiling Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
…mory-bytes-profiling Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
…mory-bytes-profiling Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
…mory-bytes-profiling Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
…mory-bytes-profiling Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
…mory-bytes-profiling Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
Contributor
|
This pull request has merge conflicts that must be resolved before it can be |
…bytes-profiling Only conflict is the import block in vllm/v1/worker/gpu_worker.py: vllm-project#57891 added `from fnmatch import filter as fnmatch_filter` at the same insertion point where this branch added `from functools import partial`. Both sides kept, order left to isort. Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
…mory-bytes-profiling Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
…mory-bytes-profiling Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
…mory-bytes-profiling Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
…bytes-profiling Upstream now passes randomize_inputs to profile_run(). The pinned-KV early return it touched is gone on this branch, so the argument lands on the one profiling call, ahead of the attention ladder. Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
…mory-bytes-profiling Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
6 tasks done
Contributor
|
This pull request has merge conflicts that must be resolved before it can be |
…bytes-profiling The only conflict is the pinned-KV early return this branch removes. vllm-project#58411 and vllm-project#58014 changed the profile_run call inside it, and both changes are already on the surviving call in the memory_profiling block, so the removal stands. Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
…fig fails The minimal KV cache the CUDA graph profiler bootstraps is built by forcing `num_gpu_blocks_override` to a block per sequence and restoring it afterwards. V1 restores it on the line after the call, so anything `get_kv_cache_config_ from_groups` raises leaves the profiling override in place, and the real KV cache sizing that follows reads it. V2 already wraps that in `try/finally`. Both now go through one helper that does the restore in a `finally`, which also drops the copy of the block-count computation each of them carried. V1's profiling teardown also walked the layers itself to detach the KV tensors and the quantized scale views. It skips a layer with no `kv_cache` attribute, so a layer holding only `_k_scale_cache`/`_v_scale_cache` keeps the profiling scales installed. It reuses `clear_layer_kv_caches()` now, which V2's teardown and the runner shutdown path already use, with the scale clearing lifted out of the `kv_cache` branch so those layers are covered too. Signed-off-by: lesj0610 <lesj0610@users.noreply.github.com>
Contributor
|
This pull request has merge conflicts that must be resolved before it can be |
…bytes-profiling The only conflict is the `vllm.v1.worker.utils` import list in gpu_model_runner.py. vllm-project#60517 dropped `is_residual_scattered_for_sp`, and this branch adds `build_minimal_kv_cache_config` and `clear_layer_kv_caches`, so the resolution keeps the two additions without the removed helper. Signed-off-by: lesj0610 <lesj0610@gmail.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
--kv-cache-memory-bytesnever checked the size it was handed. The branch returned beforememory_profilingran, logged "skipped memory profiling", and passed the number straight through. Nothing compared it against what the model needs for itself, so a size that is too large starts fine, serves short requests, and then OOMs in the middle of one once traffic reaches the activation peak.The same measurement gap exists under
gpu_memory_utilization, but it stays hidden there: the KV size is derived from a budget that is only a fraction of the device, so whatever the profile missed lands in the part the utilization fraction left unallocated. A pinned size has no such slack — the operator sized it to use what was free — so the pinned cache and the unprofiled activation peak compete for the same free memory.There is a second half to it. Even when the profile does run, it forwards with
skip_attn=True, because there is no KV cache yet to attend against. The attention activations, and the device memory those kernels take the first time they launch, are simply not in the figure. Harmless where slack absorbs it; not harmless when the cache was pinned to that figure.What this changes
Run the profile for a pinned size too, and check the value against the result. If it does not fit, fail at startup and report the largest size admitted by this profile. Refusing rather than quietly capping is deliberate: a pinned size is an explicit capacity decision, and silently reducing it hides the one number this option exists to control. Starting anyway is worse — the bytes are spent either way, and the failure just moves to a request.
When a size is pinned, extend the profile with attention: build the same minimal KV cache the CUDA graph profiling already uses, run a short ladder of batch widths through it, then tear it down. The ladder is not redundant with its widest rung, since the caching allocator keeps whole segments per size class and a served mix of widths costs more than the widest one alone. Then account for what the allocator holds beyond the live peak —
memory_profilingmeasuresallocated_bytes.all.peak, which is the right call where slack absorbs the difference and the wrong one where there is none.All of it sits inside the pinned branch; outside it, the only change is the shared minimal-KV helper below, which behaves the same when nothing fails.
The attention profile is best effort: if the minimal KV cache cannot be built, startup continues with the old estimate. Building it forces
num_gpu_blocks_overrideto one block per sequence, and V1 restored the override only on the line after the call, so a failure there left it in place for the real KV cache sizing; with a single GPU the engine reads the same config object. Both runner generations now build the minimal config through one helper that restores the override in afinally, as V2 already did. V1's profiling teardown reusesclear_layer_kv_caches(), which now clears quantized scale views independently ofkv_cache.What it deliberately does not do
The reported figure is an upper bound, not a safe setting, and the error message says so. The profile forwards against a minimal KV cache, so peaks that scale with context length are still outside it. For the serve described below, the unmargined bound was 3.53 GiB. A pinned size at that bound still OOMed under four concurrent requests, while 3.07 GiB completed the workload without failure. I did not measure where the real boundary between those two is, so no margin is applied here — the number is what the profile can account for, and the operator is told to leave room under it.
A percentage margin was the first thing I tried, and I dropped it. The shortfall scales with the model's activation footprint rather than with the cache, so a fraction of the cache size wastes memory wherever the cache is large; and a fraction fitted to a single measurement is not evidence, it just reproduces that measurement.
Test Plan
Two-rank pipeline-parallel serve with an unbalanced layer split and the KV cache pinned. The rank carrying the larger share is the one that runs out. Workload: text prompts from 1.6k to 259k tokens, images from 360 to 14k tokens, and mixed batches at concurrency 3 and 4.
Unit tests cover the override restore when the minimal config raises, and the V1 profiling teardown.
Test Result
Profiled non-KV usage on that rank, against an observed non-KV peak of 59.24 GiB under the same workload:
A pinned 4.1 GiB that used to start and OOM later is now refused at startup:
At a pinned size below the bound the same serve finishes the whole workload with no failures, including a 258,953-token prompt and four concurrent image requests.
pytest tests/v1/worker/test_cudagraph_memory_profiling.py tests/v1/worker/test_gpu_worker.py: 28 passed. With a forced failure in the V1 minimal config build,num_gpu_blocks_overrideused to keep the profiling block count after the catch; it is now restored.ruff checkandruff formatpass. Themypyhook reports five errors in this file; all five are present on the unmodified file at the same base and none are on the added lines.All results above come from local runs on one two-GPU host.