Skip to content

[KV Offload] Add per-request max_load_tokens control - #55885

Merged
orozery merged 10 commits into
vllm-project:mainfrom
albertoperdomo2:feat/opt-out-kv-load
Sep 16, 2026
Merged

orozery merged 10 commits into
vllm-project:mainfrom
albertoperdomo2:feat/opt-out-kv-load

Conversation

@albertoperdomo2

@albertoperdomo2 albertoperdomo2 commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

This PR adds the experimental per-request kv_transfer_params["max_load_tokens"], analogous to max_offload_tokens.

  • 0 disables external KV loading, so missing tokens are recomputed.
  • A positive value caps how many tokens can be loaded beyond the GPU-resident prefix. The cap is rounded down to a supported KV-cache boundary.
  • Omitting the field preserves uncapped loading. Invalid values are ignored with a warning.

Local GPU prefix-cache reuse and KV stores remain enabled.

kv_load_tiers keeps its existing semantics: it filters secondary tiers, while CPU remains available as a direct source and as the staging tier for secondary loads. skip_reading_prefix_cache is unchanged.

AI assistance (Codex) was used to inspect the implementation and draft tests and documentation.

Test Plan

  • Test max_load_tokens: 0 with synchronous and asynchronous scheduling.
  • Test positive caps, partial tails, alignment, and invalid values.
  • Verify that the store path and skip_reading_prefix_cache remain unchanged.
  • Run the focused tests on the final revision.
  • Validate the final API in a live multi-tier deployment.

Test Result

Pending.


Essential Elements of an Effective PR Description Checklist
  • Purpose
  • Test plan
  • Test results
  • Documentation update

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Claude Code Review

This pull request is from a fork — automated review is disabled. A repository maintainer can comment @claude review to run a one-time review.

@mergify

mergify Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

Documentation preview: https://vllm--55885.org.readthedocs.build/en/55885/

@mergify mergify Bot added documentation Improvements or additions to documentation kv-connector labels Sep 8, 2026
@mergify

mergify Bot commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @albertoperdomo2.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 10, 2026
@orozery

orozery commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

Thanks @albertoperdomo2 !
I think that if we want to allow disabling loading from any kind of offloading it's maybe better to introduce a new max_load_tokens (similar to max_offload_tokens).
kv_load_tiers assumes CPU is always included since all other tiers promote through it, so it's meaningless to have a non-empty list without CPU included.

@mergify mergify Bot removed the needs-rebase label Sep 10, 2026
@albertoperdomo2 albertoperdomo2 changed the title [KV Offload] Let empty kv_load_tiers disable all external loads [KV Offload] Add per-request max_load_tokens control Sep 11, 2026
@mergify

mergify Bot commented Sep 12, 2026

Copy link
Copy Markdown
Contributor

This pull request has merge conflicts that must be resolved before it can be
merged. Please rebase the PR, @albertoperdomo2.

https://docs.github.com/en/pull-requests/collaborating-with-pull-requests/working-with-forks/syncing-a-fork

@mergify mergify Bot added the needs-rebase label Sep 12, 2026
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
@albertoperdomo2

Copy link
Copy Markdown
Contributor Author

@orozery can we kick off the CI?

Comment thread tests/v1/kv_connector/unit/offloading_connector/test_scheduler.py Outdated
@orozery

orozery commented Sep 15, 2026

Copy link
Copy Markdown
Contributor

/ci run

@github-actions

Copy link
Copy Markdown

❌ This PR is 1 commit behind upstream main. Your branch must contain every commit currently on upstream main. No new CI build was started. Merge or rebase onto the latest main, then rerun /ci run. To test this branch at your own risk, use /ci run --allow-stale.

@orozery orozery added the ready ONLY add when PR is ready to merge/full CI is needed label Sep 15, 2026
@github-actions

Copy link
Copy Markdown

✅ @albertoperdomo2, CI is now available for this PR.

  • /ci run starts upstream CI; /amd-ci run starts AMD CI only.
  • Your branch must contain every commit currently on its upstream target branch. Merge or rebase onto the latest target branch, then rerun the command. Append --allow-stale to a run command to test an outdated branch at your own risk.
  • /ci retry retries failed jobs in the CI build for the current PR head. If the current head has no CI build, it starts a new CI build for the current head containing only jobs that failed in the latest earlier CI build for this PR.
  • /amd-ci retry retries failed jobs in AMD CI for the current PR head. Use /amd-ci run when the current head has no AMD CI build.
  • /ci cancel cancels scheduled or running CI builds for this PR branch; /amd-ci cancel does the same for AMD CI only.

@albertoperdomo2

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #89080 for commit 7b5253a9bf2d.

@albertoperdomo2

Copy link
Copy Markdown
Contributor Author

/ci retry

@github-actions

Copy link
Copy Markdown

✅ No failed, timed-out, or expired jobs need retrying: https://buildkite.com/vllm/ci/builds/89080

@albertoperdomo2

Copy link
Copy Markdown
Contributor Author

/ci run

@github-actions

Copy link
Copy Markdown

✅ Triggered Buildkite CI #89192 for commit a5d9908b6f90.

@orozery
orozery merged commit 75dc588 into vllm-project:main Sep 16, 2026
48 checks passed
@albertoperdomo2
albertoperdomo2 deleted the feat/opt-out-kv-load branch September 16, 2026 05:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation kv-connector ready ONLY add when PR is ready to merge/full CI is needed

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants