Repository navigation
[KV Offload] Add per-request max_load_tokens control - #55885
Conversation
|
Documentation preview: https://vllm--55885.org.readthedocs.build/en/55885/ |
|
This pull request has merge conflicts that must be resolved before it can be |
|
Thanks @albertoperdomo2 ! |
kv_load_tiers disable all external loadsmax_load_tokens control
|
This pull request has merge conflicts that must be resolved before it can be |
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
a54543d to
463fb87
Compare
|
@orozery can we kick off the CI? |
|
/ci run |
|
❌ This PR is 1 commit behind upstream |
|
✅ @albertoperdomo2, CI is now available for this PR.
|
Signed-off-by: Alberto Perdomo <aperdomo@redhat.com>
|
/ci run |
|
✅ Triggered Buildkite CI #89080 for commit |
|
/ci retry |
|
✅ No failed, timed-out, or expired jobs need retrying: https://buildkite.com/vllm/ci/builds/89080 |
|
/ci run |
|
✅ Triggered Buildkite CI #89192 for commit |
Purpose
This PR adds the experimental per-request
kv_transfer_params["max_load_tokens"], analogous tomax_offload_tokens.0disables external KV loading, so missing tokens are recomputed.Local GPU prefix-cache reuse and KV stores remain enabled.
kv_load_tierskeeps its existing semantics: it filters secondary tiers, while CPU remains available as a direct source and as the staging tier for secondary loads.skip_reading_prefix_cacheis unchanged.AI assistance (Codex) was used to inspect the implementation and draft tests and documentation.
Test Plan
max_load_tokens: 0with synchronous and asynchronous scheduling.skip_reading_prefix_cacheremain unchanged.Test Result
Pending.
Essential Elements of an Effective PR Description Checklist