Repository navigation
Conversation
Contributor
🏷️ CI GuideRuns automatically on every eligible PR before approval:
Heavy model tests:
|
This was referenced Sep 18, 2026
yhl-amd
added a commit
that referenced
this pull request
Sep 22, 2026
# Conflicts: # atom/kv_transfer/offload/chunked_scheduler.py
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
OFFLOAD_SAVE_POLICYand all save-queue round-robin state/branchesOFFLOAD_MAX_PENDING_SAVESguard, counting a DSV4 PAGE+SLOT checkpoint as one worker operationBlockManagerand enforce a real physical PAGE-block budget:OFFLOAD_SAVE_MAX_PINNED_RATIO(default0.20, maximum0.30)OFFLOAD_SAVE_MAX_PINNED_BLOCKSmin(floor(local_blocks × ratio), absolute_limit)This ports the reusable priority/budget work into #2246 without merging the #2250-only native DSV4/LMCache-MP state scheduler changes.
Why
The original DSV4 fix bounded the number of scheduler-dispatched saves, eliminating worker-side rejection storms, but a count-only queue still lets a few large finished requests retain an unbounded fraction of the GPU KV pool. It also treats every save as equally valuable.
The admission layer bounds actual scheduler-local physical blocks and spends that budget on the saves with the highest expected reuse value. A newly valuable prefix can replace lower-value work only while that work is still committed and undispatched. Victims are simulated as a complete set before state is mutated, so a candidate that still cannot fit evicts nothing.
Round-robin admission has been removed rather than retained as a compatibility mode. This avoids maintaining two different completion, deferral, and cleanup semantics: every chunked offload save now follows the same
candidate -> committed -> inflightlifecycle.The budget is DP-local. TP ranks shard the contents of the same logical block IDs, so the capacity is neither aggregated across DP pools nor multiplied by TP world size.
Existing DSV4 admission result
The count-bound change already removed the observed save rejection/rollback storm in the TP8 A/B run:
After that fix: 417 scheduler save operations, 10,659,072 saved tokens, zero
max_pending_savesrejections, zero PAGE watermark rollbacks, and pending saves mean/peak 0.38/2. The run recorded no LMCache loads, so it validates removal of the save storm rather than a CPU-load speedup.Validation
python3 -m py_compile: passedgit diff --check: passedThe current local host does not provide CPU PyTorch, so the full repository suite is left to PR CI; the targeted connector tests were run with import-only dependency stubs.