[HiCache] write_back policy refinement - #29817
Conversation
There was a problem hiding this comment.
Code Review
This pull request removes the deprecation warning for the write_back policy in the cache controller, cleans up unused variables, and introduces a staging mechanism for write-back eviction in the hierarchical radix cache. However, a critical race condition was identified in the new eviction path: device memory is freed immediately after queuing the write-back, before the asynchronous copy operation completes. This can lead to data corruption, and it is recommended to defer freeing the device memory until after the write-back synchronization is finished in the flush_staged function.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
…last-leaf fix Upstream sgl-project#29860 (a375e9f, merged). Adapted to our fork's inline _evict_swa in schedule_batch.py (upstream lives in mem_cache/common.py:free_swa_out_of_window_slots). Before: _evict_swa evicted up to pre_len - sliding_window_size - page_size under the assumption that "extra page keeps the frontier below the insert boundary". The env flag SGLANG_OPT_SWA_EVICT_DROP_PAGE_MARGIN toggled whether to apply the extra -page_size subtraction. After: gate on tree_cache.is_chunk_cache() (replicated tree-cache property, uniform across ranks) instead of the env flag: - chunk-cache: no radix tree -> no tombstone-leaf concern; evict up to the window boundary (pre_len - sliding_window_size). - radix: keep max(window, page). The trailing floor page-aligns the frontier, and subtracting at least one page keeps the frontier below the insert boundary (page_floor(seq_len)) so the last leaf is never all-tombstone. This is the case sgl-project#29860 was fixing. The env var SGLANG_OPT_SWA_EVICT_DROP_PAGE_MARGIN is now unused but left in environ.py for backward compatibility (harmless no-op). P2 batch, other PRs skipped this round: - sgl-project#27550 fix(hiradix): wait for extra pool IO - target code path (the completed+pool_transfers_done branch in can_terminate_prefetch) was simplified away by our PR sgl-project#27010 port; the pool_transfers_done invariant is now enforced through the ack-queue ordering instead. - sgl-project#28422 decode-hicache _storage_hit_query - the "pre-query" feature sgl-project#28422 patches does not exist in our fork. - sgl-project#29887 [PP] get kv_buffer_shape - target file (eager_runner.py) does not exist in our fork. - sgl-project#29817 write_back policy refinement - refactor, not a bug fix; our fork's evict() has diverged from the upstream shape and porting cleanly is out of scope for this batch. - sgl-project#28614 remove large host mem constraint - depends on sgl-project#29817's evict refactor for the bulk of its diff, and the standalone piece flips prefetch_capacity_limit from max(0, 0.8*(host-device)) to 0.5*host, a subtle memory-budget semantics change we won't ship silently. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Motivation
Modifications
Accuracy Tests
Speed Tests and Profiling
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): ❌ Run #28541126226
Latest PR Test (Extra): ❌ Run #28541125898