Skip to content

[HiCache] write_back policy refinement - #29817

Merged
xiezhq-hermann merged 3 commits into
mainfrom
hicache-write-back
Jul 2, 2026
Merged

xiezhq-hermann merged 3 commits into
mainfrom
hicache-write-back

Conversation

@xiezhq-hermann

@xiezhq-hermann xiezhq-hermann commented Jul 1, 2026

Copy link
Copy Markdown
Collaborator

Motivation

Modifications

Accuracy Tests

Speed Tests and Profiling

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #28541126226
Latest PR Test (Extra): ❌ Run #28541125898

@xiezhq-hermann xiezhq-hermann self-assigned this Jul 1, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request removes the deprecation warning for the write_back policy in the cache controller, cleans up unused variables, and introduces a staging mechanism for write-back eviction in the hierarchical radix cache. However, a critical race condition was identified in the new eviction path: device memory is freed immediately after queuing the write-back, before the asynchronous copy operation completes. This can lead to data corruption, and it is recommended to defer freeing the device memory until after the write-back synchronization is finished in the flush_staged function.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread python/sglang/srt/mem_cache/hiradix_cache.py Outdated
@stmatengss stmatengss self-assigned this Jul 1, 2026
@xiezhq-hermann
xiezhq-hermann merged commit f19246e into main Jul 2, 2026
265 of 312 checks passed
@xiezhq-hermann
xiezhq-hermann deleted the hicache-write-back branch July 2, 2026 19:16
hzwzwzw added a commit to hzwzwzw/sglang that referenced this pull request Jul 13, 2026
…last-leaf fix

Upstream sgl-project#29860 (a375e9f, merged). Adapted to our
fork's inline _evict_swa in schedule_batch.py (upstream lives in
mem_cache/common.py:free_swa_out_of_window_slots).

Before: _evict_swa evicted up to pre_len - sliding_window_size - page_size
under the assumption that "extra page keeps the frontier below the insert
boundary". The env flag SGLANG_OPT_SWA_EVICT_DROP_PAGE_MARGIN toggled
whether to apply the extra -page_size subtraction.

After: gate on tree_cache.is_chunk_cache() (replicated tree-cache
property, uniform across ranks) instead of the env flag:
- chunk-cache: no radix tree -> no tombstone-leaf concern; evict up to
  the window boundary (pre_len - sliding_window_size).
- radix: keep max(window, page). The trailing floor page-aligns the
  frontier, and subtracting at least one page keeps the frontier below
  the insert boundary (page_floor(seq_len)) so the last leaf is never
  all-tombstone. This is the case sgl-project#29860 was fixing.

The env var SGLANG_OPT_SWA_EVICT_DROP_PAGE_MARGIN is now unused but
left in environ.py for backward compatibility (harmless no-op).

P2 batch, other PRs skipped this round:
- sgl-project#27550 fix(hiradix): wait for extra pool IO - target code path (the
  completed+pool_transfers_done branch in can_terminate_prefetch) was
  simplified away by our PR sgl-project#27010 port; the pool_transfers_done
  invariant is now enforced through the ack-queue ordering instead.
- sgl-project#28422 decode-hicache _storage_hit_query - the "pre-query" feature
  sgl-project#28422 patches does not exist in our fork.
- sgl-project#29887 [PP] get kv_buffer_shape - target file (eager_runner.py) does
  not exist in our fork.
- sgl-project#29817 write_back policy refinement - refactor, not a bug fix; our
  fork's evict() has diverged from the upstream shape and porting
  cleanly is out of scope for this batch.
- sgl-project#28614 remove large host mem constraint - depends on sgl-project#29817's evict
  refactor for the bulk of its diff, and the standalone piece flips
  prefetch_capacity_limit from max(0, 0.8*(host-device)) to 0.5*host,
  a subtle memory-budget semantics change we won't ship silently.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Chronostasys pushed a commit to MindLab-Research/sglang that referenced this pull request Aug 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants