Skip to content

[Feature] Add SLRU eviction policy & fix RadixCache hit_count bug - #18843

Merged
ispobock merged 14 commits into
sgl-project:mainfrom
liubiyongge:feat/slru-eviction
Mar 8, 2026
Merged

ispobock merged 14 commits into
sgl-project:mainfrom
liubiyongge:feat/slru-eviction

Conversation

@liubiyongge

Copy link
Copy Markdown
Contributor

Motivation

In RAG (Retrieval-Augmented Generation) and batch processing scenarios, the system often faces "scan" workloads: large volumes of one-time requests (e.g., retrieving many documents). Under the default LRUStrategy, these one-time requests flush valuable hot data (e.g., shared System Prompts, multi-turn dialog history) out of the KV cache. This leads to frequent re-computation (Cache Miss) and high tail latency.

While implementing SLRU (Segmented LRU) to solve this, I discovered a critical bug in RadixCache: the node.hit_count was never incremented during match_prefix or insert operations. This rendered any frequency-based policies (like the existing LFU) ineffective.

Modifications

  1. New Strategy (sglang/srt/mem_cache/eviction_policy.py):

    • Implemented SLRUStrategy (Segmented LRU).
    • Probationary Segment: Items with hit_count < k (default 2). Evicted first.
    • Protected Segment: Items with hit_count >= k. Protected from scan noise.
  2. Critical Bug Fix (sglang/srt/managers/router/radix_cache.py):

    • Added logic to increment node.hit_count in match_prefix (on cache hit) and insert (on new node creation). This fixes the underlying mechanism for LFU/SLRU.
  3. Bug Fix (sglang/srt/server_args.py):

    • Fixed a typo: add_radix_eviction_policy_choices -> radix_eviction_policy. This fixes an AttributeError when accessing ServerArgs.radix_eviction_policy.
    • Registered 'slru' as a valid choice for the CLI argument --radix-eviction-policy.

Accuracy Tests

This PR affects the cache eviction policy and memory management logic. It does not modify the model architecture, kernel implementation, or arithmetic precision. Therefore, the model output accuracy remains unaffected.

Benchmarking and Profiling

I performed a "Scan Attack" benchmark to verify Scan Resistance.

  • Hardware: NVIDIA A10 (24GB)
  • Constraint: mem_fraction_static=0.78 (Simulating memory pressure with ~11k token capacity).
  • Workload:
    • Hot Data: 1 VIP prompt (~1800 tokens), accessed 10 times (Warmup).
    • Scan Noise: 2000 unique requests (~800 tokens each) sent sequentially to flush the cache.

Results:

Metric LRU (Baseline) SLRU (Ours) Improvement
Warmup Latency 72.18 ms 71.84 ms ~0%
After Scan Latency 549.75 ms (Evicted) 73.06 ms (Survived) 7.6x Faster
Cache Status Miss (Recompute) Hit (Protected) ✅ Success

The results demonstrate that SLRU effectively isolates hot data from scan traffic, maintaining consistent low latency (TTFT).

Checklist

Review Process

  1. Ping Merge Oncalls to start the PR flow. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • /tag-run-ci-label, /rerun-failed-ci, /tag-and-rerun-ci
  4. After green CI and required approvals, ask Merge Oncalls to merge.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @liubiyongge, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly improves cache management for RAG and batch processing by introducing the Segmented LRU (SLRU) eviction policy. This policy is designed to prevent 'scan' workloads from flushing valuable hot data, thereby reducing cache misses and tail latency. A crucial underlying bug in the RadixCache's hit_count mechanism, which previously hampered the effectiveness of frequency-based policies, has also been fixed. These changes collectively lead to more robust and efficient memory utilization in scenarios with mixed access patterns.

Highlights

  • New Eviction Policy: Implemented a new SLRUStrategy (Segmented LRU) eviction policy to enhance cache performance, particularly for 'scan' workloads in RAG and batch processing, by protecting frequently accessed data.
  • RadixCache Hit Count Fix: Resolved a critical bug in RadixCache where node.hit_count was not being incremented, which previously rendered frequency-based eviction policies (like LFU and the new SLRU) ineffective.
  • CLI Argument Update: Updated server_args.py to correctly register and allow selection of the new slru eviction policy via command-line arguments, fixing a potential AttributeError.

🧠 New Feature in Public Preview: You can now enable Memory to help Gemini Code Assist learn from your team's feedback. This makes future code reviews more consistent and personalized to your project's style. Click here to enable Memory in your admin console.

Changelog
  • python/sglang/srt/mem_cache/evict_policy.py
    • Added the SLRUStrategy class, which defines a segmented LRU eviction logic based on a protected threshold and last access time.
  • python/sglang/srt/mem_cache/radix_cache.py
    • Imported the new SLRUStrategy.
    • Integrated the SLRUStrategy into the RadixCache initialization when the 'slru' policy is selected.
    • Added logic to increment child.hit_count within the _match_prefix_helper method during cache hits, ensuring frequency tracking is accurate.
  • python/sglang/srt/server_args.py
    • Updated the RADIX_EVICTION_POLICY_CHOICES list to include 'slru' as a valid option for the radix eviction policy.
Activity
  • No specific activity (comments, reviews, or progress updates) has been recorded for this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

The pull request introduces the SLRU (Segmented LRU) eviction policy to improve scan resistance in the KV cache and aims to fix a bug where hit_count was not being incremented. While the addition of the SLRU strategy and the increment logic in _match_prefix_helper are correct, the implementation of frequency tracking is incomplete. Specifically, the hit_count is not inherited during node splits, and it is not incremented during insertions. These omissions will lead to inaccurate frequency data, particularly for shared prefixes which are critical for the effectiveness of SLRU and LFU policies. Additionally, some user-facing error messages and help texts were not updated to include the new policy.

Comment thread python/sglang/srt/mem_cache/radix_cache.py Outdated
Comment thread python/sglang/srt/mem_cache/radix_cache.py
- Added `SLRUStrategy` (Segmented LRU) to `eviction_policy.py` for scan resistance.
- Fixed a critical bug in `radix_cache.py` where `hit_count` was never incremented during match/insert.
- Fixed a typo in `server_args.py` (`add_radix_eviction_policy_choices` -> `radix_eviction_policy`) causing AttributeError.
- Registered 'slru' as a valid choice for `--radix-eviction-policy`.

Benchmarks show SLRU reduces tail latency by 7.6x (550ms -> 73ms) under heavy RAG scan workloads compared to LRU.

@hzh0425 hzh0425 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good

Could you add a ci accuracy test ?

@hzh0425

hzh0425 commented Feb 23, 2026

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Feb 23, 2026
liubiyongge and others added 3 commits February 23, 2026 22:03
- Create unit test to verify SLRU eviction policy works correctly
- Test includes setup and execution phases to validate that:
  * High frequency access keys are retained in cache
  * Low frequency access keys are evicted when capacity is exceeded
- Use realistic cache sizes to trigger actual eviction behavior
- Validate through device indices presence in match results
Comment thread python/sglang/srt/mem_cache/radix_cache.py Outdated
  For partial matches (prefix_len < len(child.key)), only the newly created
  split node should have its hit_count incremented, not the child node.
  Previously, hit_count was incremented unconditionally before checking
  the match type.
Comment thread python/sglang/srt/mem_cache/radix_cache.py Outdated
Comment thread python/sglang/srt/mem_cache/radix_cache.py Outdated
liubiyongge and others added 2 commits March 1, 2026 21:47
  Avoid self-referencing hit_count inflation where chunked requests
  increment hit_count on nodes they created in previous chunks. Add a
  new  helper method that conditionally updates
  hit_count based on the chunked flag.
@hzh0425

hzh0425 commented Mar 5, 2026

Copy link
Copy Markdown
Collaborator

@liubiyongge

Copy link
Copy Markdown
Contributor Author

@liubiyongge could you help check the failde ci

https://github.com/sgl-project/sglang/actions/runs/22570301442/job/65376731131?pr=18843#step:5:5358

@hzh0425

The test test_hicache_storage_file_backend.py::TestHiCacheStorageAccuracy::test_eval_accuracy failed in CI, but passed locally. Despite a careful code review, the issue remains unclear.

Uploading image.png…

@ispobock
ispobock merged commit cc73355 into sgl-project:main Mar 8, 2026
154 of 165 checks passed
liubiyongge added a commit to liubiyongge/sglang that referenced this pull request Mar 13, 2026
Wangzheee pushed a commit to Wangzheee/sglang that referenced this pull request Mar 21, 2026
JustinTong0323 pushed a commit to JustinTong0323/sglang that referenced this pull request Apr 7, 2026
fzyzcjy added a commit to fzyzcjy/sglang that referenced this pull request Jun 1, 2026
PR sgl-project#18843 added the chunked guard so a chunked request does not inflate
hit_count on radix nodes it created, but only the middle-chunk insert
(scheduler.py) passed chunked=True. The final prefill chunk caches the
unfinished request via batch_result_processor without the flag, so its
prefix nodes were bumped once here and again by cache_finished_req,
yielding hit_count=2. Pass chunked=True so each node is bumped exactly
once on completion.
fzyzcjy added a commit to fzyzcjy/sglang that referenced this pull request Jun 1, 2026
…ng it

Revert the earlier chunked=True at the prefill-completion cache site in
batch_result_processor: that branch runs for every request finishing prefill
(chunked final chunk and non-chunked single prefill alike), so forcing
chunked=True suppressed the legitimate hit_count bump on the general prefill
path, not just chunked self-inserts.

A request inserts its prompt prefix twice -- via cache_unfinished_req when
prefill completes and via cache_finished_req on finish -- while decode nodes are
inserted once, so prompt nodes settle at hit_count 2 and decode nodes at 1,
identically for chunked and non-chunked prefill. sgl-project#18843's middle-chunk
chunked=True still prevents the unbounded early-node inflation.

Replace the old all==1 expectation with side-by-side chunked and non-chunked
tests asserting the hit_count value set is exactly {1, 2}.
Chronostasys pushed a commit to MindLab-Research/sglang that referenced this pull request Aug 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants