Skip to content

[HiCache] Account for reclaimable page cache in cgroup headroom - #41181

Closed
paulzhang-tm wants to merge 2 commits into
sgl-project:mainfrom
paulzhang-tm:codex/hicache-cgroup-page-cache
Closed

paulzhang-tm wants to merge 2 commits into
sgl-project:mainfrom
paulzhang-tm:codex/hicache-cgroup-page-cache

Conversation

@paulzhang-tm

@paulzhang-tm paulzhang-tm commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

HiCache can reject startup or unnecessarily shrink its automatically sized host pool when clean file cache remains charged to a container or an ancestor cgroup. The cgroup-aware check introduced in #40135 uses limit - usage, but Linux can reclaim clean cached file pages to satisfy new allocations. Those pages can remain charged after the job that read them exits.

For example, a 900 GiB cgroup with 850 GiB charged usage, including 600 GiB of clean file cache, previously reported only 50 GiB of headroom. This change estimates 650 GiB, still capped by host available memory and all applicable ancestor limits.

Modifications

  • Read memory.stat at each constrained cgroup and subtract clean file-LRU pages from charged usage: max(0, inactive_file + active_file - dirty - writeback). Use the hierarchical total_* counters for cgroup v1.
  • Keep shared memory charged: it lives on the anonymous LRU, including shared anonymous mappings used by HiCache. Keep dirty and writeback pages charged as well.
  • Preserve the full-usage fallback when memory.stat is absent, and clamp inconsistent counter snapshots. Use one adjusted usage reading for both memory.max and memory.high at each level.
  • Extend the existing CPU tests for clean cache, dirty/writeback pages, shared memory, ancestor cache, cgroup-v1 hierarchical counters, memory.high, and counter clamping.

This changes the headroom estimate; allocation behavior, safety reserves, and sizing options remain unchanged.

Accuracy Tests

All 14 host-memory unit tests passed:

PYTHONPATH=python python test/registered/unit/mem_cache/test_host_memory.py -v

All applicable pre-commit checks passed for the two changed files:

pre-commit run --files python/sglang/srt/mem_cache/host_memory.py test/registered/unit/mem_cache/test_host_memory.py

No model-output changes; no GPU/server launch was run.

Speed Tests and Profiling

Not run: the change is limited to host-memory sizing during initialization.

Checklist

  • Format code with pre-commit.
  • Add focused CPU regression tests.
  • Update code documentation; no user-facing options change.
  • Provide accuracy and speed benchmarks (not applicable to this accounting fix).
  • Follow the SGLang code style guidance.

CI States

Latest PR Test (Base): ✅ Run #36398928434
Latest PR Test (Extra): ✅ Run #36398927898
Latest PR Test (AMD ROCm 10): ❌ Run #36398928378

Exclude clean file LRU pages from charged usage when estimating host
memory available under cgroup limits. Keep dirty, writeback, and shared
memory charged, and preserve ancestor and host-availability bounds.

Extend the existing host-memory tests for both cgroup versions, ancestor
page cache, memory.high, and inconsistent counter snapshots.

Co-authored-by: James Sun <jamessyt@gmail.com>
@github-actions github-actions Bot added the hicache Hierarchical Caching for SGLang label Sep 24, 2026
@alphabetc1 alphabetc1 added run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) labels Sep 28, 2026
@alphabetc1

Copy link
Copy Markdown
Collaborator

/rerun-test -c

@github-actions

github-actions Bot commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test -c:

🚀 ubuntu-latest (1 test): ✅ View workflow run

cd test/ && python3 registered/unit/mem_cache/test_host_memory.py

Venkat2811 added a commit to datacrunch-research/sglang that referenced this pull request Oct 3, 2026
…stream sgl-project#42039)

Squash of the kimi-k3 integration PR #12 (layer 5a): upstream sgl-project#42039 applied verbatim (host_memory.py and its unit test), nothing else.

Since sgl-project#40135 the host-pool budget is limit - memory.current, and memory.current includes the page cache charged to the cgroup, including the checkpoint the server just read. This counts that cache as reclaimable headroom so an explicitly sized HiCache pool is not refused at startup.

Retire when sgl-project#42039 merges. If sgl-project#41181 lands instead (same lines, different text), drop this layer on the rebase.

Upstream: sgl-project#42039

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@alphabetc1

Copy link
Copy Markdown
Collaborator

please resolve the conflicts @paulzhang-tm

@alphabetc1

Copy link
Copy Markdown
Collaborator

fixed by #42420

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bypass-fail-fast CI: a failing job no longer aborts its siblings (lint still gates) hicache Hierarchical Caching for SGLang run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants