Skip to content

Fixed incorrect host memory counting - #42039

Closed
nvjullin wants to merge 1 commit into
sgl-project:mainfrom
nvjullin:correct-host-memory-cap
Closed

nvjullin wants to merge 1 commit into
sgl-project:mainfrom
nvjullin:correct-host-memory-cap

Conversation

@nvjullin

@nvjullin nvjullin commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

Since #40135, the HiCache host-pool budget is bounded by the cgroup headroom limit - memory.current. memory.current includes page cache charged to the cgroup: the checkpoint the server just read (size of the model) and, in containerized Slurm jobs, the unpacked container image (58 GiB in a fresh job). That cache is freely reclaimable, and host MemAvailable (psutil's available) already counts it as available, but the cgroup bound counts it as used. Right after a large model loads, the check under-reports allocatable memory by roughly the checkpoint size and rejects pools the kernel would provide by reclaiming the cache. For example, GLM-5.2-NVFP4 at tp1 x pp4 on one GB300 node with --hicache-size 135 fails on every rank:

ValueError: Not enough host memory available. Requesting 135.00 GB but only have 122.49 GB free.

The same recipe passes or fails depending on the node, because the charged cache differs per job.

Modifications

  • _cgroup_memory_headroom now counts usage at each limited cgroup level as memory.current - (active_file + inactive_file), clamped at 0. On cgroup v1 it is memory.usage_in_bytes - (total_active_file + total_inactive_file). These counters cover the pages on the kernel's file-backed reclaim lists, which is what MemAvailable counts as reclaimable, so the host and cgroup bounds agree. Shared memory and tmpfs sit on the anonymous lists and stay counted as used. memory.stat's file and cache counters would wrongly count them as free.
  • Added a unit test covering v2, v1 (hierarchical total_* counters rather than the local ones), a shared-memory decoy, and the clamp. It fails on the old code.

GPU-free validation in a Slurm job (cgroup v2, 350 GiB limit, no swap), comparing the old and new code side by side, in GiB:

Step psutil available Old cgroup headroom New cgroup headroom
Baseline 358.1 323.7 348.7
Read 151 GiB of checkpoint shards 347.7 173.1 348.4
Add an 8 GiB file in /dev/shm 339.6 165.1 340.4

mlock stood in for pinned host memory: locking 315.4 GiB then succeeded, although the old bound allowed only 173.1 GiB. The kernel reclaimed the cache (175.3 to 22.8 GiB) and there was no out-of-memory kill. While the memory was held, the new headroom read 33.3 GiB against a predicted 32.0 GiB.

Accuracy Tests

N/A: model outputs are unaffected.

Speed Tests and Profiling

N/A: the only cost is one extra memory.stat read per limited cgroup level at host-pool init.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

🤖 Generated with Claude Code


CI States

Latest PR Test (Base): ✅ Run #36844629362
Latest PR Test (Extra): ⚠️ Not enabled -- add run-ci-extra label to opt in.
Latest PR Test (AMD ROCm 10): ❌ Run #36844629350

@nvjullin nvjullin changed the title fixed incorrect host memory counting Fixed incorrect host memory counting Oct 1, 2026
@github-actions github-actions Bot added the hicache Hierarchical Caching for SGLang label Oct 1, 2026
@nvjullin

nvjullin commented Oct 1, 2026

Copy link
Copy Markdown
Contributor Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Oct 1, 2026
@nvpohanh

nvpohanh commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

cc @xiezhq-hermann for vis

@nvpohanh

nvpohanh commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

All NV pipelines have passed. @xiezhq-hermann could you approve this? thanks!

@nvpohanh

nvpohanh commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Or @ispobock @alphabetc1 could you review/approve this? Because is this, we can't use the latest nightly container for some runs. Thanks!

@nvpohanh nvpohanh added the highest-priority CI: all three control labels, plus never batch-cancelled or stale-closed label Oct 2, 2026
@xiezhq-hermann xiezhq-hermann self-assigned this Oct 2, 2026
Venkat2811 added a commit to datacrunch-research/sglang that referenced this pull request Oct 3, 2026
…stream sgl-project#42039)

Squash of the kimi-k3 integration PR #12 (layer 5a): upstream sgl-project#42039 applied verbatim (host_memory.py and its unit test), nothing else.

Since sgl-project#40135 the host-pool budget is limit - memory.current, and memory.current includes the page cache charged to the cgroup, including the checkpoint the server just read. This counts that cache as reclaimable headroom so an explicitly sized HiCache pool is not refused at startup.

Retire when sgl-project#42039 merges. If sgl-project#41181 lands instead (same lines, different text), drop this layer on the rebase.

Upstream: sgl-project#42039

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@nvpohanh

nvpohanh commented Oct 5, 2026

Copy link
Copy Markdown
Collaborator

superseded by #42420

@nvpohanh nvpohanh closed this Oct 5, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hicache Hierarchical Caching for SGLang highest-priority CI: all three control labels, plus never batch-cancelled or stale-closed run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants