Repository navigation
Conversation
nvjullin
requested review from
Ying1123,
alphabetc1,
hanming-lu,
hnyls2002,
huangtingwei9988,
hzh0425,
ispobock,
merrymercy,
xiezhq-hermann and
yizhang2077
as code owners
October 1, 2026 09:43
Contributor
Author
|
/tag-and-rerun-ci |
Collaborator
|
cc @xiezhq-hermann for vis |
Collaborator
|
All NV pipelines have passed. @xiezhq-hermann could you approve this? thanks! |
nvpohanh
approved these changes
Oct 2, 2026
Collaborator
|
Or @ispobock @alphabetc1 could you review/approve this? Because is this, we can't use the latest nightly container for some runs. Thanks! |
Venkat2811
added a commit
to datacrunch-research/sglang
that referenced
this pull request
Oct 3, 2026
…stream sgl-project#42039) Squash of the kimi-k3 integration PR #12 (layer 5a): upstream sgl-project#42039 applied verbatim (host_memory.py and its unit test), nothing else. Since sgl-project#40135 the host-pool budget is limit - memory.current, and memory.current includes the page cache charged to the cgroup, including the checkpoint the server just read. This counts that cache as reclaimable headroom so an explicitly sized HiCache pool is not refused at startup. Retire when sgl-project#42039 merges. If sgl-project#41181 lands instead (same lines, different text), drop this layer on the rebase. Upstream: sgl-project#42039 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
3 of 5 tasks
Collaborator
|
superseded by #42420 |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
Since #40135, the HiCache host-pool budget is bounded by the cgroup headroom
limit - memory.current.memory.currentincludes page cache charged to the cgroup: the checkpoint the server just read (size of the model) and, in containerized Slurm jobs, the unpacked container image (58 GiB in a fresh job). That cache is freely reclaimable, and hostMemAvailable(psutil'savailable) already counts it as available, but the cgroup bound counts it as used. Right after a large model loads, the check under-reports allocatable memory by roughly the checkpoint size and rejects pools the kernel would provide by reclaiming the cache. For example, GLM-5.2-NVFP4 attp1 x pp4on one GB300 node with--hicache-size 135fails on every rank:The same recipe passes or fails depending on the node, because the charged cache differs per job.
Modifications
_cgroup_memory_headroomnow counts usage at each limited cgroup level asmemory.current - (active_file + inactive_file), clamped at 0. On cgroup v1 it ismemory.usage_in_bytes - (total_active_file + total_inactive_file). These counters cover the pages on the kernel's file-backed reclaim lists, which is whatMemAvailablecounts as reclaimable, so the host and cgroup bounds agree. Shared memory and tmpfs sit on the anonymous lists and stay counted as used.memory.stat'sfileandcachecounters would wrongly count them as free.total_*counters rather than the local ones), a shared-memory decoy, and the clamp. It fails on the old code.GPU-free validation in a Slurm job (cgroup v2, 350 GiB limit, no swap), comparing the old and new code side by side, in GiB:
available/dev/shmmlockstood in for pinned host memory: locking 315.4 GiB then succeeded, although the old bound allowed only 173.1 GiB. The kernel reclaimed the cache (175.3 to 22.8 GiB) and there was no out-of-memory kill. While the memory was held, the new headroom read 33.3 GiB against a predicted 32.0 GiB.Accuracy Tests
N/A: model outputs are unaffected.
Speed Tests and Profiling
N/A: the only cost is one extra
memory.statread per limited cgroup level at host-pool init.Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ci🤖 Generated with Claude Code
CI States
Latest PR Test (Base): ✅ Run #36844629362⚠️ Not enabled -- add
Latest PR Test (Extra):
run-ci-extralabel to opt in.Latest PR Test (AMD ROCm 10): ❌ Run #36844629350