[HiCache] Add fast_file local storage backend (direct host-buffer I/O, parallel reads, background eviction) - #39880
Open
xiezhq-hermann wants to merge 1 commit into
Open
xiezhq-hermann wants to merge 1 commit into
xiezhq-hermann wants to merge 1 commit into
Conversation
A portable local-filesystem L3 backend that keeps HiCacheFile's on-disk page format and adds direct page-first host-buffer I/O (readv/writev, no staging copy), a bounded worker pool for parallel reads and existence checks, atomic page publication, background LRU eviction between two watermarks with an optional free-space floor, an optional positive metadata cache, and an exact per-model/per-layout namespace directory. Engine touch points are small: factory registration (constructed with the host pool), the controller's batch_get_v1/batch_set_v1 dispatch list, the server-argument choice, and a shared HiCacheFile._build_config_suffix() helper so both file backends keep one format. The storage root comes from the storage_dir extra-config key, then SGLANG_HICACHE_FILE_BACKEND_STORAGE_DIR, then /tmp/hicache; no new environment variable. Also drops the downstream-only storage backend wiring from benchmark/hicache/bench_buffer_mode.py and makes its multiturn client connect to 127.0.0.1 explicitly (localhost resolves to ::1 on IPv6-preferring hosts while the server binds 127.0.0.1).
xiezhq-hermann
requested review from
JustinTong0323,
Ying1123,
alphabetc1,
hanming-lu,
hnyls2002,
huangtingwei9988,
hzh0425,
ispobock,
merrymercy,
sogalin,
wisclmy0611,
yizhang2077 and
zijiexia
as code owners
September 17, 2026 02:21
Contributor
|
Preview deployment for your docs. Learn more about Mintlify Previews.
💡 Tip: Enable Automations to automatically generate PRs for you. |
Collaborator
Author
|
/tag-and-rerun-ci |
6 tasks
4 of 5 tasks
detain
added a commit
to detain/sglang
that referenced
this pull request
Sep 28, 2026
Brings in 113 upstream commits (7ad55e4..d34f7b2). Upstream now contains three stack PRs as merged: sgl-project#40501, sgl-project#40175, sgl-project#33778; every line their squashes add is already present in the stack. Conflicts resolved: - arg_groups/fields/memory.py, managers/cache_controller.py: union of the stack's fast_file backend (sgl-project#39880) and upstream's tensorcast backend. - layers/attention/qsa/mqa.py: keep the stack's SM120 Triton imports and fp8 _scoring_dtype path; adopt upstream's ROCm 16-wide head alignment (sgl-project#38875) and its is_hip import. - models/qwen4_exp.py: all six hunks are stack-only additions (PLE host staging, CUDA-graph prewarm, replay prepare) against upstream's final sgl-project#40501; kept ours. The merged file equals the pre-merge stack version. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
HiCacheFileis a deliberately minimal reference L3 backend: every page goes through a staging tensor, reads are serial, writes are not atomic, and there is no way to bound how much disk it uses. That makes it fine for demos but not for a real local NVMe tier. Networked backends (mooncake, hf3fs, nixl, ...) fill that gap only when their infrastructure is available.This PR adds
fast_file, a portable local-filesystem backend that keeps the reference backend's on-disk format (one raw<key><suffix>.binpage per file, same key suffix, so the two can read each other's pages) and adds what a local tier needs to keep up with serving:readv/writevstraight between page-first host pool buffers and the page file (no staging copy), with a staged fallback for other layouts and for logical pools;max_size), an optional filesystem free-space floor (min_free_space), an optional positive metadata cache, and startup cleanup of abandoned temporary files;clear()never touch another deployment's files.Usage
--hicache-storage-backend fast_file \ --hicache-storage-backend-extra-config '{"storage_dir":"/mnt/nvme/hicache","max_size":"3Ti","min_free_space":"4Ti","read_workers":8}'The storage root comes from the
storage_dirextra-config key, then the existingSGLANG_HICACHE_FILE_BACKEND_STORAGE_DIR, then/tmp/hicache. No new environment variable is introduced; every other knob is an extra-config key documented in the module docstring offast_file_store.py.Engine touch points
The engine changes are intentionally tiny: the backend factory registers
fast_file(constructed with the host pool so the namespace can include its layout), the cache controller dispatches it through thebatch_get_v1/batch_set_v1path, the server argument gains the choice, andHiCacheFilegets a shared_build_config_suffix()helper so both backends keep one file format.HiCacheFileitself and its evictor are otherwise untouched.Side pools (SWA, Mamba, draft, ...) go through the v2 interface with direct I/O where the pool exposes page buffers. Rank-sharded side pools under a rank-replicated (MLA) KV namespace at TP > 1 are rejected at registration because the shared namespace cannot key them per rank; bounded eviction with side pools works but evicts KV pages and sidecar files independently (a warning is logged once).
Also included: the OSS buffer-mode bench no longer hardcodes a downstream storage backend, and its multiturn client now connects to
127.0.0.1explicitly (its defaultlocalhostresolves to::1on IPv6-preferring hosts while the server binds127.0.0.1).Test plan
test/registered/unit/mem_cache/test_hicache_fast_file_unit.py(new, CPU): 53 tests covering direct and staged transfers, format compatibility withHiCacheFile, namespaces, cap and watermark eviction, free-space floor, metadata cache, corrupt/missing page handling, concurrent same-key writes, metrics, and side-pool rules. Passed together with the existingtest_hicache_file_lru_unit.py(87 tests) against this tree.test/registered/hicache/test_hicache_storage_file_backend.py: newTestHiCacheStorageFastFilePageFirstDirectIOclass (page_first_direct, direct io backend, TP=2) including a byte-identical-generation-after-reload check; passed locally with a public 1.5B model.sglang.test.kl_test_utilsflow, buffer_only + write_through): prefill/decode KL after an L3 reload is identical to thefilebackend to six significant digits on a dense GQA model, including prompts up to 16k tokens and a bounded-capacity flood, and also identical tofileon an internal SWA-hybrid model at TP=4 with direct I/O on both the KV and SWA pools. GSM8K cold vs L3-warm scores are identical.benchmark/hicache/bench_buffer_mode.pyon a public 1.5B model with local NVMe,filevsfast_file, buffer_only mode: identical hit rates; long-prefix replay from L3 1.6x to 6x faster with 3x to 15x lower tail latency; at 256 concurrent conversations thefilebackend could not drain its write backlog whilefast_filecompleted. A CPU microbenchmark on a real page-first MHA host pool (8.4 MB pages, page cache dropped) reads at 3.7 GB/s with one worker and 20 GB/s with eight versus 0.34 GB/s through the staged path, with writes 2.6 vs 3.1 to 3.3 GB/s.Original commits
7127bf772e78578ca925d63c8b8f37085579218aCI States
Latest PR Test (Base): ❌ Run #35174115270
Latest PR Test (Extra): ❌ Run #35174115041
Latest PR Test (AMD ROCm 10): ❌ Run #35174115268