Skip to content

[HiCache] Add fast_file local storage backend (direct host-buffer I/O, parallel reads, background eviction) - #39880

Open
xiezhq-hermann wants to merge 1 commit into
mainfrom
hicache-fast-file-backend
Open

xiezhq-hermann wants to merge 1 commit into
mainfrom
hicache-fast-file-backend

Conversation

@xiezhq-hermann

@xiezhq-hermann xiezhq-hermann commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

HiCacheFile is a deliberately minimal reference L3 backend: every page goes through a staging tensor, reads are serial, writes are not atomic, and there is no way to bound how much disk it uses. That makes it fine for demos but not for a real local NVMe tier. Networked backends (mooncake, hf3fs, nixl, ...) fill that gap only when their infrastructure is available.

This PR adds fast_file, a portable local-filesystem backend that keeps the reference backend's on-disk format (one raw <key><suffix>.bin page per file, same key suffix, so the two can read each other's pages) and adds what a local tier needs to keep up with serving:

  • vectored readv/writev straight between page-first host pool buffers and the page file (no staging copy), with a staged fallback for other layouts and for logical pools;
  • a bounded worker pool for parallel page reads and existence checks;
  • atomic publication (temporary file + rename), so a reader never sees a partial page;
  • background LRU eviction between two watermarks (max_size), an optional filesystem free-space floor (min_free_space), an optional positive metadata cache, and startup cleanup of abandoned temporary files;
  • an exact per-model / per-parallel-layout / per-host-layout namespace directory under the storage root, so the startup scan, eviction and clear() never touch another deployment's files.

Usage

--hicache-storage-backend fast_file \
--hicache-storage-backend-extra-config '{"storage_dir":"/mnt/nvme/hicache","max_size":"3Ti","min_free_space":"4Ti","read_workers":8}'

The storage root comes from the storage_dir extra-config key, then the existing SGLANG_HICACHE_FILE_BACKEND_STORAGE_DIR, then /tmp/hicache. No new environment variable is introduced; every other knob is an extra-config key documented in the module docstring of fast_file_store.py.

Engine touch points

The engine changes are intentionally tiny: the backend factory registers fast_file (constructed with the host pool so the namespace can include its layout), the cache controller dispatches it through the batch_get_v1/batch_set_v1 path, the server argument gains the choice, and HiCacheFile gets a shared _build_config_suffix() helper so both backends keep one file format. HiCacheFile itself and its evictor are otherwise untouched.

Side pools (SWA, Mamba, draft, ...) go through the v2 interface with direct I/O where the pool exposes page buffers. Rank-sharded side pools under a rank-replicated (MLA) KV namespace at TP > 1 are rejected at registration because the shared namespace cannot key them per rank; bounded eviction with side pools works but evicts KV pages and sidecar files independently (a warning is logged once).

Also included: the OSS buffer-mode bench no longer hardcodes a downstream storage backend, and its multiturn client now connects to 127.0.0.1 explicitly (its default localhost resolves to ::1 on IPv6-preferring hosts while the server binds 127.0.0.1).

Test plan

  • test/registered/unit/mem_cache/test_hicache_fast_file_unit.py (new, CPU): 53 tests covering direct and staged transfers, format compatibility with HiCacheFile, namespaces, cap and watermark eviction, free-space floor, metadata cache, corrupt/missing page handling, concurrent same-key writes, metrics, and side-pool rules. Passed together with the existing test_hicache_file_lru_unit.py (87 tests) against this tree.
  • test/registered/hicache/test_hicache_storage_file_backend.py: new TestHiCacheStorageFastFilePageFirstDirectIO class (page_first_direct, direct io backend, TP=2) including a byte-identical-generation-after-reload check; passed locally with a public 1.5B model.
  • L3 KL consistency (the storage-hit KL client in sglang.test.kl_test_utils flow, buffer_only + write_through): prefill/decode KL after an L3 reload is identical to the file backend to six significant digits on a dense GQA model, including prompts up to 16k tokens and a bounded-capacity flood, and also identical to file on an internal SWA-hybrid model at TP=4 with direct I/O on both the KV and SWA pools. GSM8K cold vs L3-warm scores are identical.
  • benchmark/hicache/bench_buffer_mode.py on a public 1.5B model with local NVMe, file vs fast_file, buffer_only mode: identical hit rates; long-prefix replay from L3 1.6x to 6x faster with 3x to 15x lower tail latency; at 256 concurrent conversations the file backend could not drain its write backlog while fast_file completed. A CPU microbenchmark on a real page-first MHA host pool (8.4 MB pages, page cache dropped) reads at 3.7 GB/s with one worker and 20 GB/s with eight versus 0.34 GB/s through the staged path, with writes 2.6 vs 3.1 to 3.3 GB/s.
  • Not run here: AMD/NPU CI lanes.

Original commits

  • 7127bf772e
  • 78578ca925
  • d63c8b8f37
  • 085579218a

CI States

Latest PR Test (Base): ❌ Run #35174115270
Latest PR Test (Extra): ❌ Run #35174115041
Latest PR Test (AMD ROCm 10): ❌ Run #35174115268

A portable local-filesystem L3 backend that keeps HiCacheFile's on-disk
page format and adds direct page-first host-buffer I/O (readv/writev, no
staging copy), a bounded worker pool for parallel reads and existence
checks, atomic page publication, background LRU eviction between two
watermarks with an optional free-space floor, an optional positive
metadata cache, and an exact per-model/per-layout namespace directory.

Engine touch points are small: factory registration (constructed with the
host pool), the controller's batch_get_v1/batch_set_v1 dispatch list, the
server-argument choice, and a shared HiCacheFile._build_config_suffix()
helper so both file backends keep one format. The storage root comes from
the storage_dir extra-config key, then SGLANG_HICACHE_FILE_BACKEND_STORAGE_DIR,
then /tmp/hicache; no new environment variable.

Also drops the downstream-only storage backend wiring from
benchmark/hicache/bench_buffer_mode.py and makes its multiturn client
connect to 127.0.0.1 explicitly (localhost resolves to ::1 on
IPv6-preferring hosts while the server binds 127.0.0.1).
@mintlify

mintlify Bot commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Preview deployment for your docs. Learn more about Mintlify Previews.

Project Status Preview Updated
lmsysorg 🟢 Ready View Preview Sep 17, 2026, 2:24 AM

💡 Tip: Enable Automations to automatically generate PRs for you.

@github-actions github-actions Bot added documentation Improvements or additions to documentation hicache Hierarchical Caching for SGLang labels Sep 17, 2026
@xiezhq-hermann

Copy link
Copy Markdown
Collaborator Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 17, 2026
detain added a commit to detain/sglang that referenced this pull request Sep 28, 2026
Brings in 113 upstream commits (7ad55e4..d34f7b2). Upstream now
contains three stack PRs as merged: sgl-project#40501, sgl-project#40175, sgl-project#33778; every line
their squashes add is already present in the stack.

Conflicts resolved:
- arg_groups/fields/memory.py, managers/cache_controller.py: union of the
  stack's fast_file backend (sgl-project#39880) and upstream's tensorcast backend.
- layers/attention/qsa/mqa.py: keep the stack's SM120 Triton imports and
  fp8 _scoring_dtype path; adopt upstream's ROCm 16-wide head alignment
  (sgl-project#38875) and its is_hip import.
- models/qwen4_exp.py: all six hunks are stack-only additions (PLE host
  staging, CUDA-graph prewarm, replay prepare) against upstream's final
  sgl-project#40501; kept ours. The merged file equals the pre-merge stack version.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
staging - docs — a71c2814 Deployed Sep 17, 2026 by mintlify[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation hicache Hierarchical Caching for SGLang run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant