Skip to content

feat(hicache): restore exact prefixes before continuation arrival - #42434

Open
mks-hash wants to merge 3 commits into
sgl-project:mainfrom
mks-hash:feat/hicache-toolgap-prefetch
Open

mks-hash wants to merge 3 commits into
sgl-project:mainfrom
mks-hash:feat/hicache-toolgap-prefetch

Conversation

@mks-hash

@mks-hash mks-hash commented Oct 3, 2026 •

Copy link
Copy Markdown

Problem and behavior

A continuation normally starts its file L3 restore only when /generate arrives.
This adds exposed restore latency even if an orchestrator already knows the exact
prefix during a tool call. This experimental endpoint restores that prefix into
resident host L2 before request arrival; ordinary prefix matching, H2D and
generation consume it later. Matching early continuations join the existing
restore without duplicate backend reads.

POST /hicache/prefetch supports submit/status/cancel, one active operation and
32 recent outcomes. Existing CacheRequestHandle, storage workers, terminal ACK
and insert_host machinery own cleanup. TTL cancels pending work; successful KV
remains ordinary evictable L2, with no pin or lease. Backend replacement rejects
new submit; fixed model/tokenizer/storage namespace required.

Prerequisite / review order

Depends on #42149, which remains a separate bugfix PR. Both branches now use
upstream main 7a719e9a65e7f52b012ab37211b5f174e5d3d38d. The first commit is exactly
the refreshed prerequisite c1098230d7ec1da4881324e48c7d7d5c29728268; the two following
feature/benchmark commits have unchanged patches (git range-diff equality).
GitHub's main comparison includes the prerequisite until it merges. Review its
lifecycle changes in #42149 and the two subsequent feature commits separately.
The standalone ToolGap versions and their measured runtime SHA remain pinned;
this upstream rebase does not rewrite release patches or performance evidence.

Feature-only runtime: 341 additions in five files (335 measured, plus six lines
rejecting backend changes), with regression tests and a benchmark harness.
No scheduler/cache ownership redesign, Req fields, kernels, TreeCore algorithm,
storage format or distributed protocol changes.

Independent experimental release with patches, chart and raw evidence:
https://github.com/mks-hash/toolgap

Validation

  • Clean upstream SHA 80bb3fb6511ac421ed3b4309067681392203f3bd + prerequisite + feature: patch checks and 27 CPU/controller/API tests passed.
  • Measured implementation: NVIDIA L4, driver580.178.04, CUDA13.0, PyTorch2.13.0+cu130; 45 real-model trials, all identical32 output IDs.
  • Stage2A foundation: 41 passed / six SWA-only skips; real L3→L2→H2D, deterministic generation and finite-I/O recovery.
  • Every B consumed4080 storage tokens; every C consumed4080 host tokens; initial device hits0; duplicate backend pages0, relevant host evictions0.
  • At gaps≥500ms all nine C trials published before continuation submission; no proactive H2D. No-continuation experiment wasted111.56MiB ordinary L2, reclaimed by normal cache flush to baseline.
  • Historical rebase to 4ab720e6: 27 CPU/controller/API tests passed; its subsequent GPU smoke is recorded below. Current-base validation is separate.
  • Release backend guard and reference-release regressions were CPU-tested; the assembled release Dockerfile and GPU demo were not newly built/executed.
  • Applicable commit hooks passed; Rust hooks had no matching files, so no Rust validation is claimed.

Median continuation TTFT, ms (n=3 per cell):

Tool gap ms A recompute B request-time C proactive
0 572 502 490
100 569 514 438
500 577 528 166
1000 568 582 166
3000 571 508 163

Qwen2.5-1.5B-Instruct, 4096 inputs /4080 restored, BF16, temperature0, seed42,
32 outputs, identical generation flags. Local file copying may warm OS cache.
Zero-gap benefit is not convincing; short-gap variance is visible. This is a
scoped capability demonstration, not a universal latency or significance claim.
Raw evidence preserves the resumed matrix:22+23 valid trials with identical five
runtime file hashes; prototype/failed harness trials are excluded.

Scope

Single worker/trajectory, FULL resident cache, TP1/PP1/DP1, Python TreeCore,
fixed text-generation model and file storage. No SWA, distributed/multi-node,
LoRA/speculation/multimodal, proactive H2D, timeout reclamation, generic retry
protocol or permanent blocked-I/O recovery. No production-readiness claim.


CI States

Latest PR Test (Base): ❌ Run #37224417551
Latest PR Test (Extra): ❌ Run #37224417475
Latest PR Test (AMD ROCm 10): ❌ Run #37224417564

Earlier interface / overlap audit (4ab720e, 2026-10-04)

The audited base was 4ab720e6557b44d07bde471ed52a178795298f1a. No equivalent pre-arrival resident file-KV submit/status/cancel API found in inspected main; existing external-corpus control is ngram speculation, not KV restore. This is not a claim about all open PRs/external projects. Of the five feature runtime paths, only scheduler.py changed since pinned base. The prefetch/query/read/terminal-result/abort/match/load_back AST bodies and file backend are unchanged.

Targeted regression: 162 passed, 2 CUDA-only skips, 2601 subtests passed across11 files. Combined run had one Gloo loopback setup failure; the exact fixture passed standalone with permitted loopback access. No unresolved code-test failure. Distributed exception recovery remains out of scope.

Related main changes include generation checkpoint unification, host sizing/rank reads and opt-in unified-memory token-major/strided transfers. Benchmark did not enable unified-memory. The separately approved latest-main GPU smoke below now validates the checkpoint/generation path. Pinned ToolGap v0.1 and its recorded45-trial benchmark remain unchanged.

GPU smoke on previous upstream base — PASS

Main 4ab720e6557b44d07bde471ed52a178795298f1a, feature d1140b5ec5c320a440fc35b0704f78dbdff31449. Full Python/test source overlay from this commit, same pinned Qwen model and CUDA image. L4 / driver580.178.04 / CUDA13.0 / PyTorch2.13.0+cu130.

  • 29 passed, zero failed/skipped, including the two formerly skipped CUDA transfer tests.
  • One A/B/C trial each at500ms gap: all32 output IDs matched producer. B consumed4080 storage tokens; C consumed4080 host tokens; device hits0. Each B/C read255 unique pages /116981760 bytes, no duplicates.
  • C published L2 before continuation submission; normal H2D/generation succeeded. Terminal inflight0/cleanupfalse; wasted-prefix flush restored139520-token baseline.
  • A575.05ms, B537.18ms, C182.85ms are a single sanity run, not a replacement for the pinned45-trial performance evidence. No new statistical claim.
  • Exit0; VM and disk deleted. Entire session≤18.12min /~$0.27 estimate. No architecture/image change or second GPU execution.

Smoke report, raw evidence. ToolGap v0.1 tag/base and its45-trial data remain unchanged.

Current-main rebase and validation (7a719e9, 2026-10-04)

Base 7a719e9a65e7f52b012ab37211b5f174e5d3d38d; feature head
441815a5 (prerequisite c1098230, feature 0d3b37e8, benchmark 441815a5).
The prerequisite was adapted to the new unified physical-transfer exception
handling (#39479) and prefetch retirement helper (#41453). It preserves the new
helper and unsupported-consumer fallback. No new API/scheduler/ownership design
was needed. No /hicache/prefetch or proactive_prefetch equivalent was found in
this fetched main's SGLang runtime; that bounded audit does not cover every open PR.

  • Lifecycle/control/API pytest selection: 37 passed, including inherited
    lifecycle cases. Combined with ToolGap client/admission/CLI tests:
    102 passed, 22 subtests passed, zero failures/errors/skips.
  • Prerequisite existing CPU selection: 128 passed / 39 CUDA-fixture skips;
    added physical-transfer/dispatch/host-assembler checks: 64 passed / 2 CUDA skips.
  • Separate existing two-rank Gloo lifecycle passed; distributed exception recovery
    remains out of scope. Applicable changed-file hooks and whitespace checks passed.
  • Local tiny CPU host-pool admission-floor accommodations were external to the
    PR. Production source and transport were not replaced by mocks in cache fixtures.
  • No GPU smoke or benchmark was run on this refreshed base. The preceding
    29-test GPU smoke belongs to 4ab720e6 / d1140b5e; performance numbers belong
    to the pinned 45-trial implementation. Fresh physical H2D changes remain without
    new local GPU validation. No new performance claim or paid run accompanies this rebase.

@github-actions github-actions Bot added documentation Improvements or additions to documentation hicache Hierarchical Caching for SGLang unified-radix-cache labels Oct 3, 2026
@mks-hash
mks-hash force-pushed the feat/hicache-toolgap-prefetch branch from ea426ac to d1140b5 Compare October 3, 2026 23:08
@mks-hash
mks-hash marked this pull request as ready for review October 3, 2026 23:39
@mks-hash
mks-hash force-pushed the feat/hicache-toolgap-prefetch branch from d1140b5 to 441815a Compare October 4, 2026 18:25

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation hicache Hierarchical Caching for SGLang unified-radix-cache

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant