Repository navigation
Conversation
mks-hash
force-pushed
the
feat/hicache-toolgap-prefetch
branch
from
October 3, 2026 23:08
ea426ac to
d1140b5
Compare
mks-hash
marked this pull request as ready for review
October 3, 2026 23:39
mks-hash
requested review from
CatherineSue,
JustinTong0323,
Ying1123,
alphabetc1,
hanming-lu,
hnyls2002,
huangtingwei9988,
hzh0425,
ispobock,
merrymercy,
slin1237,
xiezhq-hermann and
yizhang2077
as code owners
October 3, 2026 23:39
mks-hash
force-pushed
the
feat/hicache-toolgap-prefetch
branch
from
October 4, 2026 18:25
d1140b5 to
441815a
Compare
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem and behavior
A continuation normally starts its file L3 restore only when
/generatearrives.This adds exposed restore latency even if an orchestrator already knows the exact
prefix during a tool call. This experimental endpoint restores that prefix into
resident host L2 before request arrival; ordinary prefix matching, H2D and
generation consume it later. Matching early continuations join the existing
restore without duplicate backend reads.
POST /hicache/prefetchsupports submit/status/cancel, one active operation and32 recent outcomes. Existing CacheRequestHandle, storage workers, terminal ACK
and insert_host machinery own cleanup. TTL cancels pending work; successful KV
remains ordinary evictable L2, with no pin or lease. Backend replacement rejects
new submit; fixed model/tokenizer/storage namespace required.
Prerequisite / review order
Depends on #42149, which remains a separate bugfix PR. Both branches now use
upstream main
7a719e9a65e7f52b012ab37211b5f174e5d3d38d. The first commit is exactlythe refreshed prerequisite
c1098230d7ec1da4881324e48c7d7d5c29728268; the two followingfeature/benchmark commits have unchanged patches (
git range-diffequality).GitHub's main comparison includes the prerequisite until it merges. Review its
lifecycle changes in #42149 and the two subsequent feature commits separately.
The standalone ToolGap versions and their measured runtime SHA remain pinned;
this upstream rebase does not rewrite release patches or performance evidence.
Feature-only runtime: 341 additions in five files (335 measured, plus six lines
rejecting backend changes), with regression tests and a benchmark harness.
No scheduler/cache ownership redesign, Req fields, kernels, TreeCore algorithm,
storage format or distributed protocol changes.
Independent experimental release with patches, chart and raw evidence:
https://github.com/mks-hash/toolgap
Validation
80bb3fb6511ac421ed3b4309067681392203f3bd+ prerequisite + feature: patch checks and 27 CPU/controller/API tests passed.4ab720e6: 27 CPU/controller/API tests passed; its subsequent GPU smoke is recorded below. Current-base validation is separate.Median continuation TTFT, ms (n=3 per cell):
Qwen2.5-1.5B-Instruct, 4096 inputs /4080 restored, BF16, temperature0, seed42,
32 outputs, identical generation flags. Local file copying may warm OS cache.
Zero-gap benefit is not convincing; short-gap variance is visible. This is a
scoped capability demonstration, not a universal latency or significance claim.
Raw evidence preserves the resumed matrix:22+23 valid trials with identical five
runtime file hashes; prototype/failed harness trials are excluded.
Scope
Single worker/trajectory, FULL resident cache, TP1/PP1/DP1, Python TreeCore,
fixed text-generation model and file storage. No SWA, distributed/multi-node,
LoRA/speculation/multimodal, proactive H2D, timeout reclamation, generic retry
protocol or permanent blocked-I/O recovery. No production-readiness claim.
CI States
Latest PR Test (Base): ❌ Run #37224417551
Latest PR Test (Extra): ❌ Run #37224417475
Latest PR Test (AMD ROCm 10): ❌ Run #37224417564
Earlier interface / overlap audit (4ab720e, 2026-10-04)
The audited base was
4ab720e6557b44d07bde471ed52a178795298f1a. No equivalent pre-arrival resident file-KV submit/status/cancel API found in inspected main; existing external-corpus control is ngram speculation, not KV restore. This is not a claim about all open PRs/external projects. Of the five feature runtime paths, only scheduler.py changed since pinned base. The prefetch/query/read/terminal-result/abort/match/load_back AST bodies and file backend are unchanged.Targeted regression: 162 passed, 2 CUDA-only skips, 2601 subtests passed across11 files. Combined run had one Gloo loopback setup failure; the exact fixture passed standalone with permitted loopback access. No unresolved code-test failure. Distributed exception recovery remains out of scope.
Related main changes include generation checkpoint unification, host sizing/rank reads and opt-in unified-memory token-major/strided transfers. Benchmark did not enable unified-memory. The separately approved latest-main GPU smoke below now validates the checkpoint/generation path. Pinned ToolGap v0.1 and its recorded45-trial benchmark remain unchanged.
GPU smoke on previous upstream base — PASS
Main
4ab720e6557b44d07bde471ed52a178795298f1a, featured1140b5ec5c320a440fc35b0704f78dbdff31449. Full Python/test source overlay from this commit, same pinned Qwen model and CUDA image. L4 / driver580.178.04 / CUDA13.0 / PyTorch2.13.0+cu130.Smoke report, raw evidence. ToolGap v0.1 tag/base and its45-trial data remain unchanged.
Current-main rebase and validation (7a719e9, 2026-10-04)
Base
7a719e9a65e7f52b012ab37211b5f174e5d3d38d; feature head441815a5(prerequisitec1098230, feature0d3b37e8, benchmark441815a5).The prerequisite was adapted to the new unified physical-transfer exception
handling (#39479) and prefetch retirement helper (#41453). It preserves the new
helper and unsupported-consumer fallback. No new API/scheduler/ownership design
was needed. No
/hicache/prefetchorproactive_prefetchequivalent was found inthis fetched main's SGLang runtime; that bounded audit does not cover every open PR.
lifecycle cases. Combined with ToolGap client/admission/CLI tests:
102 passed, 22 subtests passed, zero failures/errors/skips.
added physical-transfer/dispatch/host-assembler checks: 64 passed / 2 CUDA skips.
remains out of scope. Applicable changed-file hooks and whitespace checks passed.
PR. Production source and transport were not replaced by mocks in cache fixtures.
29-test GPU smoke belongs to
4ab720e6/d1140b5e; performance numbers belongto the pinned 45-trial implementation. Fresh physical H2D changes remain without
new local GPU validation. No new performance claim or paid run accompanies this rebase.