Repository navigation
[HiCache][perf]: pipeline the HiCache ack-count all_reduce across scheduler steps - #39903
Closed
alphabetc1 wants to merge 2 commits into
Closed
alphabetc1 wants to merge 2 commits into
alphabetc1 wants to merge 2 commits into
Conversation
alphabetc1
requested review from
Ying1123,
hanming-lu,
hnyls2002,
huangtingwei9988,
hzh0425,
ispobock,
merrymercy,
xiezhq-hermann and
yizhang2077
as code owners
September 17, 2026 04:56
alphabetc1
force-pushed
the
fix-hicache-async-ack-sync
branch
from
September 18, 2026 01:36
2d12536 to
f19d5c9
Compare
4 of 5 tasks
4 of 5 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
With
--enable-hierarchical-cache,UnifiedRadixCache.check_hicache_eventsruns once per scheduler step and MIN-reduces the local write/load ack counts (plus storage queue sizes) across TP ranks with a blocking glooall_reduce. It sits inget_next_batch_to_run, right before the next forward is launched, so its full latency lands on the scheduler critical path every step.The cost is the host's gloo latency, which varies a lot between machines. Measured with Inkling-Small-NVFP4 (mxfp8 KV, TP=4, 128 concurrent lanes, 16 turns):
A with_stack trace on the B200 host attributes 1121 ms of 400 steps (2.8 ms/step) to
_all_reduce_attn_groups -> all_reduce -> wait. The HiCache backup stream itself is <1% busy on the GPU, and write-back policy, the Rust tree core, explicit NUMA binding and a smaller host pool do not move the number. Only the per-step collective does.Modifications
Pipeline the collective instead of blocking on it:
check_hicache_events, issue the MINall_reduceof the local counts withasync_op=Trueand keep theWork.wait()on it (normally long complete), pop the acks and storage entries it agreed on, and only then issue the next reduce. Counting after popping matters: counting before would include the acks just popped and over-report on the next step (covered bytest_pops_previous_counts_then_reduces_the_remainder).pp_size == 1. CP+TP (two groups) and PP pipelines keep the blocking_sync_hicache_ready_counts.SGLANG_ENABLE_HICACHE_ASYNC_ACK_SYNC=0restores the blocking path everywhere.Behavior change: acks are processed one step later than before (host-hit prefill starts one step later; a backed-up node stays locked one step longer). The per-step call count is unchanged and symmetric across ranks.
Accuracy Tests
Host-hit correctness: shape-matched L1-hit vs host-restore probe (same request served once as a 4096-token device hit and once as a 4096-token host restore, extend 128), bf16 KV: identical text, max per-token logprob diff 0.0. The pipelined path only changes when acks are popped, not what is written or read.
test/registered/unit/mem_cache/test_hiradix_pp_sync_drain.py: two new cases for the pipelined path (pop-then-reduce order,wait()before use); the PP fixture gains the two new fields.test_unified_radix_cache_unittest.py(python and rust tree core),test_hicache_staged_write_back_dispatch.py,test_unified_cache_linker.py, and thetest/registered/unit/mem_cachedirectory pass on H200.Speed Tests and Profiling
Inkling-Small-NVFP4, mxfp8 KV, TP=4, 128 concurrent lanes x 16 turns (about 55 s timed window), same box and session, one server launch per arm:
Checklist
CI States
Latest PR Test (Base): 🚫 Run #35299352621
Latest PR Test (Extra): 🚫 Run #35299352387
Latest PR Test (AMD ROCm 10): ❌ Run #35299352539