Skip to content

[HiCache][perf]: pipeline the HiCache ack-count all_reduce across scheduler steps - #39903

Closed
alphabetc1 wants to merge 2 commits into
sgl-project:mainfrom
alphabetc1:fix-hicache-async-ack-sync
Closed

alphabetc1 wants to merge 2 commits into
sgl-project:mainfrom
alphabetc1:fix-hicache-async-ack-sync

Conversation

@alphabetc1

@alphabetc1 alphabetc1 commented Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

With --enable-hierarchical-cache, UnifiedRadixCache.check_hicache_events runs once per scheduler step and MIN-reduces the local write/load ack counts (plus storage queue sizes) across TP ranks with a blocking gloo all_reduce. It sits in get_next_batch_to_run, right before the next forward is launched, so its full latency lands on the scheduler critical path every step.

The cost is the host's gloo latency, which varies a lot between machines. Measured with Inkling-Small-NVFP4 (mxfp8 KV, TP=4, 128 concurrent lanes, 16 turns):

host 4-rank gloo all_reduce L2 (HiCache) vs L1
4x GB300 (Grace) 0.5 ms -2%
4x B200 (Xeon 6960P, 2 sockets) 3.2 ms -21%

A with_stack trace on the B200 host attributes 1121 ms of 400 steps (2.8 ms/step) to _all_reduce_attn_groups -> all_reduce -> wait. The HiCache backup stream itself is <1% busy on the GPU, and write-back policy, the Rust tree core, explicit NUMA binding and a smaller host pool do not move the number. Only the per-step collective does.

Modifications

Pipeline the collective instead of blocking on it:

  • At the end of check_hicache_events, issue the MIN all_reduce of the local counts with async_op=True and keep the Work.
  • On the next step, wait() on it (normally long complete), pop the acks and storage entries it agreed on, and only then issue the next reduce. Counting after popping matters: counting before would include the acks just popped and over-report on the next step (covered by test_pops_previous_counts_then_reduces_the_remainder).
  • Every count is a MIN over what each rank had ready one step earlier; acks and storage-queue entries only accumulate between two steps, so popping them a step later is safe. The reclaim digest is captured at issue time and checked against the same tensor on consume.
  • The pipelined path needs exactly one collective per step: TP (or a single CP) group with pp_size == 1. CP+TP (two groups) and PP pipelines keep the blocking _sync_hicache_ready_counts. SGLANG_ENABLE_HICACHE_ASYNC_ACK_SYNC=0 restores the blocking path everywhere.

Behavior change: acks are processed one step later than before (host-hit prefill starts one step later; a backed-up node stays locked one step longer). The per-step call count is unchanged and symmetric across ranks.

Accuracy Tests

Host-hit correctness: shape-matched L1-hit vs host-restore probe (same request served once as a 4096-token device hit and once as a 4096-token host restore, extend 128), bf16 KV: identical text, max per-token logprob diff 0.0. The pipelined path only changes when acks are popped, not what is written or read.

  • test/registered/unit/mem_cache/test_hiradix_pp_sync_drain.py: two new cases for the pipelined path (pop-then-reduce order, wait() before use); the PP fixture gains the two new fields.
  • test_unified_radix_cache_unittest.py (python and rust tree core), test_hicache_staged_write_back_dispatch.py, test_unified_cache_linker.py, and the test/registered/unit/mem_cache directory pass on H200.

Speed Tests and Profiling

Inkling-Small-NVFP4, mxfp8 KV, TP=4, 128 concurrent lanes x 16 turns (about 55 s timed window), same box and session, one server launch per arm:

arm tok/s vs L1
L1 11,546
L2, blocking reduce (before) 9,303 / 9,287 -20.8%
L2, pipelined reduce (this PR) 11,473 / 11,458 -0.6% / -0.8%

Checklist


CI States

Latest PR Test (Base): 🚫 Run #35299352621
Latest PR Test (Extra): 🚫 Run #35299352387
Latest PR Test (AMD ROCm 10): ❌ Run #35299352539

@alphabetc1
alphabetc1 force-pushed the fix-hicache-async-ack-sync branch from 2d12536 to f19d5c9 Compare September 18, 2026 01:36
@alphabetc1 alphabetc1 added run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) labels Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hicache Hierarchical Caching for SGLang run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) unified-radix-cache

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant