Skip to content

[HiCache][DSV4] Complete SWA pressure and host-pool fixes - #659

Merged
HanHan009527 merged 7 commits into
bytedance/deepseek_v4from
codex/dsv4-hicache-swa-performance
Jul 23, 2026
Merged

HanHan009527 merged 7 commits into
bytedance/deepseek_v4from
codex/dsv4-hicache-swa-performance

Conversation

@luoroger37

@luoroger37 luoroger37 commented Jul 23, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

This draft carries the complete P0 merge train for the current DSV4 deployment matrix:

  1. [sgl] Window-aware LRU refresh for SWA prefix cache in unified cache sgl-project/sglang#26615 (6965fe0eec) — window-aware SWA LRU refresh.
  2. Add Inkling model support sgl-project/sglang#31681 (02236fa38c) — reusable scheduler semantics for shrinking chunked prefill under SWA pressure.
  3. [HiCache] Optimize HiCache host pool free-list release sgl-project/sglang#30658 (44e3dd2713) — lazy host-pool free-list release, adapted to this fork's legacy/custom host pools.
  4. [sgl] proactively release out-of-window SWA slots after chunked prefill sgl-project/sglang#27402 + Fix SWA eviction tombstoning the last leaf sgl-project/sglang#29860 (3c1f9eafa5, a375e9f3da) — eagerly release out-of-window SWA slots using the final page-aware/tombstone-safe upstream boundary.

The full sgl-project#31681 commit also contains unrelated model code, so this PR carries its complete reusable scheduler hunk rather than importing unrelated files.

P1 items such as sgl-project#29310 and the sgl-project#27091/sgl-project#29460 metadata migration are intentionally excluded. The unmerged Decode retraction snapshot candidate is also excluded from the final diff.

Motivation

The existing DSV4 HiCache correctness backports made L2 recovery safer, but several P0 performance gaps remained:

  • Under high QPS/MC, SWA ancestry refresh, all-or-nothing prefill admission, repeated host free-list concatenation, and retaining already out-of-window SWA slots add scheduler/allocator pressure and degrade TTFT.
  • The shared upstream SWA eviction helper must preserve this fork's [HiCache][DSV4] Release prompt-prefix SWA under EIC + fix paged host accounting #655 EIC ownership rule: ordinary radix caches own protected-prefix SWA, while EIC has no SWA radix-tree state and may release out-of-window prefix SWA.

Changes

Prefill SWA/HiCache pressure path

  • Refresh SWA LRU only through the active sliding-window ancestry.
  • Shrink chunked prefill to a safe page-aligned size when current SWA capacity cannot admit the full chunk.
  • Queue released host slots and merge them lazily only when allocation needs them.
  • Share upstream free_swa_out_of_window_slots() between scheduler decode eviction and UnifiedRadixCache unfinished-request insertion.
  • Use the final Fix SWA eviction tombstoning the last leaf sgl-project/sglang#29860 boundary:
    • chunk cache retains one sliding window;
    • radix cache retains max(sliding_window_size, page_size) so the last leaf cannot become all-tombstone;
    • swa_evict_floor remains a hard lower bound.
  • Preserve cache-specific prefix ownership:
    • Unified/Radix uses the upstream default and keeps cache_protected_len as the scheduler eviction floor;
    • EIC explicitly selects release_cache_protected_prefix through its existing swa_evict_release_prefix=True policy because it has no SWA radix-tree state and reloads SWA from host storage.
  • Keep the UnifiedRadix unfinished-request optimization opt-in exactly like upstream with:
SGLANG_OPT_UNIFIED_CACHE_FREE_OUT_OF_WINDOW_SLOTS=1

Correctness invariants

  • SWA chunk shrinking never consumes the reserved sliding window, host-hit window, allocator page slack, or protected prefix.
  • Ordinary radix eviction never advances to the page-aligned insert frontier and never frees tree-owned protected-prefix SWA.
  • EIC may release only page-aligned, out-of-window request SWA; full KV remains separately owned and SWA is restored from its host backup on reuse.
  • Released host slots count as available but are not reused until allocation merges them under the allocator lock.
  • No attention kernel, KV byte format, C128 state addressing, Decode retraction, or HiCache L2 completion/lock ordering is changed.

Validation

Passed locally:

Not run locally because the isolated host Python has neither torch nor pytest:

  • test_prefill_adder.py
  • test_memory_pool_host_release.py
  • test_swa_eviction_boundary.py
  • test_unified_radix_cache_unittest.py

Before marking ready, run repository CI and production-shaped validation:

Prefill

  • DSV4 Flash FP8, PP4 x TP2 x CP2
  • UnifiedRadixTree + HiCache L2, write_through, direct + page_first_direct
  • chunked prefill 8192, swa_full_tokens_ratio=0.1
  • SGLANG_OPT_UNIFIED_CACHE_FREE_OUT_OF_WINDOW_SLOTS=1
  • high-QPS/MC TTFT, throughput, scheduler CPU time, host allocator CPU time
  • cold vs forced-L2-hit first-token logits/KL

EIC compatibility

  • EIC + DSV4 SWA under a non-zero cache_protected_len
  • verify out-of-window prompt-prefix SWA is released without changing full-KV ownership
  • verify prefix reuse restores SWA from EIC host backup
  • confirm SWA pool occupancy does not grow monotonically under multi-turn churn

CI States

Latest PR Test: Not run yet
Latest PR Test (Extra): ⚠️ Not enabled — add run-ci-extra label to opt in.

Backport upstream sgl-project#26615 so SWA prefix matches and inserts refresh only the active sliding-window ancestry instead of keeping the full walked prefix hot.

Upstream-Commit: 6965fe0
Adapt the scheduler-side SWA pressure handling merged with upstream sgl-project#31681. Charge host hits separately, preserve the sliding-window reservation, and shrink chunked prefill to a page-aligned SWA capacity instead of rejecting the request indefinitely.

Upstream-Commit: 02236fa
Adapt upstream sgl-project#30658 to the legacy host-pool layout used by deepseek_v4. Defer free-list concatenation until allocation needs released slots for the base, Mamba, V4 logical-anchor, and V4 paged pools.

Upstream-Commit: 44e3dd2
@gemini-code-assist

Copy link
Copy Markdown

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@luoroger37 luoroger37 changed the title [HiCache][DSV4] Backport SWA pressure and host-pool performance fixes [HiCache][DSV4] Complete P0 SWA pressure and decode retraction fixes Jul 23, 2026
@luoroger37 luoroger37 changed the title [HiCache][DSV4] Complete P0 SWA pressure and decode retraction fixes [HiCache][DSV4] Complete SWA pressure and decode retraction fixes Jul 23, 2026
@luoroger37 luoroger37 changed the title [HiCache][DSV4] Complete SWA pressure and decode retraction fixes [HiCache][DSV4] Complete P0 SWA pressure and host-pool fixes Jul 23, 2026
@luoroger37 luoroger37 changed the title [HiCache][DSV4] Complete P0 SWA pressure and host-pool fixes [HiCache][DSV4] Complete SWA pressure and host-pool fixes Jul 23, 2026
@luoroger37
luoroger37 marked this pull request as ready for review July 23, 2026 06:33
@gemini-code-assist

Copy link
Copy Markdown

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

@HanHan009527
HanHan009527 merged commit 08290e3 into bytedance/deepseek_v4 Jul 23, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants