Skip to content

[AMD] Enable staged HiCache write-back for DeepSeek V4 - #35043

Open
AMD-yanfeiwang wants to merge 3 commits into
sgl-project:mainfrom
AMD-yanfeiwang:amd/rocm-dsv4-hicache-staged-writeback
Open

AMD-yanfeiwang wants to merge 3 commits into
sgl-project:mainfrom
AMD-yanfeiwang:amd/rocm-dsv4-hicache-staged-writeback

Conversation

@AMD-yanfeiwang

@AMD-yanfeiwang AMD-yanfeiwang commented Aug 16, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

DeepSeek-V4 hybrid HiCache uses specialized paged, compress-state, and DSA indexer host pools. Their staged page-first write-back remained gated to native CUDA even though the shared JIT kernel supports HIP. Simply opening that gate is not sufficient on ROCm: on MI355X with ROCm 7.2.26015, mmap-backed memory registered with HIP has a device alias different from its CPU data_ptr(), so the existing PF-to-LF GPU load kernel faults while dereferencing the CPU virtual address.

Modifications

  • Enable staged write-back on ROCm for DeepSeekV4PagedHostPool, DeepSeekV4StateHostPool, and DSAIndexerPoolHost.
  • Use HIP runtime H2D copies for page-first loads from registered host storage, allowing the runtime to resolve the GPU-visible host alias. CUDA retains the existing AOT kernel path.
  • Cache converted CPU page-row indices across the layers of one L2 load, with tensor identity/version invalidation.
  • Validate ROCm HiCache host-pool capabilities during controller construction and dynamic sidecar registration. kernel + page_first now fails at startup if any pool lacks staged JIT write-back, and kernel + layer_first is rejected because its GPU kernel dereferences host CPU virtual addresses. The existing runtime checks remain as defense in depth.
  • Add byte-exact D2H/H2D round-trip coverage for C4/C128-style paged rows, compress-state ring rows, and DSA indexer rows, including in-place index mutation.

Scope and index routing

This PR supports the DeepSeek-V4 kernel + page_first host pools listed above; it is not an allocator-wide fix for every kernel that can dereference registered mmap storage. Generic MLA/MHA page-first loads and ROCm kernel layer-first transfers require separate safe paths; the latter is rejected at startup by this PR.

HybridCacheController already routes write-back indices per pool through supports_per_pool_backup_indices: pools with staged JIT keep CPU host indices as required by staged_write_back.cuh, while non-staged pools move their indices to the GPU. This PR marks the three DeepSeek-V4 pools as staged-capable and completes that existing contract; it does not replace the controller with a new group-wide routing rule.

Accuracy Tests

MI355X, ROCm 7.2.26015:

  • Registered mmap pointer probe: CPU range 0x71c149b50000..0x71c149b58000; hipHostGetDevicePointer() returned 0x717fd0020000; hipPointerGetAttributes() reported the CPU host pointer and the distinct device pointer.
  • The parent PF-to-LF kernel faulted at 0x71c149b53000, inside that CPU range, confirming that the GPU kernel dereferenced the CPU VA rather than the registered device alias.
  • Staged GPU suite: 15 passed.
  • HostPoolGroup dispatch suite: 19 passed, 1 skipped, including startup rejection and dynamic-sidecar validation cases.
  • DSV4 HiSparse suite: 16 passed, 2 skipped.
  • DSA host-pool suite: 1 passed, 2 skipped, 2 subtests passed.

Real DeepSeek-V4-Pro TP8 service smokes (kernel/page_first, 32K L1, HiCache ratio 2):

  • write_through operator-path smoke: every completed insertion eagerly backed up the prefix, deliberately forcing staged D2H without depending on eviction victim selection. Five distinct 8K prefixes filled/evicted L1; replaying the first returned cached_tokens_details={device: 0, host: 7936}. The initial and host-tier replay output token were both 65.
  • write_back policy smoke on PR head 1dea66d: after three cold fills, a seed-0 replay returned cached_tokens_details={device: 7936, host: 0}, proving the prefix remained device-only before eviction. Two additional cold fills exceeded L1; replaying seed 1 returned cached_tokens_details={device: 0, host: 7936}, proving eviction-triggered backup and load-back. The initial and host-tier replay output token were both 65.
  • Both services started successfully on 8x MI355X. Live JIT produced staged modules for 65,536-byte C4 rows, 8,448-byte indexer rows, and 2,048-byte C128 rows.
  • Final log scans found no memory-access fault, traceback, runtime error, OOM, write-back drop, or host-pressure backup failure.

Speed Tests and Profiling

Direct-A / staged candidate / Direct-B on one idle MI355X, C4-style rows, 8 layers, 1/4/16/64 pages, 40 repetitions per shape:

  • D2H synchronized completion latency: 73.37% lower geomean.
  • H2D synchronized completion latency: 8.98% lower geomean.
  • D2H host submission latency: 84.02% lower geomean.
  • H2D host submission latency: 23.23% lower geomean.

Additional 61-layer scale check on the same MI355X, C4-style rows, 1/4/16/64 pages, 10 repetitions per shape:

  • D2H synchronized completion latency: 90.24% lower geomean; every shape improved by 87.27%-91.42%.
  • D2H host submission latency: 97.57% lower geomean; every shape improved by 90.67%-99.02%.
  • H2D synchronized completion latency: 8.64% lower geomean; per-shape reductions were 24.01% / 3.35% / 2.68% / 2.51%.
  • H2D host submission latency: 14.94% lower geomean, but the 16/64-page shapes were 1.99% / 3.61% higher. The H2D submission benefit therefore does not hold uniformly at larger page counts.

Checklist

  • Format code with pre-commit.
  • Add and run focused unit/kernel tests.
  • Validate a real TP8 host-tier hit.
  • Provide paired performance results.
  • Run repository CI.

CI States

Latest PR Test (Base): ❌ Run #32138879425
Latest PR Test (Extra): ❌ Run #32138879088

@github-actions github-actions Bot added the hicache Hierarchical Caching for SGLang label Aug 16, 2026
@AMD-yanfeiwang
AMD-yanfeiwang marked this pull request as ready for review August 17, 2026 01:56

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

amd hicache Hierarchical Caching for SGLang jit-kernel sgl-kernel

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant