Skip to content

[AMD] Fix registered HiCache host pointer aliases - #35233

Merged
HaiShaw merged 4 commits into
sgl-project:mainfrom
AMD-yanfeiwang:amd/rocm-hicache-host-pointer-alias
Sep 14, 2026
Merged

HaiShaw merged 4 commits into
sgl-project:mainfrom
AMD-yanfeiwang:amd/rocm-hicache-host-pointer-alias

Conversation

@AMD-yanfeiwang

@AMD-yanfeiwang AMD-yanfeiwang commented Aug 18, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

Registered HiCache host allocations can have different CPU and GPU virtual addresses. This is observable on MI355X, where hipDeviceAttributeCanUseHostPointerForRegisteredMem is false. Passing the CPU VA to a custom kernel causes a GPU memory fault even though runtime copy APIs can still accept that address.

The affected paths include AOT/JIT HiCache transfers, Mamba transfers, HiSparse copies, and host-pool pointer tables. CPU-side storage and disaggregation metadata must continue to expose the original CPU VA.

Modifications

  • Add a target-device-aware get_device_accessible_ptr AOT API using cudaHostGetDevicePointer / hipHostGetDevicePointer.
  • Resolve direct registered-host tensor operands at AOT and JIT kernel launch boundaries while preserving the caller's current device.
  • Build CUDA/ROCm host-pool kernel pointer tables from device-accessible aliases for MHA, MLA, K-only, DSV4 paged/state, and DSA pools.
  • Keep raw CPU addresses for unregister, contiguous-buffer metadata, page metadata, and storage/disaggregation consumers.
  • Preserve unpinned, NPU, and MUSA pointer behavior.
  • Add registered-mmap regressions for pointer domains, AOT/JIT transfers, and CUDA graph capture/replay.

Accuracy Tests

ROCm 7.2 / MI355X (gfx950), Docker image lmsysorg/sglang-rocm:v0.5.17-rocm720-mi35x-20260813:

  • Fresh ROCm AOT build: passed.
  • test_hicache_page_first_write_back.py: 15 passed.
  • test_transfer_mamba.py: 16 passed.
  • test_hisparse.py: 16 passed, 2 skipped.
  • test_mem_pool_host.py: 9 passed.
  • Registered-mmap post-rebase regressions: 3 passed.
  • Native HIP runtime comparison: alias matches hipHostGetDevicePointer; wrong-device input is rejected; current device is restored.
  • Repository pre-commit hooks over the rebased PR diff: passed.

CUDA validation in lmsysorg/sglang:v0.5.17-cu130 (compile-only because this host has no NVIDIA GPU):

  • HiCache, Mamba, and HiSparse JIT modules compiled and loaded for SM90.
  • Modified AOT transfer.cu and common_extension.cc translation units compiled for SM90.

Speed Tests and Profiling

Not run. This is a correctness fix for registered-host pointer address domains; it does not change kernel bodies or transfer algorithms.

Checklist


CI States

Latest PR Test (Base): ❌ Run #34735132165
Latest PR Test (Extra): ❌ Run #34735132093
Latest PR Test (AMD ROCm 10): ❌ Run #34735132166

Comment thread python/sglang/kernels/jit/include/sgl_kernel/utils.cuh Outdated
Comment thread python/sglang/kernels/jit/include/sgl_kernel/utils.cuh Outdated
@seokwoosong

Copy link
Copy Markdown
Contributor

I independently reproduced the same registered-host pointer-domain issue on NVIDIA/WSL2 (RTX 5090, CUDA 13.0) while testing HiCache.

Replaying KV data from registered mmap host memory through the JIT load kernel caused a CUDA illegal memory access error, even after cudaMemcpyBatchAsync had been disabled as a separate WSL workaround.

The failure was resolved by passing the mapped device alias returned by cudaHostGetDevicePointer() to the JIT kernel instead of the registered host CPU address. One sequential and six concurrent host-cache replays then completed with matching outputs and zero request retractions. Two focused registered-mmap Mamba/KV transfer tests also passed.

This was an independent validation of the same fix mechanism, not a test of the exact PR head.
The WSL batch-copy fallback was a separate workaround. These results provide NVIDIA runtime evidence consistent with the pointer-domain issue and fix mechanism described in this PR.

Comment thread python/sglang/kernels/jit/csrc/kvcacheio/hisparse.cuh Outdated
Comment thread python/sglang/kernels/jit/include/sgl_kernel/runtime.cuh Outdated

@DarkSharpness DarkSharpness left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGMT on kernel side.

@yichiche

Copy link
Copy Markdown
Collaborator

@AMD-yanfeiwang Do you think this PR could solve the segmentation fault when we enable:
--hicache-io-backend kernel
--hicache-mem-layout page_first
in InferenceX agent mode?
Ref: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34615847734/job/103318404674

'hipModuleUnload(module)' failed with 'hipErrorIllegalAddress'
Error code 700
Error code 700
Error code 700
[2026-09-11 23:29:06] No live scheduler processes found; skipping py-spy and CUDA coredump.

@AMD-yanfeiwang

AMD-yanfeiwang commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor Author

@AMD-yanfeiwang Do you think this PR could solve the segmentation fault when we enable: --hicache-io-backend kernel --hicache-mem-layout page_first in InferenceX agent mode? Ref: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34615847734/job/103318404674

'hipModuleUnload(module)' failed with 'hipErrorIllegalAddress'
Error code 700
Error code 700
Error code 700
[2026-09-11 23:29:06] No live scheduler processes found; skipping py-spy and CUDA coredump.

@yichiche Yes, Qwen3.5 is covered at the code-path level. Its GQA layers use the MHA host pool, while its linear-attention states use MambaPoolHost; this PR fixes registered-host pointer aliases in both paths.

However, I'd still suggest performing end-to-end testing specifically for Qwen 3.5 to ensure thorough validation.

@yichiche

yichiche commented Sep 14, 2026 •

Copy link
Copy Markdown
Collaborator

@AMD-yanfeiwang Thanks for the kind explanation. I'm just curious, I wasn't able to reproduce this error locally. Are you aware if the permissions of a cluster or local node could also cause different behavior?

@AMD-yanfeiwang

Copy link
Copy Markdown
Contributor Author

@AMD-yanfeiwang Thanks for the kind explanation. I'm just curious, I wasn't able to reproduce this error locally. Are you aware if the permissions of a cluster or local node could also cause different behavior?

@yichiche It is unlikely to be a cluster permission issue; permission problems would normally make hipHostRegister fail explicitly.

@jiejingzhangamd

jiejingzhangamd commented Sep 14, 2026 •

Copy link
Copy Markdown
Contributor

@yichiche I did the verify with your config and AgentX test case, and not found that crash, detail is below:

End-to-end Qwen3.5 validation requested above: PASS with main + #35233 alone

Tested on MI355X / ROCm 7.2 using lmsysorg/sglang-rocm:v0.5.19-rocm720-mi35x-20260911, with the engine source replaced by upstream main at 3eeb7d37f plus #35233 head c29e6054f; the ROCm AOT extension was rebuilt from that source.

Configuration matched the failing InferenceX arm: Qwen3.5-397B-A17B-MXFP4, TP2, AITER unified attention, EAGLE MTP (steps=3, topk=1, draft tokens=4), FP8 KV, page size 16, and HiCache write_through + kernel + page_first.

Results:

  • Both ranks allocated the KV and Mamba host pools and attached KV + MAMBA to UnifiedRadixCache.
  • The same 80-token startup warmup completed; /health reached 200.
  • A real chat generation completed (32 prompt + 42 completion tokens).
  • A c=28 burst completed 28/28 requests with HTTP 200.
  • HiCache counters confirmed the kernel path actually ran: KV backup 128 tokens, Mamba backup 2 slots, 98,738,176 bytes D2H, 0 dropped tokens.
  • No HIP code 700, illegal-memory-access, scheduler exception, or unload fault appeared.
  • Focused GPU regressions also passed: test_hicache_page_first_write_back.py + test_transfer_mamba.py: 31 passed.

This confirms #35233 alone fixes the linked InferenceX crash; #37152 is not required for correctness. The InferenceX recipe should be rerun with an image containing #35233.

@HaiShaw

HaiShaw commented Sep 14, 2026

Copy link
Copy Markdown
Collaborator

base-c-test-8-gpu-h20 (1) timed out previously, irrelevant to this PR. All other looks okay.

@HaiShaw
HaiShaw merged commit 0163f8f into sgl-project:main Sep 14, 2026
398 of 452 checks passed
xiaobochen-amd added a commit to xiaobochen-amd/sglang that referenced this pull request Sep 17, 2026
…ool (sgl-project#37152)

Squashed from sgl-project/sglang PR sgl-project#37152 (open, not yet merged).

Three changes on top of sgl-project#35233 (already in this base as 0163f8f):
  - pick_group_bytes() picks the widest of 128/64/32/16 B that divides the
    element and splits across lanes into a 4/8/16 B package, so element
    sizes not divisible by 128 (MLA's 576 B fp8 row) can use the JIT
    transfer kernels. Narrow rounds sit behind #ifdef USE_ROCM; the #else
    path reduces to the prior 128 B rule, so CUDA behaviour is unchanged.
  - can_use_hicache_jit_kernel() screens on the same rule via
    _tiles_across_lanes() instead of element_size % 128.
  - MHATokenToKOnlyPoolHost.can_use_jit now admits HIP, not CUDA only.

The ROCm block quota is NOT changed by this PR. DEFAULT_BLOCK_QUOTA is
already 32 on HIP and 2 on CUDA at this base (kvcache/hicache.py:21); the PR
only carries it as context. An earlier version of this message credited the
PR with a 2 -> 16 change, which is wrong on both counts.

Original commits:
  3a82fcf [ROCm] Let MLA fp8 rows reach the HiCache read JIT
  ab96bbb [ROCm] Enable HiCache JIT transfer kernels for MHATokenToKOnlyPoolHost
  dddcda8 [ROCm] Keep CUDA on the 128 B round, and pin the lane count the Python screen mirrors
  2efb2ab [ROCm] Test that the copy-round screen agrees with the kernel rule
  339c0fa [ROCm] Clarify HiCache logical copy groups
(two "Merge branch 'main' into rocm-hicache" commits dropped; applied as the
 PR's net diff against its fork point 923e4a5)

Co-authored-by: Xiaobo Chen <xiaobo.chen@amd.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

amd hicache Hierarchical Caching for SGLang jit-kernel memory-pool run-ci CI: run the baseline test suite on this PR sgl-kernel

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants