Conversation
0895d71 to
017d840
Compare
|
I independently reproduced the same registered-host pointer-domain issue on NVIDIA/WSL2 (RTX 5090, CUDA 13.0) while testing HiCache. Replaying KV data from registered mmap host memory through the JIT load kernel caused a CUDA illegal memory access error, even after cudaMemcpyBatchAsync had been disabled as a separate WSL workaround. The failure was resolved by passing the mapped device alias returned by cudaHostGetDevicePointer() to the JIT kernel instead of the registered host CPU address. One sequential and six concurrent host-cache replays then completed with matching outputs and zero request retractions. Two focused registered-mmap Mamba/KV transfer tests also passed. This was an independent validation of the same fix mechanism, not a test of the exact PR head. |
b878239 to
55e1262
Compare
DarkSharpness
left a comment
There was a problem hiding this comment.
LGMT on kernel side.
|
@AMD-yanfeiwang Do you think this PR could solve the segmentation fault when we enable: |
@yichiche Yes, Qwen3.5 is covered at the code-path level. Its GQA layers use the MHA host pool, while its linear-attention states use MambaPoolHost; this PR fixes registered-host pointer aliases in both paths. However, I'd still suggest performing end-to-end testing specifically for Qwen 3.5 to ensure thorough validation. |
|
@AMD-yanfeiwang Thanks for the kind explanation. I'm just curious, I wasn't able to reproduce this error locally. Are you aware if the permissions of a cluster or local node could also cause different behavior? |
@yichiche It is unlikely to be a cluster permission issue; permission problems would normally make hipHostRegister fail explicitly. |
|
@yichiche I did the verify with your config and AgentX test case, and not found that crash, detail is below: End-to-end Qwen3.5 validation requested above: PASS with main + #35233 alone Tested on MI355X / ROCm 7.2 using Configuration matched the failing InferenceX arm: Qwen3.5-397B-A17B-MXFP4, TP2, AITER unified attention, EAGLE MTP (steps=3, topk=1, draft tokens=4), FP8 KV, page size 16, and HiCache Results:
This confirms #35233 alone fixes the linked InferenceX crash; #37152 is not required for correctness. The InferenceX recipe should be rerun with an image containing #35233. |
|
|
…ool (sgl-project#37152) Squashed from sgl-project/sglang PR sgl-project#37152 (open, not yet merged). Three changes on top of sgl-project#35233 (already in this base as 0163f8f): - pick_group_bytes() picks the widest of 128/64/32/16 B that divides the element and splits across lanes into a 4/8/16 B package, so element sizes not divisible by 128 (MLA's 576 B fp8 row) can use the JIT transfer kernels. Narrow rounds sit behind #ifdef USE_ROCM; the #else path reduces to the prior 128 B rule, so CUDA behaviour is unchanged. - can_use_hicache_jit_kernel() screens on the same rule via _tiles_across_lanes() instead of element_size % 128. - MHATokenToKOnlyPoolHost.can_use_jit now admits HIP, not CUDA only. The ROCm block quota is NOT changed by this PR. DEFAULT_BLOCK_QUOTA is already 32 on HIP and 2 on CUDA at this base (kvcache/hicache.py:21); the PR only carries it as context. An earlier version of this message credited the PR with a 2 -> 16 change, which is wrong on both counts. Original commits: 3a82fcf [ROCm] Let MLA fp8 rows reach the HiCache read JIT ab96bbb [ROCm] Enable HiCache JIT transfer kernels for MHATokenToKOnlyPoolHost dddcda8 [ROCm] Keep CUDA on the 128 B round, and pin the lane count the Python screen mirrors 2efb2ab [ROCm] Test that the copy-round screen agrees with the kernel rule 339c0fa [ROCm] Clarify HiCache logical copy groups (two "Merge branch 'main' into rocm-hicache" commits dropped; applied as the PR's net diff against its fork point 923e4a5) Co-authored-by: Xiaobo Chen <xiaobo.chen@amd.com>
Motivation
Registered HiCache host allocations can have different CPU and GPU virtual addresses. This is observable on MI355X, where
hipDeviceAttributeCanUseHostPointerForRegisteredMemis false. Passing the CPU VA to a custom kernel causes a GPU memory fault even though runtime copy APIs can still accept that address.The affected paths include AOT/JIT HiCache transfers, Mamba transfers, HiSparse copies, and host-pool pointer tables. CPU-side storage and disaggregation metadata must continue to expose the original CPU VA.
Modifications
get_device_accessible_ptrAOT API usingcudaHostGetDevicePointer/hipHostGetDevicePointer.Accuracy Tests
ROCm 7.2 / MI355X (
gfx950), Docker imagelmsysorg/sglang-rocm:v0.5.17-rocm720-mi35x-20260813:test_hicache_page_first_write_back.py:15 passed.test_transfer_mamba.py:16 passed.test_hisparse.py:16 passed, 2 skipped.test_mem_pool_host.py:9 passed.3 passed.hipHostGetDevicePointer; wrong-device input is rejected; current device is restored.CUDA validation in
lmsysorg/sglang:v0.5.17-cu130(compile-only because this host has no NVIDIA GPU):transfer.cuandcommon_extension.cctranslation units compiled for SM90.Speed Tests and Profiling
Not run. This is a correctness fix for registered-host pointer address domains; it does not change kernel bodies or transfer algorithms.
Checklist
CI States
Latest PR Test (Base): ❌ Run #34735132165
Latest PR Test (Extra): ❌ Run #34735132093
Latest PR Test (AMD ROCm 10): ❌ Run #34735132166