Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
21 commits
Select commit Hold shift + click to select a range
009e640
[AMD] Enable dsv4 paged HiSparse JIT tests on ROCm
amd-danli103 Jun 24, 2026
50b65ab
[AMD] Enable unified-KV HiSparse on ROCm
amd-danli103 Jun 24, 2026
4d29f16
[AMD] Add unified-KV HiSparse pool unit tests
amd-danli103 Jun 24, 2026
669e571
[AMD] Add unified-KV HiSparse eval tests
amd-danli103 Jun 24, 2026
7c9e37e
Shrink unified-KV device C4 to the host-offloaded budget so freed GPU…
amd-danli103 Jul 8, 2026
92d0caf
Merge branch 'main' into hisparse-dsv4pro-unified-kv-rocm-clean
amd-danli103 Aug 21, 2026
2b3177e
[AMD] Keep unified-KV HiSparse off the dsv4 paged copy path
amd-danli103 Aug 21, 2026
407e9d8
Merge branch 'main' into hisparse-dsv4pro-unified-kv-rocm-clean
amd-danli103 Aug 25, 2026
e8e4256
ci(amd): run dsv4 hisparse unified-KV eval suites in nightly
amd-danli103 Aug 26, 2026
502bc15
test(hisparse): align unified-KV eval env with dsv4 siblings
amd-danli103 Aug 26, 2026
1c00899
Merge branch 'main' into hisparse-dsv4pro-unified-kv-rocm-clean
amd-danli103 Sep 3, 2026
8013b9a
Merge branch 'main' into hisparse-dsv4pro-unified-kv-rocm-clean
amd-danli103 Sep 9, 2026
bd3c069
test(hisparse): put unified-KV tests on the new registered-test taxonomy
amd-danli103 Sep 9, 2026
f37eeab
Merge branch 'main' into hisparse-dsv4pro-unified-kv-rocm-clean
amd-danli103 Sep 9, 2026
007bddd
Merge branch 'main' into hisparse-dsv4pro-unified-kv-rocm-clean
amd-danli103 Sep 15, 2026
08cfec4
fix(hisparse): refuse unified-KV fp8 on the ROCm HiSparse path
amd-danli103 Sep 15, 2026
1b848af
fix(hisparse): keep unified-KV C4 pool after compressed init
amd-danli103 Sep 15, 2026
e5a4d38
Merge branch 'main' into hisparse-dsv4pro-unified-kv-rocm-clean
amd-danli103 Sep 21, 2026
4d8394f
fix(hisparse): tolerate missing unified_hisparse in compressed pool init
amd-danli103 Sep 21, 2026
7801c71
Merge branch 'main' into hisparse-dsv4pro-unified-kv-rocm-clean
amd-danli103 Sep 23, 2026
df3308a
test(hisparse): pass GSM8K host without URL scheme
amd-danli103 Sep 23, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
58 changes: 58 additions & 0 deletions .github/workflows/nightly-test-amd.yml
Original file line number Diff line number Diff line change
Expand Up @@ -79,6 +79,7 @@ on:
- nightly-8-gpu-mi35x-deepseek-v4-pro
- nightly-8-gpu-mi35x-deepseek-v4-pro-mtp
- nightly-8-gpu-mi35x-deepseek-v4-pro-dspark
- nightly-8-gpu-mi35x-deepseek-v4-pro-hisparse-unified
# 8-GPU Kimi-K2.6 (MI30x)
- nightly-8-gpu-kimi-k26
# 8-GPU Kimi-K3 (MI35x only - native MXFP4 needs gfx95x)
Expand Down Expand Up @@ -1603,6 +1604,62 @@ jobs:
python3 run_suite.py --hw amd --suite nightly-amd-8-gpu-mi35x-deepseek-v4-pro-dspark --nightly --timeout-per-file 7200 ${{ (github.event_name == 'schedule' || inputs.continue_on_error) && '--continue-on-error' || '' }}
echo "$(<github_summary.md )" >> $GITHUB_STEP_SUMMARY || true

# HiSparse needs its own server args (hisparse_config with host_to_device_ratio),
# so it cannot share a server with the non-HiSparse pro tests.
nightly-8-gpu-mi35x-deepseek-v4-pro-hisparse-unified:
name: ${{ format('nightly-8-gpu-mi35x-deepseek-v4-pro-hisparse-unified ({0}, linux-mi35x-gpu-8)', matrix.rocm_version) }}
strategy:
fail-fast: false
matrix:
rocm_version: ${{ fromJson(inputs.rocm_version && inputs.rocm_version != 'all' && format('["{0}"]', inputs.rocm_version) || '["rocm10", "rocm724", "rocm720"]') }}
if: (github.repository == 'sgl-project/sglang' || github.event_name == 'pull_request') && (!(inputs.job_filter || inputs.job_select) || (inputs.job_filter || inputs.job_select) == 'all' || contains(format(',{0},', inputs.job_filter || inputs.job_select), ',nightly-8-gpu-mi35x-deepseek-v4-pro-hisparse-unified,'))
runs-on: linux-mi35x-gpu-8
steps:
- name: Checkout code
uses: actions/checkout@v4
with:
ref: ${{ inputs.ref || github.sha }}

- name: Ensure VRAM is clear
run: bash scripts/ci/amd/ensure_vram_clear.sh rocm

- name: Setup docker (${{ matrix.rocm_version }})
run: |
touch github_summary.md
bash scripts/ci/amd/amd_ci_start_container.sh --rocm-version ${{ matrix.rocm_version }}
env:
GITHUB_WORKSPACE: ${{ github.workspace }}
ENABLE_CACHE_HOST: "1"

- name: Install dependencies
run: |
# --skip-test-time-deps: GSM8K doesn't need lmms-eval / human-eval.
bash scripts/ci/amd/amd_ci_install_dependency.sh --skip-test-time-deps
bash scripts/ci/amd/amd_ci_exec.sh pip install tabulate

- name: Accuracy Test MI35x ROCm (8-GPU DeepSeek-V4-Pro FP4 HiSparse unified-KV, TP8)
timeout-minutes: 300
run: |
> github_summary.md # Clear summary file
echo "## HiSparse unified-KV (TP8)" >> github_summary.md
bash scripts/ci/amd/amd_ci_exec.sh -w /sglang-checkout/test \
-e SGLANG_MOE_COPY_WEIGHT_VIEWS_BEFORE_H2D=1 \
-e GITHUB_STEP_SUMMARY="/sglang-checkout/github_summary.md" \
python3 run_suite.py --hw amd --suite nightly-amd-8-gpu-mi35x-deepseek-v4-pro-hisparse-unified --nightly --timeout-per-file 7200 ${{ (github.event_name == 'schedule' || inputs.continue_on_error) && '--continue-on-error' || '' }}
echo "$(<github_summary.md )" >> $GITHUB_STEP_SUMMARY || true

- name: Accuracy Test MI35x ROCm (8-GPU DeepSeek-V4-Pro FP4 HiSparse unified-KV, TP8 + DP8)
if: ${{ !cancelled() }}
timeout-minutes: 300
run: |
> github_summary.md # Clear summary file
echo "## HiSparse unified-KV (TP8 + DP8)" >> github_summary.md
bash scripts/ci/amd/amd_ci_exec.sh -w /sglang-checkout/test \
-e SGLANG_MOE_COPY_WEIGHT_VIEWS_BEFORE_H2D=1 \
-e GITHUB_STEP_SUMMARY="/sglang-checkout/github_summary.md" \
python3 run_suite.py --hw amd --suite nightly-amd-8-gpu-mi35x-deepseek-v4-pro-hisparse-unified-dp --nightly --timeout-per-file 7200 ${{ (github.event_name == 'schedule' || inputs.continue_on_error) && '--continue-on-error' || '' }}
echo "$(<github_summary.md )" >> $GITHUB_STEP_SUMMARY || true

# ==============================================================================
# 8-GPU Kimi-K2.6 (MI30x)
#
Expand Down Expand Up @@ -2354,6 +2411,7 @@ jobs:
- nightly-8-gpu-mi35x-deepseek-v4-pro
- nightly-8-gpu-mi35x-deepseek-v4-pro-mtp
- nightly-8-gpu-mi35x-deepseek-v4-pro-dspark
- nightly-8-gpu-mi35x-deepseek-v4-pro-hisparse-unified
# 8-GPU Kimi-K2.6 (MI30x)
- nightly-8-gpu-kimi-k26
# 8-GPU Kimi-K3 (MI35x only - native MXFP4 needs gfx95x)
Expand Down
18 changes: 7 additions & 11 deletions python/sglang/srt/arg_groups/hisparse_hook.py
Original file line number Diff line number Diff line change
Expand Up @@ -105,21 +105,17 @@ def validate_hisparse(server_args: ServerArgs) -> None:
# DSv4 hisparse handles its own dtype/backend pairing elsewhere; the dtype-
# aware checks below only apply to the DSA hisparse path.
if is_hip and is_v4_hisparse:
# TEMPORARY GUARD: DSv4 HiSparse is not supported on the unified-KV path.
# In unified-KV mode c4_kv_pool is None, so DeepSeekV4HiSparseTokenToKVPoolAllocator
# cannot attach and pool init dies with a cryptic AssertionError. Fail fast
# at startup with a clear message instead. Remove once unified-KV HiSparse lands.
# bf16 unified-KV HiSparse is wired (shrunk C4 device region + host
# mirror). fp8 is a two-buffer layout and still needs its own adapter.
from sglang.kernels.ops.attention.dsv4.unified_kv_kernels.env_gate import (
is_unified_kv_triton,
is_unified_kv_fp8,
)

if is_unified_kv_triton():
if is_unified_kv_fp8():
raise ValueError(
"--enable-hisparse is not supported with the unified-KV path on ROCm"
"(SGLANG_HACK_FLASHMLA_BACKEND=unified_kv_triton) for DeepSeek-V4: "
"HiSparse currently requires the separate packed KV layout. "
"Either set SGLANG_HACK_FLASHMLA_BACKEND=triton, or run without "
"--enable-hisparse."
"--enable-hisparse is not supported with unified-KV fp8 on ROCm "
"(SGLANG_DSV4_UNIFIED_KV_FP8=1). Unset that env to use bf16 "
"unified-KV HiSparse, or run without --enable-hisparse."
)
return

Expand Down
22 changes: 18 additions & 4 deletions python/sglang/srt/layers/attention/dsv4/compressor_v2.py
Original file line number Diff line number Diff line change
Expand Up @@ -208,10 +208,24 @@ def forward_unified(
elif is_unified_kv_triton():
kv_cache = token_to_kv_pool.get_unified_kv(layer_id)
page_size = 1
out_loc = getattr(
self.forward_metadata.core_metadata.unified,
f"c{compressor.ratio}_out_loc",
)
if token_to_kv_pool.unified_hisparse and compressor.ratio == 4:
# ROCm HiSparse unified-KV: the C4 compressed region is a hot
# device buffer backed by a host cold pool. Remap the RAW
# compressed slot (core_metadata.c4_out_loc, no swa offset --
# the remap table is keyed by raw compressed index) to its hot
# device row, then add the swa_pages offset into the unified
# compressed region. Only C4 layers are offloaded; C128 (HCA)
# layers keep the dense compressed slot via the swa-offset
# unified out_loc below.
out_loc = token_to_kv_pool.c4_kv_pool._translate_loc_to_hisparse_device(
self.forward_metadata.core_metadata.c4_out_loc
)
out_loc = out_loc + token_to_kv_pool.unified_swa_pages
else:
out_loc = getattr(
self.forward_metadata.core_metadata.unified,
f"c{compressor.ratio}_out_loc",
)
if is_unified_kv_fp8():
fp8_2buff = True
kv_cache_rope = token_to_kv_pool.get_unified_kv_rope(layer_id)
Expand Down
63 changes: 49 additions & 14 deletions python/sglang/srt/managers/hisparse_coordinator.py
Original file line number Diff line number Diff line change
Expand Up @@ -170,9 +170,38 @@ def __init__(
self.token_to_kv_pool_allocator, DeepSeekV4HiSparseTokenToKVPoolAllocator
)
self.is_m3_hisparse = isinstance(kvcache, MiniMaxSparseKVPool)
# unified-KV HiSparse (ROCm) keeps the compressed C4 hot/cold data in
# the bf16 unified layout, so swap-in/backup use the generic linear MLA
# path instead of the FP8 page-padded dsv4 path.
self.is_unified_hisparse = False
if self.is_dsv4_hisparse:
from sglang.srt.mem_cache.deepseek_v4_memory_pool import (
HiSparseUnifiedC4DevicePool,
)

self.is_unified_hisparse = isinstance(
self.token_to_kv_pool_allocator.hisparse_kvcache,
HiSparseUnifiedC4DevicePool,
)
# Byte layout of the host/device C4 buffers, as opposed to the pool
# structure that is_dsv4_hisparse selects: only the separate-KV path is
# page-padded, so it alone may take the dsv4 swap/copy kernels.
self.is_dsv4_paged_layout = (
self.is_dsv4_hisparse and not self.is_unified_hisparse
)
if self.is_dsv4_hisparse:
self.mem_pool_device = self.token_to_kv_pool_allocator.hisparse_kvcache
page_size = self.mem_pool_device.page_size
# Host cold-pool sizing (unified-KV / dsv4 path): the host pool is
# bound to size_full/compress_ratio -- a mirror of the GPU full-token
# c4 budget, not an independent expansion into host RAM.
# host_to_device_ratio (== c4_shrink_factor, >= 1) sets how much of
# that budget stays GPU-resident: ratio=1 keeps the full c4 pool
# on-GPU (host == device, a 1:1 mirror); ratio>1 shrinks the GPU c4
# pool (freeing memory for a larger overall token budget) while the
# host mirror still tracks the full budget. This is orthogonal to the
# per-request device_buffer_size hot window (the swap trigger), which
# is never scaled by ratio.
num_host_pages = (
self.token_to_kv_pool_allocator.size_full // self.compress_ratio
+ page_size
Expand Down Expand Up @@ -825,11 +854,17 @@ def naive_load_topk(
"""Load top-k selected tokens into device memory and return their device indices.

This is a naive per-request loop implementation for debugging/validation.
Production code uses swap_in_selected_pages (JIT CUDA kernel) instead.

Note: dsv4 hisparse is not supported — DeepSeekV4SingleKVPoolHost has no
load_to_device_per_layer and indices live in compressed space. Currently
only used as a kernel oracle in test_hisparse_unit.py (non-dsv4 path).
Production code uses swap_in_selected_pages (JIT CUDA/HIP kernel) instead.
Used as a kernel oracle in test_hisparse_unit.py.

Both the non-dsv4 and the dsv4 (DeepSeek V4 separate-KV C4) layouts are
supported. The per-request loc-selection logic below is layout-agnostic:
it indexes req_to_device_buffer / req_to_host_pool and delegates the
host->device byte copy to mem_pool_host.load_to_device_per_layer. For the
dsv4 path mem_pool_host is a DeepSeekV4PagedHostPool (which implements
load_to_device_per_layer for the token-granular, page-padded C4 value/
scale layout), and seq_lens / top_k_tokens are expressed in compressed C4
space (matching swap_in_selected_pages, called with compressed_seq_lens).

Args:
req_pool_indices: Pool indices for each request. Shape: (num_reqs,)
Expand All @@ -840,9 +875,6 @@ def naive_load_topk(
Returns:
Device KV cache indices for the selected tokens. Shape: (num_reqs, top_k)
"""
assert not self.is_dsv4_hisparse, (
"naive_load_topk is not implemented for dsv4 hisparse"
)
num_reqs = req_pool_indices.size(0)
top_k_indices = torch.full(
(num_reqs, self.top_k), -1, dtype=torch.int32, device=self.device
Expand Down Expand Up @@ -1008,11 +1040,14 @@ def _run_swap_in_kernel(
"""
num_reqs = req_pool_indices.size(0)
top_k_indices = self.top_k_device_locs_buffer[:num_reqs, : self.top_k]
swap_in_fn = (
load_cache_to_device_buffer_dsv4_mla
if self.is_dsv4_hisparse
else load_cache_to_device_buffer_mla
)

if self.is_dsv4_paged_layout:
# separate-KV: FP8 page-padded device + host C4 layout.
swap_in_fn = load_cache_to_device_buffer_dsv4_mla
else:
# unified-KV HiSparse (bf16 linear rows) and generic DSA MLA both
# use the linear swap path (stride == item_size_bytes).
swap_in_fn = load_cache_to_device_buffer_mla
plan = (
dict(
miss_src=self._miss_src[:num_reqs],
Expand Down Expand Up @@ -1097,7 +1132,7 @@ def _run_copy_only_kernel(self, num_reqs: int, skip_layer: int) -> None:
device_buffer=self.mem_pool_device.kv_buffer[skip_layer],
item_size_bytes=self.item_size_bytes,
num_blocks=self._prefetch_copy_blocks,
is_dsv4_layout=self.is_dsv4_hisparse,
is_dsv4_layout=self.is_dsv4_paged_layout,
skip_io=self.skip_io,
)

Expand Down
29 changes: 21 additions & 8 deletions python/sglang/srt/mem_cache/allocator/hisparse.py
Original file line number Diff line number Diff line change
Expand Up @@ -10,6 +10,7 @@
from sglang.srt.mem_cache.deepseek_v4_memory_pool import (
DeepSeekV4TokenToKVPool,
HiSparseC4DevicePool,
HiSparseUnifiedC4DevicePool,
)
from sglang.srt.mem_cache.hisparse_memory_pool import HiSparseDSATokenToKVPool
from sglang.srt.utils.common import get_num_new_pages
Expand Down Expand Up @@ -279,6 +280,11 @@ def __init__(
self.compress_ratio = 4

self.hisparse_kvcache = logical_attn_allocator._kvcache.c4_kv_pool
# Unified-KV (ROCm) mirrors the full C4 budget on the host cold pool;
# only it reports host-backed capacity. CUDA separate-KV is unchanged.
self._is_unified_hisparse = isinstance(
self.hisparse_kvcache, HiSparseUnifiedC4DevicePool
)
self._size_full = logical_attn_allocator.size_full
self._size_hisparse = self.hisparse_kvcache.size

Expand Down Expand Up @@ -374,10 +380,14 @@ def swa_capacity_and_available(self, *, full_capacity, swa_capacity):
)

def full_available_size(self):
return min(
self.logical_attn_allocator.full_available_size(),
self.hisparse_attn_allocator.available_size() * self.compress_ratio,
)
if not self._is_unified_hisparse:
return min(
self.logical_attn_allocator.full_available_size(),
self.hisparse_attn_allocator.available_size() * self.compress_ratio,
)
# unified-KV: cold C4 is host-resident, so capacity tracks the logical
# full-token pool; device pressure is handled by alloc_extend.
return self.logical_attn_allocator.full_available_size()

def swa_available_size(self):
return self.logical_attn_allocator.swa_available_size()
Expand Down Expand Up @@ -406,10 +416,13 @@ def free_full(self, free_indices: torch.Tensor):
self.full_free_group.append(self._copy_for_free_group(free_indices))

def available_size(self) -> int:
return min(
self.logical_attn_allocator.available_size(),
self.hisparse_attn_allocator.available_size() * self.compress_ratio,
)
if not self._is_unified_hisparse:
return min(
self.logical_attn_allocator.available_size(),
self.hisparse_attn_allocator.available_size() * self.compress_ratio,
)
# unified-KV: see full_available_size.
return self.logical_attn_allocator.available_size()

def alloc(self, need_size: int):
raise NotImplementedError(
Expand Down
Loading
Loading