Skip to content

[AMD] Enable unified-KV HiSparse on ROCm for DeepSeek-V4 - #29168

Open
amd-danli103 wants to merge 21 commits into
sgl-project:mainfrom
amd-danli103:hisparse-dsv4pro-unified-kv-rocm-clean
Open

amd-danli103 wants to merge 21 commits into
sgl-project:mainfrom
amd-danli103:hisparse-dsv4pro-unified-kv-rocm-clean

Conversation

@amd-danli103

@amd-danli103 amd-danli103 commented Jun 24, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

On ROCm, HiSparse currently runs only on the separate-KV DeepSeek-V4 path (SGLANG_HACK_FLASHMLA_BACKEND=triton, packed FP8 KV layout), enabled by #26639. The unified-KV backend (SGLANG_HACK_FLASHMLA_BACKEND=unified_kv_triton) — where compressed C4 KV lives in the unified pool's rows[swa_pages:] and the HiSparse hot buffer is a bf16 view into that region backed by a linear host cold pool — was blocked by a temporary startup guard.
This PR enables unified-KV + HiSparse on ROCm end-to-end, and adds JIT/unit/eval coverage.

Modifications

  • Enable unified-KV HiSparse on ROCm (hisparse_coordinator.py, deepseek_v4_memory_pool.py, compressor_v2.py, arg_groups/hisparse_hook.py, srt/mem_cache/allocator/hisparse.py): remove the temporary guard that blocked unified_kv_triton + HiSparse, wire the unified C4 device pool / linear host cold pool, and fix the compressed-C4 remap to account for the pre-applied swa_pages offset.
  • Enable dsv4 paged HiSparse JIT kernel tests on ROCm (test/registered/kernels/ops/kvcache/test_hisparse.py).
  • Honor the C4 shrink factor on the unified path (deepseek_v4_memory_pool.py, mem_cache/allocator/hisparse.py): upstream pool_configurator already folds host_to_device_ratio = S into bytes_per_full_token as a 1/(4S) C4 term and raises max_total_num_tokens accordingly, but the unified pool ignored it. This sizes the on-device C4 region to the shrunk budget via c4_compress_pages, and stops available_size() / full_available_size() from clamping capacity to the device C4 pool — on the unified path the cold remainder is host-resident, so the bound is the logical full-token pool. Net effect: freed VRAM actually reaches full_token. Separate-KV / CUDA behaviour is unchanged.
  • Unit tests for the unified-KV HiSparse pool (test/registered/unit/mem_cache/test_hisparse_unified_pool.py): host-mirror geometry, view aliasing of the unified compressed region, c4-only layer views, alloc/free mapping, oversubscribe and that c4_compress_pages actually shrinks the device C4 region while the c128 layers are untouched.
  • Eval tests (.../mi35x/..._unified_eval_mi35x.py and ..._unified_dp_eval_mi35x.py): GSM8K accuracy + a ~19k-token long-context swap-path retrieval test, on TP8 and on TP8+DP-attention.

Highlight — the NVIDIA CUDA forward path is unchanged. All modifications are confined to ROCm-only branches: _is_hip guards and HiSparse/unified-KV-gated code. (The JIT test change only removes skipif(is_hip()) markers, so those tests now also run on ROCm; CUDA behavior is unaffected.)

To enable this feature, you can pass e.g.

    --enable-hisparse \
    --hisparse-config '{"top_k": 1024, "device_buffer_size": 2048, "host_to_device_ratio": 4}'

in the server launch configs.

Accuracy Tests

GSM8K few-shot accuracy alone does not validate the host↔device swap path. With device_buffer_size=2048 (compressed-token budget) and GSM8K prompts of ~1k–1.7k tokens (~250–425 compressed after the 4x C4 ratio), every request stays on the fast path (host_len <= device_buffer_size → preloaded to the device hot buffer), so decode is all device-resident hits and the cold→hot swap-in (miss-copy) kernel is never exercised. GSM8K therefore validates: sparse top-k selection numerics, unified-KV pool wiring/aliasing, the compressor write path, and HIP-graph decode stability — but not swap correctness.

Note: a nonzero #cpu token in the decode log during GSM8K is the authoritative host backup of compressed C4 (every request is backed up to host), not decode-time swap-in. The swap-in (miss-copy) path is gated per-request by host_len > device_buffer_size; GSM8K's per-request footprint (~250–425 compressed) stays below device_buffer_size=2048, so it runs entirely on the preloaded hot path.

Swap correctness is established by, in increasing scope:

  1. JIT kernel tests (test/registered/kernels/ops/kvcache/test_hisparse.py): test_dsv4_swap_in_reads_paged_host_layout, *_miss_copy_layout, and the LRU/miss cases check the swap-in kernel byte-for-byte against a reference.
  2. End-to-end long-context swap test (test_b_long_context_swap): a ~19k token prompt (~4750 compressed > 2048) forces a host cold pool and a swap-in on every decode step; a passcode buried in the long context can only be retrieved if cold→hot copies land correctly. This is the authoritative end-to-end swap oracle.

Accuracy of record: The numbers reported here (TP8 0.953 / TP8 DP8 0.949) were reproduced locally on deepseek-ai/DeepSeek-V4-Pro.

Tests reproduction guidance

JIT kernel (11 passed)

HIP_VISIBLE_DEVICES=0 python -m pytest test/registered/kernels/ops/kvcache/test_hisparse.py -v

These now also run on ROCm (the skipif(is_hip()) markers were removed). The swap-in kernel cases (test_dsv4_swap_in_reads_paged_host_layout, *_miss_copy_layout, LRU/miss) compare the host→device cold-copy byte-for-byte against a reference, so a pass proves the swap-in kernel is numerically correct.

unit test (6 passed)

HIP_VISIBLE_DEVICES=0 python -m pytest test/registered/unit/mem_cache/test_hisparse_unified_pool.py -v

Validates the unified-KV pool: host-mirror geometry, the bf16 view aliasing into rows[swa_pages:], C4-only layer views, alloc/free index mapping, and oversubscribe handling.

e2e test (2 passed each)

python -m pytest -s -v test/registered/amd/accuracy/mi35x/test_deepseek_v4_pro_fp4_hisparse_unified_eval_mi35x.py
python -m pytest -s -v test/registered/amd/accuracy/mi35x/test_deepseek_v4_pro_fp4_hisparse_unified_dp_eval_mi35x.py

Each file runs two methods. What proves the feature is live:
(a) Server startup — the unified-KV HiSparse host cold pool is allocated, and the config is applied:

... enable_hisparse=True, hisparse_config='{"top_k": 1024, "device_buffer_size": 2048, "host_to_device_ratio": 2}'
... Allocating 31.35 GB host memory for V4 paged pool 'dsv4_hisparse_c4' ...

The dsv4_hisparse_c4 host pool only exists on the HiSparse path; its presence with unified_kv_triton confirms the unified hot/cold wiring is taken.
(b) test_a_gsm8k — accuracy / sparse-selection guard. Expect the method to PASS with GSM8K accuracy ≈ 0.95 (assertion gate > 0.91).
(c) test_b_long_context_swap — the swap-path guard, and the authoritative evidence that host↔device swap works. A ~19k-token prompt with a passcode buried at the top is sent; expect the method to PASS with the passcode echoed back:

long_context_swap completion='7G-ZULU-4419'

The corresponding server decode line is the smoking gun — a single long request whose compressed C4 footprint far exceeds device_buffer_size=2048:

Decode batch, #running-req: 1,, ..., #gpu token: 2049, gpu token usage: 0.00, #cpu token: 9566, cpu token usage: 0.01, cuda graph: True, ...
  • #gpu token: 2049 ≈ device_buffer_size (2048) → the device hot buffer is capped at the configured budget.
  • #cpu token: 9566 (> 0) → the remaining cold C4 tokens live in the host pool and are streamed in on every decode step.
  • The passcode can only be retrieved if those cold→hot swap-ins land correctly, so a PASS proves end-to-end swap correctness (the DP file shows the same behavior under --dp 8 --enable-dp-attention).

Speed Tests and Profiling

This is a capacity feature, not a latency one, and it is worth being explicit about the trade rather than reporting a throughput number that misses the point.
host_to_device_ratio = S shrinks the GPU-resident C4 region to 1/S of the budget and mirrors the remainder on the host. Upstream _get_bytes_per_full_token already carries the resulting 1/(4S) C4 term, so the freed VRAM shows up
directly as a larger max_total_num_tokens. #32368 measures 12.90M → 17.56M (+36%) on the same 1P1D topology.
The cost is decode-step latency: the swap-in runs on the main stream between the indexer and attention of every C4 layer. It is latency/occupancy-bound, not bandwidth-bound — LaunchKernel uses grid = batch size, so a per-rank decode batch of ~20 leaves ~9% occupancy on 256 CUs, which is why #32368's device_buffer_size sweep (2048 → 16384, ~8x fewer bytes moved) made decode monotonically worse. The levers are multi-CTA sharding (#31341) and
side-stream overlap (#28523), neither of which is in this PR's scope.
When to turn this on: HiSparse trades decode latency for token capacity, so it is a loss whenever the KV working set of your concurrent requests already fits in VRAM — you pay the swap-in cost and get nothing back. It pays off only past
the point where the non-HiSparse configuration can no longer admit the load, at which point the comparison is not "slower" but "runs at all".

CI coverage

PR-CI:

  • test/registered/kernels/ops/kvcache/test_hisparse.py (HiSparse JIT kernels; HIP skip removed so it now runs on ROCm)
  • test/registered/unit/mem_cache/test_hisparse_unified_pool.py (unified-KV C4 device pool)
    These give per-PR coverage of the unified-KV + HiSparse pool/kernel paths.

Nightly e2e (8-GPU TP8 / TP8 DP8, nightly=True):

  • nightly-amd-8-gpu-mi35x-deepseek-v4-pro-hisparse-unified
  • nightly-amd-8-gpu-mi35x-deepseek-v4-pro-hisparse-unified-dp

These two suites are registered but not yet wired into any workflow, so they do not run on PR or on the nightly schedule yet. Following the precedent of the unified-KV PR #27380, the PR-CI/nightly wiring is tracked in the AMD CI coverage issue #27521 (cc @bingxche @yctseng0211). I will add them to the issue list then.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #35813712809
Latest PR Test (Extra): ❌ Run #35813712476
Latest PR Test (AMD ROCm 10): ❌ Run #35813712792

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request implements unified-KV HiSparse support on ROCm for DeepSeek-V4, allowing the compressed C4 KV cache to reside directly within the unified pool's rows. Key updates include removing the temporary ROCm guard, introducing HiSparseUnifiedC4DevicePool to alias the unified compressed region, updating the swap-in logic to use the linear MLA path, and preventing runtime reallocations of DSA decode metadata buffers. The review feedback highlights a critical indexing bug in compressor_v2.py where unmapped slots (indicated by -1) could corrupt the SWA ring, recommends raising a RuntimeError to guard against post-capture buffer reallocations in dsa_backend.py, and suggests initializing inherited KVCache attributes in HiSparseUnifiedC4DevicePool to prevent potential AttributeErrors.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +523 to +528
out_loc = (
token_to_kv_pool.c4_kv_pool._translate_loc_to_hisparse_device(
self.forward_metadata.core_metadata.c4_out_loc
)
)
out_loc = out_loc + token_to_kv_pool.unified_swa_pages

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

If c4_out_loc contains -1 (representing an invalid or unmapped slot), _translate_loc_to_hisparse_device will return -1 (as the last element of the mapping is -1). Doing out_loc + token_to_kv_pool.unified_swa_pages will then result in -1 + unified_swa_pages. Since unified_swa_pages is positive, this becomes a valid positive index (e.g., unified_swa_pages - 1), which is the last slot of the SWA ring. This will cause the store kernel to write the compressed KV of the unmapped/invalid token into the SWA ring, corrupting the SWA ring's KV cache.

To prevent this, use torch.where to only add unified_swa_pages to valid (non-negative) indices, keeping -1 as -1.

                    out_loc = (
                        token_to_kv_pool.c4_kv_pool._translate_loc_to_hisparse_device(
                            self.forward_metadata.core_metadata.c4_out_loc
                        )
                    )
                    out_loc = torch.where(
                        out_loc >= 0,
                        out_loc + token_to_kv_pool.unified_swa_pages,
                        out_loc,
                    )

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

c4_out_loc = tl.where(c4_should_compress, raw_out_loc // 4, 0) so c4_out_loc won't contain -1. Assumption does not make sense.

Comment thread python/sglang/srt/layers/attention/dsa_backend.py Outdated
Comment on lines +317 to +321
self.data_ptrs = torch.tensor(
[x.data_ptr() for x in self.kv_buffer],
dtype=torch.uint64,
device=self.device,
)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

HiSparseUnifiedC4DevicePool inherits from HiSparseC4DevicePool (which inherits from DeepSeekV4SingleKVPool -> KVCache), but its __init__ intentionally skips calling super().__init__ to avoid allocating a second device buffer. As a result, several attributes defined in KVCache.__init__ (such as self.custom_mem_pool, self.layer_transfer_counter, self.enable_memory_saver, and self.memory_saver_adapter) are left uninitialized. If other parts of the codebase interact with HiSparseUnifiedC4DevicePool expecting it to be a fully initialized KVCache instance, this could lead to AttributeErrors. We should explicitly initialize these missing attributes to safe defaults to ensure robust compatibility.

Suggested change
self.data_ptrs = torch.tensor(
[x.data_ptr() for x in self.kv_buffer],
dtype=torch.uint64,
device=self.device,
)
self.data_ptrs = torch.tensor(
[x.data_ptr() for x in self.kv_buffer],
dtype=torch.uint64,
device=self.device,
)
self.custom_mem_pool = None
self.layer_transfer_counter = None
self.enable_memory_saver = False
from sglang.srt.mem_cache.memory_pool import get_memory_saver_adapter
self.memory_saver_adapter = get_memory_saver_adapter(False)

@amd-danli103 amd-danli103 mentioned this pull request Jun 25, 2026
2 tasks
@amd-danli103
amd-danli103 force-pushed the hisparse-dsv4pro-unified-kv-rocm-clean branch from 9337e0e to 7c21ef9 Compare June 26, 2026 03:02
@amd-danli103

Copy link
Copy Markdown
Contributor Author

Code Review

This pull request implements unified-KV HiSparse support on ROCm for DeepSeek-V4, allowing the compressed C4 KV cache to reside directly within the unified pool's rows. Key updates include removing the temporary ROCm guard, introducing HiSparseUnifiedC4DevicePool to alias the unified compressed region, updating the swap-in logic to use the linear MLA path, and preventing runtime reallocations of DSA decode metadata buffers. The review feedback highlights a critical indexing bug in compressor_v2.py where unmapped slots (indicated by -1) could corrupt the SWA ring, recommends raising a RuntimeError to guard against post-capture buffer reallocations in dsa_backend.py, and suggests initializing inherited KVCache attributes in HiSparseUnifiedC4DevicePool to prevent potential AttributeErrors.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026. For more details on the timeline and next steps, please review the Help Documentation.

Removed "preventing runtime reallocations of DSA decode metadata buffers" - this potential bug fix from this PR. Rationale: that fix is orthogonal to this feature. It will be submitted as a separate PR.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@amd-danli103

Copy link
Copy Markdown
Contributor Author

hi @1am9trash or @HaiShaw , could u pls help to trigger the CI test? Thank you!

@yctseng0211 yctseng0211 added the run-ci CI: run the baseline test suite on this PR label Jun 26, 2026

@clintg6 clintg6 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Verified all 19 tests pass locally: JIT 10/10, unit 5/5, TP8 e2e 2/2 (GSM8K 0.952), TP8+DP8 e2e 2/2 (GSM8K 0.950). Long-context swap path confirmed working.

All runtime changes are safely gated behind _is_hip / unified_hisparse so no impact on the NVIDIA path.

Gemini findings reviewed: the -1 index concern in compressor_v2.py is a false positive (c4_out_loc defaults to 0 via tl.where, not -1). The skipped super().init in HiSparseUnifiedC4DevicePool is intentional and safe as the uninitialized KVCache attributes are never accessed on this pool.

@amd-danli103 @HaiShaw

@amd-danli103

Copy link
Copy Markdown
Contributor Author

Verified all 19 tests pass locally: JIT 10/10, unit 5/5, TP8 e2e 2/2 (GSM8K 0.952), TP8+DP8 e2e 2/2 (GSM8K 0.950). Long-context swap path confirmed working.

All runtime changes are safely gated behind _is_hip / unified_hisparse so no impact on the NVIDIA path.

Gemini findings reviewed: the -1 index concern in compressor_v2.py is a false positive (c4_out_loc defaults to 0 via tl.where, not -1). The skipped super().init in HiSparseUnifiedC4DevicePool is intentional and safe as the uninitialized KVCache attributes are never accessed on this pool.

@amd-danli103 @HaiShaw

Hi @clintg6 , thank you so much for the thorough verification and the approval! The results aligned with my local results.

@amd-danli103

Copy link
Copy Markdown
Contributor Author

@amd-bot ci-status

@amd-bot

amd-bot commented Jun 29, 2026

Copy link
Copy Markdown

@amd-danli103

CI Status for PR #29168

Merge verdict: 🚫 Do not merge on CI signal — PR CI never ran. All 20 "failures" are the pr-gate job rejecting the PR for a missing run-ci label (PR currently carries only deepseek), plus the *-finish aggregators failing on the cascade. Every actual test job was skipped, so none of this PR's changed code was exercised. The PR is also reported mergeable: false (needs a rebase/conflict resolution against main).

Caution

This PR's changed code is not exercised by any PR-CI test — every test job was skipped because the gate failed. The pr-gate job across all 10 workflows exits with Missing required label 'run-ci', which fast-fails the gate and skips all downstream stages. A green-looking signal is impossible and the red Xs are not code failures. Action required: add the run-ci label (and run-ci-extra if the Extra suites are wanted), resolve the merge conflict, and re-run before judging CI.

Changed files: hisparse_hook.py (+4/-16), compressor_v2.py (+20/-4), hisparse_coordinator.py (+41/-13), deepseek_v4_memory_pool.py (+95/-0), test_hisparse.py (+0/-2), and 3 new AMD/unit tests (test_hisparse_unified_pool.py, test_deepseek_v4_pro_fp4_hisparse_unified_eval_mi35x.py, test_deepseek_v4_pro_fp4_hisparse_unified_dp_eval_mi35x.py).

Executed CI failure attribution: AMD: 0 executed failures · Others: 0 executed failures — all failures are the gate-block cascade, not executed tests. No test job ran in any workflow.

Root cause (single, shared across all workflows)

Workflow Failing jobs Test File Test Function Error Related? Why
PR #29168 (base) call-gate / pr-gate, pr-test-finish N/A N/A Missing required label 'run-ci' → exit 1 🟢 Gate config check, not code. Log
PR Test (AMD) call-gate / pr-gate, pr-test-amd-finish N/A N/A same 🟢 same gate block
PR Test ROCm 7.2 (AMD) pr-gate, pr-test-amd-rocm720-finish N/A N/A same 🟢 same gate block
PR Test Extra (AMD) pr-gate, pr-test-amd-extra-finish N/A N/A same 🟢 same gate block
PR Test Extra pr-gate, pr-test-extra-finish N/A N/A same 🟢 same gate block
PR Test (NPU / Xeon / MUSA / Arm64 / XPU) pr-gate (+musa-finish) N/A N/A same 🟢 same gate block

Every downstream stage (stage-a/b/c-*, base-a/b/c-*, extra-*, kernel/unit/multimodal tests) shows skipped — confirmed by job listings. The *-finish jobs fail only because they detect call-gate: failure.

Coverage gap — what is currently unverified

This PR adds a new memory pool (deepseek_v4_memory_pool.py) and HiSparse coordinator/compressor changes, plus three new tests intended to cover them:

  • test/registered/unit/mem_cache/test_hisparse_unified_pool.py — unit coverage for the new pool
  • test/registered/amd/accuracy/mi35x/test_deepseek_v4_pro_fp4_hisparse_unified_eval_mi35x.py and the _dp_eval_ variant — AMD MI35x accuracy

None of these ran. Until the gate passes, there is zero signal on the changed code.

What to do before merge

  1. Add the run-ci label to PR [AMD] Enable unified-KV HiSparse on ROCm for DeepSeek-V4 #29168 (the only label present is deepseek). Add run-ci-extra too if the Extra suites should run. See the CI trigger docs.
  2. Resolve the merge conflict — GitHub reports mergeable: false; rebase/merge main so CI runs against a clean tree.
  3. After re-running, confirm the new AMD MI35x accuracy + unit pool tests actually execute and pass — those are the suites that verify this PR's value. The non-AMD backends (NPU/Xeon/MUSA/Arm64/XPU) showing red here are also purely the gate cascade and can be ignored once the label is added.

Generated by amd-bot using Claude Code CLI

@amd-danli103
amd-danli103 force-pushed the hisparse-dsv4pro-unified-kv-rocm-clean branch from 7c21ef9 to 194e347 Compare June 29, 2026 03:18
@amd-danli103

Copy link
Copy Markdown
Contributor Author

/tag-and-rerun-ci

@amd-danli103

Copy link
Copy Markdown
Contributor Author

@amd-bot ci-status

@amd-bot

amd-bot commented Jun 30, 2026

Copy link
Copy Markdown

@amd-danli103

CI Status for PR #29168

Merge verdict: No executed CI failure is attributable to this PR — every red job is an unrelated pre-existing/flaky/infra failure (FP-comparator flake, perf thresholds, a pytest-missing build env, a stale test signature). The new code's unit + JIT tests passed on AMD (the target backend). BUT PR CI did not verify the actual feature: the two new end-to-end accuracy evals are nightly=True (8-GPU MI35x) and never ran, and NVIDIA base-b was fast-fail-truncated. Green here does not prove unified-KV HiSparse works on ROCm — run the nightly suites before merging on the basis of correctness.

Caution

This PR's headline value — DeepSeek-V4 Pro FP4 HiSparse unified-KV accuracy on ROCm — is exercised only by the two new tests registered nightly=True (nightly-amd-8-gpu-mi35x-deepseek-v4-pro-hisparse-unified / …-dp), which do not run on PR CI. The unit test (test_hisparse_unified_pool.py) and JIT test (jit/test_hisparse.py) ran and passed on AMD, but they only cover the memory-pool/compressor at unit level. Author must run the nightly MI35x 8-GPU eval suites to verify end-to-end accuracy before trusting a green PR run.

Warning

NVIDIA PR CI is incomplete: in run 28346854737, base-b-test-1-gpu-small (1) hit an unrelated flaky failure and fast-failed, skipping 14 downstream base-b shards. The pool unit test is also registered to NVIDIA base-b-1gpu-small and was likely in a skipped shard there (it did pass on AMD). Re-run NVIDIA base-b (or use bypass-fastfail sparingly) if NVIDIA signal is needed.

Changed files: deepseek_v4_memory_pool.py (+95), hisparse_coordinator.py, compressor_v2.py, hisparse_hook.py, +2 nightly MI35x eval tests (+466), test_hisparse_unified_pool.py (+246), jit/test_hisparse.py.

Executed CI failure attribution: AMD: 4 executed failures (0 related) · Others: 4 executed failures (0 related) · plus 2 fast-fail/gate cascades. NVIDIA base-b: 1 real failure + 14 skipped (cascade).

AMD Executed Failures

Job Test File Test Function Error Related? Why
stage-b-1gpu-small (8) test/registered/utils/test_type_based_dispatcher.py test_type_dispatcher_e2e_performance TypeError: Missing required argument 'input_embeds' 🟢 Pre-existing test out of sync with TokenizedGenerateReqInput ctor; unrelated to hisparse. A fix branch amd_fix_test_type_based_dispatcher already exists.
stage-b-1gpu-small (11) test/registered/vlm/test_vision_chunked_prefill.py (server requests) JSONDecodeError: Expecting value: line 1 column 1 🟢 VLM server returned empty body — infra/flake, no hisparse code path.
stage-b-1gpu-small-mi35x test/registered/quant/test_quark_mxfp4.py (load_weights check) ValueError: Should have found memory increase in load_weights 🟢 Quark MXFP4 quant test; unrelated to DeepSeek-V4 KV pool.
stage-b-2gpu-large (0) test/registered/perf/test_bench_serving_2gpu.py test_pp_offline_throughput_default_decode AssertionError: 5463.5 not greater than 6700 🟢 Perf-throughput threshold (Llama-3.1-8B PP); flake/regression unrelated to PR.

wait-for-stage-b-amd and pr-test-amd-finish are aggregator/gate jobs that went red because of the four rows above — not independent failures.

Other Executed Failures

Job Test File Test Function Error Related? Why
base-b-1gpu-small (1) test/registered/unit/utils/test_weight_checker_comparator.py test_chunked_result_matches_unchunked AssertionError: mean_abs_err 3.13791624e-05 != 3.13791606e-05 🟢 Flaky FP comparator (~1.8e-13 diff). Root of the 14-shard NVIDIA base-b cascade. Not hisparse.
stage-b-1gpu-xpu test/registered/xpu/test_triton_attention_backend.py (bench serving) ClientPayloadError / TransferEncodingError 400 🟢 XPU server connection drop — infra/flake, different backend.
stage-b-test-4-npu-a3 test/registered/ascend/basic_function/quant/test_npu_w4a4_quantization.py (throughput) AssertionError: 917.29 not >= 1000 🟢 Ascend NPU perf threshold; unrelated.
build-test (Arm64) test/registered/cpu/test_activation.py N/A ModuleNotFoundError: No module named 'pytest' 🟢 Arm64 build env missing pytest — infra, not PR code.

PR Test Extra and PR Test Extra (AMD) failed only at call-gate / pr-gate (exit 1) — gate cascades from the above, not independent test failures. The NPU multimodal-gen-test-2-npu-a3 job also exited 1 with no PR-related signal.

Details / what to do before merge

  • Coverage (do this): run the two new nightly suites on 8-GPU MI35x — nightly-amd-8-gpu-mi35x-deepseek-v4-pro-hisparse-unified and …-unified-dp — to actually validate accuracy of unified-KV HiSparse. PR-CI green does not cover them.
  • Unit coverage that DID run (good): test_hisparse_unified_pool.py passed on AMD (shard 13, job 83992600619); jit/test_hisparse.py passed on AMD (shard 4, job 83992600617).
  • NVIDIA completeness: if NVIDIA signal matters, re-run base-b after the flaky test_weight_checker_comparator clears, since 14 shards were skipped.
  • No action needed on any of the 8 red executed jobs — all are unrelated to this PR.

Generated by amd-bot using Claude Code CLI

@amd-danli103

Copy link
Copy Markdown
Contributor Author

/rerun-failed-ci

Lzy17 added a commit to Lzy17/sglang that referenced this pull request Aug 31, 2026
…nts)

New MI355X 2-node 1P1D leg exercising both memory features in one PD run:
HiCache on the prefill role, HiSparse on the decode role, KV transfer over MoRI,
unified-KV layout. The two features live on opposite PD roles and do not
conflict (HiCache is prefill-only, HiSparse is decode-only), so a single recipe
per model covers both. Added for all four DSV4 variants:
  mi355x-fp8/dsv4pro,  mi355x-fp8/dsv4flash,
  mi355x-fp4/dsv4pro,  mi355x-fp4/dsv4flash

Two launcher changes the recipes depend on:
- Env ordering: a recipe's prefill/decode_extra_env is now applied after the
  hardcoded DSV4 env so it can override it (needed to pin
  SGLANG_HACK_FLASHMLA_BACKEND=unified_kv_triton and the MoRI host-registration
  knobs).
- disable_radix_cache knob: --disable-radix-cache was hardcoded into all four
  COMMON_FLAGS strings; HiCache L1 *is* the radix cache and the combination is a
  hard ValueError. The recipe field disable_radix_cache: false now drives a
  RADIX_FLAG (defined before the role if/else so both branches see it under
  set -u), so HiCache can run on the prefill role.

The recipes run against a checkout tree containing three not-yet-merged PRs
(SGLANG_USE_CHECKOUT_RUNTIME=1): sgl-project#29168 (unified-KV HiSparse on ROCm), sgl-project#32368
(HiSparse PD over MoRI, based on sgl-project#29168), and sgl-project#36966 (device-alias pointer fix
so both features stop GPU-faulting at PD warmup on non-SVM host kernels). Not
wired into nightly-configs.yaml until they land.

Validation: the DSV4-Pro-FP4 config (deepseek-ai/DeepSeek-V4-Pro) was validated
end-to-end on the hand-driven path earlier (HiSparse decode GSM8K 0.960, HiCache
prefill GSM8K 0.950, both 0 GPU faults). Through the CI launcher, boot + PD KV
transfer (generate 200) + hisparse/mori-live were confirmed; the full GSM8K gate
run under checkout-runtime is pending an image-native run once the dependency PRs
merge (checkout-runtime JIT-compiles aiter on first start, which desyncs
decode/prefill startup and races the bench probe -- an artifact that disappears
with SGLANG_USE_CHECKOUT_RUNTIME=0). Accuracy gate is a placeholder pending that
number.
yuttian1 pushed a commit to yuttian1/sglang that referenced this pull request Sep 1, 2026
Follow-up to sgl-project#30315 (pool sizing) and complementary to sgl-project#29168 (unified-KV
HiSparse device pool). This adds the runtime accounting so scheduling and
allocation treat the unified SWA pool as a fixed per-request ring instead of a
linearly-consumed token pool. On the unified path SWA is addressed by
state_slot (== req_pool_idx) inside the DSV4 kernels; the paged SWA indices /
full_to_swa_index_mapping are never consumed, so accounting SWA as a linear
token pool over-throttles admission and decode retract. The real bound is
concurrency (num_req_slots), already enforced by req_to_token_pool.

All changes are gated on get_kvcache()._unified_kv; the non-unified (fp8) path
is unchanged.

- swa.py: account SWA as a fixed per-request ring slot (swa_ring_cost_tokens)
  bounded by num_req_slots (via req_to_token_pool). available_size /
  swa_available_size / new_pages_available / alloc_extend / alloc_decode take
  the unified branch and skip the vestigial paged SWA allocator + mapping.
- schedule_policy.py: rem_swa_tokens / _swa_budget_for_req use the ring-based
  swa_available_size and charge one fixed ring slot per request; skip the linear
  SWA clamp and the double-charge on chunked continuations.
- schedule_batch.py: unified-KV check_decode_mem diagnostic ([SWA-BOTTLENECK]).
- hisparse.py: forward swa_ring_cost_tokens through the DSV4-HiSparse wrapper
  (the unified HiSparse C4 device pool itself is sgl-project#29168).
- model_runner: wire req_to_token_pool into the SWA allocator.
yuttian1 pushed a commit to yuttian1/sglang that referenced this pull request Sep 1, 2026
Follow-up to sgl-project#30315 (pool sizing) and complementary to sgl-project#29168 (unified-KV
HiSparse device pool). This adds the runtime accounting so scheduling and
allocation treat the unified SWA pool as a fixed per-request ring instead of a
linearly-consumed token pool. On the unified path SWA is addressed by
state_slot (== req_pool_idx) inside the DSV4 kernels; the paged SWA indices /
full_to_swa_index_mapping are never consumed, so accounting SWA as a linear
token pool over-throttles admission and decode retract. The real bound is
concurrency (num_req_slots), already enforced by req_to_token_pool.

All changes are gated on get_kvcache()._unified_kv; the non-unified (fp8) path
is unchanged.

- swa.py: account SWA as a fixed per-request ring slot (swa_ring_cost_tokens)
  bounded by num_req_slots (via req_to_token_pool). available_size /
  swa_available_size / new_pages_available / alloc_extend / alloc_decode take
  the unified branch and skip the vestigial paged SWA allocator + mapping.
- schedule_policy.py: rem_swa_tokens / _swa_budget_for_req use the ring-based
  swa_available_size and charge one fixed ring slot per request; skip the linear
  SWA clamp and the double-charge on chunked continuations.
- schedule_batch.py: unified-KV check_decode_mem diagnostic ([SWA-BOTTLENECK]).
- hisparse.py: forward swa_ring_cost_tokens through the DSV4-HiSparse wrapper
  (the unified HiSparse C4 device pool itself is sgl-project#29168).
- model_runner: wire req_to_token_pool into the SWA allocator.
amd-danli103 and others added 7 commits September 3, 2026 01:56
Co-authored-by: Cursor <cursoragent@cursor.com>
Main now rejects new files under test/registered/amd/ and GPU suites under unit/.
HiSparse aliases bf16 rows; fp8 splits nope/rope and needs its own adapter.
Pin the live C4 out_loc remap that sits next to main's fp8_2buff store.
_init_compressed_pools was overwriting HiSparseUnifiedC4DevicePool with
kv_pools[4] (None on the unified path), so the allocator assert fired
at server start.
Keep unified-KV HiSparse C4 across DSV4.1 indexer paging and encoder
replay. Port the nightly eval job into nightly-test-amd.yml.
@amd-danli103

Copy link
Copy Markdown
Contributor Author

/rerun-failed-ci

Keep unified-KV linear swap beside MiniMax M3 HiSparse; paged FP8
dsv4 kernels stay gated on is_dsv4_paged_layout.
main's NetworkAddress.to_url() already prepends http://.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

amd deepseek memory-pool run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants