Skip to content

[AMD] Enable HiSparse on ROCm - #26639

Merged
HaiShaw merged 9 commits into
sgl-project:mainfrom
clintg6:feat/hisparse
Jun 19, 2026
Merged

HaiShaw merged 9 commits into
sgl-project:mainfrom
clintg6:feat/hisparse

Conversation

@clintg6

@clintg6 clintg6 commented May 29, 2026 •

Copy link
Copy Markdown
Contributor

Enable HiSparse for DSA models on ROCm

Summary

This change enables HiSparse for DeepSeek-style DSA models on ROCm, validated on MI355X, while keeping CUDA behavior unchanged. ROCm uses TileLang as the default HiSparse DSA backend, supports FP8 and BF16 KV cache, and can also run with AITER when the user explicitly selects the AITER DSA backend.

The main work is in three areas: making the shared HiSparse swap-in kernel correct for AMD wavefront64, adding the ROCm-specific HiSparse allocator/coordinator paths, and wiring ROCm DSA backend policy so TileLang and AITER use the KV-cache layout their kernels expect.

Why This Is Needed

HiSparse keeps a hot working set of KV cache on GPU and stores the rest in host memory. The existing implementation was built around CUDA assumptions: warp32 behavior, CUDA inline PTX copies, and allocator flows that did not cover ROCm’s page_size == 1 path.

On ROCm these assumptions can break swap-in correctness and leak temporary device mappings during decode. For DSA models, ROCm also needs a backend choice that matches the KV-cache layout. TileLang is currently the fastest validated option, so it remains the default. AITER is supported as an explicit opt-in backend for comparison and debugging.

Backend Behavior

With --enable-hisparse on ROCm, the default DSA backend is tilelang for both prefill and decode.

If the user explicitly selects:

--dsa-prefill-backend aiter --dsa-decode-backend aiter

AITER is accepted on ROCm. With BF16 KV cache, it uses the normal AITER sparse path. With FP8 KV cache, it automatically uses AITER’s BF16-Q/FP8-KV persistent sparse MLA metadata path. No extra environment flag is required.

Both TileLang and AITER keep ROCm HiSparse FP8 KV cache in the raw MLA layout:

nope 512 fp8 + rope 64 fp8 = 576

This avoids the CUDA scaled FP8 layout, which is 656 bytes per token and is incompatible with the ROCm sparse MLA views used here.

CUDA Impact

This PR is intended to be behaviorally unchanged on CUDA.

The wavefront64 kernel fixes are guarded with USE_ROCM, and the CUDA branch keeps the original assumptions: warp32 masks, CUDA inline PTX copy, and the prior LRU writeback loop. The coordinator cleanup for stale temporary HiSparse mappings is also ROCm-gated because CUDA’s swap-in path consumes only top_k_device_locs, where those stale mapping entries are harmless.

CUDA HiSparse backend selection remains dtype-specific:

bfloat16 KV -> flashmla_sparse
fp8_e4m3 KV -> flashmla_kv

Implementation

The swap-in JIT kernel now has ROCm-safe wavefront handling in python/sglang/jit_kernel/csrc/hisparse.cuh. The code uses a platform-dependent WARP_SIZE, ballot mask type, full-warp mask, and popcount helper. ROCm gets a HIP byte-copy fallback instead of CUDA inline PTX. The inclusive-scan window and large LRU writeback path are fixed under USE_ROCM so wavefront64 lanes do not write outside the current iteration’s intended shared-memory window.

The HiSparse allocator now supports ROCm’s page_size == 1 path. HiSparseTokenToKVPoolAllocator.alloc() allocates both logical and HiSparse device indices and rolls back the logical allocation if the device allocation fails. alloc_extend() no longer asserts page_size > 1, allowing ROCm decode/extend flows to use the normal allocation path.

The HiSparse coordinator now frees stale temporary device mappings during ROCm decode remap before replacing them with the reserved device-buffer slot. Without this cleanup, temporary page-size-1 mappings can leak and later corrupt swap-in lookups.

The model-runner KV-cache sizing keeps ROCm TileLang and AITER in the raw 576-wide MLA layout for FP8 KV cache. This is required for both backends and prevents the AITER sparse MLA view failure caused by the CUDA scaled layout.

The AITER DSA path now handles FP8 KV cache without user environment flags. For explicit AITER + FP8 KV, both decode and extend build AITER persistent sparse MLA metadata and pass BF16 Q with FP8 KV into mla_decode_fwd.

Tests And Validation

The packaged changes add focused unit and regression coverage:

  • HiSparse JIT tests are registered in AMD CI.
  • The swap-in JIT kernel has a multi-miss copy regression test.
  • ROCm has a large LRU writeback regression test covering the wavefront64 scan-window fix.
  • HiSparse manager unit tests are registered in AMD CI and cover the page_size == 1 allocator path and decode-remap cleanup.
  • Server-args tests cover CUDA/ROCm backend defaults, explicit AITER acceptance on ROCm, raw MLA KV-cache sizing for ROCm AITER FP8, invalid platform/backend combinations, and accepted/rejected HiSparse KV dtypes.
  • The GLM-5 HiSparse E2E test is registered for AMD CI and uses TileLang on ROCm.

Local checks run on the packaged files:

python3 -m ruff check <modified python files>
python3 -m black --check <modified python files>
python3 -m isort --check-only <modified python files>
clang-format --style=file --dry-run python/sglang/jit_kernel/csrc/hisparse.cuh

All pass.

Runtime validation on MI355X showed:

TileLang FP8 KV:
Accuracy 0.968, latency 531.376s, output throughput 80.642 token/s

AITER BF16-Q/FP8-KV persistent path:
Accuracy 0.968, latency 666.757s, output throughput 63.344 token/s

TileLang is therefore kept as the default ROCm HiSparse backend. AITER remains useful for explicit comparison and debugging, and it no longer requires a separate environment flag for FP8 KV cache.

Recommended ROCm Starting Point

For MI300X-class GPUs, the validated GLM-5 HiSparse configuration starts from:

nohup python3 -m sglang.launch_server \
  --model-path /models/GLM-5.1/ \
  --tp 8 \
  --port 8000 \
  --model-loader-extra-config '{"enable_multithread_load": true, "num_threads": 8}' \
  --trust-remote-code \
  --tool-call-parser glm47 \
  --reasoning-parser glm45 \
  --mem-fraction-static 0.65 \
  --dsa-prefill-backend tilelang \
  --dsa-decode-backend tilelang \
  --kv-cache-dtype fp8_e4m3 \
  --max-running-requests 2 \
  --watchdog-timeout 1200 \
  --skip-server-warmup \
  --enable-hisparse \
  --hisparse-config '{"top_k": 2048, "device_buffer_size": 2048, "host_to_device_ratio": 1}' \
  --disable-radix-cache \
  &

These are test/serving recommendations, not forced runtime defaults.

Notes

TileLang remains the recommended ROCm HiSparse backend because it is faster in current testing. AITER is accepted when explicitly selected and uses BF16-Q/FP8-KV persistent sparse MLA automatically for FP8 KV cache.


CI States

Latest PR Test (Base): ❌ Run #27799270152
Latest PR Test (Extra): ❌ Run #27799269960

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@HaiShaw

HaiShaw commented May 29, 2026

Copy link
Copy Markdown
Collaborator

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label May 29, 2026
@clintg6

clintg6 commented May 29, 2026

Copy link
Copy Markdown
Contributor Author

GSM8k also shows 95% accuracy on MI300X for both attention backends. @HaiShaw

@HaiShaw

HaiShaw commented Jun 3, 2026

Copy link
Copy Markdown
Collaborator

cc @Jacob0226 @Lzy17

@amd-danli103

Copy link
Copy Markdown
Contributor

hi @clintg6 , you mentioned "Both TileLang and AITER keep ROCm HiSparse FP8 KV cache in the raw MLA layout: nope 512 fp8 + rope 64 fp8 = 576", so this is for supporting DPSK-V3.2, currently NOT for V4, right?

@clintg6

clintg6 commented Jun 3, 2026

Copy link
Copy Markdown
Contributor Author

This PR enables HiSparse on ROCm for DSA models (DeepSeek V3.2, GLM-5). DeepSeek V4 HiSparse support for ROCm would be a follow-up PR once the DSv4 ROCm integration lands on main. @HaiShaw @amd-danli103

@HaiShaw

HaiShaw commented Jun 4, 2026

Copy link
Copy Markdown
Collaborator

@amd-bot ci-status

@amd-bot

amd-bot commented Jun 4, 2026

Copy link
Copy Markdown

@HaiShaw

CI Status for PR #26639

PR: [AMD] Enable HiSparse on ROCm
Changed files: python/sglang/jit_kernel/csrc/hisparse.cuh (+89/-14), python/sglang/jit_kernel/tests/test_hisparse.py (+173/-2), python/sglang/srt/arg_groups/hisparse_hook.py (+52/-22), python/sglang/srt/layers/attention/dsa_backend.py (+209/-8), python/sglang/srt/managers/hisparse_coordinator.py (+23/-10), python/sglang/srt/managers/scheduler.py (+3/-2), python/sglang/srt/mem_cache/hisparse_memory_pool.py (+17/-6), python/sglang/srt/model_executor/model_runner_kv_cache_mixin.py (+6/-5), python/sglang/srt/server_args.py (+10/-0), and 4 test files.

AMD: 2 failures (1 likely related) | Others: 1 failure (0 related)

(The pr-test-amd-finish and pr-test-finish jobs are gate aggregators that fail because of the underlying job failures below — not separate failures.)

AMD CI Failures

Job Test File Test Function Error Related? Explanation Log
stage-b-test-1-gpu-small-amd (7) test/registered/unit/managers/test_hisparse_unit.py TestHiSparseUnit.test_single_node_staging_allocates_paged_host_slots AssertionError: 4093 != 4094 🔴 Likely Directly tests HiSparse staging/paged-host-slot allocation — exactly the code path this PR modifies in hisparse_memory_pool.py, hisparse_coordinator.py, and hisparse_hook.py. The test file itself is also modified by this PR (+84/-2). Log
stage-b-test-1-gpu-small-amd (7) test/registered/perf/test_vlm_perf_5090.py test_vlm_offline_throughput AssertionError: 1120.8 not greater than 2000 🟢 Unlikely VLM offline throughput perf test on MI325; unrelated to HiSparse code paths. Looks like a perf/threshold flake on the AMD runner. Log
stage-b-test-1-gpu-small-amd-nondeterministic test/registered/models/test_vlm_models.py TestVLMModels.test_vlm_mmmu_benchmark (openbmb/MiniCPM-V-2_6) FileNotFoundError: No JSON result files found in /tmp/test_vlm_mmmu_openbmb_MiniCPM-V-2_6_* (eval crashed with IndexError: list index out of range) 🟢 Unlikely MMMU VLM benchmark for MiniCPM-V; the underlying lm-eval crash and missing JSON output are unrelated to HiSparse. Appears to be a pre-existing AMD VLM eval issue. Log

Other CI Failures

Job Test File Test Function Error Related? Explanation Log
base-c-test-8-gpu-h20 (1) test/registered/disaggregation/test_disaggregation_different_tp.py test_gsm8k AssertionError: 0.0 not greater than 0.6 (KV transfer errors: Failed to get kvcache from prefill instance, it might be dead) 🟢 Unlikely PD-disaggregation gsm8k test on NVIDIA H20. PR touches no disaggregation or KV-transfer code; score 0.0 is an infrastructure/PD-stability issue, not caused by this PR. Log

Details

🔴 test_single_node_staging_allocates_paged_host_slots (AMD) — strongest signal that this PR needs a fix. The test asserts host slots are allocated at page granularity and got 4093 != 4094 (off by one page). The PR modifies HiSparse memory-pool allocation (hisparse_memory_pool.py +17/-6), the coordinator (hisparse_coordinator.py +23/-10), and the hook (hisparse_hook.py +52/-22), and also edits this very test file (+84/-2). Suspect the staging-allocation accounting or rounding was changed slightly and the assertion expectations didn't keep up. Recommend re-checking the updated allocation logic and the new test fixtures.

🟢 The two AMD VLM failures and the H20 disaggregation failure are in code paths completely untouched by this PR and look like pre-existing flakes / infra issues. Safe to ignore from this PR's perspective, but worth a retry.

Generated by amd-bot using Claude Code CLI

@amd-danli103

Copy link
Copy Markdown
Contributor

This PR enables HiSparse on ROCm for DSA models (DeepSeek V3.2, GLM-5). DeepSeek V4 HiSparse support for ROCm would be a follow-up PR once the DSv4 ROCm integration lands on main. @HaiShaw @amd-danli103

hi @clintg6 glad to know that. Currently most of the V4 optimization jobs have been merged into main as far as I know, except for several ongoing ones. May I know if there will be a schedule/plan for your follow-up PR to support V4? I would like to avoid duplicate work if your next PR is already underway, thank you!

@HaiShaw

HaiShaw commented Jun 14, 2026

Copy link
Copy Markdown
Collaborator

@amd-bot ci-status

@amd-bot

amd-bot commented Jun 14, 2026

Copy link
Copy Markdown

@HaiShaw

CI Status for PR #26639 — [AMD] Enable HiSparse on ROCm

Merge verdict: ⚠️ Do not treat as verified yet. No PR-CI failure is clearly caused by this PR's runtime logic — the large red wall is one infra root cause (gated-HF 401) cascading through NVIDIA, plus an NPU perf-threshold flake. But the two things that matter are both red flags: (1) this PR's core ROCm/HiSparse code was not exercised by any PR-CI job (all AMD workflows startup_failure'd — a repo-wide infra outage, not your fault, but it means zero AMD signal), and (2) one CPU failure is in test_server_args.py, a file this PR modified, and deserves a look before merge.

Caution

This PR's main value (ROCm HiSparse: wavefront64 swap-in kernel in hisparse.cuh, ROCm allocator/coordinator paths in dsa_backend.py/hisparse.py, ROCm DSA backend policy) is NOT exercised by any PR-CI test that ran.

  • Both AMD workflows — PR Test (AMD) and PR Test ROCm 7.2 (AMD) — ended in startup_failure with 0 jobs launched. This is a pre-existing infra outage (the last ~15 consecutive pr-test-amd.yml runs across unrelated branches all startup_failure), NOT caused by this PR — but the practical effect is the same: the AMD-registered HiSparse tests (test/registered/jit/test_hisparse.py, test_hisparse_unit.py on stage-b-test-1-gpu-small-amd) never ran.
  • The two NEW AMD accuracy tests (test_glm51_hisparse_eval_mi30x.py, ..._mi35x.py) are registered nightly=True → they do not run on PR CI at all, by design.
  • What DID verify on CUDA/CPU: test_hisparse_unit.py ✅ passed (NVIDIA base-b 1-gpu-small shard 2), and the new TestHiSparseDsaBackendPolicy cases in test_server_args.py ✅ passed (CPU, with is_hip mocked — so they exercise policy branching only, not real HIP execution).
  • Before merge, the author must run the ROCm path on real AMD hardware: the stage-b-test-1-gpu-small-amd HiSparse suite + the glm51-hisparse nightly accuracy suites, since CI cannot currently do it.

Changed files (13): hisparse.cuh (+104/-15), dsa_backend.py (+209/-8), hisparse_coordinator.py, hisparse.py (allocator), model_runner_kv_cache_mixin.py, scheduler.py, server_args.py, hisparse_hook.py, +5 test files.

AMD: startup_failure (0 jobs ran, infra-wide) · Others: ~18 failures (0 clearly PR-caused; 1 needs author check)

AMD CI

Workflow Result Related? Why
PR Test (AMD) startup_failure — 0 jobs 🟢 No Repo-wide infra; last ~15 pr-test-amd.yml runs on unrelated branches all startup_failure. PR touches no workflow files.
PR Test ROCm 7.2 (AMD) startup_failure — 0 jobs 🟢 No Same infra outage.
PR Test Extra (AMD) gate/finish failed; extra-a-test-1-gpu-small-amd skipped 🟢 No Test job skipped (gate cascade); no AMD test actually executed.

Other CI Failures

Job Test File Test Function Error Related? Why
base-b-test-1-gpu-small (1) (root cause) test/registered/sampling/test_original_logprobs.py model load GatedRepoError: 401 on meta-llama/Llama-3.2-1B-Instruct → exit -9 🟢 No HF auth/infra. PR doesn't touch model loading, sampling, or HF auth. ~15 other base-b failures + wait-for-base-b + pr-test-finish are fast-fail cascades of this one.
build-test (xeon-gnr, base-b-test-cpu) test/registered/unit/server_args/test_server_args.py TestCudaGraphConfigDataclassAccess.test_tc_piecewise_build_config_reads_phase_config_dataclass AttributeError: module 'sglang.srt.model_executor' has no attribute 'runner_backend' (log ~line 2038) 🟡 Possibly This PR modified test_server_args.py (added top-level import torch + from ...model_runner_kv_cache_mixin import ModelRunnerKVCacheMixin). The failing test is pre-existing (not added by you) and relies on runner_backend being pre-imported when @patch resolves it via pkgutil.resolve_name. New imports can perturb import ordering. The PR's own TestHiSparseDsaBackendPolicy tests in the same file passed.
stage-b-test-1-npu-a2 (0) test/registered/ascend/basic_function/quant/test_npu_w8a8_quantization.py test_gsm8k AssertionError: throughput 503 < 700 tok/s 🟢 No NPU/Ascend perf-threshold flake; PR touches no NPU/Ascend quant code.

(pr-test-finish, pr-test-npu-finish, pr-test-extra-finish, call-gate / pr-gate rows are aggregation/cascade jobs — not independent failures.)

Details / what to do before merge

  1. Coverage (highest priority): AMD PR CI gave zero signal on the ROCm HiSparse runtime. Manually run, on AMD hardware, the stage-b-test-1-gpu-small-amd HiSparse suite (test_hisparse.py + test_hisparse_unit.py AMD registration) and the nightly-amd-accuracy-8-gpu-glm51-hisparse / nightly-amd-8-gpu-mi35x-glm51-hisparse suites. The CUDA/CPU green results only cover the manager unit logic and the is_hip-mocked policy branching — not the wavefront64 swap-in kernel or real ROCm allocator/coordinator paths. Re-running AMD CI is also blocked until the repo-wide pr-test-amd.yml startup_failure is resolved.
  2. 🟡 test_server_args.py regression check: Confirm whether the TestCudaGraphConfigDataclassAccess::test_tc_piecewise... failure reproduces on main's xeon CPU run. If it only fails on this branch, the new top-level imports changed import ordering — a robust fix is to add an explicit import sglang.srt.model_executor.runner_backend.tc_piecewise_cuda_graph_backend (so the @patch target resolves) rather than relying on incidental import side-effects.
  3. No action needed for the base-b (HF 401) cascade and the NPU perf flake — both are infra/flake, unrelated to this PR; re-run once HF auth on the runners is restored.

Generated by amd-bot using Claude Code CLI

@amd-danli103

Copy link
Copy Markdown
Contributor

@clintg6 Nice work! Curious doe Hi-Sparse with DSV4, have we tried with different: SGLANG_HACK_FLASHMLA_BACKEND = triton | unified_kv_triton

hi Hai @HaiShaw, currently this PR supports the original separate packed KV memory layout, not the unified memory layout, where the compressed region lives in the unified pool (c4_kv_pool is None in that path).
Supporting the latter one requires further adapting the new memory pool/coordinator(device hot/cold pools, write/read remap into the unified compressed region) and I think that is worth a separate follow-up PR.

I'm tracking this feature since we have a high-priority use case for customer-T. Now I already persuade them to shift to use the unified KV attention, so I'm actually developing the HiSparse feature with unified KV on top of this PR.
So I'm happy to take that follow-up. Or @clintg6, I would love to coordinate if you're looking at unified_kv_triton too, so we don't duplicate effort. Thank you!

Add a temporary guard for DSv4 HiSparse on unified-KV path.
@clintg6

clintg6 commented Jun 15, 2026

Copy link
Copy Markdown
Contributor Author

@amd-danli103 @HaiShaw Added a temporary guard in validate_hisparse: HiSparse + unified_kv_triton on DSv4 now errors at startup instead of crashing with a cryptic AssertionError. Scoped to just that combo and doesn't affect NV HiSparse path. @amd-danli103 please remove it when your unified-KV HiSparse support lands.

@amd-danli103

Copy link
Copy Markdown
Contributor

@amd-danli103 @HaiShaw Added a temporary guard in validate_hisparse: HiSparse + unified_kv_triton on DSv4 now errors at startup instead of crashing with a cryptic AssertionError. Scoped to just that combo and doesn't affect NV HiSparse path. @amd-danli103 please remove it when your unified-KV HiSparse support lands.

Thank you @clintg6 , the guard makes sence. Thanks for letting me know that!

@HaiShaw

HaiShaw commented Jun 17, 2026

Copy link
Copy Markdown
Collaborator

@amd-bot ci-status

@amd-bot

amd-bot commented Jun 17, 2026

Copy link
Copy Markdown

@HaiShaw

CI Status for PR #26639

Merge verdict: ❌ Not ready to merge. PR CI is incomplete (fast-fail cascades skipped downstream NVIDIA jobs; one AMD job dsv4-pro-fp4-amd-rocm720 is still queued). Of the executed failures, almost all are pre-existing / unrelated (AMD HiCache crash already on main, NVIDIA concurrency flakes, NPU perf threshold, XPU infra) — but one is likely caused by this PR: the CPU test_server_args.py failure. More importantly, the core HiSparse-on-ROCm feature is not exercised by PR CI at all (its e2e tests are nightly=True).

Caution

This PR's headline functionality is effectively untested by PR CI. The two new end-to-end accuracy tests (test_glm51_hisparse_eval_mi35x.py, mi30x) are registered nightly=True → they do not run on PR CI. The PR's AMD unit/jit tests (test_hisparse_unit.py, test_hisparse.py) are registered to stage-b-test-1-gpu-small-amd, but both of those shards (6, 13) are RED from a pre-existing, unrelated HiCache GPU crash, so AMD-side unit verification of the new HiSparse code is obscured. CUDA base-b ran the unit/jit tests but does not exercise the ROCm code paths this PR adds. Before merge, run the HiSparse nightly accuracy suite on MI35x/MI30x and confirm the AMD unit/jit tests pass on a clean shard.

Changed files (13): dsa_backend.py (+209/-8), hisparse.cuh (+104/-15), test_server_args.py (+139/-0), hisparse_hook.py, hisparse_coordinator.py, allocator/hisparse.py, model_runner_kv_cache_mixin.py, scheduler.py, server_args.py, 2 new nightly accuracy tests, test_hisparse.py, test_hisparse_unit.py.

Executed CI failure attribution: AMD: 3 executed failures (0 related) + 1 queued · Others: ~10 failures across NVIDIA/NPU/XPU/CPU (1 related — CPU). Fast-fail cascade jobs collapsed into their root cause; skipped downstream jobs counted as completeness gaps, not failures.

AMD Executed Failures

Job Test File Test Function Error Related? Why
stage-b-1gpu-small (13) test/registered/hicache/test_hicache_storage.py TestHiCache.test_mmlu RuntimeError: Destination indices must be a CUDA tensor 🟢 Pre-existing on main (baseline job, same test+error). Path = generic HiCache (memory_pool_host.py→transfer_kv_all_layer_lf_pf), not touched by PR.
stage-b-1gpu-small (6) test/registered/hicache/test_hicache_variants.py TestHiCacheMLA.setUpClass RuntimeError: Destination indices must be a CUDA tensor 🟢 Pre-existing on main (baseline, same).
stage-b-2gpu-large (1) test/registered/hicache/test_hicache_storage_file_backend.py TestHiCacheStorageAccuracy.test_basic_backup_and_prefetch RuntimeError: Destination indices must be a CUDA tensor 🟢 Pre-existing on main (baseline, same).

Note: these 3 pre-existing HiCache crashes (SIGQUIT) take down the entire AMD stage-b-small shards (0/11 passed), which is exactly where this PR's HiSparse unit/jit tests live — so they obscure the PR's AMD-side coverage.

Other Executed Failures

Job Test File Test Function Error Related? Why
build-test (xeon-gnr, base-b-test-cpu) test/registered/unit/server_args/test_server_args.py TestCudaGraphConfigDataclassAccess.test_tc_piecewise_build_config_reads_phase_config_dataclass AttributeError: module 'sglang.srt.model_executor' has no attribute 'runner_backend' 🔴 PR-modified file. Passes on main CPU (base-a-test-cpu ✅). PR adds top-level import torch + from ...model_runner_kv_cache_mixin import ModelRunnerKVCacheMixin, eagerly importing the model_executor package so a pre-existing test's @patch("...model_executor.runner_backend...") can no longer resolve the submodule. [MEDIUM]
base-b-1gpu-large (4) test/registered/kernels/test_mhc_kernels.py test_mhc_fused_post_pre_matches_unfused[False-1-4096] ValueError: Global server args is not set yet! 🟢 Crash in mhc.py:_prewarm_mhc_pre (not touched by PR); test-isolation issue. Root cause of the base-b fast-fail cascade.
base-b-1gpu-small (6) test/registered/core/test_srt_endpoint.py TestSRTEndpoint.test_get_server_info_concurrent AssertionError/JSONDecodeError (concurrent race in communicator.py) 🟢 Concurrency race in communicator.py (untouched); flaky. Co-root-cause of cascade.
stage-b-1-npu-a2 (0) NPU throughput bench Qwen2.5-0.5B-w8a8 AssertionError: 653.8 not >= 700 (perf threshold) 🟢 NPU backend perf flake, unrelated to ROCm.
multimodal-gen-1/2-npu-a3 NPU multimodal N/A NPU job failure 🟢 Different backend, not touched.
stage-a-1-gpu-xpu test/registered/xpu/test_xpu_basic.py TestXPUBasic.test_basic_generation ModuleNotFoundError: tvm_ffi / benchmark parse fail 🟢 XPU infra/env issue, unrelated.

Fast-fail cascade jobs (not independent failures): base-b small(7), 2gpu-large(2)/(3) — all skipped via Fast-fail: skipping — root cause job(s): wait-for-base-b, base-b-test-1-gpu-large (4), base-b-test-1-gpu-small (6).

Details / what to do before merge

  1. 🔴 Fix the CPU regression in your own test file. test/registered/unit/server_args/test_server_args.py now fails in the base-b-test-cpu suite because the new top-level imports (import torch, model_runner_kv_cache_mixin) eagerly load sglang.srt.model_executor, breaking the pre-existing test_tc_piecewise_build_config_reads_phase_config_dataclass mock-patch resolution (error even hints Did you mean: 'runner_backend_utils'?). Reproduce locally with the CPU suite; consider moving the new imports inside the test methods/class, or import sglang.srt.model_executor.runner_backend explicitly. [MEDIUM] — verify by running the base-b-test-cpu suite on this branch.

  2. 🟡 Coverage — verify HiSparse actually works on ROCm. The nightly-only accuracy tests (nightly-amd-8-gpu-mi35x-glm51-hisparse, nightly-amd-accuracy-8-gpu-glm51-hisparse) and the AMD unit/jit tests (currently blocked behind a pre-existing HiCache crash in stage-b-1gpu-small-amd) must pass before merge. Run the nightly HiSparse suites on MI35x/MI30x and confirm test_hisparse_unit.py / test_hisparse.py pass on a clean AMD shard. Green PR CI does not verify this feature.

  3. CI completeness. Re-run after the unrelated base-b root causes (mhc / srt_endpoint) clear so downstream NVIDIA stages aren't fast-fail-skipped, and wait for the queued dsv4-pro-fp4-amd-rocm720 AMD job. For a full-signal run despite unrelated flakes, the bypass-fastfail label is an option (uses more CI resources — use sparingly).

  4. The AMD HiCache, NVIDIA mhc/srt_endpoint, NPU perf, and XPU tvm_ffi failures are not caused by this PR — no action needed from the author beyond a rerun.


Generated by amd-bot using Claude Code CLI

@HaiShaw

HaiShaw commented Jun 19, 2026

Copy link
Copy Markdown
Collaborator

1test/registered/unit/server_args/test_server_args.py` was fixed (but CI [build-test (xeon-gnr, base-b-test-cpu) is stuck).

@HaiShaw
HaiShaw merged commit c436a81 into sgl-project:main Jun 19, 2026
140 of 170 checks passed
Chronostasys pushed a commit to MindLab-Research/sglang that referenced this pull request Aug 24, 2026
Co-authored-by: clintg6 <7388379+clintg6@users.noreply.github.com>
Co-authored-by: HAI <hixiao@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

jit-kernel run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants