Skip to content

[AMD][DI][CI] 10/N Add DSV4 HiSparse+HiCache 1P1D nightly recipes (FP4/FP8) x (PRO/FLASH) - #37200

Open
Lzy17 wants to merge 4 commits into
sgl-project:mainfrom
Lzy17:dsv4pro-hisparse-hicache-dici
Open

Lzy17 wants to merge 4 commits into
sgl-project:mainfrom
Lzy17:dsv4pro-hisparse-hicache-dici

Conversation

@Lzy17

@Lzy17 Lzy17 commented Aug 31, 2026 •

Copy link
Copy Markdown
Collaborator

What

Nightly CI test for DSV4 HiSparse + HiCache in PD over MoRI. One recipe per model runs HiCache on prefill and
HiSparse on decode (they don't conflict — HiCache is prefill-only, HiSparse decode-only). Covers all four
DSV4 variants: Pro/Flash × FP8/FP4.

Also:

  • Two launcher fixes the recipes need: apply a recipe's *_extra_env after the DSV4 env so it can override
    it; add a disable_radix_cache knob (--disable-radix-cache was hardcoded but HiCache needs radix on).
  • Decode uses --disable-overlap-schedule — the overlap scheduler intermittently desyncs the decode TP
    request broadcast (broadcast_pyobj), crashing the decode scheduler. Disabling it holds.

Depends on

Four things must be in the image, in order: #29168 → #32368 → #36966 (sglang), plus #37286
(MoRI bump for the ionic RoCE fix). Not wired into nightly-configs.yaml until they land.

Validation

DSV4-Pro-FP4, 2-node 1P1D, manual run (image-native sglang): HiSparse (decode) GSM8K 0.960, HiCache
(prefill) 0.950, 0 GPU faults.

CI parity note: run these with SGLANG_USE_CHECKOUT_RUNTIME=0 (image-native) once the deps are in the image.
Under checkout-runtime (SGLANG_USE_CHECKOUT_RUNTIME=1) each start JIT-compiles aiter (~20 min), which desyncs
the 8 ranks' startup and is not representative — that path is only for pre-merge testing.


CI States

Latest PR Test (Base): ❌ Run #34562375725
Latest PR Test (Extra): ❌ Run #34562375607
Latest PR Test (AMD ROCm 10): ❌ Run #34562375734

@github-actions github-actions Bot added the hicache Hierarchical Caching for SGLang label Aug 31, 2026
@Lzy17
Lzy17 force-pushed the dsv4pro-hisparse-hicache-dici branch from 3fa4246 to 0aef73a Compare August 31, 2026 09:17
@Lzy17 Lzy17 changed the title [AMD][DI][CI] Add DSV4-Pro-FP8 HiSparse+HiCache 1P1D nightly recipe [AMD][DI][CI] Add DSV4 HiSparse+HiCache 1P1D nightly recipes (4 variants) Aug 31, 2026
@Lzy17 Lzy17 changed the title [AMD][DI][CI] Add DSV4 HiSparse+HiCache 1P1D nightly recipes (4 variants) [AMD][DI][CI] 10/N Add DSV4 HiSparse+HiCache 1P1D nightly recipes (FP4/FP8) x (PRO/FLASH) Aug 31, 2026
…nts)

New MI355X 2-node 1P1D leg exercising both memory features in one PD run:
HiCache on the prefill role, HiSparse on the decode role, KV transfer over MoRI,
unified-KV layout. The two features live on opposite PD roles and do not
conflict (HiCache is prefill-only, HiSparse is decode-only), so a single recipe
per model covers both. Added for all four DSV4 variants:
  mi355x-fp8/dsv4pro,  mi355x-fp8/dsv4flash,
  mi355x-fp4/dsv4pro,  mi355x-fp4/dsv4flash

Two launcher changes the recipes depend on:
- Env ordering: a recipe's prefill/decode_extra_env is now applied after the
  hardcoded DSV4 env so it can override it (needed to pin
  SGLANG_HACK_FLASHMLA_BACKEND=unified_kv_triton and the MoRI host-registration
  knobs).
- disable_radix_cache knob: --disable-radix-cache was hardcoded into all four
  COMMON_FLAGS strings; HiCache L1 *is* the radix cache and the combination is a
  hard ValueError. The recipe field disable_radix_cache: false now drives a
  RADIX_FLAG (defined before the role if/else so both branches see it under
  set -u), so HiCache can run on the prefill role.

The recipes run against a checkout tree containing three not-yet-merged PRs
(SGLANG_USE_CHECKOUT_RUNTIME=1): sgl-project#29168 (unified-KV HiSparse on ROCm), sgl-project#32368
(HiSparse PD over MoRI, based on sgl-project#29168), and sgl-project#36966 (device-alias pointer fix
so both features stop GPU-faulting at PD warmup on non-SVM host kernels). Not
wired into nightly-configs.yaml until they land.

Validation: the DSV4-Pro-FP4 config (deepseek-ai/DeepSeek-V4-Pro) was validated
end-to-end on the hand-driven path earlier (HiSparse decode GSM8K 0.960, HiCache
prefill GSM8K 0.950, both 0 GPU faults). Through the CI launcher, boot + PD KV
transfer (generate 200) + hisparse/mori-live were confirmed; the full GSM8K gate
run under checkout-runtime is pending an image-native run once the dependency PRs
merge (checkout-runtime JIT-compiles aiter on first start, which desyncs
decode/prefill startup and races the bench probe -- an artifact that disappears
with SGLANG_USE_CHECKOUT_RUNTIME=0). Accuracy gate is a placeholder pending that
number.
@Lzy17
Lzy17 force-pushed the dsv4pro-hisparse-hicache-dici branch from 0aef73a to 20e6dd5 Compare August 31, 2026 23:27
… long context

The 256MiB host-register chunk faults on long requests: when a KV entry
spans more than one chunk, the HiSparse transfer kernel addresses from a
single contiguous host base and reads past the chunk. Short requests never
cross a chunk, so they pass. A 3.5GiB chunk keeps the KV in one chunk and
stays under the ionic ~4GB MR limit. Validated DSV4-Pro 1P1D over MoRI:
95k and 200k-char requests pass, 0 GPU faults, GSM8K 0.955.
The decode host pool has to sit on 2MiB pages. One ionic pd holds about 1M 4K
page-table entries, roughly 4 GiB; the pool is larger, so ibv_reg_mr fails with
errno:22 and the decode scheduler aborts. On 2MiB pages the same pool registers
clean, 100 GiB on one pd against 3.75 GiB on 4K.

Flash needs three settings Pro does not, one per failure it hit: a capped decode
token budget, since the pool is the budget over four and Flash budgets 26.2M; a
lower prefill HiCache ratio, since at 2 the c4 indexer pool is left 3.67 GB; and
a lower prefill memory fraction, since at 0.85 prefill runs out of VRAM and the
first lazily loaded Triton kernel has nowhere to go.

Measured, 2-node 1P1D on mia1-p02-g29 + g53, each on its own checkpoint:
fp4/dsv4pro 0.950, fp8/dsv4pro 0.945, fp4/dsv4flash 0.950, fp8/dsv4flash 0.950.
@Lzy17
Lzy17 force-pushed the dsv4pro-hisparse-hicache-dici branch from ab38f90 to bf4f861 Compare September 11, 2026 04:29
@Lzy17

Lzy17 commented Sep 11, 2026 •

Copy link
Copy Markdown
Collaborator Author

Ran all four recipes on hardware: 2-node 1P1D, MI355X, mia1-p02-g29 prefill + g53 decode, each on the checkpoint its entry names.

recipe GSM8K invalid ibv_reg_mr failures
mi355x-fp4/dsv4pro 0.950 0.000 0
mi355x-fp8/dsv4pro 0.945 0.005 0
mi355x-fp4/dsv4flash 0.950 0.000 0
mi355x-fp8/dsv4flash 0.950 0.000 0

The stack under test was #29168 → #32368 plus a host-pointer fix. We also built #32368 + #35233 on its own, with a rebuilt gfx950 sgl_kernel and no #36966 in the tree, and pro-fp4 scored 0.970 there — so #35233 covers what these recipes need, and #36966 can be closed.

What the recipes now carry, measured rather than guessed:

  • decode host pool on 2MiB pages. One ionic pd holds about 1M 4K page-table entries, roughly 4 GiB. The pool is larger, so ibv_reg_mr fails with errno:22 partway through and the decode scheduler aborts. On 2MiB pages the same pool registers clean: 100 GiB on one pd against 3.75 GiB on 4K.
  • Flash decode capped at the Pro token budget. The host pool is the budget over four, and Flash budgets 26.2M against Pro's 12.3M, so an uncapped decode wants 174 GB per rank, 1.4 TB over eight.
  • Flash prefill at --hicache-ratio 1.2. At 2 the c4 indexer pool is left 3.67 GB and prefill dies. --hicache-size is not a way out; the DSV4 path rejects it.
  • Flash prefill at mem_fraction_static 0.78. At 0.85 it is left about 0.5 GiB of VRAM and the first lazily loaded Triton kernel has no room. hipModuleLoadData then reports hipErrorNoBinaryForGpu, which is about space, not the arch.

The recipes are the run command — full flags and env for both roles — at scripts/ci/slurm/recipes/mi355x-{fp4,fp8}/dsv4{pro,flash}/1k1k/1p1d-hisparse-hicache.yaml in this PR.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hicache Hierarchical Caching for SGLang

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant