Conversation
Lzy17
force-pushed
the
dsv4pro-hisparse-hicache-dici
branch
from
August 31, 2026 09:17
3fa4246 to
0aef73a
Compare
…nts) New MI355X 2-node 1P1D leg exercising both memory features in one PD run: HiCache on the prefill role, HiSparse on the decode role, KV transfer over MoRI, unified-KV layout. The two features live on opposite PD roles and do not conflict (HiCache is prefill-only, HiSparse is decode-only), so a single recipe per model covers both. Added for all four DSV4 variants: mi355x-fp8/dsv4pro, mi355x-fp8/dsv4flash, mi355x-fp4/dsv4pro, mi355x-fp4/dsv4flash Two launcher changes the recipes depend on: - Env ordering: a recipe's prefill/decode_extra_env is now applied after the hardcoded DSV4 env so it can override it (needed to pin SGLANG_HACK_FLASHMLA_BACKEND=unified_kv_triton and the MoRI host-registration knobs). - disable_radix_cache knob: --disable-radix-cache was hardcoded into all four COMMON_FLAGS strings; HiCache L1 *is* the radix cache and the combination is a hard ValueError. The recipe field disable_radix_cache: false now drives a RADIX_FLAG (defined before the role if/else so both branches see it under set -u), so HiCache can run on the prefill role. The recipes run against a checkout tree containing three not-yet-merged PRs (SGLANG_USE_CHECKOUT_RUNTIME=1): sgl-project#29168 (unified-KV HiSparse on ROCm), sgl-project#32368 (HiSparse PD over MoRI, based on sgl-project#29168), and sgl-project#36966 (device-alias pointer fix so both features stop GPU-faulting at PD warmup on non-SVM host kernels). Not wired into nightly-configs.yaml until they land. Validation: the DSV4-Pro-FP4 config (deepseek-ai/DeepSeek-V4-Pro) was validated end-to-end on the hand-driven path earlier (HiSparse decode GSM8K 0.960, HiCache prefill GSM8K 0.950, both 0 GPU faults). Through the CI launcher, boot + PD KV transfer (generate 200) + hisparse/mori-live were confirmed; the full GSM8K gate run under checkout-runtime is pending an image-native run once the dependency PRs merge (checkout-runtime JIT-compiles aiter on first start, which desyncs decode/prefill startup and races the bench probe -- an artifact that disappears with SGLANG_USE_CHECKOUT_RUNTIME=0). Accuracy gate is a placeholder pending that number.
Lzy17
force-pushed
the
dsv4pro-hisparse-hicache-dici
branch
from
August 31, 2026 23:27
0aef73a to
20e6dd5
Compare
… long context The 256MiB host-register chunk faults on long requests: when a KV entry spans more than one chunk, the HiSparse transfer kernel addresses from a single contiguous host base and reads past the chunk. Short requests never cross a chunk, so they pass. A 3.5GiB chunk keeps the KV in one chunk and stays under the ionic ~4GB MR limit. Validated DSV4-Pro 1P1D over MoRI: 95k and 200k-char requests pass, 0 GPU faults, GSM8K 0.955.
…5GiB for long context" This reverts commit ed05ef8.
Lzy17
requested review from
Ying1123,
alphabetc1,
hanming-lu,
hnyls2002,
huangtingwei9988,
hzh0425,
ispobock,
merrymercy,
xiezhq-hermann and
yizhang2077
as code owners
September 11, 2026 04:23
The decode host pool has to sit on 2MiB pages. One ionic pd holds about 1M 4K page-table entries, roughly 4 GiB; the pool is larger, so ibv_reg_mr fails with errno:22 and the decode scheduler aborts. On 2MiB pages the same pool registers clean, 100 GiB on one pd against 3.75 GiB on 4K. Flash needs three settings Pro does not, one per failure it hit: a capped decode token budget, since the pool is the budget over four and Flash budgets 26.2M; a lower prefill HiCache ratio, since at 2 the c4 indexer pool is left 3.67 GB; and a lower prefill memory fraction, since at 0.85 prefill runs out of VRAM and the first lazily loaded Triton kernel has nowhere to go. Measured, 2-node 1P1D on mia1-p02-g29 + g53, each on its own checkpoint: fp4/dsv4pro 0.950, fp8/dsv4pro 0.945, fp4/dsv4flash 0.950, fp8/dsv4flash 0.950.
Lzy17
force-pushed
the
dsv4pro-hisparse-hicache-dici
branch
from
September 11, 2026 04:29
ab38f90 to
bf4f861
Compare
Collaborator
Author
|
Ran all four recipes on hardware: 2-node 1P1D, MI355X,
The stack under test was #29168 → #32368 plus a host-pointer fix. We also built #32368 + #35233 on its own, with a rebuilt gfx950 What the recipes now carry, measured rather than guessed:
The recipes are the run command — full flags and env for both roles — at |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Nightly CI test for DSV4 HiSparse + HiCache in PD over MoRI. One recipe per model runs HiCache on prefill and
HiSparse on decode (they don't conflict — HiCache is prefill-only, HiSparse decode-only). Covers all four
DSV4 variants: Pro/Flash × FP8/FP4.
Also:
*_extra_envafter the DSV4 env so it can overrideit; add a
disable_radix_cacheknob (--disable-radix-cachewas hardcoded but HiCache needs radix on).--disable-overlap-schedule— the overlap scheduler intermittently desyncs the decode TPrequest broadcast (
broadcast_pyobj), crashing the decode scheduler. Disabling it holds.Depends on
Four things must be in the image, in order: #29168 → #32368 → #36966 (sglang), plus #37286
(MoRI bump for the ionic RoCE fix). Not wired into
nightly-configs.yamluntil they land.Validation
DSV4-Pro-FP4, 2-node 1P1D, manual run (image-native sglang): HiSparse (decode) GSM8K 0.960, HiCache
(prefill) 0.950, 0 GPU faults.
CI parity note: run these with
SGLANG_USE_CHECKOUT_RUNTIME=0(image-native) once the deps are in the image.Under checkout-runtime (SGLANG_USE_CHECKOUT_RUNTIME=1) each start JIT-compiles aiter (~20 min), which desyncs
the 8 ranks' startup and is not representative — that path is only for pre-merge testing.
CI States
Latest PR Test (Base): ❌ Run #34562375725
Latest PR Test (Extra): ❌ Run #34562375607
Latest PR Test (AMD ROCm 10): ❌ Run #34562375734