Skip to content

[Unified Cache][9/N] add opt-in MLA load deduplication for Mooncake Linker - #39565

Merged
huangtingwei9988 merged 8 commits into
sgl-project:mainfrom
antgroup:linker_mla_dedup
Sep 20, 2026
Merged

huangtingwei9988 merged 8 commits into
sgl-project:mainfrom
antgroup:linker_mla_dedup

Conversation

@huangtingwei9988

@huangtingwei9988 huangtingwei9988 commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Add --enable-linker-mla-dedup to load replicated MLA KV caches from Mooncake on TP rank 0, then broadcast the loaded data layer by layer to the other TP ranks.

The feature is opt-in and reuses MLAHostDedupBroadcaster.

Performance

GLM-5.2 W4AFP8 on 8 × H20 96GB, TP8 / PP1, concurrency 1.
CUDA graphs enabled, chunked prefill size 8192, FP8 KV cache.

Cached prefix All-rank load TTFT Rank-0 load + broadcast TTFT TTFT reduction
128K 1.959 s 1.185 s 39.5%
256K 3.851 s 2.229 s 42.1%
512K 7.950 s 4.891 s 38.5%

Accuracy Tests

Speed Tests and Profiling

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ✅ Run #35460066451
Latest PR Test (Extra): ❌ Run #35460066360
Latest PR Test (AMD ROCm 10): ⏳ Run #35460066435

@github-actions github-actions Bot added the hicache Hierarchical Caching for SGLang label Sep 15, 2026
Co-authored-by: Zhangheng <hzh0425@apache.org>
@huangtingwei9988 huangtingwei9988 added run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) labels Sep 15, 2026
Leoyzen added a commit to Leoyzen/sglang that referenced this pull request Sep 17, 2026
[Unified Cache][9/N] add opt-in MLA load deduplication for Mooncake Linker.

Squashes 5 upstream commits (cc4431b, 455f5bb, 310b76e, 8acfd72,
d52e749 — three of them merge-main commits whose only net effect on this
PR's files is what lands below). New opt-in flag --enable-linker-mla-dedup:
with the Mooncake external linker, rank 0 loads replicated MLA KV from L3 and
broadcasts each layer to the other TP ranks instead of every rank reading the
same objects.

Files:
- python/sglang/srt/arg_groups/fields/memory.py: add enable_linker_mla_dedup.
- python/sglang/srt/arg_groups/hicache_hook.py: require the Mooncake linker.
- python/sglang/srt/mem_cache/unified_cache/linker_mla_dedup.py (new):
  LinkerMLADedupBroadcaster, adapting hybrid device-pool geometry to the
  existing MLAHostDedupBroadcaster primitive.
- python/sglang/srt/mem_cache/storage/mooncake_store/mooncake_direct_linker.py:
  LayerWiseLoadCounter on_layer_ready hook; broadcaster construction; per-batch
  logical-plan all_gather gate; rank-0-only load queueing; forward-stream
  broadcasts; broadcast-completion event tracking in num_completed_loads;
  reset/close teardown.
- test/registered/unit/server_args/test_server_args.py: flag validation test.

Manual merge notes:
- Our fork's mooncake_direct_linker.py carries fork-local behavior (fail-soft
  revalidate_load, retrying session start, SGLANG_LINKER_DEBUG_KEY probes,
  idx-checksum probe, rank-key handling). All of it is preserved untouched;
  the upstream hunks apply around it (imports, LayerWiseLoadCounter, __init__
  broadcaster wiring, num_completed_loads, start_layer_wise_loading,
  load_thread_func, reset/close).
- The upstream branch was based on main before sgl-project#39835/sgl-project#39115/sgl-project#37474; those are
  unrelated to this PR's files (verified by diffing the PR range against the
  merge base), so the squash is a faithful net port.
- Depends on MLAHostDedupBroadcaster and attn_tp_cache_group/tp_cache_group in
  CacheInitParams, both already present in our base (upstream parent == ours).
Leoyzen added a commit to Leoyzen/sglang that referenced this pull request Sep 17, 2026
[Unified Cache][9/N] add opt-in MLA load deduplication for Mooncake Linker.

Squashes 5 upstream commits (cc4431b, 455f5bb, 310b76e, 8acfd72,
d52e749 — three of them merge-main commits whose only net effect on this
PR's files is what lands below). New opt-in flag --enable-linker-mla-dedup:
with the Mooncake external linker, rank 0 loads replicated MLA KV from L3 and
broadcasts each layer to the other TP ranks instead of every rank reading the
same objects.

Files:
- python/sglang/srt/arg_groups/fields/memory.py: add enable_linker_mla_dedup.
- python/sglang/srt/arg_groups/hicache_hook.py: require the Mooncake linker.
- python/sglang/srt/mem_cache/unified_cache/linker_mla_dedup.py (new):
  LinkerMLADedupBroadcaster, adapting hybrid device-pool geometry to the
  existing MLAHostDedupBroadcaster primitive.
- python/sglang/srt/mem_cache/storage/mooncake_store/mooncake_direct_linker.py:
  LayerWiseLoadCounter on_layer_ready hook; broadcaster construction; per-batch
  logical-plan all_gather gate; rank-0-only load queueing; forward-stream
  broadcasts; broadcast-completion event tracking in num_completed_loads;
  reset/close teardown.
- test/registered/unit/server_args/test_server_args.py: flag validation test.

Manual merge notes:
- Our fork's mooncake_direct_linker.py carries fork-local behavior (fail-soft
  revalidate_load, retrying session start, SGLANG_LINKER_DEBUG_KEY probes,
  idx-checksum probe, rank-key handling). All of it is preserved untouched;
  the upstream hunks apply around it (imports, LayerWiseLoadCounter, __init__
  broadcaster wiring, num_completed_loads, start_layer_wise_loading,
  load_thread_func, reset/close).
- The upstream branch was based on main before sgl-project#39835/sgl-project#39115/sgl-project#37474; those are
  unrelated to this PR's files (verified by diffing the PR range against the
  merge base), so the squash is a faithful net port.
- Depends on MLAHostDedupBroadcaster and attn_tp_cache_group/tp_cache_group in
  CacheInitParams, both already present in our base (upstream parent == ours).
Leoyzen added a commit to Leoyzen/sglang that referenced this pull request Sep 18, 2026
[Unified Cache][9/N] add opt-in MLA load deduplication for Mooncake Linker.

Squashes 5 upstream commits (cc4431b, 455f5bb, 310b76e, 8acfd72,
d52e749 — three of them merge-main commits whose only net effect on this
PR's files is what lands below). New opt-in flag --enable-linker-mla-dedup:
with the Mooncake external linker, rank 0 loads replicated MLA KV from L3 and
broadcasts each layer to the other TP ranks instead of every rank reading the
same objects.

Files:
- python/sglang/srt/arg_groups/fields/memory.py: add enable_linker_mla_dedup.
- python/sglang/srt/arg_groups/hicache_hook.py: require the Mooncake linker.
- python/sglang/srt/mem_cache/unified_cache/linker_mla_dedup.py (new):
  LinkerMLADedupBroadcaster, adapting hybrid device-pool geometry to the
  existing MLAHostDedupBroadcaster primitive.
- python/sglang/srt/mem_cache/storage/mooncake_store/mooncake_direct_linker.py:
  LayerWiseLoadCounter on_layer_ready hook; broadcaster construction; per-batch
  logical-plan all_gather gate; rank-0-only load queueing; forward-stream
  broadcasts; broadcast-completion event tracking in num_completed_loads;
  reset/close teardown.
- test/registered/unit/server_args/test_server_args.py: flag validation test.

Manual merge notes:
- Our fork's mooncake_direct_linker.py carries fork-local behavior (fail-soft
  revalidate_load, retrying session start, SGLANG_LINKER_DEBUG_KEY probes,
  idx-checksum probe, rank-key handling). All of it is preserved untouched;
  the upstream hunks apply around it (imports, LayerWiseLoadCounter, __init__
  broadcaster wiring, num_completed_loads, start_layer_wise_loading,
  load_thread_func, reset/close).
- The upstream branch was based on main before sgl-project#39835/sgl-project#39115/sgl-project#37474; those are
  unrelated to this PR's files (verified by diffing the PR range against the
  merge base), so the squash is a faithful net port.
- Depends on MLAHostDedupBroadcaster and attn_tp_cache_group/tp_cache_group in
  CacheInitParams, both already present in our base (upstream parent == ours).
@huangtingwei9988
huangtingwei9988 merged commit e9300f6 into sgl-project:main Sep 20, 2026
167 of 200 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hicache Hierarchical Caching for SGLang run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) unified-radix-cache

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants