Repository navigation
enable hisparse prefetch utilizing IndexShare for GLM-5.2 and more - #29637
xiezhq-hermann wants to merge 2 commits into
Conversation
…exshare # Conflicts: # python/sglang/srt/managers/hisparse_coordinator.py # python/sglang/srt/model_executor/model_runner.py
|
Preview deployment for your docs. Learn more about Mintlify Previews.
💡 Tip: Enable Workflows to automatically generate PRs for you. |
There was a problem hiding this comment.
Code Review
This pull request introduces automatic shared-index prefetching in HiSparse to optimize decode performance for models that reuse anchor layer top-k selections (such as DeepSeek DSA). It refactors HiSparseCoordinator to initialize a separate prefetch stream, events, and buffers, and updates swap_in_selected_pages to asynchronously prefetch skip layers on this stream. Additionally, ModelRunner is updated to auto-enable this feature for compatible models when pipeline parallelism and speculative decoding are disabled. The reviewer feedback suggests synchronizing the newly introduced prefetch_stream in the destroy() method of HiSparseCoordinator to prevent potential race conditions or CUDA errors during teardown.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| group = self._prefetch_groups.setdefault(anchor, []) | ||
| self._prefetch_slot[i] = len(group) | ||
| group.append(i) | ||
|
|
There was a problem hiding this comment.
The newly introduced self.prefetch_stream is not synchronized in the destroy() method of HiSparseCoordinator. To prevent potential race conditions, undefined behavior, or CUDA errors when the coordinator is destroyed while prefetch operations are still in-flight, please ensure self.prefetch_stream.synchronize() is called inside destroy() if self.enable_prefetch is enabled.
|
There is a branch implementing the IO-only path (eliminating duplicated meta data processing on the skipped layers), which is potentially more performant on less powerful GPUs, can you give a try on your benchmark @huangtingwei9988: |
|
Replaced by #34329 |
Motivation
Loading KV cache from host memory on critical path has been a major overhead of HiSparse. This PR utilize ShareIndex feature introduced in GLM-5.2 to enable prefetch, which effectively hides 70% of the IO latency according to our benchmark.
Credit to an earlier PR #28523 by @huangtingwei9988 as well.
Modifications
Accuracy Tests
Speed Tests and Profiling
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): ❌ Run #28358284984
Latest PR Test (Extra): ❌ Run #28359147399