Repository navigation
[HiSparse] Support IndexCache shared layer IO overlap - #28523
huangtingwei9988 wants to merge 2 commits into
Conversation
There was a problem hiding this comment.
Code Review
This pull request introduces IndexCache prefetch capabilities to the HiSparse coordinator, enabling overlapping of host-to-device KV loads with current-layer attention/MLP computations, and integrates TVM-FFI stream synchronization. It also adds comprehensive unit tests to verify prefetching and CUDA graph replay functionality. The review feedback highlights a critical indexing bug in _local_layer_index under pipeline parallelism, and suggests defensive programming improvements to handle None values for both the configuration object and the num_hidden_layers parameter.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| def _local_layer_index(self, layer_id: int) -> Optional[int]: | ||
| if self.is_dsv4_hisparse and 0 <= layer_id < self.mem_pool_device.layer_num: | ||
| return layer_id | ||
| start_layer = getattr(self.mem_pool_device, "start_layer", 0) | ||
| if start_layer == 0 and 0 <= layer_id < self.mem_pool_device.layer_num: | ||
| return layer_id | ||
| local_layer_id = layer_id - start_layer | ||
| if 0 <= local_layer_id < self.mem_pool_device.layer_num: | ||
| return local_layer_id | ||
| return None |
There was a problem hiding this comment.
Under pipeline parallelism (PP > 1), start_layer can be non-zero. If self.is_dsv4_hisparse is True, the current implementation returns layer_id directly if 0 <= layer_id < self.mem_pool_device.layer_num. However, if start_layer > 0 and layer_id is within [0, layer_num), this will incorrectly return layer_id instead of subtracting start_layer (which would point to the wrong local layer or index out of bounds).
By simplifying the method to always subtract start_layer and check if the resulting local_layer_id is within [0, layer_num), we make the method completely robust, correct under PP, and much cleaner.
| def _local_layer_index(self, layer_id: int) -> Optional[int]: | |
| if self.is_dsv4_hisparse and 0 <= layer_id < self.mem_pool_device.layer_num: | |
| return layer_id | |
| start_layer = getattr(self.mem_pool_device, "start_layer", 0) | |
| if start_layer == 0 and 0 <= layer_id < self.mem_pool_device.layer_num: | |
| return layer_id | |
| local_layer_id = layer_id - start_layer | |
| if 0 <= local_layer_id < self.mem_pool_device.layer_num: | |
| return local_layer_id | |
| return None | |
| def _local_layer_index(self, layer_id: int) -> Optional[int]: | |
| start_layer = getattr(self.mem_pool_device, "start_layer", 0) | |
| local_layer_id = layer_id - start_layer | |
| if 0 <= local_layer_id < self.mem_pool_device.layer_num: | |
| return local_layer_id | |
| return None |
| def _cfg_get(config, name: str, default=None): | ||
| if isinstance(config, dict): | ||
| return config.get(name, default) | ||
| return getattr(config, name, default) |
There was a problem hiding this comment.
If config is None, calling getattr(config, name, default) will raise an AttributeError. Adding a defensive check for config is None makes this helper function completely null-safe and robust against missing or incomplete configurations.
| def _cfg_get(config, name: str, default=None): | |
| if isinstance(config, dict): | |
| return config.get(name, default) | |
| return getattr(config, name, default) | |
| def _cfg_get(config, name: str, default=None): | |
| if config is None: | |
| return default | |
| if isinstance(config, dict): | |
| return config.get(name, default) | |
| return getattr(config, name, default) |
| return {} | ||
|
|
||
| end_layer = start_layer + layer_num | ||
| num_hidden_layers = int(_cfg_get(config, "num_hidden_layers", end_layer)) |
There was a problem hiding this comment.
If num_hidden_layers is explicitly set to None in the configuration, _cfg_get will return None, causing int(None) to raise a TypeError. We should provide a fallback to end_layer to make this parsing robust and defensive.
| num_hidden_layers = int(_cfg_get(config, "num_hidden_layers", end_layer)) | |
| num_hidden_layers = int(_cfg_get(config, "num_hidden_layers", None) or end_layer) |
|
Replaced by #34329 |
Motivation
IndexCache currently enables top-k sharing across multiple layers. Upon retrieving the top-k indices from IndexCache, a side CUDA stream prefetches the HiSparse KV pages required for subsequent shared layers. When execution reaches a shared layer, it only needs to wait for the corresponding event, allowing the vast majority of host-to-device I/O to be overlapped with the current layer's attention or MLP computations.
Calculated based on the post-1000-token tail ITL, using no-HiSparse as the baseline:
Modifications
Accuracy Tests
Speed Tests and Profiling
Checklist
Review and Merge Process
/tag-and-rerun-ci,/tag-run-ci-label,/rerun-failed-ciCI States
Latest PR Test (Base): ❌ Run #27814926951
Latest PR Test (Extra): ❌ Run #27814926948