Repository navigation
[HiCache] Support Mamba in buffer-only host memory mode - #36345
alphabetc1 wants to merge 11 commits into
Conversation
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0e8e79dd93
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| mamba_gap_reserve = self._mamba_gap_budget_for_req(req) | ||
| total_tokens += mamba_gap_reserve |
There was a problem hiding this comment.
Reserve both Mamba slots for staged load-back
When buffer-mode admission has limited Mamba capacity, this reserves only the consuming request's slot, but _prepare_mamba_handoff also passes node_copy to cc.load, which auto-allocates a second device slot for the published tree node. Consequently a request can pass the rem_mamba_slots gate with only one recoverable slot, consume it for the request, and then have cc.load fail and discard an otherwise valid L3 hit; if the request already owns a slot, this reserves zero despite still needing the node slot. Account for the tree-node allocation as well so staged Mamba hits are not spuriously recomputed under pressure.
Useful? React with 👍 / 👎.
0e8e79d to
2be4ad8
Compare
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 2be4ad8f4d
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| handoff = self._prepare_mamba_handoff(f, req) | ||
| if handoff is None: | ||
| return _drop() | ||
| load_pools.append(handoff.request_copy) |
There was a problem hiding this comment.
Retain the request slot until the queued H2D finishes
When a staged request passes the initial budget gates but is rejected by a later check in add_one_req (for example the post-load input-token or tile gate), this queues an H2D into its Mamba slot before admission is committed. The rejection path in scheduler.py immediately returns that slot to mamba_allocator, while ready_to_load_host_cache() launches the queued copy only afterward (or during a later batch if no request was admitted), so the transfer can overwrite a slot already reassigned to another request and silently corrupt its recurrent state. The destination slot must remain owned until the load is cancelled or its completion is acknowledged.
Useful? React with 👍 / 👎.
| req.mamba_cow_src_index = None | ||
| req.mamba_needs_clear = False |
There was a problem hiding this comment.
Reset ReplaySSM cursors after restoring a checkpoint
When --enable-linear-replayssm or its speculative variant is enabled, replacing the deferred copy_from with this direct H2D leaves replayssm_write_pos and, where present, replayssm_cache_base/replayssm_is_flush unchanged. Those cursor arrays are not included in the HiCache Mamba transfer, while the ordinary MambaPool.copy_from path explicitly resets them when installing a checkpoint; a newly allocated or previously used request slot can therefore interpret stale ring entries on its first decode and produce incorrect output. Reset the ReplaySSM cursor metadata when the staged state supersedes the pending copy.
Useful? React with 👍 / 👎.
75d4c36 to
5706ad2
Compare
| prep=prep, | ||
| ) | ||
|
|
||
| def defer_staged_admission(self, req: Req, *, pool: str) -> None: |
There was a problem hiding this comment.
can you justify the change from _defer_for_capacity to this function?
There was a problem hiding this comment.
- Moved it out of the closure since add_one_req() also needs to use it now.
- Most requests here don't have staged hits, so the function is a no-op for them. Keeping "staged" in the name makes it clearer what kind of admission it handles, and it matches the related state names (_staged_admission_defers / max_staged_admission_defers).
There was a problem hiding this comment.
one more thing, admission could get deferred for reasons other than capacity (load fail). But defer_for_capacity still seems fine here. What do you think?
|
I refactored the pool management for different components in this PR: #42633 |
Motivation
--hicache-host-memory-mode buffer_onlyrejects any tree with a Mamba component, so hybrid Mamba/GDN models cannot use it. This PR lifts that fence. Part of #34899.Depends on #43138 (its commit is included here until it merges).
Changes
On top of #42633, the write path needs no Mamba code: per-pool backups already write a node's Mamba state under its last page hash. What remains is the read path, which this PR moves behind component hooks so the pipeline stays component-agnostic.
TreeComponent, no-op defaults):prepare_buffer_load_back(req, staged)returns aBufferLoadBack, the component's part in an admission-time load-back. Its parts are extra H2D destinations,finish(success), the insert fields, and the destinations the transfer ack must free. It returns None when the device cannot take the load yet, and the admission defers.validate_buffer_mode()checks the component's pools at init.MambaComponent: loads the staged state into the published node's slot and, layer-gated, into the request's own slot (a device CoW would run before the layer gate). On failure it rolls the request slot back; on success it resets the ReplaySSM cursor and drops the CoW/clear. When the tail already has a checkpoint (mamba_exist), the redundant node destination is freed at the ack. It rejects a Mamba host pool with fewer than 2 slots and int8 checkpoints.mamba_slotsintoStagedPrefetchPlanandreq.mamba_host_hit_length, and drops a staged hit that carries no Mamba state, so the request recomputes. The capacity deferral moves from a closure intodefer_staged_admission, which the adder also calls when the Mamba slots fall short._aux_loads_marginis per pool: SWA reserves a window, other pools one page.TestUnifiedHybridBufferOnlyBitExact(Inkling, 1 GPU) seedsP[:512]into L3, flushes the device tree, then measuresP[:768], so every hit comes back through the buffer-mode read path. The unit fixture's load-back requests are Mocks, so its Mamba buffer-only cases stay skipped.Results
H200,
thinkingmachines/Inkling@test:Unit tests (
unit/mem_cacheHiCache suites,unit/dllm/test_gemma4_uniform_lifecycle.py,unit/managers; Rust tree core built from source) pass the same set as main (2228 passed).test_buffer_only_rejects_mambabecomestest_buffer_only_accepts_mamba;test_mm_process_configfails on main too.Benchmark
Whether #42633's per-pool writes cost anything on a Mamba stack: arm A merges this PR onto #42633's parent (Mamba state rides the KV write), arm B onto #42633. 2xH200,
Inkling-Small-NVFP4TP2, mooncake, buffer_only. Each round writes 256 random 4096-token prompts (64 output tokens, concurrency 32), flushes, then reads them back extended by 512 tokens. Boot order A B B A, mean of rounds 2-3 of each boot (n=4 per arm):Both arms issued the same number of mooncake puts and D2H batches, with the same page-count histogram, and backed up the same number of Mamba states.
CI States
Latest PR Test (Base): ❌ Run #37814366030
Latest PR Test (Extra): ❌ Run #37814365484
Latest PR Test (AMD ROCm 10): ⏳ Run #37814366279