Skip to content

[HiCache] Support Mamba in buffer-only host memory mode - #36345

Open
alphabetc1 wants to merge 11 commits into
sgl-project:mainfrom
alphabetc1:buffer-only-mamba
Open

alphabetc1 wants to merge 11 commits into
sgl-project:mainfrom
alphabetc1:buffer-only-mamba

Conversation

@alphabetc1

@alphabetc1 alphabetc1 commented Aug 25, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

--hicache-host-memory-mode buffer_only rejects any tree with a Mamba component, so hybrid Mamba/GDN models cannot use it. This PR lifts that fence. Part of #34899.

Depends on #43138 (its commit is included here until it merges).

Changes

On top of #42633, the write path needs no Mamba code: per-pool backups already write a node's Mamba state under its last page hash. What remains is the read path, which this PR moves behind component hooks so the pipeline stays component-agnostic.

  • Component hooks (TreeComponent, no-op defaults): prepare_buffer_load_back(req, staged) returns a BufferLoadBack, the component's part in an admission-time load-back. Its parts are extra H2D destinations, finish(success), the insert fields, and the destinations the transfer ack must free. It returns None when the device cannot take the load yet, and the admission defers. validate_buffer_mode() checks the component's pools at init.
  • MambaComponent: loads the staged state into the published node's slot and, layer-gated, into the request's own slot (a device CoW would run before the layer gate). On failure it rolls the request slot back; on success it resets the ReplaySSM cursor and drops the CoW/clear. When the tail already has a checkpoint (mamba_exist), the redundant node destination is freed at the ack. It rejects a Mamba host pool with fewer than 2 slots and int8 checkpoints.
  • Pipeline: plans mamba_slots into StagedPrefetchPlan and req.mamba_host_hit_length, and drops a staged hit that carries no Mamba state, so the request recomputes. The capacity deferral moves from a closure into defer_staged_admission, which the adder also calls when the Mamba slots fall short. _aux_loads_margin is per pool: SWA reserves a window, other pools one page.
  • Test: TestUnifiedHybridBufferOnlyBitExact (Inkling, 1 GPU) seeds P[:512] into L3, flushes the device tree, then measures P[:768], so every hit comes back through the buffer-mode read path. The unit fixture's load-back requests are Mocks, so its Mamba buffer-only cases stay skipped.

Results

H200, thinkingmachines/Inkling@test:

TestUnifiedHybridBufferOnlyBitExact    storage hits [512] x 8, kl_divs 0.0 x 8
TestUnifiedHybridHiCacheBitExact       kl_divs 0.0 x 9

Unit tests (unit/mem_cache HiCache suites, unit/dllm/test_gemma4_uniform_lifecycle.py, unit/managers; Rust tree core built from source) pass the same set as main (2228 passed). test_buffer_only_rejects_mamba becomes test_buffer_only_accepts_mamba; test_mm_process_config fails on main too.

Benchmark

Whether #42633's per-pool writes cost anything on a Mamba stack: arm A merges this PR onto #42633's parent (Mamba state rides the KV write), arm B onto #42633. 2xH200, Inkling-Small-NVFP4 TP2, mooncake, buffer_only. Each round writes 256 random 4096-token prompts (64 output tokens, concurrency 32), flushes, then reads them back extended by 512 tokens. Boot order A B B A, mean of rounds 2-3 of each boot (n=4 per arm):

A B
write input tok/s 17,098 17,098
read latency mean / p50 (ms) 1,141 / 1,089 1,141 / 1,096
L3 hit tokens per round 1,048,576 / 1,028,096 identical

Both arms issued the same number of mooncake puts and D2H batches, with the same page-count histogram, and backed up the same number of Mamba states.


CI States

Latest PR Test (Base): ❌ Run #37814366030
Latest PR Test (Extra): ❌ Run #37814365484
Latest PR Test (AMD ROCm 10): ⏳ Run #37814366279

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 0e8e79dd93

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +1239 to +1240
mamba_gap_reserve = self._mamba_gap_budget_for_req(req)
total_tokens += mamba_gap_reserve

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Reserve both Mamba slots for staged load-back

When buffer-mode admission has limited Mamba capacity, this reserves only the consuming request's slot, but _prepare_mamba_handoff also passes node_copy to cc.load, which auto-allocates a second device slot for the published tree node. Consequently a request can pass the rem_mamba_slots gate with only one recoverable slot, consume it for the request, and then have cc.load fail and discard an otherwise valid L3 hit; if the request already owns a slot, this reserves zero despite still needing the node slot. Account for the tree-node allocation as well so staged Mamba hits are not spuriously recomputed under pressure.

Useful? React with 👍 / 👎.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 2be4ad8f4d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

handoff = self._prepare_mamba_handoff(f, req)
if handoff is None:
return _drop()
load_pools.append(handoff.request_copy)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Retain the request slot until the queued H2D finishes

When a staged request passes the initial budget gates but is rejected by a later check in add_one_req (for example the post-load input-token or tile gate), this queues an H2D into its Mamba slot before admission is committed. The rejection path in scheduler.py immediately returns that slot to mamba_allocator, while ready_to_load_host_cache() launches the queued copy only afterward (or during a later batch if no request was admitted), so the transfer can overwrite a slot already reassigned to another request and silently corrupt its recurrent state. The destination slot must remain owned until the load is cancelled or its completion is acknowledged.

Useful? React with 👍 / 👎.

Comment on lines +709 to +710
req.mamba_cow_src_index = None
req.mamba_needs_clear = False

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Reset ReplaySSM cursors after restoring a checkpoint

When --enable-linear-replayssm or its speculative variant is enabled, replacing the deferred copy_from with this direct H2D leaves replayssm_write_pos and, where present, replayssm_cache_base/replayssm_is_flush unchanged. Those cursor arrays are not included in the HiCache Mamba transfer, while the ordinary MambaPool.copy_from path explicitly resets them when installing a checkpoint; a newly allocated or previously used request slot can therefore interpret stale ring entries on its first decode and produce incorrect output. Reset the ReplaySSM cursor metadata when the staged state supersedes the pending copy.

Useful? React with 👍 / 👎.

@alphabetc1 alphabetc1 added bypass-fastfail run-ci-extra CI: also run the extra suite (requires run-ci) labels Sep 17, 2026
prep=prep,
)

def defer_staged_admission(self, req: Req, *, pool: str) -> None:

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

can you justify the change from _defer_for_capacity to this function?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  1. Moved it out of the closure since add_one_req() also needs to use it now.
  2. Most requests here don't have staged hits, so the function is a no-op for them. Keeping "staged" in the name makes it clearer what kind of admission it handles, and it matches the related state names (_staged_admission_defers / max_staged_admission_defers).

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

one more thing, admission could get deferred for reasons other than capacity (load fail). But defer_for_capacity still seems fine here. What do you think?

@xiezhq-hermann

Copy link
Copy Markdown
Collaborator

I refactored the pool management for different components in this PR: #42633
can you help take a look and see whether it can help simplifying your design and implementation here? tldr is with this PR now all components are managed independently, decoupling complexity. we might want to merge the mamba support rebasing on that if helpful

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bypass-fastfail hicache Hierarchical Caching for SGLang run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) unified-radix-cache

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants