Skip to content

[HiCache] Reset decode ReplaySSM ring cursor on load-back Mamba slots - #41740

Open
rodamani wants to merge 4 commits into
sgl-project:mainfrom
modal-projects:rohan/up/replayssm-load-back-cursor-reset
Open

rodamani wants to merge 4 commits into
sgl-project:mainfrom
modal-projects:rohan/up/replayssm-load-back-cursor-reset

Conversation

@rodamani

@rodamani rodamani commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

With decode ReplaySSM (--enable-linear-replayssm) and HiCache, MambaComponent.prepare_load_back allocates the request's device Mamba slot directly from the allocator. That bypasses HybridReqToTokenPool.alloc, which is where a fresh slot's replayssm_write_pos is reset; the request already holds a slot by the time alloc runs, so the reset branch is skipped. The H2D load-back writes only temporal + conv state (the host tier carries no ring state).

Cursors are reset on release only by the finish-and-insert path. Retract and abort release with is_insert=False and free the slot without touching write_pos. So a load-back can land in a slot whose previous tenant was retracted or aborted mid-decode, and the decode kernel then replays that tenant's ring on top of the loaded checkpoint (cross-request recurrent-state bleed). Cache eviction is not a source: tree-held slots enter the tree with cursor 0.

Speculative ReplaySSM (--enable-linear-replayssm-spec) keeps its cursors per request row and is not affected.

Modifications

  • Reset replayssm_write_pos on the freshly allocated load-back slot in prepare_load_back, mirroring the fresh-slot reset in HybridReqToTokenPool.alloc. No-op when decode ReplaySSM is off.
  • test/registered/unit/mem_cache/test_mamba_load_back_replayssm_reset.py.

Relationship to open PRs: #36345 resets write_pos only on its new buffer-only load-back path; #37837 resets it on the extra_buffer donate path. Neither covers prepare_load_back.

Accuracy Tests

CPU: 2 passed. With the reset removed: test_load_back_slot_starts_with_empty_ring fails (cursor stays nonzero). The test seeds the stale cursor directly; that retract/abort leave one behind is from reading the release paths. Not run end-to-end on GPU.

Speed Tests and Profiling

No hot-path change beyond the fix itself; not separately benchmarked.

Checklist

Review and Merge Process

  1. Ping Merge Oncalls to start the process. See the PR Merge Process.
  2. Get approvals from CODEOWNERS and other reviewers.
  3. Trigger CI tests with comments or contact authorized users to do so.
    • Common commands include /tag-and-rerun-ci, /tag-run-ci-label, /rerun-failed-ci
  4. After green CI and required approvals, ask Merge Oncalls or people with Write permission to merge the PR.

CI States

Latest PR Test (Base): ❌ Run #36767714730
Latest PR Test (Extra): ❌ Run #36767714122
Latest PR Test (AMD ROCm 10): ❌ Run #36767714748

prepare_load_back raw-allocs the request's device mamba slot, bypassing
HybridReqToTokenPool.alloc's fresh-slot hygiene, and the H2D load-back
writes only temporal+conv (the host tier carries no ring state). The GDN
ReplaySSM decode kernel replays ring contents whenever write_pos>0, so a
load-back slot inherited its previous tenant's ring on top of the loaded
checkpoint: cross-conversation recurrent-state bleed, observed as sticky
deterministic wrong retrievals (~10% of revisits) served through HiCache
load-back only.

Reset write_pos on the freshly allocated slot before the load-back lands,
mirroring the fresh-slot reset in HybridReqToTokenPool.alloc.

Relationship to upstream: sgl-project#36345 (open) resets write_pos
only on its new buffer-only load-back path; sgl-project#37837 (open)
resets it on the extra_buffer donate path. Neither covers prepare_load_back.

Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
@rodamani

Copy link
Copy Markdown
Contributor Author

/rerun-test -c test_retraction_mamba_backup.py test_mamba_unittest.py test_linear_replayssm_decode.py

@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test -c test_retraction_mamba_backup.py test_mamba_unittest.py test_linear_replayssm_decode.py:

🚀 ubuntu-latest (1 test): ✅ View workflow run

cd test/ && python3 registered/unit/mem_cache/test_retraction_mamba_backup.py

🚀 1-gpu-5090 (1 test): ✅ View workflow run

cd test/ && python3 registered/unit/mem_cache/test_mamba_unittest.py

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/attention/unittests/gdn/test_linear_replayssm_decode.py

⛔ test/registered/unit/mem_cache/test_mamba_load_back_replayssm_reset.py: File not found: test/registered/unit/mem_cache/test_mamba_load_back_replayssm_reset.py

@rodamani

Copy link
Copy Markdown
Contributor Author

/tag-and-rerun-ci

@github-actions github-actions Bot added the run-ci CI: run the baseline test suite on this PR label Sep 30, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

run-ci CI: run the baseline test suite on this PR unified-radix-cache

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants