Skip to content

[Mamba] Reset the ReplaySSM ring cursor on the extra_buffer donate path - #37837

Open
davidli1515 wants to merge 2 commits into
sgl-project:mainfrom
davidli1515:fix/replayssm-extra-buffer-prefix-cache
Open

davidli1515 wants to merge 2 commits into
sgl-project:mainfrom
davidli1515:fix/replayssm-extra-buffer-prefix-cache

Conversation

@davidli1515

@davidli1515 davidli1515 commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

[Mamba] Reset the ReplaySSM ring cursor on the extra_buffer donate path

Fixes #37834

Motivation

--enable-linear-replayssm requires --mamba-radix-cache-strategy no_buffer.
The guard's own comment says why:

The extra_buffer strategy donates the track snapshot via
donate_mamba_ping_pong_slot with a separate ping-pong slot swap that does NOT
route through MambaPool.copy_from, so the ReplaySSM ring cursor of the
donated/kept slot would not be reset there. Handling that donation path is a
follow-up; for now require no_buffer.

That requirement turns out to be expensive. no_buffer attaches the mamba
checkpoint at token_ids_len — prompt plus the request's own generated
tokens
— which is a depth later requests sharing only the prompt can never
reach. The mamba match validator rejects the node, and because a forward pass
resumes from a single position for all layers, the KV hit is discarded along
with it. Instrumenting the match path on a 601-token prompt whose predecessor
stored 608:

insert:  finished=True  token_ids_len=608 -> cache_len=608
match:   full_kv_hit_length=600   mamba_boundary=0   branching_seqlen=576
result:  #new-token: 601, #cached-token: 0

KV alone could have reused 600 tokens. It reused none.

Modifications

Two edits, 12 insertions / 14 deletions.

1. mem_cache/memory_pool.py — HybridReqToTokenPool.donate_mamba_ping_pong_slot

Reset replayssm_write_pos on the donated slot and on the replacement slot. The
donated checkpoint goes to the radix cache and the replacement slot starts a
fresh tracking window, so neither carries pending ring entries. This is the same
reset MambaPool.copy_from already performs on the no_buffer donate path.

2. arg_groups/attention_hook.py

Drop the guard that forced no_buffer, update the stale comment above it, and
remove the now-unused mamba_extra_buffer_of import.

Accuracy

Bit-exact output comparison is not the right bar here, for two independent
reasons:

  • ReplaySSM's reconstruction runs on tensor cores — its own kernel comment says
    "Tensor-core precision (~4e-4 TF32 / ~1e-3 bf16) is benign end-to-end
    (ReplaySSM bf16 GSM8K parity)". Upstream validates it by GSM8K parity.
  • A prefix-cache hit changes numerics by itself: resuming from a stored
    checkpoint is a different code path from recomputing, so near-tie argmax
    decisions flip. Any change that turns misses into hits will move some tokens.

So the bar is accuracy parity. Full GSM8K, 1319 questions, Qwen3.8-27B / 1×H200 /
TP=1. All arms use --disable-overlap-schedule because no_buffer asserts
against the overlap scheduler; holding it fixed keeps the strategy the only
variable.

arm config accuracy invalid cache hit
no ReplaySSM extra_buffer 0.955 0.000 88.8%
before ReplaySSM + no_buffer 0.953 0.000 92.3%
after ReplaySSM + extra_buffer 0.957 0.000 88.8%

after − before = +0.004, against ±0.028 at two standard errors for n=1319. All
three arms sit inside that band.

Performance

Same three arms, sglang.benchmark.serving, --request-rate 8 --max-concurrency 48.

rag — 2048-token shared prefix, 8 groups × 32 prompts

arm cache hit out tok/s TTFT mean TTFT p99 E2E mean
no ReplaySSM 88.2% 509.6 226.4 ms 1788.2 ms 2164.7 ms
before 65.9% 508.7 1062.6 ms 4116.2 ms 3876.4 ms
after 88.2% 509.4 228.2 ms 1789.0 ms 2157.6 ms
−78.5% −56.5% −44.3%

fewshot — 512-token shared prefix, 32 groups × 8 prompts

arm cache hit out tok/s TTFT mean TTFT p99 E2E mean
no ReplaySSM 62.2% 510.1 83.0 ms 261.5 ms 1959.6 ms
before 0.0% 508.9 159.0 ms 622.3 ms 3048.2 ms
after 62.2% 510.0 83.6 ms 302.0 ms 1982.2 ms
−47.4% −51.5% −35.0%

sharegpt — real conversations, little cross-request sharing (control)

arm cache hit out tok/s TTFT mean TTFT p99 E2E mean
no ReplaySSM 0.7% 1136.7 95.9 ms 476.6 ms 7145.5 ms
before 0.0% 1162.6 97.0 ms 477.8 ms 6929.3 ms
after 0.7% 1148.5 94.4 ms 476.8 ms 7037.4 ms
−2.6% −0.2% +1.6%

The patched arm lands on the no-ReplaySSM arm's numbers wherever prefixes are
shared, and the ShareGPT control shows all three arms within noise where they are
not — which is what confirms the gain comes from prefix reuse rather than some
other side effect.

Output throughput is flat in every row because these runs sat well below
saturation: at --request-rate 8 with 64-token outputs the offered rate is
~512 tok/s and the server delivered 508–510, while the ShareGPT rows on the same
server reach 1137–1163 tok/s. Below saturation the extra prefill work lands on
TTFT and E2E rather than on throughput.

Robustness

Ring length sweep (patched arm; the patch touches the ring cursor and L
sets the ring length, so this is where config-dependent breakage would show):

L GSM8K (200 q) cache hit errors
8 0.975 88.4% 0
16 0.970 88.4% 0
32 0.975 88.4% 0

6.5-hour soak on the patched arm, cycling the three workloads: 436 rounds,
no assertions, no tracebacks, no CUDA errors, no OOM.

A stale ring cursor fails silently and cumulatively rather than crashing, so
accuracy was re-measured on the same server instance after the soak: 0.960 over
500 questions, +0.003 against the same arm's pre-soak number.
No drift.

Checklist

  • Format the code and run the linters
  • Add unit tests / benchmarks where applicable — measured with
    sglang.test.few_shot_gsm8k and sglang.benchmark.serving; happy to add a
    regression test asserting cached_tokens > 0 under
    --enable-linear-replayssm if that is wanted
  • Update documentation — the guard comment describing the limitation is
    updated in place

Scope of verification

Being explicit about what was and was not covered:

  • Verified on v0.5.18, not on main. Both touched hunks are byte-identical
    on main, but the surrounding Req structure has since been refactored
    (req.mamba_* → req.kv.mamba_*), so CI confirmation on main would be
    worthwhile.
  • Single GPU, TP=1 only. I have not exercised the ping-pong slot allocation
    under tensor parallelism.
  • Model: Qwen3.8-27B (GDN hybrid). Not tested on Mamba-2 hybrids or KDA.
  • UnifiedRadixCache only. The legacy MambaRadixCache has the same code
    shape (mamba_component.py notes it "Mirrors
    MambaRadixCache.cache_finished_req") but was not tested.

CI States

Latest PR Test (Base): ❌ Run #33939143898
Latest PR Test (Extra): ❌ Run #33939143727
Latest PR Test (AMD ROCm 7.2): ❌ Run #33939143796

`--enable-linear-replayssm` required `--mamba-radix-cache-strategy no_buffer`
because `donate_mamba_ping_pong_slot` never reset `replayssm_write_pos`, unlike
`MambaPool.copy_from` on the no_buffer donate path.

That requirement has a large cost: `no_buffer` attaches the mamba checkpoint at
`token_ids_len` (prompt + the request's own generated tokens), a depth later
requests sharing only the prompt cannot reach. The mamba match validator then
rejects the node, and because a forward pass resumes from one position for all
layers, the KV hit is discarded with it. On a 2048-token shared prefix this
takes the prefix cache hit rate from 88.2% to 65.9% and mean TTFT from 226 ms to
1063 ms.

Reset the cursor on both the donated and the replacement slot, then drop the
guard. Measured on Qwen3.8-27B / 1xH200 / TP=1:

  - GSM8K (1319 q): 0.957 patched vs 0.953 no_buffer vs 0.955 no-ReplaySSM,
    all within +/-0.028 (2 s.e.)
  - rag workload: hit rate 65.9% -> 88.2%, mean TTFT 1063 ms -> 228 ms,
    matching the no-ReplaySSM arm's 226 ms
  - ShareGPT control (no shared prefixes): all three arms within 3%
  - L in {8, 16, 32}: no errors, hit rate stable
  - 6.5 h soak, 436 rounds: no asserts, accuracy drift +0.003

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Comment thread python/sglang/srt/arg_groups/attention_hook.py
Comment thread python/sglang/srt/mem_cache/memory_pool.py Outdated
…y, keep the KDA guard

Reset the ReplaySSM ring cursor at the two points where a slot enters a
ping-pong track buffer instead of inside donate_mamba_ping_pong_slot:

  _alloc_ping_pong_buffer     initial allocation, mirroring what alloc()
                              already does for the live decode slot
  set_mamba_ping_pong_slot    the donate swap and both lazy on-demand paths
                              (mamba_lazy_prealloc_at_boundary and
                              mamba_lazy_spec_prepare)

The invariant is that a slot entering a ping-pong track buffer always has a
zero ring cursor. The finished-request insert is then correct without a reset
of its own: it reads a slot already in the buffer, and the cursor is only
advanced for mamba_cache_indices, i.e. live decode slots. The donate-path
reset drops for the same reason.

value is a device-side id for the three install callers and a host-side -1 for
the two clear callers, so the branch is on type; comparing the tensor on the
host would force a cudaStreamSynchronize on the allocation path.

Keep requiring no_buffer for KDA, since HybridLinearAttnBackend only builds
replayssm_force_flush when not is_kda, so a KDA track snapshot would be taken
with ring entries still unfolded. The predicate matches on the config rather
than reading mamba2_cache_params.is_kda: building the cache params calls
get_parallel().attn_tp_size, and this handler runs before the process groups
exist. KimiLinearCacheParams is the only params class with is_kda == True and
has exactly two producers, KimiLinearConfig and BailingHybridConfig with
use_kda, which is what kimi_linear_config matches.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

@yuan-luo yuan-luo left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] --enable-linear-replayssm forces no_buffer, which degrades mamba prefix caching and inflates TTFT up to 4.7x

2 participants