Repository navigation
[Mamba] Reset the ReplaySSM ring cursor on the extra_buffer donate path - #37837
Open
davidli1515 wants to merge 2 commits into
Open
davidli1515 wants to merge 2 commits into
davidli1515 wants to merge 2 commits into
Conversation
`--enable-linear-replayssm` required `--mamba-radix-cache-strategy no_buffer`
because `donate_mamba_ping_pong_slot` never reset `replayssm_write_pos`, unlike
`MambaPool.copy_from` on the no_buffer donate path.
That requirement has a large cost: `no_buffer` attaches the mamba checkpoint at
`token_ids_len` (prompt + the request's own generated tokens), a depth later
requests sharing only the prompt cannot reach. The mamba match validator then
rejects the node, and because a forward pass resumes from one position for all
layers, the KV hit is discarded with it. On a 2048-token shared prefix this
takes the prefix cache hit rate from 88.2% to 65.9% and mean TTFT from 226 ms to
1063 ms.
Reset the cursor on both the donated and the replacement slot, then drop the
guard. Measured on Qwen3.8-27B / 1xH200 / TP=1:
- GSM8K (1319 q): 0.957 patched vs 0.953 no_buffer vs 0.955 no-ReplaySSM,
all within +/-0.028 (2 s.e.)
- rag workload: hit rate 65.9% -> 88.2%, mean TTFT 1063 ms -> 228 ms,
matching the no-ReplaySSM arm's 226 ms
- ShareGPT control (no shared prefixes): all three arms within 3%
- L in {8, 16, 32}: no errors, hit rate stable
- 6.5 h soak, 436 rounds: no asserts, accuracy drift +0.003
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
davidli1515
requested review from
Ying1123,
alphabetc1,
hanming-lu,
hnyls2002,
huangtingwei9988,
hzh0425,
ispobock,
merrymercy,
xiezhq-hermann and
yizhang2077
as code owners
September 3, 2026 16:21
yuan-luo
reviewed
Sep 4, 2026
yuan-luo
reviewed
Sep 4, 2026
…y, keep the KDA guard
Reset the ReplaySSM ring cursor at the two points where a slot enters a
ping-pong track buffer instead of inside donate_mamba_ping_pong_slot:
_alloc_ping_pong_buffer initial allocation, mirroring what alloc()
already does for the live decode slot
set_mamba_ping_pong_slot the donate swap and both lazy on-demand paths
(mamba_lazy_prealloc_at_boundary and
mamba_lazy_spec_prepare)
The invariant is that a slot entering a ping-pong track buffer always has a
zero ring cursor. The finished-request insert is then correct without a reset
of its own: it reads a slot already in the buffer, and the cursor is only
advanced for mamba_cache_indices, i.e. live decode slots. The donate-path
reset drops for the same reason.
value is a device-side id for the three install callers and a host-side -1 for
the two clear callers, so the branch is on type; comparing the tensor on the
host would force a cudaStreamSynchronize on the allocation path.
Keep requiring no_buffer for KDA, since HybridLinearAttnBackend only builds
replayssm_force_flush when not is_kda, so a KDA track snapshot would be taken
with ring entries still unfolded. The predicate matches on the config rather
than reading mamba2_cache_params.is_kda: building the cache params calls
get_parallel().attn_tp_size, and this handler runs before the process groups
exist. KimiLinearCacheParams is the only params class with is_kda == True and
has exactly two producers, KimiLinearConfig and BailingHybridConfig with
use_kda, which is what kimi_linear_config matches.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
4 tasks done
3 of 5 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
[Mamba] Reset the ReplaySSM ring cursor on the extra_buffer donate path
Fixes #37834
Motivation
--enable-linear-replayssmrequires--mamba-radix-cache-strategy no_buffer.The guard's own comment says why:
That requirement turns out to be expensive.
no_bufferattaches the mambacheckpoint at
token_ids_len— prompt plus the request's own generatedtokens — which is a depth later requests sharing only the prompt can never
reach. The mamba match validator rejects the node, and because a forward pass
resumes from a single position for all layers, the KV hit is discarded along
with it. Instrumenting the match path on a 601-token prompt whose predecessor
stored 608:
KV alone could have reused 600 tokens. It reused none.
Modifications
Two edits, 12 insertions / 14 deletions.
1.
mem_cache/memory_pool.py—HybridReqToTokenPool.donate_mamba_ping_pong_slotReset
replayssm_write_poson the donated slot and on the replacement slot. Thedonated checkpoint goes to the radix cache and the replacement slot starts a
fresh tracking window, so neither carries pending ring entries. This is the same
reset
MambaPool.copy_fromalready performs on the no_buffer donate path.2.
arg_groups/attention_hook.pyDrop the guard that forced
no_buffer, update the stale comment above it, andremove the now-unused
mamba_extra_buffer_ofimport.Accuracy
Bit-exact output comparison is not the right bar here, for two independent
reasons:
"Tensor-core precision (~4e-4 TF32 / ~1e-3 bf16) is benign end-to-end
(ReplaySSM bf16 GSM8K parity)". Upstream validates it by GSM8K parity.
checkpoint is a different code path from recomputing, so near-tie argmax
decisions flip. Any change that turns misses into hits will move some tokens.
So the bar is accuracy parity. Full GSM8K, 1319 questions, Qwen3.8-27B / 1×H200 /
TP=1. All arms use
--disable-overlap-schedulebecauseno_bufferassertsagainst the overlap scheduler; holding it fixed keeps the strategy the only
variable.
extra_bufferno_bufferextra_bufferafter − before = +0.004, against ±0.028 at two standard errors for n=1319. Allthree arms sit inside that band.
Performance
Same three arms,
sglang.benchmark.serving,--request-rate 8 --max-concurrency 48.rag — 2048-token shared prefix, 8 groups × 32 prompts
fewshot — 512-token shared prefix, 32 groups × 8 prompts
sharegpt — real conversations, little cross-request sharing (control)
The patched arm lands on the no-ReplaySSM arm's numbers wherever prefixes are
shared, and the ShareGPT control shows all three arms within noise where they are
not — which is what confirms the gain comes from prefix reuse rather than some
other side effect.
Output throughput is flat in every row because these runs sat well below
saturation: at
--request-rate 8with 64-token outputs the offered rate is~512 tok/s and the server delivered 508–510, while the ShareGPT rows on the same
server reach 1137–1163 tok/s. Below saturation the extra prefill work lands on
TTFT and E2E rather than on throughput.
Robustness
Ring length sweep (patched arm; the patch touches the ring cursor and
Lsets the ring length, so this is where config-dependent breakage would show):
6.5-hour soak on the patched arm, cycling the three workloads: 436 rounds,
no assertions, no tracebacks, no CUDA errors, no OOM.
A stale ring cursor fails silently and cumulatively rather than crashing, so
accuracy was re-measured on the same server instance after the soak: 0.960 over
500 questions, +0.003 against the same arm's pre-soak number. No drift.
Checklist
sglang.test.few_shot_gsm8kandsglang.benchmark.serving; happy to add aregression test asserting
cached_tokens > 0under--enable-linear-replayssmif that is wantedupdated in place
Scope of verification
Being explicit about what was and was not covered:
main. Both touched hunks are byte-identicalon
main, but the surroundingReqstructure has since been refactored(
req.mamba_*→req.kv.mamba_*), so CI confirmation onmainwould beworthwhile.
under tensor parallelism.
MambaRadixCachehas the same codeshape (
mamba_component.pynotes it "MirrorsMambaRadixCache.cache_finished_req") but was not tested.CI States
Latest PR Test (Base): ❌ Run #33939143898
Latest PR Test (Extra): ❌ Run #33939143727
Latest PR Test (AMD ROCm 7.2): ❌ Run #33939143796