Fix recurrent state loss on decode retraction - #35957
Conversation
|
/rerun-test registered/unit/mem_cache/test_decode_retraction_backup.py registered/unit/mem_cache/test_retraction_mamba_backup.py registered/scheduler/test_retract_decode.py registered/disaggregation/test_disaggregation_basic.py registered/disaggregation/test_kimi_linear_pd_dcp4.py registered/models_e2e/test_qwen3_next_models.py Retraction backup and restore now carry the recurrent state, and |
|
Results for 🚀 🚀 🚀 🚀 🚀 |
|
/rerun-failed-ci |
|
/rerun-test registered/disaggregation/test_kimi_linear_pd_dcp4.py |
|
Results for 🚀 |
…r part, sgl-project#36638) into the B1 staging line Agent B: dispatched requests are aborted on handler failure/disconnect (except BaseException as upstream), the waiter holds the ReqState from construction (no KeyError for batch requests). Not taken with reason: abort_sent dedup, rid filter in the disconnect task (would reopen weg2xsn276). sgl-project#30986 already covered by the fork, sgl-project#35957 unreachable, sgl-project#37143 not applicable (timeouts -1). Tests: test_tokenizer_manager_rid_cleanup 21 passed (hermetic, cgroup 3G).
Motivation
A retracted decode request on a model that has both sliding-window attention and recurrent state resumes on someone else's recurrent state.
Req.offload_kv_cachehandsmamba_indicesto the KV pool, andSWAKVPool.get_cpu_copyaccepts that argument and ignores it, so only the full and sliding-window components travel. The per-request slot is released on retraction and can be handed to another request, so the resumed request continues from whatever that slot now holds.HybridLinearKVPool.get_cpu_copydoes move it, so the two pools disagree behind one signature and the loss is silent.Under forced retraction, gsm8k over PD goes from 0.315 with 49.5% unparseable answers to 0.850 with none.
The startup path is what led here.
resolve_decode_retraction_backuppickshost_poolfor a hybrid SWA+SSM model and_create_unified_radix_cachethen refuses that combination, so the decode role exits withValueError: Host-pool retraction does not support Mamba models. Host-pool retraction transfers full and sliding-window components only, so the refusal stands and it is the inference that has to agree with it. Both sides came in together in #34801: the inference looks at sliding-window attention while the refusal looks at recurrent state, and a model with both falls between them. One shared predicate keeps a future state type from landing in one without the other.The blast radius is narrow, which is why the pair went unnoticed. A pure attention model resolves
host_pooland is accepted; a linear-attention model like Qwen3-Next or Kimi-Linear presents aHybridLinearKVPool, matches neither branch, falls tocpu_tensor, and carries its own state. Only a model with both traits both fails to start and, once started, loses state.Modifications
RetractionBackup.mamba_cpucarries the recurrent state when the KV pool does not.KVCache.cpu_copy_carries_mambasays which pools move it themselves;HybridLinearKVPoolsets it.uses_ssm_stateis shared by the retraction-backend inference andbuild_kv_cache, keeping those models oncpu_tensor.Accuracy
Inkling-Small over PD on 4xB200, one prefill and one decode role at TP=2 each, MXFP8 KV,
SGLANG_TEST_RETRACT=truewith interval 3, gsm8k 200 questions 10-shot. The two arms differ only in whether the recurrent state is backed up.0.850 matches what this pair reads without forced retraction (0.845-0.865). Half the answers unparseable is the signature of lost recurrent state rather than a drift in precision.
A bit-exact comparison across a retraction is not available on this model: a resumed request recomputes its boundary token through the extend path while an uninterrupted run computed it through the decode path, and the two sconv kernels are not bit-identical. Verified instead by checksumming each transferred component at backup and at restore, over the same physical slots -- full KV, the in-window sliding-window rows, and the recurrent state all round-trip identically once the state travels.
The pair test in #35840 covers this configuration.
CI States
Latest PR Test (Base): ❌ Run #32685573950
Latest PR Test (Extra): ✅ Run #32685573792
Latest PR Test (AMD ROCm 7.2): ⏳ Run #32685573957