Skip to content

Fix recurrent state loss on decode retraction - #35957

Merged
ispobock merged 4 commits into
mainfrom
fix-host-pool-retraction-inference
Aug 24, 2026
Merged

ispobock merged 4 commits into
mainfrom
fix-host-pool-retraction-inference

Conversation

@ispobock

@ispobock ispobock commented Aug 22, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

A retracted decode request on a model that has both sliding-window attention and recurrent state resumes on someone else's recurrent state. Req.offload_kv_cache hands mamba_indices to the KV pool, and SWAKVPool.get_cpu_copy accepts that argument and ignores it, so only the full and sliding-window components travel. The per-request slot is released on retraction and can be handed to another request, so the resumed request continues from whatever that slot now holds. HybridLinearKVPool.get_cpu_copy does move it, so the two pools disagree behind one signature and the loss is silent.

Under forced retraction, gsm8k over PD goes from 0.315 with 49.5% unparseable answers to 0.850 with none.

The startup path is what led here. resolve_decode_retraction_backup picks host_pool for a hybrid SWA+SSM model and _create_unified_radix_cache then refuses that combination, so the decode role exits with ValueError: Host-pool retraction does not support Mamba models. Host-pool retraction transfers full and sliding-window components only, so the refusal stands and it is the inference that has to agree with it. Both sides came in together in #34801: the inference looks at sliding-window attention while the refusal looks at recurrent state, and a model with both falls between them. One shared predicate keeps a future state type from landing in one without the other.

The blast radius is narrow, which is why the pair went unnoticed. A pure attention model resolves host_pool and is accepted; a linear-attention model like Qwen3-Next or Kimi-Linear presents a HybridLinearKVPool, matches neither branch, falls to cpu_tensor, and carries its own state. Only a model with both traits both fails to start and, once started, loses state.

Modifications

  • RetractionBackup.mamba_cpu carries the recurrent state when the KV pool does not.
  • KVCache.cpu_copy_carries_mamba says which pools move it themselves; HybridLinearKVPool sets it.
  • uses_ssm_state is shared by the retraction-backend inference and build_kv_cache, keeping those models on cpu_tensor.

Accuracy

Inkling-Small over PD on 4xB200, one prefill and one decode role at TP=2 each, MXFP8 KV, SGLANG_TEST_RETRACT=true with interval 3, gsm8k 200 questions 10-shot. The two arms differ only in whether the recurrent state is backed up.

without the backup   Accuracy 0.315   Invalid 0.495
with the backup      Accuracy 0.850   Invalid 0.000

0.850 matches what this pair reads without forced retraction (0.845-0.865). Half the answers unparseable is the signature of lost recurrent state rather than a drift in precision.

A bit-exact comparison across a retraction is not available on this model: a resumed request recomputes its boundary token through the extend path while an uninterrupted run computed it through the decode path, and the two sconv kernels are not bit-identical. Verified instead by checksumming each transferred component at backup and at restore, over the same physical slots -- full KV, the in-window sliding-window rows, and the recurrent state all round-trip identically once the state travels.

The pair test in #35840 covers this configuration.


CI States

Latest PR Test (Base): ❌ Run #32685573950
Latest PR Test (Extra): ✅ Run #32685573792
Latest PR Test (AMD ROCm 7.2): ⏳ Run #32685573957

@ispobock ispobock changed the title Fix host-pool retraction backend for hybrid SSM models Fix recurrent state loss on decode retraction Aug 24, 2026
@ispobock ispobock added run-ci CI: run the baseline test suite on this PR bypass-fastfail run-ci-extra CI: also run the extra suite (requires run-ci) labels Aug 24, 2026
@ispobock
ispobock marked this pull request as ready for review August 24, 2026 03:00
@ispobock

Copy link
Copy Markdown
Collaborator Author

/rerun-test registered/unit/mem_cache/test_decode_retraction_backup.py registered/unit/mem_cache/test_retraction_mamba_backup.py registered/scheduler/test_retract_decode.py registered/disaggregation/test_disaggregation_basic.py registered/disaggregation/test_kimi_linear_pd_dcp4.py registered/models_e2e/test_qwen3_next_models.py

Retraction backup and restore now carry the recurrent state, and uses_ssm_state replaced the inline hybrid-SSM check in build_kv_cache, so this covers the retraction path itself, the PD decode side, the HybridLinearKVPool branch that carries its own state, and hybrid-SSM pool construction.

@github-actions

github-actions Bot commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test registered/unit/mem_cache/test_decode_retraction_backup.py registered/unit/mem_cache/test_retraction_mamba_backup.py registered/scheduler/test_retract_decode.py registered/disaggregation/test_disaggregation_basic.py registered/disaggregation/test_kimi_linear_pd_dcp4.py registered/models_e2e/test_qwen3_next_models.py:

🚀 1-gpu-5090 (2 tests): ✅ View workflow run

cd test/ && python3 registered/unit/mem_cache/test_decode_retraction_backup.py
cd test/ && python3 registered/scheduler/test_retract_decode.py

🚀 ubuntu-latest (1 test): ✅ View workflow run

cd test/ && python3 registered/unit/mem_cache/test_retraction_mamba_backup.py

🚀 2-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_disaggregation_basic.py

🚀 8-gpu-b200 (1 test): ❌ View workflow run

cd test/ && python3 registered/disaggregation/test_kimi_linear_pd_dcp4.py

🚀 4-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/models_e2e/test_qwen3_next_models.py

@ispobock

Copy link
Copy Markdown
Collaborator Author

/rerun-failed-ci

@ispobock

Copy link
Copy Markdown
Collaborator Author

/rerun-test registered/disaggregation/test_kimi_linear_pd_dcp4.py

@github-actions

github-actions Bot commented Aug 24, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test registered/disaggregation/test_kimi_linear_pd_dcp4.py:

🚀 8-gpu-b200 (1 test): ✅ View workflow run

cd test/ && python3 registered/disaggregation/test_kimi_linear_pd_dcp4.py

@ispobock
ispobock merged commit 54ec2c4 into main Aug 24, 2026
329 of 367 checks passed
@ispobock
ispobock deleted the fix-host-pool-retraction-inference branch August 24, 2026 16:11
efschu pushed a commit to efschu/htsglang that referenced this pull request Sep 24, 2026
…r part, sgl-project#36638) into the B1 staging line

Agent B: dispatched requests are aborted on handler failure/disconnect
(except BaseException as upstream), the waiter holds the ReqState from
construction (no KeyError for batch requests). Not taken with reason:
abort_sent dedup, rid filter in the disconnect task (would reopen
weg2xsn276). sgl-project#30986 already covered by the fork, sgl-project#35957 unreachable, sgl-project#37143
not applicable (timeouts -1).
Tests: test_tokenizer_manager_rid_cleanup 21 passed (hermetic, cgroup 3G).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bypass-fastfail memory-pool run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant