[mem_cache] Drop the torch.unique sync from the SWA page expansion - #37463
Conversation
|
/tag-and-rerun-ci |
|
/rerun-test registered/unit/mem_cache/test_swa_unittest.py registered/unit/mem_cache/test_swa_eviction_boundary.py registered/radix_cache/test_swa_radix_cache_kl.py registered/radix_cache/unified_radix_tree/test_unified_radix_cache_kl_swa.py registered/sessions/test_streaming_session_swa.py registered/sessions/test_streaming_session_swa_extra.py registered/chunked_prefill/test_scripted_swa_1gpu.py |
|
Results for 🚀 🚀 |
|
The commit containing this change passed all the basic CI: https://github.com/sgl-project/sglang/actions/runs/33547072775 |
free_swa's whole-page coverage withouttorch.unique(its data-dependent output shape synchronizes the scheduler stream): keep the duplicate page entries and let the paged free's own page dedup collapse them. End state is unchanged; adebug_modereference check compares against the unique-based page set on every real input.CI States
Latest PR Test (Base): 🚫 Run #33545635926
Latest PR Test (Extra): ❌ Run #33545635640
Latest PR Test (AMD ROCm 7.2): 🚫 Run #33545635902