Skip to content

Fix buffer-mode idle tracking and VLM memory sizing - #37567

Merged
merrymercy merged 6 commits into
mainfrom
fix/buffer-idle-vlm-memory-reserve
Sep 3, 2026
Merged

merrymercy merged 6 commits into
mainfrom
fix/buffer-idle-vlm-memory-reserve

Conversation

@merrymercy

@merrymercy merrymercy commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

Buffer-only HiCache can report idle while write-through or load-back work is still in flight, allowing lifecycle transitions before staging activity drains. Post-capture KV sizing also needs fixed runtime headroom for multimodal execution instead of the standard capacity-scaled heuristic.

Modifications

  • Include write-through and buffer load-back operations in buffer-pipeline idle detection.
  • Reserve an additional fixed 8 GiB for multimodal runtime allocations when post-capture KV sizing is active; preserve the existing heuristic outside that mode and exclude decode-only workers.
  • Add focused regression coverage for both behaviors.
  • Reuse the current tree snapshot path, which already carries storage sharding context.

Accuracy Tests

Not applicable; this does not change model numerics.

  • LD_PRELOAD=/usr/lib64/libnuma.so.1 uv run --no-sync pytest -q test/registered/unit/server_args/test_server_args.py test/registered/unit/mem_cache/test_unified_radix_cache_unittest.py::TestBufferModePipelineIdle — 193 passed, 20 subtests passed.

Speed Tests and Profiling

Not applicable; no inference hot path is changed.

Original commits

  • 80ee123f2
  • f91e7de8f
  • 71238b1c6

Checklist

  • Format code with pre-commit.
  • Add unit tests.
  • Documentation is not required for this internal behavior change.
  • Accuracy and speed benchmarks are not applicable.
  • Follow the SGLang code style guidance.

Review and Merge Process

Please follow the standard SGLang review and merge process.


CI States

Latest PR Test (Base): 🚫 Run #33789910520
Latest PR Test (Extra): ❌ Run #33789910083
Latest PR Test (AMD ROCm 7.2): 🚫 Run #33789910446

Track every in-flight buffer-mode transfer before reporting the pipeline idle, and reserve fixed runtime headroom for multimodal models when post-capture KV sizing is active.

Co-authored-by: Yongji Wu <30348494+libertyeagle@users.noreply.github.com>

Co-authored-by: Lianmin Zheng <lianminzheng@gmail.com>
@merrymercy

Copy link
Copy Markdown
Contributor Author

/tag-and-rerun-ci

@github-actions github-actions Bot added hicache Hierarchical Caching for SGLang run-ci CI: run the baseline test suite on this PR labels Sep 2, 2026

@merrymercy merrymercy left a comment

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

approve

@merrymercy

Copy link
Copy Markdown
Contributor Author

/rerun-test test/registered/mem_cache/test_post_capture_kv_sizing.py::TestPostCaptureKVSizing test/registered/hicache/test_hicache_storage_runtime_attach_detach.py::TestUnifiedRadixCacheStorageRuntimeAttachDetach.test_runtime_attach_detach

@github-actions

github-actions Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/mem_cache/test_post_capture_kv_sizing.py::TestPostCaptureKVSizing test/registered/hicache/test_hicache_storage_runtime_attach_detach.py::TestUnifiedRadixCacheStorageRuntimeAttachDetach.test_runtime_attach_detach:

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/mem_cache/test_post_capture_kv_sizing.py TestPostCaptureKVSizing

🚀 2-gpu-h100 (1 test): ❌ View workflow run

cd test/ && python3 registered/hicache/test_hicache_storage_runtime_attach_detach.py TestUnifiedRadixCacheStorageRuntimeAttachDetach.test_runtime_attach_detach

Comment thread python/sglang/srt/arg_groups/memory_hook.py
Comment thread python/sglang/srt/arg_groups/memory_hook.py
@merrymercy

Copy link
Copy Markdown
Contributor Author

/rerun-test test/registered/mem_cache/test_post_capture_kv_sizing.py::TestPostCaptureKVSizing test/registered/hicache/test_hicache_storage_runtime_attach_detach.py::TestUnifiedRadixCacheStorageRuntimeAttachDetach.test_runtime_attach_detach

@github-actions

github-actions Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/mem_cache/test_post_capture_kv_sizing.py::TestPostCaptureKVSizing test/registered/hicache/test_hicache_storage_runtime_attach_detach.py::TestUnifiedRadixCacheStorageRuntimeAttachDetach.test_runtime_attach_detach:

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/mem_cache/test_post_capture_kv_sizing.py TestPostCaptureKVSizing

🚀 2-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/hicache/test_hicache_storage_runtime_attach_detach.py TestUnifiedRadixCacheStorageRuntimeAttachDetach.test_runtime_attach_detach

Comment thread python/sglang/srt/arg_groups/memory_hook.py Outdated
@merrymercy

Copy link
Copy Markdown
Contributor Author

/rerun-test test/registered/mem_cache/test_post_capture_kv_sizing.py::TestPostCaptureKVSizing test/registered/hicache/test_hicache_storage_runtime_attach_detach.py::TestUnifiedRadixCacheStorageRuntimeAttachDetach.test_runtime_attach_detach

@github-actions

github-actions Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/mem_cache/test_post_capture_kv_sizing.py::TestPostCaptureKVSizing test/registered/hicache/test_hicache_storage_runtime_attach_detach.py::TestUnifiedRadixCacheStorageRuntimeAttachDetach.test_runtime_attach_detach:

🚀 1-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/mem_cache/test_post_capture_kv_sizing.py TestPostCaptureKVSizing

🚀 2-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/hicache/test_hicache_storage_runtime_attach_detach.py TestUnifiedRadixCacheStorageRuntimeAttachDetach.test_runtime_attach_detach

Comment thread python/sglang/srt/arg_groups/memory_hook.py
@merrymercy

Copy link
Copy Markdown
Contributor Author

/rerun-test test/registered/vlm/test_vision_openai_server_a.py::TestLlavaServer.test_single_image_chat_completion test/registered/vlm/test_vision_chunked_prefill.py::TestVisionChunkedPrefill.test_chunked_prefill test/registered/hicache/test_hicache_storage_file_backend.py::TestHiCacheStoragePageFirstDirectIO.test_basic_backup_and_prefetch

@github-actions

github-actions Bot commented Sep 3, 2026 •

Copy link
Copy Markdown
Contributor

Results for /rerun-test test/registered/vlm/test_vision_openai_server_a.py::TestLlavaServer.test_single_image_chat_completion test/registered/vlm/test_vision_chunked_prefill.py::TestVisionChunkedPrefill.test_chunked_prefill test/registered/hicache/test_hicache_storage_file_backend.py::TestHiCacheStoragePageFirstDirectIO.test_basic_backup_and_prefetch:

🚀 1-gpu-h100 (2 tests): ✅ View workflow run

cd test/ && python3 registered/vlm/test_vision_openai_server_a.py TestLlavaServer.test_single_image_chat_completion
cd test/ && python3 registered/vlm/test_vision_chunked_prefill.py TestVisionChunkedPrefill.test_chunked_prefill

🚀 2-gpu-h100 (1 test): ✅ View workflow run

cd test/ && python3 registered/hicache/test_hicache_storage_file_backend.py TestHiCacheStoragePageFirstDirectIO.test_basic_backup_and_prefetch

Comment thread python/sglang/srt/arg_groups/memory_hook.py
@merrymercy
merrymercy merged commit 05dbe64 into main Sep 3, 2026
108 of 132 checks passed
@merrymercy
merrymercy deleted the fix/buffer-idle-vlm-memory-reserve branch September 3, 2026 20:56
StevenChenSE pushed a commit to StevenChenSE/sglang that referenced this pull request Sep 6, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

hicache Hierarchical Caching for SGLang run-ci CI: run the baseline test suite on this PR

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant