Repository navigation
Conversation
Signed-off-by: luozijian <luozijian0924@gamil.com>
|
The required CI jobs are currently gated because this PR does not yet have the |
UnifiedSWAKVPool.__init__ sets mem_usage = 0.0 (unified_memory_pool.py:1482, "cosmetic; UnifiedKVPool logs the real size"), so every reader of `sglang:kv_cache_memory_usage_gb` sees zero: the metric, the server-info payload, and the scheduler's load inquirer all go through `token_to_kv_pool_allocator.get_kvcache().mem_usage` (scheduler.py:1206/4877, load_inquirer.py:143). The KV pool is this stack's only hard capacity limit, so the blind spot hid it. The authoritative byte count is already there: UnifiedKVPool takes total_bytes (summing the full/swa/mamba sub-pools) and UnifiedSWAKVPool keeps that buffer on self.unified_buffer. The adapter writes it back as mem_usage from a wrapped constructor, for whichever pool family get_kvcache() returns. Idempotent, and it logs instead of raising when a pool carries no total_bytes, so the next boot says which path it took. Lands the same number as upstream sgl-project/sglang#37935 without forking the pool classes; #37935 is not in our pinned e59e6eb. Verified on stubs (48 GiB -> mem_usage 48.0; missing total_bytes stays 0 and does not raise). Needs an image rebuild + cold boot to take effect: adapters are COPYed, not mounted.
Preserve current per-ratio pools, unified FP8 RoPE storage, and request-window allocations. Finalize indexer memory usage after backend-specific buffers are allocated. Add CPU regression coverage for current layouts and constructor memory usage. Validated with 49 focused tests plus 27 subtests, changed-file pre-commit hooks, compilation, and diff checks. AI-assisted refresh.
|
Refreshed this PR against For a small follow-up, I’d like to check the DeepSeek-V4 sizing-to-allocation contract: page/alignment overhead, per-ratio KV and indexer buffers, and compressor-state lifetimes. The first deliverable would be a buffer map and small CPU constructor tests, with behavior changes only for a reproduced mismatch. Tensor storage, VMM-reserved address space, physical backing, and occupied slots would be kept distinct. This would stay clear of the active hybrid-SWA post-capture, host-memory, and draft-budget changes in #41961, #42039, and #38203. Does that scope fit the maintainers’ priorities? This refresh and investigation use AI assistance; the validation above is automated and maintainer review is still needed. |
UnifiedSWAKVPool.__init__ sets mem_usage = 0.0 (unified_memory_pool.py:1482, "cosmetic; UnifiedKVPool logs the real size"), so every reader of `sglang:kv_cache_memory_usage_gb` sees zero: the metric, the server-info payload, and the scheduler's load inquirer all go through `token_to_kv_pool_allocator.get_kvcache().mem_usage` (scheduler.py:1206/4877, load_inquirer.py:143). The KV pool is this stack's only hard capacity limit, so the blind spot hid it. The authoritative byte count is already there: UnifiedKVPool takes total_bytes (summing the full/swa/mamba sub-pools) and UnifiedSWAKVPool keeps that buffer on self.unified_buffer. The adapter writes it back as mem_usage from a wrapped constructor, for whichever pool family get_kvcache() returns. Idempotent, and it logs instead of raising when a pool carries no total_bytes, so the next boot says which path it took. Lands the same number as upstream sgl-project/sglang#37935 without forking the pool classes; #37935 is not in our pinned e59e6eb. Verified on stubs (48 GiB -> mem_usage 48.0; missing total_bytes stays 0 and does not raise). Needs an image rebuild + cold boot to take effect: adapters are COPYed, not mounted.
Motivation
Fixes #37852.
DeepSeek-V4 allocates KV storage through several layout-specific pools, but none of the pools populated the inherited
mem_usagefield. As a result,sglang:kv_cache_memory_usage_gband the scheduler internal state reported0for a running DeepSeek-V4 server.Modifications
mainwithout rewriting the original commit or discarding upstream storage changes.mem_usagefrom allocated tensors; finalize indexer accounting after backend-specific buffers have been allocated.Tests
Linux CPU validation of the refreshed changes with Python 3.12, Torch 2.14.0+cpu, and Triton 3.8.0:
test_dsv4_memory_usage.py,test_dsv4_compressed_pools.py,test_dsv4_unified_fp8_pool.py,test_dsv4_c4_state_lifecycle.py, andtest_dsv4_compress_write_pad.py.41cbe65de0:mem_usageis zero despite allocated KV tensors.pre-commithooks passed on the two changed files, including Ruff, isort, codespell, and CI registration checks.git diff --checkpassed.GPU/NPU execution was not run locally. Standard CI execution still needs the repository's normal maintainer authorization.
AI assistance
This change was prepared and refreshed with AI assistance. The current refresh has the automated validation listed above and remains subject to maintainer review.
CI States
Latest PR Test (Base): ❌ Run #36871532610
Latest PR Test (Extra): ❌ Run #36871532099
Latest PR Test (AMD ROCm 10): ❌ Run #36871532887