Skip to content

fix: report DeepSeek-V4 KV cache memory usage - #37935

Open
LOGO127 wants to merge 2 commits into
sgl-project:mainfrom
LOGO127:codex/sglang-37852-kv-mem-usage
Open

LOGO127 wants to merge 2 commits into
sgl-project:mainfrom
LOGO127:codex/sglang-37852-kv-mem-usage

Conversation

@LOGO127

@LOGO127 LOGO127 commented Sep 4, 2026 •

Copy link
Copy Markdown

Motivation

Fixes #37852.

DeepSeek-V4 allocates KV storage through several layout-specific pools, but none of the pools populated the inherited mem_usage field. As a result, sglang:kv_cache_memory_usage_gb and the scheduler internal state reported 0 for a running DeepSeek-V4 server.

Modifications

  • Add tensor-backed byte accounting to the single KV, indexer, and unified KV pools.
  • Count all active per-ratio KV/indexer pools and compression-state storage, including low-ratio layouts, unified FP8 RoPE buffers, and request-window state/workspace storage.
  • Merge current upstream main without rewriting the original commit or discarding upstream storage changes.
  • Initialize per-pool and top-level mem_usage from allocated tensors; finalize indexer accounting after backend-specific buffers have been allocated.

Tests

Linux CPU validation of the refreshed changes with Python 3.12, Torch 2.14.0+cpu, and Triton 3.8.0:

  • 49 tests and 27 subtests passed across test_dsv4_memory_usage.py, test_dsv4_compressed_pools.py, test_dsv4_unified_fp8_pool.py, test_dsv4_c4_state_lifecycle.py, and test_dsv4_compress_write_pad.py.
  • The nine-test accounting module also passed when executed directly, as in the registered CPU suite.
  • The real-constructor regression fails on unchanged upstream 41cbe65de0: mem_usage is zero despite allocated KV tensors.
  • All applicable pre-commit hooks passed on the two changed files, including Ruff, isort, codespell, and CI registration checks.
  • Compilation and git diff --check passed.

GPU/NPU execution was not run locally. Standard CI execution still needs the repository's normal maintainer authorization.

AI assistance

This change was prepared and refreshed with AI assistance. The current refresh has the automated validation listed above and remains subject to maintainer review.


CI States

Latest PR Test (Base): ❌ Run #36871532610
Latest PR Test (Extra): ❌ Run #36871532099
Latest PR Test (AMD ROCm 10): ❌ Run #36871532887

Signed-off-by: luozijian <luozijian0924@gamil.com>
@LOGO127

LOGO127 commented Sep 4, 2026

Copy link
Copy Markdown
Author

The required CI jobs are currently gated because this PR does not yet have the run-ci label. I verified the change locally with targeted tensor-accounting checks, formatting, compilation, and diff validation; the repository's full pytest path is Linux/Triton-dependent and cannot be collected on Windows. Could a maintainer please add the standard CI label or start the normal CI flow for this PR? Thank you.

ntxf31415 pushed a commit to ntxf31415/DeepSeek-v4.1-Flash-DGX-Sparks that referenced this pull request Sep 15, 2026
UnifiedSWAKVPool.__init__ sets mem_usage = 0.0 (unified_memory_pool.py:1482,
"cosmetic; UnifiedKVPool logs the real size"), so every reader of
`sglang:kv_cache_memory_usage_gb` sees zero: the metric, the server-info
payload, and the scheduler's load inquirer all go through
`token_to_kv_pool_allocator.get_kvcache().mem_usage` (scheduler.py:1206/4877,
load_inquirer.py:143). The KV pool is this stack's only hard capacity limit,
so the blind spot hid it.

The authoritative byte count is already there: UnifiedKVPool takes total_bytes
(summing the full/swa/mamba sub-pools) and UnifiedSWAKVPool keeps that buffer
on self.unified_buffer. The adapter writes it back as mem_usage from a wrapped
constructor, for whichever pool family get_kvcache() returns. Idempotent, and
it logs instead of raising when a pool carries no total_bytes, so the next boot
says which path it took.

Lands the same number as upstream sgl-project/sglang#37935 without forking the
pool classes; #37935 is not in our pinned e59e6eb.

Verified on stubs (48 GiB -> mem_usage 48.0; missing total_bytes stays 0 and
does not raise). Needs an image rebuild + cold boot to take effect: adapters
are COPYed, not mounted.
Preserve current per-ratio pools, unified FP8 RoPE storage, and request-window allocations. Finalize indexer memory usage after backend-specific buffers are allocated.

Add CPU regression coverage for current layouts and constructor memory usage. Validated with 49 focused tests plus 27 subtests, changed-file pre-commit hooks, compilation, and diff checks.

AI-assisted refresh.

LOGO127 commented Oct 1, 2026

Copy link
Copy Markdown
Author

Refreshed this PR against 41cbe65de0; current head is dd503a8f91. Repository lint passes, and local Linux CPU validation passes 49 tests plus 27 subtests. Runtime CI still needs the normal maintainer authorization (run-ci); GPU/NPU execution has not been validated locally.

For a small follow-up, I’d like to check the DeepSeek-V4 sizing-to-allocation contract: page/alignment overhead, per-ratio KV and indexer buffers, and compressor-state lifetimes. The first deliverable would be a buffer map and small CPU constructor tests, with behavior changes only for a reproduced mismatch. Tensor storage, VMM-reserved address space, physical backing, and occupied slots would be kept distinct. This would stay clear of the active hybrid-SWA post-capture, host-memory, and draft-budget changes in #41961, #42039, and #38203.

Does that scope fit the maintainers’ priorities? This refresh and investigation use AI assistance; the validation above is automated and maintainer review is still needed.

ntxf31415 added a commit to ntxf31415/DeepSeek-v4.1-Flash-DGX-Sparks that referenced this pull request Oct 4, 2026
UnifiedSWAKVPool.__init__ sets mem_usage = 0.0 (unified_memory_pool.py:1482,
"cosmetic; UnifiedKVPool logs the real size"), so every reader of
`sglang:kv_cache_memory_usage_gb` sees zero: the metric, the server-info
payload, and the scheduler's load inquirer all go through
`token_to_kv_pool_allocator.get_kvcache().mem_usage` (scheduler.py:1206/4877,
load_inquirer.py:143). The KV pool is this stack's only hard capacity limit,
so the blind spot hid it.

The authoritative byte count is already there: UnifiedKVPool takes total_bytes
(summing the full/swa/mamba sub-pools) and UnifiedSWAKVPool keeps that buffer
on self.unified_buffer. The adapter writes it back as mem_usage from a wrapped
constructor, for whichever pool family get_kvcache() returns. Idempotent, and
it logs instead of raising when a pool carries no total_bytes, so the next boot
says which path it took.

Lands the same number as upstream sgl-project/sglang#37935 without forking the
pool classes; #37935 is not in our pinned e59e6eb.

Verified on stubs (48 GiB -> mem_usage 48.0; missing total_bytes stays 0 and
does not raise). Needs an image rebuild + cold boot to take effect: adapters
are COPYed, not mounted.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] sglang:kv_cache_memory_usage_gb reports 0 for DeepSeek-V4 pools

1 participant