Skip to content

[mem_cache][11/N] refactor: extract KVCache and BaseSWAKVPool into pool/base.py - #35647

Open
alphabetc1 wants to merge 2 commits into
sgl-project:mainfrom
alphabetc1:refactor/mem-cache-pool-base
Open

alphabetc1 wants to merge 2 commits into
sgl-project:mainfrom
alphabetc1:refactor/mem-cache-pool-base

Conversation

@alphabetc1

@alphabetc1 alphabetc1 commented Aug 20, 2026 •

Copy link
Copy Markdown
Collaborator

Motivation

Part of #25371. Mechanical Move. First PR of the mem_cache/pool/ layer.

pool/ is the last of the three layer packages and the only one still getting worse.
Since the restructure was agreed, memory_pool.py went from 2257 to 5042 lines and from
11 to 17 classes, because there was nowhere else for a device pool to go. This creates
the package and moves the base layer into it, so every later family move has a target.

Modifications

New mem_cache/pool/base.py, holding only what every device pool derives from:

Symbol From
KVWriteLoc, unwrap_write_loc memory_pool.py
KvBufferDesc memory_pool.py
KVCache (ABC) memory_pool.py
BaseSWAKVPool (ABC) base_swa_memory_pool.py (deleted)
GB memory_pool.py

base_swa_memory_pool.py is deleted: it was a 29-line file holding one cross-family
ABC, sitting under the SWA feature's name while SWAKVPool, DeepSeekV4TokenToKVPool
and the disagg paths all depend on it. It belongs next to KVCache.

No re-export shim: all 44 call sites are updated in this PR, matching how allocator.py
was deleted outright rather than left forwarding.

RadixAttention moves under TYPE_CHECKING in the new module. It is only ever an
annotation (layer: RadixAttention), and the file has from __future__ import annotations; keeping it at runtime made pool/base.py -> layers.radix_attention ->
... -> forward_context -> pool/base.py a genuine import cycle, because pool/base.py
is imported far earlier in the chain than memory_pool.py was. This mirrors how
LayerDoneCounter was already handled.

Two now-unused imports (abc, KVWriteLoc) drop out of memory_pool.py. Both remaining
KVWriteLoc mentions there are in comments.

Accuracy Test

Mechanical move, so the bar is byte-level equality:

  • The KVWriteLoc / unwrap_write_loc / KvBufferDesc / KVCache block is
    byte-identical to the original memory_pool.py region.
  • The BaseSWAKVPool class body is byte-identical to base_swa_memory_pool.py.
  • Only the new module's import header is hand-written.

Verified on an H200 devbox (PYTHONPATH shadowing the image's copy):

              PR branch:   1746 passed, 1074 skipped, 201 subtests passed
  pristine main baseline:  1746 passed, 1074 skipped, 201 subtests passed

Identical command (test/registered/unit/mem_cache/ plus
spec/test_resolve_swa_kv_pool.py, --ignore on test_umbp_store.py which needs the
mori package that is absent from the image) run against both trees on the same box --
zero delta, so the 1074 skips are pre-existing and none were introduced here. MRO was
asserted intact (KVCache in MHATokenToKVPool.__mro__,
BaseSWAKVPool in SWAKVPool.__mro__), and the re-run after the ruff import cleanup gave
the same counts.

Benchmark and Profiling Results

Not applicable -- no runtime behavior changes.

Checklist

  • Format your code according to the Code Formatting with Pre-Commit.
  • Add unit tests as outlined in the Running Unit Tests.
  • Update documentation / docstrings / example tutorials as needed, according to Writing Documentation.
  • Provide throughput / latency benchmark results and accuracy evaluation results as needed, according to Benchmark and Profiling.
  • For reviewers: If you haven't made any contributions to this PR and are only assisting with merging the main branch, please remove yourself as a co-author when merging the PR.
  • Please feel free to join our Slack channel at https://slack.sglang.ai to discuss your PR.

Independent of the other in-flight #25371 PRs -- touches no file that #35306, #35643 or
#35644 touch. Once this lands, #35638 can enable its pool role and move the device
pools from its _SHRINKING_MODULES pin into _LEGACY_HOMES.

🤖 Generated with Claude Code


CI States

Latest PR Test (Base): ❌ Run #34434260912
Latest PR Test (Extra): ✅ Run #34439738060
Latest PR Test (AMD ROCm 10): ❌ Run #34434260836

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: e6bbc4efc0

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

Comment on lines +14 to 16
from sglang.srt.mem_cache.pool.base import (
unwrap_write_loc,
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Update the mocked module tree for the new pool import

When test/manual/minimax_m3/test_npu_memory_pool.py runs _load_npu_memory_pool_module(), it stubs sglang.srt.mem_cache as a plain ModuleType and provides only the old memory_pool.unwrap_write_loc. Executing this newly added import therefore raises ModuleNotFoundError: 'sglang.srt.mem_cache' is not a package before any of the standalone NPU pool tests run. The loader needs to stub sglang.srt.mem_cache.pool.base and its unwrap_write_loc symbol, or otherwise load the real package.

Useful? React with 👍 / 👎.

@alphabetc1 alphabetc1 added run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci) labels Aug 20, 2026
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@alphabetc1
alphabetc1 force-pushed the refactor/mem-cache-pool-base branch from e6bbc4e to 2e551e7 Compare September 2, 2026 10:51
@github-actions github-actions Bot added hicache Hierarchical Caching for SGLang memory-pool labels Sep 2, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

apple-silicon blackwell SM100/SM120 bypass-fastfail deepseek hicache Hierarchical Caching for SGLang memory-pool mthreads npu run-ci CI: run the baseline test suite on this PR run-ci-extra CI: also run the extra suite (requires run-ci)

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant