Skip to content

feat(services): lmcache-server kind and lmcache-mp connector - #507

Merged
ishandhanani merged 3 commits into
NVIDIA:mainfrom
sammshen:lmcache-server
Sep 29, 2026
Merged

ishandhanani merged 3 commits into
NVIDIA:mainfrom
sammshen:lmcache-server

Conversation

@sammshen

@sammshen sammshen commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

services[].type: lmcache-server plus connector: lmcache-mp. Full recipes in examples/features/lmcache-server.yaml and examples/features/lmcache-server-disagg.yaml.

Aggregated: one server per worker node, every vLLM worker offloads to it. connector: lmcache-mp implies the service; the entry is only needed to pass args.

engine:
  type: vllm
  connector: lmcache-mp

services:
  - name: lmcache
    type: lmcache-server
    args: [--l1-size-gb, "100", --chunk-size, "256", --max-workers, "2"]

Disaggregated, prefill only: NIXL for P/D, LMCache on the prefill side via a hand-written MultiConnector passed as the role's connector.

roles:
  prefill:
    args:
      connector: '{"kv_connector":"MultiConnector","kv_role":"kv_both","kv_connector_extra_config":{"connectors":[{"kv_connector":"NixlConnector","kv_role":"kv_both"},{"kv_connector":"LMCacheMPConnector","kv_connector_module_path":"lmcache.integration.vllm.lmcache_mp_connector","kv_role":"kv_both","kv_connector_extra_config":{"lmcache.mp.host":"tcp://localhost","lmcache.mp.port":8750}}]}}'

services:
  - name: lmcache
    type: lmcache-server
    placement:
      node: prefill
    args: [--l1-size-gb, "100", --chunk-size, "256"]

What it renders (srtctl dry-run):

lmcache type=lmcache-server placement=workers start=before_workers critical=true
  command: lmcache server --host 0.0.0.0 --port 8750 --http-host 0.0.0.0 --http-port 8751 --l1-size-gb 100 --chunk-size 256 --max-workers 2
--kv-transfer-config '{"kv_connector":"LMCacheMPConnector","kv_connector_module_path":"lmcache.integration.vllm.lmcache_mp_connector","kv_role":"kv_both","kv_connector_extra_config":{"lmcache.mp.host":"tcp://localhost","lmcache.mp.port":8750}}'

LMCache must be installed in the job container.

@functionstackx functionstackx left a comment •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Hi @sammshen any chance you get ask ur claude to get this working for ATOM too or we can do this on an follow up PR? my understanding is that the ATOM folks like using LMCache too

+viz @cquil11

"lmcache-mp": KVConnector(
"LMCacheMPConnector",
module_path="lmcache.integration.vllm.lmcache_mp_connector",
extra_config={"lmcache.mp.host": "tcp://localhost", "lmcache.mp.port": LMCACHE_SERVER_PORT},

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"tcp://localhost", "lmcache.mp.port"

are these configurable or always static?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

these are static right now

@cquil11

cquil11 commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

This overall lgtm but while we're at it we should add support for connecting to atom

@cquil11

cquil11 commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

@sammshen SemiAnalysisAI#32 take a look here

@cquil11

cquil11 commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

@sammshen on my fork, I added support for sglang and atom ootb. I have tested atom in-tree as well as your vllm + lmcache mp integration. here are the results
Tested on hardware through InferenceX, applied as a patch:

Path Hardware Run Result
ATOM extra-kv-connectors (Mooncake + lmcache_offload) MI355X 1P1D, conc 256 (InferenceX#3543) 36461236354 Passed
lmcache-server + vLLM LMCacheMPConnector MI300X TP8, conc 16 (InferenceX#3545) 36460859852 Passed

Not yet tested on hardware: SGLang enable-lmcache, and ATOM lmcache_mp.

@cquil11

cquil11 commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

actually change of plans -- let's just merge this and then enable atom and sglang in a follow up PR.

Add services[].type: lmcache-server, which runs the LMCache multiprocess
server on every worker node before workers and gates on GET /healthcheck,
and connector: lmcache-mp, which points vLLM's LMCacheMPConnector at that
node-local server. The kind fixes only the ports srtctl owns; recipe args
are appended verbatim.
@@ -0,0 +1,88 @@
# An LMCache DRAM tier on the prefill side of a disaggregated vLLM job.

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

was this tested?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes

@codecov-commenter

codecov-commenter commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.50000% with 1 line in your changes missing coverage. Please review.
⚠️ Please upload report for BASE (main@f5508d5). Learn more about missing BASE report.

Files with missing lines Patch % Lines
src/srtctl/services/lmcache_server.py 94.44% 1 Missing ⚠️
Additional details and impacted files
@@           Coverage Diff           @@
##             main     #507   +/-   ##
=======================================
  Coverage        ?   83.17%           
=======================================
  Files           ?      152           
  Lines           ?    21914           
  Branches        ?        0           
=======================================
  Hits            ?    18226           
  Misses          ?     3688           
  Partials        ?        0           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

Comment thread examples/features/lmcache-server-disagg.yaml Outdated
Comment thread src/srtctl/backends/vllm.py
"lmcache",
"server",
"--host",
"0.0.0.0",

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking: the connector only reaches the RPC server over tcp://localhost, and upstream's --host default is localhost (lmcache/v1/multiprocess/config.py). Binding the ZMQ port on 0.0.0.0 lets anything on the network reach the cache for no benefit. Suggest 127.0.0.1 here. --http-host (line 47) does need to stay 0.0.0.0, because the readiness probe checks http://<node>:8751/healthcheck.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit

"0.0.0.0",
"--http-port",
str(LMCACHE_HTTP_PORT),
*service.args,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Non-blocking: the ports are fixed and the kind has no options. If someone passes --port / --http-port / --host in args, argparse keeps the last value, so the server moves. The connector's extra config and the readiness probe still point at 8750/8751, and nothing reports an error. Two things would help:

  • In validate(), reject those flags in args, or accept them only through options so the connector and the probe can read the same value.
  • Refuse placement.per: worker. It would start several servers on 8750 on the same node.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit / overengineering / hypotehtical situation that isn't actually feasible

Comment thread examples/features/lmcache-server.yaml Outdated
@devin-ai-integration

Copy link
Copy Markdown
Contributor

Review summary. The inline comments cover the code-specific points. Two more blocking items that don't belong to a single line:

  1. Launch snapshots are missing. This branch is based on f5508d5, from before chore: agent harness - split CLAUDE.md, REVIEW.md, skills, design-rule/doc/launch-snapshot checks, blocking ty #526. After merging main locally, make check passes lint, ty and schema-docs, and 3345 tests pass, but 3 tests fail in tests/test_launch_snapshots.py: there is no snapshot for either new example. Please rebase, run make snapshots, and commit tests/snapshots/launch/features__lmcache-server.txt and features__lmcache-server-disagg.txt.

  2. Cite the upstream source (REVIEW.md: "Upstream behavior is assumed, not cited"). I checked the behavior against the lmcache==0.5.5 wheel and it matches:

    • the lmcache = lmcache.cli.main:main entry point;
    • --host/--port/--http-host/--http-port, with defaults localhost/5555/0.0.0.0/8080, in lmcache/v1/multiprocess/config.py;
    • GET /healthcheck, which returns 503 until the engine is initialized (lmcache/v1/multiprocess/http_apis/info_api.py);
    • the lmcache.mp.host/lmcache.mp.port keys (lmcache/integration/vllm/lmcache_mp_connector.py).

    Please name the LMCache version the vllm-lmcache image ships and link these files in the description.

On testing: the thread shows a hardware run of the agg path only (vLLM + LMCacheMPConnector, MI300X TP8). I don't see a run of lmcache-server-disagg.yaml. Please either run it or mark it untested in the description.

devin-ai-integration Bot and others added 2 commits September 29, 2026 05:05
… through connector, set PYTHONHASHSEED on the service

- KVConnector.service_type: the lmcache-mp row implies an lmcache-server on the nodes of the roles that use it (resolved via kv_connector_for_mode, so role overrides count); a declared lmcache-server entry takes over.
- lmcache-server-disagg: roles.prefill.args.connector carries the MultiConnector JSON, so prefill gets one --kv-transfer-config instead of two.
- Both examples set PYTHONHASHSEED in the service env; services do not inherit the top-level environment.
- Launch snapshots for both examples.

Co-Authored-By: ishan <idhanani@nvidia.com>
A direct vllm serve aggregate worker dropped every KV connector, so
roles.agg.args.connector (for example lmcache-mp offload to the
lmcache-server service) never reached --kv-transfer-config. It now gets
the connector its role names; the engine-wide default stays a
prefill/decode setting.
@ishandhanani
ishandhanani merged commit 6adea77 into NVIDIA:main Sep 29, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

5 participants