feat(services): lmcache-server kind and lmcache-mp connector - #507
Conversation
96f5a53 to
19fe712
Compare
| "lmcache-mp": KVConnector( | ||
| "LMCacheMPConnector", | ||
| module_path="lmcache.integration.vllm.lmcache_mp_connector", | ||
| extra_config={"lmcache.mp.host": "tcp://localhost", "lmcache.mp.port": LMCACHE_SERVER_PORT}, |
There was a problem hiding this comment.
"tcp://localhost", "lmcache.mp.port"
are these configurable or always static?
There was a problem hiding this comment.
these are static right now
|
This overall lgtm but while we're at it we should add support for connecting to atom |
|
@sammshen SemiAnalysisAI#32 take a look here |
|
@sammshen on my fork, I added support for sglang and atom ootb. I have tested atom in-tree as well as your vllm + lmcache mp integration. here are the results
Not yet tested on hardware: SGLang |
|
actually change of plans -- let's just merge this and then enable atom and sglang in a follow up PR. |
Add services[].type: lmcache-server, which runs the LMCache multiprocess server on every worker node before workers and gates on GET /healthcheck, and connector: lmcache-mp, which points vLLM's LMCacheMPConnector at that node-local server. The kind fixes only the ports srtctl owns; recipe args are appended verbatim.
19fe712 to
68deba6
Compare
| @@ -0,0 +1,88 @@ | |||
| # An LMCache DRAM tier on the prefill side of a disaggregated vLLM job. | |||
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #507 +/- ##
=======================================
Coverage ? 83.17%
=======================================
Files ? 152
Lines ? 21914
Branches ? 0
=======================================
Hits ? 18226
Misses ? 3688
Partials ? 0 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
| "lmcache", | ||
| "server", | ||
| "--host", | ||
| "0.0.0.0", |
There was a problem hiding this comment.
Non-blocking: the connector only reaches the RPC server over tcp://localhost, and upstream's --host default is localhost (lmcache/v1/multiprocess/config.py). Binding the ZMQ port on 0.0.0.0 lets anything on the network reach the cache for no benefit. Suggest 127.0.0.1 here. --http-host (line 47) does need to stay 0.0.0.0, because the readiness probe checks http://<node>:8751/healthcheck.
| "0.0.0.0", | ||
| "--http-port", | ||
| str(LMCACHE_HTTP_PORT), | ||
| *service.args, |
There was a problem hiding this comment.
Non-blocking: the ports are fixed and the kind has no options. If someone passes --port / --http-port / --host in args, argparse keeps the last value, so the server moves. The connector's extra config and the readiness probe still point at 8750/8751, and nothing reports an error. Two things would help:
- In
validate(), reject those flags inargs, or accept them only throughoptionsso the connector and the probe can read the same value. - Refuse
placement.per: worker. It would start several servers on 8750 on the same node.
There was a problem hiding this comment.
nit / overengineering / hypotehtical situation that isn't actually feasible
|
Review summary. The inline comments cover the code-specific points. Two more blocking items that don't belong to a single line:
On testing: the thread shows a hardware run of the agg path only (vLLM + |
… through connector, set PYTHONHASHSEED on the service - KVConnector.service_type: the lmcache-mp row implies an lmcache-server on the nodes of the roles that use it (resolved via kv_connector_for_mode, so role overrides count); a declared lmcache-server entry takes over. - lmcache-server-disagg: roles.prefill.args.connector carries the MultiConnector JSON, so prefill gets one --kv-transfer-config instead of two. - Both examples set PYTHONHASHSEED in the service env; services do not inherit the top-level environment. - Launch snapshots for both examples. Co-Authored-By: ishan <idhanani@nvidia.com>
A direct vllm serve aggregate worker dropped every KV connector, so roles.agg.args.connector (for example lmcache-mp offload to the lmcache-server service) never reached --kv-transfer-config. It now gets the connector its role names; the engine-wide default stays a prefill/decode setting.
services[].type: lmcache-serverplusconnector: lmcache-mp. Full recipes inexamples/features/lmcache-server.yamlandexamples/features/lmcache-server-disagg.yaml.Aggregated: one server per worker node, every vLLM worker offloads to it.
connector: lmcache-mpimplies the service; the entry is only needed to passargs.Disaggregated, prefill only: NIXL for P/D, LMCache on the prefill side via a hand-written
MultiConnectorpassed as the role'sconnector.What it renders (
srtctl dry-run):LMCache must be installed in the job container.