Repository navigation
feat(offload): enable LMCache offload via ATOM_KV_OFFLOAD env - #2414
Merged
Merged
Conversation
…onfig
LMCache offload is enabled through --kv-transfer-config, which launchers such
as srtctl own for P/D transfer and refuse to let a recipe override. Add
--kv-offload-config (a JSON dict, like --dspark-config / --dcp-config) and fold
it into kv_transfer_config in EngineArgs._get_engine_kwargs:
- alone: {"kv_connector":"lmcache_offload","kv_role":"offload"} plus any extra
keys (e.g. "lmcache.chunk_size");
- with a P/D connector: both behind a "multi" connector (appended when the
transfer config is already "multi");
- rejected when --kv-transfer-config already names an offload connector.
The engine path after argument parsing is unchanged.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Contributor
🏷️ CI GuideRuns automatically on every eligible PR before approval:
Heavy model tests:
|
Replace the --kv-offload-config CLI flag with environment variables. Launchers such as srtctl reserve --kv-transfer-config and serialize recipe args with str(), but pass worker env through untouched, so an env var is the robust way to add offload there. - ATOM_KV_OFFLOAD=lmcache selects the in-process lmcache_offload connector; lmcache_mp selects the standalone-server lmcache_mp connector. Unset/0/off/none keeps offload off. - ATOM_KV_OFFLOAD_EXTRA_CONFIG (JSON object) becomes the connector's kv_connector_extra_config (lmcache.<field>, lmcache.mp.*). - Composition with --kv-transfer-config is unchanged: a P/D connector and the offload connector run behind "multi", and a second offload connector is rejected. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com> Signed-off-by: yihonglie <hyi@amd.com>
valarLip
approved these changes
Sep 28, 2026
valarLip
approved these changes
Sep 28, 2026
yhl-amd
pushed a commit
that referenced
this pull request
Sep 28, 2026
Resolves docs/environment_variables.md: #2414's ATOM_KV_OFFLOAD and ATOM_KV_OFFLOAD_EXTRA_CONFIG rows join the rewritten LMCache offload table. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
yhl-amd
added a commit
to SemiAnalysisAI/InferenceX
that referenced
this pull request
Sep 28, 2026
Move the MiniMax-M3 MI355X ATOM AgentX recipe to rocm/atom-dev:nightly_202609281543, which contains ROCm/ATOM#2414, and restore the TP2 C20/C25/C30 and TP4 C40/C48 LMCache DRAM-offload arms that #3463 dropped. srtctl reserves ATOM's kv-transfer-config, so each DRAM variant enables the in-process lmcache_offload connector with ATOM_KV_OFFLOAD=lmcache and carries the legacy script's LMCache settings. 中文:将 MiniMax-M3 MI355X ATOM AgentX 配方镜像更新为包含 ROCm/ATOM#2414 的 rocm/atom-dev:nightly_202609281543,并恢复 #3463 删除的 TP2 C20/C25/C30 与 TP4 C40/C48 LMCache DRAM 卸载臂。srtctl 占用 ATOM 的 kv-transfer-config, 因此各 DRAM 变体通过 ATOM_KV_OFFLOAD=lmcache 启用进程内 lmcache_offload 连接器,并沿用旧脚本的 LMCache 设置。 Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
10 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Motivation
LMCache KV offload is enabled only through
--kv-transfer-config. Launchers that own that flag cannot add offload. For example, srtctl (NVIDIA/srt-slurm) generates--kv-transfer-configfor P/D workers and rejects it in recipe args. As a result, single-node ATOM AgentX recipes on srt-slurm (MiniMax-M3, Kimi-K3, GLM-5.2 on MI355X) had to drop all of their LMCache offload points.srtctl passes each role's
envto the worker unchanged. It also serializes recipe args withstr(), so a JSON-valued flag is fragile there. An environment variable avoids both problems.Change
Add two environment variables, registered in
atom/utils/envs.py.EngineArgs._get_engine_kwargsfolds them intokv_transfer_config, so nothing changes after argument parsing:ATOM_KV_OFFLOADkv_transfer_config0/off/nonelmcache{"kv_connector":"lmcache_offload","kv_role":"offload"}(in-process)lmcache_mp{"kv_connector":"lmcache_mp","kv_role":"offload"}(standalonelmcache server)--kv-transfer-config{"kv_connector":"multi","connectors":[<pd>, <offload>]}--kv-transfer-configalreadymulticonnectors--kv-transfer-configalready naming an offload connector (name or alias)ValueErrorATOM_KV_OFFLOAD_EXTRA_CONFIGtakes a JSON object and becomes the offload connector'skv_connector_extra_config. Both backends read it:lmcache.<field>overrides and, forlmcache_mp,lmcache.mp.*options such aslmcache.mp.port. An unknown mode, a non-object extra config, or an extra config without a mode raises an error. When offload is enabled, the resultingkv_transfer_configis logged at startup.The first revision of this PR used a
--kv-offload-configCLI flag. That flag is removed; the environment variables replace it. Without them, behavior is unchanged.docs/environment_variables.md, the offload README and the module docstring document the variables.Test
tests/test_arg_utils_spec.py: 17 environment-variable cases covering every row above, case-insensitive modes, extra-config pass-through, invalid input, and confirming--kv-offload-configis gone. Withtests/test_envs.pyandtests/entrypoints/test_metrics.py, 112 tests pass locally (CPU torch).blackis clean on the changed files.ruff checkreports only the two E402 imports thattests/test_arg_utils_spec.pyalready had.rocm/atom-dev:nightly_202609231534with this branch mounted), Llama-3.1-8B-Instruct TP1. The test sends a ~23k-token prompt, sends enough distinct prompts to evict it from the GPU prefix cache, then sends it again:ATOM_KV_OFFLOAD=lmcache: startup logs the composedkv_transfer_configandKV transfer config loaded: {'kv_connector': 'lmcache_offload', ...}. The resend loads 22,784 tokens from LMCache (atom:lmcache_loaded_tokens,atom:prefix_cache_offload_tokens) with zero GPU prefix hits. Output matches the first response, and there are no errors.ATOM_KV_OFFLOADunset: no offload connector, and everyatom:lmcache_*counter stays at 0.ATOM_KV_OFFLOAD=lmcache_mpselectslmcache_mpand passesATOM_KV_OFFLOAD_EXTRA_CONFIGthrough (lmcache.mp.port). On DeepSeek-R1-0528 TP8 it registers all ranks and saves to thelmcache server. MP data-plane reuse is out of scope for this PR: lookups there fail in the existing TP-collapse path (No GPU context found ... world size 1), andlmcache_mprequires an MLA backend. The equivalent--kv-transfer-configbehaves the same way.🤖 Generated with Claude Code