Skip to content

feat(offload): enable LMCache offload via ATOM_KV_OFFLOAD env - #2414

Merged
valarLip merged 2 commits into
mainfrom
feat/kv-offload-config
Sep 28, 2026
Merged

valarLip merged 2 commits into
mainfrom
feat/kv-offload-config

Conversation

@yhl-amd

@yhl-amd yhl-amd commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

LMCache KV offload is enabled only through --kv-transfer-config. Launchers that own that flag cannot add offload. For example, srtctl (NVIDIA/srt-slurm) generates --kv-transfer-config for P/D workers and rejects it in recipe args. As a result, single-node ATOM AgentX recipes on srt-slurm (MiniMax-M3, Kimi-K3, GLM-5.2 on MI355X) had to drop all of their LMCache offload points.

srtctl passes each role's env to the worker unchanged. It also serializes recipe args with str(), so a JSON-valued flag is fragile there. An environment variable avoids both problems.

Change

Add two environment variables, registered in atom/utils/envs.py. EngineArgs._get_engine_kwargs folds them into kv_transfer_config, so nothing changes after argument parsing:

ATOM_KV_OFFLOAD Resulting kv_transfer_config
unset / 0 / off / none unchanged
lmcache {"kv_connector":"lmcache_offload","kv_role":"offload"} (in-process)
lmcache_mp {"kv_connector":"lmcache_mp","kv_role":"offload"} (standalone lmcache server)
any mode, plus a P/D connector in --kv-transfer-config {"kv_connector":"multi","connectors":[<pd>, <offload>]}
any mode, with --kv-transfer-config already multi offload appended to connectors
any mode, with --kv-transfer-config already naming an offload connector (name or alias) ValueError

ATOM_KV_OFFLOAD_EXTRA_CONFIG takes a JSON object and becomes the offload connector's kv_connector_extra_config. Both backends read it: lmcache.<field> overrides and, for lmcache_mp, lmcache.mp.* options such as lmcache.mp.port. An unknown mode, a non-object extra config, or an extra config without a mode raises an error. When offload is enabled, the resulting kv_transfer_config is logged at startup.

The first revision of this PR used a --kv-offload-config CLI flag. That flag is removed; the environment variables replace it. Without them, behavior is unchanged. docs/environment_variables.md, the offload README and the module docstring document the variables.

Test

  • tests/test_arg_utils_spec.py: 17 environment-variable cases covering every row above, case-insensitive modes, extra-config pass-through, invalid input, and confirming --kv-offload-config is gone. With tests/test_envs.py and tests/entrypoints/test_metrics.py, 112 tests pass locally (CPU torch).
  • black is clean on the changed files. ruff check reports only the two E402 imports that tests/test_arg_utils_spec.py already had.
  • GPU (MI355X, rocm/atom-dev:nightly_202609231534 with this branch mounted), Llama-3.1-8B-Instruct TP1. The test sends a ~23k-token prompt, sends enough distinct prompts to evict it from the GPU prefix cache, then sends it again:
    • ATOM_KV_OFFLOAD=lmcache: startup logs the composed kv_transfer_config and KV transfer config loaded: {'kv_connector': 'lmcache_offload', ...}. The resend loads 22,784 tokens from LMCache (atom:lmcache_loaded_tokens, atom:prefix_cache_offload_tokens) with zero GPU prefix hits. Output matches the first response, and there are no errors.
    • ATOM_KV_OFFLOAD unset: no offload connector, and every atom:lmcache_* counter stays at 0.
    • ATOM_KV_OFFLOAD=lmcache_mp selects lmcache_mp and passes ATOM_KV_OFFLOAD_EXTRA_CONFIG through (lmcache.mp.port). On DeepSeek-R1-0528 TP8 it registers all ranks and saves to the lmcache server. MP data-plane reuse is out of scope for this PR: lookups there fail in the existing TP-collapse path (No GPU context found ... world size 1), and lmcache_mp requires an MLA backend. The equivalent --kv-transfer-config behaves the same way.

🤖 Generated with Claude Code

…onfig

LMCache offload is enabled through --kv-transfer-config, which launchers such
as srtctl own for P/D transfer and refuse to let a recipe override. Add
--kv-offload-config (a JSON dict, like --dspark-config / --dcp-config) and fold
it into kv_transfer_config in EngineArgs._get_engine_kwargs:

- alone: {"kv_connector":"lmcache_offload","kv_role":"offload"} plus any extra
  keys (e.g. "lmcache.chunk_size");
- with a P/D connector: both behind a "multi" connector (appended when the
  transfer config is already "multi");
- rejected when --kv-transfer-config already names an offload connector.

The engine path after argument parsing is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

🏷️ CI Guide

Runs automatically on every eligible PR before approval:

  • ✅ Pre Checkin: Black, Ruff, catalog schema validation, non-GPU unit tests

Heavy model tests:

  • ✅ Run after the PR is approved and Pre Checkin passes
  • ✅ Run immediately when an approval review is submitted
  • ✅ Can be requested before approval with labels
Label Tests
ci:full Run all heavy PR model tests: native ATOM, vLLM, and SGLang
ci:atom Run native ATOM model accuracy tests
ci:vllm Run ATOM vLLM OOT model accuracy tests
ci:sglang Run ATOM SGLang model accuracy tests

Heavy jobs are skipped when the PR is not approved and no matching ci:* label is present.
Add labels via the sidebar or gh pr edit 2414 --add-label <label>

Replace the --kv-offload-config CLI flag with environment variables.
Launchers such as srtctl reserve --kv-transfer-config and serialize
recipe args with str(), but pass worker env through untouched, so an
env var is the robust way to add offload there.

- ATOM_KV_OFFLOAD=lmcache selects the in-process lmcache_offload
  connector; lmcache_mp selects the standalone-server lmcache_mp
  connector. Unset/0/off/none keeps offload off.
- ATOM_KV_OFFLOAD_EXTRA_CONFIG (JSON object) becomes the connector's
  kv_connector_extra_config (lmcache.<field>, lmcache.mp.*).
- Composition with --kv-transfer-config is unchanged: a P/D connector
  and the offload connector run behind "multi", and a second offload
  connector is rejected.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Signed-off-by: yihonglie <hyi@amd.com>
@yhl-amd yhl-amd changed the title feat(offload): add --kv-offload-config independent of --kv-transfer-config feat(offload): enable LMCache offload via ATOM_KV_OFFLOAD env Sep 28, 2026
@valarLip
valarLip merged commit 40a373c into main Sep 28, 2026
30 of 59 checks passed
@valarLip
valarLip deleted the feat/kv-offload-config branch September 28, 2026 08:07
yhl-amd pushed a commit that referenced this pull request Sep 28, 2026
Resolves docs/environment_variables.md: #2414's ATOM_KV_OFFLOAD and
ATOM_KV_OFFLOAD_EXTRA_CONFIG rows join the rewritten LMCache offload table.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
yhl-amd added a commit to SemiAnalysisAI/InferenceX that referenced this pull request Sep 28, 2026
Move the MiniMax-M3 MI355X ATOM AgentX recipe to
rocm/atom-dev:nightly_202609281543, which contains ROCm/ATOM#2414, and
restore the TP2 C20/C25/C30 and TP4 C40/C48 LMCache DRAM-offload arms that
#3463 dropped. srtctl reserves ATOM's kv-transfer-config, so each DRAM
variant enables the in-process lmcache_offload connector with
ATOM_KV_OFFLOAD=lmcache and carries the legacy script's LMCache settings.

中文:将 MiniMax-M3 MI355X ATOM AgentX 配方镜像更新为包含 ROCm/ATOM#2414 的
rocm/atom-dev:nightly_202609281543,并恢复 #3463 删除的 TP2 C20/C25/C30 与
TP4 C40/C48 LMCache DRAM 卸载臂。srtctl 占用 ATOM 的 kv-transfer-config,
因此各 DRAM 变体通过 ATOM_KV_OFFLOAD=lmcache 启用进程内 lmcache_offload
连接器,并沿用旧脚本的 LMCache 设置。

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants