Skip to content

feat(agentx): run the MiniMax-M3 MI300X LMCache point on the lmcache-server service - #3545

Merged
cquil11 merged 8 commits into
mainfrom
agentx/minimaxm3-mi300x-lmcache-server
Sep 29, 2026
Merged

cquil11 merged 8 commits into
mainfrom
agentx/minimaxm3-mi300x-lmcache-server

Conversation

@cquil11

@cquil11 cquil11 commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Moves the MI300X MiniMax-M3 AgentX LMCache variant (override_tp8_c16_lmcache) to srt-slurm's lmcache-server service. srt-slurm owns startup, readiness, and cleanup; the setup script only installs dependencies.

  • Replaces eight 162 GB per-rank servers with one 1296 GB node-local server shared by all eight TP ranks, retaining --max-workers 2.
  • Preserves the 6000-second connector timeout and allows 900 seconds for service installation and readiness.
  • Aligns the recipe image with the master config's vLLM ROCm v0.30.0.

Uses the LMCache patch shared with #3543, from srt-slurm #32. Upstream: #507 and #528.

Validation

Configuration Run Result
MI300X, TP8, concurrency 16 36460859852 Passed; all eight GPUs registered, 1296 GB L1 allocated, KV stores and retrieves observed.

Throughput was about 14% below the earlier eight-server run (9,970 vs. 11,586 tok/s), with higher p90 TTFT. That baseline used vLLM v0.29.0, so the comparison does not isolate the effect of changing the server layout.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

amd-master.yaml moved minimaxm3-fp8-mi300x-vllm-agentic-mtp to
vllm/vllm-openai-rocm:v0.30.0 in #3361, but the recipe still named v0.29.0,
so the single-node adapter rejected every point.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I also looked at the new lmcache-server service binding --host 0.0.0.0 --http-host 0.0.0.0 in the vendored srt-slurm patch (line ~600) — this exposes the RPC/HTTP ports to the whole node network rather than just localhost, but that's consistent with how srt-slurm's existing service kinds bind for multi-worker access within the job's network namespace, so it doesn't look like a new exposure introduced by this PR.

Extended reasoning...

The diff migrates LMCache DRAM cache setup from a per-TP-rank shell-scripted fleet to a vendored srt-slurm lmcache-server service; it touches a shell script, a benchmark YAML config, a large vendored patch, and a changelog, with no direct auth/crypto surface. An inline finding (pre-existing set -euo pipefail nounset violation) is already being posted, so I did not approve; I additionally checked the 0.0.0.0 bind in the new service definition and found it matches the pattern of other srt-slurm service kinds rather than being a novel exposure.

…MCache point

The refreshed lmcache-server patch (SemiAnalysisAI/srt-slurm#32 at 181b2e4)
keeps a role's connector on a direct vllm serve aggregate worker, so the
variant names connector: lmcache-mp instead of hand-writing its
kv-transfer-config. The message queue timeout falls back to LMCache's
default (300 s) instead of 6000 s.
The role connector is now the lmcache-mp preset written out plus
lmcache.mp.mq_timeout 6000, so the point keeps its measured timeout
instead of falling back to LMCache's 300 s default.
…slurm v2.30.0

main bumped srt-slurm to v2.30.0 and dropped the 504 patch. The 507 patch
is regenerated on v2.30.0 from SemiAnalysisAI/srt-slurm#32 at 181b2e4, and
the #3545 perf-changelog entry is rewritten to state the benchmark change:
one lmcache-server with 1296 GB L1 instead of 8 per-rank 162 GB servers,
image v0.30.0, LMCacheMPConnector at localhost:8750 with mq_timeout 6000,
re-measured results.
@github-actions

Copy link
Copy Markdown
Contributor

@cquil11

cquil11 commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

/reuse-sweep-run 36479049525

@cquil11
cquil11 force-pushed the agentx/minimaxm3-mi300x-lmcache-server branch from 5a8ecf3 to 835de73 Compare September 29, 2026 14:19
@cquil11

cquil11 commented Sep 29, 2026

Copy link
Copy Markdown
Collaborator Author

/stage-results 36479049525

…0x-lmcache-server

# Conflicts:
#	inferencex-e2e/perf-changelog.yaml
@github-actions

github-actions Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

@cquil11 staged run 36479049525: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-09-28~r36479049525

This run remains available across future /stage-results requests. Staging the same run ID again updates its staged data. Staging workflow

@cquil11
cquil11 merged commit 9568edd into main Sep 29, 2026
25 checks passed
@cquil11
cquil11 deleted the agentx/minimaxm3-mi300x-lmcache-server branch September 29, 2026 14:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

2 participants