fix(config): annotate DSV4 TRT AgentX native KV offload - #2708
Conversation
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
Summary
Scope
This changes master-config metadata only. It does not modify benchmark recipes or runtime behavior, so there is no performance changelog entry.
Validation
Note
Low Risk
Master-config metadata only; no recipe, runtime, or benchmark-path changes.
Overview
Marks the six DeepSeek-V4 TensorRT-LLM AgentX search points as DRAM KV offload with the native backend (
native@1.3.0rc24) and setsdram-utilizationso generated CPU DRAM budget is 193 GB.Metadata-only change in
nvidia-master.yaml; recipes and runtime behavior are unchanged.Reviewed by Cursor Bugbot for commit 6404c1b. Bugbot is set up for automated code reviews on this repo. Configure here.