Add Qwen3.5 FP4 B200 AgentX MTP - #2420
Conversation
Add the initial SGLang NEXTN AgentX recipe with golden synthetic acceptance and 256k traces.\n\n中文:新增 Qwen3.5 FP4 B200 AgentX MTP 初始配置,使用 SGLang NEXTN、黄金合成接受长度和 256k 轨迹数据集。
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
Append the Qwen3.5 FP4 B200 AgentX MTP benchmark entry.\n\n中文:追加 Qwen3.5 FP4 B200 AgentX MTP 基准测试触发记录。
Probe the allocated DGXC node for the complete Qwen3.5 NVFP4 checkpoint on /raid and fall back to shared Lustre when it is absent.\n\n中文:在分配到的 DGXC 节点上探测 /raid 中完整的 Qwen3.5 NVFP4 权重;若未暂存,则回退到共享 Lustre。
Add strict two-pool HiCache accounting, request cache reporting, session-sticky attention-DP routing, and AgentX request/graph limits for targeted B200 tuning.\n\n中文:为 B200 定向调优加入严格的双池 HiCache 容量核算、请求级缓存统计、会话粘性的注意力数据并行路由,以及 AgentX 请求数与 CUDA Graph 上限。
Move the configuration into the AgentX section, document the measured cache model, and add session-sticky DEP8 no-offload probes for the throughput frontier.\n\n中文:将配置移入 AgentX 区域,记录实测缓存模型,并为吞吐量前沿加入具备会话粘性路由的 DEP8 无卸载探测点。
中文:添加 B200 TP2 c16 HiCache 对照探针,并严格使用按 GPU 数量计算的 DRAM 上限。
中文:将 B200 TP2 HiCache 验证点设为独立的 c20,避免一次性重复启动无卸载配置。
Add the isolated TP4 c96 no-offload point after c64 remained healthy, so the HBM working-set cliff can be measured before adding matched HiCache coverage. 中文:在 c64 保持健康后新增独立的 TP4 c96 无卸载测试点,以便在添加匹配的 HiCache 覆盖前测量 HBM 工作集拐点。
Remove the DEP8 search arm after its c128 probe fragmented prefix reuse, completed no profiled requests in three minutes, and delivered less than half of TP4 c64 aggregate prefill throughput. 中文:DEP8 c128 测试出现前缀复用碎片化,三分钟内没有完成任何正式请求,聚合预填充吞吐量不足 TP4 c64 的一半,因此移除该搜索分支。
Account for SGLang target KV, Mamba, and native NEXTN draft host pools as H * 31/15 per rank, with page-alignment reserve, so projected use stays within runner-generated DRAM. 中文:按每个 rank 的 H * 31/15 计入 SGLang 目标 KV、Mamba 与原生 NEXTN 草稿主机缓存,并预留页对齐空间,确保预计用量不超过 runner 生成的 DRAM 上限。
Add matched no-offload and DRAM HiCache c72 points immediately after the measured c64 KV-pressure onset. 中文:在实测 c64 KV 压力拐点之后,添加匹配的 c72 无卸载与 DRAM HiCache 对照测试点。
中文:将最新 main 合并到 B200 提交分支,并将本 PR 的性能变更记录重新追加到文件末尾。
Simplify the B200 launcher script to TP-only serving after the measured DEP arm was removed from the search space. 中文:实测 DEP 分支已从搜索空间移除,因此将 B200 启动脚本简化为仅使用 TP 的服务路径。
Replace the dominated c24-c64 no-offload tail with dense c12-c22 points and matched HiCache coverage beginning at c16. 中文:移除性能受压的 c24-c64 无卸载尾部,改为密集采样 c12-c22,并从 c16 开始提供匹配的 HiCache 覆盖。
Densify TP4 around the measured c64-c72 transition and extend the HiCache curve beyond the HBM cliff.\n\n中文:细化 B200 Qwen TP4 在实测 c64-c72 缓存拐点附近的并发点,并将 HiCache 曲线扩展到 HBM 拐点之后。
Remove no-offload c22 after live profiling placed it past the TP2 cache cliff, while retaining the HiCache arm and c20 boundary.\n\n中文:实测表明无卸载 c22 已越过 TP2 缓存拐点,因此移除该点,同时保留 HiCache 配置与 c20 边界点。
中文:在缓存抖动前收紧 B200 TP4 并发范围。
中文:B200 TP2 在并发 14 后切换至 HiCache。
中文:裁剪性能受支配的 B200 TP4 并发 66 配置。
中文:使用 SGLang 默认单 tokenizer 启动路径,避免长时间 HiCache 初始化期间的多 tokenizer 共享内存竞态。
中文:移除已被 c64 严格支配的 B200 c68 无卸载点,并让 HiCache 从缓存悬崖边界开始接管。
Use six tokenizer workers for TP4 long-context replay while retaining SGLang default single-worker startup for TP2 HiCache. 中文:TP4 长上下文回放使用 6 个 tokenizer worker;TP2 HiCache 保持 SGLang 默认单 worker 启动路径,避免共享内存初始化竞态。
Keep the measured Pareto region and validate every node-local checkpoint shard before selecting NVMe weights.\n\n中文:保留实测帕累托区间,并在选择 NVMe 权重前校验节点本地检查点的全部分片。
Remove no-offload c22 after live profiling placed it past the TP2 cache cliff, while retaining the HiCache arm and c20 boundary.\n\n中文:实测表明无卸载 c22 已越过 TP2 缓存拐点,因此移除该点,同时保留 HiCache 配置与 c20 边界点。
中文:在缓存抖动前收紧 B200 TP4 并发范围。
中文:B200 TP2 在并发 14 后切换至 HiCache。
中文:裁剪性能受支配的 B200 TP4 并发 66 配置。
中文:使用 SGLang 默认单 tokenizer 启动路径,避免长时间 HiCache 初始化期间的多 tokenizer 共享内存竞态。
中文:移除已被 c64 严格支配的 B200 c68 无卸载点,并让 HiCache 从缓存悬崖边界开始接管。
Use six tokenizer workers for TP4 long-context replay while retaining SGLang default single-worker startup for TP2 HiCache. 中文:TP4 长上下文回放使用 6 个 tokenizer worker;TP2 HiCache 保持 SGLang 默认单 worker 启动路径,避免共享内存初始化竞态。
Keep the measured Pareto region and validate every node-local checkpoint shard before selecting NVMe weights.\n\n中文:保留实测帕累托区间,并在选择 NVMe 权重前校验节点本地检查点的全部分片。
Pin the AIPerf trace idle-gap branch and set the Qwen B200 replay cap to 300 seconds.
Pin Qwen B200 AgentX to the validated AIPerf scheduling revision and run GSM8K instead of SWE-bench for the full-hour sweep. 中文:将 Qwen B200 AgentX 固定到已验证的 AIPerf 调度版本,并在一小时完整扫描中使用 GSM8K 替代 SWE-bench。
5e3d346 to
1267ca9
Compare
|
As a PR reviewer and CODEOWNER, I have reviewed this and have:
Additional detail section:
Signed: |
❌❌❌ REJECTED ❌❌❌@cquil11 — blocking: no passing sweep/evals exist on any commit currently in this PR. The linked validation and eval runs executed on commits that were later rebased out of the branch, so ✅ Check 0 (CODEOWNER): PASS — signer is an org MEMBER covered via |
中文:将已验证的 Qwen B200 扫描提交重新纳入分支历史,不更改当前文件内容。
|
/reuse-sweep-run 30758866378 中文:复用已验证的扫描运行 30758866378。 |
|
As a PR reviewer and CODEOWNER, I have reviewed this and have:
Additional detail section:
Signed: |
✅✅✅ Verdict: PASS ✅✅✅✅ Check 0 (CODEOWNER): PASS — all changed paths resolve to |
Remove the node-local Qwen3.5 FP4 checkpoint probe because validated runs consistently use the shared Lustre checkpoint. Keep exporting the resolved MODEL_PATH to the benchmark script.\n\n中文:移除未使用的 Qwen3.5 FP4 节点本地 NVMe 检查;已验证的运行均使用共享 Lustre 模型路径。继续将解析后的 MODEL_PATH 传递给基准测试脚本。
Move the changelog entry to the file tail and restore the two-space separator after PR #2420 so process_changelog sees additions only. Co-authored-by: Cursor <cursoragent@cursor.com>
* feat: add Qwen3.5 FP4 B200 AgentX MTP Add the initial SGLang NEXTN AgentX recipe with golden synthetic acceptance and 256k traces.\n\n中文:新增 Qwen3.5 FP4 B200 AgentX MTP 初始配置,使用 SGLang NEXTN、黄金合成接受长度和 256k 轨迹数据集。 * chore: trigger B200 AgentX sweep Append the Qwen3.5 FP4 B200 AgentX MTP benchmark entry.\n\n中文:追加 Qwen3.5 FP4 B200 AgentX MTP 基准测试触发记录。 * perf: expand B200 AgentX preflight range * chore: use SGLang v0.5.16 for Qwen B200 * chore: collect SGLang cache metrics on B200 * fix: remove unsupported SGLang tool-choice flag on B200 * perf: prefer local Qwen FP4 weights on B200 Probe the allocated DGXC node for the complete Qwen3.5 NVFP4 checkpoint on /raid and fall back to shared Lustre when it is absent.\n\n中文:在分配到的 DGXC 节点上探测 /raid 中完整的 Qwen3.5 NVFP4 权重;若未暂存,则回退到共享 Lustre。 * perf: prepare Qwen AgentX cache and DEP tuning Add strict two-pool HiCache accounting, request cache reporting, session-sticky attention-DP routing, and AgentX request/graph limits for targeted B200 tuning.\n\n中文:为 B200 定向调优加入严格的双池 HiCache 容量核算、请求级缓存统计、会话粘性的注意力数据并行路由,以及 AgentX 请求数与 CUDA Graph 上限。 * perf: add Qwen B200 DEP probes Move the configuration into the AgentX section, document the measured cache model, and add session-sticky DEP8 no-offload probes for the throughput frontier.\n\n中文:将配置移入 AgentX 区域,记录实测缓存模型,并为吞吐量前沿加入具备会话粘性路由的 DEP8 无卸载探测点。 * perf: add matched B200 HiCache cliff probe 中文:添加 B200 TP2 c16 HiCache 对照探针,并严格使用按 GPU 数量计算的 DRAM 上限。 * perf: isolate B200 HiCache validation point 中文:将 B200 TP2 HiCache 验证点设为独立的 c20,避免一次性重复启动无卸载配置。 * perf: extend Qwen B200 TP4 cliff probe Add the isolated TP4 c96 no-offload point after c64 remained healthy, so the HBM working-set cliff can be measured before adding matched HiCache coverage. 中文:在 c64 保持健康后新增独立的 TP4 c96 无卸载测试点,以便在添加匹配的 HiCache 覆盖前测量 HBM 工作集拐点。 * perf: drop dominated Qwen B200 DEP arm Remove the DEP8 search arm after its c128 probe fragmented prefix reuse, completed no profiled requests in three minutes, and delivered less than half of TP4 c64 aggregate prefill throughput. 中文:DEP8 c128 测试出现前缀复用碎片化,三分钟内没有完成任何正式请求,聚合预填充吞吐量不足 TP4 c64 的一半,因此移除该搜索分支。 * fix: enforce Qwen HiCache DRAM total Account for SGLang target KV, Mamba, and native NEXTN draft host pools as H * 31/15 per rank, with page-alignment reserve, so projected use stays within runner-generated DRAM. 中文:按每个 rank 的 H * 31/15 计入 SGLang 目标 KV、Mamba 与原生 NEXTN 草稿主机缓存,并预留页对齐空间,确保预计用量不超过 runner 生成的 DRAM 上限。 * perf: add B200 TP4 c72 HiCache A/B Add matched no-offload and DRAM HiCache c72 points immediately after the measured c64 KV-pressure onset. 中文:在实测 c64 KV 压力拐点之后,添加匹配的 c72 无卸载与 DRAM HiCache 对照测试点。 * refactor: remove dominated B200 DEP path Simplify the B200 launcher script to TP-only serving after the measured DEP arm was removed from the search space. 中文:实测 DEP 分支已从搜索空间移除,因此将 B200 启动脚本简化为仅使用 TP 的服务路径。 * perf: densify B200 TP2 cache cliff Replace the dominated c24-c64 no-offload tail with dense c12-c22 points and matched HiCache coverage beginning at c16. 中文:移除性能受压的 c24-c64 无卸载尾部,改为密集采样 c12-c22,并从 c16 开始提供匹配的 HiCache 覆盖。 * perf: resolve B200 Qwen cache cliff Densify TP4 around the measured c64-c72 transition and extend the HiCache curve beyond the HBM cliff.\n\n中文:细化 B200 Qwen TP4 在实测 c64-c72 缓存拐点附近的并发点,并将 HiCache 曲线扩展到 HBM 拐点之后。 * perf: prune dominated B200 TP2 point Remove no-offload c22 after live profiling placed it past the TP2 cache cliff, while retaining the HiCache arm and c20 boundary.\n\n中文:实测表明无卸载 c22 已越过 TP2 缓存拐点,因此移除该点,同时保留 HiCache 配置与 c20 边界点。 * perf: stop B200 TP4 before cache thrashing 中文:在缓存抖动前收紧 B200 TP4 并发范围。 * perf: switch B200 TP2 to HiCache after c14 中文:B200 TP2 在并发 14 后切换至 HiCache。 * perf: prune dominated B200 TP4 c66 中文:裁剪性能受支配的 B200 TP4 并发 66 配置。 * fix(agentx): use stable tokenizer startup 中文:使用 SGLang 默认单 tokenizer 启动路径,避免长时间 HiCache 初始化期间的多 tokenizer 共享内存竞态。 * perf(agentx): cap B200 no-offload frontier 中文:移除已被 c64 严格支配的 B200 c68 无卸载点,并让 HiCache 从缓存悬崖边界开始接管。 * fix(agentx): scale Qwen tokenization by topology Use six tokenizer workers for TP4 long-context replay while retaining SGLang default single-worker startup for TP2 HiCache. 中文:TP4 长上下文回放使用 6 个 tokenizer worker;TP2 HiCache 保持 SGLang 默认单 worker 启动路径,避免共享内存初始化竞态。 * perf(agentx): trim B200 cache-thrash tail Keep the measured Pareto region and validate every node-local checkpoint shard before selecting NVMe weights.\n\n中文:保留实测帕累托区间,并在选择 NVMe 权重前校验节点本地检查点的全部分片。 * fix(agentx): cap Qwen trace idle gaps Pin the AIPerf trace idle-gap branch and set the Qwen B200 replay cap to 300 seconds. * chore(aiperf): pin merged trace idle cap branch * chore(agentx): bump AIPerf trace idle cap fix Signed-off-by: Cam Quilici <cjquilici@gmail.com> * fix(agentx): update AIPerf trace-cap cleanup Signed-off-by: Cam Quilici <cjquilici@gmail.com> * chore(agentx): pin reconstruction-only idle cap Update AIPerf after removing runtime trace idle enforcement. Signed-off-by: Cam Quilici <cjquilici@gmail.com> * fix(agentx): update AIPerf warmup handoff Signed-off-by: Cam Quilici <cjquilici@gmail.com> * fix(agentx): retain baseline parents across warmup * fix(agentx): extend SGLang keep-alive * fix(agentx): pin validated AIPerf scheduling on B200 Update the Qwen3.5 B200 AgentX submission to the validated AIPerf revision containing profiling handoff anchoring, cache-bust continuity, spawn/join timing, and runtime idle caps.\n\n中文:将 Qwen3.5 B200 AgentX 提交固定到已验证的 AIPerf 版本,包含 profiling 交接锚定、cache-bust 连续性、spawn/join 时序以及运行时空闲上限修复。 * test(agentx): validate current AIPerf on Qwen B200 Pin Qwen B200 AgentX to the validated AIPerf scheduling revision and run GSM8K instead of SWE-bench for the full-hour sweep. 中文:将 Qwen B200 AgentX 固定到已验证的 AIPerf 调度版本,并在一小时完整扫描中使用 GSM8K 替代 SWE-bench。 * feat: add Qwen3.5 FP4 B200 AgentX MTP Add the initial SGLang NEXTN AgentX recipe with golden synthetic acceptance and 256k traces.\n\n中文:新增 Qwen3.5 FP4 B200 AgentX MTP 初始配置,使用 SGLang NEXTN、黄金合成接受长度和 256k 轨迹数据集。 * chore: trigger B200 AgentX sweep Append the Qwen3.5 FP4 B200 AgentX MTP benchmark entry.\n\n中文:追加 Qwen3.5 FP4 B200 AgentX MTP 基准测试触发记录。 * perf: expand B200 AgentX preflight range * chore: use SGLang v0.5.16 for Qwen B200 * chore: collect SGLang cache metrics on B200 * fix: remove unsupported SGLang tool-choice flag on B200 * perf: prefer local Qwen FP4 weights on B200 Probe the allocated DGXC node for the complete Qwen3.5 NVFP4 checkpoint on /raid and fall back to shared Lustre when it is absent.\n\n中文:在分配到的 DGXC 节点上探测 /raid 中完整的 Qwen3.5 NVFP4 权重;若未暂存,则回退到共享 Lustre。 * perf: prepare Qwen AgentX cache and DEP tuning Add strict two-pool HiCache accounting, request cache reporting, session-sticky attention-DP routing, and AgentX request/graph limits for targeted B200 tuning.\n\n中文:为 B200 定向调优加入严格的双池 HiCache 容量核算、请求级缓存统计、会话粘性的注意力数据并行路由,以及 AgentX 请求数与 CUDA Graph 上限。 * perf: add Qwen B200 DEP probes Move the configuration into the AgentX section, document the measured cache model, and add session-sticky DEP8 no-offload probes for the throughput frontier.\n\n中文:将配置移入 AgentX 区域,记录实测缓存模型,并为吞吐量前沿加入具备会话粘性路由的 DEP8 无卸载探测点。 * perf: add matched B200 HiCache cliff probe 中文:添加 B200 TP2 c16 HiCache 对照探针,并严格使用按 GPU 数量计算的 DRAM 上限。 * perf: isolate B200 HiCache validation point 中文:将 B200 TP2 HiCache 验证点设为独立的 c20,避免一次性重复启动无卸载配置。 * perf: extend Qwen B200 TP4 cliff probe Add the isolated TP4 c96 no-offload point after c64 remained healthy, so the HBM working-set cliff can be measured before adding matched HiCache coverage. 中文:在 c64 保持健康后新增独立的 TP4 c96 无卸载测试点,以便在添加匹配的 HiCache 覆盖前测量 HBM 工作集拐点。 * perf: drop dominated Qwen B200 DEP arm Remove the DEP8 search arm after its c128 probe fragmented prefix reuse, completed no profiled requests in three minutes, and delivered less than half of TP4 c64 aggregate prefill throughput. 中文:DEP8 c128 测试出现前缀复用碎片化,三分钟内没有完成任何正式请求,聚合预填充吞吐量不足 TP4 c64 的一半,因此移除该搜索分支。 * fix: enforce Qwen HiCache DRAM total Account for SGLang target KV, Mamba, and native NEXTN draft host pools as H * 31/15 per rank, with page-alignment reserve, so projected use stays within runner-generated DRAM. 中文:按每个 rank 的 H * 31/15 计入 SGLang 目标 KV、Mamba 与原生 NEXTN 草稿主机缓存,并预留页对齐空间,确保预计用量不超过 runner 生成的 DRAM 上限。 * perf: add B200 TP4 c72 HiCache A/B Add matched no-offload and DRAM HiCache c72 points immediately after the measured c64 KV-pressure onset. 中文:在实测 c64 KV 压力拐点之后,添加匹配的 c72 无卸载与 DRAM HiCache 对照测试点。 * refactor: remove dominated B200 DEP path Simplify the B200 launcher script to TP-only serving after the measured DEP arm was removed from the search space. 中文:实测 DEP 分支已从搜索空间移除,因此将 B200 启动脚本简化为仅使用 TP 的服务路径。 * perf: densify B200 TP2 cache cliff Replace the dominated c24-c64 no-offload tail with dense c12-c22 points and matched HiCache coverage beginning at c16. 中文:移除性能受压的 c24-c64 无卸载尾部,改为密集采样 c12-c22,并从 c16 开始提供匹配的 HiCache 覆盖。 * perf: resolve B200 Qwen cache cliff Densify TP4 around the measured c64-c72 transition and extend the HiCache curve beyond the HBM cliff.\n\n中文:细化 B200 Qwen TP4 在实测 c64-c72 缓存拐点附近的并发点,并将 HiCache 曲线扩展到 HBM 拐点之后。 * perf: prune dominated B200 TP2 point Remove no-offload c22 after live profiling placed it past the TP2 cache cliff, while retaining the HiCache arm and c20 boundary.\n\n中文:实测表明无卸载 c22 已越过 TP2 缓存拐点,因此移除该点,同时保留 HiCache 配置与 c20 边界点。 * perf: stop B200 TP4 before cache thrashing 中文:在缓存抖动前收紧 B200 TP4 并发范围。 * perf: switch B200 TP2 to HiCache after c14 中文:B200 TP2 在并发 14 后切换至 HiCache。 * perf: prune dominated B200 TP4 c66 中文:裁剪性能受支配的 B200 TP4 并发 66 配置。 * fix(agentx): use stable tokenizer startup 中文:使用 SGLang 默认单 tokenizer 启动路径,避免长时间 HiCache 初始化期间的多 tokenizer 共享内存竞态。 * perf(agentx): cap B200 no-offload frontier 中文:移除已被 c64 严格支配的 B200 c68 无卸载点,并让 HiCache 从缓存悬崖边界开始接管。 * fix(agentx): scale Qwen tokenization by topology Use six tokenizer workers for TP4 long-context replay while retaining SGLang default single-worker startup for TP2 HiCache. 中文:TP4 长上下文回放使用 6 个 tokenizer worker;TP2 HiCache 保持 SGLang 默认单 worker 启动路径,避免共享内存初始化竞态。 * perf(agentx): trim B200 cache-thrash tail Keep the measured Pareto region and validate every node-local checkpoint shard before selecting NVMe weights.\n\n中文:保留实测帕累托区间,并在选择 NVMe 权重前校验节点本地检查点的全部分片。 * fix(agentx): cap Qwen trace idle gaps Pin the AIPerf trace idle-gap branch and set the Qwen B200 replay cap to 300 seconds. * fix(agentx): extend SGLang keep-alive * test(agentx): validate current AIPerf on Qwen B200 Pin Qwen B200 AgentX to the validated AIPerf scheduling revision and run GSM8K instead of SWE-bench for the full-hour sweep. 中文:将 Qwen B200 AgentX 固定到已验证的 AIPerf 调度版本,并在一小时完整扫描中使用 GSM8K 替代 SWE-bench。 * Update nvidia-master.yaml * Update perf-changelog.yaml * chore(b200): remove unused Qwen NVMe probe Remove the node-local Qwen3.5 FP4 checkpoint probe because validated runs consistently use the shared Lustre checkpoint. Keep exporting the resolved MODEL_PATH to the benchmark script.\n\n中文:移除未使用的 Qwen3.5 FP4 节点本地 NVMe 检查;已验证的运行均使用共享 Lustre 模型路径。继续将解析后的 MODEL_PATH 传递给基准测试脚本。 --------- Signed-off-by: Cam Quilici <cjquilici@gmail.com>
* llm-d-vllm: optimize DSv4-Pro GB200 recipe configs Tune decode and prefill vLLM flags based on validated GB200 NVL72 benchmarks showing +2-43% tok/s/GPU and 15-25% lower TPOT: Decode (mid-curve-megamoe): - gpu-memory-utilization 0.85 -> 0.9 (larger KV cache pool) - Add max-model-len 9280 (tight ISL8192+OSL1024 bound) - max-num-seqs/batched-tokens/cudagraph-capture 512 -> 1024 - Disable NCCL symmetric memory (standard NVLink path faster) - Add no-enable-flashinfer-autotune, rust frontend Prefill (both recipes): - gpu-memory-utilization 0.9 -> 0.95 - Add max-model-len 9280, max-num-seqs 16, max-num-batched-tokens 32768 - Disable NCCL symmetric memory, enable rust frontend - Enable randomize-dp-dummy-inputs EPP (both recipes): - Switch from max-score-picker to weighted-random-picker (threshold=0.1) for better load distribution under high concurrency Low-latency decode: disable NCCL symmetric memory, add rust frontend. * Update perf-changelog.yaml * Fix llm-d sweep changelog config key * Bump version, remove stream-interval * Update perf-changelog.yaml * Fix image tag * Fix perf-changelog for PR #2498 Point pr-link at #2498 and drop inaccurate max-num-seqs 512->1024 claim. Co-authored-by: Cursor <cursoragent@cursor.com> * Fix perf-changelog append-only CI failure for PR #2498 Move the changelog entry to the file tail and restore the two-space separator after PR #2420 so process_changelog sees additions only. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(gb200): reuse guarded Enroot imports Move the shared image URI and squash import helpers before the llm-d early path, then use the same locked and atomic importer for its Quay image. Give fresh imports a private temporary Enroot runtime directory so CI does not fall back to /run/enroot. 中文:修复 GB200 启动器,让 llm-d 的 Quay 镜像复用带锁、原子替换和校验的共享导入逻辑。为全新导入创建独立的临时 Enroot 运行目录,避免 CI 回退到无写权限的 /run/enroot。 * fix(gb200): pin Quay image digest Pin the vllm0.26 OCI index so Enroot 3.5 can query it on GB200. Quay returns 404 for the tag during Enroot’s headerless permission probe, while the immutable digest returns 200 and resolves the arm64 manifest. 中文:固定 vllm0.26 的 OCI 索引摘要,使 GB200 上的 Enroot 3.5 能够正常查询镜像。Quay 对 Enroot 不带 Accept 标头的标签权限探测返回 404,而不可变摘要返回 200 并可解析 arm64 清单。 * Fix EPP startup by dropping unsupported weighted-random-picker threshold. EPP v0.9.0 rejects the threshold parameter; remove it from both GB200 recipes and update the changelog entry accordingly. Co-authored-by: Cursor <cursoragent@cursor.com> * Drop llmd-vllm low-middle sweep using low-latency recipe at c256/512. The low-latency recipe is tuned for c1 only (1P×1D); high-concurrency 1P×4D jobs were misconfigured and failing CI unrelated to recipe tuning. Co-authored-by: Cursor <cursoragent@cursor.com> * Lower mid-curve megamoe prefill GMU to avoid warmup OOM. Prefill at 0.95 left no headroom for JIT spikes during c256 warmup on GB200 CI. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cameron Quilici <cjquilici@gmail.com> Co-authored-by: Cursor <cursoragent@cursor.com>
Summary
Add Qwen3.5-397B-A17B NVFP4 AgentX coverage on B200 with SGLang native NEXTN MTP, golden synthetic acceptance, prefix caching, the 256k trace dataset, and a 300-second per-trace idle-gap cap.