Retire legacy B200 DGX Cloud and B300 CoreWeave pools - #2746
Conversation
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
1 similar comment
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 67955b2. Configure here.
67955b2 to
95d8d8e
Compare
There was a problem hiding this comment.
Beyond the inline finding, I also checked two other candidates: the doc line in experimental/CollectiveX/docs/methodology.md that got blindly find-replaced from "b200-dgxc" to "b200-nscale" — it's now factually backwards (it describes a retired pool showing degraded GDR throughput, but b200-nscale is the pool this PR keeps live), though it's historical prose rather than functional code. I also looked at launch_b200-nscale-slurm.sh silently falling back to the compat launcher for single-node/unsupported cases instead of exiting non-zero as before — a behavior change, but it appears intentional per the PR description ("retain compatibility for existing configurations").
Extended reasoning...
The inline CONFIRMED finding (stale comment pointing to the wrong launcher script for single-node minimaxm3/glm5.2 recipes) is a real but low-severity documentation/comment bug. Beyond it, I independently verified two other candidates raised during the hunt against the actual diff: (1) the methodology.md line, which was mechanically renamed from "b200-dgxc" to "b200-nscale" and is now semantically wrong given that b200-nscale is the pool being kept, not retired — this is docs-only and doesn't affect runtime, so it doesn't rise to inline-comment severity but is worth a human's eye; (2) the intentional-looking control-flow change in launch_b200-nscale-slurm.sh where three previous exit 1 error paths now silently exec into the compat launcher — this matches the PR's stated goal of routing unmigrated configs through the compat launcher, so it appears to be a deliberate design choice rather than a bug, though it does mean previously-loud failures (e.g., a genuinely unsupported model/framework combo) now proceed silently into a different code path, which could mask misconfiguration in the future. Given the hunt exited on max_bugs (not dry_streak) and there's already a confirmed inline finding, plus this PR touches CI runner infrastructure across ~28 files, a human review is warranted regardless, so I'm using the defer path to add these two ruled-out-but-notable items without duplicating the inline comment.
| fi | ||
|
|
||
| # B200: runners/launch_b200-dgxc.sh resolves the checkpoint to a cluster-local | ||
| # B200: runners/launch_b200-nscale-slurm.sh resolves the checkpoint to a cluster-local |
There was a problem hiding this comment.
🟡 Comment-rename bug: single-node model-path comments now say launch_b200-nscale-slurm.sh resolves MODEL_PATH, but that script only handles multinode dsv4/kimik3 and falls back to launch_b200-nscale-compat.sh for everything else (including these single-node minimaxm3/glm5.2 recipes). sweep:launch_b200-nscale-slurm\.sh (resolves|rewrites) [also at: experimental/CollectiveX/docs/methodology.md:128 - Blind text rename of 'b200-dgxc' to 'b200-nscale' corrupted a factual historical statement: the line now reads "The…; utils/runner_setup/RUNNER_SETUP.md:261 - Blind 'b200-dgxc'->'b200-nscale' rename left stale path facts: line 256 says pre-staged weights live under…]
Extended reasoning...
An engineer debugging a staging/path issue on these single-node B200 scripts reads the comment, opens runners/launch_b200-nscale-slurm.sh, and finds no matching model-prefix branch (it only has dsv4/kimik3), wasting time chasing the wrong launcher instead of runners/launch_b200-nscale-compat.sh, which is where the actual MODEL_PATH resolution now lives. Same wrong-file rename occurs at benchmarks/single_node/agentic/minimaxm3_fp4_b200_mtp.sh:70, benchmarks/single_node/fixed_seq_len/deprecated/minimaxm3_fp4_b200.sh:22, minimaxm3_fp4_b200_mtp.sh:25, minimaxm3_fp8_b200.sh:31, minimaxm3_fp8_b200_mtp.sh:38, and configs/deprecated/nvidia-minimaxm3-8k1k-master.yaml:1168,1194,1246.
Verification: Severity: nit (comment inaccuracy only; no runtime change). Confirmed real. Runners dispatch via bash ./runners/launch_${RUNNER_NAME%%_*}.sh (benchmark-tmpl.yml:339, profile.yml:206), so RUNNER_NAME=b200-nscale-slurm_XX -> launch_b200-nscale-slurm.sh is the entry launcher for these single-node B200 recipes. But that script does NOT resolve MODEL_PATH for glm5.2/minimaxm3: for single-node it…

Summary
cluster:b200-nscale_00through_09registrationsRuntime cleanup
Validation