Skip to content

Retire legacy B200 DGX Cloud and B300 CoreWeave pools - #2746

Merged
cquil11 merged 1 commit into
mainfrom
codex/migrate-b200-dgxc-to-nscale
Aug 26, 2026
Merged

cquil11 merged 1 commit into
mainfrom
codex/migrate-b200-dgxc-to-nscale

Conversation

@cquil11

@cquil11 cquil11 commented Aug 26, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • migrate active B200 benchmark selectors from retired generic/DGX Cloud labels to cluster:b200-nscale
  • update the Nscale runner inventory to the ten live _00 through _09 registrations
  • retain compatibility for existing active B200 configurations using Nscale storage and Slurm settings
  • remove B300 CoreWeave and B200 DGX Cloud runner groups, launcher references, and functional documentation
  • leave deprecated configs and historical changelog entries unchanged

Runtime cleanup

  • removed all ten B200 DGX Cloud runner registrations after cancelling their active workflow and waiting for leases to clear
  • stopped the retired B200 DGX Cloud runner and telemetry processes
  • confirmed there are no B300 CoreWeave runner registrations remaining

Validation

  • launcher syntax checks pass
  • 303 focused tests pass
  • no retired selectors remain in active master configs

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have。

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

1 similar comment
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have。

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 67955b2. Configure here.

Comment thread runners/launch_b200-nscale-compat.sh
@cquil11
cquil11 force-pushed the codex/migrate-b200-dgxc-to-nscale branch from 67955b2 to 95d8d8e Compare August 26, 2026 22:05

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline finding, I also checked two other candidates: the doc line in experimental/CollectiveX/docs/methodology.md that got blindly find-replaced from "b200-dgxc" to "b200-nscale" — it's now factually backwards (it describes a retired pool showing degraded GDR throughput, but b200-nscale is the pool this PR keeps live), though it's historical prose rather than functional code. I also looked at launch_b200-nscale-slurm.sh silently falling back to the compat launcher for single-node/unsupported cases instead of exiting non-zero as before — a behavior change, but it appears intentional per the PR description ("retain compatibility for existing configurations").

Extended reasoning...

The inline CONFIRMED finding (stale comment pointing to the wrong launcher script for single-node minimaxm3/glm5.2 recipes) is a real but low-severity documentation/comment bug. Beyond it, I independently verified two other candidates raised during the hunt against the actual diff: (1) the methodology.md line, which was mechanically renamed from "b200-dgxc" to "b200-nscale" and is now semantically wrong given that b200-nscale is the pool being kept, not retired — this is docs-only and doesn't affect runtime, so it doesn't rise to inline-comment severity but is worth a human's eye; (2) the intentional-looking control-flow change in launch_b200-nscale-slurm.sh where three previous exit 1 error paths now silently exec into the compat launcher — this matches the PR's stated goal of routing unmigrated configs through the compat launcher, so it appears to be a deliberate design choice rather than a bug, though it does mean previously-loud failures (e.g., a genuinely unsupported model/framework combo) now proceed silently into a different code path, which could mask misconfiguration in the future. Given the hunt exited on max_bugs (not dry_streak) and there's already a confirmed inline finding, plus this PR touches CI runner infrastructure across ~28 files, a human review is warranted regardless, so I'm using the defer path to add these two ruled-out-but-notable items without duplicating the inline comment.

fi

# B200: runners/launch_b200-dgxc.sh resolves the checkpoint to a cluster-local
# B200: runners/launch_b200-nscale-slurm.sh resolves the checkpoint to a cluster-local

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 Comment-rename bug: single-node model-path comments now say launch_b200-nscale-slurm.sh resolves MODEL_PATH, but that script only handles multinode dsv4/kimik3 and falls back to launch_b200-nscale-compat.sh for everything else (including these single-node minimaxm3/glm5.2 recipes). sweep:launch_b200-nscale-slurm\.sh (resolves|rewrites) [also at: experimental/CollectiveX/docs/methodology.md:128 - Blind text rename of 'b200-dgxc' to 'b200-nscale' corrupted a factual historical statement: the line now reads "The…; utils/runner_setup/RUNNER_SETUP.md:261 - Blind 'b200-dgxc'->'b200-nscale' rename left stale path facts: line 256 says pre-staged weights live under…]

Extended reasoning...

An engineer debugging a staging/path issue on these single-node B200 scripts reads the comment, opens runners/launch_b200-nscale-slurm.sh, and finds no matching model-prefix branch (it only has dsv4/kimik3), wasting time chasing the wrong launcher instead of runners/launch_b200-nscale-compat.sh, which is where the actual MODEL_PATH resolution now lives. Same wrong-file rename occurs at benchmarks/single_node/agentic/minimaxm3_fp4_b200_mtp.sh:70, benchmarks/single_node/fixed_seq_len/deprecated/minimaxm3_fp4_b200.sh:22, minimaxm3_fp4_b200_mtp.sh:25, minimaxm3_fp8_b200.sh:31, minimaxm3_fp8_b200_mtp.sh:38, and configs/deprecated/nvidia-minimaxm3-8k1k-master.yaml:1168,1194,1246.

Verification: Severity: nit (comment inaccuracy only; no runtime change). Confirmed real. Runners dispatch via bash ./runners/launch_${RUNNER_NAME%%_*}.sh (benchmark-tmpl.yml:339, profile.yml:206), so RUNNER_NAME=b200-nscale-slurm_XX -> launch_b200-nscale-slurm.sh is the entry launcher for these single-node B200 recipes. But that script does NOT resolve MODEL_PATH for glm5.2/minimaxm3: for single-node it…

@cquil11
cquil11 merged commit 8b79ab5 into main Aug 26, 2026
15 checks passed
@cquil11
cquil11 deleted the codex/migrate-b200-dgxc-to-nscale branch August 26, 2026 22:16
wufann pushed a commit to wufann/InferenceX that referenced this pull request Sep 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant