Skip to content

[GLM-5.2 GB200] Preserve draft precision and upgrade disaggregated image / [GLM-5.2 GB200] 保留 draft 精度并升级分离式镜像 - #3401

Merged
edwingao28 merged 11 commits into
mainfrom
fix/glm52-gb200-draft-quantization-off
Oct 7, 2026
Merged

edwingao28 merged 11 commits into
mainfrom
fix/glm52-gb200-draft-quantization-off

Conversation

@edwingao28

@edwingao28 edwingao28 commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Preserve GLM-5.2 GB200 draft precision and upgrade both disaggregated recipes to digest-pinned upstream SGLang 0.5.18 CUDA 13, removing the startup NIXL override.

Testing: Sweep 37506762704: 14 performance + 4 evals passed and artifacts reviewed at d8e7cc8; CPU CI passed. Reuse selected.

Limits: Cancellations and cleanup/identity warnings remain. Main's newer launcher is not newly measured; CODEOWNER re-review remains required.

中文

保留 GLM-5.2 GB200 draft 原始精度,将两个分离式配方升级到固定 digest 的上游 SGLang 0.5.18 CUDA 13 镜像,移除启动时的 NIXL 覆盖。

测试: 扫描 37506762704 的 14 个性能点和 4 个 eval 全部通过,并完成 d8e7cc8 的产物审阅;CPU CI 通过,已选择复用。

限制: 保留请求取消、清理和身份核对警告。同步引入的新版主分支 launcher 尚未重新实测;仍需 CODEOWNER 复审。

AI 模型: Claude Opus 5.5 (claude-opus-5-5) 负责原始实现;GPT-6(具体变体不可确认)负责更新、冲突修复、验证和委派审阅。关联 #3228;修复及配置变更。

AI model disclosure

  • Claude Opus 5.5 (claude-opus-5-5): original implementation.
  • GPT-6 (exact variant unavailable): updates, conflict repair, verification and delegated review.

Related Issue

Related to #3228.

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have completed the AI model disclosure and kept it current
  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of inferencex-e2e/perf-changelog.yaml and have not edited historical entries
  • If this PR can affect benchmark performance or adds or modifies a recipe, it carries exactly one primary sweep label (a maintainer applies it on fork PRs): full-sweep-fail-fast (recommended), full-sweep-enabled, or non-canary-full-sweep-enabled. Optional modifiers all-evals, evals-only, and agentx-fast require a primary label; the last two block reuse while applied.
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /use <run_id> (or the legacy /reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the primary sweep label will no longer automatically kick off new sweeps. Remove and re-add the primary sweep label to force a new sweep.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

@edwingao28
edwingao28 force-pushed the fix/glm52-gb200-draft-quantization-off branch from 91efe8b to e908cc4 Compare September 23, 2026 21:59
@edwingao28 edwingao28 added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 23, 2026
@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

@edwingao28 edwingao28 removed the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 25, 2026
@edwingao28 edwingao28 changed the title [GLM-5.2] Disable GB200 draft MoE quantization / 关闭 GB200 draft MoE 量化 [GLM-5.2 GB200] require PowerX for draft flag-off reruns / 关闭草稿量化并要求功耗验收 Sep 25, 2026
@edwingao28 edwingao28 changed the title [GLM-5.2 GB200] require PowerX for draft flag-off reruns / 关闭草稿量化并要求功耗验收 [GLM-5.2 GB200] disable draft FP8 conversion and require PowerX / 关闭草稿 FP8 转换并要求功耗验收 Sep 26, 2026
@edwingao28 edwingao28 added full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) engine-patch and removed full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) engine-patch labels Sep 26, 2026
@edwingao28
edwingao28 force-pushed the fix/glm52-gb200-draft-quantization-off branch from 68d3ce2 to 274c69d Compare September 27, 2026 07:55
@edwingao28 edwingao28 added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Sep 27, 2026
@edwingao28 edwingao28 changed the title [GLM-5.2 GB200] disable draft FP8 conversion and require PowerX / 关闭草稿 FP8 转换并要求功耗验收 [GLM-5.2 GB200] disable draft MoE FP8 conversion / 关闭 draft MoE FP8 转换 Sep 27, 2026
@edwingao28 edwingao28 added full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures and removed full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) labels Sep 27, 2026
@functionstackx

Copy link
Copy Markdown
Collaborator

Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding

@cursor
cursor Bot force-pushed the fix/glm52-gb200-draft-quantization-off branch 2 times, most recently from 6113b30 to fca8de5 Compare September 28, 2026 18:38
…refill

Keep the draft at shipped precision and disable only the two background UCX progress controls for disaggregated prefill. Preserve images, workload, strict synchronization and golden acceptance. Request the scoped engine-patch waiver.

保留 GLM-5.2 GB200 draft 的原始精度,仅关闭分离式 prefill 的两项 UCX 后台进展控制;保留镜像、工作负载、严格同步及 golden acceptance,并申请对应补丁例外。
@edwingao28
edwingao28 force-pushed the fix/glm52-gb200-draft-quantization-off branch from fca8de5 to 9fbe9bf Compare September 29, 2026 07:48
@edwingao28 edwingao28 changed the title [GLM-5.2 GB200] disable draft MoE FP8 conversion / 关闭 draft MoE FP8 转换 [GLM-5.2 GB200] Preserve draft precision and disable background UCX progress / [GLM-5.2 GB200] 保留 draft 精度并关闭 UCX 后台进展 Sep 29, 2026
@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

中文:同步 main,并将原有 UCX setup 选择迁移到 Python launcher。[skip-sweep]
@edwingao28
edwingao28 marked this pull request as ready for review October 2, 2026 19:23
@edwingao28
edwingao28 requested a review from a team October 2, 2026 19:23
@claude

claude Bot commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @edwingao28's task in 47s —— View job


LGTM, no blocking issues found.

I followed the setup-script choice from start to finish. lanes.config_file() passes CONFIG_FILE to setup_script(), which removes the :override suffix and looks the recipe up in recipe_setup_scripts. The four GLM-5.2 disaggregated points in nvidia-master.yaml use recipes/glm5.2/.../disagg-*-mtp-variants.yaml, and those match the keys exactly. These GLM points have no EVAL_CONFIG_FILE, so eval-only runs pick the same script. Every other recipe and framework still gets install-torchao.sh or no script, as before.

checkout.py copies glm52-gb200-nixl-prefill.sh into the srt-slurm configs/ directory. The patch inside it is guarded by SHA-256 checks before and after, so it can safely run twice, and it fails loudly if the image changes. The changelog entry is added at the end of the file, and all four config keys exist in nvidia-master.yaml.

I couldn't run pytest in this environment because running uv run needs approval.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I also checked whether the master config (nvidia-master.yaml) needed a matching update for the SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE flip — it doesn't set that env var for these four recipe keys, so there's no master-config drift to fix.

Extended reasoning...

The diff touches four GLM-5.2 GB200 recipe YAMLs (compliant 1->0 flip of a prohibited draft-precision flag), a new hash-guarded NIXL conn.py patch script plus launcher plumbing (lanes.py/submit.py) to select it for two disaggregated recipes, a pending-approval engine-patch waiver doc, new driver tests, and a perf-changelog append. An inline finding already flags a policy-violating Chinese description line in the new perf-changelog entry, so a human look is warranted regardless; I additionally verified there's no corresponding nvidia-master.yaml entry for the flag that would need updating in lockstep, ruling out that specific cross-file consistency concern.

Comment thread inferencex-e2e/perf-changelog.yaml Outdated
Comment on lines +9247 to +9249
- "Keep the GLM-5.2 NextN/MTP draft at shipped precision and disable background UCX progress for GB200 disaggregated prefill to avoid the observed event-arm crash; images, topology, workload and golden acceptance are unchanged."
- "使 GLM-5.2 NextN/MTP draft 保持原始发布精度,并关闭 GB200 分离式 prefill 的 UCX 后台进展以规避已观察到的 event-arm 崩溃;镜像、拓扑、工作负载和 golden acceptance 不变。"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3401

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) This new perf-changelog entry adds a Chinese description line, which AGENTS.md:126 explicitly prohibits for new entries (English-only, no bilingual descriptions), unlike the rest of the bilingual-docs policy. Fix: remove the Chinese description string and keep only the English description line in this appended entry, consistent with every other entry in the file.

Why this was flagged

AGENTS.md:126 states 'New inferencex-e2e/perf-changelog.yaml entries must be English-only. Do not add Chinese translations or bilingual descriptions.' The new entry appended at inferencex-e2e/perf-changelog.yaml:9247-9249 includes both an English description and a Chinese description ('使 GLM-5.2 NextN/MTP draft 保持原始发布精度...') under the same description: list. This is a direct violation of the file's documented English-only invariant, which the rest of the bilingual-docs policy (AGENTS.md:20) does not override since the comment at line 126 explicitly carves this file out. No other check in the diff catches this since the file is otherwise append-only/byte-sensitive and no linter is shown running on it.

Verification: nit. The diff appends a new perf-changelog entry whose description list contains both an English line and a Chinese translation line, under the PR 3401 entry in perf-changelog.yaml. AGENTS.md:126 states new entries must be English-only, with no Chinese translations or bilingual descriptions. Neither validate_perf_changelog.py nor validation.py contains any bilingual check, so the line causes no validation failure. Fix is to remove the Chinese description string, keeping only the English line.

@edwingao28

Copy link
Copy Markdown
Collaborator Author

/reuse-sweep-run 36538876933

@functionstackx functionstackx left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why nixl patch and why so many chnages outside of draft preicison

使用上游 NIXL 修复 GLM GB200 进展路径并同步主分支。
@edwingao28 edwingao28 changed the title [GLM-5.2 GB200] Preserve draft precision and disable background UCX progress / [GLM-5.2 GB200] 保留 draft 精度并关闭 UCX 后台进展 [GLM-5.2 GB200] Preserve draft precision and use upstream NIXL 1.4.0 / [GLM-5.2 GB200] 保留 draft 精度并使用上游 NIXL 1.4.0 Oct 5, 2026
@edwingao28 edwingao28 added full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) and removed full-sweep-enabled Full sweep with canary gate; matrix jobs run to completion despite failures labels Oct 5, 2026
中文:同步主分支并保留 GLM GB200 已验证配置。
@edwingao28

Copy link
Copy Markdown
Collaborator Author

/use 37366783142

@github-actions

github-actions Bot commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

@edwingao28 staged run 37366783142: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-10-05~r37366783142

This run remains available across future /stage-results requests. Staging the same run ID again updates its staged data. Staging workflow

中文:固定校验归档测试用例标识,避免并行收集随 gzip 时间戳变化。
中文:升级 GLM GB200 分离式镜像并移除启动时 NIXL 覆盖。
@edwingao28 edwingao28 removed the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Oct 6, 2026
@edwingao28 edwingao28 changed the title [GLM-5.2 GB200] Preserve draft precision and use upstream NIXL 1.4.0 / [GLM-5.2 GB200] 保留 draft 精度并使用上游 NIXL 1.4.0 [GLM-5.2 GB200] Preserve draft precision and upgrade disaggregated image / [GLM-5.2 GB200] 保留 draft 精度并升级分离式镜像 Oct 6, 2026
@edwingao28 edwingao28 added the full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended) label Oct 6, 2026
中文:同步主分支并解决 GLM GB200 审阅分支冲突。
@edwingao28

Copy link
Copy Markdown
Collaborator Author

/use 37506762704

@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

@edwingao28 staged run 37506762704: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-10-06~r37506762704

This run remains available across future /stage-results requests. Staging the same run ID again updates its staged data. Staging workflow

cquil11
cquil11 previously requested changes Oct 7, 2026
- "Run decoder SWA bounded replay with prefill graphs disabled on every point."
- "Throughput keeps the committed thinking-on golden AL 3.51 selected by the srt connector; evals use real verification."
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3696

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

why are there 2?

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

first is txt version patch to NIXL
second is upgrade SGLang image that alr includes NICL 1.4.0

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

removed first one

@functionstackx functionstackx left a comment •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

remove one of the perf changelog entries before you merge

合并 GLM GB200 的性能变更记录。
将 GLM GB200 的变更记录修复同步到最新主分支。
@edwingao28
edwingao28 dismissed cquil11’s stale review October 7, 2026 16:13

fixed as requested

@edwingao28
edwingao28 merged commit a789327 into main Oct 7, 2026
30 checks passed
@edwingao28
edwingao28 deleted the fix/glm52-gb200-draft-quantization-off branch October 7, 2026 16:13
seungrokj added a commit that referenced this pull request Oct 8, 2026
Resolve conflicts:
- inferencex-e2e/docs/configuration-procedures{,_zh}.md: accept deletion
  from main (#3789 removed the configuration procedure guides); the mono
  decode details remain in this PR's perf-changelog entry.
- inferencex-e2e/perf-changelog.yaml: keep main's new entries (#3401,
  #3692, #3685) and append this PR's entry after them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
seungrokj added a commit that referenced this pull request Oct 8, 2026
Resolve conflicts:
- inferencex-e2e/docs/MODELS{,_zh}.md: keep this PR's DeepSeek-V4.1-Flash
  row (ATOM restored on srt-slurm 2026-09-30) and main's GLM-5.3 row
  ("deprecation policy" wording).
- inferencex-e2e/docs/configuration-procedures{,_zh}.md: accept deletion
  from main (#3789 removed the configuration procedure guides).
- inferencex-e2e/perf-changelog.yaml: keep main's new entries (#3401,
  #3692, #3685) and append this PR's entry after them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

full-sweep-fail-fast Full sweep with canary gate; first failure cancels the rest of that matrix (recommended)

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

4 participants