Repository navigation
[GLM-5.2 GB200] Preserve draft precision and upgrade disaggregated image / [GLM-5.2 GB200] 保留 draft 精度并升级分离式镜像 - #3401
Conversation
|
Thanks for the contribution!
中文感谢你的贡献!
|
91efe8b to
e908cc4
Compare
|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=37506762704 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=37506762704 |
68d3ce2 to
274c69d
Compare
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
6113b30 to
fca8de5
Compare
…refill Keep the draft at shipped precision and disable only the two background UCX progress controls for disaggregated prefill. Preserve images, workload, strict synchronization and golden acceptance. Request the scoped engine-patch waiver. 保留 GLM-5.2 GB200 draft 的原始精度,仅关闭分离式 prefill 的两项 UCX 后台进展控制;保留镜像、工作负载、严格同步及 golden acceptance,并申请对应补丁例外。
fca8de5 to
9fbe9bf
Compare
|
Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge |
中文:同步 main,并将原有 UCX setup 选择迁移到 Python launcher。[skip-sweep]
|
Claude finished @edwingao28's task in 47s —— View job LGTM, no blocking issues found. I followed the setup-script choice from start to finish.
I couldn't run |
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline finding, I also checked whether the master config (nvidia-master.yaml) needed a matching update for the SGLANG_NVFP4_CKPT_FP8_NEXTN_MOE flip — it doesn't set that env var for these four recipe keys, so there's no master-config drift to fix.
Extended reasoning...
The diff touches four GLM-5.2 GB200 recipe YAMLs (compliant 1->0 flip of a prohibited draft-precision flag), a new hash-guarded NIXL conn.py patch script plus launcher plumbing (lanes.py/submit.py) to select it for two disaggregated recipes, a pending-approval engine-patch waiver doc, new driver tests, and a perf-changelog append. An inline finding already flags a policy-violating Chinese description line in the new perf-changelog entry, so a human look is warranted regardless; I additionally verified there's no corresponding nvidia-master.yaml entry for the flag that would need updating in lockstep, ruling out that specific cross-file consistency concern.
| - "Keep the GLM-5.2 NextN/MTP draft at shipped precision and disable background UCX progress for GB200 disaggregated prefill to avoid the observed event-arm crash; images, topology, workload and golden acceptance are unchanged." | ||
| - "使 GLM-5.2 NextN/MTP draft 保持原始发布精度,并关闭 GB200 分离式 prefill 的 UCX 后台进展以规避已观察到的 event-arm 崩溃;镜像、拓扑、工作负载和 golden acceptance 不变。" | ||
| pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3401 |
There was a problem hiding this comment.
🟡 (optional) This new perf-changelog entry adds a Chinese description line, which AGENTS.md:126 explicitly prohibits for new entries (English-only, no bilingual descriptions), unlike the rest of the bilingual-docs policy. Fix: remove the Chinese description string and keep only the English description line in this appended entry, consistent with every other entry in the file.
Why this was flagged
AGENTS.md:126 states 'New inferencex-e2e/perf-changelog.yaml entries must be English-only. Do not add Chinese translations or bilingual descriptions.' The new entry appended at inferencex-e2e/perf-changelog.yaml:9247-9249 includes both an English description and a Chinese description ('使 GLM-5.2 NextN/MTP draft 保持原始发布精度...') under the same description: list. This is a direct violation of the file's documented English-only invariant, which the rest of the bilingual-docs policy (AGENTS.md:20) does not override since the comment at line 126 explicitly carves this file out. No other check in the diff catches this since the file is otherwise append-only/byte-sensitive and no linter is shown running on it.
Verification: nit. The diff appends a new perf-changelog entry whose description list contains both an English line and a Chinese translation line, under the PR 3401 entry in perf-changelog.yaml. AGENTS.md:126 states new entries must be English-only, with no Chinese translations or bilingual descriptions. Neither validate_perf_changelog.py nor validation.py contains any bilingual check, so the line causes no validation failure. Fix is to remove the Chinese description string, keeping only the English line.
|
/reuse-sweep-run 36538876933 |
functionstackx
left a comment
There was a problem hiding this comment.
why nixl patch and why so many chnages outside of draft preicison
使用上游 NIXL 修复 GLM GB200 进展路径并同步主分支。
中文:同步主分支并保留 GLM GB200 已验证配置。
|
/use 37366783142 |
|
@edwingao28 staged run 37366783142: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-10-05~r37366783142 This run remains available across future |
中文:固定校验归档测试用例标识,避免并行收集随 gzip 时间戳变化。
中文:升级 GLM GB200 分离式镜像并移除启动时 NIXL 覆盖。
中文:同步主分支并解决 GLM GB200 审阅分支冲突。
|
/use 37506762704 |
|
@edwingao28 staged run 37506762704: https://inferencemax-app-git-staging-semianalysisai.vercel.app/inference?i_dates=2026-10-06~r37506762704 This run remains available across future |
| - "Run decoder SWA bounded replay with prefill graphs disabled on every point." | ||
| - "Throughput keeps the committed thinking-on golden AL 3.51 selected by the srt connector; evals use real verification." | ||
| pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/3696 | ||
|
|
There was a problem hiding this comment.
first is txt version patch to NIXL
second is upgrade SGLang image that alr includes NICL 1.4.0
There was a problem hiding this comment.
removed first one
合并 GLM GB200 的性能变更记录。
将 GLM GB200 的变更记录修复同步到最新主分支。
Resolve conflicts:
- inferencex-e2e/docs/configuration-procedures{,_zh}.md: accept deletion
from main (#3789 removed the configuration procedure guides); the mono
decode details remain in this PR's perf-changelog entry.
- inferencex-e2e/perf-changelog.yaml: keep main's new entries (#3401,
#3692, #3685) and append this PR's entry after them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Resolve conflicts:
- inferencex-e2e/docs/MODELS{,_zh}.md: keep this PR's DeepSeek-V4.1-Flash
row (ATOM restored on srt-slurm 2026-09-30) and main's GLM-5.3 row
("deprecation policy" wording).
- inferencex-e2e/docs/configuration-procedures{,_zh}.md: accept deletion
from main (#3789 removed the configuration procedure guides).
- inferencex-e2e/perf-changelog.yaml: keep main's new entries (#3401,
#3692, #3685) and append this PR's entry after them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Description
Preserve GLM-5.2 GB200 draft precision and upgrade both disaggregated recipes to digest-pinned upstream SGLang 0.5.18 CUDA 13, removing the startup NIXL override.
Testing: Sweep 37506762704: 14 performance + 4 evals passed and artifacts reviewed at
d8e7cc8; CPU CI passed. Reuse selected.Limits: Cancellations and cleanup/identity warnings remain. Main's newer launcher is not newly measured; CODEOWNER re-review remains required.
中文
保留 GLM-5.2 GB200 draft 原始精度,将两个分离式配方升级到固定 digest 的上游 SGLang 0.5.18 CUDA 13 镜像,移除启动时的 NIXL 覆盖。
测试: 扫描 37506762704 的 14 个性能点和 4 个 eval 全部通过,并完成
d8e7cc8的产物审阅;CPU CI 通过,已选择复用。限制: 保留请求取消、清理和身份核对警告。同步引入的新版主分支 launcher 尚未重新实测;仍需 CODEOWNER 复审。
AI 模型: Claude Opus 5.5 (
claude-opus-5-5) 负责原始实现;GPT-6(具体变体不可确认)负责更新、冲突修复、验证和委派审阅。关联 #3228;修复及配置变更。AI model disclosure
claude-opus-5-5): original implementation.Related Issue
Related to #3228.
Type of Change
Checklist
inferencex-e2e/perf-changelog.yamland have not edited historical entriesfull-sweep-fail-fast(recommended),full-sweep-enabled, ornon-canary-full-sweep-enabled. Optional modifiersall-evals,evals-only, andagentx-fastrequire a primary label; the last two block reuse while applied.OWNER/MEMBER/COLLABORATOR) has commented/use <run_id>(or the legacy/reuse-sweep-run) on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the primary sweep label will no longer automatically kick off new sweeps. Remove and re-add the primary sweep label to force a new sweep.