Conversation
|
Thanks for the contribution!
中文感谢你的贡献!
|
| - "Rely on the pinned SGLang image's defaults for JIT norm, Top-K v2, SWA leaf splitting, unified-cache out-of-window reclamation, and custom all-reduce v2 instead of overriding them in the recipes." | ||
| - "新增 B200 Dynamo+SGLang AgentX TP8 聚合式 HiCache 配方,并发为 1、4、8;使用 lmsysorg/sglang:nightly-dev-20260916-c9a8fba9、检查点内置 DSpark K=6、整节点 CPU/DRAM 分配,以及按数据点设置的请求数/CUDA graph 上限。" | ||
| - "依赖固定 SGLang 镜像对 JIT norm、Top-K v2、SWA leaf splitting、unified-cache 窗口外槽位回收和 custom all-reduce v2 的默认行为,不在配方中覆盖这些设置。" | ||
| pr-link: TBD |
There was a problem hiding this comment.
🔴 Both new changelog entries use pr-link: TBD (lines 8600 and 8611) instead of the required XXX placeholder, so the changelog gate CI will reject this PR. validate_added_pr_link in infx/workflows/validate_perf_changelog.py:134-145 only accepts PR_LINK_PLACEHOLDERS = {"XXX", ".../pull/XXX"} or the exact https://github.com/SemiAnalysisAI/InferenceX/pull/<PR#> URL; "TBD" matches neither, so validate_added_pr_link raises ChangelogValidationError("new PR entry must use ... or an XXX placeholder; found 'TBD'"). …
Why this was flagged
…Fix: replace both pr-link: TBD occurrences with pr-link: XXX so infx.workflows.validate_perf_changelog (run in .github/workflows/run-sweep.yml:117-125) and the later PR-number substitution in infx/workflows/prepare_perf_changelog_merge.py pass, covering both new entries at perf-changelog.yaml:8600 and perf-changelog.yaml:8611.
The trigger is CI running infx.workflows.validate_perf_changelog (invoked from .github/workflows/run-sweep.yml:117-125, 'Validate perf-changelog matrix') on this PR's diff of perf-changelog.yaml. validate_added_pr_link (infx/workflows/validate_perf_changelog.py:134-145) is called for each appended entry with its pr-link string. It checks CANONICAL_PR_LINK regex or membership in PR_LINK_PLACEHOLDERS = {"XXX", "https://github.com/SemiAnalysisAI/InferenceX/pull/XXX"} (lines 20-24). 'TBD' is neither, so it raises ChangelogValidationError with message "new PR entry must use 'https://github.com/SemiAnalysisAI/InferenceX/pull/' or an XXX placeholder; found 'TBD'", failing the workflow that main() (line 336) turns into exit code 1. On base main, entries follow…
Verification: normal (mechanism corrected) — Both new appended entries end with pr-link: TBD (the added block at perf-changelog.yaml:8587+, agg entry ~line 8600 and disagg entry ~line 8611). The repo convention/tooling requires the literal XXX placeholder until the PR exists (CONTRIBUTING.md:180; .github/workflows/claude.yml:235). "TBD" is not accepted: PR_LINK_PLACEHOLDERS = {"XXX",… | Severity:…
| - "Use automatic golden acceptance-length selection and rely on the pinned SGLang image's defaults for SWA leaf splitting, unified-cache out-of-window reclamation, and custom all-reduce v2." | ||
| - "新增 B200 Dynamo+SGLang AgentX DEP8/DEP8 HiCache 配方:1P1D 并发 64、128,以及 2P1D 并发 256;decode max-running-requests 分别限制为 128、256、512;1P1D 的 decode swa-full-tokens-ratio 为 0.01,2P1D c256 为 0.005 且 memory fraction 为 0.91。" | ||
| - "使用自动黄金接受长度选择,并依赖固定 SGLang 镜像对 SWA leaf splitting、unified-cache 窗口外槽位回收和 custom all-reduce v2 的默认行为。" | ||
| pr-link: TBD |
There was a problem hiding this comment.
🟡 (optional) The new perf-changelog entry for dsv4-fp4-b200-dynamo-sglang-agentic-disagg misstates the tuned decode value: it says "decode swa-full-tokens-ratio 0.01 for 1P1D" but the actual decode arg in both 1P1D recipes is 0.02 (0.01 is the prefill value). Fix: correct the description to state decode swa-full-tokens-ratio 0.02 for 1P1D (prefill uses 0.01), so the changelog matches disagg-b200-1p1d-dep8-dep8-c64-mtp-kvoffload.yaml and disagg-b200-1p1d-dep8-dep8-c128-mtp-kvoffload.yaml.
Why this was flagged
perf-changelog.yaml is the record maintainers rely on to know what tuning a recipe uses without opening the YAML. The new entry (perf-changelog.yaml line ~8611) claims decode swa-full-tokens-ratio 0.01 for the 1P1D recipes, but disagg-b200-1p1d-dep8-dep8-c64-mtp-kvoffload.yaml:194 and disagg-b200-1p1d-dep8-dep8-c128-mtp-kvoffload.yaml:194 both set decode swa-full-tokens-ratio: 0.02; 0.01 is actually the prefill value at line 122 in each file. A maintainer trusting the changelog to diff or reproduce the decode tuning gets the wrong number, since nothing cross-checks changelog prose against recipe args.
Verification: nit. The changelog description is factually wrong but nothing functional breaks. In disagg-b200-1p1d-dep8-dep8-c64-mtp-kvoffload.yaml the prefill worker (disaggregation-mode: prefill at line 117) sets swa-full-tokens-ratio: 0.01 at line 122, while the decode worker (disaggregation-mode: decode at line 190) sets swa-full-tokens-ratio: 0.02 at line 194.… | nit. The changelog entry at…
There was a problem hiding this comment.
[by Codex] Fixed in 96dadb0. The changelog now states the actual 1P1D values: prefill swa-full-tokens-ratio: 0.01 and decode swa-full-tokens-ratio: 0.02; the 2P1D c256 decode value remains 0.005.\n\n
中文
\n\n已在 96dadb0 中修复。changelog 现准确记录 1P1D:prefill 为swa-full-tokens-ratio: 0.01,decode 为 swa-full-tokens-ratio: 0.02;2P1D c256 的 decode 值仍为 0.005。\n\n|
View unofficial run (performance): https://inferencex.semianalysis.com/inference?unofficialRun=36517054431 View unofficial run (accuracy): https://inferencex.semianalysis.com/evaluation?unofficialRun=36517054431 |
|
As a PR reviewer and CODEOWNER, I have reviewed this and have:
Additional detail section:
Signed: |
❌❌❌ REJECTED ❌❌❌@kedarpotdar-nv — two blocking issues at ❌ Check 4 (Reuse command): FAIL — No authorized reuse command has been posted on this PR. The only comments are two bot notices and the sign-off; an ❌ Check 13 (Draft runs as shipped): FAIL —
Passed and not applicable checks✅ Check 0 (CODEOWNER): PASS — ✅ Check 1 (Passing sweep on in-PR commit): PASS — head ✅ Check 2 (Evals pass): PASS — ➖ Check 3 (Recipe linked/merged/complete): N/A — disaggregated/multi-node submission; both entries are ✅ Check 5 (Latest checklist template): PASS — all 14 items plus the merged-upstream sub-item of the current ✅ Check 6 (Upstream image / engine-first): PASS — image is upstream ✅ Check 7 (No deprecated models/scenarios): PASS — ✅ Check 8 (No architecture hacks): PASS — no ✅ Check 9 (Spec-decode via chat template): PASS — recipes run ✅ Check 10 (No engine patches): PASS — no ✅ Check 11 (Agentic spec-decode golden AL): PASS — ➖ Check 12 (Append-only): N/A — neither new Assessed commit: |
|
Sorry, over the weekend, there was 2 major refactors to clean up the technical debt accumalated over the past 11 months of moving at the speed of light. We don't see any major refactors in the forthseeable future besides cleaning up AMD multinode AgentX pile of bash. As much, due to the refactors, u would need to ask your agent to rebase from remote main@latest. Thank you in advance for ur understanding |
将 B200 AgentX 配方与最新主分支对齐,移除不必要的环境变量和合成接受长度覆盖,并恢复 DSpark draft 的原始激活精度。
96dadb0 to
05dfe3f
Compare
…ult-envs-v2 # Conflicts: # inferencex-e2e/perf-changelog.yaml
修复重构后未设置的 NSCALE_MODEL_ROOT,恢复已验证的 B200 DeepSeek-V4-Pro-0813 检查点路径。
Summary
This PR updates the six DSV4-Pro B200 Dynamo+SGLang AgentX points already merged in #3257, now under
inferencex-e2e/after the repository refactor. It keeps aggregate TP8 at c1/c4/c8, 1P1D DEP8/DEP8 at c64/c128, and 2P1D DEP8/DEP8 at c256.enable-w4a4-mxfp4-megamoefrom both disaggregated serving roles. This avoids lowering the bundled DSpark draft's activations below the shipped FP8 path, as identified by the CODEOWNER verifier./scratch/models/DeepSeek-V4-Pro-0813path on the NScale launcher. The refactor leftNSCALE_MODEL_ROOTunset and caused the first rebased canary to fail before model load.Validation
main.AI model disclosure
OpenAI GPT-5 (Codex) prepared the original PR. Codex prepared this rebase and validation; the exact model identifier for this update is not exposed by the runtime and could not be verified. No delegated agents were used for this update.
中文
摘要
本 PR 更新已通过 #3257 合并的六个 DSV4-Pro B200 Dynamo+SGLang AgentX 数据点。仓库重构后,相关文件位于
inferencex-e2e/。保持聚合式 TP8 并发 1/4/8、1P1D DEP8/DEP8 并发 64/128,以及 2P1D DEP8/DEP8 并发 256。enable-w4a4-mxfp4-megamoe,避免内置 DSpark draft 的激活精度低于发布时的 FP8 路径;此问题由 CODEOWNER 验证器指出。/scratch/models/DeepSeek-V4-Pro-0813路径。重构后NSCALE_MODEL_ROOT未设置,导致重整基底后的首个 canary 在模型加载前失败。验证
main的性能变更日志追加验证通过。AI 模型披露
OpenAI GPT-5(Codex)准备了原始 PR。Codex 完成了本次重整基底及验证;运行环境未提供本次更新所用模型的准确标识,因此无法核实。此次更新未使用委派代理。