Qwen3.5 FP4 GB300 disaggregated Dynamo-SGLang: update SGLang to v0.5.14-cu130 and NVFP4-V2 / Qwen3.5 FP4 GB300 分离式 Dynamo-SGLang:将 SGLang 更新至 v0.5.14-cu130 并切换到 NVFP4-V2 - #2238
Conversation
… NVFP4-V2 checkpoint
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
| image: lmsysorg/sglang:nightly-dev-cu13-20260624-b2c8f7a2 | ||
| model: nvidia/Qwen3.5-397B-A17B-NVFP4 | ||
| image: lmsysorg/sglang:v0.5.14-cu130 | ||
| model: nvidia/Qwen3.5-397B-A17B-NVFP4-V2 |
There was a problem hiding this comment.
what is difference between nvidia/Qwen3.5-397B-A17B-NVFP4 & nvidia/Qwen3.5-397B-A17B-NVFP4-v2
functionstackx
left a comment
There was a problem hiding this comment.
Are yall seeing lots of customer traction for qwen3.5 hence why modelopt team has pritotized and used their limited engineering bandwidth on an v2 quant on qwen3.5?
+viz @Ankur-singh
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=29473335540 |
|
/reuse-sweep-run 29473335540 |
|
As a PR reviewer and CODEOWNER, I have reviewed this and have:
Additional detail section:
Signed: |
✅✅✅ Verdict: PASS ✅✅✅✅ Check 0 (CODEOWNER): PASS — |
|
recipes PR merged, should this go in? |
Summary
qwen3.5-fp4-gb300-dynamo-sglangfromlmsysorg/sglang:nightly-dev-cu13-20260624-b2c8f7a2to the upstreamlmsysorg/sglang:v0.5.14-cu130image.nvidia/Qwen3.5-397B-A17B-NVFP4tonvidia/Qwen3.5-397B-A17B-NVFP4-V2, and updates the GB300 launcher model path accordingly.Topologies
Changed paths
configs/nvidia-master.yamlrunners/launch_gb300-nv.shbenchmarks/multi_node/srt-slurm-recipes/sglang/qwen3.5/gb300-fp4/8k1k/disagg/stp/perf-changelog.yamlValidation
38d77d67b74131d77a2ec4ffeda6f200820a1bbf; all four applicable multi-node throughput jobs and all four matching multi-node eval jobs ran successfully.lmsysorg/sglang:v0.5.14-cu130andnvidia/Qwen3.5-397B-A17B-NVFP4-V2.Open review item
CHANGES_REQUESTEDreview asks for customer-traction evidence and the rationale for prioritizing the V2 quantization checkpoint. That rationale is not present in the current PR body or GitHub review thread.中文说明
qwen3.5-fp4-gb300-dynamo-sglang使用的镜像从lmsysorg/sglang:nightly-dev-cu13-20260624-b2c8f7a2更新为上游镜像lmsysorg/sglang:v0.5.14-cu130。nvidia/Qwen3.5-397B-A17B-NVFP4切换为nvidia/Qwen3.5-397B-A17B-NVFP4-V2,并同步更新 GB300 启动器中的模型路径。拓扑
变更路径
configs/nvidia-master.yamlrunners/launch_gb300-nv.shbenchmarks/multi_node/srt-slurm-recipes/sglang/qwen3.5/gb300-fp4/8k1k/disagg/stp/下的 4 个配方 YAMLperf-changelog.yaml验证
38d77d67b74131d77a2ec4ffeda6f200820a1bbf上成功完成;4 个适用的多节点吞吐量任务和 4 个对应的多节点评估任务均已实际运行并通过。lmsysorg/sglang:v0.5.14-cu130和nvidia/Qwen3.5-397B-A17B-NVFP4-V2。待确认审阅问题
CHANGES_REQUESTED审阅要求补充客户使用情况,以及优先投入 V2 量化检查点的依据。现有 PR 描述和 GitHub 审阅线程中尚未提供相关说明。