Skip to content

Add DSV4 B200 disaggregated Dynamo SGLang MTP configuration / 新增 DSV4 B200 分离式 Dynamo SGLang MTP 配置 - #2554

Merged
adibarra merged 3 commits into
mainfrom
dsv4-fp4-b200-dynamo-sglang-mtp-nscale
Aug 12, 2026
Merged

Add DSV4 B200 disaggregated Dynamo SGLang MTP configuration / 新增 DSV4 B200 分离式 Dynamo SGLang MTP 配置#2554
adibarra merged 3 commits into
mainfrom
dsv4-fp4-b200-dynamo-sglang-mtp-nscale

Conversation

@RohitNagraj

@RohitNagraj RohitNagraj commented Aug 11, 2026

Copy link
Copy Markdown
Collaborator

Description

Add a DSV4 B200 disaggregated Dynamo SGLang MTP configuration for 8k/1k.

  • Add five checked-in srt-slurm recipes with EAGLE speculative decoding and chat-formatted benchmark inputs.
  • Use the staged DSV4 checkpoint native quantization metadata and the checkpoint-compatible MoE backend.
  • Use NIXL with UCX RC transport for KV state transfer.
  • Route multi-node runs through the dedicated B200 launcher and pin srt-slurm to the validated revision.

Validation:

  • Generated all five throughput configurations and four evaluation configurations.
  • Validated launcher shell syntax, YAML and srtctl schemas, recipe/config consistency, chat-template usage, and changelog ordering.
  • Passed the matrix-logic test suite.

中文说明

新增 DSV4 在 B200 上的分离式 Dynamo SGLang MTP 8k/1k 配置。

  • 新增五个已检入的 srt-slurm recipe,启用 EAGLE 投机解码和聊天模板格式的基准测试输入。
  • 使用已暂存 DSV4 检查点原生的量化元数据,并选择与检查点兼容的 MoE 后端。
  • 使用 NIXL 和 UCX RC 传输 KV 状态。
  • 通过专用 B200 启动器路由多节点作业,并将 srt-slurm 固定到已验证的版本。

验证:

  • 已成功生成五个吞吐量配置和四个评估配置。
  • 已验证启动器 Shell 语法、YAML 与 srtctl schema、recipe 与配置一致性、聊天模板使用情况以及变更日志顺序。
  • 已通过矩阵逻辑测试套件。

Related Issue

N/A

Type of Change

  • Bug fix
  • New feature
  • Configuration change
  • Documentation update
  • Other (please describe)

Checklist

  • I have tested my changes locally
  • I have updated documentation if necessary
  • For every change that can affect benchmark performance and every recipe addition or modification, I have appended a new entry to the physical end of perf-changelog.yaml and have not edited historical entries
  • Before merging via reuse, an authorized maintainer (OWNER/MEMBER/COLLABORATOR) has commented /reuse-sweep-run on this PR. Do this only once there is a final full sweep that is all green with evals passing, since after this comment the sweep label will no longer automatically kick off new sweeps. Remove and re-add the label to force one.

新增 DSV4 B200 nscale UCX 验证配置。
@RohitNagraj
RohitNagraj force-pushed the dsv4-fp4-b200-dynamo-sglang-mtp-nscale branch from 328eff9 to 0555089 Compare August 11, 2026 01:19
@RohitNagraj
RohitNagraj marked this pull request as ready for review August 11, 2026 01:20
@RohitNagraj
RohitNagraj requested a review from a team August 11, 2026 01:20
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@claude

claude Bot commented Aug 11, 2026

Copy link
Copy Markdown
Contributor

Claude finished @RohitNagraj's task in 6m 51s —— View job


PR Review — DSV4 FP4 B200 nscale Dynamo-SGLang MTP UCX canary

  • Gather context and read the full diff
  • Review recipe YAML (disagg-b200-mid-curve-1p2d-dep8-dep8-mtp.yaml)
  • Review configs/nvidia-master.yaml changes (image accessibility, config key format)
  • Review perf-changelog.yaml (entry position, present alongside master config change)
  • Review runners/launch_b200-nscale-slurm.sh (enroot import, routing logic)
  • Post review feedback

Summary

No blocking issues found. Two 🟡 warnings posted inline, both on the new recipe: UCX_PROTO_EMULATION_ENABLE does not appear to be a real UCX config variable (the canary's "host-staged protocol emulation" knob may be a no-op — host staging is still forced by UCX_IB_GPU_DIRECT_RDMA=no + cuda_copy, and the already-enabled UCX_PROTO_INFO=y output will show what was actually selected), and cpus-per-task: "144" looks copied from Grace-based GB200/B300 recipes and may be unsatisfiable on the x86 nscale B200 nodes, where no existing recipe sets that directive.

Checks that passed: perf-changelog.yaml updated alongside the master config and appended at the end of the file; image is a public Docker Hub reference matching the recipe's container:; launcher keeps the enroot import docker:// reproducibility pattern and its new framework gate is logically correct (SPEC_DECODING is exported by benchmark-multinode-tmpl.yml); the recipe includes use_chat_template: true for MTP; the dsv4 prefix, dynamo-sglang framework value, and config schema all match established entries; the dynamo hash in the recipe matches the router version in the master config.

Caveat: this sandbox has no network access, so I could not verify externally that the Docker Hub tag nightly-dev-cu13-20260710-cfc66e05 exists, that the pinned srt-slurm commit 04e87fcc is reachable from main (the --single-branch main clone requires it), or the UCX variable question above — the PR's stated manifest/matrix validations presumably cover the first two.


# Keep RC for inter-node transport, but stage CUDA buffers through host memory.
UCX_TLS: "rc,cuda_ipc,cuda_copy,self"
UCX_IB_GPU_DIRECT_RDMA: "no"
UCX_PROTO_EMULATION_ENABLE: "y"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 WARNING: UCX_PROTO_EMULATION_ENABLE does not appear to be a real UCX configuration variable
Why it matters: Mainline UCX exposes UCX_PROTO_ENABLE / UCX_PROTO_INFO, and proto-v2 "emulation" protocols are selected automatically when a transport lacks a capability — I can find no PROTO_EMULATION_ENABLE knob in the UCX config table. Unknown UCX_* vars are silently ignored (at most a "unused env var" warning), so the run will proceed, but the canary's stated third knob ("enable host-staged protocol emulation") would be a no-op and conclusions drawn from this arm could be misattributed. Host staging is still forced here by UCX_IB_GPU_DIRECT_RDMA=no + cuda_copy, so the experiment isn't broken — it just may not test what the PR description claims. (Same variable appears in the decode block at line 77.)
Fix: Verify the variable exists in the container's UCX build (ucx_info -c | grep -i proto), or rely on UCX_PROTO_INFO=y output (already enabled) to confirm which protocol was actually selected, and drop the variable if it's not recognized. Note: I could not verify against UCX source/docs from this sandbox (network access blocked), so please double-check before drawing conclusions from the run.

Comment on lines +17 to +19
sbatch_directives:
cpus-per-task: "144"
mem: "0"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 WARNING: cpus-per-task: "144" looks copied from GB200/B300 recipes and may be unsatisfiable on nscale B200 nodes
Why it matters: 144 matches GB200/GB300 Grace nodes (2×72 cores) — see e.g. disagg-gb200-mid-curve-1p1d-dep8-dep16-mtp.yaml, which this recipe otherwise mirrors. The B200 nscale nodes are x86 (the sibling dsv4-fp4-b200-dynamo-vllm-mtp recipes use the vllm-x86_64 image), and none of the existing B200 nscale recipes set cpus-per-task (they use sbatch_directives: segment: "1" instead, e.g. benchmarks/multi_node/srt-slurm-recipes/vllm/deepseek-v4/8k1k/disagg-b200-1p1d-dep8-dep8-mtp.yaml:15). If the nscale nodes have fewer than 144 cores available per task, sbatch rejects the job with "Requested node configuration is not available" and the canary never starts.
Fix: Confirm the nscale B200 node core count supports --cpus-per-task=144, or drop the directive (and consider whether segment: "1" from the existing nscale recipes is needed) so Slurm uses its defaults.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. This is a "[Do Not Merge]" canary PR whose content is byte-for-byte identical to commit 0555089, already merged to main separately — the same recipe, master-config entry, changelog entry, and launcher gating were already landed and validated. No new risk introduced.

What was reviewed: the new SGLang recipe's UCX/host-staging env vars for prefill vs. decode, the launcher's expanded framework/spec-decoding gate logic in launch_b200-nscale-slurm.sh, and the srt-slurm checkout/recipe-copy branching for the dynamo-sglang path. The flagged missing SGLANG_OPT_USE_CUSTOM_ALL_REDUCE_V2=1 in prefill_environment was checked against sibling multi-node DSv4 recipes — CAR_V2 is single-node only and correctly omitted here since prefill runs on 1 node with DEP8.

Extended reasoning...

Overview

The PR adds one new benchmark recipe YAML for a DeepSeek-V4-Pro FP4 Dynamo-SGLang MTP run on B200 nscale (1P/2D DEP8, host-staged UCX transport), a corresponding entry in configs/nvidia-master.yaml, a perf-changelog.yaml entry, and an extension of runners/launch_b200-nscale-slurm.sh to route this new framework/spec-decoding combination and pin the srt-slurm revision used for it.

Security risks

None. This is CI/benchmark infrastructure config — no auth, crypto, or user-facing surface. The launcher changes only affect environment-variable gating and which pinned public repo revision is checked out for benchmark execution.

Level of scrutiny

Low-to-moderate is appropriate: this is a benchmark-recipe addition following an extremely well-established pattern in this repo (dozens of similar recipes exist), and it is explicitly marked '[Do Not Merge]' — a canary intended only to validate the config via CI, not to be merged into main as-is. Notably, the entire diff (all 4 files) is byte-for-byte identical to commit 0555089, which already landed on main from the same author. That means this exact change has already been reviewed/merged through a separate path, materially reducing residual risk here.

Other factors

The bug-hunting pass found no issues, and the one candidate concern raised (missing CAR_V2 flag in the new prefill_environment) was checked against sibling recipes and confirmed to be correct as written — CAR_V2 is intentionally single-node-only and prefill here runs on a single node. The launcher's new gating conditional was traced against SPEC_DECODING, which is a standard workflow-exported env var used identically across many other launcher scripts, so no wiring risk there either.

@github-actions

Copy link
Copy Markdown
Contributor

扩展 DSV4 B200 Dynamo-SGLang MTP 配置,补充五个分离式拓扑并保持配置与 recipe 一致。
@RohitNagraj RohitNagraj changed the title [Do Not Merge] Add DSV4 FP4 B200 nscale Dynamo-SGLang MTP UCX canary / 添加 DSV4 FP4 B200 nscale Dynamo-SGLang MTP UCX 验证配置 Add DSV4 B200 disaggregated Dynamo SGLang MTP configuration / 新增 DSV4 B200 分离式 Dynamo SGLang MTP 配置 Aug 11, 2026
@github-actions

Copy link
Copy Markdown
Contributor

@Ankur-singh

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run 31451666262

@Ankur-singh

Copy link
Copy Markdown
Collaborator

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this. https://github.com/SemiAnalysisAI/InferenceX/actions/runs/31451666262
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this. https://github.com/SemiAnalysisAI/InferenceX/actions/runs/31451666262
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified against the current MODELS.md that this PR does not submit a deprecated model, scenario, or model-scenario combination.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; FLOPs at lower precisions is fine, given that the config passes private evals. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.

Additional detail section:

Validation and eval evidence. Run 31451666262 ran on the exact PR head 8c15a0af, attempt 1, and settled success with 19 success / 11 skipped / 0 failures. All five multi-node 8k1k / benchmark jobs executed and passed — one per submitted recipe — alongside four green multi-node eval / jobs and a successful collect-evals. setup and canary-select both succeeded, so the matrix genuinely executed rather than being skipped under a reuse gate. The skipped lanes (single-node, agentic, agentic eval, eval /) have no matrix entries for this multi-node config.

DISAGG master↔recipe parity. Re-derived independently for all five variants rather than taken from any creator packet. Every nvidia-master.yaml variant agrees with its CONFIG_FILE= recipe, including the naming convention where dep denotes dp-attention enabled and tp denotes it disabled:

conc master prefill master decode recipe
1 1 x TP8/EP1 1 x TP8/EP1 disagg-b200-low-latency-1p1d-tp8-tp8-mtp
32, 64, 128 1 x TP8/EP8, dp-attn 6 x TP8/EP1 disagg-b200-low-latency-1p6d-dep8-tp8-mtp
256, 1024 1 x TP8/EP8, dp-attn 1 x TP8/EP8, dp-attn disagg-b200-mid-curve-1p1d-dep8-dep8-mtp
256 1 x TP8/EP8, dp-attn 2 x TP8/EP8, dp-attn disagg-b200-mid-curve-1p2d-dep8-dep8-mtp
6144 5 x TP8/EP8, dp-attn 3 x TP8/EP8, dp-attn disagg-b200-mid-curve-5p3d-dep8-dep8-mtp

Recipe resource blocks agree with the master worker counts on B200's 8-GPU nodes in every case (workers x 8 GPUs = nodes x 8; e.g. 5P3D declares prefill_nodes: 5 / prefill_workers: 5 / gpus_per_prefill: 8 and decode_nodes: 3 / decode_workers: 3 / gpus_per_decode: 8). KV transport is consistent end to end: master declares kv-p2p-transfer: nixl and all five recipes set disaggregation-transfer-backend: nixl on both the prefill and decode halves. Note that concurrency 256 is deliberately submitted twice under different topologies (1P1D and 1P2D), which is Pareto exploration rather than a duplicate.

Speculative decoding and chat template. All five recipes run EAGLE MTP with speculative-num-steps: 3, speculative-eagle-topk: 1, speculative-num-draft-tokens: 4, and — checked explicitly rather than assumed — every one of the five sets use_chat_template: true in its benchmark block. That matters for this submission specifically: acceptance length here is measured rather than simulated, so without chat-formatted prompts the AL distribution would be taken on raw-token inputs and overstate the speedup. The golden-AL item is not applicable because this is a fixed-seq-len 8k/1k submission rather than an agentic workload; no SGLANG_SIMULATE_ACC_* knobs appear anywhere in the diff, which is correct for this scenario.

Model and scenario scope. MODELS.md lists DeepSeek-V4-Pro as active for Single-turn 8k1k, with Single-turn 1k1k recorded as deprecated; this submission is 8k1k only and adds no 1k1k lane. MODELS.md also records the engine expectation for this model as the native/upstream vLLM and SGLang engines with native MTP, which is exactly what these recipes exercise on the upstream lmsysorg/sglang:nightly-dev-cu13-20260710-cfc66e05 image.

Single-node recipe publication — not applicable. All five recipes are multi-node srt-slurm recipes under benchmarks/multi_node/srt-slurm-recipes/** (multinode: true, disagg: true), so the single-node vLLM-recipes / SGLang-cookbook publication requirement does not apply to this PR.

No engine or serving-stack patching. The diff contains no .patch, git apply, sed -i, site-packages edit, monkey-patch, --hf-overrides, or forked/rebuilt engine wheel, and the pinned upstream SGLang image runs as shipped. launch_b200-nscale-slurm.sh extends its allow-list to admit dsv4 + fp4 + dynamo-sglang + mtp and clones NVIDIA/srt-slurm pinned to commit 04e87fcc505d6d851451781a5499ca19a02ec2b4, overlaying this repo's recipe YAMLs into its recipes/ tree — harness configuration and the established pattern for srt-slurm lanes. The runtime Dynamo install is the standing convention for dynamo-* configs.

Signed: Ankur-singh

@Klaud-Cold

Copy link
Copy Markdown
Collaborator

✅✅✅ Verdict: PASS ✅✅✅

✅ Check 0 (CODEOWNER): PASS — Ankur-singh is a named owner of configs/nvidia-master.yaml; all other changed paths fall to the catch-all, which a recognized CODEOWNER satisfies.
✅ Check 1 (sweep on in-PR commit): PASS — PR head 8c15a0af carries 5 green executed multi-node 8k1k / and 4 green executed multi-node eval / check-runs from run 31451666262 (the single-node */ / eval / lanes have no matrix entries for this multi-node config).
✅ Check 2 (evals pass): PASS — downloaded agg_eval_all.json: GSM8K em_strict 0.964–0.968 across the four eval'd topologies, above the dsv4 bar of 0.91 (utils/evals/thresholds.yaml), on the same image lmsysorg/sglang:nightly-dev-cu13-20260710-cfc66e05 as this PR's config.
➖ Check 3 (recipe link): N/A — disaggregated/multi-node submission (benchmarks/multi_node/srt-slurm-recipes/**, multinode: true, disagg: true, dynamo-sglang); the recipe-link requirement applies to single-node recipes only.
✅ Check 4 (reuse command): PASS — /reuse-sweep-run 31451666262 posted by Ankur-singh (COLLABORATOR).
✅ Check 5 (latest checklist): PASS — sign-off matches the current template; the two unchecked items (agentic golden-AL, single-node recipe publication) are both explained as N/A in the additional detail section.
✅ Check 6 (upstream image / ordering): PASS — image is upstream lmsysorg/sglang:nightly-dev-cu13-20260710-cfc66e05; engine-first ordering satisfied by existing dsv4 B200 vLLM/SGLang entries (dsv4-fp4-b200-sglang, dsv4-fp4-b200-vllm-mtp).
✅ Check 7 (deprecated models): PASS — MODELS.md lists dsv4 Single-turn 8k1k as active (only 1k1k is deprecated); this PR is 8k1k only.
✅ Check 8 (no architecture hacks): PASS — no --hf-overrides / model-config edits; env knobs and swa-full-tokens-ratio are runtime/memory tuning, not FLOPs reduction.
✅ Check 9 (spec-decode chat template): PASS — all five MTP recipes set use_chat_template: true in their benchmark block.
✅ Check 10 (no engine patches): PASS — no patching of the pinned SGLang image; the launcher change is harness plumbing (pinned srt-slurm clone, recipe overlay, standing runtime Dynamo install for dynamo-* lanes).
➖ Check 11 (agentic golden AL): N/A — fixed-seq-len 8k1k, not agentic; correctly runs measured (not synthetic) acceptance — the master entry sets no SYNTHETIC_ACCEPTANCE and no SGLANG_SIMULATE_ACC_* appears in the recipes.

@adibarra
adibarra merged commit 1446647 into main Aug 12, 2026
30 checks passed
@adibarra
adibarra deleted the dsv4-fp4-b200-dynamo-sglang-mtp-nscale branch August 12, 2026 18:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Development

Successfully merging this pull request may close these issues.

4 participants