Skip to content

[NV] Add GB300 AgentX Qwen3.5 recipes - #2121

Merged
Ankur-singh merged 17 commits into
mainfrom
nv-qwen35-agentx-gb300
Jul 30, 2026
Merged

[NV] Add GB300 AgentX Qwen3.5 recipes#2121
Ankur-singh merged 17 commits into
mainfrom
nv-qwen35-agentx-gb300

Conversation

@csahithi

@csahithi csahithi commented Jul 8, 2026

Copy link
Copy Markdown
Collaborator
  • 9 Qwen3.5-397B-A17B-NVFP4 GB300 sglang AgentX recipes (agg + disagg pareto)
  • configs/nvidia-master.yaml: qwen3.5-fp4-gb300-dynamo-sglang{,-agentic-agg,-agentic-disagg}
  • perf-changelog.yaml: agentic pareto entry
  • runners/launch_gb300-nv.sh: qwen3.5 launcher branch + dynamo-wheels cache

@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

2 similar comments
@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@csahithi
csahithi force-pushed the nv-qwen35-agentx-gb300 branch 2 times, most recently from b90cee7 to 7dc6bac Compare July 8, 2026 20:55
@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown
Contributor

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Additional findings (outside current diff — PR may have been updated during review):

  • 🔴 configs/nvidia-master.yaml:10940-10969 — 9 new sweep entries in configs/nvidia-master.yaml (agentic-agg + agentic-disagg, lines 10940–11085) declare spec-decoding: none, but every one of the 9 corresponding recipe files enables NEXTN MTP (speculative-algorithm: NEXTN + speculative-num-steps: 3 + eagle-topk/num-draft-tokens; the recipe filenames also encode -mtp-). The engine runs MTP correctly because srtctl applies the recipe directly, but SPEC_DECODING propagates verbatim into RESULT_FILENAME (_spec-none_… in artifact names, benchmark-multinode-tmpl.yml:225) and into the aggregated result JSON's spec_decoding field (utils/agentic/aggregation/process_agentic_result.py:139), so downstream dashboards / Pareto comparisons will silently misclassify these MTP runs as non-MTP. Fix: change spec-decoding: nonespec-decoding: "mtp" on all 9 entries in this block (matches the file's existing MTP convention — every other qwen3.5/dsr1/dsv4 MTP entry uses mtp; the parallel dsv4 agentic entry using none is self-consistent because its vLLM recipes have no speculative-* fields).

    Extended reasoning...

    What the bug is

    All 9 newly added sweep entries under qwen3.5-fp4-gb300-dynamo-sglang-agentic-agg and qwen3.5-fp4-gb300-dynamo-sglang-agentic-disagg (configs/nvidia-master.yaml lines 10943, 10957, 10986, 11000, 11014, 11029, 11044, 11059, 11074) set:

    spec-decoding: none

    But every one of the 9 corresponding recipe files under benchmarks/multi_node/srt-slurm-recipes/sglang/qwen3.5/gb300-fp4/agentic/*.yaml enables NEXTN MTP:

    speculative-algorithm: NEXTN
    speculative-num-steps: 3
    speculative-eagle-topk: 1
    speculative-num-draft-tokens: 4

    The recipe filenames themselves also encode -mtp- (e.g. agg-gb300-tp4-c1-mtp-hicache-jid2191933.yaml, disagg-gb300-1p1d-tp4-tp4-c32-mtp-hicache-convaware-jid2212783.yaml).

    Why the engine still runs MTP but the metadata is wrong

    srtctl apply reads the recipe file directly, so the engine gets --speculative-algorithm NEXTN … on its command line and NEXTN MTP runs as intended. The spec-decoding field in nvidia-master.yaml is not a live engine setting — it feeds SPEC_DECODING into the workflow env, which has two downstream consumers, and both are silently corrupted by this mislabel:

    1. Artifact filenames.github/workflows/benchmark-multinode-tmpl.yml:225 bakes _spec-${{ env.SPEC_DECODING }}_ directly into RESULT_FILENAME. So uploaded result JSONs land as …_spec-none_conc*.json despite the run actually using MTP.

    2. Aggregated result JSON bodyutils/agentic/aggregation/process_agentic_result.py:139 writes "spec_decoding": os.environ.get("SPEC_DECODING", "none") verbatim into the aggregated result JSON. Downstream dashboards / Pareto plots that filter or facet on spec_decoding will silently misclassify these MTP runs as non-MTP.

    (Some verifiers also flagged utils/compare_results.py:51 reading result['spec_decoding'] for baseline lookup — worth double-checking, but the two consumers above are sufficient on their own.)

    Convention check

    Every other MTP recipe in nvidia-master.yaml (dsr1-mtp, dsv4-b200-vllm-mtp, qwen3.5 fixed-seq-len MTP tiers, etc.) uses spec-decoding: "mtp". The one other agentic entry that legitimately uses spec-decoding: nonedsv4-fp4-gb300-dynamo-vllm-agentic — is self-consistent because its vLLM recipes at benchmarks/multi_node/srt-slurm-recipes/vllm/deepseek-v4/agentic/*.yaml contain no speculative-* fields. So this is a real inconsistency introduced by this PR, not the file's existing style.

    Step-by-step proof (agentic-agg conc=1)

    1. Sweep matrix expands the spec-decoding: none / conc-list: [1] entry into a GHA job with SPEC_DECODING=none, CONFIG_FILE=recipes/sglang/qwen3.5/gb300-fp4/agentic/agg-gb300-tp4-c1-mtp-hicache-jid2191933.yaml.
    2. srtctl apply -f $CONFIG_FILE reads the recipe, which contains speculative-algorithm: NEXTN + speculative-num-steps: 3 — the sglang server starts with NEXTN MTP enabled. ✅ Engine correct.
    3. benchmark-multinode-tmpl.yml:225 computes RESULT_FILENAME=…_spec-none_conc1_… from env.SPEC_DECODING=none. ❌ Artifact filename claims non-MTP.
    4. process_agentic_result.py:139 reads os.environ.get("SPEC_DECODING", "none")"none" and writes "spec_decoding": "none" into the aggregated result JSON. ❌ Result body claims non-MTP.
    5. A downstream MTP-vs-non-MTP Pareto plot that groups by spec_decoding puts this run in the non-MTP bucket, distorting the frontier.

    Fix

    Mechanical 9-line change — replace spec-decoding: none with spec-decoding: "mtp" on all 9 new entries in this block (lines 10943, 10957, 10986, 11000, 11014, 11029, 11044, 11059, 11074). Nothing else in the recipe files needs to change; the change is purely to keep the master-config label truthful to what the recipe actually runs.

Comment thread perf-changelog.yaml Outdated
Comment on lines +4366 to +4374
- config-keys:
- qwen3.5-fp4-gb300-dynamo-sglang-agentic-agg
- qwen3.5-fp4-gb300-dynamo-sglang-agentic-disagg
description:
- "Add Qwen3.5-397B-A17B-NVFP4 FP4 GB300 SGLang AgentX Pareto-frontier benchmarks via the srtctl/dynamo stack"
- "agg: single aggregated node (TP4, MTP/NEXTN, hierarchical KV cache), concurrencies [1, 96]"
- "disagg: 1P1D-4P1D (TP4/DEP4 prefill, TP4/DEP4/DEP8 decode, MTP/NEXTN, hicache, conv-aware routing) over 7 Pareto points at concurrency 32-384"
- "Image: lmsysorg/sglang:nightly-dev-cu13-20260624-b2c8f7a2; runner: gb300-nv"
pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2121

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 The new perf-changelog entry for qwen3.5-fp4-gb300-dynamo-sglang-agentic-{agg,disagg} (perf-changelog.yaml:4366-4373) will block the changelog gate on two independent grounds: (1) it's missing the required pr-link field, and (2) it's inserted mid-file rather than appended to the end. Fix: move the 8-line block to the end of the file and add pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2121 (or the XXX placeholder while WIP).

Extended reasoning...

What breaks. utils/validate_perf_changelog.py (invoked from .github/workflows/run-sweep.yml on every PR push) enforces three independent checks that this new entry violates. Any one of them is enough to fail the changelog gate; the entry fails all three.

1. Missing pr-link — schema-level rejection. In utils/matrix_logic/validation.py:683, ChangelogEntry declares pr_link: str = Field(alias="pr-link") with no default and extra="forbid". parse_changelog (validate_perf_changelog.py:128-133) calls ChangelogEntry.model_validate(entry) on every entry; Pydantic raises ValidationError on the new entry, which is re-raised as ChangelogValidationError('… fails ChangelogEntry validation'). Even if that were bypassed, validate_added_pr_link (lines 144-160) rejects it a second time: str(entry.get('pr-link') or '') == '', which is neither in PR_LINK_PLACEHOLDERS = {'XXX', 'https://github.com/SemiAnalysisAI/InferenceX/pull/XXX'} nor equal to the canonical PR URL, so it raises ChangelogValidationError('new PR entry must use …').

2. Mid-file insertion — positional rejection. The new 8-line block sits at lines 4366-4373, between an existing pull/1931 entry and the pre-existing minimaxm3-fp4-mi355x-vllm-mtp entry — with ~30 more historical entries after it. compare_entries (lines 163-208) walks base_entries and head_entries in lockstep by index; at the insertion index, head_entries[i] is the new qwen3.5 entry but base_entries[i] is the pre-existing minimaxm3 entry, and without_pr_link differs, so it raises ChangelogValidationError('entry N changed; existing entries are immutable except for pr-link-only corrections').

3. Byte-level prefix check. validate_raw_change (lines 211-235) additionally requires head_raw.startswith(base_raw) when there are additions. A mid-file insertion breaks that even before compare_entries runs, raising 'appended entries changed historical perf-changelog.yaml bytes; restore the base file byte-for-byte and append at the end'. AGENTS.md line 78 codifies the same rule: 'new entries MUST be appended to the END, never inserted in the middle or prepended.'

Step-by-step proof.

  1. CI runs python3 utils/validate_perf_changelog.py --base-ref origin/main --head-ref HEAD.
  2. parse_changelog reads perf-changelog.yaml and iterates entries. When it reaches the qwen3.5 entry, ChangelogEntry.model_validate({config-keys: [...], description: [...]}) fails because pr-link is required and absent → ChangelogValidationError → exit non-zero → the Validate perf-changelog check turns red → PR cannot merge. If pr-link were the only issue, adding pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2121 at the end of the block would clear check 1.
  3. Even then, validate_raw_change computes base_raw = git show origin/main:perf-changelog.yaml and head_raw = the current file; because ~30 entries follow the insertion point in head_raw but not at the same byte offset in base_raw, head_raw.startswith(base_raw) is false → raises. Only after moving the 8-line block to the file end (after the current last entry) does this pass.
  4. compare_entries then compares base and head entries index-by-index; index-of-first-difference is now len(base_entries), the appended entry is treated as an addition, validate_added_pr_link re-validates the pr-link, and the gate passes.

Fix. Two edits:

  • Move the qwen3.5-fp4-gb300-dynamo-sglang-agentic-{agg,disagg} entry (perf-changelog.yaml:4366-4373) to the end of perf-changelog.yaml, after the current last entry.
  • Add pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2121 to the moved block (or pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/XXX for early WIP).

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Resolved both the above comments

@csahithi
csahithi force-pushed the nv-qwen35-agentx-gb300 branch from 80d1f5c to ae49f13 Compare July 8, 2026 21:33
@github-actions

github-actions Bot commented Jul 8, 2026

Copy link
Copy Markdown
Contributor

speculative-algorithm: NEXTN
speculative-num-steps: 3
speculative-eagle-topk: 1
speculative-num-draft-tokens: 4

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

started about discussion about AgentX MTP +viz @kedarpotdar-nv

@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown
Contributor

@cquil11 cquil11 added the agentx AgentX benchmarks, recipes, and infrastructure label Jul 9, 2026 — with ChatGPT Codex Connector
@csahithi
csahithi force-pushed the nv-qwen35-agentx-gb300 branch from bd940fb to e571715 Compare July 9, 2026 22:53
@github-actions

github-actions Bot commented Jul 9, 2026

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

@csahithi csahithi changed the title [WIP] Add GB300 AgentX Qwen3.5 recipes [NV] Add GB300 AgentX Qwen3.5 recipes Jul 10, 2026
@Ankur-singh

Copy link
Copy Markdown
Collaborator

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this.
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; FLOPs at lower precisions is fine, given that the config passes private evals. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that the single-node recipes are similar to the official vLLM recipes and/or theSGLang cookbook:
    • If they are not, I have verified that a PR has been opened in vLLM recipe repo or SGLang repo and linked it below in the additional detail section:
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.

Additional detail section:

  • PR validation: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/29056354557 / PR 验证:https://github.com/SemiAnalysisAI/InferenceX/actions/runs/29056354557
  • Evals: not applicable to this AgentX-only submission; the eval jobs are intentionally skipped. / 评估:本提交仅包含 AgentX 基准测试,因此不适用;评估任务按设计跳过。
  • This AgentX multi-node submission includes disaggregated configurations, so no single-node recipe update is required. / 本 AgentX 多节点提交包含分离式配置,因此无需更新单节点 recipe。
  • AgentX uses /v1/chat/completions, satisfying the MTP chat-template requirement. / AgentX 使用 /v1/chat/completions,满足 MTP 的 chat template 要求。
  • The golden acceptance length is applied using real-draft-token, preserving real draft-token timing under the AgentX fairness convention. / 按照 AgentX 公平性约定,固定 acceptance length 使用 real-draft-token 模式,保留真实 draft token 的时序。

Signed: ankur-singh

@Klaud-Cold

Copy link
Copy Markdown
Collaborator

✅✅✅ Verdict: PASS ✅✅✅

✅ Check 0 (CODEOWNER): PASS — @Ankur-singh is a listed owner of configs/nvidia-master.yaml (the only specifically-owned changed path); remaining paths fall to the catch-all, which a recognized CODEOWNER satisfies.
✅ Check 1 (sweep on in-PR commit): PASS — full sweep 29056354557 ran on the PR head f644fd7 itself; all 9 executed multi-node agentic / benchmark jobs concluded success. single-node */ / eval / lanes were skipped because they don't apply to this agentic-only config set (run-sweep.yml hardcodes run-eval: false for agentic lanes), not because of a reuse-skip.
➖ Check 2 (evals pass): N/A — evals are structurally never run for agentic scenarios (verified in run-sweep.yml; find_reusable_sweep_run.py accepts bmk_agentic_* sources without evals). Sign-off leaves the eval box unchecked with this reasoning, per the checklist's exception item.
➖ Check 3 (recipe link): N/A — exclusively multi-node/disagg-stack submission (benchmarks/multi_node/srt-slurm-recipes/**, entries multinode: true, framework dynamo-sglang); the recipe-link requirement covers single-node recipes only.
✅ Check 4 (reuse command): PASS — /reuse-sweep-run posted by Ankur-singh (COLLABORATOR).
✅ Check 5 (latest template): PASS — all current PR_REVIEW_CHECKLIST.md items present; the two unchecked items (evals, recipe link) are each justified in the additional detail section.
✅ Check 6 (upstream image / engine-first): PASS — both new entries pin upstream lmsysorg/sglang:nightly-dev-cu13-20260709-074bb928; the engine is upstream SGLang (dynamo is only the disagg frontend; the agg entry serves via a pure sglang frontend), and a same-model/SKU qwen3.5-fp4-gb300-dynamo-sglang entry already exists. Informational: the perf-changelog entry cites image ...20260624-b2c8f7a2 while the config pins ...20260709-074bb928 — worth aligning, not blocking.
✅ Check 7 (no architecture hacks): PASS — no --hf-overrides/model-config edits; all flags are parallelism/quant/kernel selection. Informational: SGLANG_SIMULATE_ACC_LEN=3.39 (match-expected, real-draft-token) pins spec-decode acceptance length per the declared AgentX fairness convention; draft and verify computation still execute, so no architecture FLOPs are removed.
✅ Check 8 (spec-decode chat template): PASS — all MTP/NEXTN configs benchmark through build_replay_cmd, which sets --endpoint /v1/chat/completions --endpoint-type chat (benchmarks/benchmark_lib.sh:1448-1449).

@Ankur-singh

Copy link
Copy Markdown
Collaborator

@functionstackx @adibarra

@xinli-sw

Copy link
Copy Markdown
Collaborator

@cquil11 @functionstackx please review and sign off, thanks!

csahithi added 2 commits July 20, 2026 12:23
# Conflicts:
#	configs/nvidia-master.yaml
#	perf-changelog.yaml
#	runners/launch_gb300-nv.sh
…-per-node)

The agentic DRAM-offload matrix logic requires available-cpu-dram-mib for any
runner used by a kv-offloading: dram (hicache) config. gb300-nv had no hardware
entry, so generate_sweep_configs test-config crashed for the gb300 agentic keys
(qwen3.5 here, and dsv4 already on main). GB300 NVL72 uses the same 2x Grace /
4-GPU-per-node topology and LPDDR5X capacity as GB200, so mirror gb200-nv's
available host DRAM (860_160 MiB) and gpus-per-node (4).
@github-actions

Copy link
Copy Markdown
Contributor

csahithi added 3 commits July 24, 2026 10:30
Agentic sweep results are ingested by the dedicated trigger-agentic-ingest
job (repository dispatch to InferenceX-app, main-push only), not through
collect-results, whose gate intentionally excludes the agentic sweeps. This
reverts the qwen branch's addition so run-sweep.yml matches main.
@github-actions

Copy link
Copy Markdown
Contributor

@github-actions

Copy link
Copy Markdown
Contributor

csahithi added 2 commits July 26, 2026 18:38
…erArgs fix)

SGLang 20260724 (>= #31811) makes ServerArgs read-only after resolution, which
crashed the Dynamo worker at startup (bare server_args.incremental_streaming_output
assignment). Pin Dynamo to the ai-dynamo/dynamo#12081 head commit (8e4bfa69), a
source build that sets incremental_streaming_output on parsed_args *before*
resolution, so no post-resolution mutation is attempted. Swap the 7 disagg recipes'
dynamo wheel 1.3.0.dev1 -> hash 8e4bfa69 and bump the disagg router version to match.
@github-actions

Copy link
Copy Markdown
Contributor

csahithi added 3 commits July 27, 2026 10:52
…n routing

Switch the 7 disagg recipes from the legacy nvext conv-aware routing (which the
newer Dynamo no longer honors) to header-based session affinity:
- benchmark.env: AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING=1 ->
  AIPERF_HTTP_X_DYNAMO_SESSION_ID_FROM_CORRELATION_ID=true
- frontend: add nginx_session_affinity + nginx_session_affinity_header X-Dynamo-Session-ID
- frontend args: router-reset-states -> router-session-affinity-ttl-secs 3600
Skip the legacy nvext conv-aware CLI routing when a recipe opts into the
X-Dynamo-Session-ID header path (AIPERF_HTTP_X_DYNAMO_SESSION_ID_FROM_CORRELATION_ID=true);
the existing AIPERF_USE_DYNAMO_CONV_AWARE_ROUTING opt-out and default are unchanged.
@github-actions

Copy link
Copy Markdown
Contributor

1 similar comment
@github-actions

Copy link
Copy Markdown
Contributor

csahithi added 2 commits July 28, 2026 11:49
…6 -> 512

The c128 (1P1D DEP4/DEP4) decode server crashed on a DeepEP low-latency dispatch
assertion (x.size(0) <= num_max_dispatch_tokens_per_rank, deep_ep.cpp:1262): under
DP-attention load imbalance one DP rank's token count (amplified by MTP draft
tokens) exceeded the 256 buffer. Double the DeepEP dispatch buffer to 512 (still
well under sglang's 1024 cap) to give headroom for the imbalanced+MTP worst case.
@github-actions

Copy link
Copy Markdown
Contributor

@Ankur-singh

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run 30391410523

@Ankur-singh

Copy link
Copy Markdown
Collaborator

As a PR reviewer and CODEOWNER, I have reviewed this and have:

  • Verified that as of the moment of typing this, this is the latest version of PR_REVIEW_CHECKLIST.md
  • Verified that the general code quality meets the InferenceX standard and does not make the code quality any worse.
  • Verified that this PR has passed PR validation. Please link to GitHub Action workflow that shows this. — https://github.com/SemiAnalysisAI/InferenceX/actions/runs/30391410523
  • Verified that this PR passes evals. Please link to GitHub Action workflow that shows this.
  • Verified that speculative decoding PRs uses chat templates to align the AL distribution to real world
  • For agentic workloads: verified that speculative-decoding configs (EAGLE / MTP / draft models) run with simulated synthetic acceptance, with the acceptance-length value taken from the committed golden AL curve in golden_al_distribution/ for that model, thinking mode, and draft length. A submission may choose any supported draft length, but it may not substitute a different acceptance target.
  • Verified that the model architecture isn't changed with benchmark hacks like using --hf-overrides to skipping indexer for every x layers on models that don't natively support this. As a general rule, we won't accept optimizations that reduces the number of model architecture FLOPs. Anything that makes that same computation run faster is fair game; FLOPs at lower precisions is fine, given that the config passes private evals. As an general north star princple, we should only use optimizations which is used in production by customers that care about accuracy
  • If an company claims that they support vLLM/SGLang as first class LLM inference engines on their hardware, I have verified that the respective vLLM submission made using upstream https://hub.docker.com/u/vllm docker repo, upstream SGLang https://hub.docker.com/u/lmsysorg docker repo. The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet as supported by vLLM/SGLang community maintainers
  • If an company claims that they support vLLM/SGLang as first class upstream in-tree LLM inference engines on their hardware, I have have verified that the respective vLLM/SGLang submission has been made before additional frameworks (TRT-LLM, ATOM, etc.). The only exceptions are for new hardware, such as MI455X UALoE72, Vera Rubin NVL72, Rubin NVL8, etc., and for new model architectures where there is an actual reason why vLLM/SGLang does not fundamentally support them yet.
  • Verified that every single-node vLLM/SGLang recipe in this PR is documented in the official vLLM recipes and/or the SGLang cookbook:
    • I linked the corresponding upstream PR in the vLLM recipe repo or SGLang repo and verified that it is MERGED before this InferenceX PR merges. An opened, draft, or closed-without-merge upstream PR does not satisfy this requirement. If the matching recipe was already published, I linked the published recipe/cookbook page in the additional detail section below.
  • Verified that this PR does not patch the inference engine or serving stack — the pinned image must run as shipped. This covers .patch files / git apply / patch, inline patches embedded in benchmark scripts (e.g. a python3/sed heredoc that rewrites installed engine sources before serving), in-place edits of site-packages, monkey-patching, overwriting container files, and installing forked/rebuilt engine wheels on top of the pinned image. The only exception is a patch covered by a filled-out waiver at docs/waiver/<PR_NUMBER>.md — named after the PR that introduces the patch and filed in that same PR, stating what is patched, why the unmodified upstream image cannot run this benchmark, the upstream PR/issue link, and the removal plan — which I have linked below in the additional detail section.
  • If any of the above criteria cannot reasonably be satisfied, I have provided additional reasoning below.

Additional detail section:

  • Scope: adds the first AgentX (agentic-coding) submission for Qwen3.5-397B-A17B-NVFP4 on GB300 — two new master-config keys, qwen3.5-fp4-gb300-dynamo-sglang-agentic-agg (TP4 aggregated, concurrencies 1 and 96) and qwen3.5-fp4-gb300-dynamo-sglang-agentic-disagg (1P1D-4P1D, TP4/DEP4 prefill and TP4/DEP4/DEP8 decode, 7 Pareto points at concurrency 32-384) — backed by 9 new benchmarks/multi_node/srt-slurm-recipes/sglang/qwen3.5/gb300-fp4/agentic/ recipes. Supporting changes: a cluster:gb300-nv hardware entry in configs/runners.yaml, a persistent dynamo source-build wheel cache plus a qwen3.5 branch in runners/launch_gb300-nv.sh, and one backwards-compatible guard in benchmarks/benchmark_lib.sh.
  • Validation: Run Sweep 30391410523 ran on the current PR head 220701e8bc3b91b95fa8e4e4597b4c0ccb8efcfa itself (not a reused run: setup and canary-select both executed) and is green, with all 9 executed multi-node agentic / benchmark jobs successful — one per Pareto point across both new config keys. Every job ran on this PR's pinned image lmsysorg/sglang:nightly-dev-cu13-20260724-433429b1.
  • Evals (unchecked): not applicable — evals are structurally never scheduled for agentic-coding scenarios, so agentic eval /, eval /, and collect-evals are skipped in the sweep. I am deliberately leaving this box unchecked rather than citing a top-level green run, because no eval job actually executed for these configs.
  • Agentic synthetic acceptance (golden AL): all 9 recipes set SGLANG_SIMULATE_ACC_LEN: '3.39' with SGLANG_SIMULATE_ACC_METHOD: match-expected and SGLANG_SIMULATE_ACC_TOKEN_MODE: real-draft-token in the server environment (the decode_environment block for the disaggregated recipes). 3.39 is exactly the committed golden value in golden_al_distribution/qwen3.5_mtp.yaml for qwen3.5-397b-a17b-nvfp4, thinking_on, MTP level 3 — matching the recipes' declared draft length (speculative-algorithm: NEXTN, speculative-num-steps: 3, speculative-eagle-topk: 1, speculative-num-draft-tokens: 4, i.e. 3 draft tokens plus the bonus verification token). No acceptance target was substituted. This is the same mode-consistent convention as the already-merged DeepSeek-V4 GB300 agentic recipes, which pin 2.49 = dsv4_mtp.yaml thinking_on level 3.
  • Chat template (spec decode): the agentic replay path issues chat-completions requests — build_replay_cmd sets --endpoint /v1/chat/completions and --endpoint-type chat (benchmarks/benchmark_lib.sh:1770-1771 at this head), so the AL distribution is aligned to real-world chat usage.
  • No architecture hacks: no --hf-overrides and no model-config rewriting. All server flags are parallelism, quantization, attention/MoE kernel-backend, and KV/hierarchical-cache selection. Synthetic acceptance is the prescribed AgentX fairness mechanism, not a FLOPs reduction: the draft and verification steps still execute.
  • Upstream image and framework ordering: both new entries pin the upstream lmsysorg/sglang:nightly-dev-cu13-20260724-433429b1 image from the official SGLang Docker repo, and the engine is upstream SGLang (dynamo supplies only the disaggregated frontend/router). A same-model, same-SKU SGLang entry (qwen3.5-fp4-gb300-dynamo-sglang) already exists; no additional framework such as TRT-LLM or ATOM is introduced ahead of it.
  • Recipe item (unchecked): not applicable — this PR contains no single-node recipes. Every recipe added is multi-node/disaggregated under benchmarks/multi_node/srt-slurm-recipes/, and both config entries are multinode: true. The upstream recipe/cookbook requirement covers single-node recipes only.
  • No engine patching: no .patch files, git apply, sed -i, site-packages edits, monkey-patching, or rebuilt engine wheels — the pinned SGLang image runs as shipped. Disclosed for completeness: the recipes set dynamo.install: true with a pinned dynamo.hash, which builds the dynamo frontend/router inside the container at runtime. That is the disaggregated serving frontend rather than the inference engine, and it is the established pattern already used by the merged DeepSeek-V4 GB300 agentic recipes; launch_gb300-nv.sh adds only a persistent build cache for it. No waiver is required.
  • Launcher change: the qwen3.5 agentic branch clones upstream NVIDIA/srt-slurm at tag v1.0.22 rather than a personal fork, because the two features previously requiring the fork (srtctl apply --no-preflight and srun_options propagation through benchmark_stage) are now upstream. --no-preflight is needed because the model resolves to the compute-node-local NVMe path /scratch/models/..., which the GHA runner pod cannot stat.

Signed: Ankur-singh

@Klaud-Cold

Copy link
Copy Markdown
Collaborator

✅✅✅ Verdict: PASS ✅✅✅

✅ Check 0 (CODEOWNER): PASS — @Ankur-singh is a listed owner of configs/nvidia-master.yaml (the only specifically-owned changed path); all other paths fall to the catch-all, which a recognized CODEOWNER satisfies.
✅ Check 1 (sweep on in-PR commit): PASS — run 30391410523 executed on the current PR head 220701e8 itself (setup/canary-select ran; not a reuse-skip); all 9 multi-node agentic / benchmark jobs concluded success, and it produced all 9 bmk_agentic_* result artifacts the reuse path requires. single-node *//eval / lanes were skipped because no such configs exist in this PR, not due to reuse.
➖ Check 2 (evals pass): N/A — evals are structurally never scheduled for agentic lanes (run-sweep.yml hardcodes run-eval: false for sweep-agentic and sweep-multi-node-agentic; find_reusable_sweep_run.py accepts bmk_agentic_* runs without eval artifacts). Sign-off leaves the eval box unchecked with exactly this reasoning. All 9 jobs ran the PR's pinned image lmsysorg/sglang:nightly-dev-cu13-20260724-433429b1.
➖ Check 3 (recipe link): N/A — exclusively multi-node/disagg-stack submission (all recipes under benchmarks/multi_node/srt-slurm-recipes/**; both master-config entries multinode: true, framework dynamo-sglang); the recipe-link requirement covers single-node recipes only.
✅ Check 4 (reuse command): PASS — /reuse-sweep-run 30391410523 posted by Ankur-singh (COLLABORATOR), pinning the green run on this head.
✅ Check 5 (latest template): PASS — every current PR_REVIEW_CHECKLIST.md item is present; the two unchecked items (evals, recipe link) are each justified in the additional detail section per the checklist's exception item.
✅ Check 6 (upstream image / engine-first): PASS — both new entries pin upstream lmsysorg/sglang:nightly-dev-cu13-20260724-433429b1 on GB300; the engine is upstream SGLang (dynamo supplies only the disagg frontend/router; the agg recipes serve via a pure sglang frontend), a same-model/SKU qwen3.5-fp4-gb300-dynamo-sglang entry already exists, and no non-SGLang framework is introduced. The prior changelog/config image mismatch is resolved at this head.
✅ Check 7 (no architecture hacks): PASS — no --hf-overrides/model-config rewrites; all server flags are parallelism, quantization, kernel-backend, and KV/hicache selection. Synthetic acceptance (SGLANG_SIMULATE_ACC_LEN) is the sanctioned AgentX mechanism — draft and verify FLOPs still execute.
✅ Check 8 (spec-decode chat template): PASS — the agentic replay path issues chat-completions requests: build_replay_cmd sets --endpoint /v1/chat/completions --endpoint-type chat (benchmarks/benchmark_lib.sh).
✅ Check 9 (no engine patches): PASS — no .patch/git apply/sed -i/site-packages edits or rebuilt engine wheels; dynamo.install (pinned hash) builds the disagg frontend, not the SGLang engine, matching the merged DeepSeek-V4 GB300 agentic pattern, and the launcher now clones upstream NVIDIA/srt-slurm v1.0.22.
✅ Check 10 (golden AL): PASS — all 9 recipes pin SGLANG_SIMULATE_ACC_LEN: '3.39' (match-expected, real-draft-token) in the server/decode environment; 3.39 equals golden_al_distribution/qwen3.5_mtp.yaml qwen3.5-397b-a17b-nvfp4 / thinking_on / level 3, matching the declared draft length (speculative-num-steps: 3). No synthetic-acceptance knobs on non-agentic configs.

@Ankur-singh

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run

@Klaud-Cold

Copy link
Copy Markdown
Collaborator

✅✅✅ Verdict: PASS ✅✅✅

✅ Check 0 (CODEOWNER): PASS — @Ankur-singh is a listed owner of configs/nvidia-master.yaml; all other changed paths fall to the catch-all, which a recognized CODEOWNER satisfies.
✅ Check 1 (Passing sweep on in-PR commit): PASS — run 30391410523 executed on 220701e8 (still in this PR; head 64729c6 is only a main-merge by the reuse tooling): all 9 multi-node agentic / jobs green, setup/canary-select executed (real run, not a skip/reuse).
➖ Check 2 (Evals pass): N/A — multi-node agentic evals are structurally unschedulable (utils/matrix_logic/generate_sweep_configs.py skips agentic entries with a prefill topology: "Multi-node agentic eval is unsupported"); no eval job exists for these configs, the sign-off honestly leaves that box unchecked with this reason, and every executed job ran the PR's pinned image lmsysorg/sglang:nightly-dev-cu13-20260724-433429b1.
➖ Check 3 (Upstream recipe link): N/A — disaggregated/multi-node submission (all recipes under benchmarks/multi_node/srt-slurm-recipes/**, both master entries multinode: true); the recipe-link requirement applies to single-node recipes only.
✅ Check 4 (Reuse command): PASS — authorized /reuse-sweep-run 30391410523 posted by Ankur-singh (COLLABORATOR).
✅ Check 5 (Latest checklist): PASS — all current-template items present; the two unchecked items (evals, single-node recipe link) are each explained in the additional detail section, as the template's final item allows.
✅ Check 6 (Upstream image / engine-first): PASS — both new entries pin upstream lmsysorg/sglang:nightly-dev-cu13-20260724-433429b1; engine is upstream SGLang (dynamo is only the disagg frontend/router) and a same-model GB300 SGLang-engine entry (qwen3.5-fp4-gb300-dynamo-sglang) predates this PR.
✅ Check 7 (No architecture hacks): PASS — no --hf-overrides/model-config rewriting anywhere in the diff; all server flags are parallelism, quantization (modelopt_fp4 on an NVFP4 model), kernel-backend, and cache selection.
✅ Check 8 (Spec-decode chat template): PASS — the agentic replay path issues chat-completions requests (benchmarks/benchmark_lib.sh:1770-1771: --endpoint /v1/chat/completions, --endpoint-type chat).
✅ Check 9 (No engine patches): PASS — no .patch/git apply/sed -i/site-packages edits/rebuilt engine wheels; dynamo.install: true (pinned hash) builds only the dynamo frontend/router, the same established pattern as the merged DeepSeek-V4 GB300 agentic recipes — the pinned SGLang engine image runs as shipped.
✅ Check 10 (Golden AL simulated acceptance): PASS — all 9 recipes pin SGLANG_SIMULATE_ACC_LEN: '3.39' with match-expected/real-draft-token in the serving env (aggregated_environment/decode_environment), and 3.39 equals the committed golden value in golden_al_distribution/qwen3.5_mtp.yaml for qwen3.5-397b-a17b-nvfp4 thinking_on at MTP level 3, matching the recipes' speculative-num-steps: 3; no synthetic-acceptance knobs appear on any non-agentic config.

@Ankur-singh
Ankur-singh merged commit 648d857 into main Jul 30, 2026
28 checks passed
@Ankur-singh
Ankur-singh deleted the nv-qwen35-agentx-gb300 branch July 30, 2026 00:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

agentx AgentX benchmarks, recipes, and infrastructure full-sweep-enabled NVIDIA

Projects

Development

Successfully merging this pull request may close these issues.

7 participants