Skip to content

feat(recipes): cherry-pick Solar Open2 250B NVFP4 B200 recipes (#14376) - #14912

Merged
ynpandey-nv merged 1 commit into
ai-dynamo:release/1.4.1-solar-open2-250b-post.1from
ynpandey-nv:ynpandey-nv/cherrypick-14376-solar-open2-250b
Sep 16, 2026
Merged

ynpandey-nv merged 1 commit into
ai-dynamo:release/1.4.1-solar-open2-250b-post.1from
ynpandey-nv:ynpandey-nv/cherrypick-14376-solar-open2-250b

Conversation

@ynpandey-nv

Copy link
Copy Markdown
Contributor

Summary

Cherry-picks merged #14376 commit (0bd20593b1f248a5a7f877ada9c4ebb6194eca01) onto release/1.4.1-solar-open2-250b-post.1.

  • Adds Solar Open2 250B NVFP4 aggregated (4x B200) and disaggregated (8x B200) vLLM recipes, docs, and AIPerf jobs
  • Uses a signed-off cherry-pick commit

Verification

  • Changed paths are byte-identical to main
  • git diff --check passes

…d recipes for B200 (ai-dynamo#14376)

Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com>
Signed-off-by: Yogendra Pandey <ypandey@nvidia.com>
@ynpandey-nv
ynpandey-nv requested review from a team as code owners September 16, 2026 02:32
@copy-pr-bot

copy-pr-bot Bot commented Sep 16, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

@ynpandey-nv
ynpandey-nv deployed to external_collaborator September 16, 2026 02:32 — with GitHub Actions Active
@ynpandey-nv
ynpandey-nv deployed to external_collaborator September 16, 2026 02:32 — with GitHub Actions Active
@github-actions

Copy link
Copy Markdown
Contributor

👋 Hi ynpandey-nv! Thank you for contributing to ai-dynamo/dynamo.

Just a reminder: The NVIDIA Test Github Validation CI runs an essential subset of the testing framework to quickly catch errors.Your PR reviewers may elect to test the changes comprehensively before approving your changes.

🚀

@github-actions github-actions Bot added feat external-contribution Pull request is from an external contributor labels Sep 16, 2026
@dynamo-ops

Copy link
Copy Markdown
Contributor

Unsigned Commits Detected

The following commits are not GPG-signed and must be signed before CI can run:

Commit Reason
9b4b200 unsigned

Please sign your commits and push again. See the GitHub docs on commit signature verification for help.

@github-actions github-actions Bot added the documentation Improvements or additions to documentation label Sep 16, 2026
@ynpandey-nv
ynpandey-nv merged commit 08d7f10 into ai-dynamo:release/1.4.1-solar-open2-250b-post.1 Sep 16, 2026
19 of 20 checks passed
resources:
requests:
storage: 1200Gi
# Edit this

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove this comment: the placeholder value your-storage-class-name already makes the required edit self-evident, and the recipe documentation gives the same instruction.

🤖 AI Fix

Delete the # Edit this comment.

- --router-track-prefill-tokens
- --router-kv-overlap-score-credit-decay=0.1
- --router-min-initial-workers=2
- --dyn-chat-processor=vllm

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Selecting the vLLM chat processor makes the frontend construct a vLLM model/tokenizer from the discovered worker model's local directory. The workers mount that directory under /shared-model-cache, but this frontend mounts only /tmp; with offline Hugging Face enabled, model discovery fails with MDC local_dir ... not populated and chat completions never become available. The disaggregated frontend has the same omission.

🤖 AI Fix

Mount the shared-model-cache PVC at /shared-model-cache in the Frontend component of both deployment manifests.

- name: runtime-tmp
emptyDir:
sizeLimit: 20Gi
- name: Worker

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The worker pod template requests only the generic NVIDIA GPU resource, so on a mixed-GPU cluster Kubernetes may schedule this B200/NVFP4 profile onto another GPU type. That makes the advertised recipe fail or run with unsupported kernels instead of selecting the required B200 hardware.

🤖 AI Fix

Add a required B200 node selector or node affinity to the Worker pod template.

- name: runtime-tmp
emptyDir:
sizeLimit: 20Gi
- name: PrefillWorker

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Neither prefill nor decode worker constrains scheduling to B200 nodes; their generic GPU and RDMA requests can be satisfied by other GPU node types in a mixed cluster. The B200 NVFP4 configuration can therefore be placed on unsupported hardware and fail to start.

🤖 AI Fix

Add a required B200 node selector or node affinity to both PrefillWorker and DecodeWorker pod templates.

This branch was successfully deployed

1 active deployment
external_collaborator — 9b4b200d Deployed Sep 16, 2026 by ynpandey-nv via ok-to-test #18636
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation external-contribution Pull request is from an external contributor feat size/XL

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants