Repository navigation
feat(recipes): cherry-pick Solar Open2 250B NVFP4 B200 recipes (#14376) - #14912
Conversation
…d recipes for B200 (ai-dynamo#14376) Signed-off-by: Sandhya Rani Narravula <snarravula@nvidia.com> Signed-off-by: Yogendra Pandey <ypandey@nvidia.com>
|
👋 Hi ynpandey-nv! Thank you for contributing to ai-dynamo/dynamo. Just a reminder: The 🚀 |
Unsigned Commits DetectedThe following commits are not GPG-signed and must be signed before CI can run:
Please sign your commits and push again. See the GitHub docs on commit signature verification for help. |
08d7f10
into
ai-dynamo:release/1.4.1-solar-open2-250b-post.1
| resources: | ||
| requests: | ||
| storage: 1200Gi | ||
| # Edit this |
There was a problem hiding this comment.
Remove this comment: the placeholder value your-storage-class-name already makes the required edit self-evident, and the recipe documentation gives the same instruction.
🤖 AI Fix
Delete the # Edit this comment.
| - --router-track-prefill-tokens | ||
| - --router-kv-overlap-score-credit-decay=0.1 | ||
| - --router-min-initial-workers=2 | ||
| - --dyn-chat-processor=vllm |
There was a problem hiding this comment.
Selecting the vLLM chat processor makes the frontend construct a vLLM model/tokenizer from the discovered worker model's local directory. The workers mount that directory under /shared-model-cache, but this frontend mounts only /tmp; with offline Hugging Face enabled, model discovery fails with MDC local_dir ... not populated and chat completions never become available. The disaggregated frontend has the same omission.
🤖 AI Fix
Mount the shared-model-cache PVC at /shared-model-cache in the Frontend component of both deployment manifests.
| - name: runtime-tmp | ||
| emptyDir: | ||
| sizeLimit: 20Gi | ||
| - name: Worker |
There was a problem hiding this comment.
The worker pod template requests only the generic NVIDIA GPU resource, so on a mixed-GPU cluster Kubernetes may schedule this B200/NVFP4 profile onto another GPU type. That makes the advertised recipe fail or run with unsupported kernels instead of selecting the required B200 hardware.
🤖 AI Fix
Add a required B200 node selector or node affinity to the Worker pod template.
| - name: runtime-tmp | ||
| emptyDir: | ||
| sizeLimit: 20Gi | ||
| - name: PrefillWorker |
There was a problem hiding this comment.
Neither prefill nor decode worker constrains scheduling to B200 nodes; their generic GPU and RDMA requests can be satisfied by other GPU node types in a mixed cluster. The B200 NVFP4 configuration can therefore be placed on unsupported hardware and fail to start.
🤖 AI Fix
Add a required B200 node selector or node affinity to both PrefillWorker and DecodeWorker pod templates.
Summary
Cherry-picks merged #14376 commit (
0bd20593b1f248a5a7f877ada9c4ebb6194eca01) ontorelease/1.4.1-solar-open2-250b-post.1.Verification
maingit diff --checkpasses