fix(install): restore DGX Station GB300 express setup - #6875
Conversation
Signed-off-by: San Dang <sdang@nvidia.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
📝 WalkthroughWalkthroughManaged vLLM now checks Docker image storage before pulling its pinned image, removes Hugging Face model-cache sizing and gating, documents exact override behavior, and broadens DGX Station platform detection. ChangesManaged vLLM storage preflight
DGX Station platform detection
Estimated code review effort: 4 (Complex) | ~45 minutes Possibly related issues
Possibly related PRs
Suggested labels: 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Code Coverage OverviewLanguages: TypeScript TypeScript / code-coverage/pluginThe overall coverage remains at 96%, unchanged from the TypeScript / code-coverage/cliThe overall coverage in the Show a code coverage summary of the most impacted files.
Updated |
|
@coderabbitai review |
✅ Action performedReview finished.
|
PR Review Advisor — InformationalAdvisor assessment: Informational / high confidence Model lanes
Nemotron output stays in workflow artifacts and does not change the assessment above. E2E guidanceAdvisory only. E2E / PR Gate selects and runs jobs independently. Recommended E2E: 1 optional E2E recommendation
This automated review informs maintainers. Warnings and suggestions do not require a response. A maintainer decides whether to merge. |
|
🌿 Preview your docs: https://nvidia-preview-pr-6875.docs.buildwithfern.com/nemoclaw |
## Summary Add v0.0.83 release notes to `docs/about/release-notes.mdx` for pre-tag release prep. ## Source Summary - #6773 -> `docs/about/release-notes.mdx`: Shared inference route changes are explicit and fail-safe; status shows recorded route, live route, and drift. - #6875 -> `docs/about/release-notes.mdx`: DGX Station GB300 express setup restored; vLLM storage preflight narrowed. - #6770 -> `docs/about/release-notes.mdx`: Risky Spark vLLM server warning during onboarding. - #6856 -> `docs/about/release-notes.mdx`: Re-onboard reuse preserves tier-default brave/tavily presets. - #6867 -> `docs/about/release-notes.mdx`: Unreachable custom endpoint routed through transport-recovery path. - #6860 -> `docs/about/release-notes.mdx`: Rebuild preflight uses model-aware token field for o-series/GPT-5. - #6845 -> `docs/about/release-notes.mdx`: Corporate CA anchored for image build TLS. - #6833 -> `docs/about/release-notes.mdx`: SSH ControlMaster-delegated forwards recognized in fallback. - #6837 -> `docs/about/release-notes.mdx`: Hermes light skin writes via stdin on macOS. ## Type of Change - [ ] Code change (feature, bug fix, or refactor) - [ ] Code change with doc updates - [x] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Quality Gates - [ ] Tests added or updated for changed behavior - [ ] Existing tests cover changed behavior — justification: - [x] Tests not applicable — justification: doc-only release notes - [x] Docs updated for user-facing behavior changes - [ ] Docs not applicable — justification: - [ ] Sensitive paths changed - [ ] Non-success, skipped, or missing CI check accepted by maintainer ## Verification - [x] PR description includes the DCO sign-off declaration and every commit appears as Verified in GitHub - [x] Normal pre-commit, commit-msg, and pre-push hooks passed - [x] `npm run docs` passes with 0 errors Signed-off-by: Jessica Yaunches <jyaunches@nvidia.com> Signed-off-by: Jessica Yaunches <jyaunches@nvidia.com>
<!-- markdownlint-disable MD041 --> ## Summary DGX Station now uses the existing express-install and onboarding FSM to offer a one-confirmation managed-vLLM install. The express default is the canonical pinned NVIDIA Nemotron 3 Ultra 550B recipe; `--station-deepseek` selects the existing DeepSeek V4 Flash recipe for demos. This PR does not add a parallel launcher, a local-machine image dependency, or a new network mode. The previously proposed `experimental-single-user` profile has been removed because its qualified Docker config ID was not published as a registry manifest. Supersedes #6881 with a clean history after #6875 merged; repository policy disables force-pushing the original PR branch. ## Changes - Detect DGX Station in the existing installer path and offer express setup with no follow-up model/configuration choices after confirmation. - Select Nemotron 3 Ultra by default for Station express setup while preserving `--station-deepseek` as the explicit DeepSeek V4 Flash override. - Keep managed vLLM on NemoClaw's existing Docker bridge topology: `--ipc=host`, explicit `-p 8000:8000`, no `--network host` override. - Pin Ultra to: - model `nvidia/NVIDIA-Nemotron-3-Ultra-550B-A55B-NVFP4` - revision `183968f87ae4cedce3039313cac1fd43d112c578` - served identity `nvidia/nemotron-3-ultra-550b-a55b` - context length `262144` - runtime `vllm/vllm-openai@sha256:0fec7ec5f3e6bc168e54899935fb0557da908a4832a1dbc88e2debcf2f889416` - 150 GiB CPU offload, 16 GiB shared memory, memlock/stack ulimits, MTP, `nemotron_v3`, and `qwen3_coder` - Add the approximately 352 GB Hugging Face cache preflight, post-image-pull capacity recheck, a 3600-second Ultra startup timeout, managed-container ownership protection, and expected-versus-detected model handling when port 8000 is occupied. Inconclusive model-cache probes now require explicit interactive confirmation and fail closed in non-interactive setup unless the exact disk-space override is set. - Normalize the canonical Ultra served alias back to the registered installer slug before managed-vLLM selection. Validate explicit Station-only and conflicting flags before license state, Docker setup, OpenShell build dependencies, or any other host mutation. - Require every effective managed-vLLM runtime to use a pullable immutable `repository@sha256:<manifest>` reference. Bare Docker image/config IDs and mutable tags fail before callbacks, prompts, pulls, or container launch. - Keep explicit pulls against the immutable digest even on cache hits; download and long-lived containers use `--pull=never` afterward so Docker cannot substitute another image. - Update Station express, managed-vLLM, storage, security, and Deferred-validation documentation. ## Distribution and Network Boundary All four shipped managed-vLLM refs were resolved directly from their registries without pulling layers. Each returned HTTP 200 and a `Docker-Content-Digest` equal to the requested digest: - `vllm/vllm-openai@sha256:0fec7ec5f3e6bc168e54899935fb0557da908a4832a1dbc88e2debcf2f889416` — multi-arch index containing Linux ARM64 and AMD64. - `nvcr.io/nvidia/vllm@sha256:9204569b17ee4c0eff75194b8e6e458479c8aee18953b5ab9cf359fcdac659e2` — Linux ARM64. - `nvcr.io/nvidia/vllm@sha256:447995cbb57e6c7cf792cab95e9852e5f62b5fb6d2f39e030fa4eda9a54eadb4` — Linux ARM64. - `nvcr.io/nvidia/vllm@sha256:7be6c2f676c36059a494fe17254e69ae5c677535ba6191044e5fc8e42a91c773` — Linux AMD64. The Station Ultra runtime follows the same network boundary as standard managed vLLM. `0.0.0.0` is inside the container network namespace and Docker publishes only port 8000. Because Docker's default publication can bind on all host interfaces, the existing default-deny firewall guidance still applies; this PR introduces no additional host-network exception. ## Runtime Selection | Station path | Selection | Runtime image | Network | |---|---|---|---| | Express default | Nemotron 3 Ultra 550B | published immutable Docker Hub digest above | existing bridge + `-p 8000:8000` | | `--station-deepseek` | DeepSeek V4 Flash | published immutable NGC digest above | existing bridge + `-p 8000:8000` | | Interactive managed vLLM | existing Station registry/default behavior | published immutable registry digest | existing bridge + `-p 8000:8000` | There is no installer-selectable experimental/local-only profile in this PR. A future qualified single-user recipe can be proposed only after its exact runtime is published as a pullable immutable manifest and integrated through this same registry/FSM path. ## Type of Change - [ ] Code change (feature, bug fix, or refactor) - [x] Code change with doc updates - [ ] Doc only (prose changes, no code sample modifications) - [ ] Doc only (includes code sample changes) ## Quality Gates - [x] Tests added or updated for changed behavior - [x] Docs updated for user-facing behavior changes - [x] Sensitive paths changed (preflight, onboarding, inference, and container launch) - [x] Product/design scope is being coordinated directly with the PM team; code/security review remains requested on the exact head. - [ ] Non-success, skipped, or missing CI check accepted by maintainer — no waiver requested. ## Verification Exact local head: `7624d02c7da6d96bb49058bd49474941740e9bd1`. - [x] Commit and pre-push hooks passed, including repository checks, formatting, lint, ShellCheck, secret scan, installer env-var documentation, and CLI typecheck. - [x] Cumulative focused changed-surface verification: 342 passed, 1 existing skip across installer, vLLM registry/runtime/storage, onboarding FSM, Hermes config/dashboard, CLI dispatch, and docs-contract suites; the latest storage/ordering remediation subset is 144 passed, 1 existing skip. - [x] `npm run typecheck` passed. - [x] `npm run docs:strict` passed with zero errors and two existing Fern warnings. - [x] `npm run check:installer-hash`, `bash -n install.sh scripts/install.sh`, and `shellcheck install.sh scripts/install.sh` passed. - [x] Synthetic merge-tree comparison against the pre-experiment boundary plus current merged-main state found only the intended alias normalization, canonical command assertions, occupied-port assertion, and registry-digest enforcement; canonical Ultra and `--station-deepseek` behavior are unchanged by the cleanup. - [x] All shipped managed-vLLM image manifests resolve remotely at their exact pinned digests. - [ ] Fresh physical DGX Station qualification remains tracked by the existing Deferred platform status; this PR does not claim to advance that status. - [x] No secrets, API keys, credentials, local image IDs, or host-network runtime overrides are committed. --- Signed-off-by: Aaron Erickson <aerickson@nvidia.com> --------- Signed-off-by: Aaron Erickson <aerickson@nvidia.com>
<!-- markdownlint-disable MD041 --> ## Summary The copyable starter prompt now guides a non-technical user through an end-to-end NemoClaw installation one question at a time, without leaving interactive terminal menus running or exposing credentials. It retains the DGX Station Express mapping while adding complete platform readiness, provider, messaging, approval, sudo, Ollama, credential-helper, and documentation-discovery guidance. ## Changes - Point coding agents to official Markdown documentation examples for OpenClaw, Hermes, and LangChain Deep Agents Code, and suggest the NemoClaw docs MCP server when supported. - Collect the operating system, agent, readiness evidence, provider, model, sandbox, web search, messaging, policy, credential, download, administrator-access, and final-install decisions one at a time. - Reproduce Express outcomes non-interactively: use the installed release's maintained Spark model, pin the Station Nemotron Ultra recipe with its approximately 352 GB and Deferred-validation warnings, and preserve the Windows WSL path. - Offer existing vLLM, platform- and agent-eligible Ollama, managed vLLM, OpenRouter, hosted providers, Model Router, and compatible endpoints without starting duplicate local servers. - Preserve the immutable credential-helper and form pins, complete one-time URL, same-port loopback SSH forwarding, single-submission boundary, approved absolute command, account-home scope, and verified-installer requirement. - Define safe sudo behavior, first-build messaging configuration, policy and integration ordering, separate download/notice/final approvals, and outcome verification. - Add regression coverage for credential URL handling, sudo behavior, Ollama eligibility, Express model behavior, approval timing, provider mappings, documentation links, and Deep Agents selection. - [#6875](#6875) -> `docs/resources/starter-prompt.md`: Reflect DGX Station GB300 firmware detection in the starter decision flow. - [#6883](#6883) -> `docs/resources/starter-prompt.md`: Preserve the merged Station Nemotron Ultra Express selectors and explicit non-interactive equivalent. ## Type of Change - [ ] Code change (feature, bug fix, or refactor) - [ ] Code change with doc updates - [ ] Doc only (prose changes, no code sample modifications) - [x] Doc only (includes code sample changes) ## Quality Gates - [x] Tests added or updated for changed behavior — starter-prompt contracts cover credential, sudo, Ollama, provider, Express, approval, documentation-link, and agent-selection behavior. - [ ] Existing tests cover changed behavior — justification: - [ ] Tests not applicable — justification: - [x] Docs updated for user-facing behavior changes - [ ] Docs not applicable — justification: - [x] Sensitive paths changed (security, policy, credentials, preflight, onboarding, inference, runner, sandbox, or messaging) - [x] Sensitive-path review completed or maintainer-approved waiver recorded — implementation-backed review found no remaining must-fix findings; helper pins, focused tests, and docs validation pass. - [ ] Non-success, skipped, or missing CI check accepted by maintainer — check name, approval link, and follow-up issue: ## Verification - [x] PR description includes a `Signed-off-by:` line and every commit appears as `Verified` in GitHub - [x] Normal `pre-commit`, `commit-msg`, and `pre-push` hooks passed, or `npm run check:diff` passed when hooks were skipped or unavailable - [x] Targeted behavior tests pass for the current change set, or tests are marked not applicable above — `npx vitest run test/starter-prompt-docs.test.ts test/changelog-docs.test.ts` passed 18 tests in 2 files. - [ ] Applicable broad gate passed — `npm test` for broad runtime/test-harness changes; `npm run check` for repo-wide validation/coverage changes — not run for this documentation-only change. - [x] Quality Gates section completed with required justifications or waivers - [x] No secrets, API keys, or credentials committed - [ ] `npm run docs` builds without warnings (doc changes only) — passed with zero errors and two existing Fern warnings. - [x] Doc pages follow the [style guide](https://github.com/NVIDIA/NemoClaw/blob/main/docs/CONTRIBUTING.md) (doc changes only) - [ ] New doc pages include SPDX header and frontmatter (new pages only) --- Signed-off-by: Miyoung Choi <miyoungc@nvidia.com> Signed-off-by: San Dang <sdang@nvidia.com> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Reworked the starter onboarding prompt into a stricter step-by-step flow with one-question-at-a-time sequencing, including standardized goal/agent selection and streamlined docs entry points. * Expanded “Express Install” paths (Windows WSL and DGX Spark/Station) and updated platform readiness checks, runtime/provider selection, and local model guidance. * Significantly tightened security for credentials and SSH tunnels with immutable trust-boundary rules, preview/edit/confirm behavior, and stricter policy/approval/checklists. * **Tests** * Updated and expanded starter-prompt documentation tests for redacted key placeholders and stronger security/eligibility/onboarding wording assertions. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Signed-off-by: Carlos Villela <cvillela@nvidia.com> --------- Signed-off-by: Miyoung Choi <miyoungc@nvidia.com> Signed-off-by: San Dang <sdang@nvidia.com> Signed-off-by: Carlos Villela <cvillela@nvidia.com> Co-authored-by: San Dang <sdang@nvidia.com> Co-authored-by: Carlos Villela <cvillela@nvidia.com>
Summary
DGX Station GB300 OEM systems now enter the DGX Station express-install path. Managed vLLM storage preflight is also narrowed to the Docker image pull: it blocks only for a verified shortage, recognizes both default Linux socket spellings, and no longer aborts express onboarding when capacity is inconclusive.
Proof of test
Happy case
Shortage of space
Related Issue
Closes #6757.
Closes #6858.
Changes
StationandGB300as DGX Station for express install./run/docker.sockand/var/run/docker.sock, and honor Docker's documentedDOCKER_CONTEXTprecedence.Type of Change
Quality Gates
Verification
Signed-off-by:line and every commit appears asVerifiedin GitHubpre-commit,commit-msg, andpre-pushhooks passed, ornpm run check:diffpassed when hooks were skipped or unavailablenpm run typecheck:clipassed.npm testfor broad runtime/test-harness changes;npm run checkfor repo-wide validation/coverage changes — command/result:npm run docsbuilds without warnings (doc changes only) — build passed with two pre-existing Fern warnings and no errors.Signed-off-by: San Dang sdang@nvidia.com
Summary by CodeRabbit
Summary by CodeRabbit
New Features
Bug Fixes
--yes/ disk-override handling so only explicitly verified cases can proceed.Documentation
Tests