diff --git a/docs/reference/inference-profiles.md b/docs/reference/inference-profiles.md index eb518b94ec5..86d246313b4 100644 --- a/docs/reference/inference-profiles.md +++ b/docs/reference/inference-profiles.md @@ -66,8 +66,7 @@ If validation fails, the wizard does not continue to sandbox creation. ## Local Providers -Local providers are available behind the `NEMOCLAW_EXPERIMENTAL=1` gate. -These use the same routed `inference.local` pattern, but the upstream runtime is local to the host. +Local providers use the same routed `inference.local` pattern, but the upstream runtime runs on the host rather than in the cloud. - Local Ollama - Local NVIDIA NIM diff --git a/spark-install.md b/spark-install.md index 1eeb1187e84..c5f2fd0a21b 100644 --- a/spark-install.md +++ b/spark-install.md @@ -2,19 +2,147 @@ > **WIP** — This page is actively being updated as we work through Spark installs. Expect changes. +## Prerequisites + +- **Docker** (pre-installed, v28.x) +- **Node.js 22** (installed by the install.sh) +- **OpenShell CLI** (installed via the Quick Start steps below) +- **NVIDIA API Key** from [build.nvidia.com](https://build.nvidia.com) — prompted on first run + ## Quick Start ```bash -# Clone and install +# Install OpenShell: +curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh + +# Clone NemoClaw: git clone https://github.com/NVIDIA/NemoClaw.git cd NemoClaw -sudo npm install -g . -# Spark-specific setup (configures Docker for cgroup v2, then runs normal setup) -nemoclaw setup-spark +# Spark-specific setup (For details see [What's Different on Spark](#whats-different-on-spark)) +sudo ./scripts/setup-spark.sh + +# Install NemoClaw using the NemoClaw/install.sh: +./install.sh + +# Alternatively, you can use the hosted install script: +curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash +``` + +## Verifying Your Install + +```bash +# Check sandbox is running +nemoclaw my-assistant connect + +# Inside the sandbox, talk to the agent: +openclaw agent --agent main --local -m "hello" --session-id test +``` + +## Uninstall (perform this before re-installing) + +```bash +# Uninstall NemoClaw (Remove OpenShell sandboxes, gateway, NemoClaw providers, related Docker containers, images, volumes and configs) +nemoclaw uninstall +``` + +## Setup Local Inference (Ollama) + +Use this to run inference locally on the DGX Spark's GPU instead of routing to cloud. + +### Verify the NVIDIA Container Runtime + +```bash +docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi ``` -That's it. `setup-spark` handles everything below automatically. +If this fails, configure the NVIDIA runtime and restart Docker: + +```bash +sudo nvidia-ctk runtime configure --runtime=docker +sudo systemctl restart docker +``` + +### Install Ollama + +```bash +curl -fsSL https://ollama.com/install.sh | sh +``` + +Verify it is running: + +```bash +curl http://localhost:11434 +``` + +### Pull and Pre-load a Model + +Download Nemotron 3 Super 120B (~87 GB; may take several minutes): + +```bash +ollama pull nemotron-3-super:120b +``` + +Run it briefly to pre-load weights into unified memory, then exit: + +```bash +ollama run nemotron-3-super:120b +# type /bye to exit +``` + +### Configure Ollama to Listen on All Interfaces + +By default Ollama binds to `127.0.0.1`, which is not reachable from inside the sandbox container. Configure it to listen on all interfaces: + +> **Note:** `OLLAMA_HOST=0.0.0.0` exposes Ollama on your network. If you're not on a trusted LAN, restrict access with host firewall rules (`ufw`, `iptables`, etc.). + +```bash +sudo mkdir -p /etc/systemd/system/ollama.service.d +printf '[Service]\nEnvironment="OLLAMA_HOST=0.0.0.0"\n' | sudo tee /etc/systemd/system/ollama.service.d/override.conf + +sudo systemctl daemon-reload +sudo systemctl restart ollama +``` + +Verify Ollama is listening on all interfaces: + +```bash +sudo ss -tlnp | grep 11434 +``` + +### Install OpenShell and NemoClaw + +```bash +# If the OpenShell and NemoClaw are already installed, uninstall them. A fresh NemoClaw install will run onboard with local inference options. +nemoclaw uninstall + +# Install OpenShell and NemoClaw +curl -LsSf https://raw.githubusercontent.com/NVIDIA/OpenShell/main/install.sh | sh +curl -fsSL https://www.nvidia.com/nemoclaw.sh | bash +``` + +When prompted for **Inference options**, select **Local Ollama**, then select the model you pulled. + +### Connect and Test + +```bash +# Connect to the sandbox +nemoclaw my-assistant connect +``` + +Inside the sandbox, first verify `inference.local` is reachable directly (must use HTTPS — the proxy intercepts `CONNECT inference.local:443`): + +```bash +curl -sf https://inference.local/v1/models +# Expected: JSON response listing the configured model +# Exits non-zero on HTTP errors (403, 503, etc.) — failure here indicates a proxy routing regression +``` + +Then talk to the agent: + +```bash +openclaw agent --agent main --local -m "Which model and GPU are in use?" --session-id test +``` ## What's Different on Spark @@ -42,28 +170,6 @@ Failed to start ContainerManager: failed to initialize top level QOS containers **Fix**: `setup-spark` sets `"default-cgroupns-mode": "host"` in `/etc/docker/daemon.json` and restarts Docker. This makes all containers use the host cgroup namespace, which is what k3s needs. -## Prerequisites - -These should already be on your Spark: - -- **Docker** (pre-installed, v28.x) -- **Node.js 22** — if not installed: - - ```bash - curl -fsSL https://deb.nodesource.com/setup_22.x | sudo -E bash - - sudo apt-get install -y nodejs - ``` - -- **OpenShell CLI**: - - ```bash - ARCH=$(uname -m) # aarch64 on Spark - curl -fsSL "https://github.com/NVIDIA/OpenShell/releases/latest/download/openshell-linux-${ARCH}" -o /usr/local/bin/openshell - chmod +x /usr/local/bin/openshell - ``` - -- **NVIDIA API Key** from [build.nvidia.com](https://build.nvidia.com) — prompted on first run - ## Manual Setup (if setup-spark doesn't work) ### Fix Docker cgroup namespace @@ -93,12 +199,6 @@ sudo usermod -aG docker $USER newgrp docker # or log out and back in ``` -### Then run the onboard wizard - -```bash -nemoclaw onboard -``` - ## Known Issues | Issue | Status | Workaround | @@ -109,22 +209,6 @@ nemoclaw onboard | Image pull failure (k3s can't find built image) | OpenShell bug | `openshell gateway destroy && openshell gateway start`, re-run setup | | GPU passthrough | Untested on Spark | Should work with `--gpu` flag if NVIDIA Container Toolkit is configured | -## Verifying Your Install - -```bash -# Check sandbox is running -openshell sandbox list -# Should show: nemoclaw Ready - -# Test the agent -openshell sandbox connect nemoclaw -# Inside sandbox: -nemoclaw-start openclaw agent --agent main --local -m 'hello' --session-id test - -# Monitor network egress (separate terminal) -openshell term -``` - ## Architecture Notes ```text