Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
68 commits
Select commit Hold shift + click to select a range
e324c5c
fix(onboard): confirm readiness after runtime commit
jyaunches Aug 31, 2026
dd0a9ff
fix(onboard): confirm GPU runtime readiness
jyaunches Aug 31, 2026
dd0a669
fix(onboard): scope readiness probes to gateway
jyaunches Aug 31, 2026
661786d
fix(onboard): scope readiness lists to gateway
jyaunches Aug 31, 2026
9f3629d
test(onboard): isolate scoped Hermes readiness
jyaunches Aug 31, 2026
8507e63
fix(onboard): clarify retained sandbox recovery
jyaunches Aug 31, 2026
e990fca
fix(onboard): reject mutable-name portable cleanup
jyaunches Aug 31, 2026
347f252
test(e2e): expect gateway-scoped readiness probes
jyaunches Aug 31, 2026
89b851a
test(security): expect scoped readiness commands
jyaunches Aug 31, 2026
474b30a
docs(onboard): describe readiness phase handling
jyaunches Aug 31, 2026
008e834
docs(onboard): make debounce removal evidence explicit
jyaunches Aug 31, 2026
030f18c
merge(main): refresh readiness handoff fix
jyaunches Aug 31, 2026
6f27475
refactor(onboard): consolidate readiness probes
jyaunches Aug 31, 2026
9a472e2
merge(main): refresh readiness handoff fix
jyaunches Aug 31, 2026
f9dd609
merge(main): refresh readiness handoff fix
jyaunches Aug 31, 2026
4b04b86
fix(onboard): close readiness recovery gaps
jyaunches Aug 31, 2026
383cfe8
fix(onboard): block on recovery persistence
jyaunches Aug 31, 2026
73b3f38
test(onboard): reject unscoped readiness probes
jyaunches Aug 31, 2026
a5ab267
merge: refresh main after review
jyaunches Aug 31, 2026
7d26771
test(onboard): remove unused CLI recovery input
jyaunches Aug 31, 2026
d1a395a
refactor(onboard): reuse readiness waiter
jyaunches Aug 31, 2026
10014df
merge(main): refresh readiness handoff fix
jyaunches Aug 31, 2026
65fe317
fix(onboard): isolate publication wait clock
jyaunches Aug 31, 2026
2308f78
docs(onboard): remove stale readiness override
jyaunches Aug 31, 2026
1219882
merge: resolve conflicts with main
github-actions[bot] Aug 31, 2026
b3bc573
fix(onboard): preserve committed readiness recovery
jyaunches Aug 31, 2026
c5b71b3
fix(onboard): align readiness authority checks
jyaunches Sep 1, 2026
7d28354
merge: synchronize main
jyaunches Sep 1, 2026
aa28948
test(onboard): expose persistence failure assertions
jyaunches Sep 1, 2026
70ab052
docs: explain retained sandbox recovery
jyaunches Sep 1, 2026
e2eb009
Merge branch 'main' into codex/fix-create-dashboard-readiness-handoff
prekshivyas Sep 1, 2026
04ffb8a
fix(onboard): close readiness review gaps
prekshivyas Sep 1, 2026
92bc83f
docs: clarify retained recovery evidence
jyaunches Sep 1, 2026
f7186c3
fix(onboard): stop on publication probe errors
jyaunches Sep 1, 2026
f721f64
merge: synchronize with main
prekshivyas Sep 1, 2026
f5dc929
merge: reconcile upstream branch updates
prekshivyas Sep 1, 2026
37ce787
fix(onboard): close retained recovery gaps
prekshivyas Sep 1, 2026
7c398b6
fix(onboard): preserve APF recovery persistence
prekshivyas Sep 1, 2026
29dc79d
merge: synchronize latest main
prekshivyas Sep 1, 2026
e9bd3ee
fix(onboard): close remaining recovery review gaps
prekshivyas Sep 1, 2026
b2c0271
merge: synchronize latest main
prekshivyas Sep 1, 2026
657df80
fix(onboard): align retained recovery evidence
prekshivyas Sep 1, 2026
dd1cdc2
fix(onboard): share post-create readiness deadline
prekshivyas Sep 1, 2026
3286fb5
fix(onboard): complete shared readiness review
prekshivyas Sep 1, 2026
93f84e2
fix(onboard): close final advisor findings
prekshivyas Sep 1, 2026
8bc3dd1
refactor(onboard): move create cleanup ownership
prekshivyas Sep 1, 2026
249896e
fix(onboard): isolate portable operation recovery
prekshivyas Sep 1, 2026
6b969be
fix(onboard): honor shared readiness deadline
prekshivyas Sep 1, 2026
9ec635c
fix(onboard): register recovery retry lazily
prekshivyas Sep 1, 2026
d925401
test(onboard): prove lazy recovery retry registration
prekshivyas Sep 1, 2026
0df0e81
refactor(onboard): inline sandbox publication wait
prekshivyas Sep 1, 2026
f25e9b3
test(onboard): keep identity deadline proof behavioral
prekshivyas Sep 1, 2026
b257062
refactor(onboard): centralize readiness ownership
prekshivyas Sep 1, 2026
aeba4fe
fix(onboard): preserve resumed readiness deadline
prekshivyas Sep 1, 2026
1752bfd
test(e2e): execute managed startup probe
prekshivyas Sep 1, 2026
692e3cb
fix(e2e): bind cleanup to durable sandbox identity
prekshivyas Sep 1, 2026
87fb85d
fix(e2e): refuse mutable-name sandbox deletion
prekshivyas Sep 1, 2026
950b63d
fix(onboard): scope recovery evidence to gateway
prekshivyas Sep 1, 2026
f6738ad
fix(onboard): preserve post-commit readiness budget
prekshivyas Sep 1, 2026
23e2b65
fix(onboard): retain cleanup recovery evidence
prekshivyas Sep 1, 2026
2399ade
merge(main): resolve onboarding handoff
prekshivyas Sep 1, 2026
89ea9ee
merge: resolve conflicts with main
github-actions[bot] Sep 1, 2026
c9d2f92
merge(main): refresh PR base
prekshivyas Sep 1, 2026
bdf22b7
merge(pr): preserve concurrent branch update
prekshivyas Sep 1, 2026
59ea6e2
test(onboard): consolidate identity deadline coverage
prekshivyas Sep 1, 2026
0cebdf3
test(onboard): consolidate create handoff coverage
prekshivyas Sep 1, 2026
8417d00
fix(onboard): persist verified readiness recovery
prekshivyas Sep 1, 2026
a17dd2c
merge(main): refresh readiness handoff
prekshivyas Sep 2, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions ci/source-architecture-budget.json
Original file line number Diff line number Diff line change
Expand Up @@ -18,7 +18,7 @@
"src/lib/core/ports.ts": 90,
"src/lib/core/shell-quote.ts": 28,
"src/lib/core/url-utils.ts": 30,
"src/lib/core/wait.ts": 36,
"src/lib/core/wait.ts": 37,
"src/lib/credentials/store.ts": 45,
"src/lib/inference/config.ts": 30,
"src/lib/messaging/channels/index.ts": 25,
Expand Down Expand Up @@ -59,7 +59,7 @@
},
"allowedCycles": [],
"maxRootFiles": {
"src/lib/onboard": 306,
"src/lib/onboard": 303,
"src/lib/actions": 18,
"src/lib/actions/sandbox": 182,
"src/lib/state": 39,
Expand Down
10 changes: 7 additions & 3 deletions docs/inference/configure-inference-timeouts.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -20,7 +20,7 @@ Use the error location to select the correct setting.
|---|---|---|
| `NEMOCLAW_AGENT_TIMEOUT` | OpenClaw per-request inference | `600` seconds |
| `NEMOCLAW_LOCAL_INFERENCE_TIMEOUT` | Ollama, vLLM, NIM, and compatible-endpoint onboarding validation paths that read this setting | `180` seconds |
| `NEMOCLAW_SANDBOX_READY_TIMEOUT` | Image build, gateway upload, and in-sandbox boot after creation | `180` seconds |
| `NEMOCLAW_SANDBOX_READY_TIMEOUT` | OpenShell publication, durable identity settlement, and executable readiness after creation; a new full budget starts after a managed runtime commit | `180` seconds |
| `NEMOCLAW_GATEWAY_RECOVERY_WAIT_SECONDS` | OpenShell command re-registration after policy application, plus gateway health and re-registration during managed OpenClaw or Hermes recovery | `30`, `90`, or `120` seconds, depending on the recovery phase |

The readiness timeout does not govern inference requests or provider validation.
Expand Down Expand Up @@ -94,14 +94,18 @@ This variable does not extend the later sandbox-readiness wait.

## Increase the Sandbox Readiness Timeout

Raise `NEMOCLAW_SANDBOX_READY_TIMEOUT` when onboarding creates the sandbox but image build, upload, or boot exceeds 180 seconds.
This can occur during a first run with cold caches or on a remote VM over a slow link.
Raise `NEMOCLAW_SANDBOX_READY_TIMEOUT` when OpenShell needs more than 180 seconds after the create command returns to publish the sandbox, settle its durable identity, or reach executable `Ready`.
After a managed runtime commit, NemoClaw starts a new full budget with this value to verify that the same sandbox returns to executable `Ready`. Local-inference validation does not consume the post-commit budget.
This setting does not extend image build or gateway upload time.

```bash
export NEMOCLAW_SANDBOX_READY_TIMEOUT=600
$$nemoclaw onboard
```

If either readiness deadline expires, NemoClaw preserves the sandbox and records recovery evidence when possible.
Follow [Recover a retained sandbox](../../reference/commands#recover-a-retained-sandbox) before reusing the sandbox name.

## Increase the Recovery Wait

Set `NEMOCLAW_GATEWAY_RECOVERY_WAIT_SECONDS` when OpenShell needs more than 120 seconds to re-register the sandbox after onboarding applies policy presets.
Expand Down
6 changes: 3 additions & 3 deletions docs/reference/commands.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -3954,7 +3954,7 @@ The following environment variables tune onboard-time and recovery wall-clock li
| `NEMOCLAW_OLLAMA_PULL_TIMEOUT` | `1800` (30 minutes) | Wall-clock timeout for `ollama pull` during onboard, in seconds. Accepts integer or float values. Already-downloaded layers are kept; re-running the pull resumes them. |
| `NEMOCLAW_HF_DOWNLOAD_STALL_TIMEOUT` | `600` (10 minutes) | Maximum silence between Hugging Face download output during onboard. A positive finite value in seconds overrides the default, up to the Node.js timer limit of about 24.8 days. Blank, invalid, non-positive, sub-millisecond, and oversized values use the default. This is not a total download limit. Increase it only when a working download can produce no output for ten minutes. |
| `NEMOCLAW_LOCAL_INFERENCE_TIMEOUT` | `180` | Wall-clock timeout for the inference-server validation probe during onboard, in seconds. Raise on slow networks or for very large prompts. |
| `NEMOCLAW_SANDBOX_READY_TIMEOUT` | `180` | Wall-clock timeout for post-create readiness, in seconds. Raise the timeout when the managed-image pull, explicit custom image build, gateway upload, or in-sandbox boot exceeds the default (typical on 70B+ models, first-time gateway uploads over slow links, or DGX Station / remote-VM first runs). Ordinary onboarding deletes the partially created sandbox when the deadline expires and prints the retry hint. Portable OpenClaw onboarding instead preserves the sandbox when NemoClaw cannot verify its runtime identity. |
| `NEMOCLAW_SANDBOX_READY_TIMEOUT` | `180` | Wall-clock deadline for initial post-create publication, durable identity settlement, and executable readiness. After a managed runtime commit, NemoClaw starts a new full deadline with this value to verify that the same sandbox returns to executable `Ready`; local-inference validation does not consume that new budget. A readiness failure preserves the sandbox. Follow the [identity-bound retained-sandbox recovery procedure](#recover-a-retained-sandbox) before reusing its name. |
| `NEMOCLAW_SANDBOX_READY_ERROR_DEBOUNCE` | `30` | Consecutive `Error`-phase polls the post-create readiness wait tolerates before treating `Error` as terminal. Polling starts at 250ms and backs off to a 2-second cap, while `NEMOCLAW_SANDBOX_READY_TIMEOUT` remains the overall deadline. The gateway can briefly report a just-created sandbox in `Error` while it re-registers the sandbox (seen on DGX Spark); the debounce lets that transient recover to `Ready`. Every terminal observation outside the `Error` phase, including one with no reported phase, fails immediately. Set to `1` to restore fast-fail on the first `Error` poll. |
| `NEMOCLAW_GATEWAY_RECOVERY_WAIT_SECONDS` | `30`, `90`, or `120`, depending on the recovery phase | Wall-clock timeout for OpenShell command re-registration after policy application, plus gateway health and re-registration during managed OpenClaw or Hermes recovery. A valid finite, nonnegative value overrides the internal budget for the current recovery phase. |

Expand Down Expand Up @@ -3989,11 +3989,11 @@ $$nemoclaw <sandbox-name> recover

</AgentOnly>

If the Ollama pull or post-create readiness timeout fires, onboarding emits the elapsed budget plus a hint to raise the relevant variable. The Ollama pull preserves its partial download for the next attempt. The ordinary post-create readiness wait deletes the orphaned sandbox first so the next `$$nemoclaw onboard` starts without that partially created sandbox.
If the Ollama pull or post-create readiness timeout fires, onboarding emits the elapsed budget plus a hint to raise the relevant variable. The Ollama pull preserves its partial download for the next attempt. A post-create readiness failure preserves the sandbox and records recovery evidence when possible. The sandbox name remains blocked until the identity-bound retained-sandbox recovery procedure succeeds.

<AgentOnly variant="openclaw">

For portable OpenClaw onboarding, NemoClaw instead leaves the sandbox in place when it cannot verify the exact runtime identity. Inspect it with `openshell sandbox list` and `$$nemoclaw <name> status`, then follow the recovery guidance from `status`.
Portable OpenClaw onboarding also preserves the sandbox when NemoClaw cannot verify the exact runtime identity. Inspect it with `openshell sandbox list` and `$$nemoclaw <name> status`, then follow the recovery guidance from `status`.

</AgentOnly>

Expand Down
41 changes: 24 additions & 17 deletions docs/reference/troubleshooting.mdx
Original file line number Diff line number Diff line change
Expand Up @@ -2002,24 +2002,38 @@ This is a separate budget from `NEMOCLAW_LOCAL_INFERENCE_TIMEOUT`. It covers the

<AgentOnly variant="openclaw,hermes">

For a newly created OpenClaw or Hermes sandbox, `Ready` is not the final acceptance signal. Within this same budget, NemoClaw also requires OpenShell to return a durable sandbox ID and accept `openshell sandbox exec --name <sandbox> -- true`. NemoClaw keeps waiting only when OpenShell returns its exact `sandbox is not ready` response. A missing or malformed ID, or another command failure, stops the wait. Ordinary onboarding then follows the failed-creation cleanup path. Portable OpenClaw onboarding preserves the sandbox as described below.
For a newly created OpenClaw or Hermes sandbox, `Ready` is not the final acceptance signal. Within the initial post-create budget, NemoClaw also requires OpenShell to return a durable sandbox ID and accept `openshell sandbox exec --name <sandbox> -- true`. NemoClaw keeps waiting only when OpenShell reports exact sandbox absence or its exact `sandbox is not ready` response. A missing or malformed ID, or another command failure, stops the wait. NemoClaw preserves the sandbox, saves create-attempt evidence when possible, and requires identity-bound recovery instead of ordinary failed-creation cleanup.

After a managed runtime commit, NemoClaw starts a new full `NEMOCLAW_SANDBOX_READY_TIMEOUT` budget to verify that the same durable sandbox returns to executable `Ready`. The local-inference validation timeout remains separate and does not consume this post-commit budget.

If onboarding reports that the managed runtime commit completed but the same sandbox did not return to executable `Ready`, stop before retrying.
NemoClaw keeps the sandbox, prints its create-attempt label and a one-way durable identity fingerprint, and does not start dashboard forwarding.
It saves that evidence in the retained recovery record when persistence succeeds.
Do not use OpenShell mutable-name deletion.
When NemoClaw confirms that it saved the record and the record contains a durable identity fingerprint, run `$$nemoclaw <sandbox-name> destroy`; NemoClaw verifies the retained sandbox identity before cleanup. Follow [Recover a retained sandbox](commands#recover-a-retained-sandbox) for the result-specific recovery steps.
If the saved record has no durable identity fingerprint, preserve the terminal output and give the create-attempt label to an OpenShell administrator so they can identify and remove the exact sandbox.
If NemoClaw reports that it could not save the record, preserve the terminal output and ask an OpenShell administrator to identify the exact sandbox from gateway or controller evidence; the recovery-only session remains blocked until NemoClaw can save the durable recovery record.

</AgentOnly>

The 180-second default fits typical workstations but can be exceeded when:
The 180-second default can be exceeded after the create command returns when:

- The owning gateway needs longer to publish the new sandbox and its durable identity.
- The in-sandbox agent, policy, or model runtime needs longer to become executable.
- A managed runtime commit needs longer to re-register the same sandbox as executable `Ready`.

- The host is building or uploading the sandbox image for the first time (cold caches, slow link).
- The selected model is large (70B+ parameters or 4-bit/8-bit quantisations that take time to memory-map).
- Onboarding runs on a remote VM where image upload to the gateway streams over the network (for example DGX Station first-run installer).
For a post-create readiness failure, complete the retained-sandbox identity-bound cleanup above before you retry the same sandbox name.
If that cleanup remains unresolved, use a different explicit sandbox name for a separate onboarding run.

Raise the budget before re-running onboard:
After identity-bound cleanup succeeds, raise the budget and rerun onboarding:

```bash
export NEMOCLAW_SANDBOX_READY_TIMEOUT=600
$$nemoclaw onboard
$$nemoclaw onboard --name <sandbox-name>
```

The variable accepts seconds and applies to the readiness wait only. When the ordinary create deadline expires, NemoClaw tries to delete the partially created sandbox. After successful cleanup, the output ends with `Retry: $$nemoclaw onboard`. If cleanup fails, NemoClaw instead reports that the failed sandbox could not be removed and prints `Manual cleanup: openshell sandbox delete "<name>"`.
The variable accepts seconds and sets the shared post-create readiness deadline.
NemoClaw removes the recovery state only after identity-bound cleanup succeeds.

<AgentOnly variant="openclaw">

Expand All @@ -2036,7 +2050,7 @@ openshell sandbox list
$$nemoclaw <name> status
```

If onboarding instead reports that the sandbox "did not re-register with OpenShell after policy application," the same timeout controls that post-policy command-readiness probe. Raise the budget before retrying, then inspect the same gateway and sandbox status if re-registration still fails.
If onboarding instead reports that the sandbox "did not re-register with OpenShell after policy application," `NEMOCLAW_GATEWAY_RECOVERY_WAIT_SECONDS` controls that post-policy command-readiness probe. Increase that recovery wait before retrying, then inspect the same gateway and sandbox status if re-registration still fails.

### Sandbox onboard fails with "entered Error phase before it became ready"

Expand All @@ -2048,18 +2062,11 @@ Onboarding ends with:

On a fresh onboard the OpenShell gateway can (re)start its supervisor session and re-register the just-created sandbox. During that window `openshell sandbox list` briefly reports the sandbox in the transient `Error` phase before it flips to `Ready`, as seen on DGX Spark when supervisor restart races the sandbox bootstrap.

NemoClaw polls immediately, starts retrying after 250ms, and backs off to a 2-second cap.
NemoClaw polls immediately. A single-observation wait retries after 250 ms and backs off to a two-second cap. A stable-readiness wait that requires two consecutive `Ready` observations checks every two seconds.
It tolerates 30 consecutive `Error` observations by default so this transient recovers on its own.
Only `Error` that persists through the debounce count is terminal, unless the overall `NEMOCLAW_SANDBOX_READY_TIMEOUT` deadline expires first.
Every terminal observation outside the `Error` phase, including one with no reported phase, fails immediately.

If your host needs more observations for slower re-registration, raise the debounce. Raise `NEMOCLAW_SANDBOX_READY_TIMEOUT` too if the overall deadline is too short. To fail fast on the first `Error` poll, set the debounce to `1`:

```bash
export NEMOCLAW_SANDBOX_READY_ERROR_DEBOUNCE=1
$$nemoclaw onboard
```

If the failure persists after the debounce, the sandbox is stuck. Inspect the retained diagnostics and gateway state:

```bash
Expand Down
Loading
Loading