fix(ci): switch test-e2e-sandbox to ubuntu-latest on self-hosted workflow - #2294
Conversation
…flow The snapshot rollback test (test 7 in e2e-test.sh) fails 100% on linux-amd64-cpu4 NVIDIA self-hosted runners but passes reliably on GitHub-hosted ubuntu-latest. The failure is in rollbackFromSnapshot() which silently catches an exception during renameSync/cpSync on the self-hosted runner's Docker storage driver. Confirmed by temporarily switching the runner in PR #2288 — test passed immediately on ubuntu-latest. The pr-self-hosted workflow has never had a passing run since it was added in #2121. The other E2E jobs (build-sandbox-images, test-e2e-gateway-isolation) remain on linux-amd64-cpu4 since they pass there. Signed-off-by: Julie Yaunches <jyaunches@nvidia.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (1)
📝 WalkthroughWalkthroughModified the GitHub Actions workflow to change the e2e sandbox test runner from Changes
Estimated code review effort🎯 1 (Trivial) | ⏱️ ~3 minutes Poem
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
ericksoa
left a comment
There was a problem hiding this comment.
Confirmed this is the same rollback failure blocking all PRs on pr-self-hosted. The fix is minimal and targeted — only the affected job moves to ubuntu-latest, other E2E jobs stay on the self-hosted runner. test-e2e-sandbox already passed on this PR. LGTM.
## Summary
Fix snapshot rollback failure on NVIDIA self-hosted runners and move all
E2E jobs back to self-hosted.
## Root Cause
NVIDIA self-hosted runners use Docker with the **containerd overlayfs
snapshotter** (`io.containerd.snapshotter.v1`), while GitHub-hosted
runners use the legacy **`overlay2`** driver. On the containerd
snapshotter, directories can span different overlay layers, causing
`rename(2)` to return **`EXDEV` (cross-device link not permitted)**.
`rollbackFromSnapshot()` and `cutoverHost()` in `snapshot.ts` both used
bare `renameSync()` which fails with EXDEV on these runners. The bare
`catch {}` swallowed the error and returned `false`, producing the
"Rollback returned false" failure with no diagnostic output.
This also caused a secondary **`ERR_FS_CP_EINVAL`** — since the rename
failed, `.openclaw` still existed when `cpSync` ran, and it detected it
would be copying symlink targets into themselves.
## Diagnostic Evidence
Confirmed via a matrix job across 4 runners:
| Runner | Storage Driver | `renameSync` | `cpSync` |
|--------|---------------|-------------|---------|
| `linux-amd64-cpu4` | `overlayfs` (containerd) | ❌ EXDEV | ❌
ERR_FS_CP_EINVAL |
| `linux-amd64-cpu8` | `overlayfs` (containerd) | ❌ EXDEV | ❌
ERR_FS_CP_EINVAL |
| `linux-amd64-cpu16` | `overlayfs` (containerd) | ❌ EXDEV | ❌
ERR_FS_CP_EINVAL |
| `ubuntu-latest` | `overlay2` (legacy) | ✅ | ✅ |
## Fix
Add `moveSync()` helper that tries `renameSync` first (fast same-device
path), then falls back to `cpSync` + `rmSync` on EXDEV. Both
`cutoverHost()` and `rollbackFromSnapshot()` now use `moveSync()`
instead of bare `renameSync()`.
After the fix, all runners pass:
| Runner | Before | After |
|--------|--------|-------|
| `linux-amd64-cpu4` | ❌ Rollback returned false | ✅ PASS |
| `linux-amd64-cpu8` | ❌ Rollback returned false | ✅ PASS |
| `linux-amd64-cpu16` | ❌ Rollback returned false | ✅ PASS |
| `ubuntu-latest` | ✅ PASS | ✅ PASS |
## Changes
### `nemoclaw/src/blueprint/snapshot.ts`
- Add `moveSync()` — cross-device-safe move with EXDEV fallback
- `cutoverHost()` — use `moveSync` instead of `renameSync`
- `rollbackFromSnapshot()` — use `moveSync` instead of `renameSync`
(both the archive step and the recovery path)
### `nemoclaw/src/blueprint/snapshot.test.ts`
- Add `moveSync` unit tests: same-device rename, EXDEV fallback (cpSync
+ rmSync), non-EXDEV re-throw
- Add `rmSync` to the in-memory fs mock
### `.github/workflows/pr-self-hosted.yaml`
- Move `test-e2e-sandbox` back to `linux-amd64-cpu4` (was on
`ubuntu-latest` as a workaround since PR #2294)
- Remove temporary diagnostic matrix job and `test/diag-container-fs.sh`
## Follow-up
- [ ] Phase 4: Remove duplicate `sandbox-images-and-e2e` from `pr.yaml`
(separate PR)
## Type of Change
- [x] Code change (feature, bug fix, or refactor)
## Verification
- [x] `moveSync` unit tests pass (22/22 snapshot tests green)
- [x] `test-e2e-sandbox` passes on `linux-amd64-cpu4` (NVIDIA
self-hosted)
- [x] `test-e2e-gateway-isolation` passes on `linux-amd64-cpu4`
- [x] All jobs pass on `ubuntu-latest` (no regression)
- [x] No secrets, API keys, or credentials committed
## AI Disclosure
- [x] AI-assisted — tool: Claude Code (pi agent)
---
Signed-off-by: Julie Yaunches <jyaunches@nvidia.com>
---------
Signed-off-by: Julie Yaunches <jyaunches@nvidia.com>
) ## Summary Fix snapshot rollback failure on NVIDIA self-hosted runners and move all E2E jobs back to self-hosted. ## Root Cause NVIDIA self-hosted runners use Docker with the **containerd overlayfs snapshotter** (`io.containerd.snapshotter.v1`), while GitHub-hosted runners use the legacy **`overlay2`** driver. On the containerd snapshotter, directories can span different overlay layers, causing `rename(2)` to return **`EXDEV` (cross-device link not permitted)**. `rollbackFromSnapshot()` and `cutoverHost()` in `snapshot.ts` both used bare `renameSync()` which fails with EXDEV on these runners. The bare `catch {}` swallowed the error and returned `false`, producing the "Rollback returned false" failure with no diagnostic output. This also caused a secondary **`ERR_FS_CP_EINVAL`** — since the rename failed, `.openclaw` still existed when `cpSync` ran, and it detected it would be copying symlink targets into themselves. ## Diagnostic Evidence Confirmed via a matrix job across 4 runners: | Runner | Storage Driver | `renameSync` | `cpSync` | |--------|---------------|-------------|---------| | `linux-amd64-cpu4` | `overlayfs` (containerd) | ❌ EXDEV | ❌ ERR_FS_CP_EINVAL | | `linux-amd64-cpu8` | `overlayfs` (containerd) | ❌ EXDEV | ❌ ERR_FS_CP_EINVAL | | `linux-amd64-cpu16` | `overlayfs` (containerd) | ❌ EXDEV | ❌ ERR_FS_CP_EINVAL | | `ubuntu-latest` | `overlay2` (legacy) | ✅ | ✅ | ## Fix Add `moveSync()` helper that tries `renameSync` first (fast same-device path), then falls back to `cpSync` + `rmSync` on EXDEV. Both `cutoverHost()` and `rollbackFromSnapshot()` now use `moveSync()` instead of bare `renameSync()`. After the fix, all runners pass: | Runner | Before | After | |--------|--------|-------| | `linux-amd64-cpu4` | ❌ Rollback returned false | ✅ PASS | | `linux-amd64-cpu8` | ❌ Rollback returned false | ✅ PASS | | `linux-amd64-cpu16` | ❌ Rollback returned false | ✅ PASS | | `ubuntu-latest` | ✅ PASS | ✅ PASS | ## Changes ### `nemoclaw/src/blueprint/snapshot.ts` - Add `moveSync()` — cross-device-safe move with EXDEV fallback - `cutoverHost()` — use `moveSync` instead of `renameSync` - `rollbackFromSnapshot()` — use `moveSync` instead of `renameSync` (both the archive step and the recovery path) ### `nemoclaw/src/blueprint/snapshot.test.ts` - Add `moveSync` unit tests: same-device rename, EXDEV fallback (cpSync + rmSync), non-EXDEV re-throw - Add `rmSync` to the in-memory fs mock ### `.github/workflows/pr-self-hosted.yaml` - Move `test-e2e-sandbox` back to `linux-amd64-cpu4` (was on `ubuntu-latest` as a workaround since PR NVIDIA#2294) - Remove temporary diagnostic matrix job and `test/diag-container-fs.sh` ## Follow-up - [ ] Phase 4: Remove duplicate `sandbox-images-and-e2e` from `pr.yaml` (separate PR) ## Type of Change - [x] Code change (feature, bug fix, or refactor) ## Verification - [x] `moveSync` unit tests pass (22/22 snapshot tests green) - [x] `test-e2e-sandbox` passes on `linux-amd64-cpu4` (NVIDIA self-hosted) - [x] `test-e2e-gateway-isolation` passes on `linux-amd64-cpu4` - [x] All jobs pass on `ubuntu-latest` (no regression) - [x] No secrets, API keys, or credentials committed ## AI Disclosure - [x] AI-assisted — tool: Claude Code (pi agent) --- Signed-off-by: Julie Yaunches <jyaunches@nvidia.com> --------- Signed-off-by: Julie Yaunches <jyaunches@nvidia.com>
Summary
Switch
test-e2e-sandboxfromlinux-amd64-cpu4(NVIDIA self-hosted) toubuntu-latest(GitHub-hosted) in thepr-self-hostedworkflow. The snapshot rollback test fails 100% on the self-hosted runner but passes reliably on GitHub-hosted runners.Related Issue
Unblocks all PRs gated on
pr-self-hosted— the workflow has never had a passing run since it was added in #2121.Changes
.github/workflows/pr-self-hosted.yaml: Changetest-e2e-sandboxruns-onfromlinux-amd64-cpu4toubuntu-latestbuild-sandbox-images,test-e2e-gateway-isolation) remain onlinux-amd64-cpu4since they pass thereDiagnosis
The
rollbackFromSnapshot()function innemoclaw/src/blueprint/snapshot.tsdoesrenameSync+cpSyncon/sandbox/.openclaw(which contains symlinks to.openclaw-data). On the self-hosted runner's Docker environment, one of these operations throws — but the barecatch {}swallows the error and returnsfalse, producing the "Rollback returned false" failure with no diagnostic output.Confirmed in PR #2288 by temporarily switching the runner —
test-e2e-sandboxpassed immediately onubuntu-latest.Type of Change
Verification
npx prek run --all-filespassesnpm testpassesAI Disclosure
Signed-off-by: Julie Yaunches jyaunches@nvidia.com
Summary by CodeRabbit