Skip to content

w759: fix the Windows EPERM and VM 'Domain is already active' test flakes (and main's red vault-rename check) - #252

Merged
bryding merged 3 commits into
mainfrom
w759-test-flakes
Oct 9, 2026
Merged

bryding merged 3 commits into
mainfrom
w759-test-flakes

Conversation

@bryding

@bryding bryding commented Oct 9, 2026

Copy link
Copy Markdown
Contributor

Request: w759

TL;DR: Two flakes fixed at their cause, plus a third that also made local Windows runs red. Main had gone red from two merges colliding, so that's fixed here too.

1. Windows EPERM removing a test machine's folder

  • Error (BEAST, different tests each run): Error: EPERM, Permission denied: \\?\…\ff-msb-XXXXXX at fs.rmSync in server/testMachine.ts:77 (TestRepos.cleanup), called from TestMachine.stop().
  • Cause, measured: I instrumented the cleanup to record live child processes when it fails. In all 6 failures traced, a git -C <tmp>\ffsb\<sandbox> status --porcelain=v2 --branch and/or git … log -1 was still running, started 130–280 ms earlier by the daemon's pool (readGitStatus, machine/sandboxes.ts:705 and the other gitStatus callers). git -C makes the worktree its current folder, and Windows refuses to delete a folder a live process is in. The daemon's stop waits for none of these. A plain retry 100 ms later succeeded every time.
  • Why CI doesn't see it, sourced: CI's Windows jobs run Node 24.21.0 (Found in cache @ …\node\24.21.0). From that release, rmSync retries Windows' access-denied (std::errc::permission_denied) and actually waits between tries (fs: treat std::errc::permission_denied as EPERM error nodejs/node#64698, backport b269616936). BEAST runs nvm4w Node 23.7.0. Its can_omit_error covers EBUSY/EMFILE/ENFILE/ENOTEMPTY/EPERM but not permission_denied, so maxRetries: 5 never retries this error. Its Windows sleep is also Sleep(i * retryDelay / 1000), which is 0 ms. The race is the same on CI; CI's Node just hides it.
  • Fix (no retry added): fakePoolDeps tracks the git/copy/remove it starts and gets a close(): once closing, it starts nothing new, and it waits for what is still running. TestMachine.stop() calls it before deleting the folder.

2. A sandbox PR the daemon's next report took away (4 of 10 runs on BEAST)

server/providerLedger.test.ts "ledger check both ways" (nothing changed, nothing sent 1 !== 0) and server/devRequests.test.ts w278 (no "dev_update" from the portal in 5000 ms) wrote a sandbox's PR straight into the portal's machine record. mergeSandboxes (server/machines.ts:226) replaces a machine sandbox's git with the daemon's on its next report, so the PR vanished whenever a report landed mid-test. New TestMachine.gitLook(name, extra): the daemon's own git look reports the PR, and the tests wait for the portal to have it. The board digest (server/intake.ts:85) compares only verdict/watch/ids, so later reports with the same branch and PR change nothing.

3. VM e2e "Domain is already active"

  • Error: error: Domain is already active right after fff-vm: 10/12 … in the second ("idempotent") host install. That is install.sh:617: [ "$(dom_state)" = running ] || virsh start.
  • Cause: dom_state was v domstate … | head -n 1 || echo missing. In libvirt 12 tools/vsh.c, vshPrintVa does fputs + fflush per print: virsh prints running\n, then after the command a separate \n (vshPrintExtra(ctl, "\n")). When head had already exited, that second write killed virsh with SIGPIPE, pipefail failed the pipeline, and || echo missing added a line. The answer was running\nmissing, so the running VM was started again.
  • Evidence: all 3 occurrences in the repo's CI history are this: run 37894701861 (w744: fff-vm vault-sync no longer removes the claude-tokens/ pool entries; --guest-only keeps the host scripts in step #245, 2026-10-09), run 37540521850 (main, 2026-10-06) and run 37422622262 (2026-10-06). All were on the ubuntu-26.04 host, in the second install, at step 10. The vm-logs artifacts show libvirtd: End of file while reading data: Input/output error in the same second each time, which is a client that died without closing its connection. In 37540521850 the watch timer ran 3 s earlier and had finished, which rules out a race with fff-vm watch. Why only the 26.04 host is a guess (a different head or scheduler timing there).
  • Fix: dom_state reads virsh's whole answer (s=$(v domstate …) || missing, first line). deploy/vm/test/fff-vm-nightly.test.sh case 10 holds the fake virsh's blank line back 0.3 s. That makes the old pipeline fail every time (measured locally: it printed running\nmissing), and the new one prints running.
  • "no alert about the refused size" (ci-vm-e2e.sh:690, from w753): none of the 44 failed and 36 cancelled VM runs on record contain it, so it can't be confirmed from logs. It can share this cause: the nightly's [ "$(dom_state)" = running ] || { … nothing to do; return 0; } (fff-vm:269) would skip with no alert on the same misread. This fix covers that path too.

4. Main was red (separate commit)

w748 (#247) added fffctl vault rename, and w745 (#244) fails on any unclassified vault form, so main's CI run 37907217293 failed on Windows and Ubuntu. rename is now a changes form: refused to the read-only ops worker like remove (fff-ops-priv already refuses any form it does not pin). docs/ops-worker.md lists it.

Evidence

  • Before (BEAST, Node 23.7.0): 7 of 7 full runs failed, with 1–3 EPERM each; 4 of 10 also had the PR-race failure.
  • After: 10 of 10 full runs green (1268 pass, 0 fail, 0 EPERM). After merging main: 3 of 3 green (1278 pass, 0 fail).
  • npx tsc -p . clean. shellcheck on the two changed scripts: only the existing SC1091 notes.
  • The VM fix: fff-vm-nightly.test.sh case 10 in the Lint job, and this PR's nested-VM e2e runs.

Known-flake lists: none in this repo or in the final-factory-agents skills names these. The "rerun failed" wording was only in dispatcher briefs.

🤖 Generated with Claude Code

bryding and others added 3 commits October 9, 2026 02:50
…ds virsh whole

Windows EPERM in the unit suite: TestMachine.stop() removed the machine's
temp folder while a `git -C <worktree> status`/`log` the daemon's pool had
started was still running in it (all 6 failures traced on BEAST). The daemon's
stop waits for none of them. fakePoolDeps now starts no git once closing and
stop() waits for the git still running. CI did not see it because Node 24.21
retries Windows' access-denied in rmSync (nodejs/node#64698); BEAST's 23.7
does not, and waits 0 ms between tries.

Two tests wrote a sandbox's PR into the portal's record and lost it to the
daemon's next report (providerLedger "ledger check both ways", devRequests
w278): TestMachine.gitLook makes the daemon's git look report it.

VM e2e "Domain is already active": dom_state piped virsh domstate into
head -n 1; virsh writes its trailing blank line separately, died of SIGPIPE
when head had gone, and pipefail made the answer "running\nmissing", so the
second install started the running VM. It now reads virsh's whole answer;
fff-vm-nightly.test.sh case 10 delays the blank line to prove it.

Request: w759

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
w748 (#247) added `fffctl vault rename`; w745 (#244) made
server/opsFffctlForms.test.ts fail on any vault form FFFCTL_FORMS does not
classify. Both merged, and main's CI failed on Windows and Ubuntu (run
37907217293). rename renames a vault entry, so it is a change: refused to the
read-only ops worker like remove and put (fff-ops-priv already refuses any
form it does not pin). docs/ops-worker.md lists it with the others.

Part of: w759

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@bryding
bryding merged commit 702d299 into main Oct 9, 2026
13 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant