Repository navigation
w759: fix the Windows EPERM and VM 'Domain is already active' test flakes (and main's red vault-rename check) - #252
Merged
Merged
Conversation
…ds virsh whole Windows EPERM in the unit suite: TestMachine.stop() removed the machine's temp folder while a `git -C <worktree> status`/`log` the daemon's pool had started was still running in it (all 6 failures traced on BEAST). The daemon's stop waits for none of them. fakePoolDeps now starts no git once closing and stop() waits for the git still running. CI did not see it because Node 24.21 retries Windows' access-denied in rmSync (nodejs/node#64698); BEAST's 23.7 does not, and waits 0 ms between tries. Two tests wrote a sandbox's PR into the portal's record and lost it to the daemon's next report (providerLedger "ledger check both ways", devRequests w278): TestMachine.gitLook makes the daemon's git look report it. VM e2e "Domain is already active": dom_state piped virsh domstate into head -n 1; virsh writes its trailing blank line separately, died of SIGPIPE when head had gone, and pipefail made the answer "running\nmissing", so the second install started the running VM. It now reads virsh's whole answer; fff-vm-nightly.test.sh case 10 delays the blank line to prove it. Request: w759 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
# Conflicts: # CHANGELOG.md
w748 (#247) added `fffctl vault rename`; w745 (#244) made server/opsFffctlForms.test.ts fail on any vault form FFFCTL_FORMS does not classify. Both merged, and main's CI failed on Windows and Ubuntu (run 37907217293). rename renames a vault entry, so it is a change: refused to the read-only ops worker like remove and put (fff-ops-priv already refuses any form it does not pin). docs/ops-worker.md lists it with the others. Part of: w759 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Request: w759
TL;DR: Two flakes fixed at their cause, plus a third that also made local Windows runs red. Main had gone red from two merges colliding, so that's fixed here too.
1. Windows EPERM removing a test machine's folder
Error: EPERM, Permission denied: \\?\…\ff-msb-XXXXXXatfs.rmSyncinserver/testMachine.ts:77(TestRepos.cleanup), called fromTestMachine.stop().git -C <tmp>\ffsb\<sandbox> status --porcelain=v2 --branchand/orgit … log -1was still running, started 130–280 ms earlier by the daemon's pool (readGitStatus,machine/sandboxes.ts:705and the othergitStatuscallers).git -Cmakes the worktree its current folder, and Windows refuses to delete a folder a live process is in. The daemon's stop waits for none of these. A plain retry 100 ms later succeeded every time.Found in cache @ …\node\24.21.0). From that release,rmSyncretries Windows' access-denied (std::errc::permission_denied) and actually waits between tries (fs: treatstd::errc::permission_deniedasEPERMerror nodejs/node#64698, backport b269616936). BEAST runs nvm4w Node 23.7.0. Itscan_omit_errorcovers EBUSY/EMFILE/ENFILE/ENOTEMPTY/EPERM but not permission_denied, somaxRetries: 5never retries this error. Its Windows sleep is alsoSleep(i * retryDelay / 1000), which is 0 ms. The race is the same on CI; CI's Node just hides it.fakePoolDepstracks the git/copy/remove it starts and gets aclose(): once closing, it starts nothing new, and it waits for what is still running.TestMachine.stop()calls it before deleting the folder.2. A sandbox PR the daemon's next report took away (4 of 10 runs on BEAST)
server/providerLedger.test.ts"ledger check both ways" (nothing changed, nothing sent 1 !== 0) andserver/devRequests.test.tsw278 (no "dev_update" from the portal in 5000 ms) wrote a sandbox's PR straight into the portal's machine record.mergeSandboxes(server/machines.ts:226) replaces a machine sandbox's git with the daemon's on its next report, so the PR vanished whenever a report landed mid-test. NewTestMachine.gitLook(name, extra): the daemon's own git look reports the PR, and the tests wait for the portal to have it. The board digest (server/intake.ts:85) compares only verdict/watch/ids, so later reports with the same branch and PR change nothing.3. VM e2e "Domain is already active"
error: Domain is already activeright afterfff-vm: 10/12 …in the second ("idempotent") host install. That isinstall.sh:617:[ "$(dom_state)" = running ] || virsh start.dom_statewasv domstate … | head -n 1 || echo missing. In libvirt 12tools/vsh.c,vshPrintVadoesfputs+fflushper print: virsh printsrunning\n, then after the command a separate\n(vshPrintExtra(ctl, "\n")). Whenheadhad already exited, that second write killed virsh with SIGPIPE, pipefail failed the pipeline, and|| echo missingadded a line. The answer wasrunning\nmissing, so the running VM was started again.libvirtd: End of file while reading data: Input/output errorin the same second each time, which is a client that died without closing its connection. In 37540521850 the watch timer ran 3 s earlier and had finished, which rules out a race withfff-vm watch. Why only the 26.04 host is a guess (a differentheador scheduler timing there).dom_statereads virsh's whole answer (s=$(v domstate …) || missing, first line).deploy/vm/test/fff-vm-nightly.test.shcase 10 holds the fake virsh's blank line back 0.3 s. That makes the old pipeline fail every time (measured locally: it printedrunning\nmissing), and the new one printsrunning.[ "$(dom_state)" = running ] || { … nothing to do; return 0; }(fff-vm:269) would skip with no alert on the same misread. This fix covers that path too.4. Main was red (separate commit)
w748 (#247) added
fffctl vault rename, and w745 (#244) fails on any unclassified vault form, so main's CI run 37907217293 failed on Windows and Ubuntu.renameis now achangesform: refused to the read-only ops worker likeremove(fff-ops-priv already refuses any form it does not pin).docs/ops-worker.mdlists it.Evidence
npx tsc -p .clean. shellcheck on the two changed scripts: only the existing SC1091 notes.Known-flake lists: none in this repo or in the final-factory-agents skills names these. The "rerun failed" wording was only in dispatcher briefs.
🤖 Generated with Claude Code