Skip to content

node:fs: stop the async recursive readdir walk at the first error - #40668

Merged
Jarred-Sumner merged 1 commit into
mainfrom
farm/b5b9cacc/readdir-recursive-first-error
Aug 27, 2026
Merged

Jarred-Sumner merged 1 commit into
mainfrom
farm/b5b9cacc/readdir-recursive-first-error

Conversation

@robobun

@robobun robobun commented Aug 27, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • fs.promises.readdir(dir, { recursive: true }) and the callback form never settle on a tree with two directory symlink loops and one entry that fails to open (ln -s . loop1; ln -s . loop2; ln -s bad bad). Every pool thread spins and the caller cannot catch anything. readdirSync rejects the same tree with ELOOP in milliseconds. Both async forms share AsyncReaddirRecursiveTask (node_fs_binding.rs:204).
  • Cause: a failing subtask records its error in pending_err and releases its reference (node_fs.rs:2404), but every other subtask keeps calling enqueue (node_fs.rs:2294), which schedules one new subtask per directory found. Nothing reads pending_err until subtask_count reaches zero, so the promise waits for the whole frontier. The kernel follows 40 symlinks before ELOOP, so two loops make that frontier 2^41 directories.

Fix

  • AsyncReaddirRecursiveTask gets a has_error flag, set right after the first error is recorded. enqueue then schedules nothing, and a subtask that starts after the flag is set releases its reference without opening its directory. The scan settles once the few subtasks already in flight return.
  • The rejection is the same error as before (the first one recorded). On success nothing changes: the flag is never set, so the only new cost is one relaxed atomic load per directory.
  • The subtask_count decrements are now AcqRel, so the last subtask sees the other subtasks' pending_err and queued results. The cp task already uses this ordering.
  • Verified: test/js/node/fs/fs.test.ts (readdir({recursive: true}) settles with the first error while symlink loops are still being walked, promise and callback form; on the stock binary the test times out with the child still spinning, the fixed one settles in 0.6 s under ASAN). Also ran all of fs.test.ts, readdirSync-recursive-error-leak.test.ts, dir.test.ts, fs-path-length.test.ts and node's test-fs-readdir*.js.

Background

  • The async recursive readdir is one Job whose off-thread part is shared by pool subtasks. The root directory runs in run. Each directory or symlink entry becomes a ReaddirSubtask on the work pool. subtask_count is a refcount: it starts at 1 for the root, enqueue adds one, and every subtask subtracts one when it is done. The subtask that brings it to zero calls finish_concurrently, which joins the results or keeps the error and hands the job back to the JS thread.
  • readdir_skips_subdir lists the errors that skip one entry (ENOENT, ENOTDIR, EPERM). Any other error, ELOOP included, fails the whole listing, as in the sync walker.
  • This change does not bound a tree that has only the two loops and no entry that fails early. On a stock kernel the first ELOOP appears at depth 41, so that tree is a 2^41 directory walk for readdirSync, and for node's sync and async walkers too (node v26 did not finish it here either). Only cycle detection (the dev/ino of the ancestors) would bound it. That is a semantics change, and it interacts with node:fs: skip symlink cycles in recursive readdir instead of failing with ELOOP #33432 and node:fs: match Node's recursive readdir symlink and cycle handling #31418, which make ELOOP a skipped error. It is not part of this PR.
Notes

Repro on the released build (Linux 6.17, overlayfs):

mkdir -p /tmp/t/a/b && cd /tmp/t && ln -s bad bad && ln -s . s1 && ln -s . s2
timeout -s KILL 20 bun -e 'require("fs").promises.readdir(".",{recursive:true}).then(a=>console.log("ok",a.length),e=>console.log("rej",e.code))'
# never prints, rc=137. Same with fs.readdir(".", {recursive:true}, cb).
bun -e 'try{require("fs").readdirSync(".",{recursive:true})}catch(e){console.log(e.code)}'
# ELOOP in 9 ms

With this change both async forms print rej ELOOP in about 10 ms (0.5 s for the debug build, most of it startup).

The bare two-loop tree without bad: readdirSync did not finish in 30 s on this machine. The fuzz report that saw it finish in 11 s ran on a box where the kernel returns ELOOP after 8 to 11 components, so the sync walker got its first error early there. The async walker was the only one that did not stop at that error.

USE_SYSTEM_BUN=1 bun test on the new test: the test hits the 5 s runner timeout and the runner kills the spinning child ("killed 1 dangling process"). bun bd test with this change: pass in 583 ms.


no test proof · iteration 0 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/node/fs/fs.test.ts

fs.readdir and fs.promises.readdir with recursive: true run one pool
subtask per directory. A failing subtask recorded its error and released
its reference, but every other subtask kept scheduling subtasks for the
directories it found, and the promise settled only when the whole
frontier had drained. A tree with two directory symlink loops has a
2^41 directory frontier before the kernel reports ELOOP, so the promise
never settled and every pool thread spun.

Set a flag with the first error. enqueue then schedules nothing, and a
subtask that starts after the flag is set releases its reference without
opening its directory. The subtask_count decrements are AcqRel so the
last subtask sees the other subtasks' pending_err and queued results.
@coderabbitai

coderabbitai Bot commented Aug 27, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

  • Run on-demand review

On-demand reviews are free for the next 24 days. After that, they cost $0.25 per reviewed file.

Or wait 14 minutes for your next included review.

View limit details

Limit details: You’ve used all 5 included reviews currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: df7276cd-ed32-4b74-9f95-57a1ca9a07c5

📥 Commits

Reviewing files that changed from the base of the PR and between bedc5c0 and 8def9dd.

📒 Files selected for processing (2)
  • src/runtime/node/node_fs.rs
  • test/js/node/fs/fs.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Aug 27, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: reproduced on the released build (both fs.promises.readdir and callback fs.readdir, recursive: true) with ln -s bad bad; ln -s . loop1; ln -s . loop2. Neither form settles. readdirSync rejects with ELOOP in 9 ms on the same tree.

With this change both async forms reject with ELOOP in about 10 ms. The new test in test/js/node/fs/fs.test.ts times out on the released build and passes with the debug build (583 ms under ASAN).

CI (build 106996): 179 of 181 jobs pass. The new test passes on every lane. The two red tests, test/bundler/transpiler/macro-test.test.ts (LeakSanitizer on x64-asan) and test/js/web/url/url.test.ts (IDNA table on darwin x64), are red on main too and do not touch this diff. The other failures passed on retry. Ready for review.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — focused fix with a proper hang-guard test.

Checked the refcount balance: enqueue's new early return skips both the fetch_add and the schedule, and every path through perform_work (early-exit, error, success→write_results) reaches on_subtask_done() exactly once. Traced the happens-before for the Relaxed has_error flag — whichever thread's fetch_sub(AcqRel) sees 1 synchronizes with the erroring thread's release, so finish_concurrently observes pending_err regardless of which thread finishes last. The success path is untouched (flag never set). Test mirrors the neighboring one at line ~1820 exactly (subprocess timeout + signalCode: null assertion), so the Bun.spawn timeout is a hang guard, not a time-based wait.

Extended reasoning...

Overview

The PR fixes a hang in fs.readdir/fs.promises.readdir with {recursive: true} when the tree contains symlink loops plus an entry that errors. The async walker (AsyncReaddirRecursiveTask in src/runtime/node/node_fs.rs) previously recorded the first error in pending_err but kept fanning out subtasks until the entire frontier drained — with two .-loops that's ~2^41 directories before the kernel returns ELOOP. The fix adds a has_error: AtomicBool that enqueue() checks before scheduling and perform_work() checks at entry, so once any subtask errors the frontier stops growing and the scan settles as soon as in-flight subtasks return. The three subtask_count.fetch_sub sites are factored into on_subtask_done() and the ordering is upgraded from Relaxed to AcqRel. A new test in test/js/node/fs/fs.test.ts builds the pathological tree, spawns both promise and callback forms, and asserts rejected ELOOP / exit 0 / signalCode: null.

Security risks

None. This is an internal work-scheduling short-circuit in the recursive readdir implementation; no user input parsing, no auth, no crypto, no new externally-reachable surface. The change only makes an existing error path settle sooner. The Relaxed ordering on has_error is safe because it guards no data — it's a monotonic hint; a stale false just means one more subtask runs, and the actual pending_err is published via the AcqRel refcount decrement (verified the release-sequence chain: whichever thread hits fetch_sub == 1 acquires against every prior subtask's release, so it observes pending_err and queued results regardless of interleaving).

Level of scrutiny

Medium — concurrent native code with atomics warrants a careful read, which I gave it. The change is small (~30 net lines of Rust), purely additive on the error path (success path never sets the flag, so behavior is unchanged there), and the AcqRel upgrade is strictly safer than the previous Relaxed. Refcount balance holds on every path: enqueue's early return skips both the fetch_add and the schedule (nothing to pair); the three on_subtask_done() call sites in perform_work are mutually exclusive (early-exit returns, error branch and write_results are opposite match arms). The PR description cites the cp task as prior art for the same ordering.

Other factors

The test is placed in the existing module file per convention, uses tempDir/bunExe/bunEnv, covers both sibling entry points (promise + callback), and follows the exact hang-guard pattern of the immediately preceding test in the same file (subprocess timeout: 10_000 with a signalCode: null assertion — this is a Bun.spawn kill deadline, not a bun:test timeout, so it doesn't violate the "no test timeouts" rule and satisfies REVIEW.md's "hang-guard tests assert the process exited on its own"). The PR author reports USE_SYSTEM_BUN=1 times out and bun bd test passes in 583 ms. No CODEOWNERS entry covers either changed file. No outstanding review objections in the timeline. Exit reason was dry_streak.

@Jarred-Sumner
Jarred-Sumner merged commit 09559ed into main Aug 27, 2026
10 of 12 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the farm/b5b9cacc/readdir-recursive-first-error branch August 27, 2026 23:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants