Skip to content

resolver: close the store_fd symlink-target fd when kind() errors - #36880

Open
robobun wants to merge 4 commits into
mainfrom
farm/084891e9/resolver-kind-symlink-fd-leak
Open

robobun wants to merge 4 commits into
mainfrom
farm/084891e9/resolver-kind-symlink-fd-leak

Conversation

@robobun

@robobun robobun commented Aug 4, 2026 •

Copy link
Copy Markdown
Collaborator

Sibling of #36878, same store_fd-gated close pattern, in RealFS::kind() rather than read_directory_with_iterator.

Repro

Under any mode that sets resolver.store_fd (bun --hot, bun --watch), resolve a symlink whose target can be opened but whose /proc/self/fd/N readlink fails afterwards:

LEAKED=16   # one fd per resolved symlink

The readlink failing after a successful open is rare but reachable: a FUSE mount that disappears between the two calls, a /proc that is not mounted, an ioctl-based path on a filesystem that returns EIO. The fstat and FilenameStore.append_slice calls on the same path have their own failure windows.

Cause

The POSIX is_symlink branch of RealFS::kind() opens the symlink target and installs a scopeguard that decides close-vs-store on drop:

let _guard = scopeguard::guard(file, move |file| {
    if (!store_fd || need_to_close_files) && !existing_fd.is_valid() {
        let _ = bun_sys::close(file);
    } else if bun_core::feature_flags::STORE_FILE_DESCRIPTORS {
        unsafe { (*cache_ptr).fd = file };
    }
});
let file_stat = bun_sys::fstat(*_guard)?;
symlink = bun_sys::get_fd_path(*_guard, &mut outpath)?;

With store_fd set and the fd budget unconstrained, the guard writes the fd into cache.fd. cache is a stack-local EntryCache (Copy, no Drop). On the success path cache is returned and the fd reaches the resolver's Entry cache. On the ? paths (fstat, get_fd_path, and the later FilenameStore.append_slice(symlink)?) the function returns Err, cache is dropped, and the fd inside it leaks. Entry::kind swallows the error, so the user-visible resolution still succeeds.

Fix

Track whether the fd has been published to the returned value and close on every exit before that point, the same shape as #36878. The cache.fd write moves out of the guard onto the success path (after every ?), which also removes the raw-pointer write through cache_ptr. The append_slice call moves inside the is_symlink block so the guard covers it too; it only ran for symlinks anyway. A caller-supplied existing_fd is still never closed.

Verification

test/js/bun/resolve/resolver-kind-symlink-fd-leak.test.ts builds an LD_PRELOAD shim that makes readlink("/proc/self/fd/N", ...) fail with EIO when the fd's target contains a marker string, runs bun --hot inner.ts (so resolver.store_fd = true), resolves one symlink per marker target, disarms the shim, and counts fds still pointing at a marker target.

before: FAILED=0 LEAKED=16
after:  FAILED=0 LEAKED=0

test/js/bun/resolve/resolve.test.ts, test/js/bun/resolve/import-meta.test.js, and test/cli/hot/hot.test.ts are unchanged. bun run rust:check-all passes on all 10 targets.


no test proof · iteration 0 · Platform-specific test(s) that do not run on this machine. Deferring to CI, which covers all platforms: test/js/bun/resolve/resolver-kind-symlink-fd-leak.test.ts

RealFS::kind() opens a symlink's target, then fstat()s it and reads the
fd's realpath. The scopeguard that decided close-vs-store wrote the fd
into the stack-local EntryCache whenever store_fd was set and the fd
budget allowed caching. On the success path that EntryCache is returned
and the fd reaches the resolver's Entry cache; on the fstat/get_fd_path
/append_slice error paths the EntryCache is dropped and the fd leaks.

Track whether the fd has been published to the returned value and close
on every exit before that point. The cache.fd write moves onto the
success path so the raw-pointer write through the guard is no longer
needed.
@coderabbitai

coderabbitai Bot commented Aug 4, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 9 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: c07d6f9a-ea09-49fe-9c91-679b9d184c27

📥 Commits

Reviewing files that changed from the base of the PR and between af63b8a and 97fd60e.

📒 Files selected for processing (2)
  • src/resolver/lib.rs
  • test/js/bun/resolve/resolver-kind-symlink-fd-leak.test.ts

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added the claude label Aug 4, 2026
Comment thread src/resolver/lib.rs Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — focused fd-leak fix that closes the symlink-target fd on the ? error paths where the old guard would write it into a stack-local EntryCache that's then dropped.

What was reviewed:

  • Case-analyzed the guard predicate against all four (we_opened_it × close_even_on_success) combinations — success-path behavior is preserved, error paths now close, caller-supplied existing_fd is still never closed.
  • Checked drop order: fd_published/file/we_opened_it are declared before _guard so the non-move closure's borrows outlive it.
  • Confirmed moving append_slice(symlink)? inside the is_symlink block is behavior-preserving (outer default was b"") and now covered by the guard.
  • Test: LD_PRELOAD shim correctly gates on glibc Linux, drains pipes concurrently, disarms before counting, and process.exit(0) runs synchronously before --hot can reload.
Extended reasoning...

Overview

The PR fixes an fd leak in the POSIX is_symlink branch of RealFS::kind() (src/resolver/lib.rs:1449-1491). The original scopeguard's else-arm wrote the opened fd into a stack-local EntryCache on drop; when fstat, get_fd_path, or the later FilenameStore.append_slice returned Err via ?, the function returned early, the stack EntryCache was dropped, and the fd leaked. The new code tracks publication with a Cell<bool>, moves the cache.fd write onto the success path after every ?, and closes on any exit before that point. As a side benefit the unsafe { (*cache_ptr).fd = file } raw-pointer write is gone.

Security risks

None. This is a resource-leak fix on an error path; no new inputs are parsed and no privilege boundary is crossed.

Level of scrutiny

Medium — the resolver is core infrastructure and --hot/--watch exercise this path heavily, but the change is ~30 lines confined to one function's cleanup logic. I walked all four (existing_fd valid? × close_even_on_success?) cases and confirmed the new predicate matches the old success-path behavior exactly while adding the missing close on the error paths. The non-move closure's captured locals (file, we_opened_it, close_even_on_success, fd_published) are all declared before _guard, so reverse drop order keeps them alive when the guard runs. Moving append_slice inside the block is safe because symlink was only ever non-empty inside is_symlink, and it brings that ? under the guard's coverage.

Other factors

The LD_PRELOAD fault-injection test is well-constructed: it skips on musl/non-Linux/no-cc, compiles the shim in beforeAll, spreads bunEnv, drains stdout/stderr/exited concurrently, and asserts a combined {stdout, stderr, exitCode} object. The inner script disarms the shim before counting leaked fds so the count itself isn't sabotaged, and process.exit(0) is synchronous so --hot can't reload on the SHIM_OFF write. The PR description reports before/after LEAKED=16 → 0 and green on resolve.test.ts, import-meta.test.js, hot.test.ts, and rust:check-all. The comment-cop bot's paragraph-comment flag was addressed in 63c6cf6.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't find any bugs. The fd-ownership rewrite checks out on every path I traced (existing_fd never closed; store_fd success path stores instead of closing; all ? exits now close the fd we opened), and it drops the raw-pointer write into the stack-local cache. That said, this is resolver hot-path fd-lifecycle code with a Linux-glibc-only test, so a human should sign off.

What was reviewed:

  • Enumerated all four (we_opened_it × close_even_on_success) combinations against the old guard — success-path behavior is preserved, error paths now close.
  • Confirmed need_to_close_files() already returns true when STORE_FILE_DESCRIPTORS is false, so the store-vs-close branches stay mutually exclusive.
  • Checked drop order: _guard is declared after fd_published/file, so the non-move closure's borrows are live when it runs.
  • Test: LD_PRELOAD shim disarms itself before the leak count and the assertion is on LEAKED=0, not on FAILED.
Extended reasoning...

Overview

The PR reworks the POSIX is_symlink branch of RealFS::kind() in src/resolver/lib.rs. Previously a scopeguard decided close-vs-store on drop and, in the store case, wrote the fd through a raw *mut EntryCache into a stack-local cache. On the ? early-return paths (fstat, get_fd_path, and the later append_slice) that cache was dropped and the fd inside it leaked. The new code tracks we_opened_it, close_even_on_success, and an fd_published Cell; the guard now only closes, and the cache.fd = file write moves onto the straight-line success path after every ?. The append_slice call also moves inside the symlink block so the guard covers it. A new Linux-glibc-only test uses an LD_PRELOAD shim to force readlink(/proc/self/fd/N) to fail and counts leaked fds under --hot.

Security risks

None identified. This is fd-lifecycle bookkeeping with no user-controlled input reaching a new sink. The change actually removes an unsafe raw-pointer write.

Level of scrutiny

High. Module resolution is a hot path used by every import, and fd ownership is exactly the class REVIEW.md flags as most-blocked. I traced each combination:

  • existing_fd valid → we_opened_it=false: guard never closes; cache.fd set iff STORE_FILE_DESCRIPTORS — matches old else if branch.
  • We opened it, store_fd && !need_to_close_files(): guard closes only when !fd_published (i.e., on the ? paths — the fix); on success cache.fd is set and fd_published=true keeps it open — matches old success behavior.
  • We opened it, !store_fd || need_to_close_files(): guard always closes; cache.fd never set — matches old if branch.
  • need_to_close_files() returns true when STORE_FILE_DESCRIPTORS is false (src/resolver/lib.rs:1523), so the STORE_FILE_DESCRIPTORS && !(we_opened_it && close_even_on_success) gate can never leave an fd both un-stored and un-closed.

Drop order is sound: _guard is declared after fd_published and file, so it drops first and the non-move closure's captured references are still live.

Other factors

The test is well-constructed (disarms the shim before counting, asserts a combined {stdout, stderr, exitCode} object, spreads bunEnv, chains onto any existing LD_PRELOAD), but only runs on Linux+glibc with a C compiler — macOS/Windows/musl skip it. The comment-cop bot's note about a long comment was addressed in a follow-up commit (variable names now carry the intent). The pattern mirrors sibling PR #36878. Given the critical path and platform-limited test coverage, deferring to a human reviewer rather than auto-approving.

@robobun

robobun commented Aug 4, 2026 •

Copy link
Copy Markdown
Collaborator Author

CI status on build 88792 (final): every individual test lane passed. The aggregate is red because test/js/node/worker_threads/worker-transfer-terminate-stress.test.ts hit a JSC ExceptionScope::assertNoException() SIGABRT on one debian-13-x64-asan shard; every other failure passed on retry. This diff touches only the Rust resolver's POSIX symlink branch, no JSC or worker code, and the same assertion fires on unrelated PR build 88637.

The new test (resolver-kind-symlink-fd-leak.test.ts) passed on every glibc Linux lane and is skipped elsewhere. Ready for a maintainer to merge.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant