Skip to content

Bun.Archive: extract() with a glob rejects at an entry that a system call refuses; files(glob) matches normalized names - #44731

Open
robobun wants to merge 1 commit into
mainfrom
robobun/a178f38c/archive-glob-entry-errors
Open

robobun wants to merge 1 commit into
mainfrom
robobun/a178f38c/archive-glob-entry-errors

Conversation

@robobun

@robobun robobun commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator

Related to #43132

Problem

  • archive.extract(dir, { glob }) resolves although entries were not written. extract_to_disk_filtered (src/runtime/api/Archive.rs:1377 to :1479) did continue after a failed mkdir, open, write or symlink. extract(dir) rejects.
  • archive.files(glob) matched the stored name, extract(dir, { glob }) the normalized one. On a tar -C dir . tarball, files("src/*") returns nothing.

Fix

  • With a glob, a refused directory or file stops the extraction and rejects with that error (code, syscall, path). A refused symlink rejects at the end.
  • An EACCES at open replaces the file, so a second extract over a read-only file works.
  • match_glob_patterns takes a NormalizedName. A leading ./ of a pattern is dropped.
  • Verified: test/js/bun/archive.test.ts (20 new tests, 12 fail on main). Self-reviewed: 15 concerns raised, 9 addressed, 1 in part, 5 held.

Background

  • A glob call runs extract_to_disk_filtered. Other calls run Archiver::extract_to_dir, which bun install shares.
  • The normalized name is the path that extract() writes, without ./ and //.
  • Weighed: open and write arms only (two arms keep the defect), and one filtered extractor (it changes 21 of 29 probed glob results).

Downsides

  • extract(dir, { glob }) now rejects where it resolved with a smaller count. extract(dir) still resolves at a directory or symlink that it cannot create (Notes).
  • With a glob, a read-only file in the destination is replaced, not kept.
  • Cost: files(glob) +327 instructions per entry (+6.6%), an extract() job 152 to 160 bytes, release .text +2,816 bytes.
Notes

Source of the report. An automated check found both defects. No user reported them. #43132 lists the glob skip in one row of its table. This PR changes that row: with a glob, the read-only a.txt now gets the new content.

Three questions for a maintainer.

  1. The held part. extract(dir) without a glob has the same defect for a directory entry below the top level and for every symlink entry: it resolves and counts the entry (src/libarchive/lib.rs:1628, :1635, :1136). The fix is a policy in ExtractOptions that only Bun.Archive sets, because bun install, bun create and the --compile download share that extractor. It exists as commit 6aba634052 on branch robobun/a178f38c/archive-entry-errors. It is not in this PR because it changes extract(dir) for callers that did not ask, and because it uses the per-opener policy that Bun.Archive: reject when a header read fails, with one in-memory tar reader and a named policy #43242 still asks about. Should it follow as its own PR?
  2. The interim. Until then the two call forms differ the other way on a directory or symlink entry that cannot be created: extract(dir) resolves, extract(dir, { glob }) rejects. If that is not acceptable, I move the directory and symlink arms of the glob form into the held part. Then the forms differ on 2 of 30 probed rows (by a code walk, not a build).
  3. The end state. Is it one extractor with a filter hook in Archiver::extract_to_dir, and extract_to_disk_filtered deleted? The design review set it aside for now: it changes 21 of 29 probed glob results (count, file mode, symlink rules), and its write path (read_data_into_fd) resolves after a refused write. This question does not gate this PR.

Settlement per failure. Release builds of main (bd599f5af9) and of this branch, Linux x64. 14 rows run as uid 65534 (chmod cases), 16 rows as uid 0 (cases that do not depend on the user).

main this PR
reporter script 1 (glob swallows write failures) exit 86 exit 0
reporter script 2 (files(glob) against extract({ glob })) exit 86 exit 0
rows where the glob form rejects 0 of 30 22 of 30
rows where extract(dir) rejects and the glob form resolves 10 1 (it replaces a read-only file)
rows where extract(dir) resolves and the glob form rejects 0 13

The 8 rows where the glob form still resolves: 2 controls, 1 read-only file that it now replaces, 3 symlinks whose name is taken (EEXIST, the name keeps what it has), and 2 directory entries whose name is taken by a regular file (see "Not fixed here").

What the glob form rejects with, as uid 65534:

failure rejection
destination not writable, top-level file EACCES, open, <dest>/a.txt
destination not writable, nested file first EACCES, mkdir, <dest>/sub
directory entry in a directory that is not writable EACCES, mkdir, <dest>/sub/empty
a file has the name of a parent directory ENOTDIR, open, <dest>/blocker/x.txt
a directory has the name of a file entry EISDIR, open, <dest>/taken
name of 300 bytes ENAMETOOLONG, open, <dest>/<name>
write past a file size limit EFBIG, write, <dest>/big.txt (the partial file is removed)
symlink in a directory that is not writable EACCES, symlinkat, <dest>/sub/l (the other entries are written)

path is the destination as the caller passed it, joined with the name that the failed call was given. When the directory of an entry cannot be made, the rejection is that mkdir error, not the ENOENT of the open after it. extract(dir) still rejects with Error("ReadError") and no code (#44509).

Measurements. Release builds of main (bd599f5af9) and of this branch (158f2fafcc), Linux x64. perf, valgrind and strace are not installed here, so the counts come from ptrace tools (a syscall counter and an int3 plus single-step instruction counter), objdump, nm and size.

  • Job allocation of extract(): mov $0x98,%edi before call mi_malloc on main, $0xa0 here (152 to 160 bytes, both in the 160-byte class). ExtractResult 8 to 16 bytes. A const assert keeps it at 16.
  • Successful extract(dir, { glob: "**" }), output syscalls per 1,000 entries (N = 1,000 against 2,000): top-level files openat 1,000, write 1,000, close 1,000. Nested files add mkdirat 1,000. Directories mkdirat 1,000. Symlinks symlinkat 1,000. Identical on both builds. A fixed archive of 6 entries: 95 file syscalls on both.
  • Allocator calls per entry inside the extract job: 4 (top-level file), 5 (nested file), 4 (directory), 6 (symlink). Identical on both builds.
  • Instructions per entry inside the extract job, glob form, six archive shapes: -2 to -67 against main (5,353 to 5,351 for a top-level file, 10,144 to 10,089 for a nested file). The counter repeats within about 0.1%. The path buffer is now taken once per call and not once per entry.
  • files("**"): 4,986 to 5,313 instructions per regular entry (+327, +6.6%) for 10-byte names. That is the one normalize pass. files() without a glob: 4,574 to 4,555 (no change within the spread).
  • bun install: no line of bun_libarchive or bun_install changes. extract_to_dir::<()> 6,832 bytes, create_deferred_symlinks 2,411, ExtractTarball::run 10,487, drain_callback 12,075 on both builds.
  • Release .text: 65,453,397 to 65,456,213 bytes (+2,816). The file grows by 4,096 bytes. ExtractContext job function 4,359 to 5,498 bytes, FilesContext job function 2,380 to 2,982.

Behavior changes to know.

  • extract(dir, { glob }) stops at a refused directory or file, as extract(dir) does. The entries after it are not written. Main wrote them.
  • A refused symlink does not stop the other entries. The promise rejects after the last entry with the first refused symlink. On a file system without symlinks, every regular file is still written.
  • A file that open refuses with EACCES is unlinked and created again. Main skipped the entry and kept the old content. This covers an archive that holds a read-only file twice (tar -rf), and a second extract over read-only files of the first. extract(dir) creates files writable, so it did not have the case.
  • On Windows the glob form no longer creates a file read-only when its mode lacks the owner-write bit.
  • files("./src/*") and extract(dir, { glob: "./src/*" }) now mean src/*. On main the first matched only stored ./ names, the second matched nothing.
  • files(glob) still lists an entry that extract() leaves out for its name (second byte :). A test pins it.
  • The docs recipe "validate before extraction" now checks with patterns. The old string checks rejected every entry of a tar -C dir . archive, and passed src/x\.env, which extract() writes as the hidden file src/x/.env.

Not fixed here.

Other open PRs on this function. All are mine. None converts these arms.

Self-review. A review of the first shape (three commits on top of #43242) asked for this shape: a PR against main without the extract(dir) half. It raised 15 concerns.

  • Addressed (9): a read-only file no longer fails a second extract. A refused symlink no longer cuts the tree. files(glob) keeps an entry with a long name. Two tests for the error of a parent directory. Per-platform errno values in place of a wildcard, and a reason at each skipped test. The wording of the keys of files(). The comment and the test name of NormalizedName, and a test for the : name. The docs recipe. One-line comments.
  • In part (1): the mkdir livelock. This PR links paths: stop the recursive mkdir walk when a parent exists but its child is still not found #40600 and does not copy its fix. A test here would hang until that fix lands.
  • For the held part (5): directory modes at mkdir time, the macOS errno of a trailing slash, the directory-over-a-file case, and two missing tests.

Tests run. Debug build (ASAN): test/js/bun/archive.test.ts 127 pass, 2 skip as root, 128 pass, 1 skip as uid 65534. Release build of main: 12 of the 20 new tests fail as root, 15 as uid 65534. The 5 that pass on both pin what must not change (an existing directory, a symlink whose name is taken, an entry that the glob leaves out, the : name, a long name). bun run rust:check-all: 12 targets ok. The Windows lines are type-checked only. I cannot run Windows here, so the Windows values of two tests (EISDIR, ENOENT) are a first pin for CI.


no test proof · iteration 0 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/bun/archive.test.ts

…call refuses

extract(dir, { glob }) went to the next entry when mkdir, open, write or
symlink failed for an entry, and resolved with a smaller count. Now a
refused directory or file stops the extraction, and the promise rejects
with the error of the failed call: code, syscall and path. A refused
symlink does not stop the other entries and rejects after the last one.
A file that cannot be opened for a write is replaced, so a later entry
or a second extract can follow a read-only file. On Windows a file is
no longer created read-only.

files(glob) gave the stored name of an entry to the glob, and
extract(dir, { glob }) the normalized name. match_glob_patterns now
takes a NormalizedName, which only the normalization constructs. A
leading "./" of a pattern is dropped, so both spellings select the same
entries. The keys of the Map stay the entry names.
@robobun

robobun commented Oct 8, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 6:50 PM PT - Oct 7th, 2026

❌ @robobun, your commit 158f2fa has 4 failures in Build #123805 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 44731

That installs a local version of the PR into your bun-44731 executable, so you can run:

bun-44731 --bun

@robobun

robobun commented Oct 8, 2026

Copy link
Copy Markdown
Collaborator Author

Status: ready for review. Three questions for a maintainer are at the top of the Notes in the PR body.

How I reproduced it. Linux x64, release build of main (bd599f5af9), as a user that is not root, because root ignores directory permissions.

  • extract(dir, { glob: "**" }) into a destination with mode 0555, into one whose sub directory has mode 0555, into one where a directory has the name of an entry, and under a file size limit (ulimit -f 8). Main resolves with 0, 3, 3 and 3, and files are missing. extract(dir) rejects in all four cases. This branch rejects with EACCES, EACCES, EISDIR and EFBIG.
  • files(glob) against extract(dir, { glob }) on a tarball from tar -cf x.tar -C dir .. On main the two differ for 5 of the 6 patterns of the docs. On this branch they agree for all 6.
  • The cases that do not depend on the user (a file where a directory must be, a name of 300 bytes, a write past a file size limit) are tests in test/js/bun/archive.test.ts. 12 of the 20 new tests fail on main as root, and 15 as another user.

@coderabbitai

coderabbitai Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: oven-sh/bun/.coderabbit.yaml
  • Review profile: ASSERTIVE
  • Plan: Essentials
  • Run ID: 4bccddc8-b68a-4beb-b126-17f77459bb8b
📥 Commits

Reviewing files that changed from the base of the PR and between 5749c31 and 158f2fa.

📒 Files selected for processing (5)
  • docs/runtime/archive.mdx
  • packages/bun-types/bun.d.ts
  • src/runtime/api/Archive.rs
  • src/runtime/error.rs
  • test/js/bun/archive.test.ts
💤 Files with no reviewable changes (1)
  • src/runtime/error.rs

Included review availability: This review used your included allowance. Your plan provides up to 10 included reviews per hour; 9 remain after this review.


Walkthrough

Archive glob matching now uses normalized extraction-style paths while files() retains raw UTF-8 entry names. Filtered extraction reports filesystem failures with syscall details and entry paths. Documentation and tests cover matching and extraction behavior.

Changes

Archive glob and extraction behavior

Layer / File(s) Summary
Normalize archive names for glob matching
src/runtime/api/Archive.rs, packages/bun-types/bun.d.ts, docs/runtime/archive.mdx, test/js/bun/archive.test.ts
Glob patterns ignore repeated leading ./. Filtered listing and extraction match normalized entry names, while files() map keys remain unnormalized. Tests cover path separators, exclusions, and long or unsafe names.
Propagate extraction filesystem failures
src/runtime/api/Archive.rs, src/runtime/error.rs, test/js/bun/archive.test.ts, docs/runtime/archive.mdx, packages/bun-types/bun.d.ts
Filtered extraction returns path-qualified system errors for filesystem failures. It handles file, directory, and symlink cases, removes partial files after write failures, and exposes system-error details to JavaScript. Tests cover failures, retries, and repeated extraction.

Suggested reviewers: jarred-sumner

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to 158f2

The glob matching and extraction-error changes appear mergeable after normal checks; no actionable risk remains from this review.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies both primary changes: glob extraction now rejects refused system calls, and files(glob) matches normalized names. It is longer than preferred but remains specific and rele…
Description check ✅ Passed The description explains the problem, implementation, behavior changes, limitations, verification results, and test coverage. It does not use the template headings exactly, but it provides the require…
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings marked 🟡 are optional suggestions and need no follow-up push.

Additional findings (outside the current diff — GitHub can't attach inline comments there):

  • 🟣 src/runtime/api/Archive.rs — Pre-existing: users who extract a truncated or corrupt archive with a glob still get a resolved promise with a smaller count, where extract(dir) rejects. At Archive.rs:1488-1491 a failed archive.read_data only sets write_success = false, so the partial file is unlinked and the loop moves on with no rejection. At Archive.rs:1397 a Fatal read_next_header just ends the while loop and Success(count) is returned. The docs and bun.d.ts added here say the glob form rejects when an entry cannot be written. Fix: map a negative read_data result and a Failed/Fatal read_next_header in the glob loop to ExtractResult::Err(ExtractError::ReadError), the same outcome extract_to_dir gives at src/libarchive/lib.rs:1439-1441 and :1887.

    Why this was flagged

    An archive whose body is cut short is passed to archive.extract(dest, { glob: '**' }). For an entry whose data is missing, archive.read_data at Archive.rs:1487 returns a negative libarchive status; Archive.rs:1488-1491 sets write_success = false and breaks, Archive.rs:1524 unlinks the partial file, rejection is still None so no return happens, and the loop continues to the next header. When the next header is itself unreadable, read_next_header at Archive.rs:1397 returns Fatal, succeeded() is false, the loop ends and Archive.rs:1573 returns ExtractResult::Success(count). The caller gets a resolved promise with a count lower than the archive's entry count and no error. The base branch behaves the same way at these two sites, so this is pre-existing, but the non-glob path extract_to_dir returns Err(crate::Error::Fail) for both a Fatal header at src/libarchive/lib.rs:1439-1441 and a failed data read at src/libarchive/lib.rs:1887, so extract(dest) rejects for the same input. The new rejection variable at Archive.rs:1481 only covers bun_sys::write outcomes, not the libarchive read result.

    Verification: Pre-existing: the base already continued on a failed read_data and returned Ok(count). Trigger: archive.extract(dest, { glob }) on a tarball truncated mid-entry. src/runtime/api/Archive.rs:1488-1491 sets write_success = false and no rejection, so the loop continues; line 1397 ends on Fatal and line 1573 returns Success(count). src/libarchive/lib.rs:1439-1441 and 1887 return Err(crate::Error::Fail).

Comment on lines +1461 to +1466
if let Err(err) = &opened
&& err.get_errno() == bun_sys::E::EACCES
&& bun_sys::unlinkat(dir_fd, pathname_z).is_ok()
{
opened = bun_sys::openat(dir_fd, pathname_z, flags, mode);
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Users who run extract() with a glob into a directory holding a read-only file lose that file: it is deleted and overwritten, where extract(dir) rejects with EACCES. Archive.rs:1461-1466 answers any EACCES from openat with unlinkat and a second openat, so every pre-existing file without the owner-write bit under the destination is replaced, not only an entry this call wrote. The PR text lists this as a downside; the consequence is silent data loss with no rejection. Fix: make a second extract and a duplicate read-only entry work without deleting foreign files, e.g. create files with the owner-write bit as extract_to_dir does at src/libarchive/lib.rs:1685 (| 0o666) instead of the unlink-and-retry, or limit the unlink to a name this same call already created.

Why this was flagged

A destination directory that is writable contains a file r.txt with mode 0o444 that the user or another user placed there, and the archive holds an entry r.txt; the call comes through archive.extract(dest, { glob }) into extract_to_disk_filtered (src/runtime/api/Archive.rs:1351). openat with O_WRONLY|O_CREAT|O_TRUNC at Archive.rs:1459 fails with EACCES; Archive.rs:1461-1466 then calls bun_sys::unlinkat on the name and opens again, which deletes the existing read-only file and writes the archive content in its place, and the promise resolves with the entry counted. The base branch did Err(_) => continue at this site, so the read-only file kept its content and the entry was skipped; the non-glob path extract_to_dir at src/libarchive/lib.rs:1718-1732 retries once after make_path and then returns the EACCES error, so extract(dir) rejects. The unlink is not limited to files this call created, it applies to any file that exists before the call.

Verification: src/runtime/api/Archive.rs:1459-1466: openat fails with EACCES, and the new arm deletes whatever file is at that name and recreates it with the archive's content; nothing checks that this call created the file. src/libarchive/lib.rs:1718-1733 propagates the EACCES with ?, so extract(dest) rejects. On the base branch the glob path continued, keeping the file.

Comment on lines +1569 to +1571
#[cfg(unix)]
if let Some(rejection) = refused_symlink {
return rejection;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) Users extracting with a glob onto a filesystem that cannot hold symlinks (FAT/exFAT, many CIFS and Docker Desktop bind mounts, drvfs) now get a rejection where main resolved and extract(dir) still resolves. At src/runtime/api/Archive.rs:1540-1547 an EPERM from symlinkat is retried once and then stored in refused_symlink at :1554; after every regular file has been written, :1570 returns it and the count is discarded. Fix: with a glob, treat a symlinkat refusal the way the unfiltered extractor and main do for a symlink entry, or apply one policy to both call forms; at minimum keep the Success count when the only refusals are symlinks.

Why this was flagged

Trigger: an archive with at least one symlink entry extracted via archive.extract(dir, { glob }) onto a mount whose symlinkat returns EPERM or EACCES. At src/runtime/api/Archive.rs:1540 symlinkat fails with EPERM; :1543-1547 runs make_path on the parent and retries, failing again; :1554 stores entry_error in refused_symlink. The loop keeps writing files, and at :1569-1571 the function returns the stored rejection instead of ExtractResult::Success(count). On the base commit the same arm did continue and the promise resolved with the file count; the unfiltered extract(dir) path (libarchive extract_to_dir) is unchanged in this PR and still resolves on a symlink failure. The population (CI containers on Windows hosts, USB/exFAT, SMB shares) sees every glob extract of such an archive flip from resolve to reject, with no way to learn how many entries were written.

Verification: At src/runtime/api/Archive.rs:1540 symlinkat fails; :1549-1561 only EEXIST is benign, every other errno is stored in refused_symlink at :1554; the loop keeps writing files, then :1569-1572 returns it, discarding count. On base (bd599f5a) the same path did continue, so the promise resolved with the symlink skipped. extract(dir) still resolves: src/libarchive/lib.rs:1136-1147 matches Err(_) with a warn and continues.

Comment on lines +1470 to +1474
return match parent {
// The open fails because this mkdir failed.
Some((parent_dir, Err(mkdir_err))) => {
ExtractResult::entry_error(&mkdir_err, root, parent_dir)
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) Users whose glob extract fails at a nested mkdir get an error path naming a directory that was never attempted, not the one the kernel refused. At src/runtime/api/Archive.rs:1473 (and :1436, :1557) entry_error overwrites the path that make_path already put in the error with the whole entry or parent name, but make_path walks up on ENOENT and the refused mkdirat is an ancestor component. Fix: build the rejection from the path the mkdir error carries (mkdir_err.path, joined with root) at the three make_path sites, falling back to the entry name only when that path is empty. Example: dest is 0o555 and the archive holds sub/deep/x.txt; the error says code: EACCES, syscall: mkdir, path: <dest>/sub/deep while the call that failed was mkdir <dest>/sub.

Why this was flagged

A user calls archive.extract(dest, { glob: "**" }) where dest is not writable (mode 0o555) and the archive has the file entry sub/deep/x.txt with no directory entries. mkdir_recursive_at_mode (src/sys/lib.rs:2407-2421) tries mkdirat(dir, "sub/deep"), gets ENOENT, and make_path_with steps to the previous component; mkdirat(dir, "sub") fails with EACCES and check_p! stores path = "sub" in the returned error. Archive.rs:1472-1474 calls ExtractResult::entry_error(&mkdir_err, root, parent_dir) with parent_dir = "sub/deep", and entry_error (Archive.rs:736-741) calls err.with_path(&path), discarding the "sub" the error carried. The promise rejects with path: <dest>/sub/deep, although the refused call was for <dest>/sub. The same override happens at Archive.rs:1436 for a directory entry and at Archive.rs:1557 for a symlink parent. The base branch swallowed these failures entirely, so this is the first time the path is surfaced; the PR's own tests only cover one-level parents, where the last component is also the refused one, so the suite does not catch the ancestor case.

Verification: mkdir_recursive_at_mode (src/sys/lib.rs:2402-2423) walks mkdirat(dir, "sub/deep") -> ENOENT -> mkdirat(dir, "sub") -> EACCES, so mkdir_err.path == "sub". src/runtime/api/Archive.rs:1472-1473 does ExtractResult::entry_error(&mkdir_err, root, parent_dir) with parent_dir = "sub/deep", and entry_error calls err.with_path(&path), which replaces path wholesale.

Comment on lines +1410 to +1412
let Some(name) = NormalizedName::new(raw_pathname, &mut normalized_buf[..]) else {
continue;
}
let pathname_z: &bun_core::ZStr = bun_paths::resolve_path::normalize_buf_z::<
bun_paths::platform::Posix,
>(raw_pathname, &mut normalized_buf[..]);
};

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟣 pre-existing, not blocking: Pre-existing: users extracting with a glob still get a resolved promise, with a smaller count, when an entry name is too long for the path buffer, where the PR's purpose is to reject at every entry that was not written. At Archive.rs:1410-1412 NormalizedName::new returns None for a stored name of MAX_PATH_BYTES or more and the loop does continue, so the entry is silently dropped with no rejection and no count. Fix: treat a name that does not fit as a refused entry and return an ENAMETOOLONG rejection (code, syscall, path) like the other arms, so no entry is skipped without an error; the files(glob) path keeps spilling to a Vec since it writes nothing.
A small fix can ride a push you are already making; otherwise a short reply is enough.

Why this was flagged

An archive with a pax-extended name of 4096 or more bytes (Linux MAX_PATH_BYTES; the test at test/js/bun/archive.test.ts 'files(glob) lists an entry with a name longer than a path' builds such a 10,000-byte name) is passed to archive.extract(dest, { glob: '**' }), reaching extract_to_disk_filtered. At Archive.rs:1410 NormalizedName::new(raw_pathname, &mut normalized_buf[..]) returns None because stored.len() >= buf.len() (Archive.rs:1294-1296), and Archive.rs:1411 does continue. The entry is never written, never counted, and no error is produced; the promise resolves with the count of the other entries. The base branch did the same continue at the old raw_pathname.len() >= normalized_buf.len() check, so this is pre-existing, but it is the same class of swallowed failure the PR rewrites this function to remove, and a 300-byte name now rejects with ENAMETOOLONG while a 4096-byte name still vanishes silently. No safeguard reports it: the Directory, File and SymLink arms are never reached for that entry.

Verification: Pre-existing. Triggered when an archive entry's stored name is >= MAX_PATH_BYTES and the user calls extract(dest, { glob }). At src/runtime/api/Archive.rs:1410-1412 NormalizedName::new returns None and the loop does continue, so the entry is neither written, counted, nor turned into a rejection, and the promise resolves with a smaller count. The base branch had the identical skip (base Archive.rs:1347).

Comment on lines +1435 to 1438
if let Err(err) = dir_fd.make_path(pathname) {
return ExtractResult::entry_error(&err, root, pathname);
}
count += 1;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟣 pre-existing, not blocking: Pre-existing: users extracting with a glob still get a resolved promise, with the entry counted, when a directory entry's name is taken by a regular file in the destination. At Archive.rs:1435 dir_fd.make_path maps mkdir's EEXIST to Ok without checking that the existing node is a directory (src/sys/lib.rs:2418), so the arm counts it at :1438 and nothing is created. Fix: after make_path on a directory entry, verify the node at pathname is a directory (fstatat) and otherwise reject through entry_error with the mkdir EEXIST or ENOTDIR, while still tolerating an existing directory, which the 'directory entry whose directory exists' test relies on.
A small fix can ride a push you are already making; otherwise a short reply is enough.

Why this was flagged

The destination holds a regular file named d (left by a previous extract or by the user) and the archive holds a directory entry d/ with no file entries below it, passed to archive.extract(dest, { glob: '**' }) which runs extract_to_disk_filtered. At src/runtime/api/Archive.rs:1435 dir_fd.make_path("d") calls mkdir_recursive_at_mode (src/sys/lib.rs:2402); mkdirat returns EEXIST for the file and src/sys/lib.rs:2418 maps it to MakePathStep::Exists, so make_path_with (src/paths/component_iterator.rs:212-216) returns Ok(()). The arm then does count += 1 at Archive.rs:1438 and the promise resolves with a count that includes a directory that does not exist. The base branch behaves the same (its EEXIST arm at the old :1375 was dead because make_path never surfaces EEXIST), so this is pre-existing; but the PR's stated purpose is that a refused mkdir rejects, and the PR notes list this row as not fixed. A file entry below d would reject with ENOTDIR at open, but a directory-only entry is silent.

Verification: pre-existing. Trigger: the destination holds a regular file d and the archive holds a directory entry d/, extracted with a glob. mkdir_recursive_at_mode maps EEXIST to Ok at src/sys/lib.rs:2418 without checking the node type, so at src/runtime/api/Archive.rs:1435-1438 count += 1 runs and the promise resolves while d remains a regular file. On the base the same route applied.

Comment on lines +1549 to +1552
match created {
Ok(()) => count += 1,
// The name is taken. It keeps what it has, and the entry is not counted.
Err(err) if err.get_errno() == bun_sys::E::EEXIST => {}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟣 pre-existing, not blocking: Users re-extracting with a glob after an archive changed a name from a regular file to a symlink keep the stale file and get a resolved promise, while a file entry whose name is taken is replaced. At src/runtime/api/Archive.rs:1552 an EEXIST from symlinkat is swallowed and the entry is neither counted nor reported, so the symlink's target is never applied. Fix: handle a taken symlink name the way the file arm does at :1455-1462 (unlink and recreate, or at least surface the EEXIST through refused_symlink) so every taken-name case on the glob path has one behaviour.
A small fix can ride a push you are already making; otherwise a short reply is enough.

Why this was flagged

Trigger: extract(dir, { glob }) where dir already holds a regular file, directory, or older symlink at a symlink entry's name; entry point ExtractContext::do_run -> extract_to_disk_filtered. At src/runtime/api/Archive.rs:1540 symlinkat returns EEXIST; :1543 does not retry (EEXIST is not EPERM/ENOENT); :1552 matches EEXIST and does nothing, so count is not incremented and refused_symlink is untouched; :1573 returns Success. The file arm in the same PR replaces a taken name: O_TRUNC at :1453 and the EACCES unlink-and-reopen at :1455-1462. The base commit also continued on EEXIST, but the PR's stated contract is that a refused syscall rejects, and it rewrote this arm while keeping the one silent exception. A second extract over an existing tree (the exact scenario the PR added the EACCES replacement for) therefore leaves a stale regular file where the archive now has a symlink, with no signal. Remedy: unlink and recreate on EEXIST, or report it through refused_symlink.

Verification: Pre-existing, acknowledged in diff. Trigger: extract(dir, { glob }) into a destination that already has a regular file at a symlink entry's name. In the src/runtime/api/Archive.rs symlink arm EEXIST neither counts nor sets refused_symlink, so the stale file is left in place and the symlink's target never applied. The base's filtered path did exactly the same.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants