Skip to content

drop unused libarchive read formats and filters - #39484

Merged
Jarred-Sumner merged 1 commit into
mainfrom
ali/size-libarchive-formats
Aug 18, 2026
Merged

Jarred-Sumner merged 1 commit into
mainfrom
ali/size-libarchive-formats

Conversation

@alii

@alii alii commented Aug 18, 2026 •

Copy link
Copy Markdown
Member

Problem

  • bun only ever registers the tar reader (archive_read_support_format_tar / _gnutar) and the gzip read filter (TarballStream.rs, runtime/api/Archive.rs, cli/pack_command.rs, libarchive/lib.rs), but every libarchive reader and read filter still ends up in the binary.
  • The reason is TarballStream::open_archive (src/install/TarballStream.rs), which calls archive_read_set_format and archive_read_append_filter to skip bidding. Upstream, those two functions register the requested format/filter themselves, through archive_read_support_format_by_code() and a switch over every archive_read_support_filter_*(), so one reference from bun keeps the 7zip, rar, zip, iso9660, xz, zstd, ... readers alive through --gc-sections.

Fix

  • patches/libarchive/select-registered-only.patch: archive_read_set_format and archive_read_append_filter no longer register anything; they only select among the formats/filters the caller already registered, and fail with "Format is not registered" / "Filter is not registered" otherwise. The slot lookup loops are upstream's, unchanged.
  • scripts/build/deps/libarchive.ts: drop the reader and read filter sources bun never registers, plus archive_ppmd8 and the blake2 reference implementations that only the zip and rar5 readers used. format_empty, format_raw, ppmd7 and filter_program stay because archive_match.c, archive_write_set_format_7zip.c and archive_read_append_filter.c still reference them by name.
  • TarballStream::open_archive now registers gzip itself before archive_read_append_filter (it already registered tar before archive_read_set_format). The comment there explains the ordering constraint: tar has to be registered before read_set_options, and set_format has to come after it, because archive_set_format_option() writes NULL to a->format when it is done dispatching.
  • Stripped binary size, measured by CI against the build of this branch's merge-base (ff5fc1f, build 100261): linux-x64 -272 KB, linux-aarch64 -320 KB, darwin-aarch64 -259 KB, darwin-x64 -278 KB, windows-x64 -288 KB, windows-aarch64 -292 KB, musl/android/freebsd between -258 KB and -310 KB. The CI size annotation shows a larger number because it compares against a newer main that grew in between.
  • Verified on a debug build of this branch: nm on the binary shows only the tar/gnutar/gzip registration functions plus the three kept-by-reference stubs; test/cli/install/bun-install-streaming-extract.test.ts (the only path that uses the patched functions; with the read_support_filter_gzip() line removed, its streaming case fails with "Fail extracting tarball", so the existing suite covers the new ordering requirement), bun-install-tarball-integrity.test.ts, bun-pack.test.ts, bun-add.test.ts (includes an uncompressed .tar), and test/js/bun/archive.test.ts all pass.

Background

  • libarchive reads an archive through a chain of "read filters" (decompressors such as gzip) feeding a "format" reader (tar, zip, ...). A program registers the candidates it wants with archive_read_support_*(), and on open libarchive "bids": each registered candidate inspects the first bytes and the best match wins.
  • Bidding needs read-ahead, which the streaming bun install extractor cannot provide before the first HTTP chunk arrives, so TarballStream uses archive_read_append_filter(GZIP) and archive_read_set_format(TAR) to pin the chain up front instead. Those are the only two libarchive entry points bun uses that pick a reader by code, and they are what this patch changes.
  • libarchive is built as part of bun's own ninja graph (scripts/build/deps/libarchive.ts lists the source files directly), so removing a file from SOURCES removes it from the link; the patch is what makes that possible without undefined references.

@coderabbitai

coderabbitai Bot commented Aug 18, 2026

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 49 minutes

Limit details: You’ve used all 1 included review currently available under your plan. You completed 74 included PR reviews in the past 7 days; at that activity level, included reviews refill at 1 review per hour.

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: ff011947-e706-4f2d-bf98-4e3a463001c4

📥 Commits

Reviewing files that changed from the base of the PR and between 258517a and 53d5224.

📒 Files selected for processing (3)
  • patches/libarchive/select-registered-only.patch
  • scripts/build/deps/libarchive.ts
  • src/install/TarballStream.rs

Comment @coderabbitai help to get the list of available commands.

@alii

alii commented Aug 18, 2026

Copy link
Copy Markdown
Member Author

@robobun adopt

@robobun

robobun commented Aug 18, 2026 •

Copy link
Copy Markdown
Collaborator

Adopted and verified: debug build of this branch links with only the tar/gnutar/gzip readers left, and the streaming extract, tarball integrity, pack, add and Bun.Archive suites pass (removing the new gzip registration makes the streaming test fail, so the ordering is covered). CI was green (https://buildkite.com/bun/bun/builds/100323); against this branch's own merge-base the stripped binaries shrink by 258 to 320 KB per platform, details in the PR description.
Merged as 2219e22.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it patches vendored libarchive semantics and prunes the compiled source list for the bun install extraction path, a human look would still be worthwhile.

What was reviewed:

  • Confirmed TarballStream is the only caller of archive_read_set_format/archive_read_append_filter, and it now registers gzip/tar before selecting them; other libarchive read users (Archive.rs, pack_command.rs, libarchive/lib.rs) already register tar/gnutar/gzip explicitly and don't use the patched entry points.
  • Traced the patched archive_read_set_format.c slot-lookup loop — a->format is set to &a->formats[0] before the name check, so removing the auto-register call cannot introduce a null deref; an unregistered format cleanly returns ARCHIVE_FATAL.
  • Checked the removed sources (ppmd8, blake2, filter/format _all/_by_code, etc.) against the remaining files' references; the kept-but-unused entries (format_empty/format_raw/ppmd7/filter_program) match what the surviving objects still name.
Extended reasoning...

Overview

This PR shrinks the linked libarchive surface by (a) adding patches/libarchive/select-registered-only.patch, which strips the auto-registration calls out of archive_read_set_format and archive_read_append_filter so they only pick among formats/filters the caller already registered, (b) dropping ~20 read-format/filter source files plus ppmd8/blake2*_ref from the DirectBuild source list in scripts/build/deps/libarchive.ts, and (c) adding an explicit read_support_filter_gzip() call in TarballStream::open_archive before archive_read_append_filter(GZIP) to satisfy the new select-only semantics. The Rust change is one functional line plus a consolidated comment.

Security risks

Archive extraction during bun install is security-relevant, but this change removes attack surface (unused parsers for rar/7zip/cab/iso/xar/zip/etc. are no longer compiled in) rather than adding any. The patched C functions do not touch untrusted input — they operate on caller-supplied integer codes. No new parsing or filesystem behavior is introduced.

Level of scrutiny

Medium-high. This is a semantic patch to vendored C plus a hand-curated compile list, and it sits directly under the bun install streaming-extract hot path. Per the repo guidance, dependency/vendoring changes warrant a maintainer look. The failure mode of a mistake here is either a link error (loud, caught by CI) or the streaming extractor failing to open archives for tarballs above the 2 MiB streaming threshold (should be caught by the install tarball tests the author ran, but only if those tests exceed the threshold).

Other factors

I cross-checked the patch against the upstream archive_read_set_format.c/archive_read_append_filter.c at the pinned commit: the slot/bidder lookup loops are unchanged and safely handle the "not registered" case (they return ARCHIVE_FATAL with the new error string, no null deref). archive_read_append_filter_program_signature still references archive_read_support_filter_program, which is intentionally kept in SOURCES. The removed blake2/ppmd8 objects were only reachable from the rar5 reader, which is also removed. No new automated test is added — coverage relies on the existing streaming-extract and install-tarball suites, which is reasonable for a build-config change but is another reason a human should confirm CI is green across all targets before merge.

@robobun

robobun commented Aug 18, 2026

Copy link
Copy Markdown
Collaborator
Updated 8:42 PM PT - Aug 17th, 2026

✅ @alii, your commit 53d522480ab6be095241fb70f16f344e01ce25f0 passed in Build #100323! 🎉


🧪   To try this PR locally:

bunx bun-pr 39484

That installs a local version of the PR into your bun-39484 executable, so you can run:

bun-39484 --bun

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants