Skip to content

bundler: keep sqlite imports external when the loader comes from the extension or loader map - #38275

Open
robobun wants to merge 7 commits into
mainfrom
farm/89088782/sqlite-loader-map-external
Open

robobun wants to merge 7 commits into
mainfrom
farm/89088782/sqlite-loader-map-external

Conversation

@robobun

@robobun robobun commented Aug 13, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • bun build --target bun with the sqlite loader selected by extension or loader map (--loader .db:sqlite, bunfig [loader] ".db" = "sqlite", Bun.build({ loader: { ".db": "sqlite" } }), or importing a .sqlite file without an attribute) emits the build machine's absolute path:
    var app_default = import.meta.require("/home/me/project/app.db", { type: "sqlite" }).db;
    The bundle (or --compile executable) opens that path wherever it runs, so it breaks once moved. The same import written as import db from "./app.db" with { type: "sqlite" } is emitted unchanged and left external.
  • Cause: resolve_import_records (src/bundler/bundle_v2.rs) only externalizes a record whose loader an attribute set, in a check that runs before resolution. A loader that comes from the file's extension is only known after resolution, so the database becomes a module in the graph and the Loader::Sqlite arm of src/bundler/ParseTask.rs builds the module from source.path.text, which is absolute. The plugin fallback resolver (run_resolver, used when an onResolve plugin matched but returned nothing) and the in-memory files path pick the loader separately and had the same gap.
  • Two printer gaps made externalizing alone insufficient or unsafe:
    • print_require_or_import_expr (src/js_printer/lib.rs) printed external require() and import() without the record's loader; only import statements got with { type } back. The cjs output format turns every import into require(), so even the attribute form printed __toESM(require("./app.db")) there, which loads a .db file as JavaScript.
    • resolve_import_records stored the loader it resolved for every bundled file on the import record. Where the linker later turns such a record external, that loader was printed on an import of something else: an import redirected through a module.exports = require(x) shim to an external x printed import * as y from "x" with { type: "js" } (on main and 1.4.0; x resolving to a .ts file at runtime then fails to parse), and with the require()/import() printing added here, --splitting would have printed import("./data-HASH.js", { with: { type: "json" } }) on every cross-chunk dynamic import (found in review; a JS chunk then fails to parse as JSON).

Fix

  • ImportRecord::loader now means one thing: a loader the runtime must apply to the import as printed. It is set by the parser from a type attribute (unchanged) and by the bundler when it leaves a sqlite import to the runtime; resolve_import_records no longer stores the loader it resolves for a bundled file there. Nothing else read that value: the linker, metafile, HTML output and dev server all take a bundled file's loader from the input file, and builds that go through onResolve plugins already ran with it unset. This fixes the shim redirect leak and makes the --splitting chunk rewrite inert, without touching the linker.
  • leaves_sqlite_import_to_runtime / mark_sqlite_import_external (bundle_v2.rs): after the resolved loader is known, a Loader::Sqlite record from an import statement, require() or import() is flagged IS_EXTERNAL_WITHOUT_SIDE_EFFECTS, given loader = Sqlite, and skipped, exactly like the attribute case. It runs on all four resolution paths (disk and files, directly and via the plugin fallback). The record still holds the specifier as written, so the output reads ./app.db relative to the bundle. Not applied when an onLoad or native onBeforeParse plugin filter matches the resolved file (the plugin keeps receiving it, as for any other loader; src/jsc/bindings/JSBundlerPlugin.cpp gains JSBundlerPlugin__anyOnBeforeParseMatches, the existing filter matcher applied to the native plugin list, because the onLoad matcher does not see those filters), for targets other than bun (the parse task still reports To use the "sqlite" loader, set target to "bun"), or in the dev server, which resolves imports itself and is unchanged.
  • Printer: an external require() gets , { type: "x" } and an external import() without user options gets , { with: { type: "x" } } when the record carries a loader, for the bun platform only (the only consumer of these forms; import.meta.require and import() in Bun accept them, and the code this replaces already emitted the require form). Import statements already did this; the three sites now share one import_attribute_type table, replacing two copies with byte-identical output.
  • Why this is the right behavior: the docs for the loader say the database "stays external to the bundle", which is what the attribute form has always done. A module in the graph cannot be made portable, since the non-embedded loader deliberately does not copy the file, so the only path the parse task could print is the source path. Leaving the import to the runtime with the type re-attached makes the loader-map output identical to the attribute output in every output format, including extensions the runtime would not recognize on its own (.db). sqlite_embedded is untouched; unused imports are still dropped (the record is external without side effects, matching the NoSideEffectsPureData the bundled module used to get).
  • Visible side effect: an external import that carries an attribute keeps its type in cjs output (__toESM(require("./db.sqlite", { type: "sqlite" }))), and a shim redirected to an external no longer prints the shim's loader. Output for records without an attribute is otherwise unchanged.
  • Not changed: an entry point that is itself a database, or a database a plugin's onResolve returned a path for, still goes through the parse task and prints the absolute path; there is no import specifier to keep in those cases.
  • Verified with test/bundler/bundler_bun.test.ts. The sqlite loader selected by extension or loader map block keeps a different database at build time and next to the bundle, so the program's output says which file the bundle opened: import (esm and cjs), require() (minified), import(), an unused import being dropped, --loader, bunfig [loader], Bun.build({ loader }) alone and combined with a passthrough onResolve plugin and/or files, an onLoad plugin taking precedence, plus the --target node error. test/bundler/native-plugin.test.ts gets two cases: a native plugin filtering .sqlite still receives the file (its counter sees the contents), and one filtering .ts leaves the import external (fails on the released build with the absolute path). The type attributes on imports the linker makes external block pins the shim redirect (import statement, require(), import(), run against a TypeScript external) and --splitting cross-chunk import() of .json and .ts files. 11 of the 15 new tests fail on the released binary (the sqlite ones with the absolute path, the redirect one with with { type: "js" }); the file passes with the debug build.
  • Also run with the debug build: bundler_edgecase, bundler_plugin, bundler_plugin_chain, native-plugin, bundler_files, bundler_barrel, bundler_splitting, bundler_compile_splitting, bundler_cjs, bundler_cjs2esm, bundler_loader, bun-build-api, metafile, bundler_html, html-import-manifest, esbuild/{loader,default,importstar,splitting}, bundler_compile -t sqlite, transpiler/transpiler.test.js, bake/dev/{bundle,plugins}, test/js/bun/import-attributes; native-plugin (full file); cargo clippy on bun_ast, bun_js_printer, bun_bundler; clang-format on the C++ file. Manually confirmed a --compile build with --loader .db:sqlite opens the database next to the executable.

Related

Background

  • Loader map: the extension to loader table built from --loader, bunfig [loader] and Bun.build({ loader }). Path::loader() (src/resolver/lib.rs) consults it and then falls back to Loader::from_string(extension), which is why a .sqlite extension selects the sqlite loader with no configuration and .db does not.
  • sqlite vs sqlite_embedded loaders: both make import db from "./x" evaluate to a bun:sqlite Database. sqlite opens the file at runtime and does not copy it; sqlite_embedded (with { type: "sqlite", embed: "true" }) copies it into the output directory or the compiled executable and references the copy.
  • Import record: the bundler's entry for one import. source_index points at the bundled module once one exists; a record without one is printed back out as an import of its path. IS_EXTERNAL_WITHOUT_SIDE_EFFECTS additionally lets tree shaking drop it when nothing uses the binding. The linker repurposes records in two places: a cross-chunk dynamic import under --splitting is pointed at the chunk file, and an import of a one-line module.exports = require(x) shim is redirected to x.
  • Re-attached attributes: for --target bun the printer adds with { type: "x" } to import statements from record.loader, so the runtime uses the loader the build decided on. The runtime accepts the same information on require(id, { type }) and import(id, { with: { type } }).
  • Resolution paths: resolve_import_records handles imports the bundler resolves itself; when an onResolve plugin's filter matches but the callback returns nothing, run_resolver does the same work after the round trip to the JS thread; both check Bun.build({ files }) (in-memory files) before the disk resolver.
Earlier version of this PR

The first push only added the externalization in resolve_import_records and keyed the new require()/import() printing on record.loader while resolution still stored every resolved loader there. Review found that --splitting output for target bun then carried { with: { type: "json" } } on cross-chunk dynamic imports and failed at runtime, and that shim redirects leaked the shim's loader (already the case for import statements on main). The second commit changes what the field holds instead, moves the sqlite check into a helper that also covers the plugin fallback and files paths, and adds the redirect and splitting tests. Later commits make a matching onLoad plugin, and then a matching native onBeforeParse plugin, take precedence over the externalization, which the first version skipped.


no test proof · iteration 0 · Platform-specific test(s) that do not run on this machine. Deferring to CI, which covers all platforms: test/bundler/native-plugin.test.ts

@coderabbitai

coderabbitai Bot commented Aug 13, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 27 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 59a7605e-fe84-45e7-976b-68605fba50ae

📥 Commits

Reviewing files that changed from the base of the PR and between 9c629a4 and 53210a0.

📒 Files selected for processing (6)
  • src/ast/import_record.rs
  • src/bundler/bundle_v2.rs
  • src/js_printer/lib.rs
  • src/jsc/bindings/JSBundlerPlugin.cpp
  • test/bundler/bundler_bun.test.ts
  • test/bundler/native-plugin.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Aug 13, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 9:57 AM PT - Aug 14th, 2026

✅ @robobun, your commit 53210a0823f3a7cbda6bce1cabebf3e41cad60d0 passed in Build #95932! 🎉


🧪   To try this PR locally:

bunx bun-pr 38275

That installs a local version of the PR into your bun-38275 executable, so you can run:

bun-38275 --bun

@robobun

robobun commented Aug 13, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: ready for review (head 53210a0).

Reproduced on release 1.4.0 and current main with bun build ./entry.mjs --target bun --loader .db:sqlite --outdir out (entry imports ./app.db): out/entry.js contains import.meta.require("/abs/path/app.db", { type: "sqlite" }). The same happens with bunfig [loader], Bun.build({ loader }), with a passthrough onResolve plugin or files in the mix, and for a plain .sqlite import with no attribute. The attribute form emits import db from "./app.db" with { type: "sqlite" } instead.

With this branch all of those emit the attribute form's output. Changes since the first push, all from review: resolved loaders are no longer stored on import records (they leaked onto --splitting chunk imports with the new printing, and already leaked onto shim redirects on main), the externalization covers the plugin fallback and files resolution paths, and a matching onLoad or native onBeforeParse plugin takes precedence over it. test/bundler/bundler_bun.test.ts: 11 of the 15 new cases fail on the released binary; test/bundler/native-plugin.test.ts: 1 of the 2 new cases does. All pass on the debug build.

CI for this head: every lane that has run is green (177 jobs); the only things outstanding are the two darwin 14 aarch64 - test-bun jobs, which have not been picked up by a macOS runner for several hours and keep being auto-retried, and a handful of tests marked flaky by CI itself (passed on retry or alone; none in the bundler). Nothing in the diff is waiting on a fix.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. Because it changes bundler output semantics (records that were previously bundled as modules are now left external) and adds a second argument to every external require()/import() printed for the bun target, a human look would still be worthwhile.

What was reviewed:

  • Traced resolve_import_records to confirm import_record.path still holds the original specifier (not the resolved absolute path) at the new continue, and that user---external and dev-server paths branch off before the loader is assigned.
  • Verified the import_attribute_type refactor is byte-identical to both replaced match tables, including the FP::Json special case.
  • Confirmed the new printer branches sit after the record.source_index.is_valid() early return, so bundled records never reach them; only attribute-set or newly-externalized sqlite records carry a loader there.
  • Checked __toESM(...) wrapping and ws! minified output against the test assertions.
Extended reasoning...

Overview

This PR fixes bun build --target bun emitting the build machine's absolute path when the sqlite loader is selected by extension or loader map (rather than by with { type: "sqlite" }). Three files: src/bundler/bundle_v2.rs adds a post-resolution check that externalizes Loader::Sqlite records (mirroring the pre-resolution attribute check at line ~6087); src/js_printer/lib.rs extracts two duplicate loader→type-string match tables into a shared import_attribute_type helper and teaches the external require() and import() printer branches to emit the type argument (import statements already did this); test/bundler/bundler_bun.test.ts adds 9 tests covering the variant matrix.

Security risks

None identified. The change affects bundler output for a Bun-specific loader; no auth, crypto, or untrusted-input parsing is involved.

Level of scrutiny

High. resolve_import_records is the bundler's central import-resolution loop, and print_require_or_import_expr shapes every external require/import in bundled output. The bundle_v2 change alters which records enter the module graph — a semantic shift, even though it aligns loader-map behavior with the documented attribute behavior. The printer change is not sqlite-specific: any external record with record.loader set (attribute-carrying imports that were also matched by --external, for instance) now gets a second argument on require()/import() under --target bun. I traced that such records are rare (bundled records return early at source_index.is_valid(), and post-resolution non-sqlite records always get a source index), and the emitted form is what Bun's runtime already accepts, but this is the kind of broad-reaching output change a maintainer should sign off on.

Other factors

The change is well-executed: the new check mirrors the existing attribute case with tighter guards (target.is_bun(), dev_server.is_none(), ImportKind::Stmt|Require|Dynamic), the refactor deduplicates two byte-identical tables per the repo's own guidance, and test coverage hits import/require/dynamic-import × esm/cjs × minified, tree-shaking of unused imports, all three loader-map surfaces (CLI flag, bunfig, Bun.build), and the --target node error path. The build-time vs runtime database trick makes the tests prove which file the bundle actually opens. The PR description lists an extensive set of adjacent bundler suites that were run. Nothing here looks wrong to me — deferring purely because the blast radius (bundler resolution + printer output for all bun-target externals) warrants a maintainer's eyes.

…extension or loader map

A file whose loader is selected as "sqlite" by the loader map (--loader
.db:sqlite, bunfig [loader], Bun.build({ loader })) or by a .sqlite
extension was bundled as a module that called import.meta.require() with
the absolute path of the database on the build machine. The import
attribute form of the same loader keeps the import external instead.

Externalize the record in resolve_import_records for that case as well,
and have the printer re-attach the type attribute on external require()
and import() calls (it already did so for import statements), so cjs
output and the require()/import() forms load the file with the sqlite
loader regardless of its extension.
…rinted back

resolve_import_records stored the loader it picked for every resolved
file on the import record. The printer re-emits a record's loader as a
type attribute for target bun, so the two places where the linker later
turns such a record external (dynamic imports rewritten to chunks under
--splitting, and imports redirected through a `module.exports =
require(x)` shim to an external x) printed the bundled file's loader on
an import of something else. For import statements this was already
happening (`import * as x from "ext" with { type: "js" }`); with the
previous commit it would also have reached require() and import().

The record now only carries a loader set by an import attribute or by
the bundler when it deliberately leaves the import to the runtime; the
loader of a bundled file stays with the input file, which is the only
thing the rest of the bundler reads. The sqlite externalization moves
into one helper and also covers the plugin fallback resolver and
in-memory `files`, which pick loaders on their own paths.
@robobun
robobun force-pushed the farm/89088782/sqlite-loader-map-external branch from c917ccd to 6aadb2d Compare August 14, 2026 09:29
Comment thread src/ast/import_record.rs Outdated
Comment thread src/bundler/bundle_v2.rs Outdated
Comment thread src/js_printer/lib.rs Outdated
Comment thread src/ast/import_record.rs Outdated
Comment thread src/bundler/bundle_v2.rs Outdated
Comment thread src/ast/import_record.rs
Comment thread src/bundler/bundle_v2.rs Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline findings, also checked: the two removed import_record.loader = Some(...) writes — remaining reads in src/bundler/ either unwrap_or_else to the path-based loader (6204/6566/2265/2499) or are on ParseTask/result objects, so no consumer loses the loader; the HTML-entrypoint guard at 6491 already treats None as HTML. An export ... from './x.sqlite' with the loader-map case was also examined and does not lose the type attribute in esm output.

Extended reasoning...

The two inline findings are a nit (onLoad bypass now matches the attribute form) and a pre-existing sibling under --splitting; neither is introduced by this PR. I separately grepped src/bundler/ for reads of import_record.loader to check the description's claim that nothing depended on the removed writes: all remaining reads fall back to path.loader(&options.loaders) when unset, and metafile/linker/input-file loader come from the parse result, not the record. This is still a semantic redefinition of a shared field across four resolution paths plus new printer output for external require()/import(), so a human pass on the bundler side is appropriate.

Comment thread src/bundler/bundle_v2.rs Outdated
Comment thread src/js_printer/lib.rs
… externalizing the import

An onLoad plugin whose filter matches the resolved file used to get the
file through the parse task; keep that ahead of the sqlite
externalization, as for any other loader. The decision moves into a
method that checks the plugin filters, and marking the record is a
separate helper.
Comment thread src/bundler/bundle_v2.rs Outdated
Comment thread src/bundler/bundle_v2.rs Outdated
Comment thread src/bundler/bundle_v2.rs Outdated
Comment thread src/bundler/bundle_v2.rs Outdated
…te import from being externalized

Native plugins keep their filters in a separate list that the onLoad
check does not see. Expose a match over that list and consult it too,
so the file still reaches the parse task that runs the plugin.
Comment thread src/bundler/bundle_v2.rs

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed the latest push (53210a0) and didn't find any bugs — the onBeforeParse gap from my last comment is addressed with a per-filter match (JSBundlerPlugin__anyOnBeforeParseMatches) and covered by two new tests in native-plugin.test.ts. Given the scope — redefining what ImportRecord::loader may hold, edits on four resolution paths in bundle_v2.rs, new C++ FFI, and overlap with #35692's approach to the shim-redirect leak — a human look is still worthwhile.

What was checked this run:

  • NativePluginList shares fileNamespace/namespaces/group() with NamespaceList, so the templated anyMatchesForNamespace instantiates; FilterRegExp::match already locks per call, same thread pattern as the existing has_any_matches.
  • filter_input allocates fresh BunStrings per call, so the two consecutive plugin checks each transfer their own copies.
  • import_attribute_type output is byte-identical to the two loader→string tables it replaces.
  • The outstanding comment-cop note on bundle_v2.rs:5982 is a two-line doc comment, not a workaround.
Extended reasoning...

Overview

The PR makes bundler-selected sqlite imports (via .sqlite extension, --loader, bunfig, or Bun.build({ loader })) print as external with { type: "sqlite" } imports instead of bundling a module holding the build machine's absolute path. It touches src/ast/import_record.rs (doc comment tightening the loader field's contract), src/bundler/bundle_v2.rs (new leaves_sqlite_import_to_runtime/mark_sqlite_import_external called on all four resolution paths; removes the two writes that stored a bundled file's resolved loader on the record), src/js_printer/lib.rs (adds { type } / { with: { type } } to external require()/import(); consolidates two loader→attribute tables into one import_attribute_type), src/jsc/bindings/JSBundlerPlugin.cpp (templated anyMatchesForNamespace + new JSBundlerPlugin__anyOnBeforeParseMatches FFI), and ~330 lines of new tests across bundler_bun.test.ts and native-plugin.test.ts.

Since my last inline comment at 11:14 UTC, commit 53210a0 added the native-plugin filter check I asked for — the tighter per-filter variant rather than the coarse has_on_before_parse_plugins() fallback — plus a shared filter_input helper on the Rust side and two tests exercising both the match and no-match cases.

Security risks

None identified. No auth, crypto, or untrusted-input parsing is touched. The new FFI (JSBundlerPlugin__anyOnBeforeParseMatches) reads plugin filter lists that are populated at setup time and follows the same locking pattern as the existing JSBundlerPlugin__anyMatches.

Level of scrutiny

High. This is bundler-core: it changes what ImportRecord::loader is allowed to hold (from "whatever loader was resolved" to "only a loader the runtime must apply to the printed import"), which is an invariant every future writer of that field must respect. The four resolution-path edits in bundle_v2.rs are structurally similar but not mechanical, and the printer now emits attributes on require()/import() where it previously did not. The change is well-argued and thoroughly tested (15 new cases, 11 fail on the released binary), but it is not the kind of simple/obvious change auto-approval is meant for.

Other factors

  • The PR description notes #35692 fixes the shim-redirect symptom by copying the loader at the redirect, and says the two approaches are compatible. A maintainer may still want to decide whether both should land or whether this PR's at-source fix supersedes #35692.
  • The one remaining unresolved bot comment (comment-cop on bundle_v2.rs:5982) targets a two-line doc comment on leaves_sqlite_import_to_runtime; the author already pushed back on identical notes and I don't consider it blocking.
  • All three of my earlier inline concerns (onLoad precedence, the pre-existing user-options --splitting sibling, and onBeforeParse precedence) have been addressed or explicitly deferred with a note in the description.

@robobun

robobun commented Aug 20, 2026

Copy link
Copy Markdown
Collaborator Author

One more symptom that the require() hunk in print_require_or_import_expr covers, for the record. The attribute form also loses its type in cjs output on main, and --bytecode implies cjs:

bun -e 'const {Database}=require("bun:sqlite"); const db=new Database("my.db"); db.run("create table t(x)"); db.run("insert into t values (1)")'
printf 'import db from "./my.db" with { type: "sqlite" };\nconsole.log(db.query("select x from t").get());\n' > index.ts
bun build --target=bun --format=cjs ./index.ts --outfile out.cjs && bun out.cjs

On main (34cbb9a40b) the bundle contains var import_my = __toESM(require("./my.db")); and fails at runtime with error: Expected ";" but found "format" at my.db:1:8, because the runtime loads the file with the default loader. --format=esm prints with { type: "sqlite" } for the same input. The record reaches the external require() branch with loader = Sqlite set by the parser, so the new second argument is what fixes it.

A note on coverage: bun/sqlite-extension-import-cjs uses db.sqlite. The runtime picks the sqlite loader for that extension on its own, so the run step of that test passes with or without the second argument. Only the toContain assertion pins it. A cjs case that imports a .db file (attribute form or loader map) fails at runtime without the second argument, so it would pin the behavior and not only the printed text.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant