Skip to content

bun build --compile: embed text imports as pre-encoded strings - #40177

Merged
Jarred-Sumner merged 7 commits into
mainfrom
farm/6720e7a2/compile-text-modules
Aug 23, 2026
Merged

Jarred-Sumner merged 7 commits into
mainfrom
farm/6720e7a2/compile-text-modules

Conversation

@robobun

@robobun robobun commented Aug 23, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • import md from "./x.md" (text loader) in a bun build --compile executable became a JS module with a string literal. The binary carried the text as module source and, with --bytecode, again as a UTF-16 constant in the bytecode cache. At load, JSC parsed the module and copied the string onto the heap. A 3 MB text import cost 6 to 9 MB of binary with --bytecode and 13 to 19 MB of RSS.
  • The draft on jarred/compile-text-modules-draft had the right shape but every text import came back as an empty Module {}: ModuleLoader.cpp had no case for the ExportDefaultObject tag when it comes from Bun__fetchBuiltinModule, so it wrapped empty source.

Fix

  • ParseTask (compile mode, bun target) registers a text file as an embedded asset and makes the module export default require("<bunfs path>"). to_bytes writes the bytes as a WTF::StringImpl body: 8-bit when the text is ASCII, UTF-16LE at an even offset otherwise (encode_text_module, the same split as String::clone_utf8). fetch_builtin_module answers with a JSString over the section bytes (ExportDefaultObject): no parse, no bytecode, no copy.
  • ModuleLoader.cpp handles ExportDefaultObject from the builtin probe in both the require() path (the value is module.exports) and the ESM path (synthetic module with default).
  • The call is require(...) (the napi loader's call target), not import.meta.require, so --bytecode (CommonJS output) compiles it. The lazy-export ESM path now imports the runtime __require for such a call. This also fixes import addon from "./x.node" in ESM output, which threw ReferenceError: __require is not defined.
  • Verified: test/bundler/bundler_compile.test.ts (compile/TextImport*, 5 new tests: encodings, require, import(), --bytecode in both formats, same basenames, browser chunk) and test/bundler/bundler_loader.test.ts (napi ESM). Also the rest of bundler_compile, bundler_loader, compile-asset-bunfs, bundler_html_server, bundler_compile_splitting, bun-build-compile, text-loader.

Background

  • A compiled executable embeds a standalone module graph: a section with every bundled file's name, bytes, and an Encoding tag. At startup Graph::from_bytes maps it, and the resolver answers /$bunfs/root/... specifiers from it. File::to_wtf_string wraps Latin-1 bytes in a zero-copy ExternalStringImpl.
  • JSC stores strings as Latin-1 or UTF-16 code units, never UTF-8. Handing it the exact width is what makes zero-copy possible. Encoding::Utf16 takes the value of the never-written Utf8 variant, so an older runtime reads it through its plain-copy arm instead of an invalid discriminant.
  • A lazy export is a module whose whole body is export default <expr>. The linker prints it as ESM (var x_default = expr) or, when require()d, as CommonJS (module.exports = expr). ERequireCallTarget prints as require in CommonJS output and as the runtime's __require (import.meta.require for the bun target) otherwise.
  • ExportDefaultObject is the ResolvedSource tag the file loader already uses: the runtime hands over a JSValue and the module loader exports it as default.
Notes

Measurements, 3 MB text import, delta over a tiny import. "After" RSS is the debug+ASAN build, so it is an upper bound.

binary, no bytecode binary, --bytecode RSS
ASCII, before 3.0 MB 6.0 MB +12.6 MB
ASCII, after 3.0 MB 3.0 MB +1.0 MB
mixed unicode, before 4.3 MB 8.8 MB +18.9 MB
mixed unicode, after 4.6 MB 4.6 MB +0.9 MB

The mixed case is mostly ASCII with a non-ASCII character on every line. UTF-16 doubles it, which is slightly larger than the escaped literal without --bytecode, and half of it with. Text that is only non-ASCII in the Latin-1 range (café) is stored as UTF-16 too; that matches what clone_utf8 would build at runtime and keeps the encoder to one simdutf pass.

Decisions on the open points of the draft:

  • --asset-naming does not apply to text modules; they always use [name]-[hash].[ext]. With a user template without [hash], two same-named text files would have shared one path and one importer would silently have read the other's text. Two files with the same name and the same bytes print the same path; to_bytes now skips the repeated path instead of tripping debug_assert_eq!(graph.files.count(), modules.len()).
  • Bun.embeddedFiles does not list text modules: the docs example serves every entry as a static route, and these bytes are not the file's UTF-8. fs.readFileSync on the hashed bunfs path still returns the encoded bytes (documented in docs/bundler/executables.mdx). compile: store module text in the width JSC loads it in #38801 adds a File::utf8_contents() view for the same problem on JS modules; extending it to text modules is a natural follow-up.
  • Browser chunks of a full-stack executable keep the inline literal (topts.target.is_bun() gate). The client transpiler inherits compile_mode.
  • Invalid UTF-8 decodes to U+FFFD like TextDecoder. The bundler's literal path produced ÿþ plus a NUL for the same bytes.
  • A .txt passed as an entry point compiles and runs (exit 0), as before.
  • --compile --bytecode with the sqlite embedded loader still fails with "import.meta is only valid inside modules" (pre-existing; that loader keeps import.meta.require(path, { type: "sqlite" })).

Overlap: #38801 also adds Encoding::Utf16 = 2 and bun_core::String::create_static_external_utf16 for JS modules with non-ASCII text. Whichever lands second rebases onto the same two definitions (this PR no longer adds a Latin-1 narrowing pass).

The fail-before of compile/TextImport* is the invalid check and then the missing <name>-<hash>.txt entry in the bunfs root. compile/TextImportSameBasename and compile/TextImportClientChunkStaysInline pass before and after; they guard the two new failure modes.

Debug-build-only failures in the suites above that are not related: compile/HelloWorldWithProcessVersionsBun (version string -debug suffix) and text-loader > dynamic-import reloaded 10000 times (12 s under ASAN against a 5 s timeout).


no test proof · iteration 1 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/bundler/bundler_compile.test.ts

@robobun

robobun commented Aug 23, 2026

Copy link
Copy Markdown
Collaborator Author

Status: implemented on top of jarred/compile-text-modules-draft (fa2e6da), tested locally on Linux x64 with the debug build, waiting for CI (Windows and macOS alignment paths run there).

To see the runtime behavior by hand:

printf 'em \xe2\x80\x94 dash\n' > note.txt
echo 'import t from "./note.txt"; console.log(typeof t, JSON.stringify(t));' > app.ts
bun build --compile --bytecode ./app.ts --outfile app && ./app
# string "em — dash\n"

Before this branch the same --bytecode build stored the text twice and copied it to the heap on load; the require()/ESM path for the ExportDefaultObject tag did not exist, so the draft printed Module {}.

@coderabbitai

coderabbitai Bot commented Aug 23, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Warning

Review limit reached

Your included review limit has been reached.

You’re in a promotional period — use the checkbox below to run this review for free:

  • Run review for free

On-demand reviews are free for the next 28 days. After that, they cost $0.25 per reviewed file.

How can I continue?

Run this review now using the option above, or comment @coderabbitai review --use-credits.

You can also wait for the limit to reset (next review available in 11 minutes), then comment @coderabbitai review or push new commits to the PR.

An organization admin can change what happens after included review limits in Billing.

How do review limits work?

CodeRabbit enforces per-developer PR review limits within each organization.

For paid Pro and Pro+ reviews, CodeRabbit uses a developer's included PR review attempts over the past 7 days to set the current hourly allowance. At typical activity levels, the full plan allowance applies. Higher sustained activity can lower the allowance until earlier attempts leave the 7-day window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: fd399aee-ac6d-4b7d-a121-aff0f0040757

📥 Commits

Reviewing files that changed from the base of the PR and between 4db7194 and 0599ea6.

📒 Files selected for processing (1)
  • src/standalone_graph/StandaloneModuleGraph.rs

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 65be02ff-3f78-450e-8e48-1b22f7f4ccad

📥 Commits

Reviewing files that changed from the base of the PR and between ceff847 and 4db7194.

📒 Files selected for processing (4)
  • docs/bundler/executables.mdx
  • src/runtime/jsc_hooks.rs
  • src/standalone_graph/StandaloneModuleGraph.rs
  • test/bundler/bundler_compile.test.ts

Included review availability: Your plan provides up to 5 included reviews per hour; 0 remain after this review.


Walkthrough

Text imports now embed content in standalone executables, encode non-Latin-1 text as UTF-16, resolve as JavaScript string exports, and remain inline for other targets. Bundler asset naming, runtime require() handling, documentation, and regression tests were updated.

Changes

Text asset embedding

Layer / File(s) Summary
Standalone text encoding and serialization
src/standalone_graph/StandaloneModuleGraph.rs
Standalone serialization detects text modules, encodes Latin-1 or UTF-16 content, excludes text modules from Bun.embeddedFiles, and suppresses duplicate paths.
Bundler asset registration and output
src/bundler/ParseTask.rs, src/bundler/bundle_v2.rs, src/bundler/linker_context/generateCodeForLazyExport.rs
Text-loader parsing registers standalone assets. Shared helpers generate asset keys and require() expressions. Output naming and runtime __require handling support these assets.
Runtime string module resolution
src/bun_core/string/mod.rs, src/runtime/jsc_hooks.rs, src/jsc/bindings/ModuleLoader.cpp
Standalone text contents become JavaScript strings. CommonJS and ESM loading resolve the strings through default exports.
Text import behavior validation and documentation
test/bundler/bundler_compile.test.ts, test/bundler/bundler_loader.test.ts, docs/bundler/executables.mdx
Documentation and tests cover loaders, encodings, empty and invalid files, asset naming, embedded-file behavior, browser inlining, and runtime require() handling.

Suggested reviewers: jarred-sumner, dylan-conway

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely describes embedding text imports as pre-encoded strings during compiled builds.
Description check ✅ Passed The description explains the problem, implementation, design decisions, and verification results in sufficient detail.

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Aug 23, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 2:10 AM PT - Aug 23rd, 2026

✅ @robobun, your commit 0599ea6486f645aa55e0261b97c2518e54012e0a passed in Build #104050! 🎉


🧪   To try this PR locally:

bunx bun-pr 40177

That installs a local version of the PR into your bun-40177 executable, so you can run:

bun-40177 --bun

Jarred-Sumner and others added 2 commits August 23, 2026 07:46
A `Loader::Text` import in a standalone executable used to become a JS
module with a string literal. The binary carried the text as module
source and, with --bytecode, again as a UTF-16 constant in the bytecode
cache, and the runtime copied it onto the heap on load.

The bundler now emits the text as an embedded asset in compile mode.
`StandaloneModuleGraph::to_bytes` stores it as a WTF::StringImpl body:
Latin-1 when every code point fits, UTF-16LE at an even offset otherwise.
The chunk holds `require("/$bunfs/root/<name>-<hash>.txt")`, and
`fetch_builtin_module` answers with a JSString over the section bytes
(ExportDefaultObject), so there is no parse, no bytecode, and no copy.

Also:
- ModuleLoader.cpp handles the ExportDefaultObject tag from
  Bun__fetchBuiltinModule in both the require() and the ESM path.
- The lazy-export ESM path imports the runtime `__require` when the
  export is a `require(<asset>)` call. `import addon from "./x.node"`
  in ESM output threw "__require is not defined" before.
- Text modules ignore --asset-naming so two same-named files never share
  a path. `to_bytes` skips a repeated path instead of miscounting.
- `Bun.embeddedFiles` does not list text modules.
- Browser chunks of a full-stack executable keep the inline literal.

Co-authored-by: Jarred Sumner <jarred@jarredsumner.com>
@robobun
robobun force-pushed the farm/6720e7a2/compile-text-modules branch from 3b37588 to 9ce8bb2 Compare August 23, 2026 07:48
Comment thread src/bun_core/string/mod.rs
Comment thread src/bundler/ParseTask.rs Outdated
Comment thread src/bundler/ParseTask.rs Outdated
Comment thread src/bundler/ParseTask.rs Outdated
Comment thread src/bundler/ParseTask.rs Outdated
Comment thread src/bundler/bundle_v2.rs Outdated
Comment thread src/bundler/bundle_v2.rs Outdated
Comment thread src/bundler/linker_context/generateCodeForLazyExport.rs Outdated
Comment thread src/jsc/bindings/ModuleLoader.cpp Outdated
Comment thread src/jsc/bindings/ModuleLoader.cpp Outdated
Comment thread src/runtime/jsc_hooks.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/bundler/linker_context/generateCodeForLazyExport.rs
Comment thread src/standalone_graph/StandaloneModuleGraph.rs
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
@Jarred-Sumner

Copy link
Copy Markdown
Collaborator

@robobun fix clippy

@robobun

robobun commented Aug 23, 2026

Copy link
Copy Markdown
Collaborator Author

Fixed in ceff847: clippy flagged manual_is_multiple_of (twice), cast_ptr_alignment on the *const u8 -> *const u16 view in File::to_wtf_string, and chunks_exact_to_as_chunks in encode_text_module. The pointer cast keeps an #[expect] with the alignment reason, the other three use is_multiple_of and as_chunks_mut::<2>. cargo clippy --workspace --no-deps is clean locally.

Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
Comment thread src/runtime/jsc_hooks.rs
UTF-8 to UTF-16 already went through simdutf; the Latin-1 decision and
the narrowing copy were scalar loops over every code unit, and the
UTF-16 write was a per-unit loop. Bind simdutf's
convert_utf16le_to_latin1_with_errors (one pass, fails at the first unit
above U+00FF) and memcpy the UTF-16 units.

Drop the unused specifier/source_url refs from the text module
ResolvedSource: the ExportDefaultObject consumers never deref them.
Comment thread src/bun_core/string/immutable.rs Outdated
Comment thread src/simdutf_sys/simdutf.rs Outdated
Comment thread src/standalone_graph/StandaloneModuleGraph.rs Outdated
to_utf16_alloc already answers the width question: None means the bytes
are ASCII and stay 8-bit, Some means UTF-16. Drop the Latin-1 narrowing
pass and the simdutf binding it needed; this is the same split
String::clone_utf8 makes at runtime.
Comment thread src/standalone_graph/StandaloneModuleGraph.rs
@Jarred-Sumner

Copy link
Copy Markdown
Collaborator

@robobun fix conflicts

…-text-modules

# Conflicts:
#	src/standalone_graph/StandaloneModuleGraph.rs
@robobun

robobun commented Aug 23, 2026

Copy link
Copy Markdown
Collaborator Author

Merged main in 0599ea6. The conflict was with #39912 (region layout of the payload): text modules now take their string-body encoding in the contiguous source-text run, and the duplicate-path skip uses the new module_dest_path helper. Compile tests, the new madvise tests, and clippy pass locally.

@Jarred-Sumner
Jarred-Sumner merged commit 0091072 into main Aug 23, 2026
10 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the farm/6720e7a2/compile-text-modules branch August 23, 2026 10:13
robobun added a commit that referenced this pull request Aug 23, 2026
ParseTask.rs: #40177 moved the embedded-asset require into
require_embedded_asset and routes compiled text imports through it; the helper
and the text branch build their nodes with Loc::new(0) like the rest of the
synthesized ASTs here.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants