Repository navigation
bun build --compile: embed text imports as pre-encoded strings - #40177
Conversation
|
Status: implemented on top of To see the runtime behavior by hand: printf 'em \xe2\x80\x94 dash\n' > note.txt
echo 'import t from "./note.txt"; console.log(typeof t, JSON.stringify(t));' > app.ts
bun build --compile --bytecode ./app.ts --outfile app && ./app
# string "em — dash\n"Before this branch the same |
|
Warning Review limit reachedYour included review limit has been reached. You’re in a promotional period — use the checkbox below to run this review for free:
On-demand reviews are free for the next 28 days. After that, they cost $0.25 per reviewed file. How can I continue?Run this review now using the option above, or comment You can also wait for the limit to reset (next review available in 11 minutes), then comment An organization admin can change what happens after included review limits in Billing. How do review limits work?CodeRabbit enforces per-developer PR review limits within each organization. For paid Pro and Pro+ reviews, CodeRabbit uses a developer's included PR review attempts over the past 7 days to set the current hourly allowance. At typical activity levels, the full plan allowance applies. Higher sustained activity can lower the allowance until earlier attempts leave the 7-day window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (1)
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (4)
Included review availability: Your plan provides up to 5 included reviews per hour; 0 remain after this review. WalkthroughText imports now embed content in standalone executables, encode non-Latin-1 text as UTF-16, resolve as JavaScript string exports, and remain inline for other targets. Bundler asset naming, runtime ChangesText asset embedding
Suggested reviewers: 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
|
Updated 2:10 AM PT - Aug 23rd, 2026
✅ @robobun, your commit 0599ea6486f645aa55e0261b97c2518e54012e0a passed in 🧪 To try this PR locally: bunx bun-pr 40177That installs a local version of the PR into your bun-40177 --bun |
A `Loader::Text` import in a standalone executable used to become a JS
module with a string literal. The binary carried the text as module
source and, with --bytecode, again as a UTF-16 constant in the bytecode
cache, and the runtime copied it onto the heap on load.
The bundler now emits the text as an embedded asset in compile mode.
`StandaloneModuleGraph::to_bytes` stores it as a WTF::StringImpl body:
Latin-1 when every code point fits, UTF-16LE at an even offset otherwise.
The chunk holds `require("/$bunfs/root/<name>-<hash>.txt")`, and
`fetch_builtin_module` answers with a JSString over the section bytes
(ExportDefaultObject), so there is no parse, no bytecode, and no copy.
Also:
- ModuleLoader.cpp handles the ExportDefaultObject tag from
Bun__fetchBuiltinModule in both the require() and the ESM path.
- The lazy-export ESM path imports the runtime `__require` when the
export is a `require(<asset>)` call. `import addon from "./x.node"`
in ESM output threw "__require is not defined" before.
- Text modules ignore --asset-naming so two same-named files never share
a path. `to_bytes` skips a repeated path instead of miscounting.
- `Bun.embeddedFiles` does not list text modules.
- Browser chunks of a full-stack executable keep the inline literal.
Co-authored-by: Jarred Sumner <jarred@jarredsumner.com>
3b37588 to
9ce8bb2
Compare
|
@robobun fix clippy |
|
Fixed in ceff847: clippy flagged |
UTF-8 to UTF-16 already went through simdutf; the Latin-1 decision and the narrowing copy were scalar loops over every code unit, and the UTF-16 write was a per-unit loop. Bind simdutf's convert_utf16le_to_latin1_with_errors (one pass, fails at the first unit above U+00FF) and memcpy the UTF-16 units. Drop the unused specifier/source_url refs from the text module ResolvedSource: the ExportDefaultObject consumers never deref them.
to_utf16_alloc already answers the width question: None means the bytes are ASCII and stay 8-bit, Some means UTF-16. Drop the Latin-1 narrowing pass and the simdutf binding it needed; this is the same split String::clone_utf8 makes at runtime.
|
@robobun fix conflicts |
…-text-modules # Conflicts: # src/standalone_graph/StandaloneModuleGraph.rs
ParseTask.rs: #40177 moved the embedded-asset require into require_embedded_asset and routes compiled text imports through it; the helper and the text branch build their nodes with Loc::new(0) like the rest of the synthesized ASTs here.
Problem
import md from "./x.md"(text loader) in abun build --compileexecutable became a JS module with a string literal. The binary carried the text as module source and, with--bytecode, again as a UTF-16 constant in the bytecode cache. At load, JSC parsed the module and copied the string onto the heap. A 3 MB text import cost 6 to 9 MB of binary with--bytecodeand 13 to 19 MB of RSS.jarred/compile-text-modules-drafthad the right shape but every text import came back as an emptyModule {}:ModuleLoader.cpphad no case for theExportDefaultObjecttag when it comes fromBun__fetchBuiltinModule, so it wrapped empty source.Fix
ParseTask(compile mode, bun target) registers a text file as an embedded asset and makes the moduleexport default require("<bunfs path>").to_byteswrites the bytes as aWTF::StringImplbody: 8-bit when the text is ASCII, UTF-16LE at an even offset otherwise (encode_text_module, the same split asString::clone_utf8).fetch_builtin_moduleanswers with a JSString over the section bytes (ExportDefaultObject): no parse, no bytecode, no copy.ModuleLoader.cpphandlesExportDefaultObjectfrom the builtin probe in both therequire()path (the value ismodule.exports) and the ESM path (synthetic module withdefault).require(...)(the napi loader's call target), notimport.meta.require, so--bytecode(CommonJS output) compiles it. The lazy-export ESM path now imports the runtime__requirefor such a call. This also fixesimport addon from "./x.node"in ESM output, which threwReferenceError: __require is not defined.test/bundler/bundler_compile.test.ts(compile/TextImport*, 5 new tests: encodings,require,import(),--bytecodein both formats, same basenames, browser chunk) andtest/bundler/bundler_loader.test.ts(napi ESM). Also the rest ofbundler_compile,bundler_loader,compile-asset-bunfs,bundler_html_server,bundler_compile_splitting,bun-build-compile,text-loader.Background
Encodingtag. At startupGraph::from_bytesmaps it, and the resolver answers/$bunfs/root/...specifiers from it.File::to_wtf_stringwraps Latin-1 bytes in a zero-copyExternalStringImpl.Encoding::Utf16takes the value of the never-writtenUtf8variant, so an older runtime reads it through its plain-copy arm instead of an invalid discriminant.export default <expr>. The linker prints it as ESM (var x_default = expr) or, whenrequire()d, as CommonJS (module.exports = expr).ERequireCallTargetprints asrequirein CommonJS output and as the runtime's__require(import.meta.requirefor the bun target) otherwise.ExportDefaultObjectis theResolvedSourcetag the file loader already uses: the runtime hands over a JSValue and the module loader exports it asdefault.Notes
Measurements, 3 MB text import, delta over a tiny import. "After" RSS is the debug+ASAN build, so it is an upper bound.
--bytecodeThe mixed case is mostly ASCII with a non-ASCII character on every line. UTF-16 doubles it, which is slightly larger than the escaped literal without
--bytecode, and half of it with. Text that is only non-ASCII in the Latin-1 range (café) is stored as UTF-16 too; that matches whatclone_utf8would build at runtime and keeps the encoder to one simdutf pass.Decisions on the open points of the draft:
--asset-namingdoes not apply to text modules; they always use[name]-[hash].[ext]. With a user template without[hash], two same-named text files would have shared one path and one importer would silently have read the other's text. Two files with the same name and the same bytes print the same path;to_bytesnow skips the repeated path instead of trippingdebug_assert_eq!(graph.files.count(), modules.len()).Bun.embeddedFilesdoes not list text modules: the docs example serves every entry as a static route, and these bytes are not the file's UTF-8.fs.readFileSyncon the hashed bunfs path still returns the encoded bytes (documented indocs/bundler/executables.mdx). compile: store module text in the width JSC loads it in #38801 adds aFile::utf8_contents()view for the same problem on JS modules; extending it to text modules is a natural follow-up.topts.target.is_bun()gate). The client transpiler inheritscompile_mode.TextDecoder. The bundler's literal path producedÿþplus a NUL for the same bytes..txtpassed as an entry point compiles and runs (exit 0), as before.--compile --bytecodewith the sqlite embedded loader still fails with "import.meta is only valid inside modules" (pre-existing; that loader keepsimport.meta.require(path, { type: "sqlite" })).Overlap: #38801 also adds
Encoding::Utf16 = 2andbun_core::String::create_static_external_utf16for JS modules with non-ASCII text. Whichever lands second rebases onto the same two definitions (this PR no longer adds a Latin-1 narrowing pass).The fail-before of
compile/TextImport*is theinvalidcheck and then the missing<name>-<hash>.txtentry in the bunfs root.compile/TextImportSameBasenameandcompile/TextImportClientChunkStaysInlinepass before and after; they guard the two new failure modes.Debug-build-only failures in the suites above that are not related:
compile/HelloWorldWithProcessVersionsBun(version string-debugsuffix) andtext-loader > dynamic-import reloaded 10000 times(12 s under ASAN against a 5 s timeout).no test proof · iteration 1 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/bundler/bundler_compile.test.ts