Skip to content

standalone_graph: zero-copy source string for compiled JS modules when contents are ASCII - #31557

Merged
Jarred-Sumner merged 3 commits into
mainfrom
claude/standalone-graph-encoding-fix
May 29, 2026
Merged

Jarred-Sumner merged 3 commits into
mainfrom
claude/standalone-graph-encoding-fix

Conversation

@dylan-conway

@dylan-conway dylan-conway commented May 29, 2026 •

Copy link
Copy Markdown
Member

from_bytes() was hardcoding encoding: Encoding::Binary for every embedded file instead of reading the per-module value serialized by to_bytes(), so File::to_wtf_string() always took the clone_utf8 path — a fresh heap copy of the entire bundled source at module load — instead of the zero-copy create_static_external path that wraps the kernel-mmapped .bun section directly.

Honoring the serialized encoding alone is unsafe: to_bytes() previously tagged JS as Latin1 purely by loader type, but --banner / --footer / hashbang and client-side (target=browser) chunks are concatenated verbatim as UTF-8 by postProcessJSChunk, so non-ASCII bytes there would mojibake under a Latin-1 ExternalStringImpl (e.g. --footer 'console.log("résumé")' → résumé). This PR gates the Latin1 tag in to_bytes() on strings::first_non_ascii(buf_bytes).is_none() — a build-time SIMD scan — so the runtime only takes the zero-copy path when it's actually safe, and falls back to clone_utf8 (today's behaviour) otherwise.

The printer escapes non-ASCII for server-side JS by default, so the common case (no banner/footer, no client chunks imported on the server) gets the zero-copy path. On a bun build --compile --bytecode binary with a 40 MB bundle, that avoids one 40 MB allocation + walk at startup and 40 MB of pinned heap that should have stayed evictable file-backed pages.

Adds compile/FooterNonAsciiUTF8 and compile/BannerNonAsciiUTF8 tests that fail (mojibake) without the ASCII guard.

(Supersedes the .zig-only #29319, which targets the no-longer-compiled reference implementation and gates on side == .server rather than verifying the bytes — that still mojibakes a non-ASCII --footer on a server chunk.)

@robobun

robobun commented May 29, 2026 •

Copy link
Copy Markdown
Collaborator
Updated 2:46 AM PT - May 29th, 2026

❌ @autofix-ci[bot], your commit ce4bd3d has 2 failures in Build #58915 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 31557

That installs a local version of the PR into your bun-31557 executable, so you can run:

bun-31557 --bun

@coderabbitai

coderabbitai Bot commented May 29, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 1c1d8976-6ec0-4784-ae01-7e7719fd9176

📥 Commits

Reviewing files that changed from the base of the PR and between e961c51 and ce4bd3d.

📒 Files selected for processing (2)
  • src/standalone_graph/StandaloneModuleGraph.rs
  • test/bundler/bundler_compile.test.ts

Walkthrough

Deserialization now restores each embedded File.encoding from the serialized CompiledModuleGraphFile.encoding. Serialization emits Latin1 for JS-like loaders only when the output buffer contains no non-ASCII bytes; otherwise it uses Binary. A bundler test verifies compiled output preserves UTF-8 non-ASCII bytes.

Changes

Encoding preservation and emission

Layer / File(s) Summary
Preserve module encoding through graph deserialization
src/standalone_graph/StandaloneModuleGraph.rs
When building embedded Files in StandaloneModuleGraph::from_bytes, File.encoding is taken from CompiledModuleGraphFile.encoding instead of being hardcoded to Encoding::Binary.
Selective Latin1 emission for JS-like loaders
src/standalone_graph/StandaloneModuleGraph.rs
StandaloneModuleGraph::to_bytes sets CompiledModuleGraphFile.encoding to Encoding::Latin1 for Loader::{Js,Jsx,Ts,Tsx} only if the emitted buffer has no non-ASCII bytes; otherwise it uses Encoding::Binary. Non-JS loaders remain Binary.
Bundler compile test for UTF-8 output
test/bundler/bundler_compile.test.ts
Adds a test matrix running bun build --compile with --banner and --footer, executes the compiled binary, and asserts stdout equals the expected UTF-8 string containing non-ASCII characters.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title accurately describes the main change: restoring zero-copy source string handling for compiled JS modules when ASCII, which directly aligns with the core fix in the changeset.
Description check ✅ Passed The description comprehensively explains the bug, the fix with the ASCII guard, performance impact, and references added tests, fully covering the required template sections.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Comment @coderabbitai help to get the list of available commands and usage tips.

@github-actions

Copy link
Copy Markdown
Contributor

This PR may be a duplicate of:

  1. fix(compile): avoid heap-copying $bunfs module sources #29319 - Fixes the same encoding: Encoding::Binary hardcoding bug in from_bytes(), plus additionally corrects client-chunk encoding in to_bytes() and adds a test

🤖 Generated with Claude Code

Comment thread src/standalone_graph/StandaloneModuleGraph.rs
@dylan-conway dylan-conway changed the title standalone_graph: read encoding from serialized record instead of hardcoding Binary standalone_graph: zero-copy source string for compiled JS modules when contents are ASCII May 29, 2026
@dylan-conway

Copy link
Copy Markdown
Member Author

Addressed in c48a2bb — gated Latin1 on strings::first_non_ascii at build time and added --footer/--banner non-ASCII tests that fail (mojibake) without the guard. Thanks for the catch on Zig parity; the Zig fromBytes does indeed default .encoding rather than read it.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — the ASCII guard in to_bytes() addresses the mojibake concern, and the new --banner/--footer tests lock it in.

Extended reasoning...

Overview

This PR touches src/standalone_graph/StandaloneModuleGraph.rs (two small hunks) and adds two regression tests in test/bundler/bundler_compile.test.ts. The Rust change (1) makes from_bytes() read the per-module encoding from the serialized record instead of hardcoding Encoding::Binary, and (2) makes to_bytes() tag JS modules as Latin1 only when strings::first_non_ascii(buf_bytes).is_none() — i.e., the final chunk bytes are verifiably all-ASCII. Together this lets File::to_wtf_string() take the zero-copy create_static_external path over the mmapped .bun section instead of heap-cloning the source at module load.

Security risks

None identified. No untrusted input handling, auth, or crypto is involved. The Latin1 external-string path wraps bytes that are already mmapped read-only into the process; the ASCII guard runs at build time on bundler output, not on runtime user input. The fallback path (Binary → clone_utf8) is the pre-PR behavior.

Level of scrutiny

Moderate. from_bytes() runs at startup of every bun build --compile binary and to_wtf_string() feeds source code to JSC, so a wrong encoding tag would mojibake or SyntaxError user code. However, the change is ~15 lines, the new behavior is gated on a byte-level SIMD ASCII check (not a loader-type heuristic), and any non-ASCII byte falls back to the existing safe path — so the worst case is identical to today's behavior. ASCII bytes are bit-identical under Latin-1 and UTF-8, so the zero-copy wrap is sound when the guard passes.

Other factors

I flagged the original mojibake regression (--banner/--footer/hashbang concatenated verbatim as UTF-8) on the first revision; the author addressed it in c48a2bb exactly as suggested — gating on actual byte contents rather than loader type — and added compile/FooterNonAsciiUTF8 and compile/BannerNonAsciiUTF8 tests that exercise both 2-byte (résumé) and 3-byte (こんにちは) UTF-8 sequences and would print mojibake without the guard. The bug-hunting pass on the updated revision found nothing. CODEOWNERS does not cover these paths. robobun shows build-rust failures on c48a2bb, but autofix.ci pushed ce4bd3d afterward; CI gates will block merge if anything remains red.

@Jarred-Sumner
Jarred-Sumner merged commit 496ce7e into main May 29, 2026
76 of 78 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the claude/standalone-graph-encoding-fix branch May 29, 2026 17:50
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants