Skip to content

printer: decode ill-formed UTF-8 as U+FFFD when quoting byte strings - #35306

Closed
robobun wants to merge 7 commits into
mainfrom
farm/86c4cb2d/text-loader-utf8-fffd
Closed

robobun wants to merge 7 commits into
mainfrom
farm/86c4cb2d/text-loader-utf8-fffd

Conversation

@robobun

@robobun robobun commented Jul 23, 2026 •

Copy link
Copy Markdown
Collaborator

The text loader (import t from "./f.txt" / with { type: "text" }) decodes invalid UTF-8 as Latin-1, while every other reader in Bun returns U+FFFD for the same bytes.

Repro

printf 'a\xff\xfeb' > t.txt
bun -e 'import t from "./t.txt"; console.log([...t].map(c=>c.charCodeAt(0)))'
# before: [97, 255, 254, 98]
# after:  [97, 65533, 65533, 98]
bun -e 'console.log([...await Bun.file("t.txt").text()].map(c=>c.charCodeAt(0)))'
# [97, 65533, 65533, 98]

Cause

The text loader wraps the raw file bytes in E::String, which the printer serializes via the Encoding::Utf8 path of write_pre_quoted_string_inner. That path used the WTF-8 stepper (wtf8_byte_sequence_length_with_invalid + decode_wtf8_rune_t with zero = 0). For ill-formed input it

  • widened an invalid lead byte or lone continuation byte to its Latin-1 code point (0xFF -> \xFF),
  • returned 0 for a bad multibyte sequence and still advanced by the lead-byte-implied width, so the following byte(s) were dropped (a\xC2 b -> [97, 0, 98], the space is gone), and
  • let WTF-8-encoded surrogates through verbatim (ED A0 80 -> \uD800).

Fix

Switch the Encoding::Utf8 arm to strings::convert_utf8_bytes_into_utf16, the same WHATWG decoder Bun.file().text() / to_utf16_alloc already use, which yields U+FFFD with the WebKit maximal-subpart advancement. On replacement.fail the loop emits \uFFFD (ascii-only output) or the raw EF BF BD bytes directly and continues, so neither the raw-byte fast path nor the \xHH branch ever sees the invalid input. Valid sequences are unchanged (regression-guarded in the new test).

Both the runtime loader and bun build share this printer path, so a bundled artifact now agrees too.

Verification

bun bd test test/js/bun/util/text-loader-invalid-utf8.test.ts

7 of the 10 cases fail on main and all 10 pass with this change; the 3 valid-UTF-8 cases guard against regressions.

In the bundler's non-ascii-only output the old path wrote the raw invalid byte into the JS source ("a<0xFF><0xFE>b"), which browsers reject as an invalid token; the output is now well-formed UTF-8.

Fixes #12981


[review] gate passed · iteration 1 · 6 files touched

fails on main (without fix)
ASAN without fix: 9 FAILED
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/bundler/bundler_loader.test.ts "test/js/bun/util/text-loader-invalid-utf8.test.ts"
bun test v1.4.0 (ee9d06e06)

test/bundler/bundler_loader.test.ts:
(pass) bundler > bun loader > bun/loader-yaml-file [901.97ms]
(pass) bundler > bun loader > bun/loader-text-file [540.52ms]
(pass) bundler > bun loader > bun/loader-json-file [522.15ms]
(pass) bundler > bun loader > bun/loader-toml-file [521.77ms]
(pass) bundler > bun loader > bun/loader-text-file [503.02ms]
(pass) bundler > node loader > bun/loader-yaml-file [517.93ms]
(pass) bundler > node loader > bun/loader-text-file [509.17ms]
(pass) bundler > node loader > bun/loader-json-file [524.89ms]
(pass) bundler > node loader > bun/loader-toml-file [521.33ms]
(pass) bundler > node loader > bun/loader-text-file [517.51ms]
(pass) bundler > bun/loader-text-file [836.54ms]
runtime failed file: /tmp/bun-build-tests/bun-sTOWF4/bun/loader-text-file-invalid-utf8-bun/out.js
stdout output:
[97,255,254,98]
---
expected stdout:
[97,65533,65533,98]
---
1843 |               console.log(`---`);
1844 |  
... (truncated)

release without fix: all passed
bun test v1.4.0-canary.1 (d7725e17b)

test/bundler/bundler_loader.test.ts:
(pass) bundler > bun loader > bun/loader-yaml-file [23.74ms]
(pass) bundler > bun loader > bun/loader-text-file [14.50ms]
(pass) bundler > bun loader > bun/loader-json-file [13.80ms]
(pass) bundler > bun loader > bun/loader-toml-file [14.13ms]
(pass) bundler > bun loader > bun/loader-text-file [13.74ms]
(pass) bundler > node loader > bun/loader-yaml-file [13.70ms]
(pass) bundler > node loader > bun/loader-text-file [16.36ms]
(pass) bundler > node loader > bun/loader-json-file [15.01ms]
(pass) bundler > node loader > bun/loader-toml-file [16.88ms]
(pass) bundler > node loader > bun/loader-text-file [14.11ms]
(pass) bundler > bun/loader-text-file [14.92ms]
(pass) bundler > bun/loader-text-file-invalid-utf8-bun [14.15ms]
(pass) bundler > bun/loader-text-file-invalid-utf8-browser [14.59ms]
(pass) bundler > bun/loader-json-proto-key-is-own-property [15.08ms]
(pass) bundler > bun/loader-toml-proto-key-is-own-property [14.81ms]
(pass) bundler > bun/loader-yaml-proto-key-is-own-property [14.24ms]
(pass) bundler > bun/loader-jsonc-proto-key-is-own-property [16.74ms]
(pass) bundler > bun/loader-json5-p
... (truncated)
passes on PR (with fix)
ASAN with fix: all passed
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/mechgate.xml" test/bundler/bundler_loader.test.ts "test/js/bun/util/text-loader-invalid-utf8.test.ts"
bun test v1.4.0 (ee9d06e06)

test/bundler/bundler_loader.test.ts:
(pass) bundler > bun loader > bun/loader-yaml-file [1005.05ms]
(pass) bundler > bun loader > bun/loader-text-file [515.89ms]
(pass) bundler > bun loader > bun/loader-json-file [521.20ms]
(pass) bundler > bun loader > bun/loader-toml-file [509.75ms]
(pass) bundler > bun loader > bun/loader-text-file [499.00ms]
(pass) bundler > node loader > bun/loader-yaml-file [523.40ms]
(pass) bundler > node loader > bun/loader-text-file [522.23ms]
(pass) bundler > node loader > bun/loader-json-file [524.00ms]
(pass) bundler > node loader > bun/loader-toml-file [526.27ms]
(pass) bundler > node loader > bun/loader-text-file [569.31ms]
(pass) bundler > bun/loader-text-file [842.38ms]
(pass) bundler > bun/loader-text-file-invalid-utf8-bun [558.93ms]
(pass) bundler > bun/loader-text-file-invalid-utf8-browser [519.04ms]
(pass) bundler > bun/loader-json-proto-key-is-own-property [512.86ms]
(pass) bundler >
... (truncated)

release with fix: all passed
$ bun scripts/build.ts --profile=release
[configured] bun-profile → bun (stripped) in 653ms (unchanged)
ninja: Entering directory `/workspace/bun/build/release'
[0/4] cargo bun_bin → libbun_rust.a (--target x86_64-unknown-linux-gnu)

  nightly-2026-07-20-x86_64-unknown-linux-gnu unchanged - rustc 1.99.0-nightly (9f36de775 2026-07-19)

�[1m�[92m    Blocking�[0m waiting for file lock on build directory
�[1m�[92m   Compiling�[0m bun_core v0.0.0 (/workspace/bun/src/bun_core)
�[1m�[92m   Compiling�[0m bun_errno v0.0.0 (/workspace/bun/src/errno)
�[1m�[92m   Compiling�[0m bun_ptr v0.0.0 (/workspace/bun/src/ptr)
�[1m�[92m   Compiling�[0m bun_boringssl_sys v0.0.0 (/workspace/bun/src/boringssl_sys)
�[1m�[92m   Compiling�[0m bun_safety v0.0.0 (/workspace/bun/src/safety)
�[1m�[92m   Compiling�[0m bun_zlib_sys v0.0.0 (/workspace/bun/src/zlib_sys)
�[1m�[92m   Compiling�[0m bun_cares_sys v0.0.0 (/workspace/bun/src/cares_sys)
�[1m�[92m   Compiling�[0m bun_zstd v0.0.0 (/workspace/bun/src/zstd)
�[1m�[92m   Compiling�[0m bun_picohttp v0.0.0 (/workspace/bun/src/picohttp)
�[1m�[92m   Compiling�[0m bun_output v0.0.0 (/workspace/bun/src/output)
�[1m�[92m   Compiling�[0m bun_clap v0.0.
... (truncated)
diff hotspot
src/bun_core/string/immutable.rs                  |  6 +-
 src/bun_core/string/immutable/unicode.rs          |  8 ++-
 src/bun_core/string/mod.rs                        | 41 +++++++----
 src/js_printer/lib.rs                             | 84 +++++++++--------------
 test/bundler/bundler_loader.test.ts               | 15 ++++
 test/js/bun/util/text-loader-invalid-utf8.test.ts | 54 +++++++++++++++
 6 files changed, 139 insertions(+), 69 deletions(-)

gate history · 4 passed · 0 rejected · iteration 1

evidence per changed file
file                                               reads  edits  tests
src/bun_core/string/immutable.rs                       5      1      0
src/bun_core/string/immutable/unicode.rs               6      3      0
src/bun_core/string/mod.rs                             3      3      0
src/js_printer/lib.rs                                  7      3      0
test/bundler/bundler_loader.test.ts                    1      2      0
test/js/bun/util/text-loader-invalid-utf8.test.ts      2      3      0

The text loader feeds raw file bytes into an E::String, which the
printer serializes via the Encoding::Utf8 path of
write_pre_quoted_string_inner. That path used the WTF-8 stepper
(wtf8_byte_sequence_length_with_invalid + decode_wtf8_rune_t with
zero=0), which for ill-formed input:

  * widened an invalid lead or lone continuation byte to its Latin-1
    code point (0xFF -> \xFF),
  * returned 0 for a bad multibyte sequence and still advanced by the
    lead-byte-implied width, dropping the following byte(s), and
  * let WTF-8-encoded surrogates through verbatim.

So importing a .txt file disagreed with Bun.file().text(),
fs.readFileSync(p, 'utf8'), and TextDecoder on the same bytes.

Switch the Utf8 arm to strings::convert_utf8_bytes_into_utf16, the
same WHATWG decoder Bun.file().text() already uses, which yields
U+FFFD with maximal-subpart advancement. On replacement.fail emit
\uFFFD (ascii-only output) or the raw EF BF BD bytes and continue,
so neither the raw-byte fast path nor the \xHH branch ever sees the
invalid input. Valid sequences keep the existing behaviour.
@github-actions

Copy link
Copy Markdown
Contributor

Found 1 issue this PR may fix:

  1. Uncaught SyntaxError when importing certain PDFs via text loader #12981 - Importing PDFs via the text loader (with { type: 'text' }) causes SyntaxError due to invalid UTF-8 bytes being emitted as broken JS string literals; this PR replaces those bytes with U+FFFD, producing valid output.

If this is helpful, copy the block below into the PR description to auto-close this issue on merge.

Fixes #12981

🤖 Generated with Claude Code

@coderabbitai

coderabbitai Bot commented Jul 23, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Changes

The UTF-8 decoder is made publicly accessible, quoted-string printing adopts WHATWG-style malformed UTF-8 replacement handling, and text-loader tests compare results with TextDecoder and filesystem decoding.

Invalid UTF-8 decoding

Layer / File(s) Summary
Expose UTF-8 decoder
src/bun_core/string/immutable/unicode.rs, src/bun_core/string/immutable.rs
convert_utf8_bytes_into_utf16 becomes public, and the existing re-export list is reformatted.
Update core string printing
src/bun_core/string/mod.rs
Core quoted-string printing uses WHATWG-style replacement decoding for malformed UTF-8 and emits escaped or raw U+FFFD output.
Apply replacement decoding and validate text loading
src/js_printer/lib.rs, test/js/bun/util/text-loader-invalid-utf8.test.ts, test/bundler/bundler_loader.test.ts
JavaScript string printing adopts replacement decoding, valid UTF-8 is asserted for literal printing, and tests cover malformed byte sequences across text loading and bundling paths.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed The change directly addresses #12981 by making the text loader emit U+FFFD and adding tests for invalid UTF-8.
Out of Scope Changes check ✅ Passed The formatting and visibility tweaks support the UTF-8 fix and no unrelated code changes are evident.
Title check ✅ Passed The title clearly summarizes the main change: decoding ill-formed UTF-8 as U+FFFD in the printer.
Description check ✅ Passed The description explains the change and includes verification details, though it uses custom headings instead of the template.

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/bun_core/string/immutable/unicode.rs (1)

413-424: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Document preconditions now that this crosses a crate boundary.

convert_utf8_bytes_into_utf16 now becomes part of js_printer's public contract via strings::convert_utf8_bytes_into_utf16, but its preconditions (non-empty slice, first byte non-ASCII) are only enforced by an unconditional unreachable!() and a debug-only debug_assert!. Any future caller that violates them will panic in release builds too. A short doc comment stating the precondition would make the widened surface safer to consume.

📝 Suggested doc comment
+/// Decodes one UTF-8 sequence starting at `bytes[0]` into UTF-16.
+/// Preconditions: `bytes` is non-empty and `bytes[0] >= 0x80` (non-ASCII lead byte).
+/// Violating either precondition panics (`unreachable!()`/`debug_assert!`).
 pub fn convert_utf8_bytes_into_utf16(bytes: &[u8]) -> UTF16Replacement {
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/bun_core/string/immutable/unicode.rs` around lines 413 - 424, Add a
concise doc comment to convert_utf8_bytes_into_utf16 documenting that bytes must
be non-empty and begin with a non-ASCII UTF-8 byte, since these are required
preconditions for callers across the crate boundary.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@test/js/bun/util/text-loader-invalid-utf8.test.ts`:
- Around line 4-7: Update the invalid-UTF-8 test in entry.ts to invoke
Bun.file(f).text() and assert its decoded result matches the imported text,
alongside the existing fs.readFileSync comparison. Keep the header comment’s
stated parity coverage accurate.
- Around line 47-50: Update the test assertion around JSON.parse(stdout) to
check exitCode and stderr first, or guard JSON parsing so subprocess crashes
surface the captured stderr diagnostics. Preserve the existing parsed JSON
assertions for successful execution and continue requiring an empty stderr and
exit code 0.
- Around line 1-53: Move the UTF-8 decoding cases and assertions from the
standalone suite into the existing text-loader suite in text-loader.test.ts,
reusing its setup and helpers where applicable. Keep a separate file only if one
case is required as a tracked regression; otherwise remove the standalone test
file while preserving coverage for all listed invalid and valid sequences.

---

Outside diff comments:
In `@src/bun_core/string/immutable/unicode.rs`:
- Around line 413-424: Add a concise doc comment to
convert_utf8_bytes_into_utf16 documenting that bytes must be non-empty and begin
with a non-ASCII UTF-8 byte, since these are required preconditions for callers
across the crate boundary.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: f1b1f6a8-d77c-4495-89c5-73ee4a7a2615

📥 Commits

Reviewing files that changed from the base of the PR and between 892b1da and c9aea5e.

📒 Files selected for processing (4)
  • src/bun_core/string/immutable.rs
  • src/bun_core/string/immutable/unicode.rs
  • src/js_printer/lib.rs
  • test/js/bun/util/text-loader-invalid-utf8.test.ts

Comment thread test/js/bun/util/text-loader-invalid-utf8.test.ts
Comment thread test/js/bun/util/text-loader-invalid-utf8.test.ts Outdated
Comment thread test/js/bun/util/text-loader-invalid-utf8.test.ts
…ile().text() parity and assert stderr before JSON.parse
@robobun

robobun commented Jul 23, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 2:51 PM PT - Jul 23rd, 2026

❌ @robobun, your commit ee9d06e has 1 failures in Build #78861 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 35306

That installs a local version of the PR into your bun-35306 executable, so you can run:

bun-35306 --bun

Comment thread test/js/bun/util/text-loader-invalid-utf8.test.ts
Comment thread src/js_printer/lib.rs
…pre_quoted_string

The bun_core copy (used by format_escapes, quote_for_json, snapshot keys,
sourcemap JSON) is documented as 'Same algorithm as
bun_js_printer::write_pre_quoted_string'; keep that true by giving its
Utf8 arm the same convert_utf8_bytes_into_utf16 replacement decode.
Comment thread src/js_printer/lib.rs
…UTF-8

The Encoding::Utf8 path now does a strict WHATWG decode, so the
WTF-8-shaped is_valid_wtf8 guard (which accepted ED A0 80..ED BF BF)
documented an invariant the decoder no longer honours. Replace it with
strings::is_valid_utf8 and note that raw text-loader bytes reach the
printer via print_string_characters_e_string without this assert.
Comment thread src/js_printer/lib.rs Outdated
Comment thread test/js/bun/util/text-loader-invalid-utf8.test.ts
Comment thread test/js/bun/util/text-loader-invalid-utf8.test.ts Outdated
…r comment

Add itBundled cases for target=bun (ascii-only output, \uFFFD arm) and
target=browser (raw EF BF BD arm) so both replacement.fail branches and
the bun build entry point for #12981 have fail-before coverage. Drop the
bug-history sentence from the runtime test header and replace it with the
issue URL.
@robobun

robobun commented Jul 23, 2026

Copy link
Copy Markdown
Collaborator Author

CI on build 78861 is green for this diff. The new tests (text-loader-invalid-utf8.test.ts and the two bun/loader-text-file-invalid-utf8-* bundler cases) pass on every lane.

The one red lane is test/js/node/test/parallel/test-https-server-connections-checking-leak.js on debian 13 aarch64 with panic: Failed to start File Watcher: EAGAIN, which also hit debian 13 x64 on the previous run (build 78847). That is inotify resource exhaustion on the runner and unrelated to this change (which only touches the UTF-8 string-quoting path in the printer). The other failure (watch-many-dirs.test.ts) passed on retry.

Ready for review/merge.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

All five earlier inline comments have been addressed and no new issues were found. Deferring to a human for final sign-off since this reworks the Encoding::Utf8 arm of write_pre_quoted_string_inner (the printer hot path shared by the runtime module loader and bun build) and its bun_core::printer sibling, and tightens the print_string_literal_utf8 debug_assert from WTF-8 to strict UTF-8.

What was reviewed:

  • Traced convert_utf8_bytes_into_utf16 through valid/invalid/overlong/surrogate/truncated inputs — replacement.len never exceeds remaining bytes, so clamped_width and the raw-byte fast path at lib.rs:824 stay in-bounds.
  • Confirmed the bun_core::printer::write_pre_quoted_string copy now matches (used by quote_for_json, sourcemap chunk, snapshot keys, format_escapes).
  • Checked test coverage: 10-case parity test against TextDecoder/readFileSync/Bun.file().text(), plus itBundled for both target: bun (ascii_only) and target: browser (raw EF BF BD) arms.
Extended reasoning...

Overview

Switches the Encoding::Utf8 arm of both bun_js_printer::write_pre_quoted_string_inner and bun_core::printer::write_pre_quoted_string from the WTF-8 stepper (wtf8_byte_sequence_length_with_invalid + decode_wtf8_rune_t) to strings::convert_utf8_bytes_into_utf16 (the WHATWG maximal-subpart decoder already used by to_utf16_alloc / Bun.file().text()). Ill-formed sequences now emit U+FFFD (as \uFFFD in ascii-only output or raw EF BF BD otherwise) instead of Latin-1 widening / NUL-with-over-advance / passing WTF-8 surrogates through. Also promotes convert_utf8_bytes_into_utf16 from pub(super) to pub, replaces the hand-rolled is_valid_wtf8 guard on print_string_literal_utf8 with strings::is_valid_utf8, and adds runtime + bundler tests.

Security risks

None identified. The change moves from a permissive decoder (which could emit raw invalid bytes into bundled JS, causing browser SyntaxError) to a strict WHATWG replacement decoder. No new untrusted-input parsing; the unreachable!() on empty bytes in convert_utf8_bytes_into_utf16 is guarded by the first < 0x80 check at both call sites (i < n ⇒ text[i..] non-empty).

Level of scrutiny

Medium-high. write_pre_quoted_string_inner runs for every string literal the printer emits (runtime module loading + bun build), and the bun_core sibling backs sourcemap JSON, snapshot keys, and format_escapes. The tightened debug_assert!(strings::is_valid_utf8(str)) on print_string_literal_utf8 is a new invariant that could fire on a caller not audited here — the author traced ~18 call sites and ran the transpiler/bundler_string/bundler_loader suites under debug, but this is the kind of change a maintainer should confirm.

Other factors

Every prior review round was addressed: the sibling copy was fixed (a571bfb), the debug_assert tightened (8140ce3), bug-history comments trimmed (5853f83, ee9d06e), and itBundled coverage added for both ascii_only arms (ee9d06e). Fail-before/pass-after was demonstrated under debug+ASAN. I verified replacement.len from convert_utf8_bytes_into_utf16_with_length is always ≤ the actual remaining bytes for non-fail results (padded zeros fail the continuation check first), so the downstream &text[i..i + clamped_width] slice at lib.rs:824 stays in-bounds. The separate-test-file placement was justified (existing text-loader.test.ts has an unrelated timeout under debug+ASAN).

@robobun

robobun commented Aug 13, 2026 •

Copy link
Copy Markdown
Collaborator Author

#38253 fixes the same text-loader symptom at the loader instead of in the printer: the text and md loaders validate the file with simdutf and, only when it is ill-formed, re-encode it through the same decoder before building the E::String.

The reason for the different layer: the 8-bit E::String payload is WTF-8 by contract. The JSON5 and JSON parsers encode a lone surrogate from an escape like "\ud800" with encode_wtf8_rune into an 8-bit string, and bun build currently prints it back as \uD800 (checked on main with a .json5 and a .json file). Replacing failed sequences inside write_pre_quoted_string_inner would turn that into \uFFFD\uFFFD\uFFFD, since convert_utf8_bytes_into_utf16 rejects ED A0..BF. Decoding at the loader leaves the printer's WTF-8 round-trip intact. Whichever approach is preferred, only one of the two PRs should land.

@robobun

robobun commented Sep 7, 2026

Copy link
Copy Markdown
Collaborator Author

Closing in favor of #41789. It changes the same arm of write_pre_quoted_string_inner but keeps decode_wtf8_rune_t, so a WTF-8 surrogate from a JSON "\ud800" escape still prints as \uD800, and it matches the bun_core copy of the loop that #40718 fixed. This branch also conflicts with main now.

@robobun robobun closed this Sep 7, 2026
@robobun

robobun commented Sep 7, 2026

Copy link
Copy Markdown
Collaborator Author

A correction to my comment above, with more data. The printer sink has more producers than the text and md loaders, so a printer-level repair is worth having next to #38253 rather than instead of it:

  • bun build of a .json, .jsonc, .json5 or .yaml import whose strings are not valid UTF-8 prints the same "Jos\x00\x00z" for Latin-1 Jos\xe9 P\xe9rez (stock 1.4.x; the runtime import of the same files is fine because expr_to_js goes through create_utf8_for_js). The JSON parser stores the raw source slice for a string without escapes (parse_string_utf8_at, src/parsers/json_stage2.rs:508), so the unvalidated bytes reach write_pre_quoted_string_inner.
  • The package.json editors (bun pm pkg set, bun add / bun remove, bun pm version, bun init) rewrite a Latin-1 manifest the same way through print_json. print_string_characters_utf8 passes json = false, so the \x00 escape lands in the .json file and node / npm then refuse to parse it.

For those doors the printer is the one shared point, but the repair has to stay WTF-8 aware: replace invalid lead bytes, stray continuation bytes, overlong and truncated sequences and anything above U+10FFFF with one U+FFFD per maximal subpart, and keep accepting ED A0..BF xx so an escaped lone surrogate from JSON / JSON5 ("a\ud800b" in a package.json currently round-trips as \uD800 through bun pm pkg set) is not turned into three U+FFFD. With that adjustment the two PRs are complementary: #38253 gives text and md imports exact TextDecoder parity (including U+FFFD x3 for a surrogate-looking byte run in a file, which a WTF-8 aware printer cannot do by design) and keeps invalid bytes out of the AST, and this PR closes the sink for every other producer.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Uncaught SyntaxError when importing certain PDFs via text loader

2 participants