Skip to content

Fix panic when bundling CSS that contains invalid UTF-8 - #32795

Closed
robobun wants to merge 8 commits into
mainfrom
farm/5cc40f1a/css-invalid-utf8
Closed

robobun wants to merge 8 commits into
mainfrom
farm/5cc40f1a/css-invalid-utf8

Conversation

@robobun

@robobun robobun commented Jun 26, 2026 •

Copy link
Copy Markdown
Collaborator

What

bun build (CLI and Bun.build) crashes with panic: unreachable on a CSS file containing a single invalid UTF-8 byte inside an @import:

printf '@import url("./x\xe2y.css");' > x.css
bun build x.css --outdir=/tmp/out
# panic: unreachable

Same crash from Bun.build({ entrypoints: ["x.css"], throw: false }), from a JS or HTML entry that imports the CSS file, and from @-x url(a\xE2b) tok; (a url() token in an at-rule prelude). A file like this comes from a truncated or wrong-encoding asset; the build should report an error, not abort.

Top of the stack:

core::option::expect_failed
BabyString::in                     src/ast/lib.rs
Log::add_resolve_error_with_level  src/ast/lib.rs
BundleV2::resolve_import_records   src/bundler/bundle_v2.rs

Why

The CSS tokenizer takes &[u8] and hands out raw sub-slices of the source as token values (upstream rust-cssparser takes &str, so it never sees ill-formed input). The raw specifier bytes land in an ImportRecord, resolution fails, and the resolve error formatter renders the message with bstr::BStr's lossy Display (invalid bytes become U+FFFD) while BabyString::in searches for the original raw bytes inside that lossy text and expect()s the hit.

The JS lexer, the HTML loader, and the runtime resolver already sanitize their specifiers to U+FFFD. CSS was the only producer of raw bytes.

Fix

  • Add bun_core::strings::replace_invalid_utf8(bytes, arena): SIMD-validate the source (simdutf) and, only when it is invalid, copy it into the parse arena with each ill-formed sequence replaced by U+FFFD, returning the arena-lifetime slice. This is the decode step from css-syntax 3.2 and matches what a browser does with the same stylesheet.
  • Call it at both places CSS source enters the parser: the bundler's ParseTask (before the Source is built, so source.contents, token positions, error line text, and source maps all index the same buffer) and the non-bundling transpile path.
  • BabyString::in degrades to an empty specifier instead of panicking when the needle is not found. After the loader fix every current producer is valid UTF-8, so this is hardening of the error path rather than the fix itself; a miss there should never abort a build.

Tests

test/bundler/css/invalid-utf8.test.ts, five cases: the CLI repro, the Bun.build API form, the at-rule url() form, an invalid byte outside any import (asserting the emitted CSS is well-formed UTF-8, an input that also used to crash), and an escaped invalid byte (the escape decoder used to advance by the encoded width of U+FFFD and silently drop the two source bytes after it, changing selectors and eating closing quotes). All five fail on the current build and pass with the fix.

Not changed here: ResolveMessage's JS getters decode non-ASCII message text as Latin-1 (log.specifier is "./café.js" for import "./café.js" today, no invalid UTF-8 involved). That is a separate pre-existing issue, so the API test asserts the ASCII affixes of the specifier.

@coderabbitai

coderabbitai Bot commented Jun 26, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 20bb8552-3876-4ccc-a645-5a862be63073

📥 Commits

Reviewing files that changed from the base of the PR and between 038075c and 9789336.

📒 Files selected for processing (2)
  • src/bun_core/string/immutable.rs
  • test/bundler/css/invalid-utf8.test.ts

Walkthrough

Adds replace_invalid_utf8 to sanitize CSS byte slices with U+FFFD replacement, wires it into two bundler CSS parsing paths, fixes a missing-substring fallback in BabyString::in, and adds tests for malformed UTF-8 in CSS imports, url() tokens, and stylesheet content.

Changes

CSS Invalid UTF-8 Handling

Layer / File(s) Summary
replace_invalid_utf8 implementation and export
src/bun_core/string/immutable.rs, src/bun_core/lib.rs
Adds replace_invalid_utf8 using utf8_chunks() and arena allocation; re-exports it via bun_core::strings.
CSS sanitization in ParseTask and transpiler
src/bundler/ParseTask.rs, src/bundler/transpiler.rs
ParseTask conditionally replaces invalid UTF-8 for CSS before building Source; build_css_output sanitizes bytes before StyleSheet::parse.
BabyString::in panic fix
src/ast/lib.rs
Replaces expect("unreachable") with a let Some(...) else branch that returns BabyString::new(0, 0) when index_of returns None.
Invalid UTF-8 CSS bundler tests
test/bundler/css/invalid-utf8.test.ts
Adds concurrent tests for malformed UTF-8 in @import specifiers, url() tokens, general stylesheet content, and escaped bytes; asserts U+FFFD replacement and correct error/success outcomes.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed Clear and specific: it summarizes the main fix, the CSS invalid UTF-8 panic.
Description check ✅ Passed The PR explains the bug, cause, fix, and tests, though it uses custom headings instead of the template's exact section names.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Jun 26, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 6:26 AM PT - Jun 29th, 2026

❌ @robobun, your commit 9789336 has 1 failures in Build #66729 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 32795

That installs a local version of the PR into your bun-32795 executable, so you can run:

bun-32795 --bun

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't find any issues — the fix is well-scoped and thoroughly tested — but it touches the bundler's core parse path and changes BabyString::in's panic semantics, so it's worth a human look.

Extended reasoning...

Overview

This PR fixes a panic: unreachable when bundling CSS containing invalid UTF-8 bytes. It adds bun_css::replace_invalid_utf8 (SIMD-validated fast path, arena-allocated U+FFFD substitution on the cold path) and calls it at the two CSS entry points (ParseTask::parse_worker and Transpiler::parse). It also hardens BabyString::in to return an empty span instead of panicking when the needle isn't found in the formatted parent message. Four new tests cover CLI, Bun.build, at-rule url(), and non-import invalid bytes.

Security risks

None identified. This is input sanitization that brings CSS in line with the JS/HTML loaders, which already replace invalid UTF-8. The change reduces attack surface by eliminating a panic on malformed input rather than introducing one.

Level of scrutiny

Moderate-to-high. The logic itself is small and the common-case path is zero-cost (is_valid_utf8 returns the original slice unchanged), but the call sites are in the bundler's per-file parse pipeline (ParseTask.rs ~line 2300, transpiler.rs), which is production-critical and runs on every CSS file. The BabyString::in change also affects error reporting for all loaders, not just CSS — turning an expect("unreachable") into a silent empty-span fallback is a deliberate trade-off (don't abort a build while reporting an error) that a maintainer should ratify.

Other factors

  • The PR description is unusually thorough, with stack trace, root-cause analysis, and explicit reasoning for each change.
  • The sanitization is correctly placed before source is built in ParseTask, so token positions / error line text / source maps all index the same buffer — this is the subtle part and it looks right.
  • bump (the per-task arena) is in scope at the ParseTask call site and outlives the parse, so the arena allocation in the cold path is sound.
  • No CODEOWNERS match the touched files.
  • The bug-hunting system found no issues.

Deferring because the bundler parse path and a cross-loader error-reporting helper are core enough that a human familiar with the arena/lifetime story should confirm.

@robobun

robobun commented Jun 27, 2026 •

Copy link
Copy Markdown
Collaborator Author

CI status for the current head (9789336, build 66729): 282 jobs passed, the Windows lanes are green, and every lane that runs the new tests (including the ASAN shards) passes. The red lanes are unrelated to this diff:

  • darwin 26 aarch64: the agent's artifact download of the built binary timed out after 120s before any test ran, on both attempts. The same agent pool has done this on every build of this branch; the job needs a Buildkite-side retry (my token cannot retry a single job).
  • alpine 3.23 x64 and x64-baseline: test/js/node/test/parallel/test-net-connect-memleak.js, a GC-sensitive node:net port that did not fail on the previous build of this same branch and has no overlap with this change (CSS source decoding and the resolve error formatter).

The diff is limited to the CSS loader path, bun_core::strings::replace_invalid_utf8, and the resolve error formatter.

@Jarred-Sumner Jarred-Sumner left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It needs to use the arena lifetime.

@robobun

robobun commented Jun 27, 2026

Copy link
Copy Markdown
Collaborator Author

Fixed in 2d63fb9: the replacement buffer is now built in the parse arena (ArenaVec::with_capacity_in + leak(), which returns the arena-lifetime slice) instead of filling a heap Vec and copying it into the arena afterwards. The valid-UTF-8 fast path still returns the input slice unchanged.

@robobun

robobun commented Jun 27, 2026

Copy link
Copy Markdown
Collaborator Author

The missing decode step has a second, uglier symptom than the panic: silent output corruption through the escape decoder. This PR already fixes it. Writing it up here so the case gets pinned by a test.

The escape decoder eats source bytes after an ill-formed sequence

Tokenizer::consume_char advances the cursor by len_utf8(decoded). When the byte after a \ is not well-formed UTF-8, next_char returns U+FFFD, so the cursor advances by len_utf8(U+FFFD) = 3 instead of the 1 byte of the ill-formed sequence it replaced, and the 2 source bytes that follow are dropped. Upstream rust-cssparser never hits this because its input is &str; the &[u8] port kept the arithmetic without that precondition. Idents, quoted strings, and unquoted url() tokens all route escapes through it.

On current main, all three build with exit 0 and no diagnostics:

printf '.a { content: "\\\xc3x AFTER"; }' > in.css && bun build in.css
#   .a { content: "\uFFFDAFTER"; }                 the `x ` after the escape is gone

printf '.a { content: "\\\xc3"; color: green; }' > in.css && bun build in.css
#   .a { content: "\uFFFD color: green; }"; }      the closing quote was eaten, so the
#                                                  rest of the rule is now string text

printf '.x\\\xc3yz { color: red; }' > in.css && bun build in.css
#   .x\uFFFD { color: red; }                       the class name lost its `yz`

The last two change which rules exist and which selectors they target, with no warning.

Why this PR covers it

replace_invalid_utf8 turns the lone ill-formed byte into a real 3-byte U+FFFD before tokenizing, so the decoder's 3-byte advance is then correct. I checked by feeding a stock build the pre-decoded bytes (byte for byte what the tokenizer sees after this change): all three come out right (content: "\uFFFDx AFTER", then content: "\uFFFD"; color: green;, then .x\uFFFDyz).

I also looked at whether the tokenizer needs its own fix on top of this (advance by the source length rather than by len_utf8 of the decoded value). It does not: after this change I could not find a path that reaches the tokenizer with ill-formed bytes. Raw file contents, plugin onLoad results, and data: URLs all funnel through the two call sites patched here, and the remaining Parser constructors (the cssInternals test hooks and Bun.color) take JS strings, which String::to_utf8 has already made well-formed. Decoding once at the boundary is the right shape; a tokenizer-side change would be unreachable code.

Suggested test

The escape case is the one shape where the tokenizer decodes the bad bytes instead of slicing past them, and the only one that corrupts structure rather than content, so it is worth a case in test/bundler/css/invalid-utf8.test.ts:

test.concurrent("escaped invalid byte does not swallow the bytes after it", async () => {
  using dir = tempDir("css-invalid-utf8-escape", {});
  // `\` followed by a lone 0xC3 lead byte. The escape must decode to U+FFFD and
  // consume exactly that byte; it used to consume three (the UTF-8 length of
  // U+FFFD), eating ident characters, string content, or the closing quote.
  writeFileSync(
    join(String(dir), "in.css"),
    Buffer.concat([
      Buffer.from(".x\\"),
      Buffer.from([0xc3]),
      Buffer.from('yz { content: "\\'),
      Buffer.from([0xc3]),
      Buffer.from('x AFTER"; }\n.a { content: "\\'),
      Buffer.from([0xc3]),
      Buffer.from('"; color: green; }\n'),
    ]),
  );

  await using proc = Bun.spawn({
    cmd: [bunExe(), "build", "./in.css", "--outdir=out"],
    env: bunEnv,
    cwd: String(dir),
    stdout: "pipe",
    stderr: "pipe",
  });
  const [stderr, exitCode] = await Promise.all([proc.stderr.text(), proc.exited]);
  expect(stderr).not.toContain("error:");
  expect(exitCode).toBe(0);

  const out = await Bun.file(join(String(dir), "out", "in.css")).text();
  // Selector: `yz` used to be eaten out of the class name.
  expect(out).toContain(".x\uFFFDyz");
  // String content: the `x ` after the escape used to be eaten.
  expect(out).toContain('content: "\uFFFDx AFTER"');
  // The closing quote used to be eaten, absorbing `; color: green; }` into the string.
  expect(out).toContain('content: "\uFFFD";');
  expect(out).toContain("color: green;");
});

On main all four toContain assertions fail; with the input decoded they pass (verified against the pre-decoded equivalent as above).

Comment thread test/bundler/css/invalid-utf8.test.ts

@Jarred-Sumner Jarred-Sumner left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Use the existing code to do this string with replacement. Do not roll your own.

@robobun

robobun commented Jun 28, 2026

Copy link
Copy Markdown
Collaborator Author

Done in 2f5e509. The replacement now goes through the existing converters: strings::to_utf16_alloc(code, false, false), which substitutes U+FFFD for invalid sequences (the same call the TextDecoder and Blob.text paths use), then strings::to_utf8_alloc, with the result copied into the parse arena. No local decoding loop left.

Comment thread src/css/css_parser.rs Outdated
Comment thread test/bundler/css/invalid-utf8.test.ts
Comment thread test/bundler/css/invalid-utf8.test.ts Outdated
@robobun
robobun force-pushed the farm/5cc40f1a/css-invalid-utf8 branch from 8547dc7 to 9f14380 Compare June 28, 2026 07:27
Comment thread src/ast/lib.rs Outdated
robobun added 7 commits June 29, 2026 09:39
The CSS tokenizer hands out raw sub-slices of the source as token
values and assumes byte positions are char boundaries, but the source
bytes were never validated. An @import specifier containing an invalid
byte reached the bundler's resolve error formatter, which expects to
find the raw specifier inside the lossily formatted message text and
panicked (unreachable) when it was not there.

Decode CSS sources before parsing, replacing invalid sequences with
U+FFFD the way a browser decodes a stylesheet, and make the error
formatter degrade to an empty specifier instead of panicking.
Build the replacement buffer with ArenaVec::with_capacity_in and return
it with leak(), instead of filling a heap Vec and copying it into the
arena afterwards.
Use strings::to_utf16_alloc (which substitutes U+FFFD for invalid
sequences when fail_if_invalid is false, the same call TextDecoder and
Blob.text use) followed by strings::to_utf8_alloc, instead of a local
utf8_chunks loop.
strings::replace_invalid_utf8 takes the arena and returns an
arena-lifetime slice: the input when it is already valid, else an arena
copy with each ill-formed sequence replaced by U+FFFD. The CSS parse
sites call it directly.

Add an escape-decoder case to the test: an escaped ill-formed byte used
to advance the tokenizer by the encoded width of U+FFFD, dropping the
two source bytes after it.
@robobun
robobun force-pushed the farm/5cc40f1a/css-invalid-utf8 branch from 7d43702 to 038075c Compare June 29, 2026 09:43

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/bun_core/string/immutable.rs`:
- Around line 1566-1572: The allocation in the invalid UTF-8 repair path only
reserves space for one replacement marker, so repeated invalid chunks in the
utf8_chunks loop can force reallocations; update the sizing logic in the string
repair routine to compute the full repaired length up front and pass that exact
capacity to ArenaVec::with_capacity_in. Use the existing invalid-chunk scan in
the code path around bytes.utf8_chunks() and UNICODE_REPLACEMENT_STR to
determine the total number of replacements before building out.

In `@test/bundler/css/invalid-utf8.test.ts`:
- Around line 109-115: In the success-path assertions in invalid-utf8.test.ts,
move the proc.exited/exitCode check to the end so the emitted CSS/file content
assertions run first and surface the more specific diff on failure. Update the
affected test blocks around the stdout/stderr checks and Bun.file(out/in.css)
validation, and apply the same ordering to the other matching case noted in the
review so exitCode is always asserted last.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: ed721635-92f7-4dec-aed9-d10d69760b53

📥 Commits

Reviewing files that changed from the base of the PR and between bb32d24 and 038075c.

📒 Files selected for processing (6)
  • src/ast/lib.rs
  • src/bun_core/lib.rs
  • src/bun_core/string/immutable.rs
  • src/bundler/ParseTask.rs
  • src/bundler/transpiler.rs
  • test/bundler/css/invalid-utf8.test.ts

Comment thread src/bun_core/string/immutable.rs Outdated
Comment thread test/bundler/css/invalid-utf8.test.ts
@robobun

robobun commented Jun 30, 2026

Copy link
Copy Markdown
Collaborator Author

Two more reproducers for the missing decode step came out of fuzzing raw bytes into the CSS path. They hit a different symptom than the resolve-error panic this PR started from: on an asserts build, a UTF-8 continuation byte at a token start (lead bytes like the 0xE2/0xE9 used in this PR's tests happen to be tolerated by the tokenizer) aborts on the char-boundary assertions in Tokenizer::get_position and Parser::position:

panic: assertion failed: strings::is_on_char_boundary(self.src, self.position)
panic: assertion failed: strings::is_on_char_boundary(self.input.tokenizer.src, self.input.tokenizer.position)
printf '.a { color: red \xAF}\n' > bad1.css && bun build ./bad1.css
printf '@property --x { syntax: "<length>"; inherits: false; initial-value:\xBD1px; }\n' > bad2.css && bun build ./bad2.css

I verified both against this branch at 9789336: the first now builds with U+FFFD in the output and the second reports the normal Unexpected token diagnostic, so the decode at the two ingress points covers them. I had an equivalent fix in flight and am not opening a second PR for the same root cause. The only pieces of it this PR does not have, in case they are worth folding in, are on farm/d725d202/css-non-utf8-ingress:

  • tests for the two shapes above, plus one for bun build --no-bundle (the build_css_output path this PR also patches, which has no test in this PR yet)
  • an is_valid_utf8 check at the bun_css entry points (StyleSheet::parse_with, StyleAttribute::parse) that returns a new ParserError::invalid_utf8, so any future caller that skips the decode gets an error in release builds instead of tripping debug assertions or slicing at non-character boundaries

@robobun

robobun commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator Author

One more way to reach the char-boundary assertion this PR fixes, in case it is useful when this gets rebased: test/js/bun/css/css-fuzz.test.ts trips it on a debug build of current main (the file is skipped in CI, so it only shows up locally).

The fuzz inputs themselves are valid UTF-8. What happens is that the fuzz tests run past their timeouts and keep rewriting the shared invalid.css while a later iteration's build is reading it, so the read returns the old contents followed by the tail of the new contents, and the ".test{content:\"\uFFFD\uFFFD\"}" input gets cut in the middle of a 3-byte sequence. The resulting file reproduces deterministically with a debug build:

printf '.test{color:red;;;}\xbf\xbd"}' > x.css && bun build x.css
# panic: assertion failed: strings::is_on_char_boundary(self.src, self.position)
#   Tokenizer::get_position (css_parser.rs:4154) <- Parser::state <- StyleSheetParser::next

Same shape as the \xAF / \xBD cases above (a continuation byte where a token starts), so decoding the source at ingress covers it and nothing separate is needed in the tokenizer. The shared fixture file in the fuzz test is being handled on its own in #38491 and #38500.

@robobun

robobun commented Sep 7, 2026

Copy link
Copy Markdown
Collaborator Author

Superseded by #41801, which ports this change onto current main. This branch is about 1960 commits behind and no longer merges: UNICODE_REPLACEMENT_STR and the public ArenaVec::leak it used are gone, and bun pm diff added a third CSS parse site since. #41801 keeps the shape settled in review here (strings::replace_invalid_utf8(bytes, arena) in bun_core::strings, arena-lifetime result, no UTF-16 round trip), covers the output cases from the later comments (raw bytes in strings and idents, --minify, --no-bundle, the CSS chunk of an HTML entry), and adds a debug assertion in the tokenizer for the precondition.

@robobun robobun closed this Sep 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants