Skip to content

sourcemap: count an ill-formed UTF-8 lead byte as one byte so the line break after it is not skipped - #38593

Open
robobun wants to merge 4 commits into
mainfrom
farm/4b7abeb3/sourcemap-invalid-utf8-lead-byte
Open

robobun wants to merge 4 commits into
mainfrom
farm/4b7abeb3/sourcemap-invalid-utf8-lead-byte

Conversation

@robobun

@robobun robobun commented Aug 14, 2026

Copy link
Copy Markdown
Collaborator

Problem

  • A byte that looks like the start of a multi-byte UTF-8 sequence but is not followed by continuation bytes (a Latin-1 encoded license header is the usual way to get one) makes the source map line counters skip the bytes after it. When one of those bytes is the line break, every mapping from there to the end of the file is one line off. It shows up in bun build --sourcemap output and in the line numbers bun run puts in error.stack.
  • Repro: printf '/*! caf\xe9\nb */\nconsole.log(1);\nconsole.log(2);\n' > in.js && bun build ./in.js --outdir out --sourcemap=external. The output has console.log(1) on its 4th line, but mappings is ;AACA;AAAA,QAAQ,IAAI,CAAC;AACb,... (5 line groups, the statement in the 3rd), and the original line recorded for it is line 1 instead of line 2. With a well-formed é the same file produces ;AAEA;AAAA;AAAA,QAAQ,... (6 groups, original line 2). At runtime, a file starting with /* caf\xe9\nb */ reports an error created on line 3 at line 2.
  • Cause: both loops advance by the width the lead byte declares (wtf8_byte_sequence_length_with_invalid, 3 for 0xE9) whether or not the decode succeeded. update_generated_line_and_column_slow in src/sourcemap/Chunk.rs did i += len after decode_wtf8_rune_t had returned the replacement character, and LineOffsetTable::generate_in in src/sourcemap/LineOffsetTable.rs did remaining = &remaining[cp_len..] the same way, so the \n inside the supposed 3-byte sequence never reaches the line terminator arm of either match.
  • The bundler chains the per-file map chunks (each file's mappings start where the previous file's line count left off), so a file that counts one of its own line breaks short also pulls every file printed after it up a line.

Fix

  • Both loops decode through strings::CodepointIterator instead of their own wtf8_byte_sequence_length_with_invalid + decode_wtf8_rune_t copies, and advance by the width it reports. An ill-formed sequence comes back as one U+FFFD, one byte wide, so the line break is looked at on its own on the next iteration. Well-formed input decodes exactly as before.
  • Why this is the right count: the line break after an ill-formed byte is a line break to everything that reads the file. JSC gets the transpiled code through String::fromUTF8ReplacingInvalidSequences, which turns the bad byte into U+FFFD and keeps the \n, so the positions it reports are what these tables have to line up with; esbuild, where both loops come from, walks the bytes with Go's range decoding, which also steps one byte past an invalid one. CodepointIterator is also what LineColumnOffset::advance already uses, and that is the third walk feeding the same line separators (the glue between files in a bundle, and the positions of the placeholder substitutions in src/bundler/Chunk.rs), so the three walks now count the same bytes the same way by construction. Open PR lexer: decode ill-formed UTF-8 in JS source as U+FFFD #38262 refines CodepointIterator's ill-formed handling to one U+FFFD per maximal subpart; these loops pick that up as is, and since a line break is never part of such a subpart the line count is the same either way.
  • The .min(remaining.len()) clamp from Don't panic generating a sourcemap for a source ending in a truncated UTF-8 sequence #32774 lives inside the iterator now; the six truncated-tail cases it added still pass.
  • Tests, all of which fail on the current build:
    • test/js/bun/sourcemap/internal-sourcemap-roundtrip.test.ts: bun build --sourcemap=external on Latin-1 files with an ill-formed sequence (a bare 0xE9, a bare 2-byte lead, three bytes of a 4-byte sequence) right before a line break, once in a legal comment (kept in the output, so both loops see it) and once in a plain comment (dropped, so only the original-side table does), decoding the mappings and checking the line each following statement maps back to. The source has blank lines between the statements that the output does not, so the two off-by-ones cannot cancel out. One more case puts the comment in a dependency and checks the entry point's statement is still placed right.
    • test/js/bun/sourcemap/internal-sourcemap.test.ts: the same two comment shapes through bun run and error.stack (unfixed they report line 2 and line 5 for an error on line 3), plus an ill-formed byte inside a template literal with, on the error's line, a string the printer re-emits with a well-formed U+FFFD, which still has to count as one column (3:26; unfixed 2:26).
  • Also ran against the debug build: the rest of test/js/bun/sourcemap/, bundler_comments, the source map tests in bundler_edgecase, the bun build --compile source map tests, test/js/node/module/*sourcemap*, test/cli/test/coverage.test.ts, and test/bake/dev/{sourcemap,server-sourcemap}.test.ts.
  • sourcemap: fix the CRLF peek when counting generated lines #38449 fixes a separate bug (the CRLF peek) in the same generated-side loop; the two changes are on different lines and independent.

Background

  • A source map entry pairs a position in the printed output (generated line and column) with a position in the source file (original line and column). Bun builds the two sides separately. LineOffsetTable::generate walks the source file once and records where every line starts, plus a byte-to-UTF-16-column table for lines containing non-ASCII bytes; add_source_mapping turns a token's byte offset into an original line and column with it. The builder in Chunk.rs is called once per mapped token and walks whatever the printer emitted since the previous call, bumping the generated line at each line terminator. The same builder produces bun build's VLQ maps and the compact maps the runtime uses to rewrite stack traces.
  • wtf8_byte_sequence_length_with_invalid(b) returns the number of bytes a UTF-8 sequence starting with b would have (2 for 0xC0..0xDF, 3 for 0xE0..0xEF, 4 for 0xF0..0xF7), judging by the lead byte alone. decode_wtf8_rune_t then checks that the following bytes are continuation bytes and returns a caller-supplied sentinel when they are not. The bug was advancing by the first number after the second call had said the sequence was not there.
  • strings::CodepointIterator (src/bun_core/string/immutable.rs) is the shared way of walking raw bytes as code points: it returns the code point and the number of bytes it took up, reporting an ill-formed sequence as U+FFFD with a width of 1, so a caller never steps over bytes that were not part of a valid sequence.
  • Ill-formed bytes get into the printed output through legal comments (/*! ... */, copied verbatim, and the only place one can sit right before a line break) and regex literals (which cannot contain a line break); strings and template literals are re-encoded by the printer and other comments are dropped. They get into the source-side table from anywhere in the file.
Repro output before and after
$ printf '/*! caf\xe9\nb */\nconsole.log(1);\nconsole.log(2);\n' > in.js
$ bun build ./in.js --outdir out --sourcemap=external
$ cat -A out/in.js
// in.js$
/*! cafM-i$
b */$
console.log(1);$
console.log(2);$

# mappings before: 5 groups, console.log(1) in the 3rd although it is on the 4th line,
# recorded as original line 1 although it is on line 2
;AACA;AAAA,QAAQ,IAAI,CAAC;AACb,QAAQ,IAAI,CAAC;
# mappings after: identical to what the same file with a well-formed é produces
;AAEA;AAAA;AAAA,QAAQ,IAAI,CAAC;AACb,QAAQ,IAAI,CAAC;

# Plain comment (dropped from the output): only the source-side table is affected.
$ printf '/* caf\xe9\nb */\nconst err = new Error("x");\nconsole.log(err.stack.split(String.fromCharCode(10))[1]);\n' > rt.js
$ bun rt.js
    at /tmp/rt.js:2:17     # before (the Error is on line 3)
    at /tmp/rt.js:3:17     # after

# Legal comment (kept in the output): both sides are affected. The blank lines
# keep the two off-by-ones from cancelling out.
$ printf '/*! caf\xe9\nb */\nconst err = new Error("x");\n\n\nconsole.log(err.stack.split(String.fromCharCode(10))[1]);\n' > rt2.js
$ bun rt2.js
    at /tmp/rt2.js:5:17    # before (the Error is on line 3)
    at /tmp/rt2.js:3:17    # after

LineOffsetTable::generate (original side) and the builder's
update_generated_line_and_column (generated side) advanced by the width a
lead byte declares even when the decode failed, so a Latin-1 byte right
before a line break swallowed the line break and every mapping after it
was one line off. Both loops now decode through strings::CodepointIterator,
which yields one U+FFFD one byte wide for an ill-formed sequence, the same
as LineColumnOffset::advance.
@coderabbitai

coderabbitai Bot commented Aug 14, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 1 minute

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 1660fd72-1ce7-4299-9fdd-e09140c769bb

📥 Commits

Reviewing files that changed from the base of the PR and between eabb96d and a492bab.

📒 Files selected for processing (4)
  • src/sourcemap/Chunk.rs
  • src/sourcemap/LineOffsetTable.rs
  • test/js/bun/sourcemap/internal-sourcemap-roundtrip.test.ts
  • test/js/bun/sourcemap/internal-sourcemap.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Aug 14, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: fix and tests pushed, review feedback addressed, waiting on CI.

  • Reproduced on the released 1.4.0 build: bun build --sourcemap=external on a file whose legal comment has a Latin-1 0xE9 right before a line break emits 5 line groups instead of 6, and bun run on a file starting with /* caf\xe9\nb */ reports an error created on line 3 at line 2 (commands and output in the PR description).
  • The new cases in test/js/bun/sourcemap/internal-sourcemap-roundtrip.test.ts (7) and test/js/bun/sourcemap/internal-sourcemap.test.ts (3) all fail on the released build and pass with this branch's debug build; the rest of test/js/bun/sourcemap/ and the related bundler, node:module, coverage and dev server source map suites pass as well.
  • Review follow-ups: the bun build helper in the roundtrip test drains stdout too (34634fc), and the two comments added in src/sourcemap/ are down to one line each (a492bab). No code change since the first push.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Beyond the inline nit, I also checked: CodepointIterator::next never sets cursor.width = 0 (wtf8_byte_sequence_length returns ≥ 1 for every byte and the ASCII fast path sets 1), so neither loop can stall; the fresh-Cursor-per-iteration pattern with width defaulted to 0 reads at exactly i/offset as intended; and the sentinel change in LineOffsetTable (invalid decode was 0, now U+FFFD) only adds columns_for_non_ascii entries for lines with ill-formed bytes — which is a correction, not a regression.

Extended reasoning...

The fix replaces two hand-rolled WTF-8 decode loops with the shared strings::CodepointIterator, which is exactly what REVIEW.md's "grep for the in-tree helper" rule asks for and what LineColumnOffset::advance already does. I traced next() in src/bun_core/string/immutable.rs and wtf8_byte_sequence_length / decode_wtf8_rune_t to confirm width is always ≥ 1 (so no infinite-loop risk from the while i < n / while !remaining.is_empty() heads), that continuation bytes 0x80-0xBF and 0xF8-0xFF decode identically to before (len==1 → T::from(p[0])), and that well-formed multibyte sequences still return the same codepoint and width. The u32 casts on i/offset are pre-existing constraints (both loops already stored offsets as u32). The one behavioral side effect beyond the fix itself — LineOffsetTable now sees c = 0xFFFD instead of c = 0 for a bad decode, so c > 0x7F becomes true and the line's columns_for_non_ascii table is populated — makes columns after an ill-formed byte match what JSC actually sees, so it's a secondary correctness improvement. This is a well-tested, correctly-layered fix, but it rewrites decode logic in two hot builders and interacts with two other open PRs (#38262, #38449), so a human look is still worthwhile.

Comment thread test/js/bun/sourcemap/internal-sourcemap-roundtrip.test.ts Outdated
Comment thread src/sourcemap/Chunk.rs Outdated
Comment on lines +537 to +540
// Same decoder as `LineOffsetTable::generate` and `LineColumnOffset::advance`:
// a lead byte whose continuation bytes are missing (a Latin-1 byte in a legal
// comment) is one U+FFFD, one byte wide, so a line terminator right after it is
// still seen below instead of being skipped with the width the lead byte declared.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you need a paragraph-long comment to justify why the workaround is OK, the code is wrong — fix the code

Comment thread src/sourcemap/LineOffsetTable.rs Outdated
Comment on lines +189 to +194
// `CodepointIterator` is also what `LineColumnOffset::advance` and the
// generated-side walk in `Chunk.rs` use, so all three agree on what a byte
// sequence is: a lead byte whose continuation bytes are missing (Latin-1 text)
// is one U+FFFD, one byte wide, and the bytes after it (a line terminator, say)
// are looked at on their own rather than skipped as part of the width the lead
// byte declared.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you need a paragraph-long comment to justify why the workaround is OK, the code is wrong — fix the code

Comment thread src/sourcemap/Chunk.rs Outdated
Comment on lines +537 to +538
// An ill-formed sequence decodes as one U+FFFD one byte wide, so a line
// terminator right after a bad lead byte still reaches the match below.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you need a paragraph-long comment to justify why the workaround is OK, the code is wrong — fix the code

Comment thread src/sourcemap/LineOffsetTable.rs Outdated
Comment on lines +189 to +190
// Same decoder as the generated-side walk in `Chunk.rs`: an ill-formed sequence
// is one U+FFFD one byte wide, so a line terminator right after it is not skipped.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If you need a paragraph-long comment to justify why the workaround is OK, the code is wrong — fix the code

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — the earlier stdout-drain nit and the comment-cop length complaints have both been addressed, and no new issues were found on re-review.

Checked that CodepointIterator::next (src/bun_core/string/immutable.rs:106) sets width = 1 and c = U+FFFD on decode failure, so both loops now advance one byte and hit the following \n on the next iteration; the while i < n / !remaining.is_empty() guards make the discarded iter.next() return unreachable-false. The sentinel change in LineOffsetTable (0 → U+FFFD) now trips the c > 0x7F non-ASCII column-tracking arm for ill-formed bytes — traced that through the columns_for_non_ascii extend and add_source_mapping lookup and it produces the same column, so it's a no-op for correctness. The removed .min(remaining.len()) clamp is reproduced inside next() at line 126, and the six truncated-tail cases from #32774 are still in the file.

Extended reasoning...

Overview

Two source-map byte walkers — update_generated_line_and_column_slow in src/sourcemap/Chunk.rs and LineOffsetTable::generate_in in src/sourcemap/LineOffsetTable.rs — are switched from hand-rolled wtf8_byte_sequence_length_with_invalid + decode_wtf8_rune_t to the shared strings::CodepointIterator. The old code advanced by the lead byte's declared width whether or not the decode succeeded, so a Latin-1 byte like 0xE9 (declared width 3) right before \n swallowed the newline and shifted every mapping below by one line. Ten new tests cover both walks independently (legal comment kept in output, plain comment dropped), three ill-formed shapes (bare 2-/3-byte lead, truncated 4-byte), the multi-file chaining case, and the runtime error.stack path.

Security risks

None. Input is source bytes already being parsed; the change replaces one bounded decoder with another. The removed .min(remaining.len()) truncation clamp lives inside CodepointIterator::next (take = (len as usize).min(tail.len())), so no new out-of-bounds read is introduced.

Level of scrutiny

Moderate. The two touched functions feed every source map Bun emits (bundler and runtime stack traces), but the diff is ~30 native lines that swap a bespoke decode loop for the crate's shared iterator — exactly the "grep for the in-tree helper before hand-writing anything" pattern REVIEW.md asks for, and the same iterator LineColumnOffset::advance already uses for the third walk over the same bytes. I verified: the loop guard makes the discarded iter.next() bool unreachable; Cursor { i, ..Default } gives width = 0 so pos = i on entry; the ASCII fast path is preserved inside next()'s first < 0x80 branch (#[inline(always)], so it folds in as before). The LineOffsetTable sentinel changing from 0 to U+FFFD means an ill-formed byte now sets byte_offset_to_first_non_ascii, but since width = 1 and it counts one column, the recorded per-byte columns match what the old byte-offset path would have produced — verified by walking the columns_for_non_ascii extend arithmetic for a single 0xE9 at offset 5.

Other factors

My prior nit (undrained stdout in the bundle helper) was fixed in 34634fc and is visible in the current diff. The comment-cop length flags on both .rs files were addressed in 92d29b1/a492bab — the added comments are now one line each. Test coverage is thorough per REVIEW.md's variant-matrix rule (three ill-formed shapes × two comment kinds, plus the dependency-chaining and template-literal/U+FFFD-column cases), the fixtures use blank lines so the two off-by-ones can't cancel, and the PR notes the tests fail on the released build. _slow is #[cold] and generate_in runs once per source, so the fresh-Cursor-per-iteration pattern (vs. reusing one) is not on a hot path and the 8-byte struct is trivially SROA'd.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant