Skip to content

Don't panic generating a sourcemap for a source ending in a truncated UTF-8 sequence - #32774

Merged
Jarred-Sumner merged 2 commits into
mainfrom
farm/33e8ee25/fix-sourcemap-truncated-utf8
Jun 26, 2026
Merged

Jarred-Sumner merged 2 commits into
mainfrom
farm/33e8ee25/fix-sourcemap-truncated-utf8

Conversation

@robobun

@robobun robobun commented Jun 26, 2026

Copy link
Copy Markdown
Collaborator

Repro

A source file whose last bytes are a truncated multi-byte UTF-8 sequence (a lead byte that declares more bytes than the file has left) panics every sourcemap-printing path: bun build --sourcemap=inline|external, plain bun <file> (the runtime transpiler builds the same line table for stack-trace remapping), and --hot rebuilds.

printf 'console.log(1);//\xf0' > t.js
bun build --sourcemap=external --outdir=out t.js
panic: range start index 4 out of range for slice of length 1

at LineOffsetTable::generate, via Chunk::add_source_mapping / js_printer::print_stmt.

bun build without --sourcemap handles the same file fine, since the lexer already treats a truncated trailing sequence as EOF. Only the sourcemap line table is unguarded. Files ending mid-sequence show up in practice: truncated copies or downloads, an editor or CI job dying mid-write, or a file being rewritten while a build is in flight.

Other truncated tails hit the same bug: \xe2, \xc3, \xc1, \xe0\x81, \xf0\x9f\x92.

Cause

LineOffsetTable::generate_in (src/sourcemap/LineOffsetTable.rs) walks the source one codepoint at a time. For each lead byte it gets the declared sequence width from wtf8_byte_sequence_length_with_invalid. The codepoint decode already clamps that width to the bytes remaining (.min(remaining.len())), but two other uses of it do not:

  • the advance, remaining = &remaining[cp_len..], so a lone 0xF0 at EOF indexes [4..] of a 1-byte slice (this is the reported panic)
  • the offset passed to index_of_newline_or_non_ascii_check_start, which does &slice[offset..] internally; this is the panic site for overlong truncated leads such as 0xC1 and 0xE0 0x81, whose zero-padded decode lands in the ASCII fast path

Every other caller of wtf8_byte_sequence_length_with_invalid in the tree already clamps both the decode and the advance: the lexer steppers and URL percent-encoder in bun_core::strings, the string escapers in js_printer, and the markdown ANSI renderer. (Chunk::update_generated_line_and_column_slow has a superficially similar unclamped advance, but it walks the printer output, which only grows by whole codepoints; the lexer strips a truncated trailing sequence from the source before anything reaches the printer, so it is not reachable and is left unchanged.)

Fix

Compute the clamped width once, cp_len = (len_ as usize).min(remaining.len()), and use it for all three: the decode slice, the skip offset, and the advance. This is the same clamped_width idiom the string escapers in js_printer and bun_core::string already use.

For any non-truncated input cp_len == len_, so behavior is unchanged; the two can only differ within len_ - 1 bytes of the end of the source.

Verification

test/js/bun/sourcemap/internal-sourcemap-roundtrip.test.ts gains a test.each over six truncated tails (0xF0, 0xE2, 0xC3, 0xC1, 0xE0 0x81, 0xF0 0x9F 0x92), chosen so the set covers both panic sites. Each spawns bun build --sourcemap=external on the malformed source and asserts exit code 0 plus the exact emitted sources and mappings. The mappings are identical across every variant because the console.log(1); payload is byte-identical and the truncated bytes live in a stripped comment.

All six fail on an unpatched build with the panic above; all pass with the fix, and the rest of the file still passes.

…ence

LineOffsetTable::generate_in computed each codepoint's width from the lead
byte's declared length. The decode already clamped that width to the bytes
remaining, but the advance and the SIMD-skip offset did not, so a source
whose final bytes are a truncated multi-byte sequence panicked:

  panic: range start index 4 out of range for slice of length 1

Compute the clamped width once and use it for the decode slice, the skip
offset, and the advance, matching the clamped_width idiom already used by
the string escapers.
@coderabbitai

coderabbitai Bot commented Jun 26, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 1f9e3277-f5d0-439f-af50-cb34ba7fbfd7

📥 Commits

Reviewing files that changed from the base of the PR and between 0589548 and 2485321.

📒 Files selected for processing (2)
  • src/sourcemap/LineOffsetTable.rs
  • test/js/bun/sourcemap/internal-sourcemap-roundtrip.test.ts

Walkthrough

Source-map generation now limits UTF-8 multibyte decoding to the remaining input bytes, and a new roundtrip test verifies that truncated or invalid trailing UTF-8 sequences still produce the same sourcemap output.

Changes

UTF-8 sourcemap truncation

Layer / File(s) Summary
Clamp UTF-8 decode length
src/sourcemap/LineOffsetTable.rs
generate_in now uses the clamped remaining-byte count for multibyte decoding and for the newline/non-ASCII fast-skip helper.
Roundtrip UTF-8 tail test
test/js/bun/sourcemap/internal-sourcemap-roundtrip.test.ts
A concurrent parameterized test builds inputs with truncated and invalid trailing UTF-8 byte sequences and compares the generated sourcemap output across cases.
🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly states the main fix: preventing sourcemap generation panics on truncated UTF-8 tails.
Description check ✅ Passed It covers the bug, root cause, fix, and verification steps, matching the template’s intent even with different headings.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Jun 26, 2026 •

Copy link
Copy Markdown
Collaborator Author

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — minimal, well-scoped clamp fix matching the idiom used everywhere else, with regression tests for both panic sites.

Extended reasoning...

Overview

This PR fixes a panic in LineOffsetTable::generate_in (src/sourcemap/LineOffsetTable.rs) when a source file ends in a truncated multi-byte UTF-8 sequence. The fix hoists the already-present .min(remaining.len()) clamp — previously applied only to the decode slice — so that the same clamped width cp_len is used for all three consumers: the decode slice, the SIMD-skip offset passed to index_of_newline_or_non_ascii_check_start, and the loop advance remaining = &remaining[cp_len..]. A new test.each block in internal-sourcemap-roundtrip.test.ts covers six truncated-tail variants that exercise both panic sites and asserts the exact emitted mappings.

Security risks

None. This converts a slice-index panic on malformed input into graceful handling. No unsafe is touched, no new parsing surface is exposed, and for any well-formed input cp_len == len_ so behavior is byte-for-byte identical. The clamp can only differ from the declared width within the final len_ - 1 bytes of the source.

Level of scrutiny

Low-to-moderate. The Rust diff is effectively a 3-line refactor: move the min(remaining.len()) computation up and reuse the result. I verified that index_of_newline_or_non_ascii_check_start does &slice_[offset as usize..] internally (so the unclamped offset was a real second panic site), that decode_wtf8_rune_t still receives the declared len_ (correct — the decoder needs the lead-byte width; truncation is handled by the zero-padded buffer, same as before), and that the len_ == 1 ASCII fast path is unaffected since the loop guard guarantees remaining.len() >= 1. The PR description accurately notes that every other caller of wtf8_byte_sequence_length_with_invalid already uses this clamped-width idiom.

Other factors

No CODEOWNERS apply to these paths. No prior human review comments are outstanding. The bug-hunting system found nothing. The added tests are deterministic, spawn child processes (so a regression would fail the test rather than crash the runner), and pin the exact mappings output to catch any unintended sourcemap drift.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — minimal, correct clamp of the declared UTF-8 width to bytes-remaining, matching the idiom already used elsewhere in the tree, with thorough regression tests.

Extended reasoning...

Overview

This PR fixes a panic in LineOffsetTable::generate_in (src/sourcemap/LineOffsetTable.rs) when a source file ends in a truncated multi-byte UTF-8 sequence. The fix hoists the existing .min(remaining.len()) clamp — which was already applied to the decode slice — up to the definition of cp_len, so the same clamped width is used for the decode, the SIMD-skip offset passed to index_of_newline_or_non_ascii_check_start, and the loop advance remaining = &remaining[cp_len..]. Net code change is ~6 lines in one function plus a comment. A new test.each block in test/js/bun/sourcemap/internal-sourcemap-roundtrip.test.ts covers six truncated-tail variants spanning both former panic sites and asserts exit code 0 plus exact sources/mappings output.

Security risks

None. This strictly tightens a slice bound: previously an out-of-bounds slice index produced a Rust panic (safe abort, no UB); now it advances by the bytes actually available. There is no unsafe, no new I/O, and no change to what data is read or emitted for well-formed input.

Level of scrutiny

Low. The diff is tiny and mechanical: it moves an existing .min(remaining.len()) two lines earlier and reuses the result. For any non-truncated input cp_len == len_ so behavior is byte-identical; the two only diverge within the final len_ - 1 bytes of the file. The loop guard while !remaining.is_empty() guarantees cp_len >= 1, so the advance still makes progress. The decode call still passes the unclamped len_ to decode_wtf8_rune_t over a zero-padded 4-byte buffer, which is unchanged from before.

Other factors

The PR description is unusually thorough — it identifies both panic sites, explains why Chunk::update_generated_line_and_column_slow (the only other unclamped caller) is unreachable for this input, and notes that every other caller of wtf8_byte_sequence_length_with_invalid already uses this clamp idiom. The new tests are end-to-end (spawn bun build --sourcemap=external) and assert concrete output, not just absence-of-crash. The bug-hunting system found no issues, and there are no outstanding human review comments. No CODEOWNERS entry covers this path.

@robobun

robobun commented Jun 26, 2026 •

Copy link
Copy Markdown
Collaborator Author

CI status: the diff is green. Every failure on this PR's two builds is unrelated infrastructure or an external breakage.

First run (65054): 4 failures, all infrastructure. An npm registry mid-publish on react's experimental dist-tags broke the test/bake harness setup; test/js/bun/http/bun-serve-file.test.ts and fetch-file-upload.test.ts timed out on darwin-14-aarch64; bun-install.test.ts failed an assertion on win2019-x64-baseline. All four cleared on the retriggered run.

Retriggered run (65072): 2 failures.

  1. test/js/node/tls/node-tls-connect.test.ts > should have peer certificate, on win2019-x64-baseline and debian-13-x64-asan. The test opens a live TLS connection to the real bun.sh host and asserts the served certificate carries an OCSP responder URI:

    // node-tls-connect.test.ts:313
    expect(infoAccess["OCSP - URI"]).toBeDefined();  // Received: undefined

    Both lanes fail identically, so bun.sh's current certificate no longer includes an OCSP URL. (CAs have been removing OCSP responders; the test's own comment on the line above is "we just check the types this can change over time".) That is a property of an external certificate, not of code. The same failure is on the most recent builds of every other open PR (65056, 65058, 65059), so it is not specific to this branch either. node-tls-connect.test.ts needs its own fix.

  2. test/regression/issue/20965.test.ts > aborting a streaming file response mid-transfer does not leak pending_requests (server.stop resolves), a 90-second timeout on darwin-14-aarch64. I ran it locally on the exact CI commit (2485321c) under the debug ASAN build, which is far slower than the release build CI uses:

    (pass) aborting a streaming file response mid-transfer does not leak
           pending_requests (server.stop resolves) [2043.00ms]
    

    A test that finishes in 2 seconds on the slowest build but needs more than 90 seconds on a fast release build is a runner-timing flake, not a regression. darwin-14-aarch64 is the same lane that produced both test/js/bun/http timeouts on the first run, and on another PR's build from the same window (65049). The test exercises Bun.serve request abort and file streaming, which has no code path in common with this PR's one-line bounds clamp in src/sourcemap/LineOffsetTable.rs.

This PR's own test (all six variants in test/js/bun/sourcemap/internal-sourcemap-roundtrip.test.ts) passed on every lane in both builds, and neither build's annotations mention this PR's changed files anywhere. Both review bots found nothing to change. I'm not going to keep retriggering CI for unrelated failures; this is ready for a maintainer.

@Jarred-Sumner
Jarred-Sumner merged commit 990be52 into main Jun 26, 2026
70 of 78 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the farm/33e8ee25/fix-sourcemap-truncated-utf8 branch June 26, 2026 20:39
@robobun

robobun commented Jun 26, 2026

Copy link
Copy Markdown
Collaborator Author

Heads up: the sourcemap of a source with a truncated trailing UTF-8 sequence tests added in this PR fail deterministically on the Windows CI lanes. The sources entry comes back with a native path separator:

 expect(received).toMatchObject(expected)
   "sources": [
-    "../in.js",
+    "..\in.js",

It fails on all retries, so it isn't flake. First hit on windows-2019-x64 in build 65191, a branch with main merged in. Since 990be52 is on main and nothing has changed this file since, main's Windows lanes and every PR's Windows lanes will fail the same way until it's addressed.

I'm not touching it from my PR (#32733, glob, unrelated; I only hit this by merging main) because the right fix needs a call that belongs here: normalize sources to forward slashes in the sourcemap writer, or normalize in the test. Flagging it so you can take it.

Jarred-Sumner pushed a commit that referenced this pull request Jun 27, 2026
Fixes #14769

### Repro

`bun build --sourcemap` writes the host path separator into the
`sources` entries of the emitted `.js.map`, so on Windows a source one
directory above the out dir comes back as `..\in.js` instead of
`../in.js`. Source map `sources` are URLs resolved against the map's
location, where `\` is not a path separator; esbuild, which this code is
ported from, explicitly converts to forward slashes here ("Make sure to
always use forward slashes, even on Windows").

This is the bug reported in #14769: `bun build src/main.mjs
--sourcemap=linked --outdir dist` on Windows produces a map whose
`sources` all carry a `..\src\` prefix, and Chrome devtools shows the
raw `..\src\main.mjs:5` in the console instead of resolving the frame
back to `main.mjs:5`.

It is also what turned the Windows test lanes red. #32774 added the
first test in the suite to pin an exact `sources` value, and it fails
deterministically on all three Windows `test-bun` lanes (2019 x64, 2019
x64-baseline, 11 aarch64). From build 65191:

```
error: expect(received).toMatchObject(expected)

  {
    "sources": [
-     "../in.js",
+     "..\in.js",
    ],
  }

      at internal-sourcemap-roundtrip.test.ts:443:17
```

All six `sourcemap of a source with a truncated trailing UTF-8 sequence`
variants fail the same way, on every retry. Main has not completed a
Windows build since #32774 merged (each one since was superseded before
the test lanes finished), so its next completed build, and every PR that
merges main, hits this too.

### Cause

`LinkerContext::generate_source_map_for_chunk` computes each file
source's relative path with `bun_paths::resolve_path::relative_alloc`,
which uses the host separator (`platform::Auto`), and writes the result
straight into the `sources` array.

Everywhere else an output-facing path is built, the tree already
enforces forward slashes: `Path::pretty` carries a Windows debug
assertion that it contains no backslashes (`assert_pretty_is_valid`),
`dupe_alloc_fix_pretty` upholds it with `platform_to_posix_in_place`,
and the dev server's source map writer plus the standalone module graph
both normalize before emitting. The bundler's chunk source map writer is
the one place that overrides `pretty` with a raw native relative path
and skips the step. The non-file branch (`path.pretty`, virtual and
plugin sources) already holds the invariant and is untouched.

### Fix

A new `source_map_relative_path` helper computes the relative path and
calls `bun_paths::resolve_path::platform_to_posix_in_place` on it; both
`sources` call sites (the first entry and the `source_indices[1..]`
loop) go through it. The normalization is a compile-time no-op when the
host separator is `/`, so POSIX output is byte-identical; on Windows the
only change is `\` to `/` inside file-namespace `sources` entries.

### Verification

`test/js/bun/sourcemap/internal-sourcemap-roundtrip.test.ts` gains a
test that builds an entry in a nested directory with one import, so the
two `sources` entries each cross multiple separators and between them
cover both call sites, and pins them exactly:

```ts
expect(map.sources).toEqual(["../src/dep.js", "../src/nested/in.js"]);
```

Because the fixed function is a no-op on POSIX by construction, the
failure is only observable on a Windows host. The three red Windows
lanes in build 65191 above are the fail-before; the `sources:
["../in.js"]` assertions from #32774 and the new test are the regression
coverage. The full roundtrip file,
`bun-build-compile-sourcemap.test.ts`,
`compile-sourcemap-internal.test.ts`, and the `snapshotSourceMap`
bundler edge cases all pass with the change.

The `expectBundled` harness masks exactly this (`parsed.sources.map(a =>
a.replaceAll("\\", "/"))` at `expectBundled.ts:1617`), which is why no
`test/bundler/` source map test ever caught it. Left in place since it
also covers non-file `pretty` sources, which this change does not touch.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants