Skip to content

Reject sources of 2 GiB or more before parsing instead of aborting in usize2loc - #39095

Merged
Jarred-Sumner merged 3 commits into
mainfrom
farm/da4ffa31/source-too-large
Aug 15, 2026
Merged

Jarred-Sumner merged 3 commits into
mainfrom
farm/da4ffa31/source-too-large

Conversation

@robobun

@robobun robobun commented Aug 15, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • A source file of 2 GiB or more aborts the process instead of producing an error when it reaches a parser through bun build, bun run, import or Bun.Transpiler:
    • JS/TS: panic: int cast: TryFromIntError(PosOverflow) / Crashed while parsing big.js (Lexer::loc -> bun_ast::usize2loc, src/ast/lib.rs)
    • TOML: panic: source length is bounded by i32::MAX: TryFromIntError(PosOverflow) (loc_of, src/parsers/toml.rs); nothing enforced that bound
    • YAML: panic: int cast: TryFromIntError(PosOverflow) (Pos::loc, src/parsers/yaml.rs)
    • XML: panic: assertion failed: contents.len() <= i32::MAX as usize in debug builds (src/parsers/xml_index.rs:40); release builds silently saturate positions
  • Cause: every parser records positions as an i32 Loc, and Bun.{JSON5,JSONC,TOML,YAML}.parse: reject inputs of 2^31 bytes or more instead of panicking #32764 only added the length check to the Bun.*.parse JS entry points (src/runtime/api.rs). The parsers themselves accept any length, and the file-based entry points (src/bundler/ParseTask.rs, src/bundler/transpiler.rs), bunfig.toml loading, pnpm migration and Bun.Transpiler hand them whatever they read.
  • Reproduces on 1.4.0 with a sparse file (writeFileSync(f, "/*"); truncateSync(f, 2 ** 31); appendFileSync(f, "*/\nx"), then bun f.js or bun build f.js), and with 2 GiB of newlines followed by a = 1 / a: 1 for TOML and YAML. JSON and JSONC are not affected: their structural indexer already reports JSON document is too large to parse (2 GiB maximum). JSON5 has the same bug and is being fixed in Add a ratchet for expect("int cast") sites and clear json5, image codecs and elf #38936.

Fix

  • Source::check_parseable_len(log, what) in bun_ast (next to Loc and usize2loc, whose precondition it establishes) logs <what> is too large to parse (2 GiB maximum) against the source and returns Err(SourceTooLarge). The error is attributed to the file without a position: computing one would scan the oversized file to find the end of its line (that scan is why the existing JSON check, which reports at offset 0, takes seconds on a single-line file).
  • Called at the single entry point of each affected parser: Parser::init (all JS/TS parsing, including scans and the transpiler API), TOML::parse, YAML::parse and XML's parse_units (both parse and parse_utf16). The XML indexer's assertion now states the bound it actually relies on, u32: the scanner may transcode UTF-16 or Latin-1 input to UTF-8 after the entry check, which at most doubles it, and positions in transcoded input only ever live in the u32 index (Scanner::loc attaches no location to them), so an input under the limit that grows past 2 GiB when transcoded keeps parsing, as it does in release builds today. SourceTooLarge converts into each crate's already-logged SyntaxError variant, so every existing caller reports it the way it reports any other parse error: bun build prints a build error and exits 1, import rejects with a BuildMessage, Bun.Transpiler throws one, bunfig and pnpm report a parse error.
  • Why the parsers rather than the file loaders: the i32 limit is the parsers' own precondition, and they are reached from more places than the two file loaders (bunfig, pnpm, S3 XML responses, test snapshots, Bun.Transpiler, bundler plugins returning contents, and XML's Latin-1 to UTF-8 re-parse, whose input can be twice the length the API check measured). Checking at the parser covers all of them and matches what the JSON parser already does. The message wording follows the JSON one.
  • Verified:
    • test/js/bun/transpiler/source-too-large.test.ts: Bun.Transpiler with a 2 GiB buffer for js, ts, toml, yaml, xml (plus json and jsonc to pin the existing behavior), bun build of a 2 GiB .xml, and bun run of a 2 GiB .js, both sparse files. Passes with bun bd test; with USE_SYSTEM_BUN=1 (1.4.0) all three fail, and the unfixed debug build aborts on the bun build case. The fixtures' first line is a syntax error in every format so that a build without the check fails fast instead of scanning 2 GiB; the crash itself needs parseable content past 2 GiB, which is what the repros above use.
    • bun bd test on the toml, xml, yaml, resolve/{toml,yaml,xml,jsonc}, transpiler and bundler_loader suites; cargo clippy on bun_ast, bun_parsers, bun_js_parser; the source lints.
  • Related: A Loc is always a byte offset: Option<Loc> for absence, no -1 sentinel #38825 changes the representation of Loc; if it lands first, MAX_PARSEABLE_LEN and the Loc::EMPTY argument in the helper are the only two lines here that need to follow it. Two things found on the way are left for separate changes: the CSS parser has the same class of casts, and the bundler's empty fallback AST for an unparsable JS file presizes its symbol tables from the source length (correct but slow in debug builds, which is why the bun build test case uses a data-format file).

Background

  • Loc (bun_ast) is the position type stored in every AST node and diagnostic: an i32 byte offset into the source. usize2loc is the shared conversion from a parser's usize cursor; like the parsers' private equivalents it is an expect, and the binary builds with panic = "abort", so an offset past i32::MAX is a process abort.
  • Source is the path plus contents handed to every parser, whether the bytes came from a file read, a bundler plugin or a JS string. Log collects diagnostics; a parse error is logged and then signalled to the caller with a bare SyntaxError value, which is the convention the new error converts into.
  • A sparse file (truncateSync past the end) takes no disk space and reads back as NUL bytes, which is enough to exercise the length check; the tests use that so the 2 GiB fixtures cost only the memory of reading them.
Before and after

Release 1.4.0, sparse /* + 2 GiB hole + */ JS file:

$ bun block.js
panic: int cast: TryFromIntError(PosOverflow)
Crashed while parsing /tmp/repro/block.js
$ bun build block.js --outdir out
panic: int cast: TryFromIntError(PosOverflow)
Crashed while parsing /tmp/repro/block.js

Release 1.4.0, 2 GiB of newlines followed by one line:

import('./dense.toml')   ->  panic: source length is bounded by i32::MAX: TryFromIntError(PosOverflow)
import('./dense.yaml')   ->  panic: int cast: TryFromIntError(PosOverflow)

Unfixed debug build, any 2 GiB .xml: panic: assertion failed: contents.len() <= i32::MAX as usize.

With this change (debug build):

$ bun build block.js --outdir out
error: File is too large to parse (2 GiB maximum)
    at /tmp/repro/block.js
$ bun build big.toml --outdir out
error: TOML document is too large to parse (2 GiB maximum)
    at /tmp/repro/big.toml
$ bun -e "import('./block.js').catch(e => console.log(e.name, e.message))"
BuildMessage File is too large to parse (2 GiB maximum)

Bun.Transpiler.transformSync with a 2 GiB Uint8Array, per loader: js/ts File is too large to parse (2 GiB maximum), toml/yaml/xml <FORMAT> document is too large to parse (2 GiB maximum), json/jsonc unchanged (JSON document is too large to parse (2 GiB maximum)), each in about a millisecond.

Every parser records positions as an i32 Loc, and usize2loc aborts on the
first position past i32::MAX. Bun.TOML.parse and the other JS entry points
already rejected such inputs, but the parsers themselves did not, so a file
of 2 GiB or more reached through bun build, bun run, import or Bun.Transpiler
crashed the process (JS, TOML, YAML) or failed the XML indexer's length
assertion.

Add Source::check_parseable_len, which logs "<what> is too large to parse
(2 GiB maximum)" against the source (without a position, so the oversized
file is not scanned) and call it at the entry of the JS parser, TOML, YAML
and XML, so every caller of those parsers gets an error instead of a crash.
JSON and JSONC already had an equivalent check in the structural indexer.
@coderabbitai

coderabbitai Bot commented Aug 15, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 31 seconds

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 6f8e82dc-34a3-4a4d-b489-944fbdb3fe52

📥 Commits

Reviewing files that changed from the base of the PR and between 88a6398 and 6fe9599.

📒 Files selected for processing (9)
  • src/ast/lib.rs
  • src/js_parser/error.rs
  • src/js_parser/parse/parse_entry.rs
  • src/parsers/error.rs
  • src/parsers/toml.rs
  • src/parsers/xml.rs
  • src/parsers/xml_index.rs
  • src/parsers/yaml.rs
  • test/js/bun/transpiler/source-too-large.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Aug 15, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 11:05 AM PT - Aug 15th, 2026

❌ @robobun, your commit 6fe9599 has 1 failures in Build #98279 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 39095

That installs a local version of the PR into your bun-39095 executable, so you can run:

bun-39095 --bun

@robobun

robobun commented Aug 15, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: fix and tests pushed, waiting on CI.

Reproduced on 1.4.0 with sparse files (writeFileSync a short prefix, truncateSync to 2 GiB): bun big.js and bun build big.js abort with panic: int cast: TryFromIntError(PosOverflow); 2 GiB of newlines plus one line does the same for .toml and .yaml imports, and any 2 GiB .xml fails the indexer's length assertion in a debug build. With this branch all of them report ... is too large to parse (2 GiB maximum) as a normal build error.

test/js/bun/transpiler/source-too-large.test.ts fails on 1.4.0 (USE_SYSTEM_BUN=1) and passes with bun bd test.

Review so far: the XML transcoding edge case from the review is handled in f82bb18 (the indexer asserts the u32 bound it relies on); the new comments were trimmed to one line each.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM. The one inline nit is a pre-existing debug-only edge case (XML in-parser transcoding of >1 GiB inputs), not a regression introduced here, and doesn't block — the change is a strict improvement over the current abort.

Extended reasoning...

Overview

Adds a single shared guard, Source::check_parseable_len, in bun_ast next to Loc/usize2loc (whose i32 precondition it enforces), and calls it once at the entry of each affected parser: JS/TS Parser::init, TOML::parse, YAML::parse, and XML parse_units. A new zero-size SourceTooLarge error converts via From into each crate's existing already-logged SyntaxError variant, so every caller — bun build, bun run, import, Bun.Transpiler, bunfig, pnpm — surfaces it exactly as it surfaces any other parse error today. A new test file covers every loader through Bun.Transpiler plus bun build/bun run via sparse 2 GiB files.

Security risks

None. This is a length guard that turns a process abort on adversarial-size input into a recoverable, logged error. No new input parsing, no new trust boundaries, no allocation driven by the oversized length.

Level of scrutiny

The JS parser entry point is hot-path production code, but the added line is a branch that's trivially false for every input under 2 GiB — behavior for real workloads is byte-identical. The error propagation reuses the existing SyntaxError variant that every caller already handles, so no caller needs updating. add_error_fmt(Some(self), Loc::EMPTY, ...) matches how other position-less diagnostics are logged (e.g. lib.rs:1804). The message wording mirrors the JSON parser's existing check. This is a small, mechanical, well-scoped defensive fix.

Other factors

  • The inline nit about XML transcoding (transcode_utf16 / Latin-1 → UTF-8 expanding self.src past 2 GiB after the check) is real but pre-existing, debug-only, and requires >1 GiB of pathological all-high-byte content. The PR description already names the Latin-1 re-parse case explicitly; release builds are unaffected because Scanner::loc returns Loc::EMPTY when transcoded. It's a reasonable follow-up, not a blocker.
  • Test design is careful: subprocess with drained pipes, sparse files to avoid disk cost, totalmem() skip gate copied from fs-oom.test.ts, first line is a syntax error in every format so a build without the check fails fast rather than scanning 2 GiB, and json/jsonc are included to pin existing behavior. The 120s per-test timeout on the two file-based cases is justified in-comment (debug allocator bookkeeping on a 2 GiB read).
  • No CODEOWNERS cover the touched files.
  • PR description documents USE_SYSTEM_BUN=1 failure and bun bd test pass, plus clippy on the three crates.

Comment thread src/parsers/xml.rs
The parser bounds its input to i32::MAX units before building the index,
but transcoding UTF-16 or Latin-1 input to UTF-8 can grow it by up to 2x;
positions are stored as u32, so that is the bound the index's assertion
should state.
Comment thread src/ast/lib.rs Outdated
Comment thread src/ast/lib.rs Outdated
Comment thread src/ast/lib.rs Outdated
Comment thread src/ast/lib.rs Outdated
Comment thread src/parsers/xml.rs Outdated
Comment thread src/parsers/xml_index.rs Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — the XML transcoding concern from the earlier review is addressed by relaxing the indexer's debug_assert to u32::MAX (f82bb18), and the comment-cop lints are resolved (6fe9599).

Checked that check_parseable_len sits at the single entry point of each parser and is a no-op fast-path for inputs under 2 GiB.
Verified parse_utf16's caller casts source.contents to &[u16], so the byte-length check bounds the unit count too, and InputEncoding::Text skips transcoding on that path.
Confirmed 2×(i32::MAX) < u32::MAX and Scanner::loc gates i32 conversion on !transcoded, so the relaxed assertion matches the actual invariant.
The three From<SourceTooLarge> impls route to each crate's already-logged SyntaxError variant, so callers report it like any other parse error.

Extended reasoning...

Overview

Adds Source::check_parseable_len and SourceTooLarge to bun_ast (next to Loc/usize2loc, whose i32 precondition it establishes), calls it at the single entry point of each affected parser (Parser::init for JS/TS, TOML::parse, YAML::parse, XML's parse_units), converts SourceTooLarge into each crate's existing SyntaxError variant, and relaxes the XML structural indexer's debug_assert from i32::MAX to u32::MAX to reflect the bound it actually relies on after in-parser transcoding. New test file exercises all seven loaders through Bun.Transpiler, plus bun build and bun run with sparse 2 GiB files.

Security risks

None. This adds an early length rejection; it cannot expose data or bypass checks. The change strictly tightens behavior (abort → recoverable error) and only fires on inputs ≥ 2 GiB.

Level of scrutiny

Moderate. The change touches the JS parser entry point (as hot as it gets), but the added code is a single usize comparison that returns Ok(()) on every input under 2 GiB — no observable change to normal parsing. For oversized inputs, a process abort becomes a logged SyntaxError routed through existing error machinery, which every caller already handles. The XML debug_assert relaxation was the only subtle piece; the earlier review examined it in detail and the author's fix (assert u32::MAX, since transcoded positions never reach i32 conversion) is the correct one of the two options I proposed.

Other factors

  • My earlier inline finding (XML transcoding could grow past the checked length) was addressed exactly as suggested; the thread is resolved.
  • The comment-cop bot flagged multi-line comments; those were trimmed in 6fe9599 and the threads are resolved.
  • Tests cover the full loader matrix including json/jsonc (pinning existing behavior), and both file-based entry points. The 2 GiB file cases are gated on totalmem() >= 10 GiB (matching fs-oom.test.ts) and use sparse files to avoid disk cost.
  • The From<SourceTooLarge> impls in three separate error enums are mechanical and each maps to the existing SyntaxError variant, preserving the "already logged" contract.
  • CSS parser and JSON5 are noted as having the same bug class but explicitly scoped out (JSON5 is #38936); that's a stated exclusion, not an oversight.

@Jarred-Sumner
Jarred-Sumner merged commit 668caf3 into main Aug 15, 2026
10 of 11 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the farm/da4ffa31/source-too-large branch August 15, 2026 20:07
robobun added a commit that referenced this pull request Aug 16, 2026
ast/lib.rs: add_formatted_msg lost its clone flag on main while this branch
made the range optional, so the four callers take both. #39095's new
Source::check_parseable_len follows this branch as its description said it
would: the error carries no location, and the 2 GiB limit stays, now because
Range::len and reported columns are still i32 rather than Loc itself.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants