Skip to content

Buffer: match Node semantics for raw <enc>Slice / <enc>Write bindings - #33530

Merged
Jarred-Sumner merged 5 commits into
mainfrom
farm/76518de0/buffer-raw-enc-methods
Jul 7, 2026
Merged

Jarred-Sumner merged 5 commits into
mainfrom
farm/76518de0/buffer-raw-enc-methods

Conversation

@robobun

@robobun robobun commented Jul 6, 2026

Copy link
Copy Markdown
Collaborator

What

The raw Buffer.prototype.<enc>Slice / <enc>Write prototype methods (hexSlice, utf8Write, ucs2Write, base64Write, ...) are Node's undocumented but long-standing raw C++ bindings, which Bun also exposes. Bun routed them through the strict argument validator used by the documented toString() / write() wrappers, so calls that succeed on Node threw on Bun.

const b = Buffer.from("68656c6c6f21", "hex"); // 6 bytes

b.hexSlice(7);          // node: ""   bun before: RangeError[ERR_OUT_OF_RANGE]
b.hexSlice(7, 3);       // node: ""   bun before: RangeError[ERR_OUT_OF_RANGE]

const d = Buffer.alloc(9, 0xcc);
d.hexWrite("68656c6c6f", 6, 1000);
// node: returns 3, writes the 3 bytes that fit
// bun before: RangeError[ERR_BUFFER_OUT_OF_BOUNDS]

The documented wrappers (buf.toString(enc, start, end), buf.write(str, offset, length, enc)) already matched Node and are unchanged. This affects code that calls the raw methods directly, which a few performance-sensitive libraries do.

Cause

Two divergences from Node's node_buffer.cc:

  • StringSlice returns "" when start >= end, before the end <= length range check. Bun did the range check first, so any start past the end threw.
  • base64 / base64url / hex / ucs2 Write are still Node's raw binding, which clamps length to the space left in the buffer. Node only moved utf8 / latin1 / ascii Write onto a strict JS wrapper that throws ERR_BUFFER_OUT_OF_BOUNDS; Bun applied that wrapper's strictness to all seven.

Separately, parseArrayIndex used WTF's truncateDoubleToInt64, which maps NaN and Infinity to INT64_MIN. That made them look negative and throw, where Node's v8::Value::IntegerValue treats NaN as 0 and saturates at the int64 bounds.

Fix

In src/jsc/bindings/JSBuffer.cpp:

  • SliceWithEncoding short-circuits to "" on start >= end before the range check, matching Node's StringSlice.
  • Added StringWriteWithEncoding, Node's clamping StringWrite: clamp length to byteLength - offset, return 0 when no room, report a negative offset/length as ERR_OUT_OF_RANGE ("Index out of range"). base64 / base64url / hex / ucs2 / utf16le Write use it; utf8 / latin1 / ascii keep the strict wrapper.
  • parseArrayIndex now follows v8::Value::IntegerValue (NaN -> 0, saturating).

While verifying hostile coercions, found that a valueOf() which shrinks a resizable ArrayBuffer during index coercion trips a debug assertion in jsBufferToString (callers snapshot byteLength before the coercion runs user JS). The documented toString() wrapper hits it too. The clamp right below the assertions already handles the stale range correctly in release, so the assertions were wrong; removed them.

Verification

Every expectation in the new tests was checked byte-for-byte against a real Node v26.3.0 binary (including pulling Node's own internal/buffer.js source out of the binary to confirm which encodings use which path). New tests in test/js/node/buffer.test.js fail on unfixed Bun and pass with the fix; latin1Slice assertions that encoded the old throwing behavior were updated to Node's actual output. The full buffer.test.js (611 pass) and all test-buffer-*.js Node-ported suites pass.

The raw Buffer.prototype.<enc>Slice / <enc>Write prototype methods were
routed through the strict validator used by the documented toString() /
write() wrappers, so calls that succeed on Node threw on Bun:

  buf.hexSlice(pastEnd)           -> "" on Node, RangeError on Bun
  buf.hexWrite(str, off, hugeLen) -> clamped on Node, RangeError on Bun

Mirror Node's node_buffer.cc bindings instead:

- StringSlice short-circuits to "" when start >= end, before the
  end <= length range check.
- base64/base64url/hex/ucs2/utf16le Write clamp length to the space
  left and report a negative index as ERR_OUT_OF_RANGE. utf8/latin1/
  ascii Write stay on the strict wrapper Node keeps for them.
- parseArrayIndex follows v8::Value::IntegerValue (NaN -> 0, saturating)
  instead of truncateDoubleToInt64, which mapped NaN/Infinity to
  INT64_MIN and made them throw.

Also drop the length/offset assertions in jsBufferToString: a valueOf()
that shrinks a resizable ArrayBuffer during index coercion leaves the
snapshotted range stale, which aborted debug builds. The existing clamp
below already handles it.
@robobun

robobun commented Jul 6, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 6:08 PM PT - Jul 6th, 2026

@robobun, your commit 195b495 is building: #69417

@github-actions github-actions Bot added the claude label Jul 6, 2026
@coderabbitai

coderabbitai Bot commented Jul 6, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 39 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 143071fb-a743-40ae-b31c-a83d9cd4afe4

📥 Commits

Reviewing files that changed from the base of the PR and between 91feaba and 195b495.

📒 Files selected for processing (2)
  • src/jsc/bindings/JSBuffer.cpp
  • test/js/node/buffer.test.js

Walkthrough

This PR updates buffer.test.js to align Buffer.latin1Slice() behavior with Node semantics, expecting empty strings instead of range errors when start >= end. It also adds new test coverage for raw enc Slice/Write bindings, verifying clamping versus strict throw behavior across encoding families.

Changes

Buffer slice/write parity tests

Layer / File(s) Summary
latin1Slice short-circuit expectations
test/js/node/buffer.test.js
Updates Buffer.latin1Slice() tests for both Buffer and Uint8Array-backed calls so that start >= end returns an empty string instead of throwing a RangeError.
Strict and clamping write parity
test/js/node/buffer.test.js
Reorganizes *Write tests into strict (utf8Write, latin1Write, asciiWrite) versus clamping (utf16leWrite, ucs2Write, base64Write, base64urlWrite, hexWrite) groups, and adds a new raw <enc>Slice / <enc>Write bindings match Node suite covering index short-circuiting, out-of-range errors, NaN/negative handling, resizable-buffer clamping via valueOf(), and wrapper behavior regressions.

Sequence Diagram(s)

Not applicable; changes are limited to test assertions without new sequential runtime interactions.

Estimated code review effort: 2/5 (Low-Medium)

Related issues: None provided.

Related PRs: None provided.

Suggested labels: test, buffer, node-compat

Suggested reviewers: None provided.

🐰 A hop, a skip, through slices tight,
Where start meets end, no more a fright,
Clamped or thrown, each encoding knows,
Node's own rhythm, the test now shows.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly summarizes the main change: aligning raw Buffer slice/write bindings with Node semantics.
Description check ✅ Passed The description covers what changed and how it was verified, with extra context in cause/fix sections.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

Comment @coderabbitai help to get the list of available commands.

Comment thread test/js/node/buffer.test.js
Comment thread test/js/node/buffer.test.js Outdated
The raw Buffer.prototype.<enc>Slice / <enc>Write prototype methods were
routed through the strict validator used by the documented toString() /
write() wrappers, so calls that succeed on Node threw on Bun:

  buf.hexSlice(pastEnd)           -> "" on Node, RangeError on Bun
  buf.hexWrite(str, off, hugeLen) -> clamped on Node, RangeError on Bun

Mirror Node's node_buffer.cc bindings instead:

- StringSlice short-circuits to "" when start >= end, before the
  end <= length range check.
- base64/base64url/hex/ucs2/utf16le Write clamp length to the space
  left and report a negative index as ERR_OUT_OF_RANGE. utf8/latin1/
  ascii Write stay on the strict wrapper Node keeps for them.
- parseArrayIndex follows v8::Value::IntegerValue (NaN -> 0, saturating)
  instead of truncateDoubleToInt64, which mapped NaN/Infinity to
  INT64_MIN and made them throw.

Also drop the length/offset assertions in jsBufferToString: a valueOf()
that shrinks a resizable ArrayBuffer during index coercion leaves the
snapshotted range stale, which aborted debug builds. The existing clamp
below already handles it.
@robobun

robobun commented Jul 6, 2026

Copy link
Copy Markdown
Collaborator Author

Good catch on the first thread: the src/jsc/bindings/JSBuffer.cpp changes were staged locally but left out of the initial commit (the commit captured only the test file), so the PR was all-tests-no-implementation. Fixed in 3b85955: the commit now includes the full C++ change (the SliceWithEncoding short-circuit, the new StringWriteWithEncoding clamping path wired to base64/base64url/hex/ucs2/utf16le, the parseArrayIndex IntegerValue semantics, and the jsBufferToString assertion removal). Full buffer.test.js passes locally (611 pass, 0 fail) and all test-buffer-*.js node-ported suites pass.

Also applied the stderr-surfacing suggestion on the resizable-shrink spawn test so an abort diagnostic lands in the diff on regression.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found on this pass now that the C++ is in — but this adds new bounds/detach-recheck logic in JSBuffer.cpp and loosens validation on a hot memory-write path, so it's worth a human set of eyes on the StringWriteWithEncoding clamping and the removed jsBufferToString assertions.

Extended reasoning...

Overview

This PR aligns Bun's raw Buffer.prototype.<enc>Slice / <enc>Write bindings with Node's semantics. In src/jsc/bindings/JSBuffer.cpp it: (1) makes SliceWithEncoding return "" on start >= end before the range check, (2) adds a new ~55-line StringWriteWithEncoding template (clamping length to remaining space, re-checking detach after each coercion) and rewires utf16le/base64/base64url/hex Write to it, (3) replaces truncateDoubleToInt64 in parseArrayIndex with a saturating toIntegerValue (NaN→0, ±Inf saturate), (4) removes three debug ASSERTs in jsBufferToString in favor of the existing runtime clamp, and (5) tweaks the strict wrapper's default-length computation when offset is NaN. test/js/node/buffer.test.js gains ~200 lines of parity tests and updates two latin1Slice assertions that encoded the old (wrong) throwing behavior.

My previous review flagged that the C++ was missing entirely; the author pushed it in 3b85955 and applied the stderr-surfacing nit. Both prior threads are resolved. The bug-hunting pass on the current revision found nothing.

Security risks

The new StringWriteWithEncoding does user-controlled offset/length arithmetic feeding into writeToBuffer (which writes into the backing store at offset for up to maxLength bytes). The clamping looks correct — offset > byteLength throws, then maxLength = min(byteLength - offset, maxLength) after re-reading byteLength post-coercion — and detach is re-checked after each parseArrayIndex. The size_t→uint32_t narrowing into writeToBuffer is pre-existing (shared with the strict path). Removing the jsBufferToString assertions is justified by the immediately-following clamp, but it does mean a future caller that skips the clamp loses the debug tripwire. Nothing jumps out as exploitable, but this is precisely the class of code (bounds arithmetic on adversarial input, re-entrancy via valueOf(), resizable-buffer shrink) that CLAUDE.md flags as the most-blocked review category.

Level of scrutiny

High. Buffer write paths are among the hottest and most memory-safety-sensitive code in the runtime, and this change loosens validation (calls that previously threw now write). The implementation is careful and well-commented against Node source lines, and the test coverage is thorough (per-encoding matrix, NaN/Infinity/negative, resizable-shrink-during-coercion spawn test, wrapper-unchanged regression checks). But new C++ bounds logic in JSBuffer.cpp is not something I should wave through without a human reviewer.

Other factors

Test coverage is strong and the author verified byte-for-byte against Node v26.3.0. No CODEOWNERS entry matches this path. CI build #69269 was still running at last timeline update. The parseArrayIndex signature change dropped the per-call-site error message in favor of a fixed "Index out of range" — that matches Node but is a user-visible message change a human might want to confirm.

Comment thread src/jsc/bindings/JSBuffer.cpp
Comment thread src/jsc/bindings/JSBuffer.cpp Outdated
robobun added 2 commits July 6, 2026 21:07
… ran JS

The offset/length coercions only run user-overridable code (valueOf /
Symbol.toPrimitive / toString) for object arguments. Gate the isDetached
re-check on that so the common path (string + numeric args) skips it, and
re-read byteLength only after a coercion that could have resized the view.
@robobun

robobun commented Jul 6, 2026

Copy link
Copy Markdown
Collaborator Author

CI status: the diff is green on its own terms, the red lanes are unrelated flake.

This PR touches only src/jsc/bindings/JSBuffer.cpp and test/js/node/buffer.test.js. The full buffer.test.js passes locally (614 pass, 0 fail) and all test-buffer-*.js node-ported suites pass.

Across the last two CI runs the only red lane was :darwin: 26 aarch64 - test-bun (plus a broken infra job), and the failing tests were a different, unrelated set each time, none of them Buffer-related:

  • build 69353: bake/dev-and-prod, webview-chrome, napi, sql-onconnect-onclose-throw, bun-install-security-provider, update_interactive_install, hot/watch-many-dirs
  • build 69417: bundler/compile-windows-metadata, cron/in-process-cron, http/proxy-stress-errors, net/net-mongodb-pattern-leak, solc, fetch/fetch-leak, napi, bun-install-security-provider

The set changing run-to-run with no overlap to the changed files is the signature of pre-existing flake on that lane. I pushed one ci: retrigger already; rather than spam more, leaving this for a maintainer to merge.

@Jarred-Sumner
Jarred-Sumner merged commit 3316c48 into main Jul 7, 2026
73 of 76 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the farm/76518de0/buffer-raw-enc-methods branch July 7, 2026 02:41
robobun added a commit that referenced this pull request Jul 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants