Conversation
|
Updated 2:54 AM PT - Sep 15th, 2026
✅ @robobun, your commit 29eca57a1e3b59c56fd1e8edbabca7fedf6f6728 passed in 🧪 To try this PR locally: bunx bun-pr 32839That installs a local version of the PR into your bun-32839 --bun |
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughNodeVM cached-data serialization now uses a validated header. Script, function, and module readers validate the header before decoding. Writers emit the wrapped format. Tests cover corrupted, mismatched, and intact cached data. ChangesCached-data hardening
Suggested reviewers: 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
There was a problem hiding this comment.
Actionable comments posted: 4
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@src/jsc/bindings/NodeVM.cpp`:
- Around line 144-179: The cached-data path in unwrapCachedData and
decodeCachedData is relying on a forgeable hash in CachedDataHeader, so
untrusted input can still reach decodeCodeBlock. Replace the public
checksum-only check in createCachedDataBuffer/hashCachedDataPayload with an
unforgeable validation approach (for example, a secret-backed tag or structural
validation) and ensure decodeCachedData rejects any payload that is not verified
before decoding. Keep the fix localized around CachedDataHeader,
unwrapCachedData, and decodeCachedData so malformed cachedData cannot proceed
past validation.
- Around line 312-315: The `NodeVM::createCachedDataBuffer()` result can be null
on allocation failure, but the current `function->putDirect()` path still stores
it and sets `cachedDataProduced` to true. In the `NodeVM.cpp` bytecode caching
flow, add a null check immediately after calling `createCachedDataBuffer()` and
before the `putDirect()` calls, mirroring the existing
`RETURN_IF_EXCEPTION`-style guarded paths. If the buffer is null, skip setting
`cachedData` and ensure `cachedDataProduced` is not marked successful.
In `@src/jsc/bindings/NodeVMSourceTextModule.cpp`:
- Around line 510-514: The cached-data path in
NodeVMSourceTextModule::cachedData() needs null checks for both bytecode() and
createCachedDataBuffer(). After calling bytecode(globalObject), verify the
RefPtr is non-null before using cachedBytecode->span(), and similarly verify the
JSUint8Array* returned by createCachedDataBuffer() before storing it in
m_cachedBytecodeBuffer. Keep the existing RETURN_IF_EXCEPTION checks, but add
early returns for the nullptr cases so the function never dereferences or caches
a missing bytecode/buffer.
In `@test/js/node/vm/vm.test.ts`:
- Around line 814-821: The crash-regression assertion in the `vm.test.ts` spawn
test is only checking `stdout` and `exitCode`, which hides useful failure
diagnostics when the child aborts. Update the `Bun.spawn` result handling in
this test to include `stderr` and `signalCode` in the asserted object, using the
existing `stderr` capture from `proc.stderr.text()` so regressions surface
native crash output and termination signals in the diff.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 7cdcdd3d-5b22-4daf-86b5-f6ab81bb7043
📒 Files selected for processing (5)
src/jsc/bindings/NodeVM.cppsrc/jsc/bindings/NodeVM.hsrc/jsc/bindings/NodeVMScript.cppsrc/jsc/bindings/NodeVMSourceTextModule.cpptest/js/node/vm/vm.test.ts
|
Addressed the review feedback in f28e900:
Two suggestions were not taken, with reasoning in the threads: the header is an integrity check rather than an authentication boundary (the same contract V8 uses for cachedData, and it must stay a public deterministic function of the bytes to keep cross-process reuse working), and createCachedDataBuffer() only returns null with a pending exception, which the RETURN_IF_EXCEPTION after each call already handles. test/js/node/vm/vm.test.ts and the Node vm cachedData tests pass locally with the debug (ASAN) build. |
There was a problem hiding this comment.
♻️ Duplicate comments (1)
src/jsc/bindings/NodeVMSourceTextModule.cpp (1)
514-520: 🩺 Stability & Availability | 🟠 Major | ⚡ Quick winBuffer null check from the prior review is still missing.
The new guard handles
cachedBytecode == nullptr, butcreateCachedDataBuffer()can also returnnullptr(its[[unlikely]]allocation-failure branch) without throwing. Line 520 then caches a null buffer intom_cachedBytecodeBufferand line 523 returns it. Guard the buffer too, matchingNodeVM::createCachedData.Proposed fix
JSUint8Array* buffer = createCachedDataBuffer(globalObject, cachedBytecode->span()); RETURN_IF_EXCEPTION(scope, nullptr); + if (!buffer) [[unlikely]] { + throwVMError(globalObject, scope, "createCachedData failed"_s); + return nullptr; + } m_cachedBytecodeBuffer.set(vm, this, buffer);🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/jsc/bindings/NodeVMSourceTextModule.cpp` around lines 514 - 520, The cached bytecode path in NodeVMSourceTextModule::createCachedData still misses the null check for the JSUint8Array returned by createCachedDataBuffer(). Add a guard after creating the buffer, before m_cachedBytecodeBuffer.set, and mirror the failure handling used in NodeVM::createCachedData by throwing a VM error and returning nullptr if the buffer allocation fails. Keep the existing cachedBytecode null handling unchanged and ensure the returned value is only used when the buffer is valid.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Duplicate comments:
In `@src/jsc/bindings/NodeVMSourceTextModule.cpp`:
- Around line 514-520: The cached bytecode path in
NodeVMSourceTextModule::createCachedData still misses the null check for the
JSUint8Array returned by createCachedDataBuffer(). Add a guard after creating
the buffer, before m_cachedBytecodeBuffer.set, and mirror the failure handling
used in NodeVM::createCachedData by throwing a VM error and returning nullptr if
the buffer allocation fails. Keep the existing cachedBytecode null handling
unchanged and ensure the returned value is only used when the buffer is valid.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: e647e475-19d9-4c2b-bdb1-d71613b12109
📒 Files selected for processing (2)
src/jsc/bindings/NodeVMSourceTextModule.cpptest/js/node/vm/vm.test.ts
There was a problem hiding this comment.
Thanks — 00f190c addresses my last note (cachedDataProduced now reports false when production fails, and the test locks it in), so I have nothing further. Deferring to a human for the final sign-off since this introduces a new cachedData wire format and changes the bytecode payload's memory ownership in core JSC bindings.
Extended reasoning...
Overview
This PR hardens node:vm's cachedData handling so that corrupted, truncated, or foreign buffers are rejected (cachedDataRejected = true / ERR_VM_MODULE_CACHED_DATA_REJECTED) instead of crashing inside JSC's bytecode decoder. It does so by wrapping serialized bytecode in a 16-byte { magic, version, XXH3-64 } header on the write side (createCachedDataBuffer) and verifying it on the read side (unwrapCachedData) before calling decodeCodeBlock. The verified payload is copied into a MallocSpan<uint8_t, JSC::VMMalloc> owned by the CachedBytecode (since JSC retains the decoder for lazy function-body decoding). All three entry points — vm.Script, vm.compileFunction, vm.SourceTextModule — are updated, along with their producers. Over the review cycle the PR also picked up fixes for three pre-existing null-deref sites and a stale cachedDataProduced(true) override, each with test coverage.
Files touched: src/jsc/bindings/NodeVM.{h,cpp}, NodeVMScript.cpp, NodeVMSourceTextModule.cpp, test/js/node/vm/vm.test.ts.
Security risks
The change is strictly a hardening: it gates an unsafe decoder behind an integrity check that previously did not exist. The "forgeable checksum" objection was raised and correctly dismissed in-thread — cachedData is an integrity/compatibility contract (matching V8's design), not an authentication boundary, and any caller who can craft a header is already executing arbitrary JS in-process. The payload copy removes a potential use-after-free where JSC's lazily-decoded functions could outlive the caller's buffer. I see no new exposure introduced.
Level of scrutiny
Moderate-to-high. This is ~150 lines of native C++ in core JSC bindings that (a) defines a new persisted binary format, (b) changes memory-ownership semantics for decoded bytecode, and (c) threads through three separate code paths. None of it is mechanical. The implementation looks correct and is well-tested (corruption matrix across all three entry points in a child process, plus the existing Node test-vm-cached-data.js cross-process round-trip), but the format and ownership decisions are the kind of thing a maintainer should ratify rather than a bot.
Other factors
The author has been responsive across four review rounds; every concern I and CodeRabbit raised has been addressed or rebutted with a sound argument, and all threads are resolved. The bug-hunting system found nothing on the latest revision. CI build #65446 is in flight for the head commit. Given the scope (native bindings, new wire format, lifetime change), I'm deferring rather than auto-approving.
|
Status at 2c7c0aa: merged main through 8cc0397, which includes the WebKit bump in #43616. There was no conflict. The bump changed the bytecode decoder and the payload layout under this PR, so everything below was run again on the merge. CI is running on it. What the bump changed for this PR:
How to reproduce, one process per offset (the script is in the PR description under Notes): flip one byte of a genuine
Local verification on the merge, debug build with ASAN: The PR description is rewritten for the current state of main. It also records how Node behaves: Node rejects truncated and stale blobs, and by default it does not verify its checksum, so a flipped byte crashes Node too. oven-sh/WebKit#368 (bounds checks inside the decoder) is closed. It went against oven-sh/WebKit#677, and bounds checks do not catch a changed byte that stays in range. All review threads are resolved. |
There was a problem hiding this comment.
No issues found this pass, and all earlier findings are addressed — but this is memory-safety C++ that introduces a new on-disk cachedData format, so it's worth a human sign-off on the header design and the owned-copy lifetime.
What was reviewed:
- Header validation in
unwrapCachedData— every field checked beforedecodeCodeBlocksees the payload;memcpyavoids alignment UB; length is compared against the actual span, not trusted. - The payload copy into
MallocSpan<uint8_t, VMMalloc>so JSC's lazy decoder does not borrow the caller's buffer. - All three consumers (
vm.Script,compileFunction,SourceTextModule) route through the same wrap/unwrap pair; the produce side symmetrically prepends the header at every site. - The
getBytecodeBuffer()/cachedData()null guards and the removedcachedDataProduced(true)override — now defensive-only after eagercheckSyntaxfrom #31827 closed the known repro.
Extended reasoning...
Overview
The PR wraps JSC's serialized bytecode in a 24-byte integrity header (magic, payloadLength, sourceHash, jscVersion, XXH3-64 payloadHash) so node:vm's three cachedData consumers can reject corrupted/truncated/mismatched buffers before handing them to JSC::decodeCodeBlock, which otherwise dereferences payload-relative offsets without bounds checks and segfaults. It also copies the accepted payload into memory owned by the CachedBytecode (JSC keeps the decoder alive for lazily decoded function bodies), and picks up three same-class defensive fixes on the produce side (getBytecodeBuffer null guard, SourceTextModule::cachedData null guard, dropped unconditional cachedDataProduced(true) override plus its now-dead setter).
Security risks
The input is user-controlled bytes flowing into a bytecode decoder — the fix is strictly a hardening in that direction. The header is an integrity check, not authentication (matching V8's contract for cachedData); a caller who can forge the header can also just supply matching bytecode, which is the documented capability. I did not find a way for the new code to make things worse: validation happens before allocation of the CachedBytecode, size arithmetic is on size_t spans with no user-multiplied quantities, and payload.size() != header.payloadLength rejects both truncation and extension.
Level of scrutiny
High. This is C++ in the JSC bindings, on a path that previously crashed the process from plain JS, and it defines a persistent binary format that becomes part of the user-visible createCachedData() contract. That combination — untrusted-input handling plus a new wire format — is exactly the kind of change a maintainer should eyeball, particularly the choice of header fields (e.g. whether source.hash() and computeJSCBytecodeCacheVersion() are the right staleness keys) and the VMMalloc allocator for the owned copy.
Other factors
Over several review rounds every finding was addressed: the module-path null-deref, the third sibling in getBytecodeBuffer, the cachedDataProduced flag inconsistency, the dead setter, and the comment-cop flags. CI was green across all completed lanes at 9ed4358 per the author's status comment. The corruption-matrix test spawns a child and asserts {stdout, stderr, exitCode, signalCode}, so a regression to the old crash would fail loudly. Deferring rather than approving because the format decision and the JSC lifetime reasoning deserve a maintainer's eyes, not because anything looks wrong.
a8baf5e to
cae82eb
Compare
|
Note GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer. |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@test/js/node/vm/vm.test.ts`:
- Around line 906-924: Update corruptions() to include explicit empty and
zero-filled cached-data variants using Buffer.alloc(0) and
Buffer.alloc(good.length), alongside the existing cases, so both boundary
representations are passed through rejection validation before header decoding.
- Around line 945-970: Extend the cached-data tests in the compileFunction and
SourceTextModule blocks to create cached data from a distinct source, then
supply it to the target source. Assert compileFunction rejects the mismatched
data and still executes normally, while SourceTextModule throws
ERR_VM_MODULE_CACHED_DATA_REJECTED; retain the existing corruption and
matching-data checks.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: a4eb3345-9772-4223-9fb9-6348c07f9cf9
📒 Files selected for processing (6)
src/jsc/bindings/NodeVM.cppsrc/jsc/bindings/NodeVM.hsrc/jsc/bindings/NodeVMScript.cppsrc/jsc/bindings/NodeVMScript.hsrc/jsc/bindings/NodeVMSourceTextModule.cpptest/js/node/vm/vm.test.ts
💤 Files with no reviewable changes (1)
- src/jsc/bindings/NodeVMScript.h
There was a problem hiding this comment.
I reviewed this PR again after 8507750 and didn't find any bugs. Because it changes memory-safety C++ in the JSC bindings on a security-relevant path (user-controlled bytes reaching decodeCodeBlock), a human look would still be worthwhile.
What was reviewed:
unwrapCachedData: header is read via memcpy after a size check, every field is compared before the payload copy, and the accepted payload is copied into an ownedMallocSpansocompileFunction/SourceTextModuleno longer hand JSC a span into a stack Vector.- The provided-vs-empty gate in all three consumers now keys on presence (
cachedDataProvided/!isUndefined());validateCachedDatareturns true for a zero-length view, soBuffer.alloc(0)reaches the header check and is rejected. createCachedDataBufferwrites header+payload into a fresh Uint8Array; thepayloadLengthnarrow touint32_tis fine for realistic bytecode sizes.- The subprocess test asserts
signalCode === nulland covers the full corruption matrix (including empty/zeros and cross-source) across Script, compileFunction, and SourceTextModule.
Extended reasoning...
Overview
The PR adds a 24-byte integrity header (magic, payload length, source hash, JSC bytecode-cache version, XXH3-64 payload hash) to the buffers createCachedData()/produceCachedData emit, and a matching unwrapCachedData() that verifies every field before JSC::decodeCodeBlock sees the payload. All three cachedData consumers — new vm.Script, vm.compileFunction, new vm.SourceTextModule — route through it, and the accepted payload is copied into memory owned by the CachedBytecode instead of borrowed. The latest commit (8507750) switches the consumers from gating on !cachedData.isEmpty() to gating on whether the option was provided, so an explicitly-passed empty buffer is rejected like Node does, and extends the test matrix with empty/zero-fill variants and cross-source rejection for compileFunction/SourceTextModule.
Security risks
This is exactly the security-relevant surface: user-controlled bytes previously flowed straight into JSC's bytecode decoder, which derives pointers from self-relative offsets in the payload with no bounds check — truncation produced wild heap reads and a SIGSEGV. The fix is the right shape (validate-before-decode with a checksum, mirroring V8's SerializedCodeData header), and the validated payload is copied so the decoder's lazy function-body decoding can't outlive the caller's buffer. I did not find a bypass: length is an exact match (catches both truncation and extension), and the XXH3-64 payload hash catches bit flips that leave length intact. The header is read via memcpy after a size < sizeof(CachedDataHeader) guard, so there's no unaligned or short read.
Level of scrutiny
High. This is hand-written C++ in src/jsc/bindings/, on a path where the pre-fix behavior was a native crash from user input, and it changes ownership of a buffer that JSC retains for lazy decoding. REVIEW.md's memory-safety and untrusted-input sections apply directly. The change is well-scoped and has been through ten iterations with every prior review finding addressed, but it is not the kind of mechanical change I'd approve without a human look.
Other factors
All prior inline findings from this reviewer and CodeRabbit are marked resolved and I verified the fixes in the current diff: the redundant m_options.cachedData retention is released after unwrapCachedData copies (cachedData = {}); the dead cachedDataProduced(true) setter was removed; the getBytecodeBuffer() null guard is defensive-only after #31827 (no deterministic fixture exists). The test follows harness conventions — subprocess with concurrent pipe drain, combined {stdout, stderr, exitCode, signalCode} assertion, deterministic corruption seeding — and the PR description's evidence block shows it fails on main (SIGSEGV/assertion) and passes on the branch under both ASAN debug and release.
8507750 to
9d6db6c
Compare
There was a problem hiding this comment.
I re-reviewed after the latest push and found no new issues — every point raised in earlier rounds (stderr in the test assertion, the three null-CachedBytecode deref sites, the stale cachedDataProduced override and dead setter, the retained m_options.cachedData copy, and empty-buffer-treated-as-absent) is now addressed in the squashed commit. Given this is C++ memory-safety hardening on an adversarial-input path into JSC's decoder and changes the on-disk cachedData format, a maintainer sign-off is still worthwhile.
What was reviewed
unwrapCachedData: header fields checked in fail-closed order before any payload deref;MallocSpancopy is owned by the returnedCachedBytecode, and the sourceVectoris cleared at the one call site that retained it.- All three consumers (
vm.Script,vm.compileFunction,vm.SourceTextModule) now gate on presence (cachedDataProvided/!isUndefined()) rather than!isEmpty(), soBuffer.alloc(0)is rejected — the test's["empty", …]variant covers it. createCachedDataBuffer: header + payload written into a singlecreateUninitializedBufferallocation; every producer (createCachedData,getBytecodeBuffer,NodeVMSourceTextModule::cachedData,compileFunctionproduce path) routes through it.
Extended reasoning...
Overview
The PR wraps node:vm cachedData in a 24-byte integrity envelope (magic / payloadLength / sourceHash / jscVersion / xxhash3-64 payloadHash) so that corrupted, truncated, extended, or mismatched buffers are rejected before reaching JSC::decodeCodeBlock, which does no bounds checking and previously crashed on truncated input. It touches NodeVM.cpp/.h (the shared createCachedDataBuffer/unwrapCachedData helpers and compileFunction path), NodeVMScript.cpp/.h (vm.Script construct/produce paths, dead setter removed), NodeVMSourceTextModule.cpp (module construct + createCachedData null guard), and adds a subprocess-based corruption-matrix test to test/js/node/vm/vm.test.ts. Since my last inline comment, a cachedDataProvided flag was added so a provided-but-empty buffer is rejected rather than treated as absent, and the test gained an ["empty", Buffer.alloc(0)] variant.
Security risks
This is a memory-safety hardening on an adversarial-input path: user-supplied bytes previously flowed straight into JSC's bytecode decoder, whose self-relative-offset format enabled wild heap reads on truncated input. The envelope check is fail-closed — every mismatch returns nullptr before any payload byte influences a pointer — and the accepted payload is copied into a MallocSpan owned by the CachedBytecode so lazy decoding cannot dangle after the caller's buffer is freed or mutated. payloadLength is compared against the actual remaining span rather than trusted, and the xxhash covers only the payload the length check already bounded. I did not find a bypass, but xxhash3 is a non-cryptographic checksum (integrity, not authentication) and the decoder itself remains unhardened until the referenced WebKit-side change lands, so the residual risk is a maintainer judgment call.
Level of scrutiny
High. This is hand-written C++ in the JSC bindings, the "most-blocked category" per REVIEW.md, on a path that consumes attacker-controllable bytes and hands them to a decoder that trusts its input. It also changes user-visible API surface: the cachedData blob format gains a 24-byte header, so blobs are no longer interchangeable with other consumers of raw JSC bytecode (acceptable per the PR's reasoning — Node's contract is same-runtime-version only — but a format decision a maintainer should ratify). The Node/Web-compat and API-design sections of .claude/docs/landing-prs.md both apply.
Other factors
This PR has been through many review iterations; every concrete issue I raised across prior rounds is verifiably fixed in the current diff (stderr/signalCode surfaced in the test assertion; null-RefPtr<CachedBytecode> guarded at all three producer sites; cachedDataProduced(true) override removed and its now-dead setter deleted; m_options.cachedData cleared after the owned copy is taken; empty-buffer routed through the rejection path with test coverage). The bug-hunt exit reason was dry_streak with no findings. No CODEOWNERS entry covers these paths and there are no outstanding CHANGES_REQUESTED reviews. Several github-actions[bot] inline threads were self-resolved by the author — I could not read their content, which is one more reason a human pass is worthwhile before merge.
cachedData was passed straight to JSC::decodeCodeBlock, which follows
offsets stored inside the buffer and is only safe on an intact copy of
its own serializer's output. A corrupted or truncated buffer crashed the
process (or was silently accepted), where Node sets cachedDataRejected
and recompiles.
createCachedData() and produceCachedData now prepend a 24-byte
{ magic, payloadLength, sourceHash, jscVersion, payloadHash } header,
and new vm.Script, vm.compileFunction, and new vm.SourceTextModule
verify every field before decoding. Anything that does not match takes
the existing rejection path. The accepted payload is copied into memory
owned by the CachedBytecode because decoded functions retain the decoder
for lazy code block decoding, and the input Vector is released once the
copy exists.
A provided-but-empty cachedData buffer is now rejected like Node does,
instead of being treated as absent. NodeVMSourceTextModule::cachedData()
surfaces a serialization failure as an error instead of dereferencing a
null RefPtr.
…portability test The portability snapshot masks JSC's cacheVersion and headerChecksum at offsets 0 and 8 so it only moves when the serialized bytes do. vm cachedData now carries a 24-byte integrity header in front of that payload, which shifted the masked fields and added version-dependent fields of its own. Strip the header for the vm entries; the existing snapshot hashes match unchanged, which confirms the payload bytes are identical to before.
b15b7f5 to
d9b4996
Compare
|
The three const vm = require("node:vm");
const src = "return a + 1234;";
const a = vm.compileFunction(src, ["a"], { produceCachedData: true });
const b = vm.compileFunction(src, ["a"], { cachedData: a.cachedData });
console.log(b(1));
The test added here does not show that. Its I opened #42229 with the lifetime half on its own: the same |
…-cacheddata-validate # Conflicts: # src/jsc/bindings/NodeVM.cpp # src/jsc/bindings/NodeVM.h # src/jsc/bindings/NodeVMScript.cpp # src/jsc/bindings/NodeVMSourceTextModule.cpp
…-cacheddata-validate
Problem
new vm.Script,vm.compileFunctionandnew vm.SourceTextModulegivecachedDatatoJSC::decodeCodeBlock()unchecked. One changed byte can crash:panic(main thread): Segmentation fault at address 0x818585C26A4. ASAN:heap-buffer-overflowinVariableLengthObject<StringImpl*>::isEmpty, orASSERTION FAILED: addr >= cachedBytecodeSpan.data()inJSC::Decoder::offsetOf.8cc039765e, one process per single-byte flip of an 888-byte blob: 262 die, 590 are accepted and run the damaged bytecode, 36 are rejected.Fix
createCachedData()andproduceCachedDataprepend a 24-byte header:{ magic, payloadLength, sourceHash, jscVersion, payloadHash }(XXH3-64 of the bytecode).unwrapCachedData(), which checks each field before the decoder sees the payload. A failed check takes the existing rejection path (cachedDataRejected === true, orERR_VM_MODULE_CACHED_DATA_REJECTED), and the source compiles normally.cachedDatawas never portable between builds, so no working cache breaks.test/js/node/vm/vm.test.ts. The fix rejects 2464 of 2464 flips and 616 of 616 truncations. Alsobundler_bytecode_portable.test.tsand Node's vm cachedData tests.Background
cachedDatais a serializedUnlinkedCodeBlock: JSC bytecode not yet linked to a global object.this + offsetand reads. It is built for caches Bun writes itself.cachedDataRejected. Release Node skips the checksum unless--verify-snapshot-checksumis set, so a flipped byte crashes it too.Notes
Reproduction. One process per offset, because a crash ends the process:
Node, for comparison. V8 does not read
cachedDatawhen the same source was already compiled in the process, so the producer and the consumer must be separate processes. Node v26.3.0, a 744-byte blob written by one process, one consumer process per mutation:[2,33423364,6]in place of[2,4,6]), 23 threw a JS error, 264 crashed (V8 fatal error,SIGSEGV, or "V8 sandbox violation detected").--verify-snapshot-checksum, five of the offsets that crash or give a wrong result by default are all rejected with the right result. The flag is off by default in release builds.So Node's default catches truncated, foreign and stale blobs, and does not catch a flipped byte. This PR catches both. The cost is one XXH3-64 pass over the blob: 0.35 ms for the 11.8 MB
cachedDataoftypescript.js(Bun.hash.xxHash3, release build, this machine), next to about 10 ms fornew vm.Script(src, { cachedData })of the same file on main.Sweep, same machine, same tree. Debug build with ASAN,
Malloc=1so that ASAN sees JSC's allocations, one fresh process per mutation. "Main" is this branch withsrc/from main8cc039765e(WebKit63a807e88ce0). Only the fiveNodeVM*files differ.vm.Script, every single-byte flipsrc/(888-byte blob)cachedDataRejected === false), exit 0The 262 deaths on main: 124 engine assertions, 77 ASAN
heap-buffer-overflow, 32 ASANSEGV, 21SHOULD NEVER BE REACHED, 1 ASANheap-use-after-free, 3 killed by a 60 s bound, 4 other. The 590 accepted blobs happened to give the right result for this source. Nothing guarantees that: the engine links and runs bytecode that differs from what it wrote.This branch, all three entry points, every single-byte flip plus every truncation at 0, 4, 8, ... bytes:
new vm.Scriptvm.compileFunctionnew vm.SourceTextModuleEvery process exits 0 with the right result and an empty stderr. A rejected module means
ERR_VM_MODULE_CACHED_DATA_REJECTED.Truncation. This PR first showed the bug with truncated blobs (a partial write or a stale file), which crashed main at that time:
new vm.Script(src, { cachedData: cd.subarray(0, cd.length >> 1) }). #43616 brought in oven-sh/WebKit#707, where a payload records its size and a shorter one is a cache miss. So main now rejects truncated blobs. The same bump brought in oven-sh/WebKit#677, which removed the per-record damage checks from the decoder on purpose: "a payload is trusted as a whole". Interior damage is therefore the caller's problem, andnode:vmis the one caller that accepts bytes from user code. The header'spayloadLengthstill rejects truncation and extension before any hashing.Why the check is not in the decoder. An earlier attempt added bounds checks to
Decoder::offsetOf,ptrForOffsetFromBaseandCachedPtr::decode(oven-sh/WebKit#368). It is closed: oven-sh/WebKit#677 went the other way for decode speed, and bounds checks do not catch a flipped byte that stays in range. A checksum over the whole payload does.Header fields.
magicrejects buffers that Bun did not write.payloadLengthrejects truncation and extension.sourceHashrejects a blob for other source text.jscVersioniscomputeJSCBytecodeCacheVersion()and rejects a blob from another build.payloadHashrejects interior damage. An accepted payload is copied into memory that theCachedBytecodeowns (createOwnedCachedBytecode(), from #42229), because JSC decodes function bodies lazily from it.Empty buffer. A provided but empty
cachedDatais rejected, as in Node. Main treats it as absent.Merges with main.
--compile --bytecodeexecutables work when cross-compiled (portable JSC bytecode cache) #40270 (portable bytecode cache,getBytecodetakes aSourceCodeType):constructScriptkeys its cachedData block onscript->options().cachedDataProvidedinside main's restructured constructor. Main had added thegetBytecodeBuffer()null guard and dropped thecachedDataProduced(true)override on its own, so those parts of this PR collapsed into main's versions.bundler_bytecode_portable.test.ts): kept main's shape for both, routed throughunwrapCachedDataand thevmPayloadhelper.NodeVM::createOwnedCachedBytecode, called by each entry point. The merge keeps that helper, makes it file-local, and hasunwrapCachedDatacall it after the header checks. No entry point can hand unverified bytes to the decoder.63a807e88ce0): no conflict. The decoder and the payload layout changed under this PR, so everything above was run again on that merge.test/bundler/bundler_bytecode_portable.test.tsfingerprints vm cachedData. The vm entries strip the 24-byte header first (vmPayload). Main's snapshot, including the values Bump WebKit to 63a807e88ce0: cheaper embedded-bytecode decode and CodeBlock linking #43616 updated, matches with no change from this PR. That also shows the JSC payload behind the header is byte-identical to main's.Tests.
rejects corrupted cachedData instead of crashingintest/js/node/vm/vm.test.tsruns truncations atlen-1,len/2, 16 and 1, 20 single-byte flips spread over the buffer, an extended buffer, a filled buffer, unrelated bytes and an intact blob for other source, through all three entry points in a child process. Against main'ssrc/it fails at the first flip that main accepts (AssertionError: flip@64,false !== true). Local runs on the merge with main8cc039765e, debug build with ASAN:vm.test.ts303 pass and 0 fail,bundler_bytecode_portable.test.ts22 pass and 0 fail (with a long timeout, the debug build is slow), Node'stest-vm-cached-data.js,test-vm-createcacheddata.js,test-vm-module-cached-data.jsandtest-vm-basic.jspass.no test proof · iteration 21 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/node/vm/vm.test.ts, test/bundler/bundler_bytecode_portable.test.ts