Skip to content

CodeCache: keep unlinked code blocks decoded from cached bytecode in the in-memory cache - #513

Merged
dylan-conway merged 5 commits into
mainfrom
claude/codecache-remember-decoded
Aug 25, 2026
Merged

dylan-conway merged 5 commits into
mainfrom
claude/codecache-remember-decoded

Conversation

@dylan-conway

@dylan-conway dylan-conway commented Aug 24, 2026 •

Copy link
Copy Markdown
Member

CodeCacheMap::findCacheAndUpdateAge only populated the map on the generate (parse) path; a block decoded from a SourceProvider's cached bytecode via fetchFromDisk was returned without being added. When several globals in one VM load the same source from a bytecode cache (ShadowRealm, vm contexts, bun build --compile --bytecode executables), every lookup missed and decoded a fresh UnlinkedProgram/ModuleProgramCodeBlock + UnlinkedFunctionExecutable tree. Insert the decoded block so later lookups reuse it, matching the generate path (and, like it, only when useCodeCache()); addCache prunes as usual, so cache bounds are unchanged. A single-global program pays one map entry: the decoded tree is already kept alive by FunctionExecutable::m_topLevelExecutable → GlobalExecutable::m_unlinkedCodeBlock while any of its functions are.

This is not Bun-specific — the same asymmetry exists in stock JSC's disk-cache path (JSC_diskCachePath) — so it is left unguarded; a USE(BUN_JSC_ADDITIONS) guard could be added if we prefer to keep the fork delta explicit.

Measured in the jsc shell, x64 Linux release: a 2000-function script loaded by the main global and then by three $262.createRealm() globals, second run from JSC_diskCachePath with JSC_forceDiskCache=1:

main this PR
decodes from cached bytecode (verboseDiskCache "Cache hit") 8 3
GC heap after fullGC() 5.42 MB 3.70 MB
max RSS 49.8 MB 38.9 MB

In a compiled --bytecode executable importing a 2000-function module into 3 ShadowRealms: UnlinkedFunctionExecutable +2005/+4006/+6007 → stays at +2005; UnlinkedFunctionCodeBlock likewise stops scaling with the number of globals.

JSTests/stress through bytecode-cache-test-helper.sh, with and without --diskCachePayloadIsPersistentForTesting=1: same results as main.

…the in-memory cache

CodeCacheMap::findCacheAndUpdateAge only populated the map on the generate
(parse) path; a block decoded from a SourceProvider's cached bytecode via
fetchFromDisk was returned without being added. When several globals in one VM
load the same source from a bytecode cache (ShadowRealm, vm contexts,
bun build --compile --bytecode executables), every lookup missed and decoded a
fresh UnlinkedProgram/ModuleProgramCodeBlock + UnlinkedFunctionExecutable tree.
Insert the decoded block so later lookups reuse it, matching the generate path;
addCache prunes as usual, so cache bounds are unchanged.

This is not Bun-specific -- the same asymmetry exists in stock JSC's disk-cache
path (JSC_diskCachePath) -- so it is left unguarded.
@coderabbitai

coderabbitai Bot commented Aug 24, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

  • Run on-demand review

On-demand reviews are free for the next 26 days. After that, they cost $0.25 per reviewed file.

Or wait 52 minutes for your next included review.

View limit details

Limit details: You’ve used the included review currently available. Your 74 included PR review attempts over the past 7 days set your current allowance at 1 review per hour.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

Run ID: 682c1d9c-c4ba-421e-a410-4874f3562895

📥 Commits

Reviewing files that changed from the base of the PR and between 45040b8 and f3475ef.

📒 Files selected for processing (1)
  • Source/JavaScriptCore/runtime/CodeCache.h

Comment @coderabbitai help to get the list of available commands.

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — small, self-contained change that mirrors the existing generate-path insertion.

What was reviewed:

  • Confirmed the new addCache(key, SourceCodeValue(vm, decoded, m_age)) call matches the generate path in getUnlinkedGlobalCodeBlock (CodeCache.cpp:230), including the Options::useCodeCache() gate.
  • Checked addCache's ASSERT(isNewEntry) — safe because m_map.find(key) just returned end(), and prune() inside addCache runs before the insert.
  • Verified fetchFromDisk returns nullptr for UnlinkedEvalCodeBlock and UnlinkedFunctionExecutable, so the new insert only fires for program/module blocks as intended.
Extended reasoning...

Overview

This PR modifies CodeCacheMap::findCacheAndUpdateAge in Source/JavaScriptCore/runtime/CodeCache.h. Previously, when the in-memory map missed and fetchFromDisk decoded an unlinked code block from the provider's cached bytecode, that block was returned but never inserted into the map — so subsequent lookups for the same key (e.g. from another global/realm in the same VM) would miss again and decode a fresh copy of the entire unlinked tree. The change inserts the decoded block via addCache, exactly as the generate (parse) path already does in CodeCache::getUnlinkedGlobalCodeBlock.

Security risks

None identified. This is an in-process, same-VM caching change: the cache key (SourceCodeKey) is unchanged, the cached value is the same object that was already being returned to the caller, and the Strong<JSCell> handle keeps it GC-safe. No new data crosses trust boundaries; the decoded block was already trusted enough to be returned and linked.

Level of scrutiny

Moderate. This is core JSC runtime code, but the change is 7 lines and is a direct mirror of the established pattern at CodeCache.cpp:230 (m_sourceCode.addCache(key, SourceCodeValue(vm, unlinkedCodeBlock, m_sourceCode.age()))), gated on the same Options::useCodeCache() check. I verified: (a) addCache's ASSERT(addResult.isNewEntry) holds because find just returned end() and prune() runs before the insert; (b) the m_age argument matches what the generate path passes via age(); (c) fetchFromDisk's if constexpr returns nullptr for UnlinkedEvalCodeBlock and UnlinkedFunctionExecutable, so only program/module blocks reach the new insert; (d) the raw decoded pointer stays valid after addCache — no GC runs, and the entry is held by a Strong<>.

Other factors

The PR description includes concrete measurements (decode count 8→3, GC heap 5.42→3.70 MB, RSS 49.8→38.9 MB in a multi-realm scenario) and reports JSTests/stress via bytecode-cache-test-helper.sh passes with and without --diskCachePayloadIsPersistentForTesting=1. Cache growth is bounded by the existing prune() machinery, so memory bounds are unchanged. The bug hunting system found no issues. Given the minimal surface area, exact symmetry with existing code, and clear testing, I'm comfortable approving.

@github-actions

github-actions Bot commented Aug 24, 2026 •

Copy link
Copy Markdown

Preview Builds

Commit Release Date
f3475ef1 autobuild-preview-pr-513-f3475ef1 2026-08-25 14:49:12 UTC
7c7e09af autobuild-preview-pr-513-7c7e09af 2026-08-24 20:01:35 UTC

@claude claude Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed this PR and didn't find any bugs. The change is small and mirrors the existing generate-path insert exactly, but it alters retention behavior on a hot JSC runtime path (every top-level script/module load), and the PR leaves the USE(BUN_JSC_ADDITIONS) guard question open — a human sign-off on both would be worthwhile.

What was reviewed:

  • Confirmed the caller returns early on a non-null findCacheAndUpdateAge result, so the generate path's own addCache at CodeCache.cpp:230 is not reached — no double insert / ASSERT(isNewEntry) trip.
  • fetchFromDisk<UnlinkedFunctionExecutable> is compile-time nullptr (not UnlinkedCodeBlock-derived), so the getUnlinkedGlobalFunctionExecutable instantiation is unaffected.
  • findResult is not used after addCache, so the map mutation can't touch a stale iterator; the extra prune() inside addCache only removes, keeping isNewEntry true.
Extended reasoning...

Overview

The PR changes 7 lines in CodeCacheMap::findCacheAndUpdateAge (Source/JavaScriptCore/runtime/CodeCache.h). On an in-memory miss, the block decoded from the provider's cached bytecode via fetchFromDisk is now inserted into m_map with addCache(key, SourceCodeValue(vm, decoded, m_age)) before being returned, gated on Options::useCodeCache(). This matches how the generate path in CodeCache::getUnlinkedGlobalCodeBlock already inserts freshly-parsed blocks.

Security risks

None identified. The change adds a Strong<JSCell> reference to an already-allocated unlinked code block into an existing bounded cache; it does not parse untrusted input, cross a trust boundary, or change key derivation. Cache bounds remain governed by the existing prune() machinery.

Level of scrutiny

Medium-high. The diff is tiny and pattern-matched against existing code, but CodeCacheMap sits on the path for every program/eval/module load in JSC. A mistake here (double-insert, iterator invalidation, unbounded retention) would affect all execution. I traced the two callers in CodeCache.cpp:

  • getUnlinkedGlobalCodeBlock: on a non-null return it takes the if (unlinkedCodeBlock && useCodeCache()) early-return branch and never reaches its own addCache, so the ASSERT(addResult.isNewEntry) in addCache cannot fire from a duplicate.
  • getUnlinkedGlobalFunctionExecutable: instantiates with UnlinkedFunctionExecutable, for which fetchFromDisk's if constexpr (std::is_base_of_v<UnlinkedCodeBlock, ...>) is false → decoded is always null → no new behavior.
  • addCache calls prune() before m_map.add; prune only removes entries, and findResult (== end()) is not dereferenced afterward, so there is no iterator hazard.

Other factors

The PR description is thorough (benchmarks, JSTests/stress parity via bytecode-cache-test-helper.sh). It also explicitly leaves open whether to wrap this in USE(BUN_JSC_ADDITIONS) to keep the fork delta explicit — that is a maintainer preference call rather than a correctness question, and is one reason a human should weigh in. Given the critical-path location plus that open style question, deferring rather than auto-approving.

… pure lookup again, getUnlinkedGlobalCodeBlock decodes then addCache()s, mirroring generate then addCache()
…except for remembering what fetchFromDisk decoded
@dylan-conway
dylan-conway merged commit 1cb96a7 into main Aug 25, 2026
44 checks passed
dylan-conway added a commit to oven-sh/bun that referenced this pull request Aug 25, 2026
robobun added a commit to oven-sh/bun that referenced this pull request Aug 25, 2026
Main moved WEBKIT_VERSION to 1cb96a7b (oven-sh/WebKit#513).
oven-sh/WebKit#268 is rebased onto that commit, so its preview carries
everything main's pin has plus the two async context fixes.
robobun added a commit to oven-sh/bun that referenced this pull request Aug 26, 2026
Main moved WEBKIT_VERSION to 1cb96a7b (oven-sh/WebKit#513).
oven-sh/WebKit#268 is rebased onto that commit, so its preview carries
everything main's pin has plus the two async context fixes.
robobun added a commit to oven-sh/bun that referenced this pull request Aug 26, 2026
Main moved WEBKIT_VERSION to 1cb96a7b (oven-sh/WebKit#513).
oven-sh/WebKit#268 is rebased onto that commit, so its preview carries
everything main's pin has plus the two async context fixes.
robobun added a commit to oven-sh/bun that referenced this pull request Aug 28, 2026
Main moved WEBKIT_VERSION to 1cb96a7b (oven-sh/WebKit#513).
oven-sh/WebKit#268 is rebased onto that commit, so its preview carries
everything main's pin has plus the two async context fixes.
robobun added a commit to oven-sh/bun that referenced this pull request Aug 28, 2026
Main moved WEBKIT_VERSION to 1cb96a7b (oven-sh/WebKit#513).
oven-sh/WebKit#268 is rebased onto that commit, so its preview carries
everything main's pin has plus the two async context fixes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant