Skip to content

s3: fix double free of a multipart part's bytes on worker terminate - #43709

Open
robobun wants to merge 5 commits into
mainfrom
robobun/1bc79006/s3-multipart-part-owns-buffer
Open

robobun wants to merge 5 commits into
mainfrom
robobun/1bc79006/s3-multipart-part-owns-buffer

Conversation

@robobun

@robobun robobun commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • A streaming S3 multipart upload frees one part's bytes twice when its worker is terminated. ASAN: attempting double-free ... in thread T.. (Worker), reported in MultiPartUpload::process_multi_part (src/runtime/webcore/s3/multipart.rs:991), freed first by UploadPart::cancel (multipart.rs:273).
  • When the buffer holds exactly one part, process_multi_part gives the new part a raw pointer to the buffer's bytes. It gives the allocation up with ManuallyDrop only after enqueue_part returns Ok. An Err skips that and leaves two owners.
  • terminate() produces that Err. The client refuses the request for a stopping VM, the failure frees the part's bytes, and the error unwinds through the buffer's Drop.

Fix

  • UploadPart::data is an Rc<Vec<u8>>. get_create_part takes the bytes out of the buffer itself (all of them, or a copy of the first part), and only when the queue has a slot. The raw pointer, allocated_size, needs_clone and the ManuallyDrop are gone.
  • perform holds a count while it lends the bytes to a request, because a request that fails before it leaves frees the part inside that call.
  • Self-reviewed: 3 concerns raised, 3 addressed.
  • Verified: test/js/bun/s3/s3-upload-terminate.test.ts (20 of 20 runs abort without this change, clean with it). Also test/js/bun/s3/ and the worker_threads.test.ts teardown rows.

Background

Notes

Repro. A host process runs a loopback S3 stand-in and 4 workers. Each worker starts 8 uploads with client.write(key, new Response(stream), { partSize: 5 MiB, queueSize: 1, retry: 0 }), where the stream's first chunk is exactly one part, then posts a message. The host terminates each worker as its message arrives. The first write of an upload is the only one that sends the request that creates the upload, so the terminate has to land before it. With several uploads per worker and several workers that happens on almost every run: 20 of 20 runs of the unfixed debug build aborted (19 of 20 with 4 uploads per worker). Every run with the change was clean. The test takes 1.8 s on the debug build.

The failing path, from the ASAN report on the unfixed build. Second free: process_multi_part (multipart.rs:991, the end of the one-big-chunk block, where the local StreamBuffer drops) from process_buffered from write from NetworkSink::write from the stream pump (rsisSinkWrite), in the worker's ordinary microtask drain. First free: free_allocated_slice (:273) from UploadPart::cancel (:393) from fail (:594) from start_multi_part_request_result (:689) from Callback::fail from execute_simple_s3_request (simple_request.rs:603, the branch for a VM that sends nothing new) from enqueue_part (:905) from process_multi_part (:967). So the refused request is the one that creates the upload, for the upload's first part.

Why only that block. The other block of process_multi_part copies a slice of the buffer for the part, so the part and the buffer never share an allocation. An Err there leaves the buffer's cursor where it was, which does not matter: every Err out of enqueue_part comes from fail, which finishes the upload, and a finished upload sends nothing more.

Release builds. A release build has no such check. The second free hands mimalloc a block it already has back, and what follows is undefined. Rare segfaults of a release build under this workload are consistent with that, but this change does not prove the link.

The bytes perform lends. The first revision of this PR kept a Vec<u8> in the part, and perform passed a slice of it to execute_simple_s3_request. That call reports a request that fails before it leaves from inside itself, and the report frees the part, so the slice pointed at freed memory until the call returned. Nothing read it, and the previous raw pointer had the same shape, but only a comment said so. The bytes are counted now and perform holds a count for the length of the call. The queue's Drop frees what a part still owns, which the raw pointer never did.

What this does not change. The leak of an upload that is still open when its worker goes (#39692) is separate, which is why the test's child process runs with detect_leaks=0. Five sibling S3 tests clear HTTP_PROXY and HTTPS_PROXY but not ALL_PROXY, which the client falls back to. This PR's test clears all of them, and the siblings are a separate change.

Other suites. test/js/bun/s3/: 160 pass, and s3-list-objects.test.ts "Should fall back to NoSuchKey ..." times out on the debug build when the whole file runs. It passes alone, it exercises client.list() only, and #38353 recorded the same timeout on an unmodified debug build. worker_threads.test.ts -t "VM teardown": 3 pass. test/internal/source-lints/: 188 pass. cargo clippy -p bun_runtime --no-deps, cargo fmt --check and prettier are clean.


[human-review] gate passed · iteration 0 · 2 files touched

fails on main (without fix)
ASAN without fix: 1 FAILED
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/pr_gate.xml" "test/js/bun/s3/s3-upload-terminate.test.ts"
bun test v1.4.3 (367d939d9)

test/js/bun/s3/s3-upload-terminate.test.ts:
killed 1 dangling process
(fail) terminate() while a multipart upload enqueues a part frees the part's bytes once [5005.80ms]
  ^ this test timed out after 5000ms.

 0 pass
 1 fail
Ran 1 test across 1 file. [7.00s]
error: script "bd" exited with code 1
__F:1:S:0

release without fix: all passed
bun test v1.4.3-canary.1 (367d939d9)

test/js/bun/s3/s3-upload-terminate.test.ts:
(pass) terminate() while a multipart upload enqueues a part frees the part's bytes once [108.47ms]

 1 pass
 0 fail
 2 expect() calls
Ran 1 test across 1 file. [242.00ms]
__F:0:S:0
passes on PR (with fix)
ASAN with fix: all passed
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/pr_gate.xml" "test/js/bun/s3/s3-upload-terminate.test.ts"
bun test v1.4.3 (367d939d9)

test/js/bun/s3/s3-upload-terminate.test.ts:
(pass) terminate() while a multipart upload enqueues a part frees the part's bytes once [3218.06ms]

 1 pass
 0 fail
 2 expect() calls
Ran 1 test across 1 file. [5.46s]
__F:0:S:0

release with fix: all passed
$ bun scripts/build.ts --profile=release
[configured] bun-profile → bun (stripped) in 1460ms (unchanged)
ninja: Entering directory `/workspace/bun/build/release'
[1/213] gen generated_host_exports.rs
generated_host_exports.rs: 121 exports (host=5, lazy=10, generic=106, rust=0); 245 extern-C blocks audited
[2/213] gen cpp.rs (cppbind)
[3/213] gen JS modules (bundle-modules)
Preprocess modules (8304ms)
Bundle modules (132ms)
Postprocesss modules (250ms)
Bundle Functions (534ms)
Generate Code (66ms)

[9.30s] Bundled "src/js" for production
  2595 kb
  197 internal modules
  13 native modules
  50 internal functions across 16 files
[4/111] cxx obj/unified/UnifiedSource-src_jsc_bindings_node-0.cpp.o
[5/111] build.rs build_script_build
[6/111] cxx obj/unified/UnifiedSource-src_jsc_bindings-0.cpp.o
[7/111] cxx obj/unified/UnifiedSource-src_jsc_bindings_webcore-2.cpp.o
[8/111] cxx obj/unified/UnifiedSource-src_jsc_bindings-4.cpp.o
[9/111] cxx obj/unified/UnifiedSource-src_jsc_bindings-3.cpp.o
[10/111] cxx obj/unified/UnifiedSource-src_jsc_bindings-1.cpp.o
[11/111] cxx obj/unified/UnifiedSource-src_jsc_bindings-2.cpp.o
[12/111] cxx obj/codegen/ZigGeneratedClasses.cpp.o
[13
... (truncated)
diff hotspot
src/runtime/webcore/s3/multipart.rs        | 125 ++++++++++-------------------
 test/js/bun/s3/s3-upload-terminate.test.ts | 112 ++++++++++++++++++++++++++
 2 files changed, 156 insertions(+), 81 deletions(-)

gate history · 1 passed · 0 rejected · iteration 0

evidence per changed file
file                                        reads  edits  tests
src/runtime/webcore/s3/multipart.rs            14     26     16
test/js/bun/s3/s3-upload-terminate.test.ts      2      5     16

An upload's `process_multi_part` took the buffered bytes out of the upload
and handed the new `UploadPart` a raw pointer to them, then transferred the
allocation with `ManuallyDrop` only after `enqueue_part` returned `Ok`. An
`Err` skipped that transfer, so the part and the local `StreamBuffer` both
owned the bytes: the upload's failure freed them through the part, and the
buffer freed them again as the error unwound.

The part now holds a `Vec<u8>` it takes when it is created, so the buffer
gives the bytes up at the same point the part receives them. The queue's
`Drop` frees what a part still owns, which the raw pointer never did.
@robobun

robobun commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator Author

Status: the fix and its test are pushed. This PR is ready for review.

How I reproduced it. A host process runs a loopback S3 stand-in and 4 workers. Each worker starts 8 uploads with client.write(key, new Response(stream), { partSize: 5 MiB, queueSize: 1, retry: 0 }). The first chunk of each stream is exactly one part. The host terminates each worker when the worker reports that its uploads are started.

  • Debug ASAN build of main: 20 of 20 runs abort with AddressSanitizer: attempting double-free ... in thread T.. (Worker) in MultiPartUpload::process_multi_part.
  • The same build with this change: every run is clean.

The test is test/js/bun/s3/s3-upload-terminate.test.ts. Run it with bun bd test test/js/bun/s3/s3-upload-terminate.test.ts.

@coderabbitai

coderabbitai Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Review Change StackReview Change Stack

Understand this PR’s impact

Explore downstream dependencies and potential security impact with Blast Radius.

View blast radius →

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: oven-sh/bun/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: e881db72-4cbe-41d9-aa56-718a70a97713

📥 Commits

Reviewing files that changed from the base of the PR and between e0bca9c and 0765a97.

📒 Files selected for processing (1)
  • test/js/bun/s3/s3-upload-terminate.test.ts

Included review availability: Your plan provides up to 10 included reviews per hour; 0 remain after this review.


Walkthrough

Multipart upload parts now own payloads through reference-counted vectors. Enqueueing selects a queue slot before moving or copying data. Request and cleanup paths release payload references. A worker-termination regression test verifies clean process completion.

Changes

Multipart payload ownership

Layer / File(s) Summary
Owned part lifecycle
src/runtime/webcore/s3/multipart.rs
UploadPart stores payloads as reference-counted vectors. Cancellation, completion, failure, and drop paths release the payloads.
Deferred enqueue and buffer processing
src/runtime/webcore/s3/multipart.rs
Part creation selects a queue slot before moving complete buffers or cloning selected ranges. Full queues leave buffered data untouched.
Worker termination regression coverage
test/js/bun/s3/s3-upload-terminate.test.ts
The test terminates workers during multipart enqueueing and requires ok, empty stderr, and exit code 0 from the subprocess.

Priority: ➖ Normal

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the S3 multipart upload double-free fix that is the main change.
Description check ✅ Passed The description explains the problem, fix, affected code paths, regression test, verification results, and known limitations. It does not use the exact template headings, but it provides the required …

Comment @coderabbitai help to get the list of available commands.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline findings, I also checked the removed queue-full write-back of owned into buffered in process_multi_part — it is behavior-preserving, since take_data never runs when get_create_part finds no free slot, so the buffer is left untouched (and cursor == 0 && size() == len means take().list is exactly the part's bytes). The retry-recursion and queueSize > 1 post-failure allocation paths were examined and ruled out as regressions introduced here.

Extended reasoning...

The change replaces UploadPart's raw-pointer/allocated_size ownership with a Vec moved in via a closure that only runs once a queue slot exists, removing the ManuallyDrop and Vec::from_raw_parts free in src/runtime/webcore/s3/multipart.rs, plus a new worker-terminate ASAN repro test. It touches no auth/crypto surface. Two inline findings (a dangling body slice inside the S3 client on the stopping-VM path and an incomplete proxy env scrub in the test) remain for a human to weigh, which is why this is not an approval.

Additional findings (outside the current diff — GitHub can't attach inline comments there):

  • 🟣 src/runtime/webcore/s3/multipart.rs — A part's request body slice is left dangling inside the S3 client whenever a worker terminates mid-upload, the exact case this PR targets. perform() at multipart.rs:357 passes self.data() (a &[u8] into the JsCell's Vec) into execute_simple_s3_request; on the stopping-VM branch that call runs on_part_response synchronously, which calls free_data() at multipart.rs:303 and drops the Vec through with_mut while the caller's body borrow is still live. REVIEW.md forbids a slice outliving its memory and comment-enforced safety; the only guard is the doc comment at multipart.rs:268. Fix: perform() must not lend the Vec across a call that can free it: take the Vec out (or clone into the request's Box body first) and only free on the non-reentrant path.

    Why this was flagged

    Trigger: any streaming multipart upload whose part request starts while the VM is stopping (worker.terminate(), the PR's own test) or whose sign_request fails. Entry: process_multi_part -> enqueue_part -> part.start() -> perform() (multipart.rs:335-365). perform() builds S3RequestOptions with body: self.data() at multipart.rs:357; data() at :271 returns &[u8] derived from JsCell::get(), a shared borrow of the Vec. execute_simple_s3_request (simple_request.rs:601-608) calls callback.fail synchronously; on_part_response then calls free_data() at multipart.rs:303 (or :285), which is JsCell::set -> with_mut -> mem::replace (js_cell.rs:83-103), forming &mut Vec and deallocating the buffer while options.body in the outer frame still points at it. On the base branch the same shape existed with a raw pointer; the PR turned it into a Rust reference, so it is now an aliasing violation and a dangling reference, not just a stale raw pointer. Nothing reads options.body after the fail today, so the guard is only the comment at :268-269; any future read (logging, verbose fetch, a body-length…

    Verification: pre-existing. acknowledged in diff: multipart.rs:268-269 ("The request perform hands these bytes to copies them, and can fail the upload before it returns, which releases them") and the PR description ("Left alone: perform still lends the part's bytes to a request that can free them. Nothing reads them after") — the stated bound is accurate. Trigger: a part's request starts while the VM…

Comment thread test/js/bun/s3/s3-upload-terminate.test.ts
`UploadPart::perform` lent the part's bytes to `execute_simple_s3_request`,
which reports a request that fails before it leaves inside that call. The
report frees the part, so the lent slice pointed at freed memory until the
call returned. Nothing read it, and only a comment said so.

The bytes are now an `Rc<Vec<u8>>`. `perform` holds a count for the length of
the call, so the report drops the part's count and the bytes go when the call
returns.

The test clears `ALL_PROXY` as well, which the S3 client falls back to, and
its comment no longer says that the client ignores `NO_PROXY`.
Comment thread src/runtime/webcore/s3/multipart.rs Outdated
Comment thread src/runtime/webcore/s3/multipart.rs Outdated
Comment thread src/runtime/webcore/s3/multipart.rs Outdated
Comment thread src/runtime/webcore/s3/multipart.rs Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I re-reviewed after the second commit and found no bugs; the earlier ALL_PROXY/comment point is addressed. Because this reworks ownership of a part's bytes in a native path with synchronous re-entrant failure callbacks (and overlaps a pending rewrite of process_multi_part), a human look would still be worthwhile.

What was reviewed:

  • Every former free_allocated_slice site now calls free_data; the queue Drop releases whatever a part still holds, and cancel only frees in the Pending state as before.
  • perform's local Rc clone: execute_simple_s3_request copies the body into Box<[u8]> before returning, so the borrowed slice is never stored past the call, and the clone keeps it valid if callback.fail runs on_part_response synchronously.
  • take_data runs only after the queue slot is reserved, so the Ok(false) path leaves buffered untouched; buffered.take().list is exact because the branch requires cursor == 0 && size() == len; the slice closure's buffered.get() borrow ends before any buffered mutation.
  • Test: hermetic loopback stand-in, proxy vars scrubbed including ALL_PROXY, pipes drained concurrently, output asserted before the exit code.
Extended reasoning...

The change replaces a raw-pointer-plus-length ownership scheme for S3 multipart part bytes in src/runtime/webcore/s3/multipart.rs with JsCell<Option<Rc<Vec>>>, moves the buffer-to-part handoff into a closure that runs only once a queue slot exists, and adds a worker-terminate regression test. It touches no auth, injection, or data-exposure surface; the S3 signing path is unchanged. The diff is small and I traced every free/acquire path and the request layer's body copy, but it is native lifecycle code with synchronous re-entrant failure callbacks, the most-blocked category in this repo's review history, and the description notes another open PR rewrites the same function, so coordination is a maintainer call rather than an approval I can make alone.

One line each. The field's comment says why the bytes are counted, so
`perform` does not repeat it.
@robobun

robobun commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator Author

The review findings are addressed.

  • The bytes perform lends (77c5c0a). A part's bytes are an Rc<Vec<u8>> now, and perform holds a count for the length of the call. A request that fails before it leaves still frees the part inside that call, and the lent slice stays valid until the call returns. The other requests in this file lend buffers that belong to the upload (buffered, multipart_upload_list, path). The failure path does not release those, so they are unchanged.
  • The test's proxy scrub (77c5c0a). The test clears ALL_PROXY and all_proxy too, and its comment no longer says that the client ignores NO_PROXY. Five sibling S3 tests have the same gap. That is a separate change.
  • Multi-line comments (0c3bdf3). Each is one line now.

I ran test/js/bun/s3/s3-upload-terminate.test.ts, the S3 upload and stream tests, and the worker_threads.test.ts teardown rows again on the debug ASAN build. All pass.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I re-reviewed the current head (through 0c3bdf3) and didn't find any bugs; the proxy-env point from my earlier inline comment is addressed in 77c5c0a. Because this rewrites native memory ownership in the S3 multipart path and the PR notes #31733 rewrites the same function, a human look would still be worthwhile.

What was reviewed:

  • Traced both process_multi_part branches: take_data runs only after get_create_part has reserved a slot, so on Ok(false) the buffer is untouched and on Err the part is the sole owner of the taken Vec (the old re-append dance is correctly gone).
  • Checked the perform keepalive: execute_simple_s3_request copies options.body into task.body before the HTTP thread sees it, so the Rc clone only needs to cover the synchronous callback.fail path, which it does.
  • Confirmed no remaining references to allocated_size/free_allocated_slice/needs_clone in src/runtime/webcore/s3/, and that detect_leaks=0 in a spawned fixture with a stated reason matches existing tests (bun-write, compression, websocket).
  • Not done here: I did not build or run the new test, so the "aborts on the unfixed build" claim rests on the PR description and CI.
Extended reasoning...

The change touches one native file, src/runtime/webcore/s3/multipart.rs, replacing UploadPart's raw-pointer plus allocated_size ownership (freed via unsafe Vec::from_raw_parts and ManuallyDrop) with JsCell<Option<Rc<Vec>>> and a take_data closure that only runs once a queue slot exists, plus a new spawned ASAN fixture test that terminates workers mid-enqueue. It touches no auth, injection, or data-exposure surface; the S3 signing and request code is unchanged. The bug-hunt ran dry with no findings, the diff removes unsafe code rather than adding it, and my earlier inline finding was addressed by a follow-up commit. I deferred rather than approved because this is a memory-ownership rewrite in native code (the most-blocked category per REVIEW.md), I could not compile or run the test in this environment, and the author notes an in-flight PR (#31733) rewriting the same function that a maintainer needs to sequence.

`get_create_part` and `enqueue_part` were generic over the closure that
produced the part's bytes, so each was compiled once per call site although
almost none of their code used the closure. The bytes always come from the
upload's buffer, either all of them or a copy of the first part of them, so
`PartBytes` says which and `get_create_part` takes them itself, still only
once the queue has a slot.
@robobun

robobun commented Sep 21, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 2:37 PM PT - Sep 21st, 2026

✅ @robobun, your commit 0765a97af0635b5911037ac761dfae8dee830bb5 passed in Build #119371! 🎉


🧪   To try this PR locally:

bunx bun-pr 43709

That installs a local version of the PR into your bun-43709 executable, so you can run:

bun-43709 --bun

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed the latest revision and didn't find any bugs; since this rewrites who owns a multipart part's bytes in native S3 code, a human look at the ownership paths would still be worthwhile.

What was reviewed:

  • Every exit of enqueue_part: Whole takes buffered only after a queue slot is found (so Ok(false) leaves the buffer intact and Err leaves the part as the single owner); First(len) copies and advances the cursor only on Ok(true).
  • perform clones the Rc across execute_simple_s3_request, which copies the body into its own Box<[u8]> (simple_request.rs:677), so an in-place failure that frees the part cannot leave the request reading freed memory; sent is read before free_data on the ETag path.
  • The test's proxy scrub now covers ALL_PROXY/all_proxy, which resolves the earlier inline note. Only remaining nit: the host awaits the worker's "started" message without wiring the worker error/exit events, so a worker that throws before posting hangs to the runner timeout instead of failing clearly.
Extended reasoning...

The change replaces UploadPart's raw slice pointer plus manual Vec::from_raw_parts free with an Rc<Vec> and moves the take-from-buffer step inside get_create_part behind the queue-slot check, removing the ManuallyDrop juggling that left two owners when the request failed in place; it adds one spawned test that terminates workers mid-enqueue. It touches no auth, injection, or data-exposure surface; the sensitive surface is native memory ownership in src/runtime/webcore/s3/multipart.rs. I traced the Ok(false)/Ok(true)/Err paths for both PartBytes variants and the in-place-failure path through perform and found the ownership balanced, and the prior ALL_PROXY finding is addressed in the current test. Deferring rather than approving because this is a memory-safety fix in native runtime code that shipped with a double free before, and the unfixed-build failure claim for the test was not independently reproduced here.

… starts

The host waited only for each worker's message, so a worker that threw or
exited first left the run to the runner's timeout. Its `error` and `exit`
events now reject the wait, and the host prints the error and exits 1.
@robobun

robobun commented Sep 21, 2026

Copy link
Copy Markdown
Collaborator Author

The remaining test nit is addressed in 0765a97. The host now rejects its wait on a worker's error and exit events, prints the error and exits 1, so a worker that fails before it posts fails the run in about a second instead of at the runner's timeout. I checked both cases (a worker that throws, a worker that exits) with the debug build.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I reviewed the latest push and found no bugs; since this reworks who owns a part's bytes across a failure callback that can re-enter the upload, a human look at the ownership model would still be worthwhile.

What was reviewed:

  • get_create_part moves bytes out of buffered only after a queue slot is found, so Ok(false) leaves the buffer untouched and the removed re-append fallback had no remaining purpose; on Err the part is the single owner.
  • perform clones the Rc before execute_simple_s3_request; the request copies the body (Box::from(options.body) in simple_request.rs), so nothing references the part's bytes after the call returns and a synchronous cancel inside it cannot invalidate the slice.
  • MultiPartUpload::drop now frees data parts still hold; in-flight part requests hold a ctx ref, so that cannot race a request.
  • The test now clears every proxy variable the client consults and rejects on early worker exit, addressing the earlier inline thread.
Extended reasoning...

The change touches src/runtime/webcore/s3/multipart.rs, replacing a raw Cell<*const [u8]> plus manual Vec::from_raw_parts free with JsCell<Option<Rc<Vec<u8>>>> and a PartBytes enum, and adds test/js/bun/s3/s3-upload-terminate.test.ts. No security-sensitive surface (no auth, injection, or data-exposure paths); it is a memory-ownership fix in native runtime code. Deferring rather than approving because the correctness depends on the reentrant failure path inside execute_simple_s3_request freeing the part while the caller is mid-call, and the regression test is a probabilistic race reproduction that relies on ASAN to observe the old double free. No CODEOWNERS entry covers the changed files, and the bug hunt ran dry with no findings.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants