Skip to content

Fix is_executable_file on OHOS where open() ignores execute permission - #8

Merged
springmin merged 1 commit into
springmin:ohos-aarch64from
DongZhidong1993:ohos-aarch64-latest
Jun 9, 2026
Merged

springmin merged 1 commit into
springmin:ohos-aarch64from
DongZhidong1993:ohos-aarch64-latest

Conversation

@DongZhidong1993

Copy link
Copy Markdown

Fix is_executable_file on OHOS where open() ignores execute permission

Test: bun test test/cli/install/bun-run.test.ts -t "inside wrong script"
@autofix-troubleshooter

Copy link
Copy Markdown

Hi! I'm the autofix logoautofix.ci troubleshooter bot.

It looks like you correctly set up a CI job that uses the autofix.ci GitHub Action, but the autofix.ci GitHub App has not been installed for this repository. This means that autofix.ci unfortunately does not have the permissions to fix this pull request. If you are the repository owner, please install the app and then restart the CI workflow! 😃

@springmin
springmin merged commit f4c73ce into springmin:ohos-aarch64 Jun 9, 2026
2 of 5 checks passed
springmin pushed a commit that referenced this pull request Jun 27, 2026
…sweep (oven-sh#32729)

### Crash

```
ASSERTION FAILED: vm().currentThreadIsHoldingAPILock() => vm().heap.mutatorState() != MutatorState::Sweeping
vendor/WebKit/Source/JavaScriptCore/runtime/JSCell.cpp(179) : bool JSC::JSCell::validateIsNotSweeping() const
```

Backtrace (from a release-asan build with asserts):

```
#3  JSC::JSCell::validateIsNotSweeping()
#4  JSC::JSCell::classInfo() const
#5  WTF::uncheckedDowncast<WebCore::JSResumableFetchSink>(JSValue const&)
#6  ResumableFetchSinkPrototype__ondrainSetCachedValue
#7  bun_runtime::webcore::fetch::fetch_tasklet::FetchTasklet::ignore_remaining_response_body
#8  JSC::WeakBlock::sweep()          <- inside GC sweep (Weak finalizer)
#9  JSC::WeakSet::sweep()
#10 JSC::PreciseAllocation::sweep()
#12 JSC::Heap::finalize()
oven-sh#21 JSC::LocalAllocator::allocateSlowCase
oven-sh#23 JSC::ErrorInstance::create        <- ordinary allocation kicked off GC
```

Found by the syscall fault-injection fuzzer's client-side grammar
scenario (fetch/node:http with abort + transient errno on the client
socket). Reproduces ~4/5 under `BUN_JSC_collectContinuously=1`.

### Cause

`FetchTasklet::on_response_finalize` is the
`WeakRefOwner<FetchResponse>::finalize` callback and runs inside
`WeakBlock::sweep` while `MutatorState == Sweeping`. When the response
body is `Locked` without a pending promise or stream it calls
`ignore_remaining_response_body()`, which called:

- `ResumableSink::detach_js()`: writes the sink wrapper's cached
`ondrain` / `oncancel` / `stream` slots via the generated
`ResumableFetchSinkPrototype__*SetCachedValue` helpers. Each does
`uncheckedDowncast<JSResumableFetchSink>(thisValue)`, which reaches
`JSCell::classInfo()` and then issues a write barrier on the wrapper
cell.
- `clear_stream_handlers()`: reaches `ReadableStreamTag__tagged` ->
`object->inherits<JSReadableStream>()` (guarded today, but one boolean
away).

Calling `classInfo()` on any cell while the mutator is sweeping is
forbidden: the cell's `Structure` may already have been swept. Assert
builds catch it; release builds corrupt the heap.

### Fix

Thread a `from_finalizer` flag through `ignore_remaining_response_body`.
When `true` (the `on_response_finalize` caller) skip `detach_js()` and
`clear_stream_handlers()`; only native state is touched. The sink's
JS-side detach still happens from `clear_sink()` in
`FetchTasklet::deinit()`, which runs as an event-loop `ConcurrentTask`
outside any sweep, so nothing leaks.

The `on_stream_cancelled_callback` caller (reader `.cancel()`, runs from
JS on the event loop) passes `false` and keeps the immediate detach.

Also corrects the `ResumableSink::detach_js` doc comment that claimed
finalizer safety.

### Verification

New test at `test/js/web/fetch/fetch-response-finalizer-sweep.test.ts`:
a child process under `BUN_JSC_collectContinuously=1` does 12 iterations
of `fetch()` with a user-constructed `ReadableStream` body (so the sink
takes the JS route with a Strong `js_this`) against a raw TCP server
that sends headers + a partial chunked body and never terminates it,
then drops the `Response` unconsumed and runs `Bun.gc(true)`.

Without the fix (`bun bd`, src/ stashed):

```
exitCode: 134
stderr: ASSERTION FAILED: vm().currentThreadIsHoldingAPILock() => vm().heap.mutatorState() != MutatorState::Sweeping
```

With the fix: `stdout: "ok"`, `exitCode: 0`.

`test/js/web/fetch/fetch-backpressure.test.ts` (exercises the
`on_stream_cancelled_callback` path) passes unchanged.

---------

Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
springmin pushed a commit that referenced this pull request Jun 29, 2026
…ed (oven-sh#33016)

A backend message that fails the connection can share a TCP read with
messages that follow it. `PostgresRequest::on_data`'s message loop had
no bail-out once `fail()` had run, so the trailing messages in that read
kept being dispatched against the already-failed connection.

### Repro

A mock backend that answers the StartupMessage with one write carrying
two messages:

```
R  int32(8) int32(99)   Authentication, unrecognized type
Z  int32(5) 'I'         ReadyForQuery
```

```ts
const sql = new SQL({ url: `postgres://u@127.0.0.1:${port}/db`, max: 1, idleTimeout: 1, connectionTimeout: 5 });
await sql`select 1`.catch(() => {});
await Bun.sleep(1600);
```

### Cause

The unrecognized `Authentication` type calls `fail()`, which sets the
status to `Failed`, closes the socket, and rejects the pending requests,
but the message loop keeps going and dispatches the `ReadyForQuery` from
the same read. That calls `set_status(Status::Connected)`, which has no
guard against leaving `Failed`, so the dead connection is flipped back
to `Connected` and the `on_data` epilogue re-arms its idle timer.
uSockets frees a closed `us_socket_t` at the end of the event-loop
iteration, so when the timer later fires, `ref_and_close` reads the
freed socket:

```
ERROR: AddressSanitizer: heap-use-after-free
READ of size 1 at 0x71f2125605d2 thread T0
    #0 us_socket_is_closed                              packages/bun-usockets/src/socket.c:143:21
    #4 PostgresSQLConnection::ref_and_close             src/sql_jsc/postgres/PostgresSQLConnection.rs:1528:31
    #5 PostgresSQLConnection::fail_with_js_value        src/sql_jsc/postgres/PostgresSQLConnection.rs:726:14
    #6 PostgresSQLConnection::fail_fmt                  src/sql_jsc/postgres/PostgresSQLConnection.rs:749:14
    #7 PostgresSQLConnection::on_connection_timeout     src/sql_jsc/postgres/PostgresSQLConnection.rs:557:14
    #8 __bun_fire_timer                                 src/runtime/dispatch.rs:1020:35
0x71f2125605d2 is located 18 bytes inside of 104-byte region
freed by thread T0 here:
    #2 us_internal_free_closed_sockets                  packages/bun-usockets/src/loop.c:305:9
```

### Fix

- `PostgresRequest::on_data`: the message loop returns once the
connection's status is `Failed`. `fail()` is terminal; nothing after it
in the same read should be handled (a `DataRow`, `CommandComplete`, or
`ErrorResponse` in that position would be just as wrong as the
`ReadyForQuery`).
- `PostgresSQLConnection::set_status`: refuses to transition out of
`Failed`. The transition function owns that invariant; every other
consumer of `Status` (the timer interval, `update_has_pending_activity`,
the idempotency check in `fail_with_js_value`) already assumes `Failed`
is terminal.

### Verification

`test/js/sql/postgres-failed-connection-resurrection.test.ts` runs a
fixture against the mock backend above and lets it outlive the
idle-timer window. Without the fix the fixture dies with the ASan report
above; with it the fixture exits 0. Gated to ASan builds because the bug
is a read of freed memory, which release lanes do not detect.

The postgres fault-injection and integration suites still pass locally
(90 tests across `test/js/sql/postgres-*.test.ts`, `sql*.test.ts`,
`tls-sql.test.ts`).

### Related

- oven-sh#32861 detaches the stored socket handle in `on_close` /
`on_connect_error` so nothing can dereference the freed `us_socket_t`
regardless of how the stale read is reached. It removes the last step of
this chain from the other end; this PR stops the failed connection from
being resurrected at all.
- oven-sh#30950 guards the JS pool's `handleConnected` against the reverse
ordering within one read (a legitimately queued `onconnect` microtask
arriving after a synchronous `onclose`).
springmin pushed a commit that referenced this pull request Jul 2, 2026
…h#33186)

### Repro

```sh
printf '{"name":"x","version":"1.0.0"}' > package.json
bun pm pkg set 'contributors[0]=alice'
```

On a release build (1.4.0 and current `main`) this exits 0 and writes
freed heap bytes into `package.json` as the property key:

```json
{
"name": "x",
"version": "1.0.0",
"P\x01\x00\x00\x00tors": {
  "\x00": "alice"
}
}
```

Depending on what was in the freed allocation the result is often not
valid JSON at all. Any `bun pm pkg set` key path containing `[index]`
hits it.

Under ASAN it is a deterministic `heap-use-after-free`:

```
ERROR: AddressSanitizer: heap-use-after-free
READ of size 1
    #0 bun_js_printer::write_pre_quoted_string_inner   src/js_printer/lib.rs:1014
    #7 PmPkgCommand::save_package_json                 src/runtime/cli/pm_pkg_command.rs:909
freed by thread T0 here:
    #7  <Box<[u8]> as Drop>::drop
    #12 PmPkgCommand::set_value                        src/runtime/cli/pm_pkg_command.rs:661
previously allocated by thread T0 here:
    #10 <Box<[u8]> as From<&[u8]>>::from
    #11 PmPkgCommand::parse_key_path                   src/runtime/cli/pm_pkg_command.rs:583
```

<details>
<summary>full ASAN report</summary>

```
=================================================================
==16563==ERROR: AddressSanitizer: heap-use-after-free on address 0x73423c7c0670 at pc 0x00000f583cc5 bp 0x7fff2667e950 sp 0x7fff2667e948
READ of size 1 at 0x73423c7c0670 thread T0
    #0 0x00000f583cc4 in _RINvCs59Hqei94dXF_14bun_js_printer29write_pre_quoted_string_innerINtB2_16StdWriterAdapterQINtB2_6WriterNtB2_12BufferWriterEEKVNtNtB2_8Encoding4Utf8UECsgBGN0jRPILJ_11bun_bundler /workspace/bun/src/js_printer/lib.rs:1014:79
    #1 0x00000ebda439 in <bun_js_printer::__gated_printer::Printer<&mut bun_js_printer::Writer<bun_js_printer::BufferWriter>, false, false, false, true, false>>::print_string_characters_utf8 /workspace/bun/src/js_printer/lib.rs:2641:21
    #2 0x00000ebdb7a7 in <bun_js_printer::__gated_printer::Printer<&mut bun_js_printer::Writer<bun_js_printer::BufferWriter>, false, false, false, true, false>>::print_string_characters_e_string /workspace/bun/src/js_printer/lib.rs:4546:22
    #3 0x00000ebdb238 in <bun_js_printer::__gated_printer::Printer<&mut bun_js_printer::Writer<bun_js_printer::BufferWriter>, false, false, false, true, false>>::print_string_literal_e_string /workspace/bun/src/js_printer/lib.rs:3018:18
    #4 0x00000ebd2be2 in <bun_js_printer::__gated_printer::Printer<&mut bun_js_printer::Writer<bun_js_printer::BufferWriter>, false, false, false, true, false>>::print_property /workspace/bun/src/js_printer/lib.rs:4807:34
    #5 0x00000ebc0166 in <bun_js_printer::__gated_printer::Printer<&mut bun_js_printer::Writer<bun_js_printer::BufferWriter>, false, false, false, true, false>>::print_expr /workspace/bun/src/js_printer/lib.rs:3962:38
    #6 0x00000ee574c1 in bun_js_printer::print_json::<&mut bun_js_printer::Writer<bun_js_printer::BufferWriter>> /workspace/bun/src/js_printer/lib.rs:8071:13
    #7 0x00000c05c270 in <bun_runtime::cli::pm_pkg_command::PmPkgCommand>::save_package_json /workspace/bun/src/runtime/cli/pm_pkg_command.rs:909:25
    #8 0x00000c060cfb in <bun_runtime::cli::pm_pkg_command::PmPkgCommand>::exec_set /workspace/bun/src/runtime/cli/pm_pkg_command.rs:330:13
    #9 0x00000c05d333 in <bun_runtime::cli::pm_pkg_command::PmPkgCommand>::exec /workspace/bun/src/runtime/cli/pm_pkg_command.rs:73:32
    #10 0x00000bf919d0 in <bun_runtime::cli::package_manager_command::PackageManagerCommand>::exec /workspace/bun/src/runtime/cli/package_manager_command.rs:704:13
    #11 0x00000c3fbb87 in bun_runtime::cli::command::exec_pm /workspace/bun/src/runtime/cli/mod.rs:1591:34
    #12 0x00000c3f2b86 in bun_runtime::cli::command::start /workspace/bun/src/runtime/cli/mod.rs:1309:43
    #13 0x00000bfad16c in bun_runtime::cli::cli::start /workspace/bun/src/runtime/cli/mod.rs:573:27
    #14 0x00000bb3c034 in main /workspace/bun/src/bun_bin/lib.rs:230:5
    #15 0x77223ccc7ca7 in __libc_start_call_main csu/../sysdeps/nptl/libc_start_call_main.h:58:16
    oven-sh#16 0x77223ccc7d64 in __libc_start_main csu/../csu/libc-start.c:360:3
    oven-sh#17 0x0000099d1d1d in __wrap___libc_start_main /workspace/bun/build/debug/../../src/jsc/bindings/workaround-missing-symbols.cpp:487:12

0x73423c7c0670 is located 0 bytes inside of 12-byte region [0x73423c7c0670,0x73423c7c067c)
freed by thread T0 here:
    #0 0x000007ae192a in free crtstuff.c
    #1 0x00000bb3c5a7 in <std::alloc::System as core::alloc::global::GlobalAlloc>::dealloc /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/std/src/sys/alloc/unix.rs:48:18
    #2 0x00000bb3be9a in __rustc::__rust_dealloc /workspace/bun/src/bun_bin/lib.rs:56:15
    #3 0x00001258b05f in alloc::alloc::dealloc_nonnull /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/alloc.rs:128:14
    #4 0x0000125872fe in <alloc::alloc::Global>::deallocate_impl_runtime /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/alloc.rs:229:22
    #5 0x000012586364 in <alloc::alloc::Global>::deallocate_impl /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/alloc.rs:344:9
    #6 0x00001258d79c in <alloc::alloc::Global as core::alloc::Allocator>::deallocate /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/alloc.rs:462:23
    #7 0x000012582946 in <alloc::boxed::Box<[u8]> as core::ops::drop::Drop>::drop /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/boxed.rs:1956:24
    #8 0x000012572e44 in core::ptr::drop_in_place::<alloc::boxed::Box<[u8]>> /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/core/src/ptr/mod.rs:809:1
    #9 0x000011f8d429 in core::ptr::drop_in_place::<[alloc::boxed::Box<[u8]>]> /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/core/src/ptr/mod.rs:809:1
    #10 0x00000ef6b73a in <alloc::vec::Vec<alloc::boxed::Box<[u8]>> as core::ops::drop::Drop>::drop /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/vec/mod.rs:4258:13
    #11 0x00000ef69e64 in core::ptr::drop_in_place::<alloc::vec::Vec<alloc::boxed::Box<[u8]>>> /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/core/src/ptr/mod.rs:809:1
    #12 0x00000c061846 in <bun_runtime::cli::pm_pkg_command::PmPkgCommand>::set_value /workspace/bun/src/runtime/cli/pm_pkg_command.rs:661:5
    #13 0x00000c061038 in <bun_runtime::cli::pm_pkg_command::PmPkgCommand>::exec_set /workspace/bun/src/runtime/cli/pm_pkg_command.rs:325:13
    #14 0x00000c05d333 in <bun_runtime::cli::pm_pkg_command::PmPkgCommand>::exec /workspace/bun/src/runtime/cli/pm_pkg_command.rs:73:32
    #15 0x00000bf919d0 in <bun_runtime::cli::package_manager_command::PackageManagerCommand>::exec /workspace/bun/src/runtime/cli/package_manager_command.rs:704:13
    oven-sh#16 0x00000c3fbb87 in bun_runtime::cli::command::exec_pm /workspace/bun/src/runtime/cli/mod.rs:1591:34
    oven-sh#17 0x00000c3f2b86 in bun_runtime::cli::command::start /workspace/bun/src/runtime/cli/mod.rs:1309:43
    oven-sh#18 0x00000bfad16c in bun_runtime::cli::cli::start /workspace/bun/src/runtime/cli/mod.rs:573:27
    oven-sh#19 0x00000bb3c034 in main /workspace/bun/src/bun_bin/lib.rs:230:5
    oven-sh#20 0x77223ccc7ca7 in __libc_start_call_main csu/../sysdeps/nptl/libc_start_call_main.h:58:16

previously allocated by thread T0 here:
    #0 0x000007ae1bc8 in malloc crtstuff.c
    #1 0x00000bb3c520 in <std::alloc::System as core::alloc::global::GlobalAlloc>::alloc /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/std/src/sys/alloc/unix.rs:14:22
    #2 0x00000bb3be30 in __rustc::__rust_alloc /workspace/bun/src/bun_bin/lib.rs:56:15
    #3 0x00001258b335 in alloc::alloc::alloc /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/alloc.rs:101:9
    #4 0x000012586b81 in <alloc::alloc::Global>::alloc_impl_runtime /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/alloc.rs:210:73
    #5 0x0000125862b6 in <alloc::alloc::Global>::alloc_impl /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/alloc.rs:332:9
    #6 0x00001258d86a in <alloc::alloc::Global as core::alloc::Allocator>::allocate /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/alloc.rs:449:14
    #7 0x00001257dbd3 in <alloc::boxed::Box<[u8]>>::try_clone_from_ref_in /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/boxed.rs:881:29
    #8 0x00001257da49 in <alloc::boxed::Box<[u8]>>::clone_from_ref_in /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/boxed.rs:840:15
    #9 0x00001257d3f4 in <alloc::boxed::Box<[u8]>>::clone_from_ref /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/boxed.rs:793:9
    #10 0x000012581e34 in <alloc::boxed::Box<[u8]> as core::convert::From<&[u8]>>::from /root/.rustup/toolchains/nightly-2026-05-06-x86_64-unknown-linux-gnu/lib/rustlib/src/rust/library/alloc/src/boxed/convert.rs:77:9
    #11 0x00000c05a47f in <bun_runtime::cli::pm_pkg_command::PmPkgCommand>::parse_key_path /workspace/bun/src/runtime/cli/pm_pkg_command.rs:583:37
    #12 0x00000c061608 in <bun_runtime::cli::pm_pkg_command::PmPkgCommand>::set_value /workspace/bun/src/runtime/cli/pm_pkg_command.rs:643:30
    #13 0x00000c061038 in <bun_runtime::cli::pm_pkg_command::PmPkgCommand>::exec_set /workspace/bun/src/runtime/cli/pm_pkg_command.rs:325:13
    #14 0x00000c05d333 in <bun_runtime::cli::pm_pkg_command::PmPkgCommand>::exec /workspace/bun/src/runtime/cli/pm_pkg_command.rs:73:32
```

</details>

### Cause

`parse_key_path` returned a `Vec<Box<[u8]>>`, and `set_value` /
`set_nested` inserted those boxed segments into the manifest AST by
reference: `E::Object::put` constructs `EString::init(key)`, whose
documented contract is that `key` is arena-owned (it records the slice,
it does not copy it). The vector is a local of `set_value`, so it
dropped before `exec_set` reached `save_package_json`, and the JSON
printer then read the dangling keys.

The non-bracket path in `set_value` did not have the bug: it borrowed
its segments straight out of the argv key, which outlives the whole
command. The bracket path differed only by the unnecessary boxing.

### Fix

`parse_key_path` now returns `Vec<&[u8]>`. Every segment is a literal
sub-slice of the input key, so nothing ever needed owning. With the
boxing gone, `set_value`'s separate non-bracket branch and its
`set_nested_simple` helper (which existed only to avoid the allocation)
were exact duplicates of the bracket path, so they are deleted and all
keys route through `parse_key_path` + `set_nested`.

`set_nested_simple`'s trailing `root.put(current_key, nested)` was a
no-op: `ExprData::EObject` is a `StoreRef` handle, so mutating the copy
returned by `root.get()` already mutates the stored object, and the put
re-stores the same handle. Dropping it with the function changes nothing
(and the prior bracket path, `set_nested`, never had it).

Intentionally not changed here: `set 'contributors[0]=alice'` produces
`"contributors": {"0": "alice"}`, an object keyed by the digit string,
rather than the array npm's `pkg set` creates, and `set 'array[]=x'`
still errors with `InvalidPath` instead of appending. Both are the npm
compat gap tracked in oven-sh#22035, which is separate from the memory safety
of the key names and is not closed by this PR.

### Verification

New test in `test/cli/install/bun-pm-pkg.test.ts` reparses the written
file and asserts the exact object. Without the fix it fails on release
(`SyntaxError: JSON Parse error: Invalid escape character x`) and on the
ASAN debug build (the child aborts on the use-after-free). With the fix
the full `bun-pm-pkg.test.ts` suite passes (74 pass, 0 fail).
springmin pushed a commit that referenced this pull request Jul 15, 2026
…ry rewrite (oven-sh#34271)

`test/js/bun/util/filesystem_router.test.ts` went red on alpine x64 in
build [73276](https://buildkite.com/bun/bun/builds/73276): the `reload()
while Bun.build() resolves the same directory` subprocess segfaulted in
`bust_dir_cache_recursive`, inlined from `NonNull::new`.

## Cause

`RealFS::entries_at` (`src/resolver/lib.rs`) replaces a cached
`DirEntry` in place when the caller's resolver generation is newer than
the cached listing's. The replacement at `*e_ptr = new_entry` drops the
old `DirEntry`, which drops its `data: StringHashMap<*mut Entry>` and
frees the hashmap's bucket allocation. The function's comment says
`entries_mutex held by caller`, but that is only true on one of the five
paths that reach it: `dir_info_uncached`, when entered from
`dir_info_cached_miss`. The other callers (`finalize_result`,
`handle_esm_resolution`, `load_index_with_extension`,
`Transpiler::run_env_loader`) all reach `entries_at` after
`dir_info_cached_maybe_log` has already returned and released both
`RESOLVER_MUTEX` and `entries_mutex`.

`FileSystemRouter::reload()` and `RouteLoader::load` iterate the same
`DirEntry.data` map under `entries_mutex` (the snapshot pattern oven-sh#33056
introduced for exactly this kind of concurrent rewrite). With
`entries_at`'s rewrite unsynchronized, a `Bun.build()` on the bundler
thread can drop the map while `reload()` on the JS thread is
mid-iteration.

The generation mismatch is what makes `entries_at` enter its rewrite
branch, so the window only opens once the bundle thread has processed at
least one batch (it bumps its own generation after every queue drain);
every subsequent `Bun.build()` then re-reads any directory that
`reload()` just refreshed to generation 0.

ASAN catches it as a heap-use-after-free with the two sides of the race
laid out exactly:

```
READ of size 16 (thread T0):
  #6 HashMap::values
  #7 StringHashMap<*mut Entry>::values                       src/collections/array_hash_map.rs:1864
  #8 FileSystemRouter::bust_dir_cache_recursive              src/runtime/api/filesystem_router.rs:395
  #9 FileSystemRouter::bust_dir_cache                        src/runtime/api/filesystem_router.rs:451
  #10 FileSystemRouter::reload                               src/runtime/api/filesystem_router.rs:476

freed by thread T11 (Bundler):
  #11 drop_in_place<bun_resolver::fs_full::DirEntry>
  #12 bun_resolver::fs::RealFS::entries_at                   src/resolver/lib.rs:1639
  #13 DirInfo::get_entries_ref                               src/resolver/dir_info.rs:266
  #14 Resolver::finalize_result                              src/resolver/resolver.rs:1714
  #15 Resolver::resolve_and_auto_install                     src/resolver/resolver.rs:1485
  ...
  oven-sh#23 BundleThread::generate_in_new_thread                   src/bundler/BundleThread.rs:276

previously allocated by thread T0:
  oven-sh#17 HashMap::reserve
  oven-sh#18 Resolver::dir_info_cached_miss                         src/resolver/resolver.rs:4591
  oven-sh#19 Resolver::dir_info_cached_maybe_log                    src/resolver/resolver.rs:4201
  oven-sh#20 Resolver::read_dir_info                                src/resolver/resolver.rs:4118
  oven-sh#21 FileSystemRouter::reload                               src/runtime/api/filesystem_router.rs:492
```

(The use side is sometimes `RouteLoader::load` at
`src/router/lib.rs:816` instead; same map, same lock.)

This has been the shape of `entries_at` since the Rust port; oven-sh#33056
narrowed the race by snapshotting under the lock but assumed the rewrite
side already held it.

## Fix

`entries_at` now takes `entries_mutex` itself, matching
`read_directory_with_iterator` which already does. The one call path
that reaches it with the lock already held (`dir_info_cached_miss` ->
`dir_info_uncached` -> `parent_.get_entries_ref`) routes through a new
`entries_at_locked` / `get_entries_ref_locked` pair so the non-recursive
mutex is not re-entered. That path is the only one that passes a
non-`None` parent to `dir_info_uncached`; the other caller
(`dir_info_for_resolution`) passes `None`, so the parent branch
containing the accessor never runs there.

## Test

The existing concurrency test now awaits one `Bun.build()` first, so the
bundle thread's generation is already past zero when the concurrent
rounds start, and then runs forty reload/build rounds instead of one.
That is the shape that reaches the stale-generation rewrite at all; the
original single-round fixture usually completes with every build still
on generation 0.

The race is scheduling-dependent. Pinning the fixture to a single core
reproduces the ASAN use-after-free on roughly 3 in 10 runs against an
unpatched debug build and 0 in 15 with this change; with all 16 cores
available the unpatched build reproduces at roughly 1 in 30. The
assertions are otherwise the same as before, so the test continues to
cover the behavior oven-sh#33056 added.

Also ran the full `filesystem_router.test.ts`,
`test/bundler/bun-build-api.test.ts` (including the thousands-of-builds
test that exercises the generation path heavily),
`test/js/bun/resolve/resolve.test.ts`, `test/cli/hot/hot.test.ts`,
`test/cli/watch/watch.test.ts`, `test/bake/framework-router.test.ts`,
and `bun run rust:check-all` (10/10 targets).

<!-- robobun:evidence:begin -->

---

**no test proof** · iteration 0 · Platform-specific test(s) that do not
run on this machine. Deferring to CI, which covers all platforms:
test/js/bun/util/filesystem_router.test.ts

<!-- robobun:evidence:end -->
springmin pushed a commit that referenced this pull request Jul 20, 2026
…ven-sh#34693)

## Use-after-free in `H2FrameParser::on_native_writable`

Fleet ASAN fuzz hit (p-h2c cleartext harness, seed 1,
`server-conn.recv.*:A8`):

```
use-after-poison READ 8 (shadow f7 = user poison, HiveArray slot re-poison)
  #0 Vec::len (write_buffer)
  #2 has_backpressure          h2_frame_parser.rs:3276
  #3 on_native_writable        h2_frame_parser.rs:9605
  #4 NewSocket<true>::on_writable  socket_body.rs:894
  #8 us_internal_ssl_on_writable   bun-usockets openssl.c:1851
allocated by: HiveArray Fallback<H2FrameParser,256>, H2FrameParser::constructor
```

### Cause

`on_native_writable` loops `flush()` and checks `has_backpressure()`
between iterations. `flush()` re-enters JS via `flush_stream_queue` ->
`dispatch_write_callback` / `onStreamEnd` / `onWantTrailers`. A callback
that destroys the session reaches `detach_native_callback`, dropping the
socket's `+1` on the parser. If that was the last external ref,
`flush()`'s own keepalive is all that remains and drops on return, so
the next `has_backpressure()` reads a HiveArray slot that was just
`drop_in_place`'d and re-poisoned by `POOL.put`.

`on_native_read` already takes a `keepalive()` for exactly this reason
(h2_frame_parser.rs:9590); `on_native_writable` did not.

In release builds there is no poison: the same ordering is a silent
use-after-free in every `node:http2` server/client on a native socket.
The read of a stale `write_buffer.len()` can satisfy the loop condition
and send the next `flush()` into UAF writes on the freed parser.

### Fix

- Take a `keepalive()` for the extent of `on_native_writable`, mirroring
`on_native_read`.
- `NativeCallbacks::on_data`/`on_writable`: copy the raw `*mut
H2FrameParser` out of the enum before dispatching, so the
`JsCell<NativeCallbacks>` borrow does not span a re-entrant
`detach_native_callback` that overwrites the cell.

### Test

`test/js/node/http2/node-http2-writable-destroy-fixture.ts` reproduces
the exact fleet stack under ASAN by faulting `send`/`writev` to 0
(backpressure, arms WRITABLE), queuing a DATA frame whose write callback
runs `session.destroy()` + `Bun.gc(true)`, then clearing the fault so
the writable event drains the queue inside `on_native_writable`. Added
to `node-http2-syscall-fault.test.ts` as an ASAN-gated subprocess test.

<details><summary>Fail-before ASAN report (matches the fleet
hit)</summary>

```
==ERROR: AddressSanitizer: use-after-poison on address 0x... 
READ of size 8 at 0x... thread T0
    #2 H2FrameParser::has_backpressure h2_frame_parser.rs:3276:33
    #3 H2FrameParser::on_native_writable h2_frame_parser.rs:9605:21
    #4 NativeCallbacks::on_writable socket_body.rs:3960:20
    #5 NewSocket<false>::on_writable socket_body.rs:894:39
allocated by thread T0 here:
    ... Fallback<H2FrameParser, 256>::new_boxed hive_array.rs:668
    ... H2FrameParser::constructor h2_frame_parser.rs:9800
SUMMARY: AddressSanitizer: use-after-poison ... Vec<u8>::len
```
</details>

<!-- robobun:evidence:begin -->

---

**no test proof** · iteration 2 · Platform-specific test(s) that do not
run on this machine. Deferring to CI, which covers all platforms:
test/js/node/http2/node-http2-syscall-fault.test.ts

<!-- robobun:evidence:end -->

---------

Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
springmin pushed a commit that referenced this pull request Jul 20, 2026
Makes Bun usable inside a Windows AppContainer (lowbox token), the
sandbox used by packaged apps and embedders that launch worker processes
with `CreateAppContainerProfile` +
`PROC_THREAD_ATTRIBUTE_SECURITY_CAPABILITIES`. Depends on the libuv-side
fixes in oven-sh/libuv#7 (cherry-pick of upstream libuv/libuv#5181's
AppContainer pipe-namespace fix) and oven-sh/libuv#8 (fs stat/realpath
bounds and tty exact-fill correctness fixes), both merged into the `bun`
branch and pinned here.

The four unconditional changes below apply everywhere; the two
AppContainer-only changes are gated on
`bun_sys::windows::is_app_container()` (a cached
`GetTokenInformation(TokenIsAppContainer)` probe) and are no-ops outside
a container.

### Changes

**Resolver ancestor-directory tolerance** (cross-platform,
unconditional). `bun run <script>` failed with `error loading current
directory` when any ancestor directory on the path to cwd was
unreadable: the resolver builds a `DirInfo` for every ancestor starting
at the drive root, and a sandboxed token (or an execute-only `0o111`
unix directory, Android `/data`, etc.) denies that listing. A
permission-denied *ancestor* is now treated as an opaque empty
directory, the same treatment the existing `ENOTDIR` tolerance applies;
errors on the requested directory itself stay fatal. Also fixes oven-sh#28220
and oven-sh#30859.

**Windows `O_RDONLY` open no longer requests `FILE_WRITE_ATTRIBUTES`**
(Windows, unconditional). The `openat` base access mask unconditionally
included `FILE_WRITE_ATTRIBUTES`, so opening a file `O_RDONLY` on a tree
with an RX-only ACL grant (Program Files, read-only shares, the normal
sandbox project-tree shape) failed `EPERM`. The mask now matches libuv's
`fs__open` (`O_RDONLY` -> `GENERIC_READ` only; write modes already
include it via `GENERIC_WRITE`); `fs.futimes` continues to work via
libuv's `ReOpenFile(FILE_WRITE_ATTRIBUTES)` at futimes time. Unskips
`test-module-readonly.js`.

**Windows directory opens no longer request `FILE_ADD_FILE |
FILE_ADD_SUBDIRECTORY`** (Windows, unconditional). `NtCreateFile` with a
`RootDirectory` handle checks the target directory's ACL for child
creates and renames, not the handle's access mask, so these bits grant
nothing and only narrow where the open is admitted. Dropping them lets
`Bun.Glob`/recursive readdir/`fs.opendir` descend RX-only directories
(Program Files, read-only shares, sandboxed project trees); creating
children through the handle still works where the ACL allows it. Removes
the now-vestigial `WindowsOpenDirOptions.read_only` field.

**Named-pipe listen failures surface as Node-shaped errors** (Windows,
unconditional). They were a codeless `ERR_INVALID_ARG_TYPE` TypeError;
now an `Error` with `code`/`errno`/`syscall`/`path` set, matching the
POSIX unix-socket listen path. Fixes oven-sh#30265.

**AppContainer-only** (gated on `is_app_container()`; no-ops outside):

- `GetFinalPathNameByHandleW(VOLUME_NAME_DOS)` is denied on every handle
inside an AppContainer because the DOS-name translation opens the mount
manager. For handles on the system volume, reconstruct the DOS name as
`<system-drive>:` + the `VOLUME_NAME_NT` tail (the system directory
carries an `ALL APPLICATION PACKAGES:(RX)` ACE by Windows default, so
its device name is resolvable from any lowbox); handles on any other
volume surface the original denial. Applies to both the typed `bun_sys`
wrapper and a raw-ABI drop-in. This is what keeps Bun's resolver and
`bun install` working inside a container; user-facing `fs.realpath` goes
through libuv and is left at Node parity (fails `EPERM`).
- `Bun.Terminal` ConPTY internal pipe names: insert `LOCAL\` into the
`\\.\pipe\...` name inside a container (the only namespace an
AppContainer may create server pipes under), matching libuv's
conditional insert.

**`deps:`** pins oven-sh/libuv `f6e75a7e` (= `bun` branch after #7 and
#8). Behavioural delta at this pin: libuv's internal pipe names gain
`LOCAL\` inside a container (upstream oven-sh#5181); `uv_fs_stat` of files the
OS holds exclusively at a drive root (the `C:\pagefile.sys` class)
reports the real stats instead of `ENOENT`; `uv_fs_realpath` preserves
the real error instead of masking as `EBADF`; console line reads don't
tear characters on an exact-fill allocation and report `UV_ENOBUFS` for
allocations too small to convert into.

### Known limitations

- `"ignore"` stdio opens the `NUL` device, whose default ACL denies
AppContainer tokens; grant the device ACL to the container SIDs from an
elevated context per boot, or use `"inherit"` stdin.
- `fs.realpath` (all variants) fails `EPERM` inside a container, as it
does under Node.js; Bun's own module resolution does not go through it.
- The isolated linker is unsupported in sandboxes (its junctions are
quarantined by the kernel); use the default hoisted linker.
- A package cache primed *outside* the container is currently
re-validated as a miss inside it; prefer letting the sandboxed process
populate its own cache.

### Tests

`test/js/bun/windows/appcontainer.test.ts` launches bun inside a real
AppContainer in the regular Windows CI lanes (bun:ffi lowbox launcher,
no admin needed) and asserts the sandbox-only behaviours (piped-stdio
spawn, `LOCAL\` pipe namespace, `fs.realpath` denial, fork + IPC); hosts
that cannot run sandboxed children skip visibly.
`resolver-permission-denied-ancestor.test.ts` covers the ancestor
tolerance on unix with an execute-only directory. The glob
`scan.test.ts` RX-only case and `named-pipe-listen-error.test.ts`
error-shape assertions cover the unconditional Windows changes.

---------

Co-authored-by: robobun <117481402+robobun@users.noreply.github.com>
Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
springmin pushed a commit that referenced this pull request Jul 21, 2026
…-sh#34834)

## What

On an idle Windows box the `GetQueuedCompletionStatusEx` timeout rounds
to the ~15.6 ms system clock tick, so `setTimeout(cb, 1)` and
`Bun.sleep(1)` fire ~15 ms late unless another process happens to have
raised the tick rate (the "works when Spotify is open" heisenbug).
Node.js has the same behavior.

## How

Bumps libuv to
[oven-sh/libuv@9687330](oven-sh/libuv@9687330),
which landed [oven-sh/libuv#9](oven-sh/libuv#9)
(same approach Go's runtime uses,
[golang/go#44343](golang/go#44343)):

* At `uv__winapi_init`, dynamically load `NtCreateWaitCompletionPacket`
/ `NtAssociateWaitCompletionPacket` / `NtCancelWaitCompletionPacket`
from ntdll.
* Per loop, create a `CREATE_WAITABLE_TIMER_HIGH_RESOLUTION` waitable
timer + wait-completion-packet and stash them on
`uv__loop_internal_fields_s` (Win10 1803+ / Server 2019+; on older
Windows the handles stay NULL and `uv__poll` keeps its original GQCS ms
wait, so behavior is unchanged).
* In `uv__poll`, when `timeout > 0`: arm the waitable timer for the
deadline, associate it with the loop's IOCP, and wait in GQCS with
`INFINITE`. When the timer fires, the kernel posts a completion with
`lpOverlapped == NULL`, which the existing dequeue loop already treats
as a pure wakeup.

The bump also pulls in the intervening Windows fs correctness fixes on
the `bun` branch
([oven-sh/libuv#7](oven-sh/libuv#7),
[#8](oven-sh/libuv#8)).

## Numbers (Windows Server 2019, idle)

|                    | before    | after    |
|--------------------|-----------|----------|
| `setTimeout(cb,1)` | 15.52 ms  | 1.41 ms  |
| `Bun.sleep(1)`     | 15.62 ms  | 1.08 ms  |
| `setTimeout(cb,5)` | 15.62 ms  | 5.24 ms  |
| `setInterval(16)`  | ~28 ms    | 16.4 ms  |

`Bun.serve` hello-world throughput on Windows debug (oha, 10s, 50
concurrent): 11,484 req/s on this branch vs 11,434 req/s on main
(noise).

The added test in `setTimeout.test.js` measures the median of 50
`setTimeout(1)` samples in a subprocess and asserts it's under 8 ms
(before: 15.6 ms; after: ~1.5 ms).

Fixes oven-sh#16714
Fixes oven-sh#26965

<!-- robobun:evidence:begin -->

---

**no test proof** · iteration 1 · Platform-specific test-only change;
deferring to CI.

<!-- robobun:evidence:end -->
springmin pushed a commit that referenced this pull request Jul 29, 2026
…rier (oven-sh#36337)

`JSNativeStreamSourceAdapter::m_controller` was a
`JSC::Weak<JSReadableStreamDefaultController>`. When the native pull
promise is rejected (socket fault on a fetch body) the adapter is queued
as the `onNativePullRejected` reaction context, which roots the
**adapter** but not the **controller**: the adapter's only edge to it
was the `Weak`. `FetchTasklet` releases both native `Strong<>`s to the
body stream before that microtask drains, so a GC in between can leave
the entire consumer graph (`controller -> stream -> reader -> pipe op ->
destination -> writer -> readyPromise`) white. The subsequent error
cascade then enqueues the pipe's writes-drained shutdown deferral
against a corpse `op`, and `performPipeShutdownAction(AbortDestination)`
dereferences a swept `readyPromise`:

```
ASSERTION FAILED: result   JSObject.h(583) JSGlobalObject *JSC::JSObject::realm() const
#5  JSC::JSObject::realm()
#6  JSC::JSPromise::rejectPromise
#7  JSC::JSPromise::reject
#8  Bun::WebStreams::writableStreamDefaultWriterEnsureReadyPromiseRejected
#9  Bun::WebStreams::writableStreamStartErroring
#10 Bun::WebStreams::writableStreamAbort
#11 WebCore::performPipeShutdownAction (AbortDestination)
#12 WebCore::JSStreamPipeToOperation::onWritesFinishedForShutdown
```

On builds without the assert the same path is a silent write into
freed/reused promise memory.

## Fix

Hold `m_controller` as a visited internal field so a queued adapter
roots the controller directly. The edge is cleared on every terminal
path (`nativeSourcePullRejected`, `nativeSourceCallClose`,
`nativeSourceCancel`); `controller->algorithmContext` is cleared by
`readableStreamDefaultControllerClearAlgorithms`, so the abandoned case
is an ordinary intra-heap cycle mark-sweep collects.
`NewSource::this_jsvalue` is only `Strong` during FileReader I/O, where
pinning the consumer graph is the correct behavior anyway.

With the `Weak` gone the adapter no longer needs a destructor, so it is
now a `JSInternalFieldObjectImpl<5>`: the five JSValue members (handle,
pendingView, closer, drainValue, controller) are internal fields visited
by the base class, with typed accessors at call sites. The scalar
members (chunkSize, flag bitfield, text-decode state) stay as plain
members.

## Verification

`native-source-onclose-leak.test.ts` (the partial-read + `releaseLock`
abandonment tests for Blob/fetch/File sources) continues to pass,
confirming the cycle does not pin. `streams.test.js`,
`pipeTo-signal-leak.test.ts`, `compression.test.ts`, `blob.test.ts` all
pass.

The crash itself is 0/1800 standalone; it reproduces ~1/3 only under a
fault-injected tracer replay. `pipeTo-shutdown-gc.test.ts` exercises the
shape (native body source, socket fault mid-stream, fire-and-forget
`pipeTo` under `collectContinuously`, `AbortDestination` shutdown arm)
as a regression surface.

<!-- robobun:evidence:begin -->

---

**no test proof** · iteration 2 · Platform-specific test(s) that do not
run on this machine. Deferring to CI, which covers all platforms:
test/js/web/streams/pipeTo-shutdown-gc.test.ts

<!-- robobun:evidence:end -->
springmin pushed a commit that referenced this pull request Aug 6, 2026
…llback (oven-sh#36986)

## What

`test/js/node/async_hooks/AsyncLocalStorage-tracking.test.ts` (the
crypto-generateKeyPair fixture) fails on every Linux x64-asan run since
oven-sh#36598 landed (builds
[89023](https://buildkite.com/bun/bun/builds/89023),
[89031](https://buildkite.com/bun/bun/builds/89031)):

```
direct leak of 24b in run (src/runtime/node/node_crypto_binding.rs:85:21) +34 more
SUMMARY: AddressSanitizer: 1480 byte(s) leaked in 35 allocation(s).
  #6 EVP_PKEY_keygen vendor/boringssl/crypto/evp/evp_ctx.cc
  #7 Bun::KeyPairJobCtx::runTask src/jsc/bindings/node/crypto/CryptoGenKeyPair.cpp:23
  #8 Bun__RsaKeyPairJobCtx__runTask src/jsc/bindings/node/crypto/CryptoGenRsaKeyPair.cpp:26
```

## Cause

The 11 extern crypto job ctxs (generateKeyPair x5, sign/verify,
diffieHellman, hkdf, generatePrime, checkPrime, generateKey) completed
by invoking the JS callback from inside C++ `runFromJS` while the ctx
was still alive; the ctx was freed only after `then()` returned. A
callback that never returns (the fixture calls `process.exit(0)` inside
it) stranded everything the ctx still owned: the generated `EVP_PKEY`,
`KeyObjectData` refs, `BIGNUM`s.

The leak is pre-existing; oven-sh#36598 made it observable by routing
`OPENSSL_malloc` through libc under ASAN. Whether LSan reported the
other job types too was codegen luck (their pointers happened to be
reachable by the conservative stack scan); `generateKeyPair`'s
`EVP_PKEY` sits behind two FastMalloc indirections and was reported
deterministically.

## Fix

Make it structurally impossible for a job ctx to hold native resources
across user JS: the native side never sees the callback.

- `runFromJS` keeps its name (the JS-thread half, paired with the
work-pool half `runTask`) but no longer receives the callback. It
returns `JSCallbackArgs`, a small by-value type whose constructors are
the only producers, so bodies read `return { err };` or `return {
jsNull(), publicKey, privateKey };`. The extern "C" shims copy it
through a typed out-pointer (C linkage cannot return a class type); the
Rust side consumes it as a slice.
- The Rust `extern_crypto_job!` plumbing does, in order: run `runFromJS`
to produce the arguments, free the ctx (`ctx_deinit`), invoke the
callback. The invariant lives in one place and applies to every job
type.
- Shutdown release: a completion task enqueued but not yet dispatched
when `process.exit()` runs (exit racing the work pool) used to be
re-queued at shutdown, stranding the ctx the same way. `AnyTaskJob` now
carries an erased release entry and the shutdown release frees the job
without running its completion. A completion posted after the final
drain is not recoverable without joining the work pool (which would
block exit); `test-crypto-op-during-process-exit.js` stays in
`no-validate-leaksan.txt` for that sliver, now with an accurate comment.
- The caught-export-exception paths encoded the `JSC::Exception` cell
itself, so the callback's err argument was not the thrown Error (not
`instanceof Error`, no `code`). They now use `Exception::value()`,
matching node: JWK export of an unsupported curve surfaces
`ERR_CRYPTO_JWK_UNSUPPORTED_CURVE`.

No behavior change otherwise:

- `Bun__EventLoop__runCallback{1,2,3}` were Rust's
`EventLoop::run_callback` exported to C++. The plumbing now calls
`run_callback` directly: same enter/exit bracketing, same
pending-exception gate, same unhandled-exception reporting, same
synchronous timing. This made `runCallback1`/`runCallback3` dead (the
crypto bodies were their last callers), so their exports and
declarations are deleted; `runCallback2` stays for the webview backends.
- Callback arity is preserved per path (observable via
`arguments.length`): error paths pass 1 arg, results 2, generateKeyPair
success 3.
- Exception paths are preserved: a throw out of argument production
skips the callback and reports unhandled, as before. Each `runFromJS`
checks its `ThrowScope` after every call that can throw
(`RETURN_IF_EXCEPTION`), since the check that used to happen inside the
nested `runCallbackN` call now happens after the C++ scope destructs;
`BUN_JSC_validateExceptionChecks` verifies this on the asan lane.
- The produced `JSValue`s live on the `then()` stack frame between
production and invocation, which JSC's conservative scan covers; they
are JS-heap values, so freeing the ctx first cannot invalidate them.
- Perf: same number of FFI crossings, no allocation added.

The Rust-native crypto jobs (pbkdf2, scrypt, random) already had the
ordering property: they resolve promises or queue the callback via
nextTick, so their ctx drops before user JS runs. The
synchronous-callback extern jobs were the gap.

## Verification

New tests in `crypto.key-objects.test.ts`:

- `isASAN`-gated leak suite: children run with
`BUN_DESTRUCT_VM_ON_EXIT=1` and `detect_leaks=1` (the asan lane's
configuration) and call `process.exit(0)` from the callback of each job
type: generateKeyPair (KeyObject and encrypted PEM outputs), sign,
diffieHellman, hkdf, checkPrime, generateKey, plus an
exit-before-completion-dispatch case (busy-spin so the queued completion
is never dispatched).
- An export-error test: `generateKeyPair('ec', { namedCurve:
'secp224r1', ...jwk encodings })` asserts the callback err is
`instanceof Error` with code `ERR_CRYPTO_JWK_UNSUPPORTED_CURVE` (matches
node; fails on main, which passes the Exception cell).

Results:

- unfixed build (src stashed): both generateKeyPair leak tests fail with
the exact CI signature (`Direct leak of 24 byte(s)` in `EVP_PKEY_keygen`
via `KeyPairJobCtx::runTask`)
- fixed build: all pass, including under
`BUN_JSC_validateExceptionChecks=1`, and ec/ed25519 keypair and verify
probes run leak-clean as well
- `AsyncLocalStorage-tracking.test.ts`: 74 pass, 0 fail (all
async-context crypto fixtures, against both bun and node)
- `crypto.test.ts` (369), `crypto.key-objects.test.ts` (117), and 37
node parallel files (`test-crypto-keygen*`, `test-crypto-sign-verify`,
`test-crypto-hkdf`, `test-crypto-dh-stateless`, `test-crypto-*prime*`)
all pass

The break landed with oven-sh#36598 (which made the leak visible); oven-sh#36657
proposed clearing individual ctx fields before the callback, and this PR
supersedes that approach with the ordering guarantee in the job plumbing
instead of per-field resets.

<!-- robobun:evidence:begin -->

---

**no test proof** · iteration 0 · Platform-specific test(s) that do not
run on this machine. Deferring to CI, which covers all platforms:
test/js/node/async_hooks/AsyncLocalStorage-tracking.test.ts
test/js/node/crypto/crypto.key-objects.test.ts

<!-- robobun:evidence:end -->
springmin pushed a commit that referenced this pull request Aug 6, 2026
…llback (oven-sh#36986)

## What

`test/js/node/async_hooks/AsyncLocalStorage-tracking.test.ts` (the
crypto-generateKeyPair fixture) fails on every Linux x64-asan run since
oven-sh#36598 landed (builds
[89023](https://buildkite.com/bun/bun/builds/89023),
[89031](https://buildkite.com/bun/bun/builds/89031)):

```
direct leak of 24b in run (src/runtime/node/node_crypto_binding.rs:85:21) +34 more
SUMMARY: AddressSanitizer: 1480 byte(s) leaked in 35 allocation(s).
  #6 EVP_PKEY_keygen vendor/boringssl/crypto/evp/evp_ctx.cc
  #7 Bun::KeyPairJobCtx::runTask src/jsc/bindings/node/crypto/CryptoGenKeyPair.cpp:23
  #8 Bun__RsaKeyPairJobCtx__runTask src/jsc/bindings/node/crypto/CryptoGenRsaKeyPair.cpp:26
```

## Cause

The 11 extern crypto job ctxs (generateKeyPair x5, sign/verify,
diffieHellman, hkdf, generatePrime, checkPrime, generateKey) completed
by invoking the JS callback from inside C++ `runFromJS` while the ctx
was still alive; the ctx was freed only after `then()` returned. A
callback that never returns (the fixture calls `process.exit(0)` inside
it) stranded everything the ctx still owned: the generated `EVP_PKEY`,
`KeyObjectData` refs, `BIGNUM`s.

The leak is pre-existing; oven-sh#36598 made it observable by routing
`OPENSSL_malloc` through libc under ASAN. Whether LSan reported the
other job types too was codegen luck (their pointers happened to be
reachable by the conservative stack scan); `generateKeyPair`'s
`EVP_PKEY` sits behind two FastMalloc indirections and was reported
deterministically.

## Fix

Make it structurally impossible for a job ctx to hold native resources
across user JS: the native side never sees the callback.

- `runFromJS` keeps its name (the JS-thread half, paired with the
work-pool half `runTask`) but no longer receives the callback. It
returns `JSCallbackArgs`, a small by-value type whose constructors are
the only producers, so bodies read `return { err };` or `return {
jsNull(), publicKey, privateKey };`. The extern "C" shims copy it
through a typed out-pointer (C linkage cannot return a class type); the
Rust side consumes it as a slice.
- The Rust `extern_crypto_job!` plumbing does, in order: run `runFromJS`
to produce the arguments, free the ctx (`ctx_deinit`), invoke the
callback. The invariant lives in one place and applies to every job
type.
- Shutdown release: a completion task enqueued but not yet dispatched
when `process.exit()` runs (exit racing the work pool) used to be
re-queued at shutdown, stranding the ctx the same way. `AnyTaskJob` now
carries an erased release entry and the shutdown release frees the job
without running its completion. A completion posted after the final
drain is not recoverable without joining the work pool (which would
block exit); `test-crypto-op-during-process-exit.js` stays in
`no-validate-leaksan.txt` for that sliver, now with an accurate comment.
- The caught-export-exception paths encoded the `JSC::Exception` cell
itself, so the callback's err argument was not the thrown Error (not
`instanceof Error`, no `code`). They now use `Exception::value()`,
matching node: JWK export of an unsupported curve surfaces
`ERR_CRYPTO_JWK_UNSUPPORTED_CURVE`.

No behavior change otherwise:

- `Bun__EventLoop__runCallback{1,2,3}` were Rust's
`EventLoop::run_callback` exported to C++. The plumbing now calls
`run_callback` directly: same enter/exit bracketing, same
pending-exception gate, same unhandled-exception reporting, same
synchronous timing. This made `runCallback1`/`runCallback3` dead (the
crypto bodies were their last callers), so their exports and
declarations are deleted; `runCallback2` stays for the webview backends.
- Callback arity is preserved per path (observable via
`arguments.length`): error paths pass 1 arg, results 2, generateKeyPair
success 3.
- Exception paths are preserved: a throw out of argument production
skips the callback and reports unhandled, as before. Each `runFromJS`
checks its `ThrowScope` after every call that can throw
(`RETURN_IF_EXCEPTION`), since the check that used to happen inside the
nested `runCallbackN` call now happens after the C++ scope destructs;
`BUN_JSC_validateExceptionChecks` verifies this on the asan lane.
- The produced `JSValue`s live on the `then()` stack frame between
production and invocation, which JSC's conservative scan covers; they
are JS-heap values, so freeing the ctx first cannot invalidate them.
- Perf: same number of FFI crossings, no allocation added.

The Rust-native crypto jobs (pbkdf2, scrypt, random) already had the
ordering property: they resolve promises or queue the callback via
nextTick, so their ctx drops before user JS runs. The
synchronous-callback extern jobs were the gap.

## Verification

New tests in `crypto.key-objects.test.ts`:

- `isASAN`-gated leak suite: children run with
`BUN_DESTRUCT_VM_ON_EXIT=1` and `detect_leaks=1` (the asan lane's
configuration) and call `process.exit(0)` from the callback of each job
type: generateKeyPair (KeyObject and encrypted PEM outputs), sign,
diffieHellman, hkdf, checkPrime, generateKey, plus an
exit-before-completion-dispatch case (busy-spin so the queued completion
is never dispatched).
- An export-error test: `generateKeyPair('ec', { namedCurve:
'secp224r1', ...jwk encodings })` asserts the callback err is
`instanceof Error` with code `ERR_CRYPTO_JWK_UNSUPPORTED_CURVE` (matches
node; fails on main, which passes the Exception cell).

Results:

- unfixed build (src stashed): both generateKeyPair leak tests fail with
the exact CI signature (`Direct leak of 24 byte(s)` in `EVP_PKEY_keygen`
via `KeyPairJobCtx::runTask`)
- fixed build: all pass, including under
`BUN_JSC_validateExceptionChecks=1`, and ec/ed25519 keypair and verify
probes run leak-clean as well
- `AsyncLocalStorage-tracking.test.ts`: 74 pass, 0 fail (all
async-context crypto fixtures, against both bun and node)
- `crypto.test.ts` (369), `crypto.key-objects.test.ts` (117), and 37
node parallel files (`test-crypto-keygen*`, `test-crypto-sign-verify`,
`test-crypto-hkdf`, `test-crypto-dh-stateless`, `test-crypto-*prime*`)
all pass

The break landed with oven-sh#36598 (which made the leak visible); oven-sh#36657
proposed clearing individual ctx fields before the callback, and this PR
supersedes that approach with the ordering guarantee in the job plumbing
instead of per-field resets.

<!-- robobun:evidence:begin -->

---

**no test proof** · iteration 0 · Platform-specific test(s) that do not
run on this machine. Deferring to CI, which covers all platforms:
test/js/node/async_hooks/AsyncLocalStorage-tracking.test.ts
test/js/node/crypto/crypto.key-objects.test.ts

<!-- robobun:evidence:end -->
springmin pushed a commit that referenced this pull request Aug 16, 2026
…wo parallel vecs (oven-sh#39145)

### Problem
- `LOLHTMLContext` in `src/runtime/api/html_rewriter.rs` keeps two vecs,
`selectors` and `element_handlers`, that describe one thing: entry `i`
of each is the selector and the handler object from the same
`rewriter.on(selector, handlers)` call.
- The pairing is only held up by convention: `on_()` pushes to both,
`build_settings()` zips them back together, and a doc comment plus an
invariant comment explain it. The mordant `parallel_vecs` lint flags
this (the one baselined finding for this file).

### Fix
- Add `ElementHandlerEntry { selector, handler: Box<ElementHandler> }`
and store `element_handlers: Vec<ElementHandlerEntry>`. One vec, one
push in `on_()`, and `build_settings()` destructures each entry instead
of zipping.
- No behavior change: the same values are pushed in the same order, the
handler is still boxed (the lol-html closures built in
`build_settings()` hold raw pointers into the box, so it must not move
when the vec reallocates), and the body of the `build_settings()` loop
is unchanged. The `#[expect(clippy::vec_box)]` comes off
`element_handlers` because it is no longer a `Vec<Box<_>>`;
`document_handlers` keeps its own.
- Remove the `parallel_vecs:src/runtime/api/html_rewriter.rs` line from
`mordant-baseline.toml`.
- Tests, in `test/js/workerd/html-rewriter.test.js` (`on()
registrations`), pin down the two things this storage has to get right.
They pass before and after this change, since it is a refactor:
- Many selectors registered on one rewriter, with two rejected `on()`
calls in the middle, each still run the handlers they were registered
with, on two transforms of the same rewriter.
- `on()` called from inside a handler, often enough to reallocate the
registry while lol-html is still calling the handlers registered before
the transform started: the running transform is unaffected and the next
one picks the additions up. With the `Box` removed from
`ElementHandlerEntry` this test fails under ASAN with a
heap-use-after-free (report in the details below), so the boxing is now
covered rather than only commented.
- Verified:
- `bun bd test` on `test/js/workerd/html-rewriter.test.js` (165 tests,
including the new ones), `html-rewriter-end-error.test.ts`,
`html-rewriter-leak.test.ts`,
`test/js/web/html/html-rewriter-doctype.test.ts` and the HTMLRewriter
regression tests: all pass.
  - `cargo clippy -p bun_runtime --no-deps`: clean.
- `cargo dylint --all -p bun_runtime` with this baseline: nothing over
the baseline. The same command with the baseline line removed but the
source change stashed reports exactly the one `parallel_vecs` finding
for this file, so the removed line is the one this change fixes.
- Regenerating the baseline with `MORDANT_BASELINE_WRITE=1` also drops
two entries this PR does not touch
(`always_unwrapped_option:src/install/PackageInstall.rs`,
`narrowed_two_ways:src/runtime/node/node_crypto_binding.rs`); those
findings were already fixed on main by other changes and are left for a
separate cleanup.

### Background
- `HTMLRewriter.on(selector, handlers)` parses the CSS selector with
lol-html and wraps the JS handler object in an `ElementHandler` (the
protected `element`/`comments`/`text` callbacks). Nothing is handed to
lol-html at that point; registrations are collected in `LOLHTMLContext`,
which is shared by the rewriter and every transform it starts, because
`transform()` can run more than once.
- `build_settings()` runs at transform time and turns each registration
into a `(selector, ElementContentHandlers)` pair for lol-html. Its
closures capture a `NonNull<ElementHandler>` pointing into the heap
allocation owned by the `Box`, which is why the handler has to stay
boxed even though clippy would normally suggest otherwise. An `on()`
call after a transform has started (for example from inside a handler)
pushes onto the same vec, which is what makes the reallocation case
reachable from JS.
- `mordant-baseline.toml` is the ratchet for the mordant lint pack run
by the Rust lints workflow: it records the accepted number of findings
per (lint, file), and CI reports anything above those counts. Removing
the line here means a reintroduction of the pattern in this file would
be reported.

<details>
<summary>ASAN report from the new test with the Box removed from
ElementHandlerEntry</summary>

```
ERROR: AddressSanitizer: heap-use-after-free
READ of size 8
    #3 <ElementHandler as HandlerLike>::global            src/runtime/api/html_rewriter.rs
    #4 handler_callback::<ElementHandler, Element, ...>   src/runtime/api/html_rewriter.rs
    #5 ElementHandler::on_element                         src/runtime/api/html_rewriter.rs
    #6 build_settings::{closure#0}                        src/runtime/api/html_rewriter.rs
    #8 lol_html ContentHandlersDispatcher::handle_start_tag
freed by thread T0 here:
    #13 RawVec<ElementHandlerEntry>::grow_one
    #15 Vec<ElementHandlerEntry>::push
    oven-sh#16 HTMLRewriter::on_                                 src/runtime/api/html_rewriter.rs
```

</details>

<!-- robobun:evidence:begin -->

---

**no test proof** · iteration 2 · Platform-specific test(s) that do not
run on this machine. Deferring to CI, which covers all platforms:
test/js/workerd/html-rewriter.test.js

<!-- robobun:evidence:end -->

---------

Co-authored-by: Alistair Smith <hi@alistair.sh>
springmin pushed a commit that referenced this pull request Sep 3, 2026
…p_until (oven-sh#41149)

Fuzzilli found an ASAN stack-buffer-overflow in `run_tasks` ->
`Log::add_msg` -> `Vec::push` during runtime auto-install.

## What happened

`enqueue_dependency_to_root` blocks in `sleep_until`. `sleep_until`
ticks the JS event loop between `is_done` polls. Each poll calls
`run_tasks`, which reads `manager.log` to log manifest 4xx errors.

The event loop tick can run module transpilation or
`AsyncModule::resume_loading_module`. Both swap the VM's log pointers
(`jsc_vm.log`, `transpiler.log`, `resolver.log`, `linker.log`, `pm.log`)
and restore them on exit. They do not save from the same source:

- `resolve_maybe_needs_trailing_slash` saves `jsc_vm.log` and swaps
every pointer except `transpiler.log`.
- `transpile_source_code_inner` saves `transpiler.log` and restores
`pm.log` to it.
- `AsyncModule::resume_loading_module` saves `jsc_vm.log` but restores
`transpiler.log` to it.

When a module transpile runs during a resolve's `sleep_until` tick, its
restore sets `pm.log` to the old `transpiler.log`, not the resolve's
scoped log. The 404 diagnostics from `run_tasks` then land in the wrong
log. With more interleaving across calls `pm.log` ends up at a dead
stack `Log`, and the next `run_tasks` poll reads it.

```
READ of size 8 at 0x7ffd98a149c0 thread T0
    #0  RawVecInner::capacity
    #1  Log::add_msg (lib.rs:2369)
    #3  Log::add_error_fmt (lib.rs:2032)
    #4  run_tasks (runTasks.rs:494)
    #5  Closure::is_done (PackageManagerEnqueue.rs:534)
    #7  AnyEventLoop::tick_raw (AnyEventLoop.rs:148)
    #8  PackageManager::sleep_until (PackageManager.rs:1063)
    #9  enqueue_dependency_to_root (PackageManagerEnqueue.rs:571)
    ...
    oven-sh#18 resolve_maybe_needs_trailing_slash (VirtualMachine.rs:4259)

Address is located in stack of thread T0 in frame
    #0  to_js_host_call (host_fn.rs:677)
  [144, 152) 'scope' <== Memory access at offset 160 overflows this variable
  [176, 240) 'scope_storage'
```

## Fix

- Swap and restore `transpiler.log` with the other log pointers in
`resolve_maybe_needs_trailing_slash`. It can no longer drift from
`jsc_vm.log` across a resolve.
- Snapshot `pm.log` on entry to `enqueue_dependency_to_root`'s
`sleep_until`. Re-assert it before each `run_tasks` poll. `run_tasks`
always sees the caller's log, whatever the event loop tick did to it.

## Test

The test runs a 404 registry in the parent process on `port: 0`. The
child queues `require()` calls with `setImmediate`, so they run during
`sleep_until`'s event-loop tick and trigger
`transpile_source_code_inner`'s log swap. Without the fix, the 404
errors land in the VM log and print to stderr at exit. With the fix,
stderr is empty.

Verified on current `main` (6f27257): the test fails without the
source change and passes with it.

Supersedes oven-sh#31120, which the stale bot closed after 90 days. Same
change, rebased onto current main; moved to this branch so the fix can
be tracked.

<!-- robobun:evidence:begin -->

---

**[human-review]** gate passed · iteration 6 · 3 files touched

<details><summary>fails on main (without fix)</summary>

```console
ASAN without fix: 1 FAILED
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/pr_gate.xml" test/js/bun/resolve/resolve-autoinstall-log-dangling.test.ts
bun test v1.4.1 (a6c4cc2)

test/js/bun/resolve/resolve-autoinstall-log-dangling.test.ts:
(pass) repeated failing auto-install resolves at varying stack depth don't read a dangling pm.log [1075.94ms]
114 |     stderr: "pipe",
115 |   });
116 | 
117 |   const [stdout, stderr, exitCode] = await Promise.all([proc.stdout.text(), proc.stderr.text(), proc.exited]);
118 | 
119 |   expect(stderr).toBe("");
                       ^
error: expect(received).toBe(expected)

- ""
+ "error: GET http://localhost:37987/autoinstall-missing-pkg-0 - 404
+ 
+ error: GET http://localhost:37987/autoinstall-missing-pkg-1 - 404
+ 
+ error: GET http://localhost:37987/autoinstall-missing-pkg-2 - 404
+ 
+ error: GET http://localhost:37987/autoinstall-missing-pkg-3 - 404
+ 
+ error: GET http://localhost:37987/autoinstall-missing-pkg-4 - 404
+ 
+ error: GET http://localhost:37987/autoinstall-missing-pkg-5 - 404
+ 
+ error: GET http://localhost:37987/autoinstall-missing-pkg-6 - 404
+ 
+ error: GET http://
... (truncated)

release without fix: 1 FAILED
bun test v1.4.1-canary.1 (a6c4cc2)

test/js/bun/resolve/resolve-autoinstall-log-dangling.test.ts:
(pass) repeated failing auto-install resolves at varying stack depth don't read a dangling pm.log [22.63ms]
114 |     stderr: "pipe",
115 |   });
116 | 
117 |   const [stdout, stderr, exitCode] = await Promise.all([proc.stdout.text(), proc.stderr.text(), proc.exited]);
118 | 
119 |   expect(stderr).toBe("");
                       ^
error: expect(received).toBe(expected)

- ""
+ "error: GET http://localhost:35697/autoinstall-missing-pkg-0 - 404
+ 
+ error: GET http://localhost:35697/autoinstall-missing-pkg-1 - 404
+ 
+ error: GET http://localhost:35697/autoinstall-missing-pkg-2 - 404
+ 
+ error: GET http://localhost:35697/autoinstall-missing-pkg-3 - 404
+ 
+ error: GET http://localhost:35697/autoinstall-missing-pkg-4 - 404
+ 
+ error: GET http://localhost:35697/autoinstall-missing-pkg-5 - 404
+ 
+ error: GET http://localhost:35697/autoinstall-missing-pkg-6 - 404
+ 
+ error: GET http://localhost:35697/autoinstall-missing-pkg-7 - 404
+ 
+ error: GET http://localhost:35697/autoinstall-missing-pkg-8 - 404
+ 
+ error: GET http://localhost:35697/autoinstall-missing-pkg-9 - 
... (truncated)
```

</details>

<details><summary>passes on PR (with fix)</summary>

```console
ASAN with fix: all passed
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test "--reporter=junit" "--reporter-outfile=/tmp/pr_gate.xml" test/js/bun/resolve/resolve-autoinstall-log-dangling.test.ts
bun test v1.4.1 (a6c4cc2)

test/js/bun/resolve/resolve-autoinstall-log-dangling.test.ts:
(pass) repeated failing auto-install resolves at varying stack depth don't read a dangling pm.log [1104.72ms]
(pass) module transpile during auto-install's event-loop tick doesn't desync pm.log [689.71ms]

 2 pass
 0 fail
 8 expect() calls
Ran 2 tests across 1 file. [3.73s]
__F:0:S:0

release with fix: all passed
$ bun scripts/build.ts --profile=release
[configured] bun-profile → bun (stripped) in 590ms (unchanged)
ninja: Entering directory `/workspace/bun/build/release'
[1/141] gen generated_host_exports.rs
generated_host_exports.rs: 122 exports (host=5, lazy=10, generic=107, rust=0); 244 extern-C blocks audited
[2/141] gen JS modules (bundle-modules)
Preprocess modules (8046ms)
Bundle modules (666ms)
Postprocesss modules (969ms)
Bundle Functions (629ms)
Generate Code (40ms)

[10.36s] Bundled "src/js" for production
  2594 kb
  197 internal modules
  13 native modules
  50 internal functions across 16 files
[2/141] cargo bun_runtime → libbun_runtime.a
�[1m�[92m   Compiling�[0m bun_output_tags v0.0.0 (/workspace/bun/src/bun_output_tags)
�[1m�[92m   Compiling�[0m bun_core v0.0.0 (/workspace/bun/src/bun_core)
�[1m�[92m   Compiling�[0m bun_parsers v0.0.0 (/workspace/bun/src/parsers)
�[1m�[92m   Compiling�[0m bun_install v0.0.0 (/workspace/bun/src/install)
�[1m�[92m   Compiling�[0m bun_jsc v0.0.0 (/workspace/bun/src/jsc)
�[1m�[92m   Compiling�[0m bun_dispatch v0.0.0 (/workspace/bun/src/dispatch)
�[1m�[92m   Compiling�[0m bun_jsc_macros v0.0.0 (/workspace/bun/src/jsc_macros)

... (truncated)
```

</details>

<details><summary>diff hotspot</summary>

```
.../PackageManager/PackageManagerEnqueue.rs        |  6 ++
 src/jsc/VirtualMachine.rs                          |  5 ++
 .../resolve-autoinstall-log-dangling.test.ts       | 65 ++++++++++++++++++++++
 3 files changed, 76 insertions(+)
```

</details>

**gate history** · 1 passed · 0 rejected · iteration 6

<details><summary>evidence per changed file</summary>

```
file                                                      reads  edits  tests
src/install/PackageManager/PackageManagerEnqueue.rs           1      2     19
src/jsc/VirtualMachine.rs                                     5      3     19
…js/bun/resolve/resolve-autoinstall-log-dangling.test.ts      3      9     18
```

</details>

<!-- robobun:evidence:end -->
springmin pushed a commit that referenced this pull request Sep 30, 2026
…#43900)

### Problem
- `test/js/bun/spawn/spawn-stdio-syscall-error.test.ts` is red on
alpine: `"lost": -82000`, not `0` (build 120303). oven-sh#43739 added the case,
not the bug.
- `read_loop` (`src/io/PipeReader.rs:766`) delivers the bytes read
before a failed read, then the error. The consumer asks for more inside
that delivery, and the reader reads the fd past the error. The error
comes late, or never.

### Fix
- `read_once` sets `PosixFlags::READ_FAILED` on a fatal error, and
`begin_read` does not read while it is set. The request parks, then
`on_reader_error` rejects it. Nothing clears the flag.
- Correct: libuv `uv__read` clears `UV_HANDLE_READABLE` before it
reports a read error.
- Verified: `test/js/bun/spawn/spawn-stdio-syscall-error.test.ts`, 17
pass. The four new cases fail without the fix. Suites: Notes.
- Self-reviewed: 15 concerns raised, 14 addressed. Not taken: a flag
reset in `start()`, which nothing on main needs.

### Background
- `PosixBufferedReader` reads the fd behind subprocess stdio, file
streams and the shell. Its parent gets `on_read_chunk`, then
`on_reader_done` or `on_reader_error`.
- `FileReader` is the parent behind a `ReadableStream`. `on_read_chunk`
resolves the parked `pull()`, and the reaction runs before it returns.
- Considered: a repeating shim failure hides the lost error. An error
stored in `FileReader` first needs a callback in every parent.
- oven-sh#43920 is a newer PR with the same change. Its test trigger is here.

### Downsides
- After a read error, a stream that ended with `'end'` now ends with
`'error'`, as in Node. With no listener the process stops.
- After a read error, `bun run --filter` no longer drains that pipe at
exit.
- Reads that do not fail pay nothing: `begin_read` tests one more bit.

<details><summary>Notes</summary>

**Trace without the fix** (bun 1.4.3-canary.1+367d939d9, shim logs each
`recv()`, writer `dd bs=1025`). The debug build of main at 8d36bff
does the same: 3 of 8 runs, two with no `'error'` and `lost` -7714150:

```
recv #5 len=262144 -> 95325
recv #6 len=166819 -> EIO (injected)      same fill_scratch call as #5
JS data 95325
recv #7 len=65536 -> 65536                 pull from inside the delivery, read_into
recv #8 .. oven-sh#116                            to EOF
{"received":452025,"got":8000125,"lost":-7548100,"events":["stdout.close","close"]}
```

No `'error'` event: the reads reached EOF before `on_reader_error` ran,
and a stored error is only returned by a later pull. When the reads park
first, `on_reader_error` rejects that pull and `'error'` comes late.
That is the CI signature: the events match and `lost` is a negative
multiple of 1025.

**Why alpine.** In CI the failure is injected: the shim fails only the
Nth `recv()`, so a later `recv()` succeeds and the extra bytes show. The
failing `recv()` must follow bytes in the same wakeup. BusyBox `head`
writes 1025 bytes at a time, so a wakeup often holds a short `recv()`
and then the failing one. coreutils `head` fills the buffer in one
`recv()`. The same run fails on debian when the writer is slow. With a
writer that copies BusyBox `head` (stdio, 1024-byte buffer), the
`RECV_AT=6` case fails 7 of 40 runs on bun 1.4.3-canary (release), and
36 of 40 when the writer also spins between chunks. This branch (debug):
0 of 120, and 0 of 60 with `dd bs=1025`.

**With no shim.** A child with one AF_UNIX socket as fd 0 and fd 1
writes a line to stdout and reads stdin. The peer leaves that line
unread, sends 8192 bytes and closes. The kernel gives the child 8192
bytes, then `ECONNRESET`, then EOF.

| Runtime | stdin events |
|---|---|
| Node v26.3.0 | `'error'` `ECONNRESET` after 8192 bytes |
| bun 1.4.3-canary.1+367d939d9 | `'end'` after 8192 bytes, no error |
| this branch | `'error'` `ECONNRESET` after 8192 bytes |

**The new cases.** Each one reaches the reader from inside the delivery
in a different way. Whole-file runs of the describe block:

| Case | Entry | Release, no fix | Debug, no fix | Debug, guard in
`read_into` only | Debug, this branch |
|---|---|---|---|---|---|
| `child_process`, `'data'` listener | pull, `read_into` | fails 6 of 6
| fails 5 of 5 | passes | passes |
| `child_process`, `'readable'` and `read()` | `set_flowing(true)`,
`read` | fails 6 of 6 | fails 5 of 5 | fails 3 of 3 | passes |
| `Bun.spawn`, `lazy` | pull, `read_into` | fails 6 of 6 | fails 5 of 5
| passes | passes |
| `Bun.spawn`, reader started at spawn | pull, `read_into` | fails 6 of
6 | fails 5 of 5 | passes | passes |

- The first three use counts: `SPAWN_FAULT_RECV_CAP=4096` makes every
`recv()` short, so `fill_scratch` calls `recv()` again in the same
wakeup. `SPAWN_FAULT_RECV_EAGAIN_AT=2` ends the first read loop, so the
consumer's read parks and the next read is poll-driven.
`SPAWN_FAULT_RECV_AT` then fails after bytes in that wakeup. For the
`'readable'` case it is 19: #3 to oven-sh#18 return 64 KiB, the highWaterMark,
so the reader is stopped and `read()` starts it again.
- The fourth uses state, and comes from oven-sh#43920: the writer waits for a
line on stdin, so the first read is parked when the bytes arrive.
`SPAWN_FAULT_RECV_MID_FILL=1` fails the `recv()` that follows one that
returned bytes, and `SPAWN_FAULT_READS_AFTER` counts the `recv()` calls
after it. Without the fix it is 1.
- With a count-based trigger and the reader started at spawn, the case
passed without the fix on a debug build: the buffered reader that runs
before JS reads `.stdout` took the bytes and the error. That is why this
case uses state.
- With the BusyBox-like writer at four speeds, the whole file passes 20
of 20 on this branch.

**With the fix**, `CAP=4096 EAGAIN_AT=2 RECV_AT=5`:

```
recv #3 -> 4096, recv #4 -> 4096, recv #5 -> EIO
JS data 8192
close(fd)
JS error EIO
```

**Node.** libuv `uv__read` (`src/unix/stream.c`): on a read error other
than `EAGAIN` it clears `UV_HANDLE_READABLE | UV_HANDLE_WRITABLE`, calls
`read_cb` with the error, then stops the watcher. It calls `read_cb`
once for each `read()`, so it never holds bytes and an error from one
batch.

**Placement.** EOF and the `maxBuffer` stop have the same guard at this
site: `close_if_final` closes the reader before the final chunk is
delivered. An error cannot use it, because a closed reader with no
stored error reads as a clean end. For the same reason `READ_FAILED` is
not part of `is_done()`.

**Parents** (14 `BufferedReaderParent` implementations, what each does
in `on_reader_error`):

- 3 release the fd: `FileReader`, `SubprocessPipeReader`, `Terminal`.
- 2 drop the reader: `FileResponseStream`, shell `subproc.rs`.
- 8 only do accounting: `filter_run.rs`, `multi_run.rs`,
`lifecycle_script_runner.rs`, `security_scanner.rs`, `git_runner.rs`,
both cron jobs, test `Worker.rs`. `lifecycle_script_runner.rs` and
`cron.rs` build a new reader with `init()` for each spawn.
- 1 is shared and can restart: the shell `IOReader`.

Only two read the same reader after an error. `filter_run.rs`
`drain_and_close_pipes` reads once more at exit. That read is now a
no-op, and `deinit()` follows. Before, it could reach a second terminal
callback and decrement `remaining_fds` twice. The shell
`IOReader::start()` does not restart a reader after a failed read on
main, because the fired one-shot poll still counts as registered. The
Windows reader gets one libuv callback for each read, so it has no
bytes-then-error batch.

**Left open.**

- The flag is permanent. oven-sh#39638 and oven-sh#37901 change `IOReader::start()` to
restart the shell's stdin reader. After this PR they must clear
`READ_FAILED` there, or a `cat` that follows a stdin read error gets no
data, no EOF and no error.
- Not in this PR: a batch that stops because the buffer is full is
labelled `ReadState::Eof` when the poll event carries the hangup
(`read_state`, `None if received_hup`). `Bun.write(file, proc.stdout)`
then writes 262144 of 400000 pending bytes and resolves. It is on main
and in 1.4.3-canary, and this PR does not change it. It needs its own
change.
- Not in this PR: release the fd on a read error once, in
`PosixBufferedReader::on_error`. oven-sh#41456 names it as a follow-up. oven-sh#41420,
oven-sh#41456 and oven-sh#42150 did it for one parent each.

**Suites run on the debug build:** `test/js/web/streams/streams.test.js`
(624 pass), `test/js/node/stream/node-stream.test.js` (112 pass),
`test/js/bun/spawn/spawn-streaming-stdout.test.ts` (pass),
`test/js/node/child_process/child_process.test.ts` (81 pass, 2 fail in
my container for reasons outside this change: "spawn in the default
shell" reads an empty `$SHELL`, and "extra stdio pipes are not
double-closed on GC" needs 5.0 s in a debug build against the 5 s
timeout, its script prints `OK`). With this branch's build, the test
file of oven-sh#43920 passes 17 of 17 in 3 runs.

oven-sh#43790 is open and edits the comment above the failing case. It changes
the event order in `native-readable.ts`, not the reader.

</details>

<!-- robobun:evidence:begin -->

---

**no test proof** · iteration 0 · platform-specific test(s) that do not
run on this machine, deferring to CI, which covers all platforms:
test/js/bun/spawn/spawn-stdio-syscall-error.test.ts

<!-- robobun:evidence:end -->
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants