Skip to content

http: never pool a proxy tunnel with a fatal TLS error or pending I/O - #32742

Merged
Jarred-Sumner merged 2 commits into
mainfrom
claude/fix-proxy-tunnel-stale-ctx-uaf
Jun 26, 2026
Merged

Jarred-Sumner merged 2 commits into
mainfrom
claude/fix-proxy-tunnel-stale-ctx-uaf

Conversation

@Jarred-Sumner

Copy link
Copy Markdown
Collaborator

What does this PR do?

Fixes a use-after-free in the HTTP client's CONNECT proxy tunnel, caught by ASAN:

READ of size 8 at 0x61e00001fe80 thread T6
  #0 Option<RefPtr<ProxyTunnel>>::as_ref
  #1 proxy_tunnel::on_close            ProxyTunnel.rs:525
  #2 SSLWrapper::trigger_close_callback uws/lib.rs:833
  #3 SSLWrapper::handle_reading         uws/lib.rs:1053
  ...
freed by thread T6 here (same stack, same `handle_reading` call):
  #5 AsyncHTTP::on_async_http_callback_raw            AsyncHTTP.rs:819
  #7 HTTPClient::send_progress_update_without_stage_check
  #9 proxy_tunnel::on_data              ProxyTunnel.rs:350
  #11 SSLWrapper::trigger_data_callback uws/lib.rs:824
  #12 SSLWrapper::handle_reading        uws/lib.rs:1046

SSLWrapper::handle_reading flushes pending decrypted bytes to the data callback, then runs the close callback, guarded only by closed_notified:

  1. The flushed data callback completes a keep-alive response through the tunnel. A fatal TLS record error sets only fatal_error — none of the shutdown flags — so the wrapper passed tunnel_poolable's !is_shutdown() check and the tunnel was handed to the keep-alive pool. Nothing called wrapper.shutdown(), so closed_notified was never latched. Dispatching the final result then freed the ThreadlocalAsyncHTTP that embeds the HTTPClient.
  2. The guard (ssl.is_none() || closed_notified()) passes.
  3. trigger_close_callback() invokes on_close(handlers.ctx) with ctx pointing at the freed client.

The pooling branch is the only terminal path that doesn't go through close_proxy_tunnel(true) → wrapper.shutdown() → closed_notified, which is the latch the read loop relies on. SSLWrapper::shutdown already special-cases the close_notify flavor of this for exactly that reason; the fatal-error flavor never reaches shutdown().

The fix is one predicate: a tunnel whose wrapper has a fatal error or pending unconsumed input/output is not poolable. That routes it through the orderly teardown that latches closed_notified, and the pending-I/O half closes the same hole for a tunnel pooled from a mid-loop data callback while more decrypted bytes or queued output remain. Both are also required for the pool to be correct on its own terms — a poisoned or dirty TLS session must not be handed to the next request.

How did you verify your code works?

New regression test in test/js/bun/http/proxy.test.ts (next to the existing close_notify sibling): an HTTPS keep-alive response through a CONNECT proxy with a corrupt TLS record appended to the same TCP burst, followed by a second request that can only complete if the HTTP client thread survived the first.

Against an unfixed ASAN debug build the fixture aborts every run:

==20981==ERROR: AddressSanitizer: heap-use-after-free on address 0x61e00001fe80
READ of size 8 at 0x61e00001fe80 thread T6
...
exit=134

With this change it prints 4096 200 200 and exits 0 with no ASAN report. test/js/bun/http/proxy.test.ts (49/49), fetch-proxy-connect-tunnel-split-envelope.test.ts, fetch-proxy-tls-intern-race.test.ts, and fetch-keepalive.test.ts all pass.

A CONNECT tunnel's SSLWrapper read loop flushes decrypted bytes to
ProxyTunnel::on_data and then runs the close callback, guarded only by
`closed_notified`. When the flushed bytes complete a keep-alive response
and the next record is invalid, the fatal record error sets no shutdown
flag, so `tunnel_poolable` accepted the wrapper and the tunnel was handed
to the keep-alive pool: nothing called `wrapper.shutdown()`, so
`closed_notified` was never latched. Dispatching the final result then
freed the ThreadlocalAsyncHTTP that embeds the HTTPClient, and the
still-running read loop fired `on_close` into the freed `handlers.ctx`.

Require the wrapper to have no fatal error and no pending unconsumed
input/output before pooling. Every non-poolable terminal path goes
through `close_proxy_tunnel(true)`, which latches `closed_notified`
before the client is freed. The pending-I/O check closes the same hole
for a tunnel pooled from a mid-loop data callback while more decrypted
bytes or queued output remain. Both are also required for the pool to be
correct on its own terms: a poisoned or dirty TLS session must not be
handed to the next request.
@robobun

robobun commented Jun 26, 2026 •

Copy link
Copy Markdown
Collaborator
Updated 2:10 AM PT - Jun 26th, 2026

❌ @Jarred-Sumner, your commit 0ae24b7 has 1 failures in Build #64892 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 32742

That installs a local version of the PR into your bun-32742 executable, so you can run:

bun-32742 --bun

@github-actions

Copy link
Copy Markdown
Contributor

Found 1 issue this PR may fix:

  1. Windows: Bun 1.3.14 segfaults in long-running localhost HTTP proxy #32585 - Segfault in long-running localhost HTTP proxy using fetch() is consistent with a UAF from a poisoned CONNECT tunnel being returned to the keep-alive pool after a fatal TLS error

If this is helpful, copy the block below into the PR description to auto-close this issue on merge.

Fixes #32585

🤖 Generated with Claude Code

Comment thread src/http/lib.rs
BIO_ctrl_pending only sees encrypted bytes still in the BIOs. When
SSL_read partially returns a record (the output buffer filled mid
record), the record has already been pulled out of the rbio and its
remaining decrypted bytes are buffered inside the SSL, reported only by
SSL_pending. Without it, the pooling guard added in the previous commit
misses exactly the buffer-full flush case it exists for: both BIOs read
empty, the tunnel is pooled, and the read loop's next SSL_read returns
the rest of that record into a callback whose ctx was just freed.
@Jarred-Sumner

Copy link
Copy Markdown
Collaborator Author

Good catch — fixed in 0ae24b7. SSLWrapper::has_pending_data now also checks SSL_pending(ssl), which is exactly the state left by a partially-returned record at the buffer-full flush: the record is already out of the rbio and its remaining decrypted bytes are buffered inside the SSL, so both BIO_ctrl_pending probes read zero. (Per BoringSSL's header, for TLS that is the only place pending bytes can hide from the two BIOs — SSL_has_pending differs from SSL_pending > 0 only for DTLS.)

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The SSL_pending follow-up in 0ae24b7 addresses my earlier comment and the predicate now matches the stated invariant; I didn't find anything further, but a UAF fix in the HTTP client's TLS pooling path is worth a human sign-off.

Extended reasoning...

Overview

Tightens the tunnel_poolable predicate in src/http/lib.rs so a CONNECT proxy tunnel whose SSLWrapper has fatal_error set or any pending I/O is never returned to the keep-alive pool, routing it through the orderly teardown that latches closed_notified instead. SSLWrapper::has_pending_data() (src/uws/lib.rs) is broadened to also consult SSL_pending(ssl) so decrypted-but-unread bytes from a partially-returned record count as pending. src/boringssl_sys/boringssl.rs adds the SSL_pending FFI declaration, and a new regression test in test/js/bun/http/proxy.test.ts reproduces the original ASAN heap-use-after-free.

Security risks

The bug being fixed is itself a memory-safety / security issue: a UAF in the HTTP client thread, plus the possibility of handing a poisoned or dirty TLS session to an unrelated next request. The change is strictly a tightening — it only refuses to pool in additional cases and falls through to the existing close path — so it shouldn't introduce new exposure. SSLWrapper::has_pending_data() has no other callers, so broadening its semantics doesn't perturb anything else. No new attack surface is added.

Level of scrutiny

High. This is production-critical networking code at the intersection of TLS state-machine handling, callback re-entrancy, and connection-pool lifetime management — exactly the kind of subtle ownership reasoning where a second pair of human eyes is warranted even when the diff is small. The correctness argument hinges on the invariant that every non-poolable terminal path reaches closed_notified before handlers.ctx can go stale, and on has_pending_data() now covering all three places BoringSSL can hide bytes (rbio, wbio, SSL_pending).

Other factors

My earlier inline comment (the SSL_pending gap at the buffer-full flush) was addressed exactly as suggested in 0ae24b7, and the author's reply correctly notes that for TLS (vs DTLS) SSL_pending is the only remaining hiding place. The new test is well-constructed and the author reports it deterministically reproduces the UAF on an unfixed ASAN build. The robobun CI comment still shows a ❌ against the earlier commit 1eeb03b and hasn't refreshed for 0ae24b7, so CI status on the final revision is unconfirmed from the thread alone.

@Jarred-Sumner
Jarred-Sumner merged commit 863379d into main Jun 26, 2026
54 of 62 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the claude/fix-proxy-tunnel-stale-ctx-uaf branch June 26, 2026 07:45
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants