Skip to content

fetch: restore the one-shot gzip inflate for text() and arrayBuffer() - #43386

Merged
Jarred-Sumner merged 4 commits into
mainfrom
robobun/ed12a069/fetch-gzip-one-shot-buffered
Sep 24, 2026
Merged

Jarred-Sumner merged 4 commits into
mainfrom
robobun/ed12a069/fetch-gzip-one-shot-buffered

Conversation

@robobun

@robobun robobun commented Sep 18, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • Since fetch: bound decompressed output to reader demand #43123, a gzip response body that inflates to between 512 KB and 32 MB costs 1.5 to 2.2x the client CPU (as reported) when it is read with arrayBuffer() or text().
  • One exact-size libdeflate call used to inflate a body that arrived whole. It usually arrives before the caller has the Response, so the mode is still Flowing. decompress_bytes (src/http/InternalState.rs) budgets Flowing at 256 KB and runs zlib passes.

Fix

  • FetchTasklet starts in a new receive mode, Unclaimed: no consumer has attached. It is demand-driven, like Flowing.
  • When such a body is complete and undecoded, process_received_body (src/http/lib.rs) moves Unclaimed -> Paused and decodes nothing. The consumer that attaches resumes the transport. BufferAll gets the libdeflate call, and a reader gets budgeted zlib passes.
  • Correct because one compare-and-swap from Unclaimed decides who was first. A held body is decoded only when a consumer attaches or its connection ends (Notes).
  • Verified: 23 new cases in test/js/web/fetch/fetch-backpressure.test.ts. The debug-log case fails on main with two zlib passes. The whole file passes.

Background

  • BodyReceiveMode (src/http/Signals.rs) is the receive backpressure for a body handed to JS. Under Flowing and Paused the client decodes at most 256 KB per pass. BufferAll (.text(), .arrayBuffer()) never pauses.
  • libdeflate inflates a whole buffer into a whole buffer, so it cannot stop at a budget. zlib can, but is slower.
  • A gzip stream ends with ISIZE, the decoded size. It sizes the libdeflate output, up to 32 MB.
Notes

#43169 rewrites the same function. Please read this before choosing a merge order. #43169 (open) rewrites process_received_body and gives decompress_output_cap a parameter, so the two PRs conflict there. It does not fix this regression. I built its src/ at 31132d6 under this PR's tests: the debug-log case still fails, with the libdeflate attempt followed by a zlib pass over the body. So both changes are needed, and the one that lands second has to carry the hold through the other's process_received_body. One more case of this PR, held, and its Response is collected: the fetch is aborted, times out after 5 s on that build, in 3 of 3 runs. It takes 0.5 s on main and here. The three collected-Response cases that main already has pass on that build, so this is specific to a body that is complete and held when the Response is collected. I do not know if that is an intended change in #43169 or a hang. The test helper edits (serveConnectProxy counts CONNECTs) are the same lines in both PRs.

Where the numbers come from. Release builds of fd8422c (then main) with and without this diff, and of b52d513 (before #43123). This PR is the same diff on 367d939: it applied without changes, and none of the commits in between touch these hunks. I did not measure again after that rebase. Client CPU is process.cpuUsage() per request over keep-alive fetches of one gzip body. The origin is a separate process on other cores. The host is shared and ran at load 40 to 75. A 7-run median moved by about 13% between runs there, so only ratios inside one run are comparable.

Median us CPU per request, 9 interleaved runs x 1000 requests, arrayBuffer(), before #43123 / main / this diff:

  524,288 B (control)   215 /   221 /   234   ranges overlap
  786,432 B             316 /   419 /   286
1,048,576 B             431 /   534 /   381
1,572,864 B             552 /   751 /   547
4,194,304 B           1,904 / 2,315 / 1,949

Main is 1.22 to 1.36x here. This diff is within noise of the cost before #43123 on every size in the band. The ratios below 1.00 are noise, not a speedup: this diff does two more thread hops per response than the old code. text() gives the same shape (1 MiB 569 / 704 / 570). So does an origin that sends the head a tick before the body (1 MiB 323 / 442 / 316), which is the path where the hold happens with no callback.

A res.body reader costs the same as on main (1 MiB 629 / 585, 4 MiB 1,969 / 1,972). A 300 B keep-alive body, 15 runs x 3000 requests in two orders: 43 / 44 / 44 and 45 / 46 / 45.

I did not use bench/snippets/fetch-gzip.mjs. It runs the server in the client process and reports wall time. On this host its unchanged rows moved by 20 to 27% between binaries, so it could not resolve the effect.

What the tests prove. One case discriminates, and only in a debug build: it reads the HTTPInternalState debug log and expects one Decompressing N bytes with libdeflate line. On main it also sees Decompressing 6167 bytes and Decompressing 4583 bytes, the two zlib passes. From JavaScript the two paths deliver the same bytes by design, so nothing else can tell them apart. The other 22 cases pass on main too. They guard the new state: both framings, a buffered consumer and a reader, the body arriving with the head, after the head, and after the consumer, TLS, a CONNECT tunnel, an origin that closes the connection or the tunnel before a consumer attaches, a Response nobody reads (the process exits), and a collected Response (the fetch is aborted).

The CONNECT proxy cases clear NO_PROXY and the proxy variables for the child and assert that the proxy saw one CONNECT. An ambient NO_PROXY that lists 127.0.0.1 makes fetch() ignore its proxy option, and the case then passes without a tunnel. I observed 0 CONNECTs that way and 1 with the variables cleared.

Please check this one. handle_response_body_from_multiple_packets no longer sets is_libdeflate_fast_path_disabled after a pass over the final chunk. It sets it before a pass over a non-final chunk only. A held body needs the flag to stay clear so that the later pass can use libdeflate. I believe the rest is unchanged: decompress_bytes sets the flag itself when it enters the libdeflate block, and every later call sees a decoder that is not None. The chunked paths never set it after a final chunk.

At the end of the transport. The first version of this description said that nothing is decoded for a Response that nobody reads. That is true only while the connection is open. A review pointed out that finalize_body_on_eof decodes whatever is held with no budget, and a held body reaches it whole. Through a CONNECT tunnel the socket is never paused, so an origin that closes an idle connection gets there before any consumer. I confirmed it: a res.body reader whose tunnel closed first got its body from one Decompressing 6167 bytes with libdeflate call. Main has the same gap for the part of a body it has not decoded yet (#43123 lists it as still unbounded), and the hold never decodes more than main does. It does not close the gap. #43169 is the change that bounds it, so I did not copy that work here. The four close cases assert exact bytes, not memory.

What the hold costs. A reader of a held body gets its first chunk one thread round trip later. A held body that nobody reads keeps its connection with nothing decoded, where main keeps it with 256 KB decoded. For a gzip stream with a truthful trailer, the set of bodies is a subset of what #43123 already parks (a decoded size of 256 KB or more). The hold trusts the trailer. A stream whose trailer overstates its size is held too, where main would decode it and free the connection. That gives a server nothing new: it can already make a client park a connection with a real body of that size, which is a few KB on the wire.

Unchanged. Bodies up to 512 KB, deflate, brotli and zstd, HTTP/2 and HTTP/3 (they start Unclaimed but are never held, because their output cap is unbounded), S3 (its Store still starts Flowing), and any body that arrives in more than one read.

A gzip response body that arrives whole usually does so before the
caller has the Response. The receive mode is still Flowing then, so the
256 KB output budget from #43123 sends the body through zlib passes.
Before that budget, one exact-size libdeflate call inflated it.

FetchTasklet now starts in a new receive mode, Unclaimed: no consumer
has attached yet. When a complete gzip body whose trailer size is above
the 512 KB shared buffer and below 32 MB has not been touched by a
decoder, the HTTP thread moves Unclaimed to Paused and decodes nothing.
The consumer that attaches resumes the transport. A buffered consumer
(text, arrayBuffer) then gets the one libdeflate call. A reader gets
budgeted zlib passes as before. Nothing is decoded for a Response that
nobody reads.

The tests cover a buffered consumer and a reader, with the body arriving
with the head, after the head, and after the consumer, over TLS and
through a CONNECT proxy, for a Response nobody reads, and for a
collected Response. The CONNECT proxy cases clear NO_PROXY and the proxy
variables for the child and assert that the proxy saw one CONNECT.
@coderabbitai

coderabbitai Bot commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Review Change StackReview Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: oven-sh/bun/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: b24199fc-d79e-4359-b4c2-0427221c2a28

📥 Commits

Reviewing files that changed from the base of the PR and between 367d939 and 3ca187d.

📒 Files selected for processing (6)
  • src/http/HTTPThread.rs
  • src/http/InternalState.rs
  • src/http/Signals.rs
  • src/http/lib.rs
  • src/runtime/webcore/fetch/FetchTasklet.rs
  • test/js/web/fetch/fetch-backpressure.test.ts

Included review availability: Your plan provides up to 10 included reviews per hour; 9 remain after this review.


Walkthrough

Changes

The change adds bounded exact-size gzip inflation, introduces an Unclaimed body receive state, defers decompression while consumers are paused or held, preserves the final-chunk libdeflate path, and expands fetch backpressure tests.

Gzip backpressure and inflation

Layer / File(s) Summary
Bounded exact-size inflation
src/http/HTTPThread.rs, src/http/InternalState.rs
The shared libdeflate buffer size uses a named constant. Gzip trailer parsing is bounded. Exact-size allocation is limited to eligible bodies below 32 MiB; other cases use streaming inflation.
Unclaimed receive-state flow
src/http/Signals.rs, src/http/lib.rs, src/runtime/webcore/fetch/FetchTasklet.rs
BodyReceiveMode::Unclaimed and Store::unclaimed() track bodies awaiting a consumer. Receive transitions support this state. Body processing defers decompression while consumption is paused or held. Final whole-body processing can retain the libdeflate fast path.
Backpressure and gzip validation
test/js/web/fetch/fetch-backpressure.test.ts
Tests cover gzip framing, buffered and streaming consumers, TLS, CONNECT proxying, debug-build libdeflate calls, unconsumed responses, and aborted collected responses.

Suggested reviewers: jarred-sumner

Priority: ➖ Normal

Merge Risk: ⚪ Minimal · up to 3ca18

No concrete merge-blocking behavior is established at the current head.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly identifies the main change: restoring one-shot gzip inflation for text() and arrayBuffer().
Description check ✅ Passed The description clearly explains the problem, implementation, scope, performance impact, and verification results. It does not use the template headings, but it provides the required content in equiva…

Comment @coderabbitai help to get the list of available commands.

Comment thread src/http/InternalState.rs Outdated
Comment thread src/http/InternalState.rs Outdated
Comment thread src/http/InternalState.rs Outdated
@robobun

robobun commented Sep 18, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 5:09 PM PT - Sep 18th, 2026

✅ @robobun, your commit e98f60c7ff61a925817cef41791d5f85b430c54e passed in Build #118084! 🎉


🧪   To try this PR locally:

bunx bun-pr 43386

That installs a local version of the PR into your bun-43386 executable, so you can run:

bun-43386 --bun

Comment thread src/http/Signals.rs Outdated
Comment thread src/http/Signals.rs Outdated
Comment thread src/http/Signals.rs Outdated
Comment thread src/http/lib.rs Outdated
Comment thread src/http/lib.rs Outdated
No code changes. Each comment that spanned two or three lines now says
only what the code cannot, in one line.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline finding, I also checked the is_libdeflate_fast_path_disabled change the description flags: decompress_bytes sets the flag itself on entry to the libdeflate block (src/http/InternalState.rs:313) and any zlib pass leaves a non-None decompressor, so a second pass over a final chunk cannot re-enter the exact-size path unless the first pass decoded nothing (the held/paused case, which is the intent). I also looked at the window between reading the cap and the Unclaimed -> Paused CAS: a consumer attaching there flips the mode first, the CAS fails, and the cap is re-read, so no body is left undecoded with a waiting consumer.

Extended reasoning...

A confirmed finding is posted inline, so this note only records what else was examined. The PR description explicitly asks a reviewer to check the relocated is_libdeflate_fast_path_disabled assignment in handle_response_body_from_multiple_packets (src/http/lib.rs:4712-4716); reading decompress_bytes confirms the flag is set at src/http/InternalState.rs:313 whenever the libdeflate block is entered, and every other decode path constructs a Decompressor, which the entry condition at line 307 also rejects, so the only way to reach libdeflate on a later final-chunk pass is a body that was never decoded, i.e. the hold. The cap-read-then-CAS ordering in process_received_body (src/http/lib.rs:4120-4136) was also traced: receive_all/receive_on_demand on the JS thread move the mode away from Unclaimed before scheduling the resume, so either the CAS loses and max_output is re-read under the new mode, or it wins and the consumer's later transition out of Paused triggers the resume that decodes the pending body.

Additional findings (outside the current diff — GitHub can't attach inline comments there):

  • 🟣 src/http/lib.rs — A streaming reader of a held gzip body gets one unbudgeted chunk of up to 32 MiB when the origin closes the connection first. finalize_body_on_eof at src/http/lib.rs:2148 decodes with usize::MAX. Through a CONNECT tunnel the socket is never paused (maybe_pause_receive bails at src/http/lib.rs:4146), so an origin keep-alive timeout reaches on_close while the whole compressed body is still undecoded and Paused. Fix: on EOF for a body that is complete and held, decode under the consumer's cap (decompress_output_cap) and leave the rest pending for the pump, instead of passing usize::MAX.

    Extended reasoning...

    The finder called this pre-existing because InternalState.rs:257 is unchanged. It is, but the state that reaches it changes. On the base a held body already had 256 KB decoded and delivered; the hold now leaves the entire body undecoded (decompress_output_pending=true, nothing in decoded_body) at src/http/lib.rs:4131. Trigger: fetch through a proxy tunnel (or any path where the socket stays readable), the response body of 512 KB-32 MB decoded arrives whole, the caller obtains res.body but does not read it for longer than the origin's keep-alive idle time (commonly 5-15 s at CDNs). The origin closes; ProxyTunnel sees the close, on_close runs; in_progress is true, is_body_complete_on_close at InternalState.rs:257 returns true because decompress_output_pending && is_done. finalize_body_on_eof at InternalState.rs:275 calls process_body_buffer with usize::MAX; libdeflate inflates the full 32 MiB into decoded_body; progress_update at lib.rs:2153 hands the whole buffer to the FetchTasklet callback, which appends it to scheduled_response_buffer and the reader receives one 32 MiB chunk.…

    Verification: pre-existing. Trigger: an h1 gzip body of 512 KiB-32 MiB decoded arrives whole while the socket stays readable (CONNECT tunnel: maybe_pause_receive returns early on self.proxy_tunnel.is_some() at src/http/lib.rs:4146; or the on_data/JS-pause race on a direct socket), and the origin closes before the consumer drains it. Mechanism verified: after the hold (src/http/lib.rs:4130-4132,…

@robobun

robobun commented Sep 18, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status (head e98f60c)

How to reproduce, with a debug build:

bun bd test test/js/web/fetch/fetch-backpressure.test.ts -t 'arrived whole'
  • On main (367d939) the case res.bytes() inflates it in one libdeflate call fails. Its debug log shows Decompressing 6167 bytes with libdeflate, then the two zlib passes Decompressing 6167 bytes and Decompressing 4583 bytes.
  • On this branch the same case sees the libdeflate line only, and all 23 cases pass. The whole file passes (112 pass, 2 skip).
  • Four of the cases cover an origin that closes the connection, or the CONNECT tunnel, before a consumer attaches. They assert exact bytes, not memory. The description says what is decoded at that point and why this PR does not change it.
  • The CPU numbers and how I measured them are in the Notes of the description. I measured them on an older base with the same diff and did not measure again after the rebase.

Review state: an automated review of 3ca187d found the decode at the end of the transport, and the description and the four cases above answer it. An automated review of e98f60c reported no issues. No self-review of this diff has finished, so this PR has had no adversarial design review.

Two risks that no reviewer has examined, checked by me as the author:

  • A consumer that attaches to a fetch body without telling the producer would wait forever on a held body. Every first attach goes through on_start_buffering (Body.rs, Blob.rs for Bun.write) or on_start_streaming (to_readable_stream: .body, clone, a Response returned from Bun.serve, HTMLRewriter). The other on_receive_value assignment, in blob/write_file.rs, re-arms a Bun.write that has already attached. Main also parks every body that this PR holds, so such a consumer would wait on main too.
  • Nothing reads the initial receive mode as Flowing. HTTP/2, HTTP/3 and the HTTP/1.1 client ask only is_receive_paused or is_demand_driven. S3 and bun install still start Flowing.

The description names the refactor that I would most like a second pair of eyes on (is_libdeflate_fast_path_disabled), and what I found when I ran these tests against #43169. The merge order of the two PRs is a maintainer's call.

Four cases for a body that the client holds undecoded when the origin
ends the connection before a consumer attaches. On a direct socket the
client sees the close once the consumer resumes it. Through a CONNECT
tunnel the close arrives while the body is held, and the test waits for
the client to close its side before the consumer attaches. A buffered
consumer and a reader get the exact bytes in both.
@robobun

robobun commented Sep 18, 2026

Copy link
Copy Markdown
Collaborator Author

Thanks for the review. The finding about the end of the transport is correct, and my description claimed too much.

  • I corrected the description. It no longer says that nothing is decoded for a Response that nobody reads. A held body is decoded when a consumer attaches or when its connection ends, and at the end it is decoded whole with no budget. The Notes now say this and credit the review.
  • I added four cases (e98f60c): the origin closes the connection, or the CONNECT tunnel, before a consumer attaches. A buffered consumer and a reader get the exact bytes. For the reader through the tunnel, the debug log shows one Decompressing 6167 bytes with libdeflate call, so the test does reach the path you described. The cases assert bytes, not memory.
  • I did not change the decode at the end of the transport. Bounding it is what http: hand the undecoded rest of a response body to its consumer when the transport ends #43169 does, and a second copy here would make the conflict between the two PRs worse. A maintainer may decide otherwise.

The review says that a confirmed finding is posted inline. I see no inline comment from it on this PR. If one was intended, it did not arrive. The reviewed commit and the current head differ in comments and tests only.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

@Jarred-Sumner
Jarred-Sumner merged commit 2838e1b into main Sep 24, 2026
10 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the robobun/ed12a069/fetch-gzip-one-shot-buffered branch September 24, 2026 02:14
Jarred-Sumner added a commit that referenced this pull request Sep 24, 2026
… its consumer

main's #43386 held such a body in the HTTP client (Unclaimed -> Paused) until a
consumer attached. Here the request ends and the body goes to the consumer as a
HeldBody, so that hold is gone:

- process_received_body decodes nothing of it while no consumer has attached.
- HeldBody::decode inflates it in one exact-size libdeflate call for a consumer
  that takes the whole body. A reader gets budgeted zlib passes.
- FetchTasklet decodes nothing of a held body until a consumer attaches.
- Signals::hold_for_consumer is removed.

The test of a collected Response with a held body expected the origin to see
its connection close. The request has ended by then and the connection is back
in the pool, so the test now counts the freed tasklets and the decode passes.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants