Skip to content

fetch: reject malformed chunk-size tokens instead of reading them as zero - #34918

Merged
cirospaciari merged 5 commits into
mainfrom
farm/c6ed3847/fetch-chunked-size-strict
Sep 10, 2026
Merged

cirospaciari merged 5 commits into
mainfrom
farm/c6ed3847/fetch-chunked-size-strict

Conversation

@robobun

@robobun robobun commented Jul 21, 2026 •

Copy link
Copy Markdown
Collaborator

Repro

import net from "node:net";
const WIRE = "HTTP/1.1 200 OK\r\nTransfer-Encoding: chunked\r\n\r\n0x5\r\nhello\r\n0\r\n\r\n";
const srv = net.createServer(s => { s.on("error",()=>{}); s.once("data", () => s.end(WIRE)); });
await new Promise(r => srv.listen(0, "127.0.0.1", r));
try {
  const res = await fetch(`http://127.0.0.1:${srv.address().port}/`);
  console.log(`resolved: status=${res.status} body=${JSON.stringify(await res.text())}`);
} catch (e) { console.log("rejected:", e.code || e.name); }
srv.close(); process.exit(0);

Before: resolved: status=200 body="" (5 payload bytes silently dropped).
Node: rejected: TypeError (llhttp HPE_INVALID_CHUNK_SIZE).
After: rejected: InvalidHTTPResponse.

Cause

phr_decode_chunked in picohttpparser stops consuming the chunk-size at the first non-hex byte and falls through to CHUNKED_IN_CHUNK_EXT, which scans to LF without validating what came after the hex run. For 0x5\r\n it reads 0, sees x, breaks, skips x5 as "extension", and bytes_left_in_chunk == 0 makes it the last chunk. The response resolves with an empty body.

The already-rejected cases in the report (+5, leading non-hex, overflow) fail because _hex_count == 0 or the overflow guard fires; the hole is specifically "at least one hex digit, then garbage".

Fix

After the hex run, only ; (chunk-ext), CR, or LF are valid (RFC 9112 7.1). Anything else returns -1 and surfaces as InvalidHTTPResponse, matching llhttp strict mode. Verified against Node:

token Node Bun before Bun after
5 body="hello" body="hello" body="hello"
5;foo body="hello" body="hello" body="hello"
0x5 reject body="" reject
5g reject body="hello" reject
5 reject body="hello" reject
5.0 reject body="hello" reject

vendor/ is fetched at build time, so the change is applied as patches/picohttpparser/strict-chunk-size.patch and registered in scripts/build/deps/picohttpparser.ts. The server-side chunked parser in packages/bun-uws/src/ChunkedEncoding.h already rejects these tokens.

Verification

bun bd test test/js/web/fetch/chunked-trailing.test.js -t "rejects malformed"

6 parameterized cases fail on main (resolve instead of reject), pass with the patch; the chunk-ext positive case passes both ways as a regression guard.


[policy-decision:dep] gate passed · iteration 2 · 3 files touched

passes on PR (with fix)
Test-only change.

Debug/ASAN (expected pass):
$ bun bd test 'test/js/web/fetch/fetch-chunked-size.test.ts'
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test test/js/web/fetch/fetch-chunked-size.test.ts
bun test v1.4.3 (5f554969b)

test/js/web/fetch/fetch-chunked-size.test.ts:
(pass) fetch: chunked chunk-size token validation > rejects malformed chunk-size > token "0x5" [309.95ms]
(pass) fetch: chunked chunk-size token validation > rejects malformed chunk-size > token "5g" [24.55ms]
(pass) fetch: chunked chunk-size token validation > rejects malformed chunk-size > token "5 " [15.61ms]
(pass) fetch: chunked chunk-size token validation > rejects malformed chunk-size > token "5\t" [16.76ms]
(pass) fetch: chunked chunk-size token validation > rejects malformed chunk-size > token "5.0" [17.16ms]
(pass) fetch: chunked chunk-size token validation > rejects malformed chunk-size > token "5-" [16.78ms]
(pass) fetch: chunked chunk-size token validation > accepts well-formed chunk-size > token "5" [38.76ms]
(pass) fetch: chunked chunk-size token validation > accepts well-formed chunk-size > token "5;ext" [14.53ms]
(pass) fetch: chunked chunk-size token validation > accepts well-formed chunk-size > token "5;ext=1" [12.99ms]
(pass) fetch: chunked chunk-size token validation > accepts well-formed chunk-size > token "05" [14.30ms]
(pass) fetch: chunked chunk-size token validation > accepts well-formed chunk-size > token "A" [16.95ms]

 11 pass
 0 fail
 16 expect() calls
Ran 11 tests across 1 file. [2.82s]
Exit: 0
diff hotspot
patches/picohttpparser/strict-chunk-size.patch | 15 +++++++
 scripts/build/deps/picohttpparser.ts           |  2 +
 test/js/web/fetch/fetch-chunked-size.test.ts   | 55 ++++++++++++++++++++++++++
 3 files changed, 72 insertions(+)

gate history · 3 passed · 1 rejected · iteration 2

evidence per changed file
file                                            reads  edits  tests
patches/picohttpparser/strict-chunk-size.patch      0      1      4
scripts/build/deps/picohttpparser.ts                1      1      4
test/js/web/fetch/fetch-chunked-size.test.ts        0      1      2

…t as zero

phr_decode_chunked stopped at the first non-hex byte in the chunk-size
token and fell through to CHUNKED_IN_CHUNK_EXT, which scans to LF without
validating. A size line like "0x5\r\n" was therefore read as size 0
(last chunk) and fetch() resolved 200 with an empty body while the 5
payload bytes were still on the wire. node/llhttp reject this with
HPE_INVALID_CHUNK_SIZE.

After the hex digits, only ';' (chunk-ext), CR, or LF are valid per
RFC 9112 7.1. Anything else now returns -1, surfaced as
InvalidHTTPResponse, matching llhttp strict mode.

Applied as a patch on the vendored picohttpparser.
@robobun

robobun commented Jul 21, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: fix verified locally.

Reproduced with the standalone script in the report: system bun resolves status=200 body="" for chunk-size 0x5; with this patch it rejects with InvalidHTTPResponse, matching Node/llhttp.

bun bd test test/js/web/fetch/fetch-chunked-size.test.ts

6 malformed-token cases fail on main, pass with the patch. 5 well-formed-token cases pass on both as regression guards.

@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown
Contributor

Review Change StackReview Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: 46d73776-fd22-4ae0-b409-50ce6ae3c035

📥 Commits

Reviewing files that changed from the base of the PR and between 4ff9193 and 87f6288.

📒 Files selected for processing (3)
  • patches/picohttpparser/strict-chunk-size.patch
  • scripts/build/deps/picohttpparser.ts
  • test/js/web/fetch/fetch-chunked-size.test.ts

Included review availability: Your plan provides up to 10 included reviews per hour; 9 remain after this review.


Walkthrough

Changes

Chunked response validation

Layer / File(s) Summary
Parser delimiter validation
patches/picohttpparser/strict-chunk-size.patch
The parser rejects chunk-size tokens followed by characters other than ;, CR, or LF.
Fetch integration and coverage
scripts/build/deps/picohttpparser.ts, test/js/web/fetch/fetch-chunked-size.test.ts
The build applies the parser patch. Fetch tests cover malformed tokens, extensions, leading zeros, uppercase hexadecimal values, response bodies, and status codes.

Merge Risk: ⚪ Minimal · up to 990ab

Malformed chunk-size tokens are now rejected while valid chunk extensions remain supported. The regression tests and build pass, with no actionable merge-blocking risk identified.

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely describes the main change: rejecting malformed chunk-size tokens instead of treating them as zero.
Description check ✅ Passed The description explains the problem, cause, fix, expected behavior, affected files, and verification results. It does not use the template headings, but it provides the required information in equiva…

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Jul 21, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 1:20 PM PT - Sep 10th, 2026

❌ @cirospaciari, your commit 990ab37 has 1 failures in Build #113928 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 34918

That installs a local version of the PR into your bun-34918 executable, so you can run:

bun-34918 --bun

chunked-trailing.test.js binds its raw servers to "localhost" which
races against IPv6-first resolution in some environments. The new
chunk-size tests bind to 127.0.0.1 explicitly and live in their own
file so the gate runs only the cases that exercise this fix.
@github-actions

Copy link
Copy Markdown
Contributor

Found 1 issue this PR may fix:

  1. next start under Bun returns 200 empty bodies for server-side fetch to local PostgREST, while Node returns real data #29515 - fetch() returns 200 with empty body from PostgREST while Node returns correct data; consistent with malformed chunk-size tokens (e.g. 0x5) being silently parsed as size 0

If this is helpful, copy the block below into the PR description to auto-close this issue on merge.

Fixes #29515

🤖 Generated with Claude Code

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No bugs found. Deferring to a human because this patches vendored HTTP wire-parsing (picohttpparser phr_decode_chunked) and tightens fetch() rejection behavior — small and well-tested, but a maintainer should sign off on going RFC-strict here.

What was reviewed:

  • Patch placement against upstream picohttpparser.c at the pinned commit — the new check runs only after ≥1 hex digit and src < bufsz, so no OOB read; incremental (split-packet) decoding still works because src == bufsz exits with -2 before the new check.
  • patches: field on the Dependency type and 10+ existing users — build integration matches the established pattern.
  • Only consumer of phr_decode_chunked is the HTTP client in src/http/lib.rs; server-side chunked parsing (bun-uws/ChunkedEncoding.h) is separate and already strict.
  • net.Server has Symbol.asyncDispose so await using _s = server in the test cleans up correctly.
Extended reasoning...

Overview

Three files: a 6-line patch to vendored picohttpparser.c (phr_decode_chunked, CHUNKED_IN_CHUNK_SIZE state), a one-line patches: registration in scripts/build/deps/picohttpparser.ts, and a new parameterized test file with 6 reject + 5 accept cases. The patch adds a check that after the hex-digit run in a chunk-size token, the next byte is one of ; / CR / LF; anything else returns -1, which surfaces to JS as InvalidHTTPResponse.

Security risks

This is a strictness increase in an HTTP response parser. It closes a response-smuggling-shaped ambiguity (a token like 0x5 was read as size 0, silently truncating the body) and matches Node/llhttp strict mode. No new attack surface is introduced; the risk is compat — a server emitting non-RFC chunk-size tokens that Bun previously tolerated will now fail. Node already rejects every case in the test matrix, so any such server is already broken there. One nuance a maintainer may want to weigh: RFC 9112 §7.1.1 defines chunk-ext = *( BWS ";" ... ), and BWS is technically "MUST be accepted"; this patch (like llhttp) rejects 5 ;ext. That's a deliberate strict-mode choice, not a bug, but it's a policy call.

Level of scrutiny

Higher than the diff size suggests — it's a vendored-dep patch to the HTTP client's wire decoder, on the fetch() hot path, and it changes user-visible acceptance behavior. That said, the change is mechanically trivial (a 3-byte allowlist), RFC-grounded, Node-verified, and the test file exercises both directions.

Other factors

I verified the patch context lines against upstream source at commit 066d2b1 and confirmed: (1) buf[src] is already bounds-checked and dereferenced by decode_hex before the new check runs, (2) the incremental case (chunk-size split across reads) still returns -2 at src == bufsz before reaching the new check, so partial tokens aren't spuriously rejected, (3) phr_decode_chunked is only called from src/http/lib.rs (fetch client), so blast radius is contained. The patches mechanism is well-established (boringssl, libuv, libarchive, etc.). The test uses await using on a net.Server, which has Symbol.asyncDispose in Bun. Bug-hunting system found nothing.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No issues found. This patches the vendored picohttpparser chunked decoder (HTTP parsing hot path used by every fetch()), so a maintainer should sign off on the vendor patch and the stricter-reject compat tradeoff.

What was reviewed:

  • Verified the new check in phr_decode_chunked runs after the src == bufsz bounds guard, so buf[src] is always in-bounds; state is preserved correctly across split reads.
  • Confirmed only ;/CR/LF are accepted after the hex run per RFC 9112 §7.1, matching llhttp; chunk-ext (5;ext) still parses via the positive test cases.
  • patches: field on the Dependency follows the existing lsqpack pattern; net.Server has Symbol.asyncDispose so await using _s = server cleans up on assertion failure.
Extended reasoning...

Overview

Three files: a 6-line patch to vendored picohttpparser.c (phr_decode_chunked, CHUNKED_IN_CHUNK_SIZE state), registration of that patch in scripts/build/deps/picohttpparser.ts via the existing patches: field on the Dependency type, and a new test file with 11 it.each cases (6 malformed tokens rejected, 5 well-formed accepted). The fix: after consuming ≥1 hex digit, if the next byte is not ;, CR, or LF, return -1 instead of falling through to CHUNKED_IN_CHUNK_EXT and treating the garbage as an extension.

Security risks

This is HTTP response parsing — security-adjacent (chunked-encoding ambiguity is a known desync/smuggling vector). The change is strictly fail-closed: it rejects inputs that were previously mis-parsed with silent data loss. No new acceptance paths. I checked the read is in-bounds (the if (src == bufsz) goto Exit guard runs at the top of the same loop iteration), and traced the split-read case: if the hex run ends exactly at bufsz, the function exits with -2 and _hex_count is preserved, so the next call re-enters the loop, hits the non-hex byte, and the new check fires correctly.

Level of scrutiny

Higher than the line count suggests. This is the first patch applied to the picohttpparser vendor dep, it sits in the fetch() hot path, and it changes observable behavior (previously-resolving responses now reject). The change itself is mechanically simple and verified against the upstream source at the pinned commit, but vendored-dep patches and HTTP-parser strictness changes are the kind of thing a maintainer should explicitly approve rather than auto-merge.

Other factors

The patches: mechanism is already used by lsqpack (scripts/build/deps/lsqpack.ts:24), so the wiring is not novel. Tests use net.Server with await using for cleanup (Server.prototype[Symbol.asyncDispose] exists at src/js/node/net.ts:3318), port: 0, and cover both the new rejections and regression guards for ;-extensions and leading-zero/uppercase-hex sizes. The PR description tabulates Node vs before/after behavior. No outstanding reviewer comments.

@robobun

robobun commented Aug 16, 2026

Copy link
Copy Markdown
Collaborator Author

One more symptom of the same hole, found while fixing a node:http server bug that used to emit this wire shape. Besides the "misread as size 0" case, a garbage chunk-size line whose first byte happens to be a hex digit leaks framing bytes into the streamed body before the request fails.

Server writes (then destroys the socket):

HTTP/1.1 200 OK\r\nTransfer-Encoding: chunked\r\n\r\n1\r\nx\r\nConnection: close\r\n\r\n

Connection: close\r\n is read as chunk-size 0xC (C is a hex digit) with onnection: close taken as a chunk extension, so the following \r\n is handed to the reader as chunk data and the decoder then waits for 10 more bytes:

body chunks seen by res.body.getReader() then
bun 1.4.0 ["x\r\n"] ECONNRESET when the socket closes
node v26.3.0 ["x"] TypeError: terminated (llhttp HPE_INVALID_CHUNK_SIZE)
phr_decode_chunked with this patch "x" decoded, returns -1 both call sites in src/http/lib.rs map -1 to InvalidHTTPResponse

Checked the last row by compiling the vendored picohttpparser.c with strict-chunk-size.patch applied and feeding it those bytes directly: unpatched returns -2 with x\r\n decoded, patched returns -1 with only x decoded. So this PR covers that case too; the 5g row in the test is the same shape.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

@cirospaciari

Copy link
Copy Markdown
Member

@robobun is this fix still relevant? aka the fix already landed?

@robobun

robobun commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

@cirospaciari Still relevant. The fix has not landed on main:

  • scripts/build/deps/picohttpparser.ts on main still pins h2o/picohttpparser@066d2b1 with no patches: entry, and src/http/lib.rs still decodes chunked bodies through phr_decode_chunked.
  • The latest canary (1.4.3-canary.1+4ff919377) still prints resolved: status=200 body="" for the 0x5 repro in the PR description. Node v26.3.0 rejects it.

On matching Node 26.3 for the streaming case in the comment you linked: Node resolves fetch() with 200, the reader sees exactly ["x"], then the next read rejects. That holds whether the bad chunk-size line arrives in the same packet as 1\r\nx\r\n or in a later one. Canary today gives ["x\r\n"] (single packet) or ["x", "\r\n"] (split) and then ECONNRESET. I am verifying what this branch does end to end for both packetizations and will add that case to the test file. I will follow up here with the result.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

@cirospaciari

Copy link
Copy Markdown
Member

@robobun check CI test/js/bun/http/serve-pending-promise-abort-leak.test.ts - code 1 on :debian: 13 x64-asan

@robobun

robobun commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

@cirospaciari Checked. That failure is not from this PR. It is a pre-existing red test on the debian-13-x64-asan lane:

I have flagged the test separately so it gets looked at on its own. The other failures in builds 113925 and 113928 are the usual one-retry flakes on unrelated files (bun-install-registry, no-orphans, bun-inspector-protocol, napi/uv, shell/exec).

Still working on the Node 26.3 streaming-parity check from my previous comment; will follow up with that next.

@cirospaciari
cirospaciari merged commit e5c3e03 into main Sep 10, 2026
9 of 10 checks passed
@cirospaciari
cirospaciari deleted the farm/c6ed3847/fetch-chunked-size-strict branch September 10, 2026 20:41
@robobun

robobun commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Follow-up on Node 26.3 parity for the streaming case (1\r\nx\r\n followed by the garbage chunk-size line Connection: close\r\n\r\n). I ran the same probe against this branch (debug build of 990ab37), the latest canary of main, and Node v26.3.0. The probe reads res.body with a reader and records every chunk it sees before the read rejects.

bad line in a later packet bad line in the same packet as 1\r\nx\r\n
main (1.4.3-canary.1+4ff919377) ["x", "\r\n"], then ECONNRESET when the socket closes ["x\r\n"], then ECONNRESET
this PR ["x"], then InvalidHTTPResponse [], then InvalidHTTPResponse
Node v26.3.0 ["x"], then TypeError: terminated ["x"], then TypeError: terminated

In all three rows fetch() itself resolves with status 200. So with this PR:

  • No framing bytes leak into the body any more, and the request fails at the bad chunk-size line instead of waiting for the socket to close. That part matches Node in both packetizations.
  • One difference remains. When the last valid chunk and the bad chunk-size line arrive in the same read, Node still hands that chunk ("x") to the reader before it errors the stream. Bun drops it: both -1 arms in handle_response_body_chunked_encoding_from_{single_packet,multiple_packets} (src/http/lib.rs) return InvalidHTTPResponse without flushing the bytes that phr_decode_chunked had already decoded in that call. A non-streaming consumer (text(), json()) sees no difference, it rejects in both runtimes.

Closing that last gap means flushing the already-decoded bytes to the body stream before failing the request in those two arms. It touches the progress-then-fail ordering in on_data, so I did not want to slip it into this PR after approval. If you want exact parity there, say so and I will send it as a small follow-up with the probe above turned into a test (the split-packet case can be made deterministic by having the server send the bad line only after the client has read "x").

@cirospaciari

Copy link
Copy Markdown
Member

@robobun open a new PR to add this parity

@robobun

robobun commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

On it. I will open the follow-up PR from main and link it here.

@robobun

robobun commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator Author

Follow-up is open: #42261. It delivers the chunks ahead of a malformed chunk-size line to a streaming reader before the error, in both packetizations, and matches the Node v26.3.0 column of the table above.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants