Conversation
…hrows The native dispatch tail answers a throw by ending the uWS response of the connection. A pipelined request's response is queued, and the connection has one uWS response state, so the tail ended the response ahead of it: the client got part of that body and a close. The tail now leaves a queued response queued and marks it. When its turn comes, node:http sends it if the listener completed it. If not, the connection closes after the responses ahead of it: the socket ends as it does for a response that must close the connection, so bytes that are still buffered go out first. The turn of a destroyed queued response closes the connection the same way. destroy() there cut a large response ahead short. startPipelinedResponse() refuses a response that is not the next one in the native queue, so a dispatch that throws before node:http queues a response cannot give its turn to a later response.
|
Warning Review paused — included plan limit reachedKeep your review moving with free on-demand reviews.
On-demand reviews are free for one more day.
Reviews can continue after your included limit without a manual trigger. An admin must approve usage-based billing. Promotion and pricing detailsOn-demand reviews are free for one more day. After that, they cost $0.25 per reviewed file. Review limit detailsOr wait 5 minutes for your next included review. Limit details: You’ve used all 10 included reviews currently available. Review configuration: ⚙️ Run configurationConfiguration used: Repository: oven-sh/bun/.coderabbit.yaml Review profile: ASSERTIVE Plan: Essentials Run ID: 📒 Files selected for processing (9)
Comment |
|
Status How I reproduced it: a
The 12 new cases in Self-reviewed: 19 concerns raised, 8 needed changes (5 in the source, 3 in the tests), and all are in. The review on the PR added three findings, and all are fixed:
CI: build 118733 has no failed job (179 passed). The tests that were red in earlier builds of this PR ( PR: #43549 |
… sent without waiting for the client socket.end() at such a turn sends only a FIN when nothing is buffered, and a client that never answers it kept the connection, its queued requests and the server's pending request count. closeWhenDrained() marks the connection and runs uWS's close gate: the connection closes now when nothing is buffered, and from the drain gate when bytes of the responses ahead are still on their way. A pipelined dispatch now holds its turn in the JS queue with its native handle until its response exists. A throw before that (the constructor of a user's IncomingMessage or ServerResponse subclass) then has an entry at its turn, and the connection closes there. This replaces the order check in startPipelinedResponse(), which closed such a connection only on the next request or on the keep-alive timeout. The code that the lint rule no-duplicate-conditional-property-access rejected goes away with it.
There was a problem hiding this comment.
Beyond the inline finding, I also checked the is_queued branch of the dispatch tail in src/runtime/server/mod.rs for the IS_REQUEST_PENDING ref it no longer releases: the queued response keeps REQUEST_HAS_COMPLETED clear, so the ref is balanced by the ordinary queued-response paths (on_request_complete when JS ends it, or handle_abort_or_timeout → mark_request_as_done when closeWhenDrained closes the socket). The closeWhenDrained C++ path added in a6139fd covers both SSL variants and bails on an upgraded or closed socket.
Extended reasoning...
A confirmed finding is posted inline, so this note only records what else was examined and ruled out on this push. The Rust tail change removes the ENDED + on_request_complete balancing for the queued case; tracing it, the response stays in the same state as any pipelined response whose dispatch returned normally (no ENDED, no REQUEST_HAS_COMPLETED), so the pending ref and the server's pending-request counter are released by the same paths those responses already use. The new native closeWhenDrained is dispatched on is_ssl for both template instantiations and guards !socket || upgraded || us_socket_is_closed, matching its sibling shutdownAfterResponseDrains. The second commit addresses the two optional points raised on the prior push (socket.end() waiting on the peer FIN; the constructor-throw orphan entry) via the placeholder entry and closeWhenDrained.
drainMicrotasks() in the dispatcher also runs the tick queue, so a kick that is scheduled only at the top of a pipelined dispatch runs while the turn is still held by the native handle, and nothing starts the response afterwards. The kick is back where the response is queued. The one at the top stays for a dispatch that throws before that point. The native tail now skips its whole block for a queued response (mark_dispatch_threw_if_queued), so the existing block is unchanged. Comments are one line each.
…on failure, and check the response The last assertion restated what already resolved the awaited promise. It now checks that the stream is the one response to /second. The socket is destroyed in the finally block, so a rejection does not leave it open.
|
This PR also fixes #43512 (the request of a pipelined response destroyed while queued never emits 'aborted', 'error' or 'close'). On this branch (a88cedd) the repro from that issue prints A test for that case is on branch One gap stays: the |
…inish, lifecycle) (#43557) One pull request for the open `node:http` server pull requests. Each root cause is fixed once, and each pull request's tests are carried over. Node is the reference: every scenario was run under Node and under Bun from one script, and the outputs were compared. Fixes #4733 Fixes #18613 Fixes #40350 Fixes #43155 Fixes #43297 Fixes #43513 Fixes #43527 Fixes #25632 Fixes #31301 Fixes #43027 Fixes #43163 Fixes #43342 Fixes #43344 Fixes #43370 Fixes #43490 Fixes #43512 Fixes #43519 Each of these has a repro that is wrong on Bun 1.4.3, right on this branch, and the same as Node. | Issue | Not closed by this PR, because | | --- | --- | | #30501 (msal-node keeps Bun alive at exit) | Probably fixed. The repro copies the teardown of msal-node. The package itself was not run. | | #14430 (yarn: "does not support SSL") | Probably fixed. `response.hasOwnProperty("socket")` is now `true`. yarn itself was not run. | | #39681 (`server.setTimeout` callback after destroy) | Probably fixed. The repro is the deterministic case of #39686. The script in the issue depends on timing and on Windows. | | #43455 (`req.complete`, nine flows) | Partially addressed. Flows 1, 2, 3, 5 and 7 are fixed, and flow 6 was already right. Flow 8 (a socket timeout while the body of an Upgrade request arrives) and flow 9 (a request that `stream.pipeline()` destroyed never reports `complete`) are not. Flow 4 differs only in `_readableState.ended`. | ### What changes for users | Area | Before | After (same as Node) | | --- | --- | --- | | `req.pause()` | The socket stops at once. `req.complete` stays `false` for a small body. | The body is received until the buffer is full. Then the socket stops. | | `res.end()` before the body arrives | `req` gets `'end'` and `'close'` at once, and the body is lost | The request completes when its body really ends | | `res.destroy()` in the middle of a body | `'end'` with bytes missing | `aborted`, then `ECONNRESET` | | `socket.destroy()` inside the `'request'` listener | The body that came with the head is dropped | That body is still delivered | | `'finish'` and the `end()` callback | Fire when `end()` buffers the bytes | Fire when the last bytes have left the socket | | A response that closes the connection | The server half-closes and waits for the peer | The socket closes right behind the FIN | | `'drain'` after a later write flushed the backlog | Lost. `pipe(res)` could hang. | Emitted | | A pipelined request whose body continues after the previous response ends | Body dropped, no response, `server.close()` hangs | Delivered | | CONNECT and Upgrade tunnel sockets | Keep reading when paused or full | Stop reading. `_read()` starts them again. | | A tunnel write that waits for a drain when the client goes away | Its callback, the callbacks of the writes behind it and the `end()` callback never run | They run with an error before `'close'` | | Upgrade request with a body, paused in its listener | The body flows away | The request keeps its body | | `ws` on a reused keep-alive socket | Writes after the Upgrade could stall | Sent | | A raw `socket.write()` behind a response that still drains (the 400 for a bad pipelined request, the reply of a `'clientError'` listener) | Lands in the middle of that response | Sent after it | | `server.close()` | Could report closed while connections were open | Waits for every connection. An idle tunnel does not keep the process alive. | | `closeAllConnections()` on a listening server | Also stops the listener and destroys tunnels and WebSockets | Destroys only the HTTP connections | | `Proxy-Connection: close` (node:http only) | Ignored. The connection stays open. | Ends the connection, like `Connection: close` | | A response larger than 16 KB, also in `Bun.serve` and over TLS | Up to 4 `send()` calls for each chunk. Slower than Node in most cases. | One write for the writes of one tick. 1.1x to 2.7x the requests per second of main, and faster than Node. | | `socket.destroy()` and then `res.end()` in a listener | (this PR, earlier) `req` ended as if it were complete | `'aborted'`, then `ECONNRESET` | | `emit('connection')` or http2 `allowHTTP1`: the response ends while the listener still reads the body | The rest of the body is dropped | The body is complete | | `httpValidation: "relaxed"`, `Content-Length` or `Transfer-Encoding` in trailers | Accepted | `HPE_INVALID_CONTENT_LENGTH`, `HPE_INVALID_TRANSFER_ENCODING` | | `req.complete` inside `'connect'`, and inside `'upgrade'` without a body | `false` | `true` | | `optimizeEmptyRequests`: `socket.parser.incoming` after the response | Keeps the request alive on an idle connection | `null` | | A HEAD or OPTIONS request with `Content-Length` | The body is dropped, and `req.complete` is `true` before it comes | The request has its body | | `res.end(chunk)` after the client went away | `finished` and `writableEnded` stay `false`, no `'prefinish'` | The response ends | | An HTTP/1.0 request with an `Expect` header | `100 Continue`, `'checkContinue'`, `'checkExpectation'` or a 417 | A plain `'request'` | | The idle sweep of `close()` and `closeIdleConnections()` | Could destroy a connection that was still receiving a request, or whose response was still draining | Closes only idle connections | ### Design | Piece | What it is | | --- | --- | | Request body state | `None / Pending / Complete / Aborted / Upgraded / Detached`. Only the last chunk sets `Complete`. One function, `leave_pending`, is the only other way out of `Pending`. | | Read flow control | One path: `push()` returning false stops the socket, `_read()` starts it. Both native pause buffers are removed: no read is copied and replayed. | | "This read is parsed" signal | `notifyWhenReadParsed()` sets a uws state bit. uws delivers a `readParsed` event after the read. It replaces a `setImmediate`. | | Close during a parse | One uws bit defers a close to the end of the current message. | | Response finish | A response is finished when it has ended and the socket has fully drained. | | Idle connection | One rule, `HttpResponse::closeIfIdle()`. A connection is idle when it receives no request (head or body) and no response is in flight, queued or undrained. The sweep of `close()` and `closeIdleConnections()` both use it. | | Idle tunnel | A tunnel at read EOF with nothing left to send. uws reports it to the server through the connection filter (`-3`, `+3`, `-4`). It still counts for `'close'`, but it does not hold the event loop, like a libuv handle in that state. | | Server `'close'` | One native close promise per `listen()`. `close()` records whether its sweep left nothing open. Then a `listen()` in the same tick cannot hold `'close'` back, as in `net.Server._emitCloseIfDrained`. | | Raw socket writes | While uws holds response bytes (its buffer, the zero-copy tail of a `res.write()`, the cork buffer), a raw write goes through `AsyncSocket::write`, the path a 1xx line takes. So the order on the wire is the order of the calls. | | Upgrade verdict | One scanner and one verdict, shared by the parser and the dispatcher. | | llhttp | Updated from 9.3.0 to 9.4.2, as Node v26.5.0 vendors it, plus one local patch (see below). Node v26.5.1 and later vendor 9.4.3. That update is not in this PR. | The parser changes also tighten request framing so that it agrees with llhttp in more cases. There is no new API surface. ### A pause holds from the next read The copy of the rest of a read (`nodeHttpPausedSpill`), its replay from a posted task and the nested parse are removed. Like in Node, the rest of the read that caused a pause is still parsed, and the socket stops at the next read. usockets reads up to 512 KB in one call. libuv reads 64 KB. | One paused, unread request (client sends 64 MB) | Bytes held | | --- | --- | | Node 25.6 | 131,018 | | This pull request | 524,234 | The price is in one case. A client sends 512 KB of small pipelined requests (19,418 of them) and never reads. Each handler answers with its own 64 KB body: | Handler | Runtime | Requests dispatched | RSS | | --- | --- | --- | --- | | Answers at once | Node 26.3 | 2,425 | +177 MB | | Answers at once | main | 41 | +9 MB | | Answers at once | This PR | 2,425 | +171 MB | | Answers one tick later | Node 26.3 | 4,850 | +336 MB | | Answers one tick later | main | 2,426 | +181 MB | | Answers one tick later | This PR | 4,850 | +330 MB | Release builds on Linux x64. This PR now does what Node does. main held fewer responses, mostly for a handler that answers at once. On macOS one read can return all 512 KB. There, Bun 1.4.3 already reached +951 MB for the handler that answers one tick later, and Node reached +1,294 MB. `server.maxRequestsPerSocket` bounds it. ### Performance #### Responses larger than 16 KB are faster, and now faster than Node On main, a response that did not fit the 16 KB uWS cork buffer released the cork. After that, each piece was its own `send()`: the buffered head, the chunk-size line, the data, the `\r\n` and the last chunk. Over TLS, each 2-byte piece was also its own record. Two changes fix that, for `Bun.serve` and for node:http: | Change | Effect | | --- | --- | | A write that does not fit goes out with the cork buffer and its framing in one vectored write | No copy is added. Over TLS, the records of all the pieces share the write batch that one `SSL_write` loop already had. | | The cork buffer holds 128 KB, up from 16 KB. Only a write of 16 KB or less is copied into it, as before. | Several writes in one tick go out in one write, like in Node. A longer write still goes out without a copy. | The bytes on the wire are the same. The vectored write uses `sendmsg()` with the flags that `send()` uses. Write syscalls for one response: | Response | Node 26.3 | main | This PR | | --- | --- | --- | --- | | 4 x `res.write(16 KB)` | 1 | 16 | 1 | | 40 x `res.write(2 KB)` | 1 | 16 | 1 | | `res.end(64 KB)` | 1 | 2 | 1 | | 256 KB file, `.pipe(res)` | 4 | 16 | 5 | Throughput (req/s, the mean of 2 rounds). Node v26.3.0, main `97246d044e`, this PR `fe0ed1fbea`, with the method below: | Case | Node | main | This PR | main / Node | PR / Node | PR / main | | --- | --- | --- | --- | --- | --- | --- | | http, 4 x `res.write(16 KB)` | 17,735 | 6,983 | 19,160 | 0.39x | 1.08x | 2.74x | | https, 4 x `res.write(16 KB)` | 11,720 | 6,110 | 15,510 | 0.52x | 1.32x | 2.54x | | http, 40 x `res.write(2 KB)` | 9,879 | 6,140 | 14,980 | 0.62x | 1.52x | 2.44x | | https, 40 x `res.write(2 KB)` | 6,981 | 5,541 | 11,780 | 0.79x | 1.69x | 2.13x | | http, `res.end(64 KB)` | 18,535 | 16,528 | 19,889 | 0.89x | 1.07x | 1.20x | | http, 256 KB file `.pipe(res)` | 2,684 | 2,375 | 2,719 | 0.88x | 1.01x | 1.15x | | https, `res.end(64 KB)` | 11,894 | 14,586 | 16,426 | 1.23x | 1.38x | 1.13x | | http, GET hello (control) | 55,994 | 70,989 | 71,958 | 1.27x | 1.29x | 1.01x | main was slower than Node in six of these eight cases. This PR is faster than Node in all eight. `Bun.serve`, measured on `a771572a8d`, before the larger cork buffer (req/s, the mean of 2 rounds): | Case | main | PR | Change | | --- | --- | --- | --- | | Direct stream, 4 x 16 KB | 6,826 | 12,479 | +83% | | 64 KB string | 16,944 | 20,145 | +19% | | TLS, 64 KB string | 15,045 | 16,892 | +12% | | hello (control) | 83,957 | 83,137 | -1.0% | These runs are on loopback, where the kernel send buffer is 2.6 MB and the work of the receiver runs inside `send()`. That is the best case for fewer writes. A new connection over a real network takes about 46 KB in its first write on Linux. The rest waits in the socket buffer, as it would after separate writes. #### Small responses are unchanged A small response is already one `recvfrom` and one `sendto` on both builds. `perf` puts 66% of the time of a hello-world server in the kernel, on both builds. CI release builds on Linux x64: main `97246d044e` (the merge base) against this PR `8834cd0787`. Both use the same WebKit. The server runs on one pinned core. `oha` sends 64 connections for 5 s after a 2 s warm-up. There are 2 rounds, and the order of the builds alternates. "Change" compares the means of the two rounds. Framework servers from `bun-perf-tester` (req/s): | Server | main, round 1 | main, round 2 | PR, round 1 | PR, round 2 | Change | | --- | --- | --- | --- | --- | --- | | express | 49,803 | 51,103 | 49,968 | 50,263 | -0.7% | | fastify | 61,214 | 61,105 | 60,705 | 60,668 | -0.8% | | node:http | 70,934 | 71,405 | 73,178 | 71,053 | +1.3% | | elysia | 84,696 | 85,096 | 85,027 | 84,882 | +0.1% | | `Bun.serve` | 89,020 | 89,099 | 88,168 | 88,373 | -0.9% | node:http paths that this PR changes (req/s): | Case | main, round 1 | main, round 2 | PR, round 1 | PR, round 2 | Change | | --- | --- | --- | --- | --- | --- | | GET hello | 70,226 | 70,341 | 70,151 | 72,359 | +1.4% | | POST, 16 KB body | 47,446 | 47,789 | 48,191 | 48,785 | +1.8% | | 64 KB response in four writes | 6,885 | 6,894 | 6,834 | 6,868 | -0.6% | | Pipelined keep-alive, depth 8 | 94,063 | 93,294 | 92,909 | 93,809 | -0.3% | p99 latency (ms), the higher of the two rounds: | Server | main | PR | | --- | --- | --- | | express | 1.94 | 1.91 | | fastify | 1.52 | 1.55 | | node:http | 1.11 | 1.12 | | elysia | 0.98 | 0.96 | | `Bun.serve` | 0.82 | 0.83 | RSS (MB), one pass of 8 s of load: | Server | Build | Start | Under load | 5 s idle | 15 s idle | | --- | --- | --- | --- | --- | --- | | express | main | 39 | 94 | 61 | 57 | | express | PR | 40 | 92 | 60 | 57 | | fastify | main | 41 | 91 | 58 | 55 | | fastify | PR | 41 | 91 | 58 | 55 | | node:http | main | 20 | 64 | 43 | 40 | | node:http | PR | 20 | 65 | 45 | 41 | | elysia | main | 28 | 46 | 36 | 35 | | elysia | PR | 29 | 46 | 37 | 36 | | `Bun.serve` | main | 14 | 30 | 22 | 22 | | `Bun.serve` | PR | 14 | 30 | 22 | 22 | Every change is within 2%. fastify and `Bun.serve` hello are lower in both rounds, by about 1%. `Bun.serve` hello shows the same -1.0% in the control row above, so a small real cost there is possible. The RSS pass ran at the same time as the throughput runs, on other cores. The commits after `8834cd0787` change tests and add one version check to the node:http dispatcher. They were not measured. ### Supersedes | Theme | Pull requests | | --- | --- | | Request body | #43592 #43579 #38196 #43518 #43602 #43408 #43427 #43597 #43555 #43456 #43466 | | Tunnels | #43570 #43485. #43596 is a duplicate of #43570. | | Parser | #43182 #43161 #43326 #43327 #40505 #43363 #42532 #42194 | | Response write | #39386 #43371 #43548 #43499 #43496 #43464 #42008 #43549 | | Response finish | #40351 #43021 #41822 #43473 #42068 #35207 #43425 #43503 | | Lifecycle | #43413 #39686 #43028 #42727 #42622 #42610 #35837 #35839 #37825 #37749 #43376 #35268 | | JS API | #41691 #41738 #38036 #42462 #36527 #39718 #37964. #42947 merged on its own. | The close drain, `resetAndDestroy()`, the pending write callback handling and the response `'close'` ordering come from #42622 and #42727 by @steipete. The diagnosis and the tests for the stalled `ws` writes come from his #42610. Not included: | Pull request | Reason | | --- | --- | | #33061 | main already enforces `headersTimeout` and `requestTimeout` | | #41672 | It makes `http.createServer({ key, cert })` stop serving TLS. That needs a product decision. | | #37543 | A type refactor with no tests and no user-visible change | | #35465 | It makes `http.Server` extend `net.Server`. Only the prototype chains were joined. The `net.Server` constructor never ran, so `_handle` and `_connections` were `undefined`, and `_emitCloseIfDrained()` emitted `'close'` on a listening server. The server is backed by uWS, not `node:net`. | | The `AutoFlusher` removal in #42622 | It makes `flushHeaders()` flush at once. That is a performance change with no relation to the rest. | ### Tests | Check | Result on a debug build (macOS arm64) | Head | | --- | --- | --- | | Every test file that this PR touches (28 files) | 1,866 pass, 2 fail. The 2 failures are `serve.test.ts` "bounds memory when proxying ... to a stalled client". They fail the same way on a debug build of main. | `83af4da4a3`, run before the last commit of main came in | | `test/js/third_party/express` (9 files) and the `body-parser` test | 299 pass, 0 fail | `83af4da4a3`, run before the last commit of main came in | | Node 25.6 against Bun, 32 scenarios from two scripts (event order, framing, lifecycle) | No regression against Bun 1.4.3 | `0ff1a23f61` | | Every vendored Node `test-http-*` and `test-https-*` file, plus the `test-net-*` and `test-tls-*` files for pause, write, end and close | 535 of 537 exit 0. `test-http-agent-keepalive.js` and `test-https-timeout.js` fail on that debug build. Both pass on every CI lane. | `daee05fcfd` (before the rebase) | | The tests that depend on what the kernel takes in one send, on Windows Server 2019 x64 and Windows 11 arm64 | pass | `3eef223328` (x64), `a29289bcc1` (arm64) | | CI build 120191 (Linux, macOS and Windows, release and ASAN) | every lane passed | `daee05fcfd` (before the rebase) | Each new test fails on Bun 1.4.3, or on the commit before its fix for a fault that this branch introduced. The two tests over the limit are `node-http-connect.test.ts` ("tests should run on bun") and `node-http-syscall-fault.test.ts` ("racing a queued drain"). Each starts a debug subprocess that needs more than 5 s on this machine. Both pass on CI. ### Changes in the last push The branch is rebased on main (`daee05fcfd` was the head before). It is now linear. Four regressions against main, each with a test that fails without its fix: | Case | main | Before this push | Now (same as Node) | | --- | --- | --- | --- | | `emit('connection')` or http2 `allowHTTP1`: `res.end()` on a request that nobody reads | `'end'`, `'close'` | No events | `'end'`, `'close'` | | The same server, an unread 32 MB body | 0 bytes held | 32 MB held | 0 bytes held | | A NUL in a header value with `httpValidation: "relaxed"` (client, `HTTPParser`, `emit('connection')` server) | Accepted | The process spins forever | `HPE_INVALID_HEADER_TOKEN` | | `Connection: close`, body in the same read as the head, a 20 KB response before the body is read | `'end'` with an empty body | No events on `req` | `'end'` with the body, `'close'` | | An empty line on an idle keep-alive connection, then `server.close()` | 0 s | About 6 s | 0 s | | Fix | Where | | --- | --- | | The finish listener of a fallback connection dumps an unread request, like Node's `resOnFinish` | `http1_server_fallback.ts` | | llhttp patch: `llhttp__internal__c_test_lenient_flags_20` is false for a NUL. The relaxed state does not consume a NUL, and the next state sent it back there. 9.4.3 has the same loop. | `llhttp.c`, noted in its `README.md` | | A node:http socket that `onData` is parsing gets the close gate of `onData`, also when a large write released the cork | `HttpResponse.h` `uncorkCompletedResponse()` | | A read that starts no message leaves an idle connection idle | `HttpContext.h` `onData` | The open review threads are fixed in `8834cd0787`: `closeAllConnections()`, `Proxy-Connection: close`, five comments cut to one line, and the test of two overlapping listeners, which now waits on events. With the generation gate in `emitCloseServer` removed, that test fails in both cases. `AsyncSocketData` keeps its bools together, which takes it from 56 to 48 bytes per socket. `http.Server` no longer extends `net.Server` (see "Not included"). The special case for it in `Ipc.ts` is gone too. `child.send(msg, httpServer)` still throws `ERR_INVALID_HANDLE_TYPE`, and its test stays. <details><summary>Changes since the first revision (2988a61)</summary> Merged with main at `c8e1f6fa5b`. The one conflict was #43708 (`req.socket` emits `'end'` and `'error'`). Its state bit `HTTP_NODE_PEER_ENDED` moved to bit 22, because bit 19 is `HTTP_NODE_NOTIFY_READ_PARSED` here. Its 15 tests run in `node-http-server-abort-events.test.ts` next to the tests of this branch (103 pass). CI on `2988a610c` had ten red tests from four causes. They are fixed: - `ed882e5e99`: `write()` to a response without a body (HEAD, 204) does not wait for unsent bytes. - `811f817704`: an idle tunnel does not hold the event loop after `server.close()`. Four vendored Node tests timed out on every platform. - `0697deec11`, `d428824c08`, `7aec062d14`: the write callback tests use a body that backs up a loopback socket, and accept what Winsock does. - `c7af1c1615`: two tests from main asserted the old `close()` contract. Review findings, each reproduced against Node v26.3.0 and fixed with a test that fails without the fix: - `d4d2b783ea`, `2cc79939bb`, `45141abca9`, `9a63dbc470`, `80a23a1822`: the idle rule. A keep-alive connection is idle again when its body ends after its response. A connection that owes a queued pipelined response, that still receives a request head or body, or whose response still drains is not idle. - `87d8904aa3`, `7eca4f7806`: `close(cb)` followed by `listen()` in the same tick reports `'close'`, also for an https server whose only connection was idle. - `3fd7b25382`: a paused pipelined request behind a response that still drains stops the connection. The first revision read 512 MiB of 512 MiB into memory. - `2d1a5e2d8d`, `33e4683a8a`, `539cb41eb9`, `197c7dfc29`, `0490e3540b`: raw socket writes stay behind every unsent response byte. The cases were a CONNECT pipelined behind a response that still drains (its `200` landed at offset 2.6 MB of a 64 MiB body), a zero-length tunnel write (it hung the tunnel), the 400 replies above, the zero-copy tail of a large `res.write()`, the cork buffer, and Windows 11, where the kernel takes the whole response and refuses the next send. - `0abd39c133`: an upgrade from the request's `'end'` listener keeps the body bytes out of the WebSocket. The connection closed with 1006 right after the 101. - `5aafc61aa6`: the callback of a small `res.write()` that the kernel refuses at the uncork runs on the drain. Reproduced on Windows 11 only. - `a29289bcc1`: a tunnel write that waits for a drain settles its callbacks when the connection closes. Known differences from Node that this PR leaves: - A handler that calls `res.end()` and then `server.close()` closes its keep-alive connection at once. Node waits for the `keepAliveTimeout`. - A pipelined Upgrade behind a response that still drains is served as a plain request. - An `end()` on a tunnel with no write pending, while the response before the CONNECT still drains, closes both directions after the flush. Node half-closes. - A raw `req.socket.write(big)` and `req.socket.end()` with no `res.end()` sends every byte but no FIN. main loses bytes here. - A CONNECT socket that is given back with `server.emit('connection', socket)` answers only the first of several pipelined requests. main answers none. - A large write from an `'upgrade'` listener stalls while the body of that Upgrade request is still pending. A second `listen()` on a listening server does not throw. Both are the same on main. - A tunnel write that fails because the client went away fails its callbacks but emits no `'error'`. Node emits `ECONNRESET`. A new `'error'` could end a process that has no listener for it. </details> <!-- robobun:evidence:begin --> --- **no test proof** · iteration 8 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/web/fetch/fetch.stream.test.ts, test/js/node/url/url.test.ts, test/js/node/tls/tls-syscall-fault.test.ts, test/js/node/net/node-net-server.test.ts, test/js/node/http/node-http.test.ts, test/js/node/http/node-http-syscall-fault.test.ts, test/js/node/http/node-http-server-close-drain.test.ts, test/js/node/http/node-http-connect.test.ts, test/js/node/http/node-http-backpressure.test.ts, test/js/node/child_process/child_process_ipc_handle.test.ts, test/js/bun/http/serve.test.ts, test/js/bun/http/serve-syscall-fault.test.ts, test/js/bun/http/bun-server.test.ts <!-- robobun:evidence:end --> --------- Co-authored-by: Jarred Sumner <jarred@jarredsumner.com> Co-authored-by: autofix-ci[bot] <114827586+autofix-ci[bot]@users.noreply.github.com>
|
Confirmed on main at 5d5f03f, with a debug build of that commit.
One note on the tests: each fixture case takes 2 to 3 s under a debug build. On this machine, with a load average above 400, 4 of the 11 passed the 5 s limit when run through Nothing is left to do here. |
Problem
firstof a 10-byte body. Node v26.3.0 completes it.HttpResult::Exceptionarm insrc/runtime/server/mod.rs:1452endsnhr.raw_response. uWS keeps one response state per socket, and for a queued response it describes the response ahead. Found in the review of node:http: emit 'connect' for a CONNECT that is pipelined behind a pending response #43376.Fix
DISPATCH_THREW_WHILE_QUEUED. This covers every throw site.advanceResponsePipelinesends that response if the listener completed it, as Node does. If not,closeWhenDrained()closes the connection once the bytes of the responses ahead have left. A destroyed queued response closes the same way:socket.destroy()cut a large response ahead short.ServerResponseconstructor) then closes at its turn. Without it, a later response took that turn.test/js/node/http/node-http.test.ts, 12 new cases, 10 fail on main.Background
ServerResponsewaits insocket[kPipelinedResponses]untiladvanceResponsePipelinegives it the socket.NodeHTTPResponse(Rust) is the native handle of one response.JSNodeHTTPServerSocket(C++) holds the current one and the native queue.uncaughtExceptionhandler sees this. Without one, the process exits at the throw.Notes
Repro. One write carries
GET /firstandGET /second. The listener holds the response to/firstopen afterfirst(Content-Length 10), and throws for/second. A later tick ends/firstwith-done.first-done, connection stays openfirst, then the closefirst-done, then the close at the turn of/secondBehaviour of the response whose dispatch threw. Node keeps that response usable: a listener that ends it (before the throw, or later) gets it sent. This branch does the same until the turn of that response comes. If it is still not complete then, the connection closes after the responses ahead, and the close path aborts that response, its request, and everything queued behind it. Node assigns the socket and waits forever. Bun never waits after a throw: the tail answers a request that is not pipelined with a close at once (that path is unchanged, and #39366 changes its bytes from an empty 200 to a 500).
The close at a turn that cannot be sent (
closeWhenDrained). Three ways to close were tried.socket.destroy()(what main does for a destroyed queued response) closes at once and discards what uWS still buffers.'finish'of the response ahead fires one tick afterend(), before its bytes have left. With a 16 MB response ahead the client got 2.6 MB, on main too.socket.end()keeps those bytes, but with nothing buffered it sends only a FIN. A client that never answers it kept the connection, the queued requests and the server's pending request count, with no timer to bound it (found in review).closeWhenDrained()setsHTTP_CONNECTION_CLOSE, sends what the cork buffer holds, and runs uWS'scloseIfDoneAndMarked(). That is how uWS closes after aConnection: closeresponse: at once when nothing is buffered, otherwise from the gate inHttpContext::onWritableafter the flush. The entry that cannot be sent stays at the head of the JS queue, so nothing behind it can start, and the close path aborts it with the rest. Connections fromserver.emit('connection')keepdestroy().The turn held by the native handle. uWS counts a pipelined request and C++ appends its native response before the JS dispatcher runs. node:http queued its
ServerResponseonly after it constructed the request and the response. A user subclass ofIncomingMessageorServerResponsewhose constructor throws left a native entry with no JS entry. On main that case ends the response ahead and the connection. With only the native gate, a probe showed a response desync: the client sent/first,/second,/thirdand gotanswer-to-thirdas the second response. Now the dispatcher pushes the native handle first and replaces it with the response.advanceResponsePipelinecloses at a handle that carries the flag, and leaves any other handle in place.abortQueuedPipelinedResponsesskips handles (the native close path notifies those). An earlier version of this PR checked the order instartPipelinedResponse()instead. That closed such a connection only on the next request or on the keep-alive timeout (found in review), so it is gone.The pipeline kick. The dispatcher kicks the pipeline on the next tick when it queues a response and nothing is in flight.
drainMicrotasks()in the dispatcher also runs the tick queue, so the kick has to be scheduled after the response is queued (as on main). A second kick at the top of a pipelined dispatch covers a dispatch that throws before that point. When nothing throws,drainMicrotasks()consumes it while the handle still holds the turn, and it does nothing. A version of this PR had only the kick at the top, and the queued response was never started (found in review, now pinned by a test).Why the tail tests identity and not a flag from dispatch time.
mark_dispatch_threw_if_queued()asks the socket for its current response and compares it with this one. It runs after theuncaughtExceptionhandlers. If they advanced the pipeline so that this response is now the current one, the tail may end it, and the identity test says so. A closed or upgraded socket has no current response, and the tail keeps its old path. The existing block of the tail is unchanged, so #39366 still applies cleanly.The promise arms. A pipelined dispatch returns
undefinedfrom the JS dispatcher, soHttpResult::RejectionandPendingare not reachable for it. The gate sits in the arm that bothExceptionandRejectionshare.Not changed.
internal/http1_server_fallback(connections fed in throughserver.emit('connection')) has no native tail. A throw there behaves as in Node.mod.rs:1345). The VM is going down in that case.New tests. Eleven spawn
node-http-pipelined-throw-fixture.js, because the throw is an uncaught exception. Its client never answers a FIN, so no case can pass because the client closed the connection.'request','checkContinue','checkExpectation'listener throws: the response ahead completes. Same result as Node v26.3.0.res.end()before the throw, andres.end()after the throw but before the turn: both responses arrive. Same result as Node.ServerResponseor anIncomingMessageconstructor: the connection closes at that turn, and no later response goes out. Node ends with the same result (it answers the next request with a 400).server.close()calls back.socket._httpMessageby hand to reach that state, and Node.js v26.3.0 fails an internal assertion on that.Review. A self-review raised 19 concerns, and 8 needed changes (5 in the source, 3 in the tests). The review on this PR added three: the FIN that a client can ignore, the constructor case that waited for the keep-alive timeout, and the kick. All are in, each with a test or a probe.
Suites run with the debug (ASAN) build.
node-http.test.ts(174 pass),node-http-connect,node-http-backpressure,node-http-server-timeouts,node-http-with-ws,node-http-uaf,node-http-server-abort-events,node-http-transfer-encoding,node-http-req-socket-pause,node-http-server-socket-end-drain,node-http-pinned-write(120 pass),test/internal/oxlint-plugin-bun.test.ts, and 41 files fromtest/js/node/test/parallel(test-http-pipeline-*,test-http-many-ended-pipelines,test-http-keep-alive*,test-http-pause*,test-http-abort*,test-http-expect-*,test-http-uncaught-from-request-callback,test-http-catch-uncaughtexception,test-http-exceptions,test-http-server-capture-rejections,test-http-server-close-idle*, and others). All pass. Probes, over TCP and over TLS: 40 connections with 5 kinds of throwing listeners, 30 clients that abort while the thrown response is queued, a client that never answers the FIN withkeepAliveTimeout: 0(throw, constructor, destroyed response), and a 16 MB response ahead. Every first response completes, every request and response closes,server.close()calls back, and the process exits on its own.[human-review] gate passed · iteration 1 · 9 files touched
fails on main (without fix)
passes on PR (with fix)
diff hotspot
gate history · 2 passed · 0 rejected · iteration 1
evidence per changed file