Conversation
The native cancel hook was an empty stub, so cancel() never wrote anything: an in-flight query ran to completion and resolved with its rows, and a query cancelled before it was dispatched never settled. A backend running a query reads nothing from its connection until the query finishes, so cancel() now opens a second connection and sends a CancelRequest built from the BackendKeyData the server already handed out. The backend answers the original connection with SQLSTATE 57014, which rejects the query through the existing ErrorResponse path. A query that is queued on a connection but whose bytes are not on the wire is failed locally instead: a CancelRequest there would stop whatever the backend is actually running. A query that was never dispatched is rejected with ERR_POSTGRES_QUERY_CANCELLED.
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: oven-sh/bun/.coderabbit.yaml Review profile: ASSERTIVE Plan: Essentials Run ID: 📒 Files selected for processing (11)
Included review availability: Your plan provides up to 10 included reviews per hour; 5 remain after this review. WalkthroughThe change adds PostgreSQL query cancellation through a separate connection that sends a CancelRequest. It updates query lifecycle handling, records PostgreSQL backend key data, adds cancellation tests, and expands the SQL cancellation documentation. ChangesPostgreSQL cancellation flow
Suggested reviewers: Priority: ➖ Normal Merge Risk: ⚪ Minimal · up to PostgreSQL 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
|
Updated 7:14 PM PT - Oct 2nd, 2026
❌ @robobun, your commit 52e8500 has 2 failures in
🧪 To try this PR locally: bunx bun-pr 33370That installs a local version of the PR into your bun-33370 --bun |
|
Found 1 issue this PR may fix:
🤖 Generated with Claude Code |
|
Not adding What it does change for that issue is the workaround its author already wrote under "alternatives considered": const query = sql`SELECT pg_sleep(100)`;
request.signal.addEventListener("abort", () => query.cancel());
const result = await query.execute();On |
…running A CancelRequest names the backend process, not a statement, so it stops whatever that backend is executing: the request at the head of the connection's FIFO. do_run()'s Prepared arm writes a query's Bind/Execute as soon as can_pipeline() allows, which happens while an earlier query is still running, so cancelling the pipelined query sent a CancelRequest that killed the query ahead of it instead. Check that the request is the head of the FIFO before asking the server to stop it. A request behind the head is settled locally, but the two cases differ: a Pending request has nothing on the wire and is failed, so advance() discards it instead of writing it, while a pipelined request is already on the socket and the backend will answer it no matter what, so its promise is rejected in place and the FIFO entry stays to consume those answers in order. Marking it failed would make advance() discard it and the next BindComplete would land on the wrong request.
CI statusThe one hard failure on both builds so far is infrastructure, not this diff.
Why it is not this change:
The only other annotation on each build is a different test that CI itself tags I have not pushed a If someone with write access could re-run just the |
There was a problem hiding this comment.
Both prior concerns are addressed (the head-of-FIFO check + reject_in_place for pipelined queries, and the generalized JSDoc), and I found nothing new — deferring only because this touches native SQL/wire-protocol code and warrants a human look.
Extended reasoning...
Overview
Implements Query.cancel() for Postgres end to end: encodes the CancelRequest packet from stored BackendKeyData, opens a second plaintext connection via Bun.connect to deliver it, and settles never-dispatched queries locally. Touches 14 files across the shared JS query layer, the Postgres adapter, native Rust (PostgresSQLQuery, PostgresSQLConnection, BackendKeyData), the .classes.ts codegen input (adds a connection cached-value slot on the Postgres query wrapper), types/docs, and a new hermetic wire-server test suite.
Prior feedback
I previously flagged (🔴) that a pipelined Status::Binding query behind the FIFO head would cancel the wrong query, and (🟡) that the shared SQL.Query.cancel() JSDoc described only Postgres. Both are fixed in this revision: do_cancel() now checks is_current_request() and routes non-head on-wire requests through the new reject_in_place() (which leaves the FIFO entry so the backend's replies are still consumed in order), a pipelined-query test was added, and the JSDoc was generalized. Nothing outstanding.
Security risks
Low. The CancelRequest carries only the pid/secret the server itself issued in BackendKeyData; no credentials cross the plaintext side-connection (matches libpq's PQcancel). connectionInfo.hostname/port/path come from the user's own SQL config, not attacker input.
Level of scrutiny
High. This is production database code: it changes user-visible semantics (a former no-op now aborts server-side queries), adds a GC-visited JS value slot on a native class, and introduces reject_in_place() whose correctness depends on FIFO-ordering invariants and on Query.resolve() being a no-op on an already-rejected promise. The pipelined case in particular is subtle enough that a maintainer familiar with advance() / the request FIFO should confirm the desync reasoning.
Other factors
Test coverage is good and hermetic (scripted wire server asserting exact bytes and connection counts, plus the new pipelined variant). Not approving solely on scope/complexity — the change looks correct to me.
|
Thanks. Since the one thing you flagged for a human is the Why a pipelined request cannot just be marked
|
| Linux x64 | debian 13, ubuntu 25.04, x64-baseline, x64-asan |
| Linux aarch64 | debian 13, ubuntu 25.04 |
| musl | alpine 3.23 x64, x64-baseline, aarch64 |
| Windows | 2019 x64, 11 aarch64 |
The two red lanes are unrelated:
darwin 26 aarch64 - test-bundies inbuildkite-agent artifact downloadwith a 120s timeout beforerunTestsstarts, so it never ran a test. Reproduced identically on #68532 and #68546.windows 2019 x64-baseline - test-bunfailsnapi > napi_wrap > has the right lifetimewithCondition was not met after 100 GC attempts, a GC-timing assertion in a suite that never constructs a SQL object. The other annotation on that lane,bake/dev-and-prod.test.ts, CI already tagsflakyand retries.
CI: final state#68546 has settled at 282 passed, 2 failed. Neither failure is this diff, and I have not pushed a 1. Both builds that hit this ran on the same agent:
So this looks like one macOS agent that cannot pull artifacts, rather than a transient blip. A fresh build would most likely be scheduled straight back onto it, which is why I would rather not burn a full 287-job rebuild on it. Re-running just that job (once 2. This diff is green.
Happy to push a retrigger if you would rather roll the dice on it, but my read is that a one-job re-run is the cheaper move. |
The `connection` cached value a query keeps so that cancel() can reach the connection it ran on is a strong GC edge. `allow_gc` clears it for a request that completed, but `on_js_error` and `on_write_fail` did not, so holding a settled Query kept its whole connection reachable: socket buffers, the prepared-statement map and all. Clear it on the failure paths too, through one helper shared with `allow_gc` so the next slot added here has an obvious place to be released. `do_cancel` reads the value before it hands off to `on_js_error`, so nothing loses its connection early.
|
Pushed That slot is a strong GC edge from the query's JS wrapper to the connection's. Being straight about the evidence: I could not make this show up as an observable leak, because a The five cancel tests still pass, the full This push also starts a fresh build, which re-rolls the two unrelated red lanes from #68546 (the |
CI: three distinct failures, none from this diff#68564 ( 1. It exits before
Three for three on 2. A property get running with a pending exception, in a test that terminates a worker mid message-port transfer. I did not wave this one off:
For what it is worth, 3. This diff is greenFrom #68546, the build that ran to completion,
Locally: the five cancel tests pass, Happy to push a retrigger if you would rather roll the dice, but with |
A query that is cancelled before it is dispatched now rejects from #run() and #runAsync(), the same placement as #41492. cancel() no longer rejects the promise itself. An eager reject inside cancel() makes a cancelled query that nobody awaits an unhandled rejection, which exits the process with code 1. Carry #41492's sqlite tests so the sqlite and mysql adapters are covered in this branch too. After #41492 merges, the query.ts hunks for #run() and #runAsync() resolve as identical on rebase.
CI result for
|
This is now one pull request against
|
There was a problem hiding this comment.
Beyond the inline findings, I also checked the cancel connection's TLS handling: the ssl_mode derivation in send_cancel_request (src/sql_jsc/postgres/PostgresSQLConnection.rs:1712-1716) fails closed (a TLS session always yields Require or stricter, a plaintext one Disable), and the new ssl_config_clone hook goes through the existing deep Clone of bun_runtime::socket::SSLConfig (src/http/ssl_config.rs:457), so the session and cancel connection do not share owned C strings and there is no double free on drop.
Extended reasoning...
The change implements Query.cancel() for the Postgres client across ~1,300 lines, adding a second native connection that carries the backend cancel key, an SSLConfig clone hook across the runtime/sql_jsc FFI boundary, and a refactor of createConnection into open/ConnectionStrings. The security-sensitive surface is the cancel dial's TLS mode and certificate settings relative to the session; those were checked and found fail-closed and deep-copied. Two confirmed findings are posted inline and a third verified finding was dropped before posting, so the PR is not approvable this run; this note only records what else was examined.
Findings marked 🟡 are optional suggestions and need no follow-up push.
…e dial The cancel connection now holds the query that it stops. When it is ready to write, it looks at the query again: the packet goes out only while the query is still the head of its session and the backend has not finished it. Before, a query that ended during the TCP and TLS handshake of the cancel connection left a CancelRequest that stopped the next query of that backend. The cancel connection reads the cancel key from the session at that time and no longer keeps a copy. Tests: a cancel connection whose SSLRequest the server answers only after the query ended sends nothing, and cancel() stops a query on a real PostgreSQL server.
…adline The loop now fails after 4 s with the rows that pg_stat_activity has for the query, in place of a test timeout with no message.
|
… end of the TLS handshake The cancel connection now holds its session and its query. It sends no CancelRequest only when the session is still connected and its backend runs another query by now. A session that is closed or failed has nothing left that the packet could stop by mistake. Before, reserved.close(), which cancels its queries and closes the connection in the same tick, got no CancelRequest, and the backend ran the query to its end. For a TLS session the cancel connection now decides when the handshake has ended. Before, it decided when the handshake started, so a query that ended during the handshake still got a CancelRequest, which stopped the next query of that backend.
|
Two other open pull requests add a ConnectionFlags member at bit 6: IS_CANCEL_REQUEST in #33370 and ON_CONNECT_PENDING in #40913. Each pull request inserts its line after KEEP_ALIVE_REQUESTED, so the second one to merge gets a textual conflict there. With LIFETIME_EXCEEDED at bit 7, the result of keeping both lines has no two members on one bit.
main closes a socket that never opened with on_connect_error (#44327), so ref_and_close is main's, and the session test for that case is main's too.
…er cancels it The connection encodes the Bind of a query when it converts the parameters, and an object parameter is converted by its own code. That code can call cancel() on the query. - advance() takes a request out of `pending_requests` before it encodes the Bind, and cancel() took it out again: "pending_requests underflow". `RequestCounter::Pending` now records that a request is in that counter, so `finish_request` is idempotent for all three counters. - `encode_request` stops when the request was rejected while its parameters were encoded, and the batch leaves the write buffer. - A cancel() that arrives while `do_run` is on the stack marks the query, and `do_run` does not send it. - `PostgresSQLConnection::cancel` decides what cancel() does. The query only finds its connection. Also: - The cancel key is an `Option`. A process id of 0 is a key, and a server that sent no BackendKeyData has none. - `query.ts` is main's again. The query that is cancelled before it runs is #41492. - The docs say what cancel() does on PostgreSQL and make no claim for the other adapters.
There was a problem hiding this comment.
Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.
Beyond the inline nit, I also checked that the cancel connection's CancelTarget (the RefPtrs to the session and query) is released on every terminal path: it is taken at write time in send_startup_message, cleared in fail_with_js_value for the timeout/connect-error/close paths, and the struct's own Drop covers the rest, so the refcounts balance.
Extended reasoning...
The only finding this run is a test-cleanup nit; the native cancel-connection lifecycle in src/sql_jsc/postgres/PostgresSQLConnection.rs was read for a refcount leak on the timeout and failure paths and none was found.
|
Merge #32149 first (Downsides). Open question for a maintainer: Downsides.
Problem
Query.cancel()does nothing on PostgreSQL: a cancelledpg_sleepruns on and resolves. A comment onBun.SQL: add ability to cancel query using AbortSignal #23175 reports it.PostgresSQLQuery::do_cancel(src/sql_jsc/postgres/PostgresSQLQuery.rs) is an empty stub.Fix
PostgresSQLConnection::open, split fromcreateConnection, lets native code open a connection.PostgresSQLConnection::cancelopens a second connection to the session's peer address, encrypted when the session is. It sends aCancelRequestunless the session runs another query. The query rejects with57014.ERR_POSTGRES_QUERY_CANCELLED(reject_in_flight, sql(postgres): reject only the query whose row the client cannot decode #43187).test/js/sql/postgres-query-cancel.test.ts(31 tests, 29 fail without it), PostgreSQL 17.11 plain and TLS,test/js/sql/. Self-reviewed: 26 concerns raised, 21 addressed (Notes).Background
CancelRequestnames the backend process, so it only stops the queue head.node:netsocket: plaintext on a TLS session, no route to[::1], cancel key in JavaScript.Downsides
tx.close({timeout})leaves an unhandled57014rejection when it cancels a running query (Notes).CancelRequestthat the server handles after its query ended stops that connection's next query: 45 to 149 of 300 rounds on a loaded server, and the ROLLBACK oftx.close()in 2 of 400 (Notes). Open question: should Bun wait for the server to close the cancel connection, like libpq, pgx, pgjdbc and asyncpg?text: +7,424 bytes at 950a49c (Notes).Notes
Repro
What
cancel()does, by request statedo_cancelfinds the connection that the query was dispatched to and callsPostgresSQLConnection::cancel. The connection owns the FIFO and the request counters, so it decides. ACancelRequeststops whatever the backend process runs now, and that is the request at the head of the FIFO.Pending: no Bind, Execute or Query of the request is written.finish_requesttakes it out ofpending_requests, the count that gates pipelining at enqueue time.on_js_errormarks itFail, andadvance()discards it before it writes anything.Binding,Running,PartialResponse):send_cancel_requestopens the cancel connection. The backend answers the session with57014, and the existingErrorResponsepath rejects the query. If that error arrives while the statement is still inParsing, the statement goes back toPendingand is not failed (see below).CancelRequestwould stop the query ahead of it. Each pipelined unit has its ownFlushandSync, so the backend answers it anyway.reject_in_flightrejects the promise, setsdiscard_response, and leaves the FIFO entry in place. The connection takes the replies of that request, rows or an error, and ReadyForQuery retires it. The error has the hintThe server already received this query and still runs it. Bun discards the result., the way sql(postgres): reject only the query whose row the client cannot decode #43187 marks an undecodable row, so a caller can tell it from a query that was never sent and does not retry a write blindly.do_runis on the stack):do_cancelmarks the query, anddo_rundoes not send it (see the next section).reject_in_flightis main'son_undecodable_rowfrom #43187 under a name that fits both callers.A
cancel()from the code of a parameterThe connection turns a parameter into its wire form while it encodes the Bind. An object does that with its own code (
toString,toJSON), and that code can callcancel()on the query that it belongs to.mainis safe from this only becausecancel()does nothing there. Three changes make it safe here:RequestCounter::Pendingrecords on the request that it is inpending_requests, asNonpipelinableandPipelinedalready do for the other two counters.finish_requestreads that record, so a second call does nothing. Before,advance()took the request out of the count ahead of the encode andcancel()took it out again.encode_request, the one caller of the batch writers, fails the batch when the request was rejected while its parameters were converted.atomicallythen takes the batch out of the write buffer, so nothing of the request reaches the wire.cancel()that arrives whiledo_runis on the stack finds no connection on the query.do_cancelmarks the query, anddo_runthrowsERR_POSTGRES_QUERY_CANCELLEDand enqueues nothing.Against PostgreSQL 17.11, debug builds, with a parameter whose
toString()callscancel()on its query orclose()on its reserved connection:cancel(), first run of the statement (advance()encodes the Bind)panic: pending_requests underflowERR_POSTGRES_QUERY_CANCELLED, and the connection runs the next querycancel(), prepared statement (do_runencodes the Bind)ERR_POSTGRES_QUERY_CANCELLEDreserved.close(), first runpanic: pending_requests underflowreserved.close(), prepared statementThe last row is also a hang on the released build 367d939 (killed after 60 s):
do_runenqueues the query on the connection that its own parameter closed. Herereserved.close()cancels the query first, sodo_runstops. A parameter that closes the connection in another way than throughclose()is not covered by this PR.The cancel connection
PostgresSQLConnection::send_cancel_requestruns on the session. It builds aConnectParamsand callsPostgresSQLConnection::open(see the split below), so the cancel connection is an ordinary native connection with the flagIS_CANCEL_REQUEST. It differs from a session in five places, each behind that flag:send_startup_messagewrites theCancelRequestin place of the StartupMessage. First it asks itsCancelTarget, which holds the session and the query, whether the packet can still stop the right query (see the lateCancelRequestbelow). It reads the cancel key from the session at that time.setup_tlsdoes not callstart()for it.on_handshakedoes. So the question above is asked when the connection can write, not when the handshake starts.startdoes not restart the connection timer, so there is one budget from dial to hang-up. The budget is the connection timeout of the session, and at most 5 s.on_dataafter the write fails the connection. A server answers aCancelRequestby hanging up, so a server that answers must not turn this into a session.postgres_connectionsuse.It always ends through
fail. Soref_and_closecloses it by main's rule for a failed connection (#44327): it sends its close_notify and FIN, and does not wait for the peer. An earlier revision of this PR had a close rule of its own for it.The cancel key is an
Option: a server that sent no BackendKeyData has none, and thencancel()on a running query does nothing. A process id of 0 is a key like any other.What it takes from the session:
getpeername), or its unix socket path. A host name can resolve to another server on a second lookup. libpq'sPQcancelalso dials the stored address.tls_config.server_name), not against the dialed address. With a certificate that names onlylocalhost, a session tolocalhostconnected through::1, and its cancel connection completed TLS with SNIlocalhostand delivered theCancelRequest. A session addressed as127.0.0.1was refused with that certificate.prefer. A session without TLS makes one withdisable. TheSSL_CTXis shared (SSL_CTX_up_ref) and the TLS options are copied, so SNI and certificate verification are the session's.The cancel key never becomes a JavaScript value.
handle.cancel()returns nothing, and the adapter has no cancel method. Measured withnet.connect,net.createConnection,tls.connect,Bun.connectandnet.Socket.prototype.{connect,write,end}patched beforenew SQL(): 0 patched calls duringcancel(), and the query rejects with57014.An earlier revision sent the packet from the adapter on a
node:netsocket. On that revisionpostgres://postgres@[::1]:5432connected butcancel()did nothing (select pg_sleep(3)resolved after 3019 ms), because URL parsing keeps the brackets and only the native connect strips them. On a TLS session it sent the 16 bytes in the clear.The split of
createConnection(the former #43942)This PR was a stack of two. #43942 held the split alone, and its review is there. Its commits are in this branch unchanged (819cd9c).
call, the host function ofcreateConnection, parses its JavaScript arguments and callsPostgresSQLConnection::open(global, group, ConnectParams).openallocates the connection, dials it and wraps it for JS. It returns the dial error of uws, andcallthrows it as the same exception as before.ConnectionStrings::newbuilds the one buffer that holds user, password, database, options and path, and the five slices into it. It is the only way to make them, so a second caller cannot pass slices that point elsewhere.Three statements change for
createConnection:calllooks up the socket group beforeopenallocates the connection. It did so after the allocation. No JavaScript runs between the two.callchecks user, password, database and path for null bytes before it copies them into the buffer. It did so after the copy and then freed the buffer. The error is the same.ConnectionStrings::newwrites the terminator after each string (count_z,append_z).mainreserves that byte and does not write it, so its buffer ends with 5 bytes that were never written. The allocation has the same size. The slices do not include the terminator, so the startup message has the same bytes.Considered a second constructor for the cancel connection. It copies the 50-line struct initialiser, and the two copies drift.
Time bound
The cancel connection has one timer from its dial to its end. The budget is the connection timeout of the session, and at most 5 s.
CancelRequestand then neither answers nor hangs up: withconnectionTimeout: 1the client closes the cancel connection after 1.1 s (testthe cancel connection ends when the server stays silent after the CancelRequest, debug build).cancel(): withconnectionTimeout: 1the process exits 0.97 s after the cancel (2 runs, debug build). Since usockets: tell a socket's holder it is gone exactly once, whoever closes it and whether or not it ever opened #44327 uSockets reports the close of a socket that never opened, so this case needs no code in this PR.A statement whose Parse is cancelled
The first run of a statement with no parameters sends Parse, Bind, Execute and Sync together, and the request is the running one. A backend that waits for a lock waits inside Parse. A cancel then gives an
ErrorResponsewith no ParseComplete before it.mainfails the statement on every error inParsing, and each query that shares the statement rejects with that error.An error with SQLSTATE
57014inParsingnow sets the statement back toPending. The next query that shares it sends the Parse again. Against PostgreSQL 17.11, with another session holdingACCESS EXCLUSIVEon the table, twoselect v from ton one connection andcancel()on the first:570145701457014The rule reads the SQLSTATE only, so a
statement_timeoutinside Parse also stops failing the queries that share the statement.Known limit: a
CancelRequestthat arrives before its queryThe server ignores a cancel while it still reads the request. Bun sends the
CancelRequestat once, also when bytes of the request are still in its write buffer. PostgreSQL 17.11 on loopback, release build,select pg_sleep(1), length($2)on an idle connection with a warm statement,execute()andcancel()in the same tick, 5 runs each:5701457014cancel()300 ms later57014"Lost" means that the query ran its full second and resolved. The 1 MB and 2 MB rows are a race, so their split is not a stable number. A simple query and a prepared query with a small parameter gave 20 of 20 rejected, on the debug and on the release build.
This PR does not fix it. A wait for the write buffer to drain is wrong: the buffer also holds the bytes of pipelined queries behind the head, and the server does not read those while it runs the head. A fix has to record when the bytes of one request have left the buffer, at every place that writes a request.
A late
CancelRequest, and the open questionA
CancelRequestnames the backend process. The server acts on it after the postmaster has accepted the cancel connection and started a process for it. If the query has ended by then, the signal stops the query that the backend runs at that moment. A signal that meets an idle backend does nothing. Bun cannot take the packet back after the write.The numbers in this section are from release builds of 950a49c and earlier revisions, as named. The code of this path has not changed since.
What this PR does. The cancel connection holds the session and the query (
CancelTarget). When it can write, which on a TLS connection is the end of the handshake, it asks the session. It writes the packet unless the session is still connected and its backend no longer runs that query. In that case it hangs up and sends nothing. A session that closed gets the packet (see the next section).What this PR does not do. The session writes its next request as soon as the backend is ready, also while a cancel connection is open. The open question at the end of this section is about that.
The size of the window on this machine. PostgreSQL 17.11 on loopback, release builds, 16 cores, load average 450 to 550 from other work. A raw protocol client on
node:net, with no use ofBun.SQL, 200 cancels of a runningpg_sleep(5), 2 runs:ErrorResponseon the sessionpg_cancel_backend()on a second session, until theErrorResponsecancel()of this PR in the same measurement: a median of 46 to 95 ms in 5 runs, and 1.7 ms in the best case. So the time is the server's. It goes into the new connection and not into the signal. A server without this load has a smaller window.How often the wrong query is stopped. A loop of 300 rounds per run on a pool of 1:
A = select 1,A.cancel()in the same tick,await A, thenB = select pg_sleep(0.01). The number is how often B, which nobody cancelled, rejected with57014:maine70cca7, wherecancel()does nothingWith B already queued behind A (
select pg_sleep(0.03), simple protocol), 300 rounds per run:select 1,cancel()in the same tickpg_sleep(0.002),cancel()in the same tickpg_sleep(0.005),cancel()3 ms laterWhat the numbers say:
ssl = on, a TCP hop that holds the TLS ClientHello of the cancel connection for 300 ms,A = pg_sleep(0.15),B = pg_sleep(1)behind it,cancel()on A after 30 ms. On c55aa7a, which asked the session before the handshake, A resolved and B rejected with57014. On 950a49c both resolve. WithA = pg_sleep(3), A rejects with57014and B resolves, on both builds.ERR_POSTGRES_CONNECTION_TIMEOUTafter 5 s, behind the cancel connections of the run before. 40 connects with no cancel before them took 7 to 21 ms.tx.close(). It callscancel()on the queries of the transaction and then sends ROLLBACK on the same connection, with nothing between the two. 200 rounds each ofsql.begin(async tx => { tx.unsafe("select pg_sleep(...)").execute(); await tx.close(); })with a query of 20 ms and of 5 ms, debug build of a46f129, PostgreSQL 17.11: in 1 round of each 200 theCancelRequeststopped the ROLLBACK, andtx.close()rejected with57014.sql.beginthen sent COMMIT, the server answered it with the ROLLBACK tag because the transaction was aborted, and the next query on the connection ran in all 400 rounds. On the released build 367d939, wherecancel()does nothing,tx.close()resolved in 100 of 100 rounds.What other clients do, as read in their sources:
PQcancel(fe-cancel.c), pgxPgConn.CancelRequest(pgconn/pgconn.go) and pgjdbcQueryExecutorBase.sendQueryCancelwrite the packet and then read until the server closes the connection. Only then does the caller continue. The comment in libpq: "Without this delay, we might issue another command only to find that our cancel zaps that command instead of the one we thought we were canceling."PoolConnectionHolder.releaseinasyncpg/pool.pywaits for it), and it ends a connection whoseCancelRequestfailed.Client.canceland postgres.jscancelwrite the packet and do not hold the session. postgres.js returns a promise that settles when the cancel connection closes. Its README says that the race can cancel another query, and that this is fine for long queries.Bun.SQL: add ability to cancel query using AbortSignal #23175 the maintainer ofBun.SQLnames this hazard ("other queries that are also executing in the same connection could also be canceled"), points to the postgres.js way, and says that he wants to look for a better solution.The open question. Should a session write no new request while one of its cancel connections is open?
tx.close(). A session has at most 1 cancel connection open, so cancels cannot queue up at the server.cancel()of a running query delays the next request of that connection until the server has closed the cancel connection. On this machine that was a median of 5 to 92 ms. The limit is the budget of the cancel connection: the connection timeout of the session, and at most 5 s. A request that is already written behind the head cannot be held back and stays exposed.advance()and in the writes at enqueue time, and a release on every end of the cancel connection.A session that closes while its cancel connection dials
reserved.close()callscancel()on each query of the reserved connection and then closes the connection. The backend still runs the query, so theCancelRequesthas to go out after the session is gone. TheCancelTargetholds the session itself. A session that is not connected any more always gets the packet: no other query of this client can follow on it.A reserved connection runs
pg_sleep(6), thenreserved.close(). Time untilpg_stat_activityshows that the backend stopped the query, release builds, 2 runs each:maine70cca7c55aa7a reached the session through the link from the query to its connection. The close clears that link, so it found no session and sent nothing. 100 rounds of reserve, query and close on the debug build of 950a49c: 0 backends left in
pg_sleep, connections 3 to 3, open file descriptors 11 to 11.Known limit: cancel while only the Parse is on the wire
The first run of a statement with parameters writes Parse, Describe and Sync and waits for the parameter types before it writes Bind and Execute. The request is
Pendingin that window.cancel()rejects it locally and sends noCancelRequest: Parse normally ends within one round trip, so a cancel would mostly arrive late and stop the next query on that backend. A backend that is blocked inside Parse (a lock wait) stays blocked until the lock clears, as onmain. The connection serves the next request as soon as the backend answers.Known limit: a link-local IPv6 peer
us_socket_remote_addressreturns the peer address without its scope. For a session tofe80::1%eth0the cancel connection dialsfe80::1, which does not reach the peer, andcancel()does nothing, as onmain. Not tested: the machine has no link-local address.A query that is cancelled before it runs
Not in this PR. #41492 makes that change in
query.tsfor every adapter, with its tests. Until it merges, a query that is cancelled before it was started never settles, as onmain. An earlier revision of this PR carried the samequery.tshunk.query.tsis main's again.close()andclose({ timeout })with a running queryclose({ timeout })insrc/js/bun/sql.tswaits on the queries in its scope withPromise.all(...).finally(...). A query that rejects during the wait leaves a rejected promise with no handler. Onmainthat needs a query that fails by itself. With this PR thecancel()that the timeout calls makes a running query reject. #32149 handles that rejection for a transaction.Against PostgreSQL 17.11,
max: 1, a runningpg_sleep(3)whose own promise has a handler. The first column is the released build 367d939, the second is the debug build of a46f129:close({ timeout: 0.3 })ERR_POSTGRES_CONNECTION_CLOSED, 1 unhandled rejectionclose()ERR_POSTGRES_CONNECTION_CLOSED, 0 unhandledclose({ timeout: 0.3 })cancel()does nothing,closewaits 3 s for the query, 0 unhandled57014, 1 unhandled rejectionclose()sql.beginrejects withERR_POSTGRES_CONNECTION_CLOSED57014,sql.beginrejects the same wayIn all 8 runs the pool ran its next query. The third row is the reason for the first line of this description. On an earlier revision of this branch, with the
sql.tsdiff of #32149 applied, that case had no unhandled rejection and rolled back.tx.close()cancels every query in the scope of the transaction, and the SAVEPOINT, RELEASE and ROLLBACK TO statements are in that scope too (run_internal_transaction_sql). One that waits behind a running query is rejected and never sent. One that the backend runs gets aCancelRequest. Onmainthey ran to their end.GC edge
do_cancelneeds the connection, so the query wrapper caches the connection it was dispatched to. That slot is a strong edge.allow_gc,rejectandon_write_failclear it, so a settledQuerythat user code keeps does not pin the connection. The testa settled query does not keep its connection aliveholds 40 settled queries of 10 pools. Without the clear inrejectit finds 11 connections alive after GC.A cancel connection holds a reference to its session and to the query (
CancelTarget) fromcancel()until it writes the packet or fails.send_startup_messageandfail_with_js_valuedrop it. The longest hold is the budget of the cancel connection.Merges with main
BackendKeyDataand removed the field onPostgresSQLConnection. This PR restoresprocess_id,secret_keyand the field, becausecancel_request()needs them.wire-frames.ts: main has the Parse, Bind and ParameterDescription helpers. This PR adds onlypgCancelRequestandpgBackendKeyData.ref_and_close, and had a session test for it. Both are gone: the merge takes main'sref_and_close, and main has the test.Open PRs that touch the same code
Checked with
git merge-treeagainst a46f129, which containsmainfaac63e. Each conflict is for the PR that lands second.PostgresSQLConnection.rsand 3 inPostgresSQLQuery.rs. Both PRs giveon_undecodable_rowthe namereject_in_flight.PostgresSQLConnection.rs, next to code that this PR adds.PostgresSQLConnection.rs. It also conflicts withmainalone, intls-sql.test.ts.createConnection, which now goes throughConnectParams. Both PRs take flag bit1 << 6inConnectionFlags. Bit 7 is free. It also conflicts withmainalone.mainalone, in the close code that usockets: tell a socket's holder it is gone exactly once, whoever closes it and whether or not it ever opened #44327 rewrote.unsafewith typed ownership #40139 (typed ownership for socket, timer, sql, redis): 10 hunks inPostgresSQLConnection.rs, and hunks inPostgresSQLQuery.rs,jsc.rsandhw_exports.rs. It also conflicts withmainalone.Tests
The server is a scripted wire server. The tests assert the exact
CancelRequestbytes, and assert that none are sent where none may be sent. That needs a server that answers on demand. The last test is the exception: it runs against a real PostgreSQL server (describeWithContainer), which is the judge of the bytes.In
postgres-query-cancel.test.ts, 31 tests:cancel()on an extended and on a simple query sendsInt32(16) Int32(80877102) Int32(pid) Int32(secret)on a second connection and rejects with57014(2). It reaches a server at an IPv6 literal and on a unix socket (2).CancelRequestarrives inside TLS (4). A cancel connection that is toldNsends nothing, under prefer and under require (2). One that meets a certificate the session's CA did not sign sends nothing (1). The certificate is checked against the host name of the session (1).cancel(): the query rejects and no Bind of it reaches the server, for a first run, for a prepared statement, and withprepare: false(2).cancel()stops a runningpg_sleep, seen inpg_stat_activity, and the connection runs the next query (1).On the released build 367d939, 29 of the 31 tests fail. The other 2 pass there because that build opens no cancel connection: the one for a server with no cancel key, and the one for the dial that never completes. The scenario of the first test for the code of a parameter stops the debug build of 950a49c with
panic: pending_requests underflow. With this PR all 31 pass on the debug build (3570e33).End to end, PostgreSQL 17.11
Debug build of a46f129. The head cancel is
pg_sleep(5), cancelled after 300 ms. The times are the server's under a load average of 800:Release build of 950a49c, the same server with
ssl = onand a certificate forlocalhost.ssland the version are whatpg_stat_sslhas for the session:Cost
For a query that is never cancelled: one cached slot on the query wrapper (8 bytes), written at dispatch and cleared when the query settles. There is no new allocation, syscall, promise or host function per query.
encode_requestreads the status of the request once per Bind.size_of::<PostgresSQLQuery>()is 64 bytes. This PR adds no field to it.textmaine70cca7, 80,676,296 on 950a49c (+7,424). Not measured for the commits after it: a release build did not finish on the machine, which had a load average above 1000dataandbsssize_of::<PostgresSQLConnection>()main, 704 bytes here: 12 for the cancel key, 16 for theCancelTargetof a cancel connection, and paddingSqlRuntimeHooks)ssl_config_clone)cancel()of a running queryPostgresSQLConnection, and on a TLS session a copy of the TLS optionsPostgresSQLConnectionobjects 2 to 2,PostgresSQLQuery2 to 2, open file descriptors 10 to 10net/tls/Bun.connectcalls during onecancel()Not measured: instructions per query and syscalls per cancel.
valgrind,perfandstraceare not on the machine and cannot be installed there.MySQL
MySQL
cancel()is unchanged: it sends nothing and rejects nothing (JSMySQLQuery::do_cancelis a stub). The docs and the JSDoc ofcancel()describe PostgreSQL only, and make no claim for the other adapters.Other suites
All of
test/js/sql/on a46f129 with the debug build and a 30 s test timeout: 957 pass, 2 skip, 47 fail. 46 failures are MySQL tests that getAccess denied for user 'root'@'localhost'from the local MariaDB. The other one isjson/jsonb bind parameter does not leak the stringified payload, which timed out at 30 s under a load average of 800. Alone it passes in 10 s, 2 of 2 runs.rustfmtandprettierare clean. The two commits after 3570e33 change two comments and the cleanup of one test. They ran in CI only: the machine lost its build.Self-review
A review of 950a49c raised 26 concerns. 21 are addressed in a46f129 and the commits after it:
cancel()from the code of a parameter, and the lost cancel indo_run. The decision moved from the query toPostgresSQLConnection::cancel. The cancel key is anOption. The close of the cancel connection is main's. The release for a dial that never opens is main's (usockets: tell a socket's holder it is gone exactly once, whoever closes it and whether or not it ever opened #44327).query.tschange (it is sql: reject a query that is cancelled before it runs #41492) and the session test for the dial that never opens (main has it).prepare: false, a cancelled head with queries pipelined behind it, an error for a cancelled pipelined query, a silent server after the packet, a server with no cancel key, the GC edge.cancel()text claims nothing for MySQL any more. The survey for the open question has asyncpg and the comment onBun.SQL: add ability to cancel query using AbortSignal #23175.Not changed, with the reason:
status. It would rewrite the state machine ofadvance()onmain. The decision is now onematchin the connection, and the three counters are recorded on the request.tx.close()cancels control statements and sends ROLLBACK right behind a cancel. Measured above. The fix is the hold of the open question.cancel()rejects gives its pool slot back while the server still runs it, so the pool can hand that connection to the next caller, who then waits.mainhas the same for an undecodable row (sql(postgres): reject only the query whose row the client cannot decode #43187).fail, also after a delivery, and builds an error that nobody reads. It is one object per cancel.Alternatives that were built or weighed
node:netsocket and change nothing.cancel()stays a no-op for[::1]and for servers that only accept TLS.node:tlsto the JS socket. The adapter then has a second copy of the sslmode rules that the native connection implements.createConnectionas an extra argument. This was built and measured first (11 tests passed). The cancel key is a JS value on the way, the host name is resolved a second time, and a TLS socket stayed open when the peer went silent.no test proof · iteration 7 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/sql/postgres-query-cancel.test.ts