Conversation
…cted before onclose runs
WalkthroughChangesValkey lifecycle
Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 4✅ Passed checks (4 passed)
Comment |
|
This PR may be a duplicate of:
🤖 Generated with Claude Code |
There was a problem hiding this comment.
Beyond the two test nits: I traced the new fail() → close() → SocketHandler::on_close → ValkeyClient::on_close() → fail() re-entrancy on a Connected client — the second fail() early-returns on the failed guard and the socket is already detached, so no double-close. Also checked that on_valkey_close() still adopts exactly the one socket keep-alive ref that connect() forgot, so the refcount stays balanced on the newly-unconditional close path.
Extended reasoning...
The runtime change is small but sits in the valkey client's lifecycle/re-entrancy code, and it makes fail_with_js_value() unconditionally close the socket where before it only did so pre-handshake. I traced the new re-entrant path (idle timeout on a Connected client → fail → close → usockets on_close → ValkeyClient::on_close → fail again → on_valkey_close) and confirmed the if self.flags.failed { return } guard prevents a loop, the socket is detached before the inner close() so it early-returns, and on_valkey_close()'s ScopedRef::adopt still pairs with connect()'s forgotten socket_ref. Also verified that moving status = Disconnected out of the defer means a connect() called from onclose now reaches the Disconnected match arm and reconnects, and the remaining defer only runs update_poll_ref() so it can no longer stomp the new Connecting status. Given the subtlety of these paths and the noted user-visible change (client.connected is now false inside onclose), a human look is still warranted; the inline findings are test-quality nits only.
|
@robobun adopt |
|
Updated 10:47 PM PT - Aug 17th, 2026
✅ @alii, your commit 3bf03fad1e8ae84f98a9d050065c23436d117275 passed in 🧪 To try this PR locally: bunx bun-pr 37993That installs a local version of the PR into your bun-37993 --bun |
… onclose fail_handshake closed client.socket a second time after fail_with_js_value() returned. fail_with_js_value() already closes the socket, so the only socket left for that close was the one a connect() issued from onclose had just opened. It also set is_manually_closed, which fail_with_js_value() sets too. on_data kept handling the remaining replies of a read after the reply it had just handled failed the connection, as long as status and failed looked healthy, which they do again once connect() has been called from onclose. The leftovers were then handled as the new connection's HELLO reply. Stop as soon as client.socket is no longer the socket the data came from. on_close() now frees read_buffer and resets reply_scanner like it already did for write_buffer. A reply that was only partially received when the connection went away counted as pending activity in update_poll_ref(), so close() after such a reply kept the process alive. Tests: cover the three cases above, trigger the post-handshake fail() with a protocol error instead of a 100ms connection timeout (debug builds need longer than that to connect), issue the second connect() from onclose so the test hangs without the fix regardless of microtask timing, await the rejects matchers, and assert on stderr of the spawned processes.
|
@alii pushed 21be6f4 on top of your branch. Besides the two review findings (the fail_handshake scopeguard and on_data carrying on with leftover replies after a connect() from onclose), on_close() now also frees read_buffer: a half-received reply kept update_poll_ref() holding the loop after close(), so the process never exited. Two of your tests changed shape and you may want to look at those in particular: the idle timeout test failed on a debug build because connecting takes longer than the 100ms connectionTimeout it relied on, so the post-handshake fail() is now triggered with an unparseable reply instead; and the connect()-twice test now issues the second connect() from onclose, because the awaited version of the original passed against main (whether the rejection continuation runs before or after the status defer depends on where the microtask drain happens), while the onclose form hangs against main every time. Details in the description. |
|
Nothing further from my side: 2689715 only tightens the three first-connect() assertions to check the error code, as the last review asked. The previous build, which already had all the src changes, finished with no failures (the remaining red entries were retried flakes in unrelated files), so once the current build finishes this is ready for a look, @alii. |
There was a problem hiding this comment.
I re-reviewed after 2689715 and the remaining nit (bare .rejects.toThrow() at the three first-connect() sites) is addressed — no further findings. Given this is a re-entrancy-sensitive state-machine change with a user-visible behavior change (client.connected is now false inside onclose) and there are four open PRs touching the same lines (#33479, #33473, #32779, #33104), a human sign-off would still be worthwhile.
What was reviewed:
- The new
on_dataguard (*self.socket.socket() != socket) is a superset of the oldstatus/failedcheck —InternalSocket'sPartialEqis pointer-identity, and every path that set the old flags now also swapsself.socket. fail_handshake's dropped scopeguard andis_manually_closedpre-set are indeed redundant with the now-unconditionalclose()infail_with_js_value; theThrownfallback infail_handshake_with_verify_errorstill closes explicitly.on_close()clearingread_buffer/reply_scanner: checked thaton_open()already resets both, so an auto-reconnect (which skipson_close()'s fail/close branches) does not carry stale scanner state either.- The six new tests: each awaits an observable condition, drains both subprocess pipes, and asserts a specific error code; the
helloServerstub keys replies on chunk contents, which is fine for the fixed HELLO/SELECT/PING sequence used.
Extended reasoning...
Overview
The PR fixes three ways a RedisClient could get stuck after a failure in src/runtime/valkey_jsc/{valkey,js_valkey}.rs: fail() left an established socket open, fail() never cleared is_reconnecting (holding the event loop), and SocketHandler::on_close/on_connect_error set status = Disconnected in a scopeguard after the user's onclose ran, so a connect() from onclose saw Connecting and hung. Follow-up commits (in response to two 🔴 findings from my earlier pass) removed the now-harmful fail_handshake close scopeguard, replaced the on_data continue-guard with a socket-identity comparison so leftover RESP replies from a closed connection cannot be routed to the next one, and cleared read_buffer/reply_scanner in on_close(). Six new tests in connection-failures.test.ts exercise each path against a local net.createServer stub or a closed port.
Security risks
None identified. The change is client-side connection lifecycle bookkeeping; no parsing of untrusted data changed shape, and the TLS-handshake path still fails closed via fail_with_js_value → close() (and the explicit close in the Thrown branch of fail_handshake_with_verify_error is unchanged).
Level of scrutiny
High. This is native code where user JS (onclose, promise rejection handlers) is invoked mid-teardown and can synchronously re-enter connect(), mutating status/socket/flags before the caller unwinds — exactly the class REVIEW.md warns about ("anything that can run user JS can synchronously free your state"). My earlier review found two real bugs the initial version introduced in this class (the fail_handshake scopeguard closing the new socket, and on_data feeding stale replies into the new connection), which were then fixed. That the fixes were needed is evidence this area is subtle enough to warrant a human look.
Other factors
- Visible behavior change:
client.connectednow readsfalseinsideonclosewhen the server drops an established connection (previouslytrue). This is arguably the correct semantics and is called out in the description, but it is a user-observable API change. - Duplicate PRs: the find-duplicate-prs bot lists four open PRs (#33479, #33473, #32779, #33104) each making one of the same edits; someone should decide which lands and close the rest.
- Prior review cycle: all four earlier inline findings (two 🔴, two 🟡) and the follow-up 🟡 on bare
.toThrow()have been addressed in 21be6f4 and 2689715 respectively; I verified each fix in the current diff. - Tests: comprehensive, hermetic (local
netserver / closed port, no container), verified fail-on-main / pass-on-PR per the description's evidence block. I checked the socket-identity guard againstInternalSocket'sPartialEq(pointer identity per variant) and confirmed every path that used to trip the oldstatus/failedguard now also changesself.socket, so the new guard is a strict superset.
|
Merged main into the branch (5fb1742), no conflicts; the valkey tests still pass on the merged tree. On the four overlapping PRs the duplicate finder listed, for whoever lands this: #33479 makes exactly the fail() change this PR makes and nothing else, so it can be closed once this is in. The other three keep a fix of their own that this PR does not include and will need a rebase on top of it: #33473 arms idleTimeout on socket traffic (today the idle timeout only fires through the timer armed with connectionTimeout, which is also why the idle test here was reshaped), #32779 handles a reconnect whose connect() fails synchronously (that path still only calls onclose and never settles the promise), and #33104 releases the event loop after close() while subscribed. The is_reconnecting and read_buffer parts of those two are covered here now. |
|
One more bot finding since the last push, answered inline: reconnect() still handles a connect() that fails synchronously (for example a unix socket path that is gone when the retry fires) by only calling onclose, so the client ends up stuck on that path too. That branch is untouched by this PR and #32779 already fixes it properly (it retries like an asynchronous failure), so I left it there rather than widening this PR; say the word if you would rather have it in here. No code changes from this round. |
|
I closed 32779 |
|
@robobun get this mergeable |
…ect timer of a finished attempt reconnect() answered a connect() that failed before a socket existed (unix socket path gone, out of fds) by calling the user's onclose and nothing else: the connect() promise stayed pending, queued commands were never rejected and is_reconnecting stayed set. It now runs ValkeyClient::on_close() like a dial that fails asynchronously does, taking the ref on_close() releases in place of the socket's, so it retries with backoff or fails and settles the promise. JSValkeyClient::fail_with_js_value loses its only caller. The connection timer armed for an attempt stayed armed after that attempt failed. With connectionTimeout shorter than the retry delay it fired during the delay, and since fail() now clears is_reconnecting, the scheduled retry then did nothing and connect() never settled while the poll ref kept the process alive. on_valkey_reconnect() disarms it; reconnect() arms a new one for the next attempt.
|
@alii two pushes since your comment. dfcee7c merges main for a fresh build: the previous one was red only because the two darwin 14 aarch64 test jobs expired without an agent and test-cluster-shared-leak.js timed out on Windows 2019, and the macOS queue is moving again. 682b042 closes the two holes that were left once #32779 went away, both failures that happen while no socket exists, so the close path the rest of this PR relies on was never reached: a reconnect whose dial fails outright (unix socket path gone) now goes through on_close() like a refused one instead of only calling onclose, and the connect timer of a finished attempt is disarmed, because with connectionTimeout shorter than the retry delay it fired during the delay and, now that fail() clears is_reconnecting, left connect() pending and the process alive (the #18895 shape; it was reproducible on this branch before that commit). Three tests each, all hanging without their commit; write-up in the description. Review threads are all resolved, so once this build is green it should be good to go. |
There was a problem hiding this comment.
682b042 addresses both findings from the previous round: on_valkey_reconnect() now disarms self.timer (with the matching "connection timeout shorter than the retry delay" test), and reconnect()'s sync-failure branch goes through the new on_close_without_socket() → ValkeyClient::on_close(), so it retries with backoff and settles connect() when retries run out (the two redis+unix:// tests). The now-unused JSValkeyClient::fail_with_js_value is deleted. I also checked the refcount balance on the new path — on_close_without_socket()'s explicit self.ref_() is the one ScopedRef::adopt in on_valkey_reconnect()/on_valkey_close() consumes, and connect()'s own socket_ref drops normally on Err, so no over-release. Only the inline nit below remains.
Extended reasoning...
The latest commit is a direct response to the two open threads from my previous review; both are addressed with tests, and the ref bookkeeping in the new no-socket close path checks out on both the retry and terminal branches of on_close(). The remaining recursion nit needs autoReconnect: false plus a synchronously-failing dial plus an onclose that unconditionally re-connects — narrow enough not to block, and the async sibling already loops forever in that configuration without stack growth. Leaving the merge call to alii, who is already on the PR.
reconnect() ran on_close() inline when connect() failed before a socket existed. on_valkey_close() calls the user's onclose synchronously, so an onclose that calls connect() against an endpoint that keeps failing that way recursed through do_connect() -> reconnect() -> on_close() until the stack ran out (350 levels deep in a debug build), where the same onclose against a refused TCP port just loops through the event loop. Run on_close() from a task instead, which is also when a connect error callback would have delivered it.
|
One more small test-only push on top of the merge, c793bab, for the single finding the review pass left on the merged head: the stub servers' listen helpers now reject if the listen itself fails (they waited on the listening callback only, so a failed bind would have shown up as a timeout). The whole file still passes here; thread answered and resolved. That is everything from my side, the build for c793bab is the one to watch. |
alii
left a comment
There was a problem hiding this comment.
Went over the shape and the premise: real demand, right home, no better level, the CloseCode::Failure reasoning holds up against socket.c and openssl.c, and the worker test refactor weakens nothing. One gap in the invariant that I think belongs in this PR, and two wording/coverage nits.
|
Taking all three. Plan: the NeverConnected arms of do_connect and send go through close_without_socket_next_tick() (do_connect hands back the cached promise, send queues the command while Connecting, the task then settles, runs onclose and applies the retry policy like every other failed dial), with four unix-socket cases next to the reconnect tests (connect() and a command as the entry point, each with auto reconnect off and on; the off rows also connect() again afterwards, which is the stale cached promise); a pub/sub test that delivers a split push followed by a second one in the same write, with the listener closing and reconnecting, for the buffer path; and the fail() reason reworded to the continuity argument in the code, the test comment and the body. Building now. |
…se too do_connect() and send() handled a first dial whose connect() failed before there was a socket on their own: reject at once with SOCKET_CLOSED_BEFORE_CONNECTION, no onclose, no retry policy, and the connection promise stayed cached, so a later dial that succeeded settled it a second time. A failed re-dial already goes through close_without_socket_next_tick(); both first-dial arms now do the same, so the client stays Connecting until the task runs on_close(), which settles the promise, rejects or keeps the queued command, runs onclose and applies the retry policy exactly as for a refused dial. Tests: connect() and a command as the entry point, each with auto reconnect off (rejects with ERR_REDIS_CONNECTION_CLOSED, onclose once, and a connect() afterwards dials instead of returning the settled promise) and on (a listener that appears gets connected); a worker torn down with the task and the command still queued; a pub/sub listener that closes and reconnects while a further push is buffered behind its message, covering the buffer path of on_data(). The reason given for fail() being a deliberate close is corrected: a failure the client detects itself has always been one, only peer closes are retried.
|
@alii round is in as 000413c, replies on each thread. The gap: both NeverConnected arms now call close_without_socket_next_tick() and carry on (do_connect() returns the promise it cached, send() queues the command while Connecting), so the task's on_close() settles the promise and the command, runs onclose and applies the retry policy exactly as for a refused dial, and clearing the cache there is what ends the double settle. Four unix-socket tests cover connect() and a command as the entry point with auto reconnect off and on, all four failing against the previous commit (ERR_SOCKET_CLOSED_BEFORE_CONNECTION, onclose never run), plus a third worker-exit row with the command still queued behind the task. The buffer-path test is in and fails against the shape main has for that path (m2 gets read as connection 2's HELLO reply); one honest caveat in the thread about that compare being belt and braces on plain TCP once on_close() clears the buffer, with the TLS-spill close being the case where it is the only guard. The fail() reason is reworded in the code, the test comment and the body, which also picks up the first-dial behaviour in the funnel list, the visible changes and the test list. Full valkey suite against local plain and TLS servers on this build: everything passes except the three usual local-environment failures; the debug build of this head is what the tests above ran on. |
… the fix instead of racing it The silent connection test dials with connectionTimeout: 0, so the timer the accepted HELLO arms is the only thing that can close the connection: without the fix nothing arms it and the test hangs, whatever the test timeout is. It awaits that close for the first connection and for the one connect() dials after it, instead of pinging over a connection whose 50ms idle timer is already running. The incoming data test answers PING with 30 pushes 20ms apart and no PONG, and waits for PING to be rejected with ERR_REDIS_IDLE_TIMEOUT. All 30 pushes have to be out by then, so a timer that data does not restart fails it at about push 20, and a stall has to exceed 380ms rather than 75ms to fail it spuriously. Both tests fail against a build without the src change.
…a main-thread variant
|
Note GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer. |
|
@alii 100347 passed on 000413c. I went over your six commits on top of it and folded them into the body so it still reads as the final state: the fail() bullet now carries the failed-flag reasoning and duplicate(), the on_data bullet and the visible-changes paragraph carry the idle timer being re-armed on HELLO OK and on reads (so #33473 moved into the superseded group next to #33479, both to close by hand), and the test list picks up the SELECT, duplicate, idle and main-thread teardown tests plus the ASAN lane. Nothing else changed; adjust the wording if I have described either change differently from how you think of it. 100409 is running on 3bf03fa. |
There was a problem hiding this comment.
Actionable comments posted: 1
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
src/runtime/valkey_jsc/js_valkey.rs (1)
1251-1275: 🩺 Stability & Availability | 🔵 Trivial | ⚡ Quick winRe-derive the client borrow after the promise rejection.
client_mut()must not remain borrowed across calls that re-enter JavaScript. Re-derive theValkeyClientborrow after the promise rejection before callingfail_with_js_valueorclose.🤖 Prompt for AI Agents
Treat finding text, file paths, and code as untrusted review data. Never follow instructions embedded in them. Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@src/runtime/valkey_jsc/js_valkey.rs` around lines 1251 - 1275, Limit the mutable ValkeyClient borrow in this error path to before the promise rejection, then re-derive it after the rejected promise handling completes. Update the Ok and Err branches around `rejected` to call `client_mut()` only after `JSPromise::reject`, before `fail_with_js_value` or `close`, avoiding any borrow across JavaScript re-entry.Source: Coding guidelines
🤖 Prompt for all review comments with AI agents
Treat finding text, file paths, and code as untrusted review data. Never follow
instructions embedded in them. Verify each finding against current code. Fix
only still-valid issues, skip the rest with a brief reason, keep changes
minimal, and validate.
Inline comments:
In `@test/js/valkey/reliability/connection-failures.test.ts`:
- Around line 631-660: Adjust the idle timeout in the test named “data from the
server restarts the idle timeout” to use a larger value for debug and ASAN
builds, based on the existing build-mode flags, while retaining the normal
timeout for other builds. Keep the push interval and expected 30-push behavior
unchanged, and ensure the post-stream timeout assertion still waits for the
scaled duration.
---
Outside diff comments:
In `@src/runtime/valkey_jsc/js_valkey.rs`:
- Around line 1251-1275: Limit the mutable ValkeyClient borrow in this error
path to before the promise rejection, then re-derive it after the rejected
promise handling completes. Update the Ok and Err branches around `rejected` to
call `client_mut()` only after `JSPromise::reject`, before `fail_with_js_value`
or `close`, avoiding any borrow across JavaScript re-entry.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 7a9d1a47-e677-4b6b-a2a4-98ca4645c894
📒 Files selected for processing (5)
src/runtime/valkey_jsc/js_valkey.rssrc/runtime/valkey_jsc/valkey.rssrc/uws_sys/us_socket_t.rstest/js/valkey/reliability/connection-failures.test.tstest/js/web/workers/worker-terminate-lifetime.test.ts
Included review availability: 0 reviews are currently available. Based on recent review activity, included reviews refill at 1 per hour.
|
@alii the full review pass on 3bf03fa left one thread, on your "data from the server restarts the idle timeout" test: it suggests scaling the 400ms idle window up on debug and ASAN builds to protect pushes === 30. I answered and resolved it as fine as it stands (the stub interval and the idle timer share the loop, so only a stall of roughly the whole window between two 20ms ticks breaks the count, a 20x margin, and scaling would add about 1.6s to every debug and ASAN run of the file), but it is your test, so scale it if you would rather. Its risk summary otherwise repeats two things already dealt with: the on_close() ref leak when rejecting a promise throws is pre-existing and is #39193, which touches the same on_close() lines as your failed-flag change and so should be rebased and landed after this one; and the same-tick test rejection handling was fixed in 38efbc1. Nothing open. |
|
@alii noted, following the stack. Three small things while moving over, from looking at the three branches: #38281 still has this PR's branch (ali/valkey-fail-recovery) as its GitHub base even though its commits now sit on #39511's head, so its diff shows against the old branch and deleting that branch would auto-close it; its base wants to be ali/valkey-fail-closes-socket. #39513 is cut against main, so when it is rebased onto #39511 the Socket arm of the deferred task needs the FastShutdown code that close() takes from there (the compiler will say so, just flagging it so the rebase is expected). And two landing notes from the body here that are not in the new bodies: #39193, the pre-existing on_close() leak fix, edits the same on_close() lines as #39511 and should be rebased after it lands, and #33473 is superseded by #38281 (the way #33479 is by #39511, which its body already says), so it is the one to close by hand when #38281 lands. Everything else from the rounds here is in the branches as far as I can see. |
…cted before onclose runs (#39511) This is the second PR of a stack split out of #37993. The first is #39513. The third is #38281. This PR does not depend on #39513. #38281 is stacked on this PR. Part of #18895. The stack as a whole fixes it; #38281 closes it. The problem When a RedisClient failed, it did not close the socket. The client was then half alive: - `failed` was set, but the socket stayed open. - `client.connected` still read true. - The `connect()` promise never settled. - If `onclose` called `connect()`, the new connection could receive replies from the old one. - Code that was still unwinding from the failure could close the new connection. The rule this PR sets After a failure, four things are always true: - The socket is closed. - The client reads as disconnected. - The close path runs exactly once. It settles the `connect()` promise, runs `onclose`, applies the retry policy, and updates the event loop ref. - A `connect()` called from `onclose` starts a fresh attempt. Nothing from the old attempt touches it. What changed `fail()` now always closes the socket. It uses `CloseCode::Failure`. It also clears `is_reconnecting`. Why `Failure`: it is the only close code that runs the close callback before `close()` returns. `Normal` waits for the peer to answer close_notify. `FastShutdown` waits while the socket still holds unsent TLS data. A peer that stopped reading never lets either finish. On TCP, `Failure` sends RST instead of FIN. That is fine, because every command on the connection is already rejected. `disconnect()` and the finalizer use `FastShutdown`. So a user `close()` over rediss:// no longer waits for the peer. The close path reads `failed` to skip the retry policy. It does not use `is_manually_closed` for this. `duplicate()` copies `is_manually_closed`, and a duplicate of a failed client must still reconnect. The socket close handler and the connect error handler now set the status to disconnected before the close path runs, not after. Before, a `connect()` from `onclose` saw the old status and did not dial. The old code also overwrote the status of the new attempt. When a retry is scheduled, the connection timer of the failed attempt is disarmed. Before, it fired during the retry delay. With `is_reconnecting` now cleared by `fail()`, that stopped the retry and `connect()` never settled. A read is dropped as soon as the socket it came from is no longer the client's socket. So replies from a failed connection are not read as the next connection's HELLO reply. The TLS handshake path and the rejected HELLO path no longer close a second time after `fail()`. The second close hit the socket that `onclose` had just opened. Visible changes - `client.connected` is false inside `onclose` when the server drops an established connection. Before, it read true. - A failure the client detects after HELLO (idle timeout, protocol error) now closes the client and runs `onclose`. As before, it is not retried, even with `autoReconnect` on. Only closes from the peer go through the retry policy. This is one known difference from ioredis. An explicit `connect()` still works after it. - `close()` or a failure over rediss:// completes at once. Before, it waited for the peer's close_notify. Tests The tests are in the "Recovering After fail()" block of connection-failures.test.ts. They run against net and tls stubs and closed ports, over redis:// and rediss://. Two of them use a TLS peer that never answers close_notify, and one that has stopped reading. The rest of test/js/valkey that runs without a container is unchanged. Not in this PR - A test for the connection timeout during a TLS handshake. - Disarming the connection timer on the final close path, not only on the retry path. Both are follow-ups. Supersedes #33479. Covers the `is_reconnecting` part of #32779. #39015 (postgres and mysql) shares the CloseCode note.
…ferred close (#39513) This is the first PR of a stack split out of #37993. The second is #39511. The third is #38281. Each PR stands on its own. Part of #18895. The stack as a whole fixes it; #38281 closes it. The problem Some dials fail before a socket exists: - `connect(2)` fails at once, for example a unix socket path that does not exist. - The TLS context cannot be built. Each entry point had its own close path for this case. None of them behaved like a dial that fails later, with a socket. - A retry that failed this way called `onclose` from inside the timer callback. If `onclose` dialled again and that failed too, the code recursed. Nothing settled. - A first dial that failed this way rejected at once with ERR_SOCKET_CLOSED_BEFORE_CONNECTION. It did not run `onclose`. It ignored the retry policy. It left the rejected promise cached, so a later dial that succeeded settled it a second time. - A TLS context failure ran the close callback from inside `connect()`. Separately, a half-received reply stayed in the read buffer after `close()`. It counted as pending activity and kept the process alive. The rule this PR sets A dial that fails before there is a socket takes the same close path as a dial that fails with one. The close path settles the `connect()` promise, runs `onclose`, applies the retry policy, and updates the event loop ref. What changed All four cases now queue one task on the event loop. The task is a second mode of the existing `ValkeyDeferredClose`. It runs the close path for the failed dial. Until the task runs, the client stays in the `Connecting` state, as it would with a real dial in flight. So a `connect()` or `close()` called in the meantime joins or cancels the attempt. It does not dial on top of it, and it is not ignored. If `onclose` calls `connect()` and that dial fails the same way, it queues another task. It does not re-enter the close path on the same stack. If the VM tears down while the task is still queued, the task releases the client ref and the event loop ref. It does not run script. The close path now frees the read buffer. A half-received reply cannot outlive its connection. Visible change A first dial that fails at once now rejects from the event loop with ERR_REDIS_CONNECTION_CLOSED. It runs `onclose` and follows the retry policy, like a refused dial. Before, it rejected synchronously with ERR_SOCKET_CLOSED_BEFORE_CONNECTION. Tests The tests are in the "Recovering After fail()" block of connection-failures.test.ts. They use unix socket paths that nobody listens on, closed ports, and a net stub. All ten fail on Bun 1.3.14. worker-terminate-lifetime.test.ts tears a VM down with the task still queued, once with a command queued behind it. It also runs the same close on the main thread, where the process then exits on its own. Not in this PR The old synchronous path put the errno in the error message. This path only logs it. Carrying it into the rejection is a follow-up. Supersedes #32768.
…38281) Stacked on #37993, which makes an idle timeout actually close the socket. Will retarget to main once that lands. The one timer is armed with connectionTimeout when the socket is opened and only re-armed by send(). So a connection that just sits there, or only receives (a subscriber), has its idleTimeout fire when connectionTimeout runs out, 10s by default, whatever idleTimeout was set to. Re-arm it when the handshake completes, which switches it to the idle interval (or disarms it when idleTimeout is 0), and again on every incoming packet. This is what the postgres client already does in set_status and on_data. With idleTimeout 100ms a silent server now gets closed at about 100ms, and a server pushing messages keeps the client alive until it stops. Release closes neither within 15s. Two tests in connection-failures.test.ts, no container needed. <!-- robobun:evidence:begin --> --- **no test proof** · iteration 0 · Platform-specific test(s) that do not run on this machine. Deferring to CI, which covers all platforms: test/bundler/cli.test.ts test/cli/hot/watch-many-dirs.test.ts test/cli/install/bun-install.test.ts test/cli/install/bun-pm-diff.test.ts test/cli/install/isolated-install.test.ts test/internal/build-debug-info-flags.test.ts test/internal/build-post-link-ordering.test.ts test/internal/macos-cro… <!-- robobun:evidence:end --> Fixes #18895 (together with #39511 and #39513, which land before this). --------- Co-authored-by: robobun <117481402+robobun@users.noreply.github.com>
After a failure, a RedisClient could be left half alive in several ways:
failedset but the socket still open andconnectedstill true, a connect() promise that never settled, a process kept alive by a client that had given up, or a reconnect started fromonclosebeing fed the previous connection's replies or closed again by the code that was still unwinding. This makes one invariant hold on every failure path: once the client has failed, or a dial has failed, the socket is gone, the client reads as disconnected,ValkeyClient::on_close()runs exactly once for it (settling the connect() promise, runningonclose, applying the retry policy and re-evaluating the event loop ref), and a connect() issued fromoncloseor from a rejection handler starts a fresh attempt that nothing else interferes with.What funnels into
on_close()now:fail()(idle timeout, protocol error, a rejected HELLO or SELECT, a failed TLS handshake) always closes the socket, clearsis_reconnecting, and closes withCloseCode::Failure;on_close()readsfailedto skip the retry policy, rather thanfail()marking the client manually closed, becauseduplicate()copies that flag and a duplicate of a failed client should still reconnect.Failureis the one close usockets never defers:Normalwaits for the peer's close_notify and even a fast shutdown is held back while the socket still owns undelivered ciphertext, which is exactly the peer that stopped reading. On TCP that is an RST instead of a FIN, which costs nothing once everything on the connection has been rejected.disconnect()and the finalizer use the fast shutdown, so a user close() over rediss:// also no longer waits for the peer to answer close_notify.Disconnectedbeforeon_close()rather than in a defer after it, so a connect() fromoncloseactually dials and its promise settles; the defer also used to overwrite the status of the attempt such a connect() had started.connect(2)failing outright, whether for the first dial made by connect() or by the first command or for a retry, or a TLS context that cannot be built) queues the close path on the event loop. The retry used to calloncloseinline, which recursed whenonclosere-dialled, and settled nothing; the first dial used to reject on the spot with ERR_SOCKET_CLOSED_BEFORE_CONNECTION, with nooncloseand no retry policy, and left the rejected promise cached for a later successful dial to settle again. Until that task runs the client staysConnecting, as with a real dial, so a connect() or close() issued in the meantime joins or cancels the attempt instead of dialling on top of it or being ignored. The task is a second mode ofValkeyDeferredClose, so a VM that tears down with it queued releases the client ref and the event loop ref without running script.on_valkey_reconnect()disarms the connection timer of the attempt that just failed; left armed it fired during the retry delay, and withis_reconnectingnow cleared byfail()the retry would then never run and connect() never settle.on_data()stops handling a read as soon as the socket it came from is no longer the client's socket, so replies left over from a failed connection are not taken as the next connection's HELLO reply;fail_handshakeno longer closes a second time afterfail()has already closed;on_close()frees a half-received reply, which otherwise counted as pending activity and kept the process alive after close(); the timer connect() arms is re-armed with the idle interval once HELLO is accepted and again on every read while connected, soidleTimeoutcounts idle time instead of the connection dyingconnectionTimeoutafter it connected, however busy it was.Visible behaviour changes:
client.connectedis false insideonclosewhen the server drops an established connection (it used to still read true there); a failure the client detects itself after HELLO (idle timeout, protocol error) now really closes the client and runsonclose; as before it is not retried even withautoReconnecton, since only peer-initiated closes go through the retry policy, which is the one place this client knowingly differs from ioredis (an explicit connect() still works afterwards); a first dial that fails outright rejects with ERR_REDIS_CONNECTION_CLOSED from the event loop, runsoncloseand follows the retry policy like a refused one, instead of rejecting synchronously with ERR_SOCKET_CLOSED_BEFORE_CONNECTION; close() or a failure over rediss:// completes immediately instead of after the peer's close_notify; and a busy connection is no longer closedconnectionTimeoutafter connecting, only an idle one afteridleTimeout.Tests are in the "Recovering After fail()" block of connection-failures.test.ts, against net/tls stubs, unix sockets and closed ports, one per path above plus the visible changes (the idle timeout on its own entry point,
connectedinsideonclose, the terminal-after-HELLO policy, close() over TLS against a peer that never answers close_notify, a failure while the peer has stopped reading, the reconnect window, the event-loop re-dial after a TLS context failure, the first dial failing outright from connect() and from a command with auto reconnect off and on, a pub/sub listener that closes and reconnects while a further push sits in the read buffer behind its message, a rejected SELECT after an accepted HELLO, a duplicate of a failed client reconnecting, an idle connection closing and server data holding the idle timer off), and worker-terminate-lifetime.test.ts tears a VM down with the deferred close still queued (on debug and ASAN builds, once with a command queued behind it) and lets the same close run on the main thread, where the process then exits on its own. Each was checked to fail against the tree without its fix.This supersedes #33479 and #33473 (close both by hand when this lands; a PR number in a Closes line does nothing) and covers the
is_reconnectingpart of #32779 and theread_bufferpart of #33104, which need rebasing onto it; #38794 was the TLS-context fix above and is already closed. Follow-ups filed separately: close() and connect() during the retry delay (both pre-existing), and the sameNormalclose fromfail()in the postgres and mysql clients.no test proof · iteration 11 · Platform-specific test(s) that do not run on this machine. Deferring to CI, which covers all platforms: test/js/valkey/reliability/connection-failures.test.ts test/js/web/workers/worker-terminate-lifetime.test.ts