Skip to content

redis: arm idleTimeout on an idle timer, not the connect timer - #33473

Closed
robobun wants to merge 3 commits into
mainfrom
farm/4d72aca0/redis-idle-timeout
Closed

robobun wants to merge 3 commits into
mainfrom
farm/4d72aca0/redis-idle-timeout

Conversation

@robobun

@robobun robobun commented Jul 6, 2026 •

Copy link
Copy Markdown
Collaborator

Repro

idleTimeout: 60_000 with connectionTimeout: 250, against a scripted RESP3 server, issuing a command every 25ms:

const c = new Bun.RedisClient(url, { idleTimeout: 60_000, connectionTimeout: 250 });
c.onclose = () => console.log("onclose fired");
await c.connect();
for (let i = 0; Date.now() - t0 < 1200; i++) {
  try { await c.get("k" + i); } catch (e) { console.log(Date.now() - t0, e.code, e.message); break; }
  await Bun.sleep(25);
}
console.log({ connected: c.connected });
254 ERR_REDIS_CONNECTION_CLOSED Connection has failed
{ connected: true }

The client dies exactly connectionTimeout ms after it connected, while it is busy, and onclose never fires. With the default idleTimeout: 0 the same mis-armed timer fires too, but its handler is a no-op, which is why only opt-in users see it. Previously reported as #18897 and #19044, both closed by changing the default to 0.

Cause

There is one timer. connect() arms it with connection_timeout_ms, and nothing cancels or re-arms it afterwards. on_connection_timeout then decides what the firing means from the status at that moment:

match self.status {
    Status::Connected => // "Idle timeout reached after {idle_timeout_interval_ms}ms"
    _ => // "Connection timeout reached after {connection_timeout_ms}ms"
}

So the connect-phase timer detonates on a now-connected client and is reported as an idle timeout. Conversely, a connection that really is idle is never closed, because the idle interval is never armed.

The failure path then left the client unusable but undetectable: fail_with_js_value only closes the socket when !connection_ready(), so a connected client kept status == Connected (.connected === true), never fired onclose, never auto-reconnected, and rejected every later command with Connection has failed.

Fix

The two-phase timer mirrors what PostgresSQLConnection already does.

  • reset_idle_timeout() re-arms the timer only once the handshake completed, so traffic cannot push out the connect deadline. Called after every reply (on_data) and every command (send). The handshake reply lands in on_data too, so that is also where the connect-phase timer gets swapped for the idle timer (or cancelled, when idleTimeout is 0).
  • When the idle timer fires, the socket is closed: onclose runs, .connected goes false, pending commands reject with ERR_REDIS_IDLE_TIMEOUT. Auto-reconnect is suppressed (dropping an idle connection is deliberate, reconnecting it immediately would defeat the option), and the client reopens lazily on the next command, like a freshly constructed one.
  • The timer is cancelled when the connection closes, in on_valkey_close() and on_valkey_reconnect(), the two hooks every branch of ValkeyClient::on_close() funnels through. Left armed, it fired during the reconnect backoff and rejected the offline queue with a connectionTimeout that never applied, while setting is_manually_closed, which on_open does not clear: auto-reconnect was then off for good.
  • connect() calls the new ValkeyClient::reset_for_new_connection() before opening a socket, so a reopened connection does not inherit the closed one's is_authenticated. Without it, enqueue() saw connection_ready() and wrote the command straight to a socket that had not opened yet; on_open dropped those bytes and the promise stayed in the in-flight queue, pairing every later reply with the wrong command. That one was already reachable on main without idleTimeout, via close() plus an un-awaited connect().
  • status is set to Disconnected before on_close runs JS, so onclose handlers and rejected promises no longer observe .connected === true on a dead socket. ValkeyClient::close() already ordered it this way on its semi-socket path.

connectionTimeout now bounds exactly the connect phase, which is what it is documented to do.

Verification

test/js/valkey/valkey-timeout.test.ts (new, talks to an in-process RESP3 server so it needs no redis):

before after
busy connection outlives connectionTimeout fails at 415ms pass
idle connection closed at idleTimeout, reopens on next command onclose never fires pass
non-pipelineable command reopens after an idle close resolves with the HELLO map pass
enableAutoPipelining: false reopens after an idle close resolves with the HELLO map pass
server-initiated close leaves no timer armed rejects with a bogus ERR_REDIS_CONNECTION_TIMEOUT pass
command racing an un-awaited connect() gets its own reply resolves with the HELLO map pass
connectionTimeout still fires on a stalled handshake pass pass

The server-close test is deterministic rather than racy: the hangup (270ms), the idle deadline (300ms) and the reconnect (~320ms) are all deadlines in one timer heap, so their order holds even under a stall. 10/10 clean runs.

Also ran the docker-gated suites against a local redis (reliability/recovery.test.ts (#29925), reliability/connection-failures.test.ts, unit/*, integration/complex-operations.test.ts): 84 pass, 1 fail, and that one needs a route to 192.0.2.1 and fails identically on main in the same container.

Manual check against a real redis (pub/sub, manual close, server-side CLIENT KILL + autoReconnect, idle close + reopen)
1 get: 1 connected: true
1 onclose: { code: "ERR_REDIS_CONNECTION_CLOSED", connected: false }   # was connected: true
2 after 700ms idle, default idleTimeout: 1 connected: true             # connectionTimeout was 200
3 pubsub: hello
4 after server kill, autoReconnect get: 1 connected: true
5 idle close: { code: "ERR_REDIS_CONNECTION_CLOSED", connected: false } after 255 ms
5 reopen on next command: 1 connected: true
6 INFO after close()+connect(): own reply (correct)

@robobun
robobun requested a review from alii as a code owner July 6, 2026 11:36
@coderabbitai

coderabbitai Bot commented Jul 6, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 21 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: ec7bad19-168b-4efe-ba0e-492669206e87

📥 Commits

Reviewing files that changed from the base of the PR and between 48ff9eb and 5170d70.

📒 Files selected for processing (5)
  • docs/runtime/redis.mdx
  • packages/bun-types/redis.d.ts
  • src/runtime/valkey_jsc/js_valkey.rs
  • src/runtime/valkey_jsc/valkey.rs
  • test/js/valkey/valkey-timeout.test.ts

Comment @coderabbitai help to get the list of available commands.

@github-actions github-actions Bot added the claude label Jul 6, 2026
@robobun

robobun commented Jul 6, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 6:51 PM PT - Jul 6th, 2026

❌ @robobun, your commit 5170d70 has 2 failures in Build #69263 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 33473

That installs a local version of the PR into your bun-33473 executable, so you can run:

bun-33473 --bun

@github-actions

github-actions Bot commented Jul 6, 2026

Copy link
Copy Markdown
Contributor

Found 1 issue this PR may fix:

  1. redis.connect behave very wrong when failed to connect (see description) #18895 - A commenter reports the exact symptoms this PR fixes: with idleTimeout set, commands fail after the timeout period despite active use, and onclose is never called

If this is helpful, copy the block below into the PR description to auto-close this issue on merge.

Fixes #18895

🤖 Generated with Claude Code

@robobun

robobun commented Jul 6, 2026

Copy link
Copy Markdown
Collaborator Author

On the Fixes #18895 suggestion: I checked, and it only half-holds, so I'm leaving it off.

What this PR fixes is the symptom @unek describes in that issue's comments: with idleTimeout set, commands start failing with connection closed while the client is constantly busy, and onclose is never called. That's exactly the mis-armed timer plus the silent zombie this PR addresses.

The repro in the issue body is a different bug and still reproduces on this branch. Connecting to a closed port with connectionTimeout: 5000:

autoReconnect: true  (default) -> rejects @31292ms  ERR_REDIS_CONNECTION_CLOSED
autoReconnect: false           -> rejects @33ms     ERR_REDIS_CONNECTION_CLOSED

connectionTimeout bounds a single connect attempt, and every auto-reconnect attempt re-arms it from scratch, so across maxRetries attempts with backoff the deadline never fires and connect() looks like it hangs. Same story for the unknown-host case in the body. That's the initial-connect retry path, not the idle timer, and fixing it means deciding whether connectionTimeout should bound the whole connect() call or each attempt — a separate change.

So #18895 should stay open for the behavior in its body, and I've left the PR description alone.

Comment thread src/runtime/valkey_jsc/js_valkey.rs Outdated
Comment thread src/runtime/valkey_jsc/js_valkey.rs

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both issues I flagged earlier are now addressed (reset_for_new_connection() in connect(), timer cancellation in on_valkey_reconnect/on_valkey_close, plus regression tests for each). I didn't find anything new, but the connection state-machine and timer-lifecycle interactions here are subtle enough that a human pass is still worthwhile.

Extended reasoning...

Overview

This PR fixes RedisClient's idleTimeout option, which previously reused the connect-phase timer and so fired connectionTimeout ms after connecting regardless of activity. The fix adds reset_idle_timeout() (re-arms the timer on every on_data/send once connected), makes the idle-timeout close path set is_manually_closed/needs_to_open_socket and actually close the socket, moves status = Disconnected before JS runs in on_close/on_connect_error, cancels the stale timer in on_valkey_reconnect/on_valkey_close, and centralizes per-connection flag reset into reset_for_new_connection() called from connect(). Also touches docs (redis.mdx), types (redis.d.ts), and adds a new self-contained test file with an in-process RESP3 mock server.

Prior review follow-up

My earlier run flagged two issues on the first commit: (1) the lazy-reconnect prologue left is_authenticated stale-true, causing non-pipelineable commands to be written to a not-yet-open socket and orphaned in in_flight; (2) the idle timer wasn't cancelled on server-initiated close, so it could fire during the reconnect backoff and permanently poison auto-reconnect via is_manually_closed. The follow-up commit (5124d8b) addresses both — reset_for_new_connection() now clears all four flags and is called from connect() (covering send(), do_connect(), and reconnect() uniformly), and both on_valkey_reconnect and on_valkey_close now call disable_connection_timeout(). Each fix has a dedicated regression test in valkey-timeout.test.ts (the non-pipelineable/enableAutoPipelining: false reopen tests, the server-initiated-close test, and the un-awaited-connect race test).

Security risks

None identified. No auth, crypto, or input-validation surface is touched; the change is timer arming/cancellation and boolean flag resets on an existing state machine.

Level of scrutiny

Moderate-to-high. The Redis client's connection lifecycle is a bag of ~10 interacting boolean flags plus two timers across several close/reconnect/reopen paths, and my first pass found two real off-by-one-state bugs. The follow-up commit looks correct and is well-tested, and the bug-hunting system found nothing further, but the interaction surface (idle-timeout close → lazy reopen, server-close → auto-reconnect, connect() now resetting flags on every call including from reconnect()) is intricate enough that a human familiar with this state machine should confirm nothing else is perturbed.

Other factors

Test coverage is good: five scenarios against an in-process RESP3 server (busy-survives-connectionTimeout, idle-close-then-reopen, non-pipelineable reopen ×2, server-close-doesn't-leak-timer, handshake-timeout-still-fires, connect-race), and the PR description reports the docker-gated reliability/unit suites pass. The redundancy between do_connect()'s manual flag clears and connect()'s new reset_for_new_connection() is harmless.

@robobun

robobun commented Jul 6, 2026

Copy link
Copy Markdown
Collaborator Author

Agreed on wanting a human pass, and thanks for the two catches.

On the one loose end: the overlap between do_connect()'s flag clears and connect()'s new reset_for_new_connection() isn't removable as-is. do_connect() has an arm that never reaches connect() at all, the one where the status is still Connected/Connecting and no socket gets opened. A client that fail()ed while connection_ready() keeps its socket, so it sits at status == Connected with failed == true, and clearing failed at the top of do_connect() is the only thing that lets commands flow again on that live socket. The other two arms (needs_to_open_socket, and Disconnected -> reconnect()) both reach connect(), so for those the clears are genuinely redundant. Untangling that arm felt out of scope for a timer fix.

Worth noting for whoever picks this up: reset_for_new_connection() is really the completion of #29925's fix. That test's own header already names the problem:

Plus on reconnect, the lingering is_authenticated = true from the prior session caused the new HELLO response to be silently dropped

That was patched by clearing the flag in on_open(), which closes the window after the socket is up but leaves the window between connect() and on_open(), where enqueue() still reads connection_ready(). Moving the reset to connect() closes that second window, which is why close() + an un-awaited connect() + INFO currently returns the HELLO map on main.

robobun added 2 commits July 6, 2026 13:49
RedisClient had one timer, armed at connect with connectionTimeout and
never re-armed. What its firing meant was decided from the status at that
moment, so on a connected client it was reported as an idle timeout: any
client with idleTimeout > 0 died connectionTimeout ms after connecting no
matter how busy it was, and a connection that was genuinely idle was never
closed.

Cancel the connect-phase timer once the handshake completes, arm the idle
timer instead, and reset it on every command sent and every reply received.
When it fires, close the socket so onclose runs and the next command opens
a new connection, rather than leaving a client that reports connected but
rejects every command.

Also flip status to Disconnected before on_close runs JS: onclose handlers
and rejected promises were observing connected === true on a dead socket.
…lose

Two problems with the previous commit, both found in review.

A reopened connection kept the closed connection's is_authenticated, so
enqueue() saw connection_ready() and wrote the command straight to a socket
that had not opened yet. on_open then dropped those bytes and left the promise
in the in-flight queue, so every later reply was paired with the wrong command.
Reachable with any non-pipelineable command (INFO, MULTI, WATCH, SUBSCRIBE, ...)
or with enableAutoPipelining: false. The same state confusion was already
reachable without idleTimeout, via close() followed by an un-awaited connect().
Reset the per-connection flags in connect(), which every path that opens a
socket goes through.

The timer was not cancelled when the connection closed, so it could fire during
the reconnect delay, reject the offline queue with a "Connection timeout
reached after 10000ms" that never applied, and set is_manually_closed, which
on_open does not clear: auto-reconnect was then disabled for good. Cancel it in
on_valkey_close and on_valkey_reconnect, the two hooks every close path funnels
through. connect() arms a fresh deadline for the next attempt.
@robobun
robobun force-pushed the farm/4d72aca0/redis-idle-timeout branch from 5124d8b to 40445fb Compare July 6, 2026 13:51

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Both earlier issues are addressed — reset_for_new_connection() in connect() and the disable_connection_timeout() calls in on_valkey_close/on_valkey_reconnect look correct, and the new tests cover the variants. No new issues on this pass; leaving final sign-off to a maintainer given how many connection-state flags and timer transitions this touches.

Extended reasoning...

Overview

This PR fixes idleTimeout on Bun.RedisClient: previously the single connect-phase timer was never re-armed, so it fired connectionTimeout ms after connecting and was misreported as an idle timeout. The fix introduces reset_idle_timeout() (called from on_data and send), cancels the timer in on_valkey_close/on_valkey_reconnect, adds reset_for_new_connection() to clear per-connection flags before opening a new socket, moves status = Disconnected ahead of the JS-visible close path, and makes the idle-timeout close suppress auto-reconnect while allowing lazy reopen on the next command. Docs and .d.ts are updated, and a new 267-line test file exercises the timer/reopen matrix against an in-process RESP3 server.

Since my previous review, the author addressed both issues I raised: (1) stale is_authenticated on the lazy-reconnect path is now cleared in connect() via reset_for_new_connection(), which also fixes a pre-existing close() + un-awaited connect() off-by-one; (2) the idle timer is now cancelled on every close path via disable_connection_timeout() in both on_valkey_close() and on_valkey_reconnect(). Both fixes have dedicated tests, and the two inline threads are resolved.

Security risks

None identified. This is client-side timer and connection-state management; no auth, crypto, or untrusted-input parsing is touched. The RESP3 mock server in the test file is test-only.

Level of scrutiny

Moderate-to-high. The change is well-reasoned and well-tested, but it reworks the interaction between five connection flags (is_authenticated, is_manually_closed, failed, needs_to_open_socket, is_selecting_db_internal), two timers, and three close/reconnect entry points. My first pass found two real bugs in exactly these interactions, and the author's own follow-up notes a remaining (intentional) overlap between do_connect()'s flag clears and reset_for_new_connection(). The design decision to suppress auto-reconnect on idle-timeout (setting is_manually_closed = true) is reasonable but is a behavioral choice a maintainer should confirm.

Other factors

  • The author explicitly agreed a human pass is warranted.
  • Test coverage is thorough (7 new tests including the non-pipelineable / enableAutoPipelining: false variants, server-initiated close, and the un-awaited connect() race), and the docker-gated reliability suites were run.
  • The current bug-hunting pass found no issues.
  • The status = Disconnected reordering in SocketHandler::on_close/on_connect_error is a small but user-visible behavior change (.connected now reads false inside onclose); it's clearly correct but worth a maintainer glance.

@robobun

robobun commented Jul 6, 2026

Copy link
Copy Markdown
Collaborator Author

CI status: the diff is green; the only red lane is Buildkite infrastructure, not this change.

Both runs (builds 69033 and 69263) fail on the same lane, darwin 26 aarch64 - test-bun, and in both the failure is before any test executes:

Searching for artifacts: "**" within step: "darwin-aarch64-build-bun"
Found 2 artifacts. Starting to download to: .../bun/bun/release
Error: buildkite-agent artifact download timed out after 120s for step 'darwin-aarch64-build-bun'.
Refusing to continue with a partial download (would silently fall back to the wrong binary).

The darwin aarch64 - build-bun step itself passes both times, so the binary is produced; the test lane just cannot pull it from S3 within the 120s window. 280 lanes pass across debian/ubuntu x64 + baseline + x64-asan, alpine aarch64/x64, and windows x64/aarch64, including every lane that runs the new test/js/valkey/valkey-timeout.test.ts.

I re-triggered once already and it hit the identical timeout, so I am not going to keep re-rolling. A maintainer re-running just the failed darwin job (or a merge, since the failure cannot involve this diff) is the path forward. Review threads are all resolved and the latest Claude review signed off.

@alii

alii commented Aug 18, 2026

Copy link
Copy Markdown
Member

Superseded by #38281, which re-arms the timer with idleTimeout once HELLO is accepted and again on every read, and documents what idleTimeout counts. Tests are in test/js/valkey/reliability/connection-failures.test.ts.

@alii alii closed this Aug 18, 2026
@robobun

robobun commented Aug 18, 2026

Copy link
Copy Markdown
Collaborator Author

Agreed, thanks. I checked the merged stack (#39513, #39511, #38281) against the test file from this PR on current main (4c68990), since review here had turned up two extra bugs beyond the timer itself.

All of them are covered:

The only tests from here that fail on main are the three that expected the next command after an idle close to reopen the connection by itself. Main documents the opposite (idleTimeout closes, onclose runs, connect() reconnects), and I confirmed that contract works end to end, including a non-pipelineable command as the first one after connect(). So that is a design difference, not a gap. Nothing left from this PR.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants