Skip to content

test(net): drop 100ms setTimeout race from the net.Server listen tests - #36310

Merged
Jarred-Sumner merged 3 commits into
mainfrom
farm/659f024c/net-server-listen-test-drop-100ms-timer
Jul 29, 2026
Merged

Jarred-Sumner merged 3 commits into
mainfrom
farm/659f024c/net-server-listen-test-drop-100ms-timer

Conversation

@robobun

@robobun robobun commented Jul 29, 2026 •

Copy link
Copy Markdown
Collaborator

What does this PR do?

Fixes test/js/node/net/node-net-server.test.ts > should listen on unix domain socket going red on the alpine 3.23 lanes since #36175 (seen on main builds 84293, 84503, 84549, 84601 and ~60 branch builds).

Every test in the net.createServer listen block armed setTimeout(closeAndFail, 100) next to server.listen(). That timer was never testing listen() itself: Bun.listen binds synchronously and 'listening' is scheduled via setTimeout(emitListeningNextTick, 1, this), so the 100 ms race was against the test process's own scheduling latency. The runner already bounds each test, so the extra timer only added a flake surface (and hid the real error behind function should not have been called).

#36175 didn't touch net or this file, but it moved the allowlisted files into a single batch, so the handful of remaining serial files (this one is in excludeFiles) now run much earlier in the shard. On alpine that lands while the docker-service coordinator is still bringing up the mysql containers in the background:

t=233226  [9/282] node-net-server.test.ts
t=233436  ✗ should listen on unix domain socket [144.58ms]   ← 100 ms timer fired
t=234907  coordinator: mysql_native_password ready           ← docker init finished 1.5 s later
t=242361  [attempt #2] node-net-server.test.ts  21 pass      ← same file green once docker is idle

(from build 84601, alpine 3.23 x64 shard 019fab6d-27b7-4c39)

Change

Remove the 100 ms setTimeout(closeAndFail, ...) from the nine listen tests and route server.on('error', ...) to done(err) so a real bind failure reports its actual error. Same assertions, same code paths (listen() → 'listening' → server.address() checks); only the hand-rolled deadline that duplicated the test runner's timeout is gone. The 500 ms timers in the events block are untouched; they guard real client↔server round trips, have is_done guards, and haven't flaked.

How did you verify your code works?

  • bun bd test test/js/node/net/node-net-server.test.ts → 21 pass / 0 fail.
  • Reproduced the race locally by running the file under background CPU+disk load (4× yes, 2 GB dd): with the old timer the listen block failed 1/5 runs at function should not have been called; with this change 5/5 runs pass under the same load (including a 246 ms 'listening' that would have tripped the old 100 ms timer).

node-tls-server.test.ts has the same 100 ms pattern and is also a serial excludeFiles entry; happy to fold it in here if preferred, but it hasn't been observed red so I kept this scoped to the reported file.


[stamp-90s] gate passed · iteration 1 · 1 files touched

passes on PR (with fix)
Test-only change.

Debug/ASAN (expected pass):
$ bun bd test 'test/js/node/net/node-net-server.test.ts'
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test test/js/node/net/node-net-server.test.ts
bun test v1.4.0 (6b920f8b9)

test/js/node/net/node-net-server.test.ts:
(pass) net.createServer listen > should throw when no port or path when using options [27.65ms]
(pass) net.createServer listen > should listen on IPv6 by default [153.42ms]
(pass) net.createServer listen > should listen on IPv4 [26.37ms]
(pass) net.createServer listen > should call listening [19.34ms]
(pass) net.createServer listen > should provide listening property [22.93ms]
(pass) net.createServer listen > should listen on localhost [17.90ms]
(pass) net.createServer listen > should listen on localhost [17.45ms]
(pass) net.createServer listen > should listen without port or host [24.19ms]
(pass) net.createServer listen > should listen on unix domain socket [19.40ms]
(pass) net.createServer listen > should bind IPv4 0.0.0.0 when listen on 0.0.0.0, issue#7355 [31.32ms]
(pass) net.createServer events > should receive data [159.16ms]
(pass) net.createServer events > should call end [155.39ms]
(pass) net.createServer events > should call close [19.65ms]
(pass) net.createServer events > should call connection and drop [70.90ms]
(pass) net.createServer events > should error on an invalid port [23.08ms]
(pass) net.createServer events > should call abort with signal [25.98ms]
(pass) net.createServer events > should echo data [111.91ms]
(pass) net.createServer events > #8374 [72.98ms]
(pass) accepted socket event-loop hold matches Node (per-connection KeepAlive) > server.stop() + accepted socket.unref() lets the process exit [318.84ms]
(pass) accepted socket event-loop hold matches Node (per-connection KeepAlive) > server.unref() alone does not drop a ref'd accepted connection's hold [1548.16ms]
(pass) accepted socket event-loop hold matches Node (per-connection KeepAlive) > half-open accepted sockets after peer FIN do not busy-poll the event loop (Windows AFD DISCONNECT) [4
... (truncated)
Exit: 0
diff hotspot
test/js/node/net/node-net-server.test.ts | 117 +++++++------------------------
 1 file changed, 25 insertions(+), 92 deletions(-)

gate history · 2 passed · 0 rejected · iteration 1

evidence per changed file
file                                      reads  edits  tests
test/js/node/net/node-net-server.test.ts      2      3      0

The 'net.createServer listen' describe block armed setTimeout(closeAndFail, 100)
alongside server.listen() as a secondary deadline. Since #36175 reordered the
serial phase, this file runs early in the alpine shard while the docker
coordinator is still initializing mysql containers in the background; the
scheduling contention let the 100ms timer beat the 'listening' event and the
test failed with the uninformative 'function should not have been called'.

The timer was never exercising listen() itself (Bun.listen binds synchronously
and 'listening' is scheduled via setTimeout(...,1)); it was racing the test
process's own scheduling latency. The test runner already imposes a per-test
timeout, so remove the hand-rolled one and surface the real Error to done()
when a listen genuinely fails. Same code paths, same assertions.
@coderabbitai

coderabbitai Bot commented Jul 29, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

@robobun, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 13 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 9695fb03-3409-4b29-9be4-3a826c9ebeb4

📥 Commits

Reviewing files that changed from the base of the PR and between 59242d6 and 6b920f8.

📒 Files selected for processing (1)
  • test/js/node/net/node-net-server.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Jul 29, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 1:06 AM PT - Jul 29th, 2026

❌ @robobun, your commit 6b920f8 has some failures in Build #84795 (All Failures)


🧪   To try this PR locally:

bunx bun-pr 36310

That installs a local version of the PR into your bun-36310 executable, so you can run:

bun-36310 --bun

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — test-only de-flake that removes a redundant 100ms secondary deadline; the assertions (mustCall on the listening callback, server.address() checks) are preserved and the error path now reports the real bind error instead of function should not have been called. One inline nit on the comment block, non-blocking.

Extended reasoning...

Overview

This PR touches a single test file, test/js/node/net/node-net-server.test.ts, removing the setTimeout(closeAndFail, 100) pattern from the nine tests in the net.createServer listen block. The per-test closeAndFail closure is replaced by a shared failOnError(server, done) helper that closes the server and calls done(err), and the now-dead clearTimeout calls inside the listening callbacks are removed. No runtime code is changed.

Security risks

None — test-only change with no new inputs, no new subprocess arguments, no network endpoints beyond the existing port: 0 / unix-socket binds already present.

Level of scrutiny

Low. This is a mechanical de-flake of a test harness pattern. I checked that:

  • The property the removed timer protected ("listening fires") is still asserted via mustCall on every listening callback.
  • The error path is strictly better: done(err) surfaces the real bind failure instead of the generic mustNotCall() message, which REVIEW.md's "Never genericize a rich existing message" prefers.
  • If neither 'listening' nor 'error' fires, the test-runner default timeout still bounds the test — the 100 ms timer was duplicating that bound with a value tight enough to race against scheduling latency under CI load, exactly the setTimeout-as-deadline pattern REVIEW.md flags.
  • The events block's 500 ms timers (which guard real client↔server round-trips and have is_done guards) are intentionally left alone, and Timer is still referenced there so no dead type usage.

Other factors

The PR description carries concrete CI evidence (build numbers, timing log) and a local reproduction under synthetic load, and explicitly scopes out node-tls-server.test.ts (same pattern, not observed red). The one inline nit — the 7-line comment above failOnError is mostly incident narrative that REVIEW.md says belongs in the PR description — is stylistic and doesn't affect correctness, so I'm approving with it noted rather than blocking on it.

Comment thread test/js/node/net/node-net-server.test.ts Outdated

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — the comment trim in 7ffe330 addresses the earlier nit.

What was reviewed:

  • Confirmed the removed 100ms timer only raced scheduling latency (bind is sync, 'listening' is next-tick) and the runner's per-test timeout still bounds a real hang.
  • Traced createCallCheckCtx/mustCall in node-harness: dropping the mustNotCall destructure doesn't change the done-slot count, and failOnError → done(err) still fails the test with a better message than the old mustNotCall()().
  • All address() assertions and server.close() cleanup on both success and error paths are preserved; the untouched 500ms events-block timers keep their is_done guards.
Extended reasoning...

Overview

Test-only de-flake of test/js/node/net/node-net-server.test.ts. Nine tests in the net.createServer listen block previously armed setTimeout(closeAndFail, 100) alongside server.listen(); this PR removes those timers and replaces the per-test closeAndFail boilerplate with a shared failOnError(server, done) helper that routes server.on('error', ...) to done(err). No runtime code, no assertion changes.

Security risks

None. Test file only; no production code paths, no new external inputs.

Level of scrutiny

Low. This is a mechanical de-flake following the exact pattern REVIEW.md and CLAUDE.md prescribe: don't use setTimeout to wait for a condition, rely on the test runner's per-test timeout, and wire failure events to reject/done with the real error. The PR description root-causes the flake (the 100ms timer measured scheduling contention on alpine while docker prestart ran, not listen() itself) and reproduces it locally under load.

I checked the REVIEW.md rule about not silently weakening a safety net: the property the old timer protected — "listen resolves to 'listening' or 'error' in bounded time" — is still enforced by (a) the runner's default timeout and (b) the 'error' listener now calling done(err). All mustCall wrappers on the listen callbacks and all expect() calls on server.address() are unchanged. Reading createCallCheckCtx in test/js/node/harness.ts confirmed that removing the unused mustNotCall destructure has no side effects (it only creates a done slot when called, not when destructured), and that mustCall's internal closeTimers() never touched the old local timeout variable anyway.

Other factors

The one prior finding (bug-history comment content) was addressed in 7ffe330 with exactly the suggested trim, and the thread is resolved. The events block's 500ms timers are deliberately left alone with a stated reason (real round-trips, is_done-guarded, not observed flaking). node-tls-server.test.ts has the same pattern and is called out in the description as intentionally out of scope. bun bd test passes 21/21 per the PR.

@robobun

robobun commented Jul 29, 2026 •

Copy link
Copy Markdown
Collaborator Author

Diff is green: node-net-server.test.ts no longer appears in the alpine annotations on build 84795 (35/40 alpine shards finished, 5 still queued; it was in the flaky set on main builds 84293/84503/84549/84601 and is gone here). bun bd test is 21/21 locally and the old 100 ms timer reproduced the flake under background load.

The remaining annotations on 84795 are unrelated to this change: terminal-platform-gaps.test.ts (win2019 ConPTY regex), cpu-prof.test.ts (win2019 "No samples collected"), complex-workspace.test.ts (verdaccio install), and a handful of parallel-batch entries that all passed when re-run alone. The previous build (84734) was a fleet-wide stall where most build-cpp jobs never left scheduled, so build-bun timed out waiting on them; nothing a test-file edit can reach.

Ready for a maintainer to merge.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — the comment-trim nit from the earlier pass is addressed, and the rest is a clean de-flake.

What was reviewed:

  • Confirmed the removed 100ms timer was the only thing dropped; every mustCall wrapper and server.address() assertion is intact, so the property "listen fires the callback" is still enforced by mustCall + the runner's per-test bound.
  • failOnError routes 'error' → done(err), which is strictly better diagnostics than the old mustNotCall()() and matches REVIEW.md's "wire failure events to reject" rule.
  • Checked that Timer and mustNotCall are still used in the untouched events block, so no dead imports/destructures were left behind.
Extended reasoning...

Overview

Test-only change to test/js/node/net/node-net-server.test.ts. Nine tests in the net.createServer listen block previously armed setTimeout(closeAndFail, 100) alongside server.listen(); under CI load on the alpine lane the 100ms timer fired before 'listening' (which is scheduled via a 1ms setTimeout), producing a flake. The PR removes those timers, extracts a shared failOnError helper that routes server.on('error', ...) to done(err), and leaves the events block's 500ms round-trip timers untouched. Net: +25 / −92 lines, no runtime code changed.

Security risks

None. Pure test-file edit; no production code paths, no auth/crypto/network-facing surface touched.

Level of scrutiny

Low. This is a targeted de-flake that follows REVIEW.md's own testing guidance ("don't use setTimeout to wait for a condition"; "wire failure events to reject the awaited promise"; "when de-flaking, keep asserting the property the original assertion protected"). The removed timer was measuring test-process scheduling latency rather than listen() behavior, and the test runner's per-test bound already covers the hang case. Every existing assertion (address.address, address.port, address.family, server.listening) is preserved verbatim, and mustCall still enforces that the listen callback actually fires.

Other factors

  • The one prior review comment (mine) asked for the CI-history portion of the new comment to be trimmed; commit 7ffe330 applied exactly the suggested wording and the thread is resolved.
  • Error reporting is improved: a real bind failure now surfaces the actual error object via done(err) instead of the opaque "function should not have been called".
  • Verified no dead code introduced — the Timer type and mustNotCall are still consumed by the unchanged events describe block.
  • The PR description acknowledges node-tls-server.test.ts has the same pattern and deliberately scopes to the file that was actually flaking, which is a reasonable boundary for a test-only change.
  • 21/21 tests pass on the debug build per the evidence block.

@Jarred-Sumner
Jarred-Sumner merged commit ccc6110 into main Jul 29, 2026
53 checks passed
@Jarred-Sumner
Jarred-Sumner deleted the farm/659f024c/net-server-listen-test-drop-100ms-timer branch July 29, 2026 12:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants