Skip to content

test(worker-terminate-lifetime): exit the fetch-teardown fixture on its condition, not on a clock - #40781

Open
robobun wants to merge 2 commits into
mainfrom
farm/db5f0601/worker-exit-fetch-fixture-condition
Open

robobun wants to merge 2 commits into
mainfrom
farm/db5f0601/worker-exit-fetch-fixture-condition

Conversation

@robobun

@robobun robobun commented Aug 28, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • test/js/web/workers/worker-terminate-lifetime.test.ts "worker exit with streaming-request-body fetches whose response.body was touched" is red on the macOS x64 lane in every parallel batch since build 107620, and failed alone in 107646: expect(stderr).toBe("") received "worker exited 3\n".
  • The fixture's worker exits on a clock, setTimeout(() => process.exit(touched > 0 ? 0 : 3), 150 + (i * 13) % 60) (line 996 on main). On a macOS x64 CI host with hours of uptime, a burst of 48 loopback connections takes 130 to 800 ms to establish, so no fetch had answered when the deadline fired. The fixture came from Worker teardown: more fixes from fuzzing terminate()/process.exit() lifetimes; WebKit bump for Atomics.wait #38457.

Fix

  • The worker exits once all eight of its responses have been touched. The exit still runs from a timer callback, as before.
  • Correct because the state the test tears down (a streaming request body whose sink holds the FetchTasklet, plus a JS-touched response body) exists exactly when a response has been touched while its request body still streams. Waiting for the condition is the test's intent. The clock was a proxy for it that only held on fast hosts.
  • A fetch rejection, or no responses within half the test timeout, still exits 3 and prints the reason. A real failure is reported instead of a hang.
  • Verified: on a macOS x64 CI host (darwin-x64-matzo, CI binary from build 107609) the old fixture failed 10 of 10 runs and the new one passed 10 of 10, plus 6 of 6 copies run at once. With the producer ref in ProducerHold removed locally, both fixtures crash 3 of 3 under the debug ASAN build, so the new fixture covers the same teardown path. Whole file passes on the debug ASAN build.

Background

  • ProducerHold (src/runtime/webcore/ByteStream.rs) is the fetch tasklet's counted ref on the response stream's native source. It lets the tasklet unhook itself at worker exit without reading a JS cell that the VM's last sweep may already have destroyed.
  • The macOS x64 CI hosts are Intel minis on macOS 14. After about 20 hours of CI uptime, netstat -sp tcp showed about 137 retransmit timeouts per 100 loopback connects, and a node client against a node server saw the same delay. After the host rebooted, the same 48 connects took 10 ms. So the delay is host state, not bun.
  • Request the post-entry-point full collection asynchronously #40338 carries a variant of this hunk (exit after the first touched response, a fixed 5000 ms guard, fetch errors swallowed), coupled to a GC change. This PR is the standalone fix. Request the post-entry-point full collection asynchronously #40338 can drop its hunk on rebase.
Notes
  • Every main build since the test landed (100135 on 08-17 through 107609 on 08-28) shows the same first-response latency (130 to 330 ms) for the old fixture on the loaded host, so there is no bun regression in the 107628 window. The test was already flaky on main there (106692, 106734, 107443, 107609, all "in the parallel batch; passed alone").
  • Per-worker first-response time on the loaded host, old fixture, build 107609: 131 to 179 ms unloaded, 297 to 403 ms while another job ran. All six workers get their first response within a few ms of each other, which points at connection establishment, not the worker or the server.
  • 48 concurrent fetch() calls from the main thread to a local Bun.serve before the reboot: linux 18 ms, macOS x64 host 84 to 788 ms. Node client to Node server on the same host: 464 to 569 ms. Single connections were 1 to 17 ms there. After the reboot: 10 ms, and the old fixture passed 5 of 5. The new fixture passed 10 of 10 in both host states.
  • dns.lookup("localhost") on the host takes 3 ms, so the DNS path (DNSServiceGetAddrInfoEx) is not involved.
  • The first revision kept the per-worker (i * 13) % 60 offset, anchored on the eighth response. The self-review showed it changes nothing: all eight producer holds exist before the timer is armed, the only later stream transition (park at 256 KB buffered) is about 2.5 s away at 512 B per 5 ms, and the exit already lands at a random point in the 5 ms cadence. The second commit drops it. The unfix check was repeated with offset 0: 3 of 3 crash.
  • The unfix used for coverage: ProducerHold::hold without increment_count, and a balancing increment_count in take. Old and new fixture both hit ASSERTION FAILED: decontaminate() in StructureID::decode() 3 of 3 under bun bd.
  • The whole file on the debug ASAN build here: 24 pass, 1 fail. The failure is "terminate() while dns.lookup() is in flight does not UAF on c-ares channel teardown", an LSan report for the node_fs_binding::Binding box, not touched by this change (node:fs: mark the per-VM Binding box as LSan-ignored (fixes worker-terminate-lifetime.test.ts on main) #35159, node:fs: free the per-VM NodeFS when a worker's VM is torn down #39684 cover it).
  • Build 107618 shows the same test red in the Windows 2019 x64 parallel batch once. The same clock race applies there.

no test proof · iteration 0 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/js/web/workers/worker-terminate-lifetime.test.ts

…ts condition, not on a clock

The worker in "worker exit with streaming-request-body fetches whose
response.body was touched" exited 150 to 210 ms after it started and
reported exit code 3 when no fetch had answered by then. A burst of 48
loopback connections takes 100 ms or more to establish on the macOS x64
CI hosts, so the deadline passed before the state the test tears down
existed, and the test went red there in every parallel batch.

The worker now exits once all eight responses have been touched. The
per-worker offset is kept, anchored on that point. A fetch rejection
and a generous bound (half the test timeout) still exit 3, with the
reason on stderr, so a real failure is reported instead of a hang.
@coderabbitai

coderabbitai Bot commented Aug 28, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review limit reached

  • Run on-demand review

On-demand reviews are free for the next 23 days. After that, they cost $0.25 per reviewed file.

Or wait 2 minutes for your next included review.

View limit details

Limit details: You’ve used all 5 included reviews currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 5fbcbb7a-fe9a-49a5-9c9d-2c68e2407a4e

📥 Commits

Reviewing files that changed from the base of the PR and between 69c6138 and e5dce16.

📒 Files selected for processing (1)
  • test/js/web/workers/worker-terminate-lifetime.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Aug 28, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: reproduced on a macOS x64 CI host with the CI binary (old fixture fails 10 of 10, new one passes 10 of 10, both before and after a host reboot). Self-review done: the per-worker exit offset was dropped as a no-op.

CI: worker-terminate-lifetime.test.ts passes on every lane in builds 107691 and 107742, including both macOS x64 test shards. The red in 107742 is unrelated and pre-existing on main: test/js/web/url/url.test.ts (macOS x64, the Unicode 16 IDNA case) and test/js/node/http2/h2-conformance.test.ts (macOS 27 aarch64); both are reported for triage. Ready for review.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM — clean move from a clock-driven exit to awaiting the actual condition the test tears down.

What was reviewed:

  • Template-literal nesting: \${J}, \${i}, \${(i * 13) % 60} resolve in the -e script scope; the unescaped ${timeout / 2} correctly resolves at test-file scope where timeout is 20_000/60_000.
  • Confirmed this is a wait-for-condition fix, not a timing shuffle to dodge a UAF — the awaited state (all J responses touched while request bodies still stream) is exactly the teardown state the test exists to exercise, and the PR verifies the unfixed ProducerHold still crashes the new fixture under ASAN.
  • fetch rejections now fail loud (exit 3) instead of being swallowed; the timeout / 2 fallback is a diagnostic hang-guard, not a condition wait.
Extended reasoning...

Overview

This PR touches a single test file, test/js/web/workers/worker-terminate-lifetime.test.ts, reworking one fixture ("worker exit with streaming-request-body fetches whose response.body was touched"). The old fixture had each worker start 8 duplex fetches and then process.exit after a fixed ~150ms + jitter, exiting 3 if no response had been touched. On slow macOS x64 CI hosts the 48 loopback connections take longer than 150ms to establish, so the worker exited before the state under test existed. The new fixture hoists concurrency into J = 8, exits once all J responses are touched (keeping the per-worker jitter so exits still land at varied points in the 5ms enqueue cadence), fails loudly on fetch rejection, and adds a timeout / 2 fallback that prints which worker stalled before exiting 3.

Security risks

None. Test-only change to a spawned fixture string. No new network endpoints, no external hosts, port: 0 preserved, no credentials or filesystem writes.

Level of scrutiny

Low-to-moderate. It's a test-only de-flake with no production code touched. The two things worth checking were (1) whether this is the "changing a test so GC runs at a different time" anti-pattern from CLAUDE.md rule 15, and (2) whether the nested template-literal interpolation is correct. On (1): the change awaits the exact condition the test is designed to tear down (touched response body + still-streaming request body), and the PR description documents that removing the ProducerHold ref still crashes the new fixture 3/3 under debug ASAN — so coverage is preserved and this is not papering over a real failure. On (2): the escaping is correct — \${J}/\${i}/\${server.url} evaluate in the -e script's scope where those bindings exist, and the one unescaped ${timeout / 2} evaluates in the test file where timeout (line 12, slow ? 60_000 : 20_000) is defined.

Other factors

The change directly implements the repo's own test guidance ("do not use setTimeout to wait for a condition; await the event itself"). The old .catch(() => {}) swallow is replaced with a logged exit(3), which aligns with REVIEW.md's "wire every failure event to fail the test." The fallback setTimeout at timeout / 2 is a diagnostic hang-guard (fires well before the test-runner timeout), not a condition wait. No CODEOWNERS entry covers this path, and there are no outstanding reviewer objections in the timeline.

The teardown state is complete at the eighth touched response. The
offset could not move a stream across any state boundary, and the exit
already lands at a random point in the 5 ms enqueue cadence.
Comment thread test/js/web/workers/worker-terminate-lifetime.test.ts

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant