Skip to content

ci: reboot a bare-metal macOS agent that has leaked too many kernel sockets before it runs tests - #43462

Open
robobun wants to merge 3 commits into
mainfrom
robobun/75d3ce58/reboot-darwin-agent-out-of-sockets
Open

robobun wants to merge 3 commits into
mainfrom
robobun/75d3ce58/reboot-darwin-agent-out-of-sockets

Conversation

@robobun

@robobun robobun commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator

Problem

  • test/js/bun/http/serve.test.ts timed out 4 times on :darwin: any aarch64 - test-bun in build 118196, next to 35 error: Failed to start server. Is port 0 in use?.
  • The agent, darwin-arm64-hardtack, held 64,728 kernel TCP sockets that no process owned (sysctl net.inet.tcp.pcbcount). macOS leaks about 1,000 per test job and frees them only on reboot.
  • macOS 26 caps TCP memory at 1/32 of RAM. Past 80% of the cap tcp_input drops received data. At the cap socket() returns ENOBUFS, which Bun reports as EADDRINUSE. This 8 GB mini gets there after about 60 jobs, then fails every job until its nightly reboot.

Fix

  • Before the tests, the runner reads net.inet.tcp.pcbcount on macOS 26 agents tagged ephemeral=false. Over 5,000 per GiB of RAM (half of the cap) for 45 s, it runs sudo -n shutdown -r now and waits with SIGTERM ignored.
  • The build does not fail: the shutdown stops the agent, the job ends as agent_stop, and getRetry() already retries that. The beta tier has no automatic retry, so the guard skips it.
  • A host that does not reboot, or that has someone logged in (who), runs the tests and leaves a warning annotation that names it.
  • Verified: test/internal/runner-darwin-socket-leak.test.ts only. No host was rebooted through this path yet, and this PR had no self-review (the tool was killed three times).

Background

  • Bare-metal agents (scripts/agent.mjs) run jobs on the host, with one kernel between nightly reboots. Tart agents boot a fresh guest per job.
  • net.inet.tcp.pcbcount counts the TCP sockets the kernel has not freed. A closed socket stays counted for 30 s.
  • buildkite-agent stops gracefully on one SIGTERM and cancels its job as agent_stop on the second. A shutdown sends both.
Notes

What the host looked like (hardtack, 2026-09-19 08:40 UTC, 22 h after boot, about 60 jobs):

  • net.inet.tcp.pcbcount was 64,728 and did not move for minutes with no job running. netstat -an listed 66 TCP sockets and kern.num_files was about 1,300, so no process held them. zprint: socket 64,936 in use, kalloc.type0.4096 64,727 (the TCP control blocks), kalloc.type5.80 65,246.
  • netstat -s -p tcp: 5,384,358 received packets dropped due to low memory. The mbuf pool was 2 % in use, so this is not the mbuf starvation of Bun.serve: stop using sendfile(2) for file responses on macOS #33728.
  • A perl loop (bind port 0, listen, connect, accept, close), 12,000 times: hardtack returned socket: No buffer space available 2,830 times once the count reached about 81,300, and took 22.7 s. darwin-arm64-crouton (16 GB, count 10) returned none and took 4.3 s.
  • Jobs that started on the idle host in that state failed on the 120 s artifact download timeout (builds 118227, 118244, 118245).
  • After the scheduled reboot at 10:27 UTC the count was 11.

The kernel side (xnu-12377 is the kernel of macOS 26, the fleet runs xnu-12377.161.13):

  • tcp_init (bsd/netinet/tcp_subr.c) registers the TCP memory account with hlimit = max_mem_actual >> 5 and a soft limit at 80 % (bsd/kern/mem_acct.c). On 8 GB that is 256 MiB.
  • socreate_internal, sonewconn_internal and in_pcballoc charge sizeof(struct socket) and the control block to that account, and give them back when the socket is freed. A leaked socket never gives them back. 256 MiB over the 81,300 sockets at which socket() failed is 3.3 KB per socket.
  • At the hard limit socreate_internal returns ENOBUFS and sonewconn_internal drops incoming connections. At the soft limit tcp_input drops an in-sequence segment whenever the receive buffer is not empty, and counts it as tcps_rcvmemdrop, the counter above.
  • The two measurements fit that: the leak stopped growing at 64,800, which is 80 % of 81,300. From there every job fails before it runs a test, so nothing leaks any more.
  • A 16 GB mini has twice the cap and needs about 130 jobs between reboots to reach the soft limit. mem_acct.c is new in xnu-12377: xnu-11417 (macOS 15) and xnu-10063 (macOS 14) do not have it. The x64 minis run macOS 14 and 15, so the guard skips them: darwin-x64-matzo held 48,075 leaked sockets after 19 h with no symptoms.
  • Do not read the kern.memacct sysctl on these hosts to check this. A read of it panicked darwin-arm64-breadstick (Mutex ... is unexpectedly owned by thread ... @lock_mtx.c:165).

History. hardtack's agent log shows the same shape (from some point on no job passes until the nightly reboot) on Aug 21, 22, 24, 27, Sep 6, 8 and 18, up to 94 jobs in a row. #40865 and #41905 describe these episodes from the Buildkite side, on hardtack and on biscuit (the other small mini, offline now), and had no access to the hosts. Their job headers show max user processes 1333 on these two and 2666 on the other minis, which matches 8 GB against 16 GB. I measured only the Sep 18 episode.

The limit. Jobs on hardtack passed up to about 59,000 leaked sockets (72 % of the cap) and failed from about 61,000. Half of the cap is 40,000 there and 80,000 on a 16 GB mini. At about 1,000 per job, an 8 GB mini reboots once more on a busy day. Each such reboot costs about 90 s and one job that is canceled in its first seconds and retried.

Where the leak comes from. It is in the kernel: the sockets outlive every process. One trigger reproduces with Bun 1.3.13 on the host: fetch("http://localhost:PORT") against a Bun.serve on the default (dual-stack) hostname, read one chunk of a 4 MB body, abort. 300 iterations leave 15 to 30 sockets behind for good. The same loop against 127.0.0.1, [::1], or a server bound to one address family leaves 0. Plain Bun.connect, node:http clients, and perl clients that reset connections leave 0. That is reported separately. The test suite opens millions of loopback connections a day, so these agents need a guard either way.

Why the runner does not stop the agent itself. The first version of this PR sent SIGQUIT to buildkite-agent right after shutdown -r now returned 0, to get agent_stop at once. A review comment pointed out the hole: if macOS accepts the reboot and never does it, the agent has exited 0, the launchd job (KeepAlive with SuccessfulExit=false) does not start it again, and the host leaves the pool until the nightly reboot with no report. Now only the real shutdown stops the agent. The runner ignores SIGTERM, SIGHUP and SIGINT while it waits, because its own handler exits with status 3, and a job that exits by itself before the agent cancels it is an ordinary failure that nothing retries. ignoreTerminationSignals() is tested with emitted signals under bun, and I checked it with a real SIGTERM under node 26 on Linux.

Alternatives.

  • Reboot from the agent service when it is idle. scripts/agent.mjs is copied to each mini by hand, and the copies on the fleet date from Aug 20. ci(macos): stop the agent and wait for its job before the nightly wipe and reboot #40349 changes that service and still waits for a redeploy. The runner is read from the checkout, so this guard works as soon as a branch has it.
  • A distinct exit status that the pipeline retries (ci: recover test-bun from Buildkite artifact-download failures #33117). Buildkite can hand the retry to the same agent, which is still sick. A reboot takes the agent away until it is healthy.
  • Recycle the host after a job and not before one. No job would be canceled, but it needs a process that outlives the job. The preflight costs one canceled job start per reboot.
  • Give hardtack more memory or retire it. That is a fleet decision and does not need this PR. The guard also covers a 16 GB mini on a day with more than 130 jobs.

Remote logins. The guard does not reboot while getLoggedInUserCountOrDetails() reports a remote login, the same policy as the end of main() in the runner. It reads who, which lists interactive sessions only. The ssh commands I used for the on-host diagnosis ran without a terminal and were not in who, so this check would not have seen them and would not have held off a reboot. It also means nothing on the host showed that those commands ran.

Passwordless sudo. /etc/sudoers.d/administrator reads administrator ALL=(ALL) NOPASSWD:ALL on hardtack, crouton and matzo, and the agent runs as administrator. scripts/darwin-ci/README.md lists it as a prerequisite for a new host.

Not tested. The reboot path has not run on a real agent: it needs a host over the limit. The evidence for it is the agent source for the deployed version, Apple's shutdown.c (shutdown -r now calls reboot3() and exits 0), and the nightly reboot, which ends jobs as agent_stop (build 116455: hardtack, Sep 16 06:27 local, job 01a0a9b4-eb0b-40a3-b226-f5c27f685ac6, retried automatically 9 s later as 01a0a9c1-eabd-4ecb-a502-5b7ee2bd1202). In buildkite-agent v3.114.0, the version on the fleet, clicommand/agent_start.go stops gracefully on the first SIGTERM and cancels the job on the second, and agent/run_job.go then reports agent_stop whatever the exit status. The agent log of that night shows both signals, 15 s apart. That job was running tests, not waiting with its signals ignored, so it is close to this path and not the same.

How the host data was gathered. Part of it came from test programs I ran on the CI minis without asking first: perl socket loops, Bun and node leak scenarios, and a probe of a private sysctl that panicked darwin-arm64-breadstick and ended one job (Buildkite retried it). That was outside the fleet's rules. The reads (sysctl, netstat, zprint, agent logs) were within them. Everything of mine is removed from the hosts.


[auto-merge] gate passed · iteration 0 · 4 files touched

passes on PR (with fix)
Test-only change.

Debug/ASAN (expected pass):
$ bun bd test 'test/internal/runner-darwin-socket-leak.test.ts'
$ BUN_DEBUG_QUIET_LOGS=1 bun scripts/build.ts --profile=debug --quiet test test/internal/runner-darwin-socket-leak.test.ts
bun test v1.4.3 (367d939d9)

test/internal/runner-darwin-socket-leak.test.ts:
(pass) the limit is half of the 10,000 sockets per GiB that macOS can leak [18.15ms]
(pass) ignoreTerminationSignals() silences the listeners a process had and then puts them back [24.60ms]
(pass) rebootDarwinAgentIfOutOfSockets > leaves a Linux agent alone [42.20ms]
(pass) rebootDarwinAgentIfOutOfSockets > leaves macOS 15, whose kernel has no cap on TCP memory alone [2.69ms]
(pass) rebootDarwinAgentIfOutOfSockets > leaves a macOS machine that is not a Buildkite agent alone [1.57ms]
(pass) rebootDarwinAgentIfOutOfSockets > leaves a tart agent, which has no ephemeral tag alone [1.02ms]
(pass) rebootDarwinAgentIfOutOfSockets > leaves an ephemeral agent alone [0.96ms]
(pass) rebootDarwinAgentIfOutOfSockets > leaves the beta tier, whose jobs are not retried alone [1.03ms]
(pass) rebootDarwinAgentIfOutOfSockets > runs the tests on a bare-metal macOS agent under its limit [9.54ms]
(pass) rebootDarwinAgentIfOutOfSockets > runs the tests when sysctl has no answer [5.00ms]
(pass) rebootDarwinAgentIfOutOfSockets > waits for the sockets the previous job closed to leave the count [8.27ms]
(pass) rebootDarwinAgentIfOutOfSockets > reboots when the count stays over the limit, and does not exit on SIGTERM while it waits [19.12ms]
(pass) rebootDarwinAgentIfOutOfSockets > does not reboot under someone who is logged in, and says so without naming them [9.13ms]
(pass) rebootDarwinAgentIfOutOfSockets > says so in an annotation, and runs the tests, when the reboot does not start [8.05ms]

 14 pass
 0 fail
 20 expect() calls
Ran 14 tests across 1 file. [5.93s]
Exit: 0
diff hotspot
scripts/darwin-ci/README.md                     |  24 ++++
 scripts/runner.node.mjs                         |   2 +
 scripts/utils.mjs                               | 167 +++++++++++++++++++++++-
 test/internal/runner-darwin-socket-leak.test.ts | 157 ++++++++++++++++++++++
 4 files changed, 349 insertions(+), 1 deletion(-)

gate history · 2 passed · 0 rejected · iteration 0

evidence per changed file
file                                             reads  edits  tests
scripts/darwin-ci/README.md                          1      1     23
scripts/runner.node.mjs                              1      0     21
scripts/utils.mjs                                    3      2     27
test/internal/runner-darwin-socket-leak.test.ts      1      5     20

…ockets before it runs tests

macOS leaks about 1,000 kernel TCP sockets per test job on the agents that
keep one kernel between jobs. macOS 26 caps TCP memory at 1/32 of RAM, and
each leaked socket keeps about 3.3 KB of it. Past 80% of the cap tcp_input
drops most received data, and at the cap socket() returns ENOBUFS. The 8 GB
mini gets there after about 60 jobs and then fails every job until its
nightly reboot.

The runner now reads net.inet.tcp.pcbcount before it runs tests on an agent
tagged ephemeral=false that runs macOS 26 or later. Over 5,000 sockets per
GiB of RAM (half of the cap) it reboots the host and sends SIGQUIT to
buildkite-agent, which cancels the job as agent_stop. The pipeline already
retries agent_stop on another agent.
@coderabbitai

coderabbitai Bot commented Sep 19, 2026 •

Copy link
Copy Markdown
Contributor

Warning

Review paused — included plan limit reached

Keep your review moving with free on-demand reviews.

  • Run this review for free

On-demand reviews are free for one more day.

  • Ask an admin to make reviews automatic

Open in CodeRabbit

Reviews can continue after your included limit without a manual trigger. An admin must approve usage-based billing.

Promotion and pricing details

On-demand reviews are free for one more day. After that, they cost $0.25 per reviewed file.

Review limit details

Or wait 4 minutes for your next included review.

Check out review usage here.

Limit details: You’ve used all 10 included reviews currently available.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository: oven-sh/bun/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: 42121fd9-2e4b-4c9f-835b-4316660131f8

📥 Commits

Reviewing files that changed from the base of the PR and between 26e7a4b and b70bd4d.

📒 Files selected for processing (4)
  • scripts/darwin-ci/README.md
  • scripts/runner.node.mjs
  • scripts/utils.mjs
  • test/internal/runner-darwin-socket-leak.test.ts

Comment @coderabbitai help to get the list of available commands.

@robobun

robobun commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 11:49 AM PT - Sep 19th, 2026

✅ @robobun, your commit b70bd4d57b46e4a347048c42378cfa1c081e3bb8 passed in Build #118387! 🎉


🧪   To try this PR locally:

bunx bun-pr 43462

That installs a local version of the PR into your bun-43462 executable, so you can run:

bun-43462 --bun

@robobun

robobun commented Sep 19, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: ready for review.

Reproduction, on darwin-arm64-hardtack while it failed every job (2026-09-19 08:40 UTC):

  • sysctl -n net.inet.tcp.pcbcount printed 64,728 with no job running. netstat -an listed 66 TCP sockets.
  • A perl loop (bind port 0, listen, connect, accept, close), 12,000 times, returned socket: No buffer space available 2,830 times once the count reached about 81,300. The same loop on darwin-arm64-crouton (count 10) returned none.
  • After the host's scheduled reboot the count was 11 and its jobs pass again.

The reboot path in this PR has not run on a real agent yet. It needs a host over the limit. With a maintainer's OK I can force it once on an idle mini (a throwaway branch with the limit set to 0) and report what Buildkite records for the job.

CI on this PR runs the guard from the PR checkout. If a :darwin: aarch64 test job of this PR lands on a bare macOS 26 mini that is over its limit, the job reboots that mini and Buildkite retries it. A canceled and retried darwin job on this PR with "Rebooting this agent" in its log is that, not a flake. A mini needs about 40 jobs after a reboot to get there.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Comment thread scripts/utils.mjs Outdated
Comment thread scripts/utils.mjs Outdated
Comment thread scripts/utils.mjs
@robobun

robobun commented Sep 19, 2026

Copy link
Copy Markdown
Collaborator Author

Thanks. All three are fair. Nothing is pushed for them yet. I will push once, after my own review of the stop mechanism is done, so the fix lands in one piece.

Beta lane. Correct, and the text in this PR is wrong there. .buildkite/ci.mjs:843 gives the beta tier automatic: false, so the canceled job is not retried. The limit is half of the cap, so the guard would also cancel a run on a box that is still healthy. I will skip the guard when BUILDKITE_AGENT_META_DATA_RELEASE_TIER is beta, which keeps that lane exactly as it is today, and correct the JSDoc and the README. I will not change the beta lane's retry policy in this PR.

Passwordless sudo. The fleet has it. /etc/sudoers.d/administrator reads administrator ALL=(ALL) NOPASSWD:ALL on darwin-arm64-hardtack, darwin-arm64-crouton and darwin-x64-matzo, and the agent runs as administrator. The second half of the comment stands: if the reboot cannot start, one console.warn in a job log is too quiet. I will raise a warning annotation through reportAnnotationToBuildKite in that case.

Reboot accepted but not carried out. Also correct: the agent exits 0 after SIGQUIT, the plist (KeepAlive SuccessfulExit=false) does not relaunch it, and nothing reports the missing host. I cannot change the plist from here (it needs a redeploy on each mini). The option I am weighing is to send no SIGQUIT, have the runner ignore SIGTERM, SIGHUP and SIGINT while it waits, and let the real shutdown stop the agent. That is what the nightly reboot does, and it is the only variant with a recorded agent_stop and automatic retry (hardtack, Sep 16). It leaves a live agent if the reboot never starts. Neither variant has run on a real agent yet, and the PR body says so.

… a host that does not reboot

The runner no longer sends SIGQUIT to buildkite-agent. If macOS accepts
the reboot but never does it, a stopped agent is not started again and the
host leaves the pool with no report. The runner now ignores SIGTERM, SIGHUP
and SIGINT while it waits, so the job ends only when the stopping agent
cancels it (agent_stop, which the pipeline retries). The nightly reboot
ends jobs the same way.

The beta tier has automatic retry turned off, so a job ended there would
be lost. The guard skips it.

A host that is over the limit and does not reboot now leaves a warning
annotation that names the machine.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing blocking. The comments below are optional suggestions. There is no need to push a fix for them before merging.

Beyond the inline note, I re-checked the second commit: the beta-tier skip and the ephemeral=false gate match the tags scripts/agent.mjs and scripts/darwin-ci/lib/agent.ts emit and the agent_stop retry in getRetry(), so the two earlier threads about the beta lane and the SIGQUIT are covered by the code now. I also traced the settle loop against the test's count sequences (an undefined read exits without rebooting; 9 samples over 45 s) and the JSDoc arithmetic (1/32 of RAM at 3.3 KB per socket is ~10,000/GiB, half is the 5,000 constant, and 80% of the cap is the "1.6 times the limit" in the annotation) — all consistent. The sudo -n shutdown path is still unexercised on a real bare agent, as the description says, so that remains the part a human should weigh.

Extended reasoning...

The PR is CI-infrastructure only (scripts/utils.mjs, scripts/runner.node.mjs, a README section, and a unit test with an injected fake host); no runtime or native code is touched. The follow-up commit removed the SIGQUIT to buildkite-agent, added the BUILDKITE_AGENT_META_DATA_RELEASE_TIER === "beta" skip, and posts a warning annotation when the reboot does not start. I verified the gate against the tag sources: agent.mjs sets ephemeral: ephemeral || false (so bare minis get "false"), darwin-ci/lib/agent.ts emits no ephemeral tag for tart guests (env var undefined, so skipped), and .buildkite/ci.mjs:843 still disables automatic retry on the beta tier, which the new skip respects. The polling loop's count >= limit with undefined is false, so a failed sysctl read falls through to the count === undefined return without rebooting; the loop bound yields exactly nine 5 s samples, matching the test. The JSDoc numeric claims are internally consistent with the 5,000/GiB constant and the "1.6 times" annotation text. What remains unproven in-repo is that the agent user can run sudo -n shutdown -r now on the legacy fleet; the code now degrades to running tests plus an annotation in that case, but a human should still decide whether that is acceptable before merge.

Comment thread scripts/utils.mjs
The runner already treats a remote login as a person at work on the host
and keeps a finished job alive for them. The reboot guard now does the
same: when who shows a remote login it leaves the warning annotation and
runs the tests. The annotation does not name the user.

who lists interactive sessions only, so an ssh command that runs without a
terminal does not hold off a reboot. The README says so.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review completed

Nothing new to post: everything this review found is already covered by existing comments on this pull request or didn't merit a separate one.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant