Skip to content

ci(macos): drain the agent before the nightly wipe and reboot - #43611

Open
robobun wants to merge 4 commits into
mainfrom
robobun/3277a30b/darwin-nightly-cleanup-under-job
Open

robobun wants to merge 4 commits into
mainfrom
robobun/3277a30b/darwin-nightly-cleanup-under-job

Conversation

@robobun

@robobun robobun commented Sep 20, 2026 •

Copy link
Copy Markdown
Collaborator

Needs a maintainer decision first: #43630 (this is option 1, drain). #43629 stops the red builds without a rollout.

Problem

Fix

  • The cleanup script stops the agent service (launchctl kill SIGTERM), waits until no buildkite-agent process is left, then wipes and reboots. After two hours it reboots without the wipe. If the host is still up 600 s later, it starts the agent again.
  • agent.ts start forwards SIGTERM to buildkite-agent, then exits 0 however the agent ended, so launchd does not restart the service under the cleanup.
  • install waits for the old service to be gone between launchctl bootout and bootstrap.
  • Verified: test/internal/ci-agent.test.ts and the generated script against stubs (Notes). The launchd path was never run on a Mac. It needs a trial on one mini.

Background

  • A bare mini runs buildkite-agent as a child of node agent.ts start, the launchd service. launchd signals node, not the agent.
  • On its first SIGTERM buildkite-agent takes no new job, finishes its job and exits.
  • Since ci: content-addressed machine images #43608 every byte of scripts/agent.ts is part of the name of every CI image. This PR bakes all of them again.
Notes

What a merge changes, and where.

  • Bare Mac minis: nothing until someone runs install on them (see Deploy).
  • Linux and Windows agents: they get the new start with their next image bake, which this PR triggers. The systemd unit has KillMode=process, so systemctl stop sent SIGTERM to node only and left the agent running. With the forwarding it reaches buildkite-agent, which stops gracefully. The cloud agents are ephemeral (one job, then the machine is deleted), so nothing there stops the service in normal operation. On Windows a SIGTERM listener never fires.
  • Tart hosts: nothing. Their nightly (scripts/darwin-ci/lib/agent.ts:74-81) boots the agent out before it deletes anything.

On the hosts (read-only). last reboot on crouton shows 06:27 for every day. /Library/LaunchDaemons/com.buildkite.cleanup.plist is byte-identical (same sha1) on crouton, breadstick, hardtack, biscuit, matzo, cornbread, bagel and pretzel, and it is the template from scripts/agent.ts. The agent log of crouton for build 118830 (local time) shows that the agent already waits for its job on a first SIGTERM:

06:27:23 Received CTRL-C, send again to forcefully kill the agent(s)
06:27:23 Gracefully stopping agent. Waiting for current job to finish before disconnecting...
06:27:35 Process with PID: 26252 finished with Exit Status: 1, Signal: nil
06:27:36 Finished job 01a0bd3d-0dd3-4a6d-9c4e-2479b9612b64 for build at https://buildkite.com/bun/bun/builds/118830
06:27:36 Disconnected

The service on these hosts is node agent.mjs start (pid 341 on crouton) with buildkite-agent start as its child (pid 400). pgrep -x buildkite-agent finds the child. launchctl print system/buildkite-agent reports exit timeout = 5.

The wipe set. The old script also removed $BASE_PREFIX/{var,etc}/buildkite-agent/{builds,cache}/* and <home>/cache/*, then ran chown and chmod on the Homebrew directories. None of those paths exists on the eight minis above: every bare mini uses the ~/Library/Services/buildkite-agent layout that install creates, and the agent cache is ~/Library/Caches/buildkite-agent, which the old script did not touch either. The new script removes builds, /tmp/* and /var/tmp/*. It removes builds itself and not builds/*, so that a symlink at builds is not followed. The agent creates the directory again for its next job.

The two hour bound. Darwin test jobs have timeout_in_minutes: 45, so a drain takes less than that. The bound is for an agent that does not exit. In that case the checkout stays, and the reboot ends the job as a lost agent. A reboot also ends the script. If the script is still alive 600 s after it asked for the reboot, it starts the agent again (launchctl kickstart), so a host never stays in the pool without an agent. The script writes to ~/Library/Logs/buildkite-agent/cleanup.log, one line on a normal night.

A stop always exits 0. launchd restarts this service when it exits with anything but 0 (KeepAlive, SuccessfulExit=false). A restarted agent would take jobs while the cleanup waits for the old one. So after a forwarded signal run() returns whatever the exit status of the agent is, and start exits 0. Without a signal, a failed agent is still an error and launchd restarts it.

install on a host that already runs this. launchctl bootout returns before a service that stops gracefully is gone, and bootstrap fails with EIO until it is. install polls launchctl print system/<label> between the two (at most 60 s), as scripts/darwin-ci/lib/agent.ts does for the Tart agent.

Checks of the cleanup script. The script that install generates was extracted from agent.ts and run under /bin/sh with a PATH that holds only stubs that log their arguments. Six cases: an idle agent, a job that ends after three polls, an agent that never exits (240 polls, no rm), no agent at all, a failed shutdown (it runs reboot), and a failed shutdown and reboot. Nothing ends the script in this harness, so every case also reaches sleep 600 and the kickstart. sh -n, dash -n and bash --posix -n accept it. The generated plist parses, and its CDATA section gives back the script byte for byte. This harness is not in the PR: a similar test file was removed from #40349 by a maintainer.

The test. run() is checked with a service process that gets SIGTERM, for a command that then exits 0, exits 1, or dies of the signal: the service exits 0 each time. With run() from main the service dies of the signal (exitCode 143) and the command never sees it. A fourth case checks that a command that fails with no signal is still an error.

Deploy. On each bare mini, from a checkout of main with this change, while the agent is idle (install boots the service out, which ends a running job):

sudo env PATH=$PATH node scripts/agent.ts install

It reuses the token from the existing buildkite-agent.cfg. darwin-ci provision <host> bare ends in the same install (scripts/darwin-ci/lib/agent.ts:102-106), so a provision from a ref without this change writes the old daemon again. This is also the first time the minis run the single-file agent.mts from #43595: they still run the agent.mjs and utils.mjs copies from 2026-08-11. The pass needs an owner, and a trial on one mini first. I can run both over ssh if a maintainer says yes in #43630.

History. This PR first carried the runner change too. A review asked for the split, because the runner part needs no rollout and no decision. Supersedes #40349, which predates the TypeScript conversion in #43595. The self-review raised five points: the split, getUser() outside agent.ts, the test location, a recorded maintainer decision with a rollout owner and a trial, and the missing note about the Linux agents. All five are addressed here, in #43629 and in #43630.


no test proof · iteration 0 · platform-specific test(s) that do not run on this machine, deferring to CI, which covers all platforms: test/internal/ci-agent.test.ts

…ep the runner alive when a host goes down

The com.buildkite.cleanup daemon of the bare macOS agents removed builds/*
and rebooted at 06:27 local time, also under a running job. The script now
stops the agent service first, waits for the agent to exit (it finishes its
job and exits 0), and only then wipes. After two hours it reboots without
the wipe. `agent.ts start` passes SIGTERM on to buildkite-agent, which is
what makes the stop a graceful one.

The test runner looks the user up once instead of once per test file, and a
node test file that is gone since the tests were listed fails its own step.
Both were throws that ended a shard with exit 1 once its host had begun to
shut down, which Buildkite records as a test failure and does not retry. The
shard now ends as a lost agent, which is retried.
@robobun

robobun commented Sep 20, 2026 •

Copy link
Copy Markdown
Collaborator Author
Updated 11:57 PM PT - Sep 19th, 2026

❌ @robobun, your commit 01a20d4 has 1 failures in Build #118851 (All Failures):


🧪   To try this PR locally:

bunx bun-pr 43611

That installs a local version of the PR into your bun-43611 executable, so you can run:

bun-43611 --bun

@robobun

robobun commented Sep 20, 2026 •

Copy link
Copy Markdown
Collaborator Author

Status: waits for a maintainer decision in #43630. The code is ready for review at 143a6f3.

This PR is now the daemon change only (scripts/agent.ts). The runner change that stops the red builds moved to #43629, which needs no rollout.

Reproduction: the job log of build 118828 (host biscuit) shows Test filter ".../import-custom-condition.test.ts" had no matches at 05:27:02 UTC and spawn .../bun-profile ENOENT at 05:27:14. Builds 118829 (breadstick) and 118830 (crouton) ended the same way in the same minute. On crouton, last reboot shows 06:27 local time for every day, and the deployed com.buildkite.cleanup.plist is the template from scripts/agent.ts.

CI: this PR changes scripts/agent.ts, so its build has to bake every CI image again (#43608). In build 118947 the Debian, Alpine and Windows images were baked from this branch, their agents ran the new start, and every test lane on them passed except test/js/bun/spawn/spawn.test.ts on debian x64-asan, which also fails on main. The two Ubuntu bakes fail before a machine exists: InvalidUserID.Malformed: Invalid user id: "99720109477". The YAML writer in .buildkite/ci.ts writes the owner id 099720109477 without quotes, and the leading zero is lost. That is a bug on main, not in this diff, and a push here cannot get past it until it is fixed. The previous build of this branch (118851, before #43608) passed on every lane, Ubuntu and the darwin minis included, except test/bake/deinitialization.test.ts on alpine aarch64, which this diff does not touch.

The agent change takes effect on a bare mini only after sudo env PATH=$PATH node scripts/agent.ts install on that host while it is idle.

@coderabbitai

coderabbitai Bot commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

Review Change StackReview Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: oven-sh/bun/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: f4b24d1a-56f7-4f91-8bed-0d02b8ea329d

📥 Commits

Reviewing files that changed from the base of the PR and between 01a20d4 and 143a6f3.

📒 Files selected for processing (2)
  • scripts/agent.ts
  • test/internal/ci-agent.test.ts

Included review availability: Your plan provides up to 10 included reviews per hour; 5 remain after this review.


Walkthrough

The agent now forwards configured termination signals to child commands. macOS cleanup waits for agent exit before removing files and rebooting. Service reinstall waits for removed services. Integration tests cover signal forwarding and exit-code failures.

Changes

Agent runtime behavior

Layer / File(s) Summary
Signal forwarding and process results
scripts/agent.ts, test/internal/ci-agent.test.ts
run() forwards configured signals and removes listeners after completion. start() forwards SIGTERM. Tests cover normal, failing, and signal-terminated child commands.
macOS cleanup and service reinstall
scripts/agent.ts
Cleanup waits for agent exit before removing paths, logs output to cleanup.log, and restarts after reboot. Service reinstall waits for booted-out services to disappear.

Suggested reviewers: dylan-conway

Priority: ➖ Normal

🚥 Pre-merge checks | ✅ 4
✅ Passed checks (4 passed)
Check name Status Explanation
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Title check ✅ Passed The title clearly and concisely describes the main change: gracefully draining the macOS agent before the nightly wipe and reboot.
Description check ✅ Passed The description explains the problem, fix, scope, deployment steps, and verification results. It does not use the template headings exactly, but it provides the required information in equivalent sect…

Comment @coderabbitai help to get the list of available commands.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings marked 🟡 are optional suggestions and need no follow-up push.

Additional findings (outside the current diff — GitHub can't attach inline comments there):

  • 🔴 scripts/agent.ts — An operator re-running install on a Mac whose agent service is running now gets a failed install and a box with no agent, where the base branch replaced the service in place. launchctl bootout (scripts/agent.ts:972) sends SIGTERM to the start process, which since scripts/agent.ts:1129 forwards it and stays alive until buildkite-agent has disconnected; bootout returns before that, and the immediate launchctl bootstrap at scripts/agent.ts:973 fails with EIO, so run() throws and the cleanup plist is never (re)loaded. Fix: after bootout, wait until launchctl print system/buildkite-agent fails before bootstrapping (as scripts/darwin-ci/lib/agent.ts:113-116 does), for both plists in the loop. [also at: scripts/agent.ts:1129 - Operators re-running install on a mini that already runs this script can now get a failed deploy that leaves the agent service unloaded.]

    Extended reasoning...

    scripts/darwin-ci/lib/agent.ts:112 records the launchd behaviour this depends on: "bootout returns before a busy agent has finished its cancel grace period, and bootstrap fails with EIO until it has", and its unload() polls launchctl print for that reason. On the base branch the service's main process (node running start) had no SIGTERM listener, so bootout's SIGTERM killed node at once, launchd killed the process group, and the service was gone before bootstrap ran. With this change run() installs process.on("SIGTERM", () => child.kill("SIGTERM")) (scripts/agent.ts:75-79) and keeps waiting for close. buildkite-agent answers the forwarded SIGTERM by stopping gracefully: it logs "Gracefully stopping agent", calls the Buildkite API to disconnect, then exits; even idle that takes a network round trip, and with a job it takes until the job ends or launchd's 20 s ExitTimeOut SIGKILL. install() runs run(["launchctl", "bootstrap", "system", p]) (scripts/agent.ts:973) milliseconds after…

    Verification: normal — triggered whenever an operator re-runs sudo node scripts/agent.ts install on a bare mini whose service is already running this PR's start (the PR's own documented redeploy step, repeated for any future update; also installBareAgent() in scripts/darwin-ci/lib/agent.ts:105). Recovery is a second install run (bootout then fails harmlessly and bootstrap succeeds), but the first run…

Comment thread scripts/agent.ts
Comment thread scripts/agent.ts Outdated
…however the agent ended, start the agent again if the reboot does not come

`launchctl bootout` returns before a service that stops gracefully is gone,
and `bootstrap` fails with EIO until it is, so `install` polls
`launchctl print` between the two, as scripts/darwin-ci does.

After a forwarded signal run() no longer throws on the command's exit
status. `start` then exits 0, so launchd (KeepAlive, SuccessfulExit=false)
does not start a new agent under the cleanup that waits for the old one.

The cleanup script starts the agent again if it is still running ten
minutes after it asked for the reboot.
@robobun

robobun commented Sep 20, 2026

Copy link
Copy Markdown
Collaborator Author

On the finding about install outside the diff (bootout then bootstrap): correct, and fixed in 01a20d4. With the SIGTERM forwarding, the start process no longer dies at once on launchctl bootout, so the service can still be on its way out when bootstrap runs. install now polls launchctl print system/<label> until the service is gone (at most 60 s) before it bootstraps each plist. scripts/darwin-ci/lib/agent.ts does the same for the Tart agent and records the EIO. I could not run this against launchd: it needs a write on a Mac mini, and I do not make one without a yes from a maintainer.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I re-reviewed after 01a20d4 and found no bugs; the two points from my earlier inline comments are addressed in the code (run() returns normally after a forwarded signal so the service exits 0 whatever the agent's status, and the cleanup script kickstarts the agent if the host is still up 600 s after asking for the reboot). Because this rewrites a fleet LaunchDaemon, changes how systemctl stop/launchctl stops reach buildkite-agent, and needs a manual install on each bare mini, a human infra owner should still sign off on the rollout.

What was reviewed:

  • run() signal forwarding: listeners are removed in .finally, asked only short-circuits the throw when a forwarded signal was actually delivered, and the non-forwarded exit-7 path still throws.
  • Double-signal risk on Linux: the systemd unit uses KillMode=process, so only node gets SIGTERM and the forward does not make the agent see a second (forceful) SIGTERM.
  • getUser(): the only remaining userInfo() caller in scripts/ is the memoized wrapper itself; runner.node.ts and buildkite.ts both use it.
  • install bootout/bootstrap loop: bounded to 60 x 1 s, well above launchd's default 20 s ExitTimeOut, and the ENOENT-only rethrow in runner.node.ts keeps other read errors loud.
Extended reasoning...

Overview

The PR touches only CI tooling: scripts/agent.ts (a forward option on run(), a memoized getUser(), a rewritten macOS nightly-cleanup LaunchDaemon script, and a bootout-wait-bootstrap loop in install), the two callers in scripts/buildkite.ts and scripts/runner.node.ts that switch to getUser(), an ENOENT-tolerant readFileSync in the runner's node-test loop, and a new test/internal/ci-agent.test.ts. No runtime, native, or user-facing code changes. The second commit (01a20d4) is a direct response to the two inline findings from the prior run: run() no longer throws on a non-zero/signal exit once a forwarded signal was delivered, and the cleanup script restarts the agent after a 600 s wait if the reboot never came.

Security risks

None that reach users. The cleanup script runs as root and does rm -rf "${homePath}/builds" /tmp/* /var/tmp/*; homePath is derived from the install-time paths in agent.ts, not from external input, and the script now removes builds itself rather than builds/*, so a symlink there is not followed. The generated shell is embedded in a plist CDATA block with fixed content. The getUser() cache is a process-lifetime memo of os.userInfo(), which cannot change meaningfully for a running service.

Level of scrutiny

Medium. The code itself is small and I could trace every branch of run() and the cleanup script by reading. What warrants a human is the operational surface: the LaunchDaemon on every bare Mac mini is replaced, systemctl stop on the Linux agents now reaches buildkite-agent (previously only node was signalled, and KillMode=process meant the agent kept running), and the PR itself states that the agent half only takes effect after someone runs sudo node scripts/agent.ts install on each idle mini. Those are fleet decisions, not code-correctness questions, and CODEOWNERS does not cover scripts/, so nobody is auto-assigned.

Other factors

The bug hunt ended on a dry streak and the single ruled-out candidate (the bare catch {} around the cleanup process.kill in the test) is a nit on a best-effort teardown of an already-dead PID. I could not execute the new test in this environment (shell execution was denied), so the claim that the three forwarding variants pass rests on the PR description and on reading the fixture logic; the fixture matrix (exit 0, exit 1, death by signal, and the non-forwarded exit-7 throw) does cover each load-bearing clause of the run() change. The test file is new rather than appended to an existing one and uses *.fixture.mjs names inside a tempDir, which does not affect the test runner's discovery; the USE_SYSTEM_BUN=1 rule is not meaningful here because the code under test is plain TypeScript in scripts/, not a native change.

The cached user lookup and the node test read are in #43629, with getUser()
in scripts/buildkite.ts. This branch now changes scripts/agent.ts only: the
cleanup daemon, the SIGTERM forwarding in `start`, and the wait between
bootout and bootstrap in `install`.

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code review found no issues

No high-confidence issues detected in this change.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant