Skip to content

macOS: run the LaunchDaemon as a Standard job, not Background, so the engine's first Metal compile finishes inside ollama's GPU-discovery window (#1521) - #1525

Merged
gen16k merged 1 commit into
mainfrom
fix/1521-discovery-restart
Sep 21, 2026
Merged

gen16k merged 1 commit into
mainfrom
fix/1521-discovery-restart

Conversation

@gen16k

@gen16k gen16k commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #1521

Why

On macOS, the first engine start with a cold Metal shader cache lost the GPU. This happens after the bundled ollama moves to a new version, and whenever the cache is cold. ollama's own GPU discovery was cut off by its fixed 30 s watchdog (llama-server GPU discovery watchdog timed out), and the engine then ran as a CPU host until its next restart (total_vram="0 B", mmap off).

Cause. The LaunchDaemon plist said ProcessType=Background, and every process the daemon spawns inherits that class. So ollama serve and every llama-server ran at background QoS. A llama-server's first start has macOS's MTLCompilerService compile ggml's Metal shaders at the requester's QoS. Under Background that compile outlasted the 30 s window.

What I measured on an M4 Mac mini, from a cold cache:

  • The same bundle, started by hand from a shell, discovered the GPU in 12 s.
  • Started as the product does, discovery was cut off 3 times out of 3.
  • ollama kills the discovery process at 30 s, and the compile it had started is thrown away. So restarting straight away, which is what this PR first did, is cut off the same way (tried on hardware and dropped).

Where Background came from. It dates from the repository's first import, with the reason "tells App Nap to leave us alone". App Nap is a mechanism for apps, not for a LaunchDaemon. The Linux systemd unit and the Windows service set no lower priority.

What changed

Measurements

The product's own start, 3 runs per condition, median unless noted. No other load was running (sharing off). Power is the mean of 5 × 1 s powermetrics samples taken from 3 s into a 1024-token generation.

M4 Mac mini 16 GB (qwen3.5-4b): Background → Standard M5 Pro MacBook Pro 48 GB, on AC (qwen3.6-35b-a3b): Background → Standard
GPU discovery, cold cache 30.4 s, cut off 3/3 → 14.2 s, 0/3 30.3 s, 3/3 → 12.1 s, 0/3
Model load, cold cache 121.5 s → 24.3 s 103.1 s → 24.3 s
Model load, warm 7.36 s → 2.37 s 7.22 s → 1.84 s
Generation 26.0 → 26.8 tok/s 59.8 → 77.7 tok/s
Generation while a foreground CPU load runs 21.3 → 22.4 tok/s 39.8 → 47.5 tok/s
Foreground CPU load (one loop per core) during generation, vs. alone 10.25 / 9.84 s → 10.11 / 9.90 s +15 % → +16 % over its own alone figure
Power while generating 8.75 → 9.23 W (+5.5 %, +3 % per token) 28.2 → 38.9 W (+38 %, +8 % per token)
  • The cost is power while generating. Battery life on a Mac laptop is expected to be shorter; this was not measured on battery.
  • The foreground was not slowed more under Standard.
  • A void first attempt is not in the table: the first round's "foreground load during generation" waited for the generation as well, so it was discarded and redone.

macOS security side (Background Task Management / launchd)

  • Disposition: BTM kept the job [enabled, allowed] through every switch and reload on both Macs. There was no disallow or disable.
  • launchd: spawn type went from background (5) to daemon (3).
  • Notification: a changed plist is registered again as an updated legacy daemon item (new record), and BackgroundTaskManagementAgent prepares a user notification at that moment ("Background Items Added"). That is the same notification a fresh install shows.
  • Product path on hardware (M4 Mac mini, this branch built as 0.0.3-dev.20260922+1fda5731):
    • binaries were replaced with install -m 0755, as install.sh does;
    • waired-agent install exited 0, and the plist reads Standard and lints clean;
    • BTM showed [enabled, allowed];
    • the engine was ready at utility QoS;
    • cold-cache discovery took 12.1 s and was not cut off (Metal, Apple M4); the model loaded at +24.3 s and generation worked.
  • A trap when testing by hand: overwriting the running /usr/local/bin/waired-agent in place with cp and then running it gets that process killed by macOS (Killed: 9), because code signing is cached per file. The installer's install -m writes a new file and is not affected.

Tests and checks

  • TestRenderLaunchDaemonPlist_HappyPath pins the pair. Every darwin test in internal/platform/service passes on an M5 Pro Mac, using the cross-compiled test binary.
  • Mutation: rendering Background again fails the test (plist must set ProcessType=Standard).
  • Other checks:
    • go test ./internal/platform/service/ on Linux;
    • GOOS=darwin|windows go vet;
    • golangci-lint on Linux: 0 issues. GOOS=darwin reports one finding that is already on main (service_darwin_test.go:627, QF1001) and not from this change;
    • hostname-guard.py;
    • decision-log-guard.py.

Other OSes

  • Linux and Windows are not affected by this mechanism. Their services set no lower priority, and ollama's Vulkan and CUDA backends compile lazily, not during discovery.

  • Related risks surfaced by the same review are not in this PR. They include:

    • Windows' battery-time QoS for background services;
    • Linux's render-group access for AMD/Intel GPUs.

    They need hardware checks and owner approval before they are filed.

Refs

🤖 Generated with Claude Code

https://claude.ai/code/session_01BviBGHJqCeM88hxCRYrtAg

… engine's first Metal compile finishes inside ollama's GPU-discovery window (#1521)

The LaunchDaemon plist said ProcessType=Background, and every process
the daemon spawns inherits the job's class: ollama serve and every
llama-server ran at background QoS. macOS compiles a llama-server's
Metal shaders in MTLCompilerService at the requester's QoS, so on a
cold shader cache (a new engine version) the compile outlasted ollama's
fixed 30 s GPU-discovery watchdog. ollama kills the discovery process,
does not retry, and keeps "no GPU" for the life of the engine, so the
engine sized itself for 0 B of VRAM with mmap off until its next start.
A killed request's compile is thrown away, so restarting straight away
(the fix first tried here) is cut off the same way; that was checked on
hardware and dropped.

Measured with the product's own start, 3 runs each, cold cache:
- M4 Mac mini: discovery cut off 3/3 under Background, 14.2 s under
  Standard; cold model load 121.5 s -> 24.3 s; warm load 7.4 -> 2.4 s;
  generation 26.0 -> 26.8 tok/s; power while generating +5.5 %.
- M5 Pro MacBook Pro: cut off 3/3 -> 12.1 s; cold load 103 -> 24.3 s;
  warm load 7.2 -> 1.8 s; generation 56-60 -> 72-78 tok/s; power +38 %
  (+8 % per token).
Foreground CPU work was slowed by a concurrent generation equally under
both classes. Background Task Management kept the job [enabled,
allowed]; launchd's spawn type went from background (5) to daemon (3).

App Nap, the reason the plist said Background, applies to apps, not to
a LaunchDaemon. Linux (systemd) and Windows (SCM) set no lower priority
either. No upgrade path for older installs: pre-release (owner).

Decision record: docs/decisions/20260922/0230-launchdaemon-runs-as-a-standard-job.md
The 2026-09-21 knowledge note that left the cause open is corrected in
place.

Fixes #1521

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BviBGHJqCeM88hxCRYrtAg
Signed-off-by: gen16k <gen16k@users.noreply.github.com>
@gen16k
gen16k force-pushed the fix/1521-discovery-restart branch from 97ccf6c to 1fda573 Compare September 21, 2026 19:07
@gen16k gen16k changed the title ollama: restart the engine once when its own GPU discovery was cut off (#1521) macOS: run the LaunchDaemon as a Standard job, not Background, so the engine's first Metal compile finishes inside ollama's GPU-discovery window (#1521) Sep 21, 2026
@gen16k
gen16k merged commit dbbc650 into main Sep 21, 2026
26 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

macOS: the first start after the bundled ollama moves to a new version times out its GPU discovery, and that engine runs as a CPU host until it restarts

1 participant