macOS: run the LaunchDaemon as a Standard job, not Background, so the engine's first Metal compile finishes inside ollama's GPU-discovery window (#1521) - #1525
Merged
Conversation
… engine's first Metal compile finishes inside ollama's GPU-discovery window (#1521) The LaunchDaemon plist said ProcessType=Background, and every process the daemon spawns inherits the job's class: ollama serve and every llama-server ran at background QoS. macOS compiles a llama-server's Metal shaders in MTLCompilerService at the requester's QoS, so on a cold shader cache (a new engine version) the compile outlasted ollama's fixed 30 s GPU-discovery watchdog. ollama kills the discovery process, does not retry, and keeps "no GPU" for the life of the engine, so the engine sized itself for 0 B of VRAM with mmap off until its next start. A killed request's compile is thrown away, so restarting straight away (the fix first tried here) is cut off the same way; that was checked on hardware and dropped. Measured with the product's own start, 3 runs each, cold cache: - M4 Mac mini: discovery cut off 3/3 under Background, 14.2 s under Standard; cold model load 121.5 s -> 24.3 s; warm load 7.4 -> 2.4 s; generation 26.0 -> 26.8 tok/s; power while generating +5.5 %. - M5 Pro MacBook Pro: cut off 3/3 -> 12.1 s; cold load 103 -> 24.3 s; warm load 7.2 -> 1.8 s; generation 56-60 -> 72-78 tok/s; power +38 % (+8 % per token). Foreground CPU work was slowed by a concurrent generation equally under both classes. Background Task Management kept the job [enabled, allowed]; launchd's spawn type went from background (5) to daemon (3). App Nap, the reason the plist said Background, applies to apps, not to a LaunchDaemon. Linux (systemd) and Windows (SCM) set no lower priority either. No upgrade path for older installs: pre-release (owner). Decision record: docs/decisions/20260922/0230-launchdaemon-runs-as-a-standard-job.md The 2026-09-21 knowledge note that left the cause open is corrected in place. Fixes #1521 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01BviBGHJqCeM88hxCRYrtAg Signed-off-by: gen16k <gen16k@users.noreply.github.com>
gen16k
force-pushed
the
fix/1521-discovery-restart
branch
from
September 21, 2026 19:07
97ccf6c to
1fda573
Compare
This was referenced Sep 22, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #1521
Why
On macOS, the first engine start with a cold Metal shader cache lost the GPU. This happens after the bundled ollama moves to a new version, and whenever the cache is cold. ollama's own GPU discovery was cut off by its fixed 30 s watchdog (
llama-server GPU discovery watchdog timed out), and the engine then ran as a CPU host until its next restart (total_vram="0 B", mmap off).Cause. The LaunchDaemon plist said
ProcessType=Background, and every process the daemon spawns inherits that class. Soollama serveand everyllama-serverran at background QoS. A llama-server's first start has macOS'sMTLCompilerServicecompile ggml's Metal shaders at the requester's QoS. Under Background that compile outlasted the 30 s window.What I measured on an M4 Mac mini, from a cold cache:
Where Background came from. It dates from the repository's first import, with the reason "tells App Nap to leave us alone". App Nap is a mechanism for apps, not for a LaunchDaemon. The Linux systemd unit and the Windows service set no lower priority.
What changed
internal/platform/service/service_darwin.go: the plist'sProcessTypeisStandard(the same as leaving the key out). The comment gives the reason and points to the decision record.service_darwin_test.go: pins the key/value pair<key>ProcessType</key>\n <string>Standard</string>as a PRODUCT CONTRACT (owner decision 2026-09-22 on macOS: the first start after the bundled ollama moves to a new version times out its GPU discovery, and that engine runs as a CPU host until it restarts #1521).docs/decisions/20260922/0230-launchdaemon-runs-as-a-standard-job.md, with the measurements.docs/knowledges/20260921/2210-…).waired-agent install.engine_discovery_error) stays as the detector.Measurements
The product's own start, 3 runs per condition, median unless noted. No other load was running (sharing off). Power is the mean of 5 × 1 s
powermetricssamples taken from 3 s into a 1024-token generation.macOS security side (Background Task Management / launchd)
[enabled, allowed]through every switch and reload on both Macs. There was no disallow or disable.spawn typewent frombackground (5)todaemon (3).0.0.3-dev.20260922+1fda5731):install -m 0755, asinstall.shdoes;waired-agent installexited 0, and the plist readsStandardand lints clean;[enabled, allowed];/usr/local/bin/waired-agentin place withcpand then running it gets that process killed by macOS (Killed: 9), because code signing is cached per file. The installer'sinstall -mwrites a new file and is not affected.Tests and checks
TestRenderLaunchDaemonPlist_HappyPathpins the pair. Every darwin test ininternal/platform/servicepasses on an M5 Pro Mac, using the cross-compiled test binary.Backgroundagain fails the test (plist must set ProcessType=Standard).go test ./internal/platform/service/on Linux;GOOS=darwin|windows go vet;GOOS=darwinreports one finding that is already on main (service_darwin_test.go:627, QF1001) and not from this change;hostname-guard.py;decision-log-guard.py.Other OSes
Linux and Windows are not affected by this mechanism. Their services set no lower priority, and ollama's Vulkan and CUDA backends compile lazily, not during discovery.
Related risks surfaced by the same review are not in this PR. They include:
render-group access for AMD/Intel GPUs.They need hardware checks and owner approval before they are filed.
Refs
docs/decisions/20260922/0230-launchdaemon-runs-as-a-standard-job.mdProcessType🤖 Generated with Claude Code
https://claude.ai/code/session_01BviBGHJqCeM88hxCRYrtAg