Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
18 changes: 10 additions & 8 deletions cmd/waired-agent/gpu_topology.go
Original file line number Diff line number Diff line change
Expand Up @@ -4,14 +4,15 @@ import (
"github.com/waired-ai/waired-agent/internal/runtime/state"
)

// The GPU topology reading the daemon cannot take, read back from where
// an elevated setup left it (waired-agent#459).
// The GPU topology reading an elevated setup took, read back for a daemon
// that could not take it itself (waired-agent#459).
//
// The daemon runs as a service user with no access to /dev/dri/renderD*,
// so its own reading of "is this accelerator's memory the system's
// memory?" is almost always unknown on Linux. `sudo waired init` can
// open the node, takes the reading once and writes gpu-topology.json;
// this is the read side.
// The daemon opens /dev/dri/renderD* itself where its service user is
// in the `render` group, which the installer arranges (#1535). Where it
// is not, its own reading of "is this accelerator's memory the system's
// memory?" is unknown on Linux. `sudo waired init` can open the node,
// takes the reading once and writes gpu-topology.json; this is the read
// side, and a live reading still wins over it.
//
// Read per profiler construction rather than cached in a package
// variable, for the reason hostMemoryMeasurement is: the same process
Expand All @@ -36,7 +37,8 @@ func persistedGPUIntegration(stateDir string) func(pciID string) (bool, bool) {
}

// persistedGPUVRAM is the same record's memory reading, for the parts
// whose size the daemon cannot read either (waired-agent#1483).
// whose size the daemon may not be able to read either
// (waired-agent#1483).
func persistedGPUVRAM(stateDir string) func(pciID string) (int, bool) {
return func(pciID string) (int, bool) {
rec, err := state.ReadGPUTopology(stateDir)
Expand Down
16 changes: 9 additions & 7 deletions cmd/waired/init_gpu_topology.go
Original file line number Diff line number Diff line change
Expand Up @@ -9,15 +9,17 @@ import (
"github.com/waired-ai/waired-agent/internal/runtime/state"
)

// Taking the GPU topology reading during `sudo waired init`, because the
// daemon cannot take it (waired-agent#459).
// Taking the GPU topology reading during `sudo waired init`, for a daemon
// that cannot take it itself (waired-agent#459).
//
// On Linux the fact — is this accelerator's memory the system's memory?
// — is behind /dev/dri/renderD*, which is mode 0660 root:render while
// the unit runs as User=waired with no supplementary groups. An elevated
// setup can open it; the service never can. So the reading is taken here
// and persisted, the way host-memory.json persists a measurement the
// daemon can only take under conditions it has to arrange.
// — is behind /dev/dri/renderD*, which is mode 0660 root:render on
// Debian and Ubuntu. The installer puts the service user in `render`
// (#1535), so the daemon normally reads it live; an elevated setup can
// always open it. The reading is taken here and persisted, the way
// host-memory.json persists a measurement the daemon can only take under
// conditions it has to arrange, as the floor for a host whose service
// user is not in the group.
//
// `init` rather than the installer, because a re-setup is the supported
// way to re-take it — which is also what makes a swapped GPU stop being
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -83,6 +83,7 @@ You can add it later with `sudo waired runtimes install ollama`.
|---|---|
| Packages | `waired` and `waired-tray`, from Waired's apt repository |
| Background service | `waired-agent.service` (systemd), starts at boot |
| Service user | `waired`, added to the `render` group so the engine can use an AMD or Intel GPU |
| Settings | `/etc/waired` |
| State | `/var/lib/waired` |

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,7 @@ curl -fsSL https://github.com/waired-ai/waired-agent/releases/latest/download/in
|---|---|
| パッケージ | `waired`と`waired-tray`。Wairedのaptリポジトリから |
| バックグラウンドサービス | `waired-agent.service`(systemd)。起動時に開始 |
| サービスのユーザー | `waired`。推論エンジンがAMDやIntelのGPUを使えるように、`render`グループに追加 |
| 設定 | `/etc/waired` |
| 状態 | `/var/lib/waired` |

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -37,7 +37,7 @@ waired models ls --detail

どのGPUでモデルを動かすかはOllama自身が決め、Wairedはそれに従います。単体GPUは使われます。プロセッサーに内蔵されたGPUが使われるのは、Appleシリコン、AMD Strix Halo(Ryzen AI Max)、DGX SparkのようなNVIDIAのユニファイドメモリのパソコンだけです。それ以外のパソコン、たとえばRadeon 780MやIntel Arcのグラフィックスを持つノートパソコンでは、最初の行に推論エンジンがそのGPUを標準では使わないと表示され、モデルはプロセッサ向けに決められます。これは推論エンジンの判断で、検出の失敗ではありません。

Linuxでは、WairedはAMDのGPUを`amdgpu`ドライバーから直接読みます。AMDのGPUを検出するのに、ROCmをインストールする必要はありません。Intel Arcの単体GPUのメモリ容量は、Linuxでは管理者権限がないと読めないので、`sudo waired init`を実行したときに一度だけ読みます。それまでは最初の行にそのGPUのメモリ容量が分からないと表示され、モデルはプロセッサ向けに決められます。
Linuxでは、WairedはAMDのGPUを`amdgpu`ドライバーから直接読みます。AMDのGPUを検出するのに、ROCmをインストールする必要はありません。推論エンジンは`render`グループを通してAMDやIntelのGPUを使います。インストーラは、サービスを動かすユーザー`waired`をこのグループに追加します。そのようなGPUが見つかっているのにモデルがプロセッサで動く場合や、最初の行にIntel Arcの単体GPUのメモリ容量が分からないと表示される場合は、`id waired`を実行して`render`が含まれているか確認します。含まれていなければ、`sudo usermod -a -G render waired`を実行してサービスを再起動します。

それでも内蔵GPUを使わせたい場合は、下の`WAIRED_NVIDIA_SMI`と同じ方法でサービスに`OLLAMA_IGPU_ENABLE=1`を設定し、サービスを再起動します。そのパソコンでも、Wairedはモデルをプロセッサ向けに決めます。

Expand Down
11 changes: 7 additions & 4 deletions docs-site/src/content/docs/troubleshooting/slow-or-wrong.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,10 +59,13 @@ not use that GPU by default, and models are sized for the processor. That is
the engine's decision, not a detection failure.

On Linux, Waired reads AMD GPUs from the `amdgpu` driver directly, so ROCm
does not need to be installed for an AMD GPU to be found. An Intel Arc card's
memory size can only be read with administrator rights there, so it is read
once when you run `sudo waired init`. Until then the first line says its
memory size is not known, and models are sized for the processor.
does not need to be installed for an AMD GPU to be found. The engine reaches
an AMD or Intel GPU through the `render` group, and the installer adds
`waired`, the user the service runs as, to that group. If such a GPU is found
but models still run on the processor, or the first line says an Intel Arc
card's memory size is not known, run `id waired` and check that `render` is
listed. If it is not, run `sudo usermod -a -G render waired` and restart the
service.

To have the engine use a built-in GPU anyway, set `OLLAMA_IGPU_ENABLE=1` for
the service the same way as `WAIRED_NVIDIA_SMI` below, then restart it. Waired
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -5,6 +5,7 @@ supersedes:
superseded_by:
- docs/decisions/20260921/0300-nvidia-single-pool-parts-are-named.md
- docs/decisions/20260921/1600-amd-parts-are-named-by-their-isa-target.md
- docs/decisions/20260922/1430-linux-service-user-joins-render.md
---

# 測定の出自は、チップの粒度で事実から導く (20260920 20:00)
Expand All @@ -29,6 +30,11 @@ Accepted。waired-agent#1455、および #459 の Ask 1・2。
ディスクリートカードは ISA ターゲット(`gfx1100` など)で名指す(#1485)。
`GPU.Model` を使わない点は AMD についても変わらない。

**Consequences の「`render` グループを常時与える案は採らない」の 1 文は、
`docs/decisions/20260922/1430-linux-service-user-joins-render.md` が改めた**(#1535)。推論エンジンも同じサービスユーザーで
動き、AMD / Intel の GPU で計算するのに render ノードが要るため。
installer がサービスユーザーを `render` に入れる。残りは不変。

## Context

カタログの出自の語彙 `catalog.HostClasses` は、綴りが**ベンダ名 + メモリ容量**
Expand Down
80 changes: 80 additions & 0 deletions docs/decisions/20260922/1430-linux-service-user-joins-render.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,80 @@
---
status: accepted
supersedes:
- docs/decisions/20260920/2000-host-provenance-is-derived-at-chip-granularity.md
---

# Linux のサービスユーザーを render グループに入れる (20260922 14:30)

## Status

Accepted。waired-agent#1535。
`docs/decisions/20260920/2000-host-provenance-is-derived-at-chip-granularity.md`
の Consequences にある「`render` グループを常時与える案は採らない」の
**1 文だけ**を改める。そのほかの部分は変えない。

## Context

- #1466(#459)では、render ノードの ioctl で読む事実を扱った(AMD の FUSION、のちに
Intel の単体 GPU の VRAM)。デーモンはそのノードを開けない。そこで `sudo waired init`
が 1 回だけ読んで `runtime/gpu-topology.json` に残す形にした。理由は「変わらない
事実のために常時の特権を増やさない」だった。
- その判断は、プロファイラの読み取りだけを見ていた。推論エンジン(ollama・vLLM)は、
資格情報を変えずにデーモンの子として起動し(`internal/runtime/spawner_unix.go`)、
同じ `waired` で動く。エンジンが GPU で計算するには、次のノードを開く必要がある。
- AMD: ROCm は `/dev/kfd` と `renderD*`、Vulkan は `renderD*`。
- Intel: Vulkan で `renderD*`。
- systemd の既定の mode は 0666(v236 以降、`group-render-mode`)。しかし Debian と
Ubuntu は `-Dgroup-render-mode=0660` でビルドしている。そのためサポート対象
(Ubuntu 24.04 以降・Debian trixie 以降)では、`renderD*` と `/dev/kfd` は
`0660 root:render` になる。
- 実機で確かめた。機械は RTX PRO 4000 と AMD の iGPU を載せた Linux 機で、
Ubuntu 26.04・systemd 259。製品の ollama 0.34.0 を `waired` で起動した。
- グループが無いと、AMD の iGPU は検出に現れず、ログにも何も出ない。
- `render` を付けると、iGPU が Vulkan の装置として現れる。既定の設定では
`dropping integrated GPU` で外れ、CUDA だけで動く。これは今と同じ結果になる。
- 上流の ollama の installer は、`ollama` ユーザーを `render` と `video` に入れている
(グループがある場合)。ROCm と Intel の docs も、`render` への参加を求めている。

## Decision

- サービスユーザー `waired` を `render` に入れる。グループが無い機械では入れずに
先へ進み、グループを作ることはしない。
- .deb の postinst: configure のたびに `usermod -a -G render waired` を実行する。
upgrade のときも実行されるので、既存の install も直る。
- `waired-agent install`: `ensureGPUGroups` で同じことをする。
- `video` には入れない(オーナーの判断、2026-09-22、#1535)。`video` は
Web カメラ(`/dev/video*`)と画面出力(`/dev/dri/card*`)にも届く。
サポート対象のディストリでは、計算には要らない。
- unit に `SupplementaryGroups=` は書かない。
- 無いグループを名指すと、unit が起動しない(exit 216/GROUP)。
- systemd は `User=` の補助グループをグループ DB から付ける(systemd.exec(5))。
グループに入れて再起動すれば足りる。
- `gpu-topology.json`(`sudo waired init` が残す読み取り)は残す。
- グループがあれば、デーモンがその場で読む。その場の読み取りが優先される。
- 残した読み取りは、サービスユーザーがグループに入っていない機械のための下限になる。

## Consequences

- Linux の AMD / Intel の GPU を、エンジンが開けるようになる。
- 実機で確かめたのは、AMD の iGPU が検出に現れるところまで(既定では ollama 自身が外す)。
- 単体 GPU・Strix Halo・Intel Arc で実際に計算できるかは、#1495・#1496・#1497 の
計測で確かめる。
- デーモンは FUSION と Intel の VRAM を自分で読むようになる。多くの機械では
`gpu-topology.json` の読み取りを使わなくなる。
- 未解決の点(#1535 に記録): ollama の Vulkan は、CAP_PERFMON か root でないと
空き VRAM を読めない。`NoNewPrivileges=yes` はファイル capability も無効にする。
capability を与えるかは、別に判断する。
- Windows(LocalSystem)と macOS(root)には、同じ問題は無い。

## Refs

- https://github.com/waired-ai/waired-agent/issues/1535
- https://github.com/waired-ai/waired-agent/issues/1534
- docs/decisions/20260920/2000-host-provenance-is-derived-at-chip-granularity.md
- docs/knowledges/20260920/2100-linux-gpu-facts-the-daemon-cannot-read.md
- https://github.com/systemd/systemd/blob/e362a4efcbe7366d03f6b60959b752440fe1f69b/rules.d/50-udev-default.rules.in#L61-L63
- https://salsa.debian.org/systemd-team/systemd/-/blob/debian/master/debian/rules
- https://github.com/ollama/ollama/blob/6383a0fa9cbf97494b847226e189f6e36b401a08/scripts/install.sh#L197-L230
- https://rocm.docs.amd.com/projects/install-on-linux/en/latest/install/prerequisites.html
- internal/platform/service/service_linux.go, packaging/debian/waired/postinst
Original file line number Diff line number Diff line change
Expand Up @@ -87,6 +87,13 @@ GB10 と Grace には存在しない**(ACPI/UEFI の SBSA で起動するた
デーモンは PCI ペアをキーに読み戻す。**`render` グループを常時与える案は
採らない** — 変わらない事実のために常時の特権を増やさないため。

**訂正(20260922):** この結論は #1535 で改めた。推論エンジン(ollama)も
デーモンの子として同じ `waired` で動き、AMD / Intel の GPU で計算するのに
render ノードが要る。この節はプロファイラの読み取りだけを見ていた。
installer がサービスユーザーを `render` に入れ、永続化はグループに
入っていない機械のための下限として残る。
`docs/decisions/20260922/1430-linux-service-user-joins-render.md`

## Refs

- https://github.com/waired-ai/waired-agent/issues/459
Expand All @@ -96,3 +103,4 @@ GB10 と Grace には存在しない**(ACPI/UEFI の SBSA で起動するた
- https://github.com/lmstudio-ai/lms/issues/589
- docs/decisions/20260920/2000-host-provenance-is-derived-at-chip-granularity.md
- internal/hardware/integrated_linux.go, internal/runtime/state/gpu_topology.go
- https://github.com/waired-ai/waired-agent/issues/1535
Original file line number Diff line number Diff line change
Expand Up @@ -69,6 +69,8 @@ revision は載っていない(検証機の 13C0 も無い)。だから名
システム RAM の速度とバス幅を知りたいときはここを読むことになるが、
これだけは root が要る。render ノードの ioctl は Debian 系で 0660
root:render なので、デーモン(`User=waired`、補助グループ無し)からは開けない。
**訂正(20260922):** #1535 以後は installer がサービスユーザーを `render` に
入れるので、デーモンも開ける(`docs/decisions/20260922/1430-linux-service-user-joins-render.md`)。

### 6. Windows には同じものが無い

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -47,7 +47,12 @@ i915 の region info は 4 + 4 + 8 + 8 + 64 バイトの union で **88 バイ

どちらも `DRM_RENDER_ALLOW` だが、render ノードは Debian 系で 0660
root:render なので、デーモンからは開けない。`sudo waired init` が 1 度読み、
`gpu-topology.json` に `vram_total_mb` として残す。AMD の FUSION の読みと
`gpu-topology.json` に `vram_total_mb` として残す。
**訂正(20260922):** デーモンから開けないままでは、エンジンもこのカードを
開けず、Vulkan で使えなかった。#1535 で installer がサービスユーザーを
`render` に入れ、デーモンもエンジンも開けるようになった。永続化は、
グループに入っていない機械のための下限として残る
(`docs/decisions/20260922/1430-linux-service-user-joins-render.md`)。AMD の FUSION の読みと
同じ経路である。読めていない単体カードは「使わない GPU」の側に理由つきで置く。
エンジンは使うかもしれないので、表示は「CPU で動く」とは言わず、
「メモリ容量が分からない」と言う。
Expand Down
3 changes: 2 additions & 1 deletion internal/hardware/engine_gpus.go
Original file line number Diff line number Diff line change
Expand Up @@ -57,7 +57,8 @@ const unusedIntegratedReason = "the engine uses an integrated GPU by default onl

// unreadIntelMemoryReason is the Reason an Intel card whose memory size
// could not be read carries. On Linux the size comes from the driver's
// query on the render node, which `sudo waired init` can open.
// query on the render node, which the service opens through the render
// group (#1535) and `sudo waired init` can open regardless.
const unreadIntelMemoryReason = "its memory size could not be read, so models are not sized for it; on Linux, running `sudo waired init` reads it"

// UnusedGPU is a GPU that was detected but that the host is not
Expand Down
8 changes: 5 additions & 3 deletions internal/hardware/gpu_intel_linux.go
Original file line number Diff line number Diff line change
Expand Up @@ -19,9 +19,11 @@ import (
// i915's DRM_I915_QUERY_MEMORY_REGIONS, the same two queries Mesa's
// anv/iris use to size local memory. Both are DRM_RENDER_ALLOW, so the
// only obstacle is the node's mode (0660 root:render on Debian-family
// systems): the daemon's own reading usually fails, and `sudo waired
// init` takes it and persists it alongside the integration reading
// (cmd/waired/init_gpu_topology.go), which is how the daemon gets it.
// systems). The installer puts the service user in `render` (#1535), so
// the daemon reads it itself. `sudo waired init` also takes it and
// persists it alongside the integration reading
// (cmd/waired/init_gpu_topology.go), the floor for a host whose service
// user is not in the group.
//
// No Intel discrete card has been available to measure this against. The
// layouts below are transcribed from include/uapi/drm/xe_drm.h and
Expand Down
20 changes: 10 additions & 10 deletions internal/hardware/integrated_linux.go
Original file line number Diff line number Diff line change
Expand Up @@ -43,16 +43,16 @@ import (
// AMD acknowledges the gap: ROCm/rocm-systems#8476, "APUs are not
// identifiable through amdsmi", open.
//
// WHO CAN READ IT. The render node is mode 0660 root:render. The daemon
// runs as User=waired with no supplementary groups, so it cannot open
// it, and granting it the group standing would buy a privilege for the
// sake of a fact that never changes. Instead `sudo waired init` — which
// is already the elevated path, and already writes state that
// service_linux.go's FixStateOwnership chowns back — takes the reading
// once and persists it, the way host-memory.json persists the
// available-memory measurement. This function is still called on every
// profile: where the node does happen to open it is the better source,
// and where it does not the answer is UNKNOWN, never "discrete".
// WHO CAN READ IT. The render node is mode 0660 root:render on Debian
// and Ubuntu. The daemon runs as User=waired, which the installer puts
// in `render` (#1535): the inference engine runs as the same user and
// needs the node to compute on the GPU at all, so the daemon can read
// this itself. `sudo waired init` also takes the reading once and
// persists it, the way host-memory.json persists the available-memory
// measurement; that is the floor for a host whose service user is not
// in the group. This function is still called on every profile: where
// the node opens it is the better source, and where it does not the
// answer is UNKNOWN, never "discrete".

// Constants transcribed from include/uapi/drm/amdgpu_drm.h.
const (
Expand Down
Loading
Loading