Skip to content

hardware: take the GPU topology reading where the privileges are, and read it back where they are not (#459) - #1466

Merged
gen16k merged 1 commit into
mainfrom
feat/459-init-persists-gpu-topology
Sep 20, 2026
Merged

gen16k merged 1 commit into
mainfrom
feat/459-init-persists-gpu-topology

Conversation

@gen16k

@gen16k gen16k commented Sep 20, 2026

Copy link
Copy Markdown
Contributor

docs-not-needed: the only new user-visible string is one failure line from waired init ("Could not record this computer's GPU memory topology: …"), printed on a best-effort step that changes nothing the user chose and nothing any docs-site page describes. No command, flag, prompt or status surface changes.

Why

#1462 landed a first-class Linux reading — AMDGPU_IDS_FLAGS_FUSION via DRM_IOCTL_AMDGPU_INFO, the bit Mesa reads for has_dedicated_vram — and then could not use it.

/dev/dri/renderD* is mode 0660 root:render, and the unit runs as User=waired with no supplementary groups. Measured on this repo's own Linux host while writing #1462:

euid=1000  open /dev/dri/renderD129: permission denied   -> the product reports UNKNOWN
euid=0     device_id=0x13c0  ids_flags=0x11              -> FUSION = true

So the reading exists, is correct, and never reaches the daemon.

What this does

The answer never changes for a given part, so a standing render group membership would buy nothing a single read cannot — and would be a privilege taken for a fact that is fixed. sudo waired init is already the elevated path, already the supported way to re-run setup, and already writes state that service_linux.go's FixStateOwnership chowns back to the service user.

So init takes the reading once and writes runtime/gpu-topology.json, in the shape host-memory.json already uses for a measurement the daemon can only take under conditions it has to arrange. The daemon reads it back through a new hardware.WithPersistedIntegration seam. A re-setup re-takes it, which is also how a swapped GPU stops being described by a stale entry.

The PCI vendor:device pair is the key, not a field beside the answer. Replace the card and the old entry matches nothing, so the host falls back to unknown rather than to a description of hardware that is gone.

Three things it is careful about

  • An unelevated waired init writes nothing, rather than an empty record over a good one. A device is recorded only when both its reading and its pair are in hand — and on a host where the node will not open, neither is.
  • A live reading still wins. The persisted answer is merged first and the live one second, so the existing three-state rule (a source that knows overrides) keeps the current reading authoritative wherever there is one. The persisted answer is a floor, not a ceiling.
  • Best effort in both directions. Enrolling a device has nothing to do with this, so a write failure is printed and init continues.

Apple Silicon deliberately gets no record. The reading is certain there (the architecture says so), it needs no privilege, and there is no PCI bus to key it by. Nothing is lost: the daemon takes that one itself every time. TestGPUTopologyFrom_AppleSiliconNeedsNoRecord says so out loud, so the empty case is not read as an oversight.

The knowledge note

docs/knowledges/20260920/2100-linux-gpu-facts-the-daemon-cannot-read.md records what was measured and, more usefully, what does not work — so the next reader does not re-derive it:

  • KFD's cpu_cores_count reads 0 on this repo's own APU, which is what ROCr derives HSA_AMD_MEMORY_PROPERTY_AGENT_IS_APU from. AMD's own ROCm/rocm-systems#8476 is open on this.
  • local_mem_size is hardcoded 0 in kfd_topology.c, unchanged v5.10 → master.
  • The iGPU advertises max_link_width 16 at 16 GT/s — the same PCIe link a discrete card does.
  • amdgpu_drv.c's 76-entry AMD_IS_APU table (MIT-licensed) contains neither 0x13c0 nor 0x1586: modern parts fall through to PCI_ANY_ID/CHIP_IP_DISCOVERY and amdgpu_discovery.c settles the flag at runtime, by GC IP version or by asking the SMU for the package type.
  • ollama's sysfs ratio rule (vram <= 4 GiB && gtt >= 4*vram) classifies this iGPU right and the Strix Halo reference host wrong.
  • /proc/cpuinfo on aarch64 has no model name line at all, so CPU.Model is empty on Grace / GB10 / Jetson, and /proc/device-tree/model does not exist on GB10 or Grace either (both boot SBSA/UEFI).

Verification

  • go test ./internal/... ./cmd/...: green. New tests cover the record's round trip and its tolerance of a missing or corrupt file, IntegratedFor's PCI keying, what init writes and what it refuses to write, and the three merge cases (persisted answers where live cannot, live overrides stale, an uncovered card says nothing).
  • go vet cross-built for linux / windows / darwin: clean.
  • golangci-lint: the only findings are three pre-existing SA4023 in internal/relay/client/client.go, untouched here.
  • testnet-gate-guard.sh (78 packages classified), decision-log-guard.py: ok.

What this does not do

It does not widen UnifiedMemory — #459's three false ClassDiscrete assumptions are still open, and the budget rule that would have to travel with a widened class is still family-specific. It does not give Linux/NVIDIA a reading either: CGO_ENABLED=0 forecloses CUDA, NVML carries no such flag, and there is no such hardware in this fleet to verify against.

Refs #459, #1455

🤖 Generated with Claude Code

https://claude.ai/code/session_01BviBGHJqCeM88hxCRYrtAg

… read it back where they are not

waired-agent#1462 landed a first-class Linux reading of "are this
accelerator's memory and the OS's RAM one physical pool?" —
AMDGPU_IDS_FLAGS_FUSION via DRM_IOCTL_AMDGPU_INFO, the bit Mesa reads
for has_dedicated_vram — and then could not use it. /dev/dri/renderD* is
mode 0660 root:render and the unit runs as User=waired with no
supplementary groups, so the daemon reads "unknown" forever. Measured on
this repo's own Linux host: euid 1000 gets EACCES, euid 0 gets
ids_flags = 0x11.

The answer never changes for a given part, so a standing group
membership would buy nothing a single read cannot. `sudo waired init` is
already the elevated path, already the supported way to re-run setup,
and already writes state that service_linux.go's FixStateOwnership
chowns back to the service user. So it takes the reading once and writes
runtime/gpu-topology.json, in the shape host-memory.json already uses
for a measurement the daemon can only take under conditions it has to
arrange. A re-setup re-takes it, which is also how a swapped GPU stops
being described by a stale entry.

The PCI vendor:device pair is the KEY rather than a field beside the
answer. Replacing the card leaves the old entry matching nothing, and
the host falls back to unknown rather than to a description of hardware
that is gone.

Three things this is careful about:

  - An unelevated `waired init` writes NOTHING, rather than an empty
    record over a good one. A device is recorded only when both its
    reading and its pair are in hand.
  - A live reading still wins. The persisted answer is merged first and
    the live one second, so the existing three-state rule — a source
    that knows overrides — makes the current reading authoritative
    wherever there is one.
  - The write is best effort in both directions. Enrolling a device has
    nothing to do with this, so a failure is printed and init continues.

Apple Silicon deliberately gets no record: the reading is certain there
(the architecture says so), it needs no privilege, and there is no PCI
bus to key it by. Nothing is lost — the daemon takes that one itself
every time.

The knowledge note records what was measured and what does NOT work, so
the next reader does not re-derive it: KFD's cpu_cores_count reads 0 on
this repo's own APU (which is what ROCr derives
HSA_AMD_MEMORY_PROPERTY_AGENT_IS_APU from), local_mem_size is hardcoded
0 in kfd_topology.c, the iGPU advertises the same PCIe link as a
discrete card, amdgpu_drv.c's 76-entry AMD_IS_APU table contains neither
0x13c0 nor 0x1586 because modern parts are settled at runtime by
amdgpu_discovery.c, and ollama's sysfs ratio rule classifies this iGPU
right and the Strix Halo reference host wrong.

Refs #459, #1455

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01BviBGHJqCeM88hxCRYrtAg
Signed-off-by: gen16k <gen16k@users.noreply.github.com>
@gen16k
gen16k merged commit de9aba9 into main Sep 20, 2026
32 checks passed
gen16k added a commit that referenced this pull request Sep 22, 2026
…en an AMD or Intel GPU (#1535) (#1537)

The inference engine is the daemon's child and runs as `waired`. On
Debian and Ubuntu, /dev/dri/renderD* and /dev/kfd are 0660 root:render,
because both distributions build systemd with group-render-mode=0660.
So without the group the engine cannot see an AMD GPU (ROCm or Vulkan)
or an Intel GPU (Vulkan), and it runs on the CPU without saying why.
#1466 had chosen not to grant the group, but that choice considered
only the profiler's reading.

- The .deb postinst runs `usermod -a -G render waired` on every
  configure, when the group exists, so an upgrade fixes an existing
  install. `waired-agent install` does the same through
  ensureGPUGroups.
- `video` is left out (owner decision): it also opens cameras and the
  display, and computing does not need it.
- The unit keeps no SupplementaryGroups=, because naming an absent group
  stops the unit (216/GROUP). systemd applies a User='s groups from the
  group database.
- The gpu-topology.json reading from `sudo waired init` stays, as the
  floor for a host whose service user is not in the group.
- The new decision record partly supersedes the 2000 decision. The
  notes and comments that said the service has no groups are corrected.
- The troubleshooting and Linux install pages (en/ja) now name the
  group, with how to check it.

Checked on a Linux host with an NVIDIA card and an AMD iGPU:
- Without the group, the product's ollama does not list the iGPU at all.
- With the group, it finds the iGPU through Vulkan. With default
  settings it then drops the iGPU and serves on CUDA, as it does today.
- The postinst (run twice, as an upgrade) adds the group, and the
  restarted daemon runs with it.


Claude-Session: https://claude.ai/code/session_01BviBGHJqCeM88hxCRYrtAg

Signed-off-by: gen16k <gen16k@users.noreply.github.com>
Co-authored-by: gen16k <gen16k@users.noreply.github.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant