Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 13 additions & 1 deletion cmd/waired-agent/hardware_summary_test.go
Original file line number Diff line number Diff line change
Expand Up @@ -239,7 +239,19 @@ func TestHardwareSummaryFor_AppleSiliconBudgetSurvivesTheWireAdapter(t *testing.
// which is the whole point of recording it as a debt. RAMAvailableMeasuredAt
// (#699) was the most recent, held here between the proto PR and this one,
// and it dated the very field that sat here before it. Empty again.
var notPublishedByAgent = map[string]bool{}
var notPublishedByAgent = map[string]bool{
// waired-agent#69, held here for exactly one PR. The contract had to
// land alone (docs/decisions/20260719/0000-concurrent-proto-development.md
// §2), and hardware.GPU has no free-VRAM field to publish from until
// the per-OS reader lands with it. Until then every device reports 0,
// which hostfit reads as "no free reading" and answers with the
// total — so the budget is unchanged rather than wrong.
//
// The reader PR deletes this line. If it is still here, the contract
// is published with no producer, which is the #251 shape this map
// exists to make visible rather than to excuse.
"HardwareGPUSummary.VRAMFreeMB": true,
}

// TestHardwareSummaryFor_PublishesEveryWireField guards the bug class
// rather than the three fields: a field added to the broadcast summary
Expand Down
17 changes: 16 additions & 1 deletion docs/decisions/20260727/1830-ollama-vram-pool-nvidia-only.md
Original file line number Diff line number Diff line change
@@ -1,11 +1,26 @@
---
status: accepted
superseded_by:
- docs/decisions/20260813/1120-ollama-budget-sized-on-free-vram.md
---

# ollama の VRAM プールは proto 側で導出し、当面 NVIDIA に限る (20260727 18:30)

## Status
Accepted
Accepted(一部置換)

**§4「安全性は数値の精度ではなく accessor の床で担保する」だけ**が
`docs/decisions/20260813/1120-ollama-budget-sized-on-free-vram.md` に置き換わった。
床という構造は残り、**何に対する床か**が「単体デバイスの total」から
「単体デバイスの free(測れたとき)」に移っている。§4 の「今日の挙動より
悪くならない」保証はそこで意図的に失効する — この決定が直していた誤りは
複数枚での**過小**評価だけであり、#69 が直すのは 1 枚での**過大**評価で、
total に置いた床は後者を丸ごと飲み込むため。

§1(プールを proto 側で導出し wire に新フィールドを足さない)、§2(合算対象は
当面 NVIDIA のみ)、§3(控除はデバイス増分あたり)は**そのまま有効**。
Consequences が「#69 の担当範囲として残す。混ぜると両方おかしくなる」と
指定した分離も、そのとおりに守られている(de-rate は合算の前・デバイスごと)。

## Context

Expand Down
104 changes: 104 additions & 0 deletions docs/decisions/20260813/1120-ollama-budget-sized-on-free-vram.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,104 @@
---
status: accepted
supersedes:
- docs/decisions/20260727/1830-ollama-vram-pool-nvidia-only.md
---

# ollama の VRAM 予算は total ではなく free で測る。床の基準も free に移す (20260813 11:20)

## Status
Accepted

`docs/decisions/20260727/1830-ollama-vram-pool-nvidia-only.md` の **§4(安全性は
accessor の床で担保する)だけ**を置き換える。同決定の §1〜§3(プールを proto 側で
導出する / 合算対象は当面 NVIDIA のみ / 控除はデバイス増分あたり)はそのまま有効で、
1830 は `accepted` のまま残る。

## Context

`hostfit` はこれまで VRAM 予算を **total** で測っていた。一方 ollama の
スケジューラが見るのは **free** である。`docs/knowledges/20260727/1830-ollama-multi-gpu-placement.md`
が固定エンジン(0.31.1)のソースを読んで記録している既知のずれ、逐語:

> スケジューラが見るのは **free** メモリ、こちらが合算するのは **total**。1 枚の
> ときからある差だが、枚数分だけ拡大する(#69)。しかも
> `bestGPUGroupByAvailableMemory` には `bestSingleGPUFit` のような 80% の割引がない。

害は具体的である。ディスプレイを駆動している 8 GB カードは、コンポジタと他プロセスが
数百 MB〜数 GB を保持したまま「8 GB 使える」と値付けされる。モデルとコンテキストは
その楽観的な数字で選ばれ、実際にロードすると spill し、#621 の post-load verify が
気づいてコンテキストを縮め、エンジンを 1 回再起動する。**選定は直らない** — verify が
直せるのは窓であってモデルの選択ではない。

1830 §4 はこの床を置いた:

> `OllamaVRAMBudgetMB()` は (a) unified ホストでは `UsableVRAMMB` を必ず優先し、
> (b) **単体デバイスの値を下回らない**。……これがあるので `OllamaVRAMPoolMB` が
> どう間違っていても今日の挙動より悪くならない。

当時それは正しかった。**当時直していた誤りが「複数枚での過小評価」だけだったから**である。
床は「プール計算がどう間違っても 1 枚分は必ず確保される」ことを保証し、
waired#942 の逆向き(動くものを断る)に転ぶのを構造的に防いでいた。

**#69 が直す誤りは向きが逆で、同じ床がそれを丸ごと飲み込む。** 1 枚のホストでは
`VRAMPoolMB == 0` なので予算は `EffectiveVRAMMB`(= total)に落ち、free は式に
現れない。複数枚でも、free 合算が 1 枚の total を下回った瞬間にクランプされる。
つまり **床を total に置いたまま free を導入しても、#69 の主症例には何も起きない。**

## Decision

1. **デバイス単位の貸出可能量を `free`(測れたとき)とする。** `Device.lendableMB()` は
`VRAMFreeMB > 0 && VRAMFreeMB < VRAMTotalMB` のときだけ free を返し、それ以外は
total。`OllamaVRAMPoolMB` はこれを合算する。**de-rate は合算の前・デバイスごと**に
効く(1830 の Consequences が #69 の担当範囲として明示的に残した形)。

2. **床の基準を「1 枚の total」から「1 枚の free(測れたとき)」に移す。**
`Host.ollamaSingleDeviceMB()` を新設し、`OllamaVRAMBudgetMB` はそれを床に使う。
床という**構造は残す** — 変わるのは何に対する床かだけである。

3. **`EffectiveVRAMMB` は動かさない。** 1830 が「`min_vram_mb`、エンジン選択、vLLM の
TP=1 フォールバックは全て『1 枚の値』として書かれている」と述べたとおりで、
その判断は今も有効。free 対応は ollama 予算の経路にだけ入れる。

4. **unified-memory ホストは対象外。** `UsableVRAMMB` が既に共有プールから GPU が
wire down できる正直な上限であり、出荷済みのどの検出器も UMA デバイスの free を
報告しない。改善する余地がなく、フォールバックは推測になる。

5. **測定は 1 回きり、自分のエンジンが weights を持つ前。** ハードウェアプロファイルは
TTL で再サンプルされるので、ロード後に取った free は**自分の重みを除外する** —
再 tune のたびに空きが減って縮み続ける螺旋になり、ホストが自分の serve している
モデルを自分に課すことになる。`RAMAvailableGB` が #568 で同じ危険を同じ言葉で
名指しし、同じ答え(1 回測って永続化・ライブでは読まない)を採っている。それに倣う。

6. **`0` は「free 未測定」であって「空きゼロ」ではない。** total にフォールバックする。
free を報告しないドライバも、フィールドを知らない旧 agent も、ここに着地して
今日の予算を保つ。これが 5 と合わせて、**測っていないホストを de-rate しない**
ことを構造的に保証する。

## Consequences

- **1830 §4 の「今日の挙動より悪くならない」保証は失効する。** それが目的である。
予算は測定された free の分だけ**下がりうる**。代わりの保証は 6 の
「測っていなければ動かさない」— 悪化しうるのは、ドライバが実際に空きを報告した
ホストに限られる。
- `TestOllamaBudgetNeverShrinksTheHost`(`OllamaVRAMBudgetMB() >= EffectiveVRAMMB()` の
全数掃引)は**意図的に反転する**。新しい不変条件は
`OllamaVRAMBudgetMB() >= ollamaSingleDeviceMB()` かつ
`ollamaSingleDeviceMB() <= EffectiveVRAMMB()`。
- 過小評価に転ぶ経路が新しく開く(測定時に他プロセスが一時的に VRAM を掴んでいた
場合)。5 の「静かな瞬間に 1 回測る」がその窓を狭める唯一の手段であり、
#568 が RAM について既に受け入れたのと同じトレードオフである。
- エンジンとの整合は改善する。`bestSingleGPUFit` は free の 80% で判定するので、
free で測る予算は依然としてエンジンより**楽観的**なまま — つまりこの変更は
エンジンより厳しくなる方向には行き過ぎない。
- 1830 の Consequences が挙げた残件のうち、`#266`(カード別帯域テーブル)と
「`EffectiveVRAMMB` が列挙順の `GPUs[0]` を信じている件」(#264 調査項目 6)は
**未着手のまま**。この決定はどちらにも触れない。

## Refs
- https://github.com/waired-ai/waired-agent/issues/69
- https://github.com/waired-ai/waired-agent/issues/264
- https://github.com/waired-ai/waired-agent/issues/568 — 同型の「1 回測って永続化」
- `docs/decisions/20260727/1830-ollama-vram-pool-nvidia-only.md`(§4 を置換)
- `docs/decisions/20260727/1240-host-fit-single-source-proto-hostfit.md`
- `docs/knowledges/20260727/1830-ollama-multi-gpu-placement.md`
122 changes: 112 additions & 10 deletions proto/hostfit/hostfit.go
Original file line number Diff line number Diff line change
Expand Up @@ -303,6 +303,22 @@ type Host struct {
// asks a field added to a published struct to say so explicitly.
VRAMPoolMB int `json:"-"`

// VRAMAvailable0MB is the first GPU's free VRAM — the single-device
// mirror of VRAMPoolMB, and the only way the free reading reaches a
// SINGLE-GPU host, which is the shape waired-agent#69 actually
// reported (an 8 GB card also driving the display). A pooled host
// gets its de-rate inside VRAMPoolMB; a one-card host has no pool,
// so without this field it would keep sizing against the total.
//
// 0 means "no free reading" and leaves the total in place. Read only
// through OllamaVRAMBudgetMB, never directly — EffectiveVRAMMB stays
// the raw single-device figure that min_vram_mb, engine selection
// and vLLM's TP=1 fallback were authored against, for the same
// reason VRAMPoolMB may not widen it.
//
// json:"-" for the reason above: an input, not a payload.
VRAMAvailable0MB int `json:"-"`

// MemoryBandwidthSpecGBs is the published PEAK read bandwidth of the
// pool the weights are read from, in GB/s. 0 means "unknown", and
// that is a case this package must keep working for rather than an
Expand Down Expand Up @@ -403,6 +419,45 @@ type Device struct {
// the AMD Windows registry fallback reports devices this way — and
// such a device contributes nothing to a pool.
VRAMTotalMB int

// VRAMAvailableMB is how much of VRAMTotalMB the driver reported
// free, measured once while no engine of ours held weights. 0 means
// "no free reading" and falls back to the total, which is what every
// producer sent before the field existed
// (signer.HardwareGPUSummary.VRAMFreeMB carries the discipline and
// the reason, waired-agent#69).
//
// Named "available" rather than "free" — matching Host.RAMAvailableGB
// for the same quantity — partly to keep scripts/ci/protoconsumer
// working. That guard matches producers by field NAME, so a Device
// field spelled VRAMFreeMB and assigned right here in
// FromHardwareSummary would read as a proto-internal producer for the
// WIRE field of that name, and the real producer debt would never
// become visible in its table. The same consideration named
// LocalModelChoiceAt (waired-agent#647); the driver's own word stays
// on the wire, where nvidia-smi's memory.free is what it reports.
//
// json:"-" for the reason Host.VRAMPoolMB carries it: Device is an
// INPUT both adapters build, never a payload — the wire shape is
// signer.HardwareGPUSummary — and the additive-only guard asks a
// field added to a published struct to say so explicitly. Its
// untagged siblings predate the guard's baseline.
VRAMAvailableMB int `json:"-"`
}

// lendableMB is what this device can actually lend an engine: its free
// reading where the driver gave one, its total otherwise.
//
// The fallback is the whole safety argument. A device whose driver will
// not report free memory, and a producer that predates the field, both
// arrive here as 0 and get the total — so this rule can only ever
// de-rate a device the reader actually measured, never one it guessed
// at.
func (d Device) lendableMB() int {
if d.VRAMAvailableMB > 0 && d.VRAMAvailableMB < d.VRAMTotalMB {
return d.VRAMAvailableMB
}
return d.VRAMTotalMB
}

// OllamaVRAMPoolMB is the VRAM ollama may pool across devices, or 0
Expand Down Expand Up @@ -454,7 +509,15 @@ func OllamaVRAMPoolMB(devs []Device) int {
continue
}
n++
sum += d.VRAMTotalMB
// Free where it was measured, total otherwise. Summing free is
// what the engine itself does — availableMemoryForLoad sums
// gpu.FreeMemory — so the pool now answers the question the
// scheduler asks rather than an optimistic neighbour of it
// (waired-agent#69). The de-rate is applied per device, BEFORE
// the sum, which is what #264's decision record asks for: the
// total−free gap is per-device, so summing totals accumulates
// it once per card.
sum += d.lendableMB()
}
if n < 2 {
// Nothing to pool. Reported as "unknown" rather than as the
Expand Down Expand Up @@ -487,14 +550,21 @@ func FromHardwareSummary(hw *signer.HardwareSummary) Host {
}
if len(hw.GPUs) > 0 {
h.VRAM0MB = hw.GPUs[0].VRAMTotalMB
h.VRAMAvailable0MB = hw.GPUs[0].VRAMFreeMB
}
// Every GPU has always been on the wire; only this adapter and its
// agent-side twin threw the rest away. Nothing new has to be
// published for the pool, so the fix reaches every already-deployed
// agent the moment the control plane bumps its proto tag.
// agent-side twin threw the rest away. Nothing new had to be
// published for the pool, so that fix reached every already-deployed
// agent the moment the control plane bumped its proto tag. The free
// reading is the exception: it is a new field, so it arrives only
// from an agent new enough to measure it, and 0 keeps the total.
devs := make([]Device, 0, len(hw.GPUs))
for _, g := range hw.GPUs {
devs = append(devs, Device{Vendor: g.Vendor, VRAMTotalMB: g.VRAMTotalMB})
devs = append(devs, Device{
Vendor: g.Vendor,
VRAMTotalMB: g.VRAMTotalMB,
VRAMAvailableMB: g.VRAMFreeMB,
})
}
h.VRAMPoolMB = OllamaVRAMPoolMB(devs)
return h
Expand Down Expand Up @@ -529,20 +599,52 @@ func (h Host) EffectiveVRAMMB() int {
// the GPU can wire down, so it keeps winning.
//
// The aggregate may never come in BELOW the single-device figure. That
// is the mirror of router.VLLMVRAMBudgetMB's own floor, and here it is
// what makes this change structurally unable to regress: the worst a
// wrong pool can do is leave today's behaviour in place. It is also why
// is the mirror of router.VLLMVRAMBudgetMB's own floor. It is also why
// no "floor at the largest device" clause is needed — a host whose
// GPUs[0] is its small card gets the pool, which already exceeds the
// large one. Whether EffectiveVRAMMB itself should rank devices rather
// than trust enumeration order is a separate question, deliberately not
// answered here (waired-ai/waired-agent#264 item 6).
//
// What the floor is measured AGAINST changed with waired-agent#69: it
// is the single device's FREE figure where one was measured, not its
// total. docs/decisions/20260813/1120-ollama-budget-sized-on-free-vram.md
// records why, and revises §4 of the pool decision that set the earlier
// floor. The short version is that a floor at the total made the budget
// structurally unable to shrink, which was the point while the only
// error being corrected was an UNDER-count across cards — and is
// exactly wrong once the error being corrected is an OVER-count on one.
func (h Host) OllamaVRAMBudgetMB() int {
single := h.ollamaSingleDeviceMB()
if h.UnifiedMemory || h.VRAMPoolMB <= single {
return single
}
return h.VRAMPoolMB
}

// ollamaSingleDeviceMB is EffectiveVRAMMB as the OLLAMA path must read
// it: the same single-device budget, de-rated to what the driver
// reported free.
//
// Separate from EffectiveVRAMMB rather than folded into it, because the
// pool decision already settled that widening or narrowing THAT figure
// moves min_vram_mb, engine selection and vLLM's TP=1 fallback, all of
// which were authored against a whole card. This one moves only the
// ollama budget, which is the only consumer waired-agent#69 is about.
//
// A unified-memory host is left alone: UsableVRAMMB is already the
// honest bound on what its GPU can wire down from a shared pool, and no
// shipped detector reports a free figure for one, so there is nothing
// here to improve and a fallback to guess at.
func (h Host) ollamaSingleDeviceMB() int {
eff := h.EffectiveVRAMMB()
if h.UnifiedMemory || h.VRAMPoolMB <= eff {
if h.UnifiedMemory {
return eff
}
return h.VRAMPoolMB
if h.VRAMAvailable0MB > 0 && h.VRAMAvailable0MB < eff {
return h.VRAMAvailable0MB
}
return eff
}

// HasGPU reports whether the host has any GPU-addressable memory at
Expand Down
Loading
Loading