Skip to content

proto(hostfit): size the ollama VRAM budget on free memory, not the card's total (#69) - #772

Merged
gen16k merged 1 commit into
mainfrom
feat/69-proto-vram-free
Aug 13, 2026
Merged

gen16k merged 1 commit into
mainfrom
feat/69-proto-vram-free

Conversation

@gen16k

@gen16k gen16k commented Aug 13, 2026

Copy link
Copy Markdown
Contributor

The engine sizes placement against free VRAM — ollama 0.31.1's availableMemoryForLoad sums gpu.FreeMemory — while this repo sized its budget against the total. docs/knowledges/20260727/1830-ollama-multi-gpu-placement.md recorded the mismatch when it read the pinned engine, verbatim:

スケジューラが見るのは free メモリ、こちらが合算するのは total。1 枚のときからある差だが、枚数分だけ拡大する(#69)。しかも bestGPUGroupByAvailableMemory には bestSingleGPUFit のような 80% の割引がない。

A card that is also driving a display is therefore valued at more than it can lend. The model and its context get chosen optimistically, the load spills, and #621's post-load verify shrinks the window and restarts the engine. The selection is never corrected — that verify can only move the window, not un-pick a model.

This is the contract half, on its own, as docs/decisions/20260719/0000-concurrent-proto-development.md §2 requires. The settled field table is on #69.

Measured before designing

Three premises in the issue were stale, and the arithmetic had moved into proto/hostfit since it was written. Two things that reading turned up changed the design:

The floor would have swallowed the whole fix. OllamaVRAMBudgetMB clamps at the single-device figure. A one-card host — the shape #69 actually reports, an 8 GB card driving a display — has VRAMPoolMB == 0, so the budget falls through to EffectiveVRAMMB and a free reading never enters the arithmetic at all. De-rating only inside the pool would have been a no-op on the reported case.

A live reading would spiral. The hardware profile is re-sampled on a TTL, including after our own engine loads a model. A free figure taken then excludes our own weights, so each re-tune would see less memory and shrink further. RAMAvailableGB already names this hazard in the same words (#568):

It is measured ONCE per install/upgrade, while no engine or model is resident, and persisted — never a live reading. … A live figure would move with every resample and would count a resident model against the very host that serves it.

So the field's contract is "measured once, while nothing of ours is resident, persisted" — not a gauge.

What lands

Three deliberate non-changes:

  • EffectiveVRAMMB does not move. min_vram_mb, engine selection and vLLM's TP=1 fallback were authored against a whole card; the pool decision already settled that widening or narrowing it moves all three.
  • Unified-memory hosts are untouched. UsableVRAMMB is already the honest bound on a shared pool, and no shipped detector reports free memory for one — there is nothing to improve and a fallback to guess at.
  • A reading at or above the total is ignored. A bogus figure can only ever de-rate a device that was measured, never inflate one.

This PR is behaviourally inert

0 means "no free reading" and falls back to the total at every level. Nothing produces a non-zero value until the reader lands, so today's fleet and any driver that will not answer keep exactly today's budget. That is what lets the contract be published and tagged without waiting on the reader, while the ratchet points only at a settled surface.

Test inversion — declared, per §Test discipline

TestOllamaBudgetNeverShrinksTheHost asserted OllamaVRAMBudgetMB() >= EffectiveVRAMMB() over an exhaustive sweep. That invariant is intentionally given up here and split in two:

  • TestOllamaBudgetNeverShrinksAHostItDidNotMeasure — the half that still holds, and the important one: an unmeasured host may never be de-rated.
  • TestOllamaBudgetNeverFallsBelowWhatWasMeasured — the replacement: the budget may de-rate to what the driver reported, never past it.

The floor did not disappear; it changed what it is measured against. That is a revision of §4 only of the pool decision, recorded in docs/decisions/20260813/1120-ollama-budget-sized-on-free-vram.md with links both ways. 20260727/1830 stays accepted — §1–§3 are untouched — and its ## Status says which part moved.

The producer debt is declared in three places, not hidden

Two CI guards caught this PR publishing a contract with no writer, which is the #180/#251 shape:

  • TestHardwareSummaryFor_PublishesEveryWireField → an entry in notPublishedByAgent
  • protoconsumer → two entries in exemptions.go

The reader PR deletes all three — protoconsumer fails with "something under cmd/, internal/ now writes it — delete the entry" the moment a producer appears, so the debt cannot be quietly left standing.

One naming note: hostfit's field is spelled VRAMAvailableMB, not VRAMFreeMB, partly to keep protoconsumer honest. That guard matches producers by field name, so a Device.VRAMFreeMB assigned in FromHardwareSummary would have read as a proto-internal producer for the wire field of the same name and the debt would never have been visible in its table at all. LocalModelChoiceAt (#647) was named for the same consideration. The driver's own word stays on the wire, where nvidia-smi's memory.free is what it reports; available matches Host.RAMAvailableGB for the same quantity.

Verification

Full local gate, all green:

  • scripts/dev/ci-lint-local.sh — 20 of 20
  • proto-additive-guard.sh — OK (all published API intact, additions are omitempty); it initially rejected Device.VRAMAvailableMB for lacking an explicit tag, which is why it carries json:"-"
  • protoconsumer — OK (292 exported proto fields, 227 with a producer, 65 declared)
  • decision-log-guard.py — decision log OK: 73 件
  • gofmt -l clean · go vet ./... clean (both modules) · golangci-lint run (v2.12.2, after cache clean) → 0 issues
  • go test ./... root and proto/ — pass · go build -tags prod ./... · make verify-cross — exit 0

Byte-identity is pinned by TestHardwareSummary_VRAMFree_CanonicalJSON: an unmeasured device encodes byte-for-byte as it did before this PR, a measured one places the key between vram_total_mb and compute_cap, and a pre-addition payload parses with the field zero.

Rebased onto 09442ae.

Follow-up, in order

  1. Tag proto/v0.2.46 (automated per proto merge).
  2. Reader PR — the per-OS free-VRAM reader, the two adapters, and deletion of all three debt entries. Windows is nearly free: gpu_nvidia_windows.go:170-172 already reads NVML's {total, free, used} and discards free.
  3. CP bump in the private repo — including the strip for pollers that have not declared vram-free-v1.

docs-not-needed: wire contract and fit arithmetic only; no user-visible surface changes, and the change is inert until a producer lands.

Refs #69, #264, #568

…ard's total (#69)

The engine decides placement against free VRAM — ollama 0.31.1's
availableMemoryForLoad sums gpu.FreeMemory — while this repo sized its
budget against the total. A card that is also driving a display was
therefore valued at more than it can lend: the model and its context are
chosen optimistically, the load spills, and #621's post-load verify has
to shrink the window and restart the engine. The selection itself is
never corrected, because that verify can only move the window.

Adds the contract half, alone, as docs/decisions/20260719/0000 §2
requires:

  - signer.HardwareGPUSummary.VRAMFreeMB (vram_free_mb, omitempty),
    gated behind the new CapabilityVRAMFreeV1. The gate is not optional:
    the field is agent-reported and rides the signed NetworkMap on every
    PEER entry, so an agent that does not know it drops the key on
    canonical re-marshal and fails verification — the shape
    CapabilityRAMAvailableV1 already exists for.
  - hostfit.Device.VRAMAvailableMB and hostfit.Host.VRAMAvailable0MB,
    with Device.lendableMB() and Host.ollamaSingleDeviceMB() applying
    the de-rate per device, before the sum. VRAMAvailable0MB is what
    reaches a SINGLE-GPU host, which is the shape #69 actually reported;
    a one-card host has no pool to carry the reading.

Three things it deliberately does not do. EffectiveVRAMMB does not move:
min_vram_mb, engine selection and vLLM's TP=1 fallback were authored
against a whole card, as the pool decision recorded. Unified-memory
hosts are untouched, since UsableVRAMMB is already the honest bound and
no shipped detector reports free memory for one. And a reading at or
above the total is ignored, so a bogus figure can only ever de-rate a
device that was measured, never inflate one.

0 means "no free reading" and falls back to the total everywhere, so
this is behaviourally inert until a producer exists: today's fleet, and
any driver that will not answer, keep exactly today's budget.

The floor changed what it measures against, from the device's total to
its free figure, which is a revision of §4 of the pool decision rather
than a new rule. docs/decisions/20260813/1120 records it and links both
ways; the earlier decision stays accepted, since only §4 moved.
TestOllamaBudgetNeverShrinksTheHost is inverted accordingly and split:
an unmeasured host still may never be de-rated, and a measured one may
never fall below what was actually measured.

The producer debt is declared in three places rather than hidden —
notPublishedByAgent and two protoconsumer entries — and the reader PR
deletes them. hostfit's field is spelled "available" rather than "free"
partly to keep protoconsumer honest: that guard matches producers by
field NAME, and a Device.VRAMFreeMB assigned in FromHardwareSummary
would have read as a proto-internal producer for the wire field of the
same name, hiding the debt entirely.

Refs #69, #264

Signed-off-by: gen16k <gen16k@users.noreply.github.com>
@gen16k
gen16k merged commit 800b119 into main Aug 13, 2026
21 checks passed
gen16k added a commit that referenced this pull request Aug 13, 2026
… about local inference (#263, #225, #70, #35, #203, #69) (#773)

Five inference issues where the agent reported something untrue about local inference. Every premise was re-checked against origin/main first, and four of the five had moved — three would have produced the wrong change if implemented as written.

#263: huggingface_hub 1.x removed the [cli] extra, so the pin that was supposed to make the console script a hard guarantee resolved to plain huggingface_hub with a warning — working for exactly the reason the pin existed to stop relying on. The request now states the real requirement and the verify stage asserts the binary is there, which is the form that cannot go stale.

#225: the issue's stated cause (a PATH probe) was already fixed by #238, and its proposed signature does not exist. What remained is that this was the last engine-presence site not routed through engineInstalledOnHost, reading a 30 s-cached profile that engine_resolve.go already documents as too late for a fresh install.

#70 + #35: only two-step plans ever verified that a model reached VRAM, so a detected GPU that failed to engage kept its label while inference ran on the CPU. #35's proposed fix would have changed nothing — Accelerators has no production reader — so both close through one mechanism. Single-step plans relabel without restarting and without forcing a load. detectApple's swallowed system_profiler failure is now the warning VendorDetector's contract requires.

#203: proposals 1 and 3 were already implemented and pinned by tests citing the issue. Proposal 2 was broken on two surfaces — the wizard row flattened every failure to internal, and the boot benchmark reached no surface at all. Both are fixed; the boot result is reported but never persisted, and the two endings that are not verdicts about the host stay excluded.

#69: the reader paying the debt #772 declared. Free VRAM is read (appended to the nvidia-smi query; already in hand from NVML on Windows) and frozen after the first reading per device, because a live figure re-sampled after our own engine loads weights would exclude them and shrink the budget on every re-tune.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant