proto(hostfit): size the ollama VRAM budget on free memory, not the card's total (#69) - #772
Merged
Merged
Conversation
…ard's total (#69) The engine decides placement against free VRAM — ollama 0.31.1's availableMemoryForLoad sums gpu.FreeMemory — while this repo sized its budget against the total. A card that is also driving a display was therefore valued at more than it can lend: the model and its context are chosen optimistically, the load spills, and #621's post-load verify has to shrink the window and restart the engine. The selection itself is never corrected, because that verify can only move the window. Adds the contract half, alone, as docs/decisions/20260719/0000 §2 requires: - signer.HardwareGPUSummary.VRAMFreeMB (vram_free_mb, omitempty), gated behind the new CapabilityVRAMFreeV1. The gate is not optional: the field is agent-reported and rides the signed NetworkMap on every PEER entry, so an agent that does not know it drops the key on canonical re-marshal and fails verification — the shape CapabilityRAMAvailableV1 already exists for. - hostfit.Device.VRAMAvailableMB and hostfit.Host.VRAMAvailable0MB, with Device.lendableMB() and Host.ollamaSingleDeviceMB() applying the de-rate per device, before the sum. VRAMAvailable0MB is what reaches a SINGLE-GPU host, which is the shape #69 actually reported; a one-card host has no pool to carry the reading. Three things it deliberately does not do. EffectiveVRAMMB does not move: min_vram_mb, engine selection and vLLM's TP=1 fallback were authored against a whole card, as the pool decision recorded. Unified-memory hosts are untouched, since UsableVRAMMB is already the honest bound and no shipped detector reports free memory for one. And a reading at or above the total is ignored, so a bogus figure can only ever de-rate a device that was measured, never inflate one. 0 means "no free reading" and falls back to the total everywhere, so this is behaviourally inert until a producer exists: today's fleet, and any driver that will not answer, keep exactly today's budget. The floor changed what it measures against, from the device's total to its free figure, which is a revision of §4 of the pool decision rather than a new rule. docs/decisions/20260813/1120 records it and links both ways; the earlier decision stays accepted, since only §4 moved. TestOllamaBudgetNeverShrinksTheHost is inverted accordingly and split: an unmeasured host still may never be de-rated, and a measured one may never fall below what was actually measured. The producer debt is declared in three places rather than hidden — notPublishedByAgent and two protoconsumer entries — and the reader PR deletes them. hostfit's field is spelled "available" rather than "free" partly to keep protoconsumer honest: that guard matches producers by field NAME, and a Device.VRAMFreeMB assigned in FromHardwareSummary would have read as a proto-internal producer for the wire field of the same name, hiding the debt entirely. Refs #69, #264 Signed-off-by: gen16k <gen16k@users.noreply.github.com>
gen16k
added a commit
that referenced
this pull request
Aug 13, 2026
… about local inference (#263, #225, #70, #35, #203, #69) (#773) Five inference issues where the agent reported something untrue about local inference. Every premise was re-checked against origin/main first, and four of the five had moved — three would have produced the wrong change if implemented as written. #263: huggingface_hub 1.x removed the [cli] extra, so the pin that was supposed to make the console script a hard guarantee resolved to plain huggingface_hub with a warning — working for exactly the reason the pin existed to stop relying on. The request now states the real requirement and the verify stage asserts the binary is there, which is the form that cannot go stale. #225: the issue's stated cause (a PATH probe) was already fixed by #238, and its proposed signature does not exist. What remained is that this was the last engine-presence site not routed through engineInstalledOnHost, reading a 30 s-cached profile that engine_resolve.go already documents as too late for a fresh install. #70 + #35: only two-step plans ever verified that a model reached VRAM, so a detected GPU that failed to engage kept its label while inference ran on the CPU. #35's proposed fix would have changed nothing — Accelerators has no production reader — so both close through one mechanism. Single-step plans relabel without restarting and without forcing a load. detectApple's swallowed system_profiler failure is now the warning VendorDetector's contract requires. #203: proposals 1 and 3 were already implemented and pinned by tests citing the issue. Proposal 2 was broken on two surfaces — the wizard row flattened every failure to internal, and the boot benchmark reached no surface at all. Both are fixed; the boot result is reported but never persisted, and the two endings that are not verdicts about the host stay excluded. #69: the reader paying the debt #772 declared. Free VRAM is read (appended to the nvidia-smi query; already in hand from NVML on Windows) and frozen after the first reading per device, because a live figure re-sampled after our own engine loads weights would exclude them and shrink the budget on every re-tune.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The engine sizes placement against free VRAM — ollama 0.31.1's
availableMemoryForLoadsumsgpu.FreeMemory— while this repo sized its budget against the total.docs/knowledges/20260727/1830-ollama-multi-gpu-placement.mdrecorded the mismatch when it read the pinned engine, verbatim:A card that is also driving a display is therefore valued at more than it can lend. The model and its context get chosen optimistically, the load spills, and #621's post-load verify shrinks the window and restarts the engine. The selection is never corrected — that verify can only move the window, not un-pick a model.
This is the contract half, on its own, as
docs/decisions/20260719/0000-concurrent-proto-development.md§2 requires. The settled field table is on #69.Measured before designing
Three premises in the issue were stale, and the arithmetic had moved into
proto/hostfitsince it was written. Two things that reading turned up changed the design:The floor would have swallowed the whole fix.
OllamaVRAMBudgetMBclamps at the single-device figure. A one-card host — the shape #69 actually reports, an 8 GB card driving a display — hasVRAMPoolMB == 0, so the budget falls through toEffectiveVRAMMBand a free reading never enters the arithmetic at all. De-rating only inside the pool would have been a no-op on the reported case.A live reading would spiral. The hardware profile is re-sampled on a TTL, including after our own engine loads a model. A free figure taken then excludes our own weights, so each re-tune would see less memory and shrink further.
RAMAvailableGBalready names this hazard in the same words (#568):So the field's contract is "measured once, while nothing of ours is resident, persisted" — not a gauge.
What lands
signer.HardwareGPUSummary.VRAMFreeMB(vram_free_mb,omitempty), gated behind a newCapabilityVRAMFreeV1. The gate is not a formality: the field is agent-reported and rides the signed NetworkMap on every peer entry, so an agent that does not know it drops the key on canonical re-marshal and fails verification.CapabilityRAMAvailableV1exists for exactly this shape. My first field table said no capability was needed — that was wrong, and inference: Ollama VRAM budget uses total VRAM, not runtime-free VRAM, on discrete desktop GPUs #69 carries the correction.hostfit.Device.VRAMAvailableMBandhostfit.Host.VRAMAvailable0MB, withDevice.lendableMB()andHost.ollamaSingleDeviceMB()applying the de-rate per device, before the sum — which is what hostfit: the ollama VRAM budget is the first GPU only, so a multi-GPU host is judged as if the extra cards were not there #264's record asked inference: Ollama VRAM budget uses total VRAM, not runtime-free VRAM, on discrete desktop GPUs #69 to do rather than compensating on the pooling side.Three deliberate non-changes:
EffectiveVRAMMBdoes not move.min_vram_mb, engine selection and vLLM's TP=1 fallback were authored against a whole card; the pool decision already settled that widening or narrowing it moves all three.UsableVRAMMBis already the honest bound on a shared pool, and no shipped detector reports free memory for one — there is nothing to improve and a fallback to guess at.This PR is behaviourally inert
0means "no free reading" and falls back to the total at every level. Nothing produces a non-zero value until the reader lands, so today's fleet and any driver that will not answer keep exactly today's budget. That is what lets the contract be published and tagged without waiting on the reader, while the ratchet points only at a settled surface.Test inversion — declared, per §Test discipline
TestOllamaBudgetNeverShrinksTheHostassertedOllamaVRAMBudgetMB() >= EffectiveVRAMMB()over an exhaustive sweep. That invariant is intentionally given up here and split in two:TestOllamaBudgetNeverShrinksAHostItDidNotMeasure— the half that still holds, and the important one: an unmeasured host may never be de-rated.TestOllamaBudgetNeverFallsBelowWhatWasMeasured— the replacement: the budget may de-rate to what the driver reported, never past it.The floor did not disappear; it changed what it is measured against. That is a revision of §4 only of the pool decision, recorded in
docs/decisions/20260813/1120-ollama-budget-sized-on-free-vram.mdwith links both ways.20260727/1830staysaccepted— §1–§3 are untouched — and its## Statussays which part moved.The producer debt is declared in three places, not hidden
Two CI guards caught this PR publishing a contract with no writer, which is the #180/#251 shape:
TestHardwareSummaryFor_PublishesEveryWireField→ an entry innotPublishedByAgentprotoconsumer→ two entries inexemptions.goThe reader PR deletes all three —
protoconsumerfails with "something under cmd/, internal/ now writes it — delete the entry" the moment a producer appears, so the debt cannot be quietly left standing.One naming note: hostfit's field is spelled
VRAMAvailableMB, notVRAMFreeMB, partly to keepprotoconsumerhonest. That guard matches producers by field name, so aDevice.VRAMFreeMBassigned inFromHardwareSummarywould have read as a proto-internal producer for the wire field of the same name and the debt would never have been visible in its table at all.LocalModelChoiceAt(#647) was named for the same consideration. The driver's own word stays on the wire, wherenvidia-smi'smemory.freeis what it reports;availablematchesHost.RAMAvailableGBfor the same quantity.Verification
Full local gate, all green:
scripts/dev/ci-lint-local.sh— 20 of 20proto-additive-guard.sh—OK (all published API intact, additions are omitempty); it initially rejectedDevice.VRAMAvailableMBfor lacking an explicit tag, which is why it carriesjson:"-"protoconsumer—OK (292 exported proto fields, 227 with a producer, 65 declared)decision-log-guard.py—decision log OK: 73 件gofmt -lclean ·go vet ./...clean (both modules) ·golangci-lint run(v2.12.2, aftercache clean) →0 issuesgo test ./...root andproto/— pass ·go build -tags prod ./...·make verify-cross— exit 0Byte-identity is pinned by
TestHardwareSummary_VRAMFree_CanonicalJSON: an unmeasured device encodes byte-for-byte as it did before this PR, a measured one places the key betweenvram_total_mbandcompute_cap, and a pre-addition payload parses with the field zero.Rebased onto
09442ae.Follow-up, in order
proto/v0.2.46(automated per proto merge).gpu_nvidia_windows.go:170-172already reads NVML's{total, free, used}and discardsfree.vram-free-v1.docs-not-needed: wire contract and fit arithmetic only; no user-visible surface changes, and the change is inert until a producer lands.
Refs #69, #264, #568