runtime: strix halo on Windows keeps Vulkan on the numbers, not on ROCm's absence (#1233) - #1247
Merged
Merged
Conversation
…Cm's absence (#1233) The Windows arm said "ROCm has no Windows APU support; Vulkan is the only GPU path", in a comment and in its Reason string. At ollama 0.33.3 that is false. The Windows overlay is now rocm_v7_1 and carries gfx1151, and a Ryzen AI Max+ 395 running it reports library=ROCm compute=gfx1151 ... type=iGPU total="76.8 GiB" and serves a 21.8 GB model entirely on the GPU. It is slower, which is the reason that survives. Same model, same `-c 32768 -np 1 -b 1024 -ub 1024`, warm, ~9.9k-token prompt: Vulkan 900 tok/s prefill, 57.4 tok/s decode, 102.2 GiB exposed ROCm 819 tok/s prefill, 51.8 tok/s decode, 76.8 GiB exposed ROCm+HSA_OVERRIDE_GFX_VERSION=11.5.1 858 / 51.5 The exposed column carries as much weight as the rates: ROCm sees less of the unified pool, so a model that fits under Vulkan can fail to fit under ROCm on the same machine. No behaviour changed, and adding a ROCm step would not change any either: with both backends on disk ollama discovers each and picks Vulkan itself — four restarts, one of them with the HSA override set, all dispatched [{ID:0 Library:Vulkan}]. The step would only make the installer fetch an overlay nothing then uses, since WantsROCm is what wantROCmOverlay asks. The MAINTENANCE block above amdROCmSupportedRes now separates the half that is settled from the half that is not: gfx1151 measured, RDNA4 still open because no RX 9000 exists in reach and it fails safe meanwhile. Two smaller corrections found on the way. The overlay is ~250 MB (247 MB at 0.33.3), not the ~300 MB three comments claimed — checked against the asset. And the note records the trap that made the first experiment worthless: OLLAMA_VULKAN and HSA_OVERRIDE_GFX_VERSION do not switch the backend, so three "different" runs all served on Vulkan and returned near-identical numbers that read as "the backend does not matter here". The dispatch line tells you; the rates do not. Refs #1193 Refs waired-ai/waired#1312 Signed-off-by: gen16k <gen16k@users.noreply.github.com>
This was referenced Sep 5, 2026
gen16k
added a commit
that referenced
this pull request
Sep 6, 2026
… numbers that survive the spread The owner rejected #1247's ROCm conclusion on two grounds, and both were right. Prior verification had concluded ROCm did not work on Windows, and a 102.2 GiB / 76.8 GiB difference in exposed memory looked wrong. Reading my own logs back, the evidence for "ROCm engages" was ollama's dispatch label and /api/ps — neither of them llama.cpp's accounting. The line I had quoted, `load_tensors: ROCm_Host model buffer size`, is a HOST buffer. The device line was never read. It is read now, with a CPU-only control obtained by moving both backends out of lib/ollama: ROCm load_tensors: ROCm0 model buffer size = 21171.18 MiB, GPU 453 % Vulkan load_tensors: Vulkan0 model buffer size = 21171.18 MiB, GPU 112 % CPU only size_vram 0, GPU 0 %, prefill 158.1, decode 19.33 So both really run on the GPU. The performance claim did not survive. #1247 measured one turn per backend at num_predict 64 — a 1.2-second decode window — and the same Vulkan configuration moved 22 % between sessions under it, which is twice the gap it claimed. On the owner's suggestion the measurement was sized up: 30-36k-token prompts, num_predict 512, the backends alternated inside each run, the cold turn after each load discarded, six samples each. prefill Vulkan 876.8 (839.4-912.6) ROCm 636.3 (580.0-654.1) decode Vulkan 49.2 ( 47.9- 50.1) ROCm 43.8 ( 43.1- 44.3) Neither range overlaps. Vulkan by 37.8 % and 12.3 %, and the decode spread fell to 2.7 % — that collapse is the whole reason the gap is readable. The conclusion #1247 reached survives; the numbers it reached it with do not, and are replaced. The history is corrected in the same pass. #1247 said the arm's "ROCm has no Windows APU support" stopped being true at 0.33.3. Reading the zip central directory for five releases shows gfx1151 in the overlay back to v0.12.6, and starting v0.31.1 and v0.33.2 on the host shows both discovering ROCm exactly as 0.33.3 does. What DID change is upstream of that: the 2026-05 survey (waired's own records) found the overlay was `lib/ollama/rocm` at ROCm v6.1 and AMD documenting Ryzen AI APUs as excluded, so the claim was written against a real state of the world and the stamp went stale rather than the claim being careless. Also finishes two claims #1235 narrowed in one copy each: the cached-token comments in handlers.go and eventring.go, and the untuned-default comments in ollama.go and inference.go. The knowledge note on the second records why that one is contained — every consumer gates on ContextLength > 0, so the engine's own default reaches no advertisement, no measurement and no verdict. Refs #1247 Refs #1235 Refs waired-ai/waired#1312 Signed-off-by: gen16k <gen16k@users.noreply.github.com>
gen16k
added a commit
that referenced
this pull request
Sep 6, 2026
…t bump knows what to re-read (#1233) Two facts arrived after the performance re-measurement, and they change which sentence the arm should lead with. ROCm on gfx1151 has open upstream defects that a coding agent meets on every request. ollama/ollama#17895 has it returning wrong output above ~4k prompt tokens — "fluent, confident, wrong answers" with nothing logged, and past ~8k byte-identical replies to different prompts — reproduced across 0.32.5 to 0.32.14, on Debian and Windows, on three model families, with the bundled rocm_v7_2. ollama/ollama#17847 has it bleeding KV state between sequential requests at OLLAMA_NUM_PARALLEL=1. Both report the same machine clean on Vulkan and on CPU. Vulkan has one of its own, #17870, and the difference in kind is what decides this: it FAILS the request where ROCm answers wrongly, and a wrong answer with no error is caught by nothing this repository runs. So speed becomes the lesser reason and is labelled as what it is — a measurement of ollama's own bundled ROCm, which is a stock HIP build. Tuned gfx1151 builds exist and report the prefill split going the other way, so the figures here are a fact about the engine we ship rather than about ROCm. The rest is a standing checklist rather than a proposal. The MAINTENANCE block gains the four upstream threads and the ROCm 7.2.4 question, with the warning that 7.2.4 answers performance and not correctness — #17895 already reproduces on 7.2 — so the two must not be read as one. The new knowledge note carries the background a bump needs to interpret what it finds: which overrides are bug workarounds and which are preferences, why amdROCmSupportedRes is a download decision that Linux answers without a list, what the list actually matches when run against real SKU strings, why it cannot be derived from the overlay while hardware.GPU carries no gfx target, and the method — including the three harness traps that cost this investigation two false conclusions before it produced a true one. Nothing is recommended and nothing changes: the arm still names Vulkan, and it now says why in terms that expire when the upstream threads do. Refs #1247 Refs #1248 Refs waired-ai/waired#1312 Signed-off-by: gen16k <gen16k@users.noreply.github.com>
gen16k
added a commit
that referenced
this pull request
Sep 6, 2026
…d the next bump knows what to re-read (#1233, #1193) (#1250) * runtime: ROCm did not start working on Strix Halo at 0.33.3 — it always did (#1233) #1247 said the Windows arm's "ROCm has no Windows APU support" stopped being true at 0.33.3, because the overlay "is now rocm_v7_1 and carries gfx1151". That "now" was an inference from the comment's own version stamp, and it is wrong. The overlay asset was read for five releases — 0.31.1, 0.32.13, 0.32.15, 0.33.2 and 0.33.3. All five unpack to rocm_v7_1 and all five carry the same targets: gfx906, gfx906-xnack-, gfx1030, gfx1100, gfx1101, gfx1102, gfx1150, gfx1151, gfx1200, gfx1201. The overlay's contents are not what changed, and the stamp "Ollama 0.31.x, ROCm v6.1 overlay" was wrong about the version it named. Shipping kernels is not engaging, so that was measured too: 0.31.1 and 0.33.2 were each unpacked on the same Ryzen AI Max+ 395 with their own overlay and started with OLLAMA_IGPU_ENABLE=1 — no model needed, the startup discovery line is enough. Both emit, character for character, what 0.33.3 emits: library=ROCm compute=gfx1151 ... type=iGPU total="76.8 GiB" available="76.6 GiB" So the claim was not overtaken by an upstream change. It was false on every engine this product has pinned, at least back to the version its own stamp named. Nothing else in #1247 moves: Vulkan is still faster here, ollama still picks it by itself with both backends on disk, and the reason the arm now gives — measured faster, not ROCm absent — is still the right one. Only the history was wrong. The knowledge note takes a dated 補足 rather than a rewrite, so the record shows what was believed and what replaced it, and it carries the lesson one level up from its own §4: a version stamp in a comment is not evidence about that version. Both checks were cheap — one ranged request for a zip's central directory, one 14-second model-less server start — so there was no reason not to make them first. Refs #1247 Refs waired-ai/waired#1312 Signed-off-by: gen16k <gen16k@users.noreply.github.com> * gateway, runtime: finish the two claims the pin bump narrowed in one copy each (#1193) #1235 narrowed convert.go's "ollama has no equivalent on any surface" and ollamaContextFloor's "the pinned engine's own default is 32768". Both statements existed in more than one place, and the other copies were left saying what the measurement had just disproved. The cached-token claim also lived in internal/gateway/handlers.go (setCachedInput: "every engine but vLLM with --enable-prompt-tokens-details") and in internal/observability/ eventring.go (CachedInputTokens: "only vLLM reports the breakdown"). Both now say when that stopped being true. This matters more than the wording: CachedInputTokens feeds waired_inference_cached_input_tokens_total, whose Help already says "Only engines that report a prompt-token breakdown move this" — so ollama hosts silently started moving a counter two comments said they could not. The untuned-default claim lived in internal/runtime/ollama.go and cmd/waired-agent/inference.go, both describing the #624 boot-order gap as costing "its 32k default". 0.33.3 derives that default from the VRAM it found instead — measured 32768 on a 37.4 GiB Mac and 262144 on a 102.2 GiB Strix Halo — so the window an untuned spawn takes on a large host is now far bigger than the sentence implies. The tuned path is unaffected, which is why nothing here changes behaviour; the gap's cost simply is not what it says. Found by tracing what the two narrowed claims actually feed, rather than by grep: the question "what would change for the product" is what turned up the counter and the untuned path. Refs #1235 Refs waired-ai/waired#1312 Signed-off-by: gen16k <gen16k@users.noreply.github.com> * docs(knowledges): the engine's default window reaches nothing — every consumer gates on ContextLength > 0 (#1193) ollama 0.33.3 derives its own default context from the VRAM it sees (32768 on the 37.4 GiB / 23.8 GiB hosts, 262144 on the 102.2 GiB Strix Halo). Record why that changed only comments: the tuned path never consults the engine default, the untuned path (#624 boot-order gap) is narrow, and all five readers of ModelTuning.ContextLength return nothing when it is 0, so an untuned engine advertises no window, produces no measurement, and triggers no verdict. What is not established — an untuned spawn that the larger default makes fail — is said so. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0167iiPnKQz1qcuz2bhNBGep Signed-off-by: gen16k <gen16k@users.noreply.github.com> * runtime: measure the Strix Halo backends properly, and keep Vulkan on numbers that survive the spread The owner rejected #1247's ROCm conclusion on two grounds, and both were right. Prior verification had concluded ROCm did not work on Windows, and a 102.2 GiB / 76.8 GiB difference in exposed memory looked wrong. Reading my own logs back, the evidence for "ROCm engages" was ollama's dispatch label and /api/ps — neither of them llama.cpp's accounting. The line I had quoted, `load_tensors: ROCm_Host model buffer size`, is a HOST buffer. The device line was never read. It is read now, with a CPU-only control obtained by moving both backends out of lib/ollama: ROCm load_tensors: ROCm0 model buffer size = 21171.18 MiB, GPU 453 % Vulkan load_tensors: Vulkan0 model buffer size = 21171.18 MiB, GPU 112 % CPU only size_vram 0, GPU 0 %, prefill 158.1, decode 19.33 So both really run on the GPU. The performance claim did not survive. #1247 measured one turn per backend at num_predict 64 — a 1.2-second decode window — and the same Vulkan configuration moved 22 % between sessions under it, which is twice the gap it claimed. On the owner's suggestion the measurement was sized up: 30-36k-token prompts, num_predict 512, the backends alternated inside each run, the cold turn after each load discarded, six samples each. prefill Vulkan 876.8 (839.4-912.6) ROCm 636.3 (580.0-654.1) decode Vulkan 49.2 ( 47.9- 50.1) ROCm 43.8 ( 43.1- 44.3) Neither range overlaps. Vulkan by 37.8 % and 12.3 %, and the decode spread fell to 2.7 % — that collapse is the whole reason the gap is readable. The conclusion #1247 reached survives; the numbers it reached it with do not, and are replaced. The history is corrected in the same pass. #1247 said the arm's "ROCm has no Windows APU support" stopped being true at 0.33.3. Reading the zip central directory for five releases shows gfx1151 in the overlay back to v0.12.6, and starting v0.31.1 and v0.33.2 on the host shows both discovering ROCm exactly as 0.33.3 does. What DID change is upstream of that: the 2026-05 survey (waired's own records) found the overlay was `lib/ollama/rocm` at ROCm v6.1 and AMD documenting Ryzen AI APUs as excluded, so the claim was written against a real state of the world and the stamp went stale rather than the claim being careless. Also finishes two claims #1235 narrowed in one copy each: the cached-token comments in handlers.go and eventring.go, and the untuned-default comments in ollama.go and inference.go. The knowledge note on the second records why that one is contained — every consumer gates on ContextLength > 0, so the engine's own default reaches no advertisement, no measurement and no verdict. Refs #1247 Refs #1235 Refs waired-ai/waired#1312 Signed-off-by: gen16k <gen16k@users.noreply.github.com> * runtime: the Strix Halo arm names Vulkan for correctness, and the next bump knows what to re-read (#1233) Two facts arrived after the performance re-measurement, and they change which sentence the arm should lead with. ROCm on gfx1151 has open upstream defects that a coding agent meets on every request. ollama/ollama#17895 has it returning wrong output above ~4k prompt tokens — "fluent, confident, wrong answers" with nothing logged, and past ~8k byte-identical replies to different prompts — reproduced across 0.32.5 to 0.32.14, on Debian and Windows, on three model families, with the bundled rocm_v7_2. ollama/ollama#17847 has it bleeding KV state between sequential requests at OLLAMA_NUM_PARALLEL=1. Both report the same machine clean on Vulkan and on CPU. Vulkan has one of its own, #17870, and the difference in kind is what decides this: it FAILS the request where ROCm answers wrongly, and a wrong answer with no error is caught by nothing this repository runs. So speed becomes the lesser reason and is labelled as what it is — a measurement of ollama's own bundled ROCm, which is a stock HIP build. Tuned gfx1151 builds exist and report the prefill split going the other way, so the figures here are a fact about the engine we ship rather than about ROCm. The rest is a standing checklist rather than a proposal. The MAINTENANCE block gains the four upstream threads and the ROCm 7.2.4 question, with the warning that 7.2.4 answers performance and not correctness — #17895 already reproduces on 7.2 — so the two must not be read as one. The new knowledge note carries the background a bump needs to interpret what it finds: which overrides are bug workarounds and which are preferences, why amdROCmSupportedRes is a download decision that Linux answers without a list, what the list actually matches when run against real SKU strings, why it cannot be derived from the overlay while hardware.GPU carries no gfx target, and the method — including the three harness traps that cost this investigation two false conclusions before it produced a true one. Nothing is recommended and nothing changes: the arm still names Vulkan, and it now says why in terms that expire when the upstream threads do. Refs #1247 Refs #1248 Refs waired-ai/waired#1312 Signed-off-by: gen16k <gen16k@users.noreply.github.com> --------- Signed-off-by: gen16k <gen16k@users.noreply.github.com> Co-authored-by: gen16k <gen16k@users.noreply.github.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
gen16k
added a commit
that referenced
this pull request
Sep 6, 2026
…ents stop contradicting the file they are in (#1261) The Strix Halo arms rest on upstream facts that move without notice, and #1250 wrote the recheck list for them. Three things it did not cover. ggml-org/llama.cpp#27856 (open, checked 2026-09-06) has qwen4exp — the architecture behind the Flash-Next entry #1259 just shipped — collapsing 3.5-4x in decode once context passes ~1k on HIP/gfx1151, plateauing at 5.5-6.1 tok/s where CUDA decays only mildly. It moves with the VENDORED LLAMA.CPP version rather than with ollama's release or its ROCm overlay, so it is on a different recheck axis from the four ollama threads already listed and ollama's release notes will never mention it. It also lands on the arm nobody was watching. The Windows arm already names Vulkan, so it never reaches HIP; the LINUX arm prefers ROCm, and the #290 probe cannot see this failure — it falls back only on positive evidence of CPU residency (size_vram == 0), and a model that is on the GPU but 4x slow at depth is not that. Conservative by design, but the failure shape is outside the arm. No Linux Strix Halo is in the fleet, so this is an upstream report, not an observation here; the note says so. Two stale comments in the same file, both left by #1247 correcting one copy of a figure and not the other: - BackendROCm's doc calls the Windows ROCm overlay ~350 MiB. WantsROCm, sixty lines below it, already carries the measured 247 MB at 0.33.3, as does the knowledge note. - The backend table test names its case "vulkan only (no ROCm on Win APU)". The arm it pins says the opposite in as many words — "Vulkan, because it is FASTER here — not because ROCm is absent" — because ROCm does engage gfx1151 under Windows. The assertion was right; the name was the last copy of the claim #1233 disproved. No behaviour change: comments, one test case name, and a knowledge-note section. Refs #1233 Refs #1247 Refs #1255 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0167iiPnKQz1qcuz2bhNBGep Signed-off-by: gen16k <gen16k@users.noreply.github.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
#1233 が求めた測定です。製品の挙動は変えません — 変えるのはコメントと
Reason文字列、それに測定で分かった数字の訂正だけです。偽だった主張
Strix Halo の Windows の腕は、コメントでもユーザーに見える
Reason文字列でもROCm has no Windows APU support; Vulkan is the only GPU pathと言っていました。ollama 0.33.3 では成立しません。Windows の overlay は
lib/ollama/rocm_v7_1(コメントの刻印は v6.1)になり、rocBLAS カーネルに gfx1151 が入りました。実機のディスク上で確認しています。overlay を置いた状態でエンジンはこう報告します:Vulkan を退避させると 21.80 GB のモデルを 100% GPU で配ります(
[{ID:0 Library:ROCm}]/load_tensors: ROCm_Host//api/psはsize_vram=size)。生き残る理由 — 遅いこと
同一モデル・同一
-c 32768 -np 1 -b 1024 -ub 1024(エンジンのバッチ選択は交絡していません)・warm・約 9,910 トークン:[{ID:0 Library:Vulkan}][{ID:0 Library:ROCm}]HSA_OVERRIDE_GFX_VERSION=11.5.1Vulkan が最良の ROCm 構成に対して prefill +4.9% / decode +11.5%。
露出 VRAM の列はレートと同じだけ重いです。ROCm はユニファイドプールを 76.8 GiB しか見せないので、同じ機械で Vulkan なら載るモデルが ROCm では載らないことがあります。深い文脈での逆転は測っていませんが、未検証のリスクの向きは Vulkan 有利(KV の置き場が広い)です。
ROCm ステップを足しても何も変わらない
両方ディスクに在ると ollama は自分で選び、Vulkan を選びます — 4 回の独立再起動(
OLLAMA_VULKAN=1あり / なし / HSA override あり / 素)で全部[{ID:0 Library:Vulkan}]。腕にステップを足しても応じる backend は変わらず、WantsROCmがwantROCmOverlayの問いなので overlay を余分に落とすだけになります。試したいユーザーの道はWAIRED_OLLAMA_GPU_MODE=rocmのままです。!!! MAINTENANCEブロックを片付けた#1193 で「#1233 が測定を持つ」と書いた場所を、解けた半分と解けていない半分に分けました。gfx1151 は測定済み。RDNA4(gfx1200/1201)は
amdROCmSupportedResに RX 9000 のパターンが無く、そのカードが手元に無いので open のまま — ただし安全側に落ちます(Vulkan は動く)。あわせて「この一覧が決めるのは overlay を落とすかどうかであって、両方在れば backend はエンジンが選ぶ」ことも書き足しました。ついでに直した 2 点
~300 MBと言っていましたが、実測は 247 MB(0.33.3)。アセットに当てて確認し、~250 MBに揃えました。OLLAMA_VULKANもHSA_OVERRIDE_GFX_VERSIONも backend を切り替えません。3 通り走らせて 905/905/903・57.8/57.6/58.0 という「ほぼ同一」の数字が出て、それは「ここでは backend は関係ない」と読めてしまいます。実際は 3 回とも Vulkan が応じていました。見るべきはレートではなくrunner.inference="[{ID:0 Library:...}]"とload_tensors: <Device> model buffer size。 さらにlib/ollama/vulkanをvulkan.offにリネームしても隠せません(libdirs=ollama,vulkan.off)。ディレクトリごと外へ動かす必要があります。記録:
docs/knowledges/20260906/0430-rocm-runs-on-strix-halo-windows.md実測環境: sv-evox2(Windows 11 26200 / Ryzen AI Max+ 395 / 127.15 GB UMA / ollama 0.33.3 base + ROCm overlay)、モデルは
qwen3.6:35b-a3b-q4_K_M。Fixes #1233
Refs #1193
Refs waired-ai/waired#1312
docs-not-needed: コメント・
Reason文字列・docs/knowledges/のみで、ユーザーが読む面は変わりません(Reasonはログとデバッグ出力にしか出ず、docs-site のどこにも引用されていないことを確認済み)。🤖 Generated with Claude Code
https://claude.ai/code/session_0167iiPnKQz1qcuz2bhNBGep