Studio: report host VRAM usage when no single GPU's usage can be attributed - #8481
Conversation
…ibuted On a Windows ROCm host with two asymmetric AMD cards the System tab and the floating monitor show Unknown used VRAM on every device row and on the aggregate tile. Windows shares no key between the LUID performance counter instances and torch ordinals, so usage is paired to devices by capacity ranking and kept only when capacity forces it. On a 45 GiB plus 7.98 GiB pair nothing at or below 7.98 GiB is forced, which is idle and every small model. The pairing is unresolvable but the sum is not: over a bijection the total is the same whichever way round the counters go. Report it as vram_used_gb_aggregate, and let the tile prefer the per-device sum, fall back to the aggregate, and otherwise stay Unknown rather than 0. Per-device rows are unchanged: a value still appears only when capacity forces it.
for more information, see https://pre-commit.ci
…sible set The aggregate dropped sub-threshold counters to force a one-per-device count, the same noise filter the per-device matcher uses. That is safe there, because per-device only ever emits a value capacity FORCES, but not for a sum, which emits every counter it kept. The instance names carry no vendor, LUID or PCI key, so a counter cannot be told apart from a foreign adapter's: a card hidden by HIP_VISIBLE_DEVICES, an iGPU, an NVIDIA card in the same box, or a Basic Render Driver placeholder over the cutoff. Whenever such an adapter is busier than one visible card is idle, the filter dropped the visible card and kept the foreign one, and the total gained bytes that are on no visible card. Six such cases, worst of them a single idle card beside a busy NVIDIA card reading as 6 GiB of AMD usage. Establish the set by cardinality instead. Get-Counter returns an instance for every WDDM adapter, so a visible card is always in the list, and a list exactly as long as the visible set contains those cards and nothing else. An unexplained instance now reports Unknown, which is narrower on purpose: a host total that is confidently wrong is worse than no total. Widening it needs the counters joined on LUID or PCI bus id rather than on capacity rank.
ROCm/legacy-rocm-build#1909 was the wrong source and is dropped from the last two places it was still quoted: it names no operating system, its repro is rocm-smi on Linux, and it closed with no fix. AMD documents the behaviour directly. The hipMemGetInfo reference warns "On Windows, the free memory only accounts for memory allocated by this process and may be optimistic", and an AMD engineer gives the mechanism in ROCm/librocdxg#57: WDDM virtualises video memory, so HIP reports a per-process budget and a fresh process sees free at or near total whatever else is resident. Measured there at 24410 MiB of 24560 on a deliberately filled card. Recorded as intended Windows behaviour rather than a defect, which is worth stating. Also records why WSL needs no branch. AMD keeps the WSL2 reading consistent with native Linux, where free tracks physical residency, and sys.platform is "linux" there, so the predicate is False and the accurate figure passes through uncapped. Two fixes alongside it. The aggregate now rides only on an index set identical to the payload's devices: the utilization and visibility probes enumerate separately, and _torch_get_per_device_info drops a device whose mem_get_info raises while the aggregate's enumeration keeps it, which would inflate the percentage and floor free at 0. And the frontend precedence test set its aggregate to the per-device sum, so it passed whichever source won and pinned nothing.
|
Codex Review: Didn't find any major issues. Hooray! Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
|
Codex Review: Didn't find any major issues. Bravo. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
|
Codex Review: Didn't find any major issues. You're on a roll. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
|
Codex Review: Didn't find any major issues. More of your lovely PRs please. Reviewed commit: ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
If Codex has suggestions, it will comment; otherwise it will react with 👍. Codex can also answer questions or update the PR. Try commenting "@codex address that feedback". |
The aggregate's safety argument rests on Get-Counter emitting exactly one GPU Adapter Memory instance per WDDM adapter, and the docstring asserted that as established fact. It is not confirmed against Microsoft's counter documentation, so it now says so. The behaviour is unchanged and the assumption fails closed: a second instance for one adapter makes the counter list longer than the visible set, which returns None rather than a total. Verified against the helper directly.
|
Correcting an overclaim in my own docstring, pushed as 2b9303d. Comment-only, AST-verified, behaviour unchanged. The aggregate's safety argument rests on It now says it is an assumption, and records that it fails closed. Checked against the helper directly: An extra instance makes the counter list longer than the visible set, which returns None rather than a total, so a wrong assumption costs an Unknown and never a wrong number. That is the right direction for this helper, since the whole reason it exists is that a confidently wrong host total is worse than Unknown. Worth confirming before anyone widens this: the durable fix is joining counters to devices on LUID or PCI bus id rather than on capacity rank. |
Problem
On a Windows ROCm host with two asymmetric AMD cards the System tab and the floating monitor show
Unknown / 53.0 GiBused VRAM, withUnknown usedandUnknown freeon every device row. Reported in #7452 (Radeon PRO W7900 45 GiB plus W7500 7.98 GiB, Windows 10, ROCm 7.13).Worth stating precisely, because the issue title is shorthand: per-device totals are correct in the reporter's own screenshot (
45.0 GiB total,7.98 GiB total). Only the usage side is blank. Totals come fromtorch.get_device_properties, which is reliable. Nothing here claims to fix a total.This is a regression from the fix for #7072 (#7238), by the same reporter six days later, and it cannot be fixed by reverting: #7238 exists because the old code fabricated per-device usage.
Windows shares no key between the LUID
GPU Adapter Memorycounter instances and torch ordinals, so #7238 pairs usage to devices by capacity ranking and keeps a value only when capacity forces it. On a 45 GiB plus 7.98 GiB pair, nothing at or below 7.98 GiB is forced, so idle and every small model report unknown, and the aggregate tile, which requires every device to be known, goes blank with them.Fix
The pairing is unresolvable but the sum is not: over a bijection the total is the same whichever way round the counters go. Report it.
_rocm_windows_aggregate_used_bytesemits the visible set's total used only when the counter list is that set, established by cardinality:Get-Counterreturns an instance for every WDDM adapter, so a visible card is always somewhere in the list, and a list exactly as long as the visible set therefore contains those cards and nothing else. Plus the existing check that no usage sits above its ranked capacity./api/systemcarry it asvram_used_gb_aggregate, and only when the utilization probe's index set matches the payload's devices, so the tile cannot divide a two-card total by a one-card capacity.Per-device rows are unchanged: a value still appears only when capacity forces it.
Why not the noise filter
The first cut of this reused the sub-threshold filter that
_match_adapter_used_to_devicesapplies, dropping counters under 64 MiB to force a one-per-device count. That is safe for per-device attribution, which only ever emits a value capacity forces, but it is not safe for a sum, which emits every counter it kept.The instance names carry no vendor, LUID or PCI key, so a retained counter cannot be told apart from a foreign adapter's. Whenever a foreign adapter is busier than one visible card is idle, the filter dropped the visible card and kept the foreign one, and the host total silently gained bytes that are on no visible card. Six cases, all reproduced against the code:
HIP_VISIBLE_DEVICES)The last two are the worst: a single card's capacity admits almost any foreign reading, and the per-device path returns unknown there, so the aggregate would have been the only wrong number on the page. All six now report Unknown. A host total that is confidently wrong is worse than no host total.
The cardinality rule is narrower on purpose, and it is narrow: any unexplained adapter instance suppresses the aggregate. Widening it needs the counters joined on LUID or PCI bus id rather than on capacity rank.
Sources
hipMemGetInforeference: "On Windows, the free memory only accounts for memory allocated by this process and may be optimistic."GPU Adapter Memoryas per adapter across all processes, againstGPU Process Memorywhich is per process, and warns that summing the per-process set double-counts shared memory. There is no published reference page for the counter set, and no documented LUID to ordinal mapping.ROCm/legacy-rocm-build#1909, which the repo used to cite for this, is dropped: it names no operating system, its repro is
rocm-smion Linux, and it closed with no fix.Not fixed here
Per-device attribution on an asymmetric pair cannot be recovered from this data source. It needs a real key. Correction to an earlier draft of this description: HIP does expose one,
hipDeviceProp_t.luidis populated on Windows, and theGPU Adapter Memoryinstance names are LUID-keyed, so the join is available in principle. It is PyTorch that surfaces nothing, so reaching it means a HIP call outside torch. Worth doing separately. Note the low half of the instance LUID is 8 hex digits (0x00012FBFobserved), andphys_Nis the index inside a linked adapter, not an adapter ordinal.Two coverage caveats stated plainly. If the reporter's host carries any adapter instance beyond its two cards, the aggregate stays Unknown there and this does not fix his tile; the counter dump needed to settle that is not in the issue. And if his counter query returns nothing at all, the pre-existing fallback already reported Unknown and this changes nothing either. Neither case regresses anything.
Verification
studio/backend/tests/test_rocm_windows_vram_7452.py(9 tests) drives the reporter's figures throughget_visible_gpu_utilizationandget_gpu_utilization, covers the six wrong-number cases above, a changing instance list across polls, a WDDM spill over the smaller card, and the aggregate rule as a unit.studio/frontend/tests/rocm-windows-vram-aggregate.test.ts(4 tests) covers the tile's source selection; its precedence case previously set the aggregate equal to the per-device sum, so it passed whichever source won, and now uses a differing value.Anti-vacuity, run rather than assumed: with only
hardware.pyreverted and the tests kept, the two new soundness tests fail and pass again on restore. On the frontend, deleting the new module only yields a module-not-found error, so the check was redone by restoring the pre-PR inline logic under the same exported surface: 2 of 4 fail, and they are the two #7452 behaviours.Upgrade order tested in both directions, executed over 28 payload combinations. A new frontend against an old backend, with the key absent and as
undefined,null,NaN, the string"0.36"andInfinity, renders Unknown in every case and never leaksundefined,NaNor a fabricated 0. An old frontend against a new backend is unaffected: the pre-PR code reads onlydevices[].vram_used_gb, and there is no schema validation on/api/systemthat could reject an unknown key (no zod, ajv, io-ts or similar anywhere in the frontend), so the extra key is inert.No new dependency, no persisted state, no schema change.
test_rocm_windows_vram_7072.py(15) andtest_rocm_multi_gpu_vram_system_wide.py(27) still pass, 51 in that slice altogether. Frontendnpm testis 1957 passing, and the new file is matched by the runner's glob, verified rather than assumed.There is no AMD or Windows CI in this repository, so torch, the performance counter and the platform are mocked, and none of this is hardware validation.
Fixes #7452