Skip to content

Studio: report host VRAM usage when no single GPU's usage can be attributed - #8481

Merged
danielhanchen merged 7 commits into
mainfrom
fix/vram_7452
Aug 12, 2026
Merged

danielhanchen merged 7 commits into
mainfrom
fix/vram_7452

Conversation

@danielhanchen

@danielhanchen danielhanchen commented Aug 11, 2026 •

Copy link
Copy Markdown
Member

Problem

On a Windows ROCm host with two asymmetric AMD cards the System tab and the floating monitor show Unknown / 53.0 GiB used VRAM, with Unknown used and Unknown free on every device row. Reported in #7452 (Radeon PRO W7900 45 GiB plus W7500 7.98 GiB, Windows 10, ROCm 7.13).

Worth stating precisely, because the issue title is shorthand: per-device totals are correct in the reporter's own screenshot (45.0 GiB total, 7.98 GiB total). Only the usage side is blank. Totals come from torch.get_device_properties, which is reliable. Nothing here claims to fix a total.

This is a regression from the fix for #7072 (#7238), by the same reporter six days later, and it cannot be fixed by reverting: #7238 exists because the old code fabricated per-device usage.

Windows shares no key between the LUID GPU Adapter Memory counter instances and torch ordinals, so #7238 pairs usage to devices by capacity ranking and keeps a value only when capacity forces it. On a 45 GiB plus 7.98 GiB pair, nothing at or below 7.98 GiB is forced, so idle and every small model report unknown, and the aggregate tile, which requires every device to be known, goes blank with them.

Fix

The pairing is unresolvable but the sum is not: over a bijection the total is the same whichever way round the counters go. Report it.

  • _rocm_windows_aggregate_used_bytes emits the visible set's total used only when the counter list is that set, established by cardinality: Get-Counter returns an instance for every WDDM adapter, so a visible card is always somewhere in the list, and a list exactly as long as the visible set therefore contains those cards and nothing else. Plus the existing check that no usage sits above its ranked capacity.
  • The utilization payloads and /api/system carry it as vram_used_gb_aggregate, and only when the utilization probe's index set matches the payload's devices, so the tile cannot divide a two-card total by a one-card capacity.
  • The VRAM tile prefers the per-device sum, falls back to the aggregate, and otherwise stays Unknown. It never falls back to 0, which was the [Bug] VRAM Usage in System Tab is wrong. #7072 symptom.

Per-device rows are unchanged: a value still appears only when capacity forces it.

Why not the noise filter

The first cut of this reused the sub-threshold filter that _match_adapter_used_to_devices applies, dropping counters under 64 MiB to force a one-per-device count. That is safe for per-device attribution, which only ever emits a value capacity forces, but it is not safe for a sum, which emits every counter it kept.

The instance names carry no vendor, LUID or PCI key, so a retained counter cannot be told apart from a foreign adapter's. Whenever a foreign adapter is busier than one visible card is idle, the filter dropped the visible card and kept the foreign one, and the host total silently gained bytes that are on no visible card. Six cases, all reproduced against the code:

counters visible cards truth filtered sum
30 GiB, 5 GiB, 30 MiB 45 + 8 GiB 30.03 GiB 35.0 GiB (a card hidden by HIP_VISIBLE_DEVICES)
30 GiB, 6 GiB, 20 MiB 45 + 8 GiB 30.02 GiB 36.0 GiB (an NVIDIA card in the same box)
30 GiB, 1 GiB, 10 MiB 45 + 8 GiB 30.01 GiB 31.0 GiB (an iGPU carveout)
30 GiB, 200 MiB, 20 MiB 45 + 8 GiB 30.02 GiB 30.20 GiB (a placeholder above the cutoff)
6 GiB, 30 MiB 45 GiB 0.03 GiB 6.0 GiB (one idle card, one busy NVIDIA)
6 GiB, 30 MiB, 3 MiB 45 GiB 0.03 GiB 6.0 GiB

The last two are the worst: a single card's capacity admits almost any foreign reading, and the per-device path returns unknown there, so the aggregate would have been the only wrong number on the page. All six now report Unknown. A host total that is confidently wrong is worse than no host total.

The cardinality rule is narrower on purpose, and it is narrow: any unexplained adapter instance suppresses the aggregate. Widening it needs the counters joined on LUID or PCI bus id rather than on capacity rank.

Sources

  • AMD's hipMemGetInfo reference: "On Windows, the free memory only accounts for memory allocated by this process and may be optimistic."
  • ROCm/librocdxg#57, open. Its title is about WSL2, but in it an AMD engineer gives the mechanism: WDDM virtualises video memory, so HIP reports a per-process budget and a fresh process sees free at or near total whatever else is resident. Measured at 24410 MiB of 24560 on a deliberately filled card. Recorded there as intended Windows behaviour, not a defect.
  • Microsoft documents GPU Adapter Memory as per adapter across all processes, against GPU Process Memory which is per process, and warns that summing the per-process set double-counts shared memory. There is no published reference page for the counter set, and no documented LUID to ordinal mapping.

ROCm/legacy-rocm-build#1909, which the repo used to cite for this, is dropped: it names no operating system, its repro is rocm-smi on Linux, and it closed with no fix.

Not fixed here

Per-device attribution on an asymmetric pair cannot be recovered from this data source. It needs a real key. Correction to an earlier draft of this description: HIP does expose one, hipDeviceProp_t.luid is populated on Windows, and the GPU Adapter Memory instance names are LUID-keyed, so the join is available in principle. It is PyTorch that surfaces nothing, so reaching it means a HIP call outside torch. Worth doing separately. Note the low half of the instance LUID is 8 hex digits (0x00012FBF observed), and phys_N is the index inside a linked adapter, not an adapter ordinal.

Two coverage caveats stated plainly. If the reporter's host carries any adapter instance beyond its two cards, the aggregate stays Unknown there and this does not fix his tile; the counter dump needed to settle that is not in the issue. And if his counter query returns nothing at all, the pre-existing fallback already reported Unknown and this changes nothing either. Neither case regresses anything.

Verification

studio/backend/tests/test_rocm_windows_vram_7452.py (9 tests) drives the reporter's figures through get_visible_gpu_utilization and get_gpu_utilization, covers the six wrong-number cases above, a changing instance list across polls, a WDDM spill over the smaller card, and the aggregate rule as a unit. studio/frontend/tests/rocm-windows-vram-aggregate.test.ts (4 tests) covers the tile's source selection; its precedence case previously set the aggregate equal to the per-device sum, so it passed whichever source won, and now uses a differing value.

Anti-vacuity, run rather than assumed: with only hardware.py reverted and the tests kept, the two new soundness tests fail and pass again on restore. On the frontend, deleting the new module only yields a module-not-found error, so the check was redone by restoring the pre-PR inline logic under the same exported surface: 2 of 4 fail, and they are the two #7452 behaviours.

Upgrade order tested in both directions, executed over 28 payload combinations. A new frontend against an old backend, with the key absent and as undefined, null, NaN, the string "0.36" and Infinity, renders Unknown in every case and never leaks undefined, NaN or a fabricated 0. An old frontend against a new backend is unaffected: the pre-PR code reads only devices[].vram_used_gb, and there is no schema validation on /api/system that could reject an unknown key (no zod, ajv, io-ts or similar anywhere in the frontend), so the extra key is inert.

No new dependency, no persisted state, no schema change.

test_rocm_windows_vram_7072.py (15) and test_rocm_multi_gpu_vram_system_wide.py (27) still pass, 51 in that slice altogether. Frontend npm test is 1957 passing, and the new file is matched by the runner's glob, verified rather than assumed.

There is no AMD or Windows CI in this repository, so torch, the performance counter and the platform are mocked, and none of this is hardware validation.

Fixes #7452

danielhanchen and others added 2 commits August 11, 2026 18:38
…ibuted

On a Windows ROCm host with two asymmetric AMD cards the System tab and the
floating monitor show Unknown used VRAM on every device row and on the
aggregate tile. Windows shares no key between the LUID performance counter
instances and torch ordinals, so usage is paired to devices by capacity
ranking and kept only when capacity forces it. On a 45 GiB plus 7.98 GiB pair
nothing at or below 7.98 GiB is forced, which is idle and every small model.

The pairing is unresolvable but the sum is not: over a bijection the total is
the same whichever way round the counters go. Report it as
vram_used_gb_aggregate, and let the tile prefer the per-device sum, fall back
to the aggregate, and otherwise stay Unknown rather than 0.

Per-device rows are unchanged: a value still appears only when capacity
forces it.
chatgpt-codex-connector[bot]

This comment was marked as resolved.

…sible set

The aggregate dropped sub-threshold counters to force a one-per-device count,
the same noise filter the per-device matcher uses. That is safe there, because
per-device only ever emits a value capacity FORCES, but not for a sum, which
emits every counter it kept.

The instance names carry no vendor, LUID or PCI key, so a counter cannot be
told apart from a foreign adapter's: a card hidden by HIP_VISIBLE_DEVICES, an
iGPU, an NVIDIA card in the same box, or a Basic Render Driver placeholder over
the cutoff. Whenever such an adapter is busier than one visible card is idle,
the filter dropped the visible card and kept the foreign one, and the total
gained bytes that are on no visible card. Six such cases, worst of them a
single idle card beside a busy NVIDIA card reading as 6 GiB of AMD usage.

Establish the set by cardinality instead. Get-Counter returns an instance for
every WDDM adapter, so a visible card is always in the list, and a list exactly
as long as the visible set contains those cards and nothing else. An
unexplained instance now reports Unknown, which is narrower on purpose: a host
total that is confidently wrong is worse than no total. Widening it needs the
counters joined on LUID or PCI bus id rather than on capacity rank.
ROCm/legacy-rocm-build#1909 was the wrong source and is dropped from the last two places it
was still quoted: it names no operating system, its repro is rocm-smi on Linux,
and it closed with no fix.

AMD documents the behaviour directly. The hipMemGetInfo reference warns "On
Windows, the free memory only accounts for memory allocated by this process and
may be optimistic", and an AMD engineer gives the mechanism in
ROCm/librocdxg#57: WDDM virtualises video memory, so HIP reports a per-process
budget and a fresh process sees free at or near total whatever else is resident.
Measured there at 24410 MiB of 24560 on a deliberately filled card. Recorded as
intended Windows behaviour rather than a defect, which is worth stating.

Also records why WSL needs no branch. AMD keeps the WSL2 reading consistent with
native Linux, where free tracks physical residency, and sys.platform is "linux"
there, so the predicate is False and the accurate figure passes through uncapped.

Two fixes alongside it. The aggregate now rides only on an index set identical to
the payload's devices: the utilization and visibility probes enumerate
separately, and _torch_get_per_device_info drops a device whose mem_get_info
raises while the aggregate's enumeration keeps it, which would inflate the
percentage and floor free at 0. And the frontend precedence test set its
aggregate to the per-device sum, so it passed whichever source won and pinned
nothing.
@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Hooray!

Reviewed commit: 22ebc59edb

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. Bravo.

Reviewed commit: d6d6bf0b62

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. You're on a roll.

Reviewed commit: c96218a91f

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

@chatgpt-codex-connector

Copy link
Copy Markdown

Codex Review: Didn't find any major issues. More of your lovely PRs please.

Reviewed commit: 4e64a57e5d

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

The aggregate's safety argument rests on Get-Counter emitting exactly one
GPU Adapter Memory instance per WDDM adapter, and the docstring asserted that
as established fact. It is not confirmed against Microsoft's counter
documentation, so it now says so.

The behaviour is unchanged and the assumption fails closed: a second instance
for one adapter makes the counter list longer than the visible set, which
returns None rather than a total. Verified against the helper directly.
@danielhanchen

Copy link
Copy Markdown
Member Author

Correcting an overclaim in my own docstring, pushed as 2b9303d. Comment-only, AST-verified, behaviour unchanged.

The aggregate's safety argument rests on Get-Counter emitting exactly one GPU Adapter Memory instance per WDDM adapter, and the docstring stated that as established fact. I could not confirm it against Microsoft's counter documentation, so it should not have been written as settled.

It now says it is an assumption, and records that it fails closed. Checked against the helper directly:

two cards, one adapter emits a 2nd instance -> None
two cards, clean 1:1                        -> 46.0 GiB

An extra instance makes the counter list longer than the visible set, which returns None rather than a total, so a wrong assumption costs an Unknown and never a wrong number. That is the right direction for this helper, since the whole reason it exists is that a confidently wrong host total is worse than Unknown.

Worth confirming before anyone widens this: the durable fix is joining counters to devices on LUID or PCI bus id rather than on capacity rank.

@danielhanchen
danielhanchen merged commit 26ef15b into main Aug 12, 2026
50 of 60 checks passed
@danielhanchen
danielhanchen deleted the fix/vram_7452 branch August 12, 2026 09:02
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] Stops Reading VRAM Total / Usage on RDNA3

1 participant