Skip to content

docs(benchmarks): correct the runbook — ORT-CUDA does work here, via the conda wheel - #1601

Merged
justinchuby merged 1 commit into
mainfrom
docs/ort-cuda-conda
Aug 20, 2026
Merged

justinchuby merged 1 commit into
mainfrom
docs/ort-cuda-conda

Conversation

@justinchuby

Copy link
Copy Markdown
Owner

The Windows CUDA runbook (merged earlier today in #1487) claimed an end-to-end ORT-CUDA run was "not performed here (no CUDA-enabled ORT is installed on this box)". That was wrong, and the repo owner caught it.

onnxruntime_gpu-1.28.0 — the same version this repo pins — is already installed in the conda environment, right next to the CUDA wheels the native EP uses:

…\anaconda3\Lib\site-packages\onnxruntime\capi\
    onnxruntime.dll                   16.6 MB
    onnxruntime_providers_cuda.dll   252.9 MB

The false conclusion came from searching only target\ and the repo tree — a search that never covered the place such a package would live. "Not installed", inferred from a search that could not have found it, is not a measurement. The correction is recorded inline in the doc rather than quietly edited away, since the reasoning error is the reusable part.

What is added (§4b)

The measured recipe, plus the two failure stages — which matter because the second one looks like a model problem and is actually configuration:

  1. Without ONNX_GENAI_ORT_LIB_DIR: ort-sys loads the CPU-only ORT from target\; provider list is ["AzureExecutionProvider", "CPUExecutionProvider"].
  2. With the lib dir but without ONNX_GENAI_EP_FALLBACK=1: the CUDA EP is present, but session creation fails with "This session contains graph nodes that are assigned to the default CPU EP, but fallback to CPU EP has been explicitly disabled by the user." Some nodes legitimately have no CUDA kernel.

Both set:

ort_available_execution_providers: ["TensorrtExecutionProvider", "CUDAExecutionProvider", "CPUExecutionProvider"]

Why this matters beyond the doc

It unblocks native-vs-ORT A/B on CUDA, which had been treated as unavailable. Measured on qwen05b-fresh, 32 greedy tokens, all three arms emit identical token ids:

arm token[1] stream
native (cuda) 358 baseline
ort (cuda) 358 identical
ort (cpu) 358 identical

That also retires the open "qwen05b ORT arm mis-decodes GQA at token 1" item — there is no mis-decode. A divergence in this A/B is now a real signal rather than a suspected harness artifact.

Docs only; no code changes. Every figure quoted was measured in this session.

…the conda wheel

The Windows CUDA runbook said an end-to-end ORT-CUDA run was "not performed
here (no CUDA-enabled ORT is installed on this box)". That was wrong. The conda
environment already has `onnxruntime_gpu-1.28.0` — the same version this repo
pins — with a 253 MB `onnxruntime_providers_cuda.dll`, sitting next to the CUDA
wheels the native EP already uses.

The false conclusion came from searching only `target\` and the repo tree, i.e.
from a search that never covered the place such a package would actually live.
"Not installed" inferred from an incomplete search is not a measurement, so the
correction is recorded inline rather than quietly edited away.

Adds section 4b with the measured recipe and the two failure stages, which fail
in different-looking ways:

- Without `ONNX_GENAI_ORT_LIB_DIR`, `ort-sys` loads the CPU-only ORT it
  downloaded into `target\`; the provider list has no CUDA EP.
- With the lib dir set but without `ONNX_GENAI_EP_FALLBACK=1`, the CUDA EP is
  present but session creation fails with "graph nodes ... assigned to the
  default CPU EP, but fallback ... has been explicitly disabled". That reads
  like a model problem and is actually configuration.

This unblocks native-vs-ORT A/B on CUDA, which had been treated as unavailable.
Measured on qwen05b-fresh over 32 greedy tokens, the native CUDA arm and the ORT
CUDA arm emit identical token ids, and both match the ORT CPU arm — so a
divergence in that A/B is now a real signal rather than a harness artifact.

This also retires the "qwen05b ORT arm mis-decodes GQA at token 1" item: all
three arms agree, token[1] = 358 in each.

Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
@justinchuby
justinchuby merged commit 0da1fe0 into main Aug 20, 2026
2 checks passed
@justinchuby
justinchuby deleted the docs/ort-cuda-conda branch August 20, 2026 18:21
justinchuby added a commit that referenced this pull request Aug 20, 2026
Picks up #1586, #1601 and #1602. No conflicts; Cargo.lock was the only
file both sides touched and it auto-merged.

Audited the same way as the previous two syncs rather than by reading
the diff: every main-only file verified byte-identical to origin/main by
blob hash, and for Cargo.lock every line main added over the merge-base
confirmed present in the result, with this stack's three
onnx-runtime-memory-api entries still intact.

#1586 is the reason this sync matters beyond staying mergeable. It
corrects the macOS `struct statvfs` FFI layout, which is what made
`platform_capacity::tests::disk_capacity_is_measured_for_the_working_directory`
and `engine::tests::an_explicit_byte_limit_is_honored_without_a_device_query`
fail on every local run of this branch. Those two were the only failures
in the crate suite here, and every "559 passed / 2 failed" figure
reported against this PR was that pair. They are expected to be gone
now; verified in the following run rather than assumed.

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot-Session: c80f8522-983c-47f7-8241-2155a823aabe
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants