Repository navigation
docs(benchmarks): correct the runbook — ORT-CUDA does work here, via the conda wheel - #1601
Merged
Merged
Conversation
…the conda wheel The Windows CUDA runbook said an end-to-end ORT-CUDA run was "not performed here (no CUDA-enabled ORT is installed on this box)". That was wrong. The conda environment already has `onnxruntime_gpu-1.28.0` — the same version this repo pins — with a 253 MB `onnxruntime_providers_cuda.dll`, sitting next to the CUDA wheels the native EP already uses. The false conclusion came from searching only `target\` and the repo tree, i.e. from a search that never covered the place such a package would actually live. "Not installed" inferred from an incomplete search is not a measurement, so the correction is recorded inline rather than quietly edited away. Adds section 4b with the measured recipe and the two failure stages, which fail in different-looking ways: - Without `ONNX_GENAI_ORT_LIB_DIR`, `ort-sys` loads the CPU-only ORT it downloaded into `target\`; the provider list has no CUDA EP. - With the lib dir set but without `ONNX_GENAI_EP_FALLBACK=1`, the CUDA EP is present but session creation fails with "graph nodes ... assigned to the default CPU EP, but fallback ... has been explicitly disabled". That reads like a model problem and is actually configuration. This unblocks native-vs-ORT A/B on CUDA, which had been treated as unavailable. Measured on qwen05b-fresh over 32 greedy tokens, the native CUDA arm and the ORT CUDA arm emit identical token ids, and both match the ORT CPU arm — so a divergence in that A/B is now a real signal rather than a harness artifact. This also retires the "qwen05b ORT arm mis-decodes GQA at token 1" item: all three arms agree, token[1] = 358 in each. Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Copilot-Session: d60eb808-7cc6-4abc-b48d-2a6dd3841624
justinchuby
added a commit
that referenced
this pull request
Aug 20, 2026
Picks up #1586, #1601 and #1602. No conflicts; Cargo.lock was the only file both sides touched and it auto-merged. Audited the same way as the previous two syncs rather than by reading the diff: every main-only file verified byte-identical to origin/main by blob hash, and for Cargo.lock every line main added over the merge-base confirmed present in the result, with this stack's three onnx-runtime-memory-api entries still intact. #1586 is the reason this sync matters beyond staying mergeable. It corrects the macOS `struct statvfs` FFI layout, which is what made `platform_capacity::tests::disk_capacity_is_measured_for_the_working_directory` and `engine::tests::an_explicit_byte_limit_is_honored_without_a_device_query` fail on every local run of this branch. Those two were the only failures in the crate suite here, and every "559 passed / 2 failed" figure reported against this PR was that pair. They are expected to be gone now; verified in the following run rather than assumed. Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com> Copilot-Session: c80f8522-983c-47f7-8241-2155a823aabe
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The Windows CUDA runbook (merged earlier today in #1487) claimed an end-to-end ORT-CUDA run was "not performed here (no CUDA-enabled ORT is installed on this box)". That was wrong, and the repo owner caught it.
onnxruntime_gpu-1.28.0— the same version this repo pins — is already installed in the conda environment, right next to the CUDA wheels the native EP uses:The false conclusion came from searching only
target\and the repo tree — a search that never covered the place such a package would live. "Not installed", inferred from a search that could not have found it, is not a measurement. The correction is recorded inline in the doc rather than quietly edited away, since the reasoning error is the reusable part.What is added (§4b)
The measured recipe, plus the two failure stages — which matter because the second one looks like a model problem and is actually configuration:
ONNX_GENAI_ORT_LIB_DIR:ort-sysloads the CPU-only ORT fromtarget\; provider list is["AzureExecutionProvider", "CPUExecutionProvider"].ONNX_GENAI_EP_FALLBACK=1: the CUDA EP is present, but session creation fails with "This session contains graph nodes that are assigned to the default CPU EP, but fallback to CPU EP has been explicitly disabled by the user." Some nodes legitimately have no CUDA kernel.Both set:
Why this matters beyond the doc
It unblocks native-vs-ORT A/B on CUDA, which had been treated as unavailable. Measured on
qwen05b-fresh, 32 greedy tokens, all three arms emit identical token ids:That also retires the open "qwen05b ORT arm mis-decodes GQA at token 1" item — there is no mis-decode. A divergence in this A/B is now a real signal rather than a suspected harness artifact.
Docs only; no code changes. Every figure quoted was measured in this session.