Bump Microsoft.ML.OnnxRuntime.Gpu and Microsoft.ML.OnnxRuntimeGenAI.Cuda - #243
dependabot[bot] wants to merge 1 commit into
Conversation
Bumps Microsoft.ML.OnnxRuntime.Gpu from 1.29.0 to 1.30.0 Bumps Microsoft.ML.OnnxRuntimeGenAI.Cuda from 0.15.2 to 0.16.0 --- updated-dependencies: - dependency-name: Microsoft.ML.OnnxRuntime.Gpu dependency-version: 1.30.0 dependency-type: direct:production update-type: version-update:semver-minor - dependency-name: Microsoft.ML.OnnxRuntimeGenAI.Cuda dependency-version: 0.16.0 dependency-type: direct:production update-type: version-update:semver-minor ... Signed-off-by: dependabot[bot] <support@github.com>
How to use the Graphite Merge QueueAdd either label to this PR to merge it via the merge queue:
You must have a Graphite account in order to use the merge queue. Sign up using this link. An organization admin has enabled the Graphite Merge Queue in this repository. Please do not merge from GitHub as this will restart CI on PRs being processed by the merge queue. |
|
Important Review skippedBot user detected. To trigger a single review, invoke the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Advanced Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
Looks like these dependencies are no longer updatable, so this is no longer needed. |
Updated Microsoft.ML.OnnxRuntime.Gpu from 1.29.0 to 1.30.0.
Release notes
Sourced from Microsoft.ML.OnnxRuntime.Gpu's releases.
1.30.0
ONNX Runtime 1.30.0 expands generative AI inference, improves CPU and GPU performance, adds Go bindings, and strengthens runtime reliability. These notes cover changes since ONNX Runtime 1.29.1.
Highlights
Announcements & Compatibility
-Donnxruntime_USE_FP4_QMOE=OFF(#32096, #32163).block_size=32. Set-Donnxruntime_USE_FPA_INTB_GEMM_FULL=ONwhen building from source to retain the full kernel set, including BF16, zero-point, bias, larger-block-size, and native Hopper variants (#32324).GemmandMatMulexecution is gated on hardware acceleration. CPU-assigned FP16 nodes without a matching kernel now fall back to FP32 (#32301, #32197).Security & Reliability
Model Loading, Memory, and Input Validation
Split,Scan,GatherND,ScatterND,SpaceToDepth/DepthToSpace,Crop,Conv,Normalizer, and pooling (#29461, #31668, #32034, #32039, #32076, #32157, #32160, #32161, #32345, #32349).BifurcationDetectorinputs, generation subgraph shapes, and QEmbed segment inputs. BeamSearch buffer expansion now uses dynamic shape storage (#31648, #31701, #32009, #32078, #32144).TreeEnsemblenode references and bounded subtree comparison, rejected non-finite CPURoiAligncoordinates, and requiredImageScalerbias to match the channel count (#32031, #32043, #32011, #32002).MatMulFpQ4shape inputs, and checked MLAS blockwise quantization/dequantization index ranges (#31682, #32032, #32007).GPU Bounds and Resource Lifetimes
MatMulNBits,RemovePadding,RotaryEmbedding,SparseAttention, Whisper beam search, NMS, QDQ, andGatherElements(#31643, #31994, #31995, #31996, #31998, #32014, #32029, #32030).CudaAsyncBufferstaging storage alive across CUDA graph replay (#31968, #32121).Dependencies and Tooling
js-yaml,joi,fast-uri, and the Next.js end-to-end fixture (#32397, #32486, #32488, #32505, #32508).New Features
Core APIs & Runtime
KernelContext::GetPreallocatedOutput(#29726, #32089).EngramGateandNGramHashMapping, and expanded kernel coverage for Qwen-3.5 operators (#32268, #32106).Plugin Execution Providers
... (truncated)
1.29.1
This is a patch release on top of v1.29.0, containing GroupQueryAttention capability and KV-cache layout improvements, plugin Execution Provider performance tooling updates, and targeted graph and optimizer fixes.
GroupQueryAttention
causalattribute, with explicit handling for unsupported execution paths (#31704)attention_biaswith a sliding-window KV cache, including explicit position IDs and post-eviction bias indexing (#32302)Runtime and Performance Tools
onnxruntime_perf_testto use plugin Execution Provider device allocators for generated inputs, loaded test data, and pre-allocated outputs, avoiding unnecessary per-run host/device copies (#32244)Bug Fixes and Documentation
MulandPowpatterns (#32016)Contributors
Thanks to our 7 contributors for this release!
@adrastogi, @apsonawane, @edgchen1, @javier-intel, @jnagi-intel, @tianleiwu, @Wayne-Ch
Release highlights were drafted with AI assistance and are subject to release-team review.
Full Changelog: v1.29.0...v1.29.1
Commits viewable in compare view.
Updated Microsoft.ML.OnnxRuntimeGenAI.Cuda from 0.15.2 to 0.16.0.
Release notes
Sourced from Microsoft.ML.OnnxRuntimeGenAI.Cuda's releases.
No release notes found for this version range.
Commits viewable in compare view.
Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting
@dependabot rebase.Dependabot commands and options
You can trigger Dependabot actions by commenting on this PR:
@dependabot rebasewill rebase this PR@dependabot recreatewill recreate this PR, overwriting any edits that have been made to it@dependabot show <dependency name> ignore conditionswill show all of the ignore conditions of the specified dependency@dependabot ignore this major versionwill close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)@dependabot ignore this minor versionwill close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)@dependabot ignore this dependencywill close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)