Skip to content

Gemma4 ONNX + CUDA support: all PRs tracker #257

Description

@justinchuby

Gemma4 CUDA inference — tracking issue

Performance (H200 GPU)

Config tok/s VRAM
F16 INT4 (Q4_K_M) 176.7 9 GB
F16 onnx-standard 161.9 12 GB
F16 default 157.8 12 GB
F16 cuda (GQA) 151.4 14 GB
F32 default 143.9 22 GB
Image CUDA ~240
Audio CPU 15.1
Tri-modal (image+audio+text) 90 ~14 GB

Text generation: exact token match with HF PyTorch ✅

mobius PRs

onnxruntime PRs

onnxruntime-extensions PRs

onnxruntime-extensions Issues

onnxruntime-genai PRs

Olive

Issues

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions