Skip to content

Notable changes: April 9–25, 2026 #220

Description

@justinchuby

Summary

Over the past three weeks (April 9–25), mobius merged 50 PRs from 5 contributors (plus 2 bots). The headline additions are Gemma 4 multimodal support (text + vision + audio with MoE), significant improvements to CUDA/GPU execution provider compatibility, and a push to bring L1–L3 test coverage above 95%.


🧠 New Model Support

⚡ CUDA / GPU Inference

�� ORT GenAI Integration

✅ Quality & Testing

🛠️ Developer Experience & Code Quality


Justin's notes:

  • Olive pass support: using mobius to acquire an onnx model as an Olive pass

🔭 What's Next

  • Open source release
  • E2E Olive recipes: Expand e2e Olive recipes for a batch of models.
  • Expanding GPU validation — Continue broadening L4/L5 GPU test coverage across more model families
  • Additional multimodal architectures — Audio-to-audio, TTS, etc.
  • Support generating models from Olive quantized PyTorch weights (gptq etc.).-
  • Quantization support — Improve MXFP4/INT4 weight handling across more architectures / from gguf models
  • ORT GenAI runtime parity — Close remaining gaps in config generation for newer models
  • Performance testing by integrating with ep cert testing infra

Metadata

Metadata

Assignees

No one assigned

    Labels

    documentationImprovements or additions to documentation

    Type

    No type

    Projects

    No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions