Skip to content

Export omnilingual-asr to sherpa-onnx - #2770

Merged
csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:export-omnilingual-asr
Nov 12, 2025
Merged

csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:export-omnilingual-asr

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Nov 12, 2025 •

Copy link
Copy Markdown
Collaborator

See also

https://github.com/facebookresearch/omnilingual-asr

Screenshot 2025-11-12 at 20 19 08

Download

You can find the exported models at
Screenshot 2025-11-12 at 20 21 26

RTF test in GitHub actions (ubuntu-latest, 1 thread)

https://github.com/csukuangfj/sherpa-onnx/actions/runs/19296999620/job/55181385857#step:7:140


---test ./model.int8.onnx with ./fr.wav----iter---5
---> text is---- ne vous demandez pas ce que votre pay peut faire pour vous demandez-vous plutôt ce que vous pouvez faire pour lui
RTF: 0.34275678191740244

---test ./model.int8.onnx with ./en.wav----iter---6
---> text is---- ask not what your country can do for you ask what you can do for your country
RTF: 0.3368970705355211

---test ./model.int8.onnx with ./de.wav----iter---6
---> text is---- alles hat ein ende nur die wurst hat zwei
RTF: 0.333141361043004

---test ./model.int8.onnx with ./es.wav----iter---6
---> text is---- no preguntes que puede hacer tu país por ti pregunta que puedes hacer te un puer tu país
RTF: 0.3414773699444932

---test ./model.int8.onnx with ./fr.wav----iter---6
---> text is---- ne vous demandez pas ce que votre pay peut faire pour vous demandez-vous plutôt ce que vous pouvez faire pour lui
RTF: 0.34011135706346063

---test ./model.int8.onnx with ./en.wav----iter---7
---> text is---- ask not what your country can do for you ask what you can do for your country
RTF: 0.3376139151616368

---test ./model.int8.onnx with ./de.wav----iter---7
---> text is---- alles hat ein ende nur die wurst hat zwei
RTF: 0.33370075062068255

---test ./model.int8.onnx with ./es.wav----iter---7
---> text is---- no preguntes que puede hacer tu país por ti pregunta que puedes hacer te un puer tu país
RTF: 0.34266051487306154

---test ./model.int8.onnx with ./fr.wav----iter---7
---> text is---- ne vous demandez pas ce que votre pay peut faire pour vous demandez-vous plutôt ce que vous pouvez faire pour lui
RTF: 0.34023562524344664

---test ./model.int8.onnx with ./en.wav----iter---8
---> text is---- ask not what your country can do for you ask what you can do for your country
RTF: 0.33787408989143364

---test ./model.int8.onnx with ./de.wav----iter---8
---> text is---- alles hat ein ende nur die wurst hat zwei
RTF: 0.33406119991484967

---test ./model.int8.onnx with ./es.wav----iter---8
---> text is---- no preguntes que puede hacer tu país por ti pregunta que puedes hacer te un puer tu país
RTF: 0.3443957493617938

---test ./model.int8.onnx with ./fr.wav----iter---8
---> text is---- ne vous demandez pas ce que votre pay peut faire pour vous demandez-vous plutôt ce que vous pouvez faire pour lui
RTF: 0.34308558880444656

---test ./model.int8.onnx with ./en.wav----iter---9
---> text is---- ask not what your country can do for you ask what you can do for your country
RTF: 0.33812713399827377

---test ./model.int8.onnx with ./de.wav----iter---9
---> text is---- alles hat ein ende nur die wurst hat zwei
RTF: 0.3374609472598578

---test ./model.int8.onnx with ./es.wav----iter---9
---> text is---- no preguntes que puede hacer tu país por ti pregunta que puedes hacer te un puer tu país
RTF: 0.34255228921292896

---test ./model.int8.onnx with ./fr.wav----iter---9
---> text is---- ne vous demandez pas ce que votre pay peut faire pour vous demandez-vous plutôt ce que vous pouvez faire pour lui
RTF: 0.341598544903942

Summary by CodeRabbit

  • New Features

    • Automated end-to-end pipeline for exporting, packaging, and publishing omnilingual ASR models in ONNX format with quantization support.
  • Documentation

    • Added workflow documentation describing the omnilingual ASR export process and parameter specifications.
  • Tests

    • Added automated testing infrastructure for ONNX model validation across multiple languages.

@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Nov 12, 2025
@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello @csukuangfj, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request integrates the omnilingual-asr model into the sherpa-onnx ecosystem by providing comprehensive scripts for model conversion and validation. The primary goal is to enable the use of this powerful, multilingual Automatic Speech Recognition (ASR) model in a highly optimized ONNX format, including an INT8 quantized version, thereby facilitating faster and more efficient deployment. The included testing framework ensures the integrity and performance of the exported models across different languages.

Highlights

  • Model Export: Introduces scripts to export the omnilingual-asr model from Facebook Research to ONNX format, creating both a standard model.onnx and a quantized model.int8.onnx version for efficient inference.
  • Vocabulary Generation: A tokens.txt file is generated, containing the vocabulary used by the ASR model, which is essential for decoding the model's output.
  • Model Testing and Validation: Includes a test.py script to load and evaluate the exported ONNX models using onnxruntime, performing inference on sample audio files in multiple languages (English, German, Spanish, French) and calculating the Real-Time Factor (RTF).
  • Documentation: Adds a README.md file within the scripts/omnilingual-asr directory, providing an introduction to the export process and referencing a GitHub Actions workflow for usage.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/export-omnilingual-asr-to-onnx.yaml
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Nov 12, 2025 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

Walkthrough

Introduces complete end-to-end automation for exporting omnilingual ASR models to ONNX format, consisting of a GitHub Actions workflow, Python export script with metadata and quantization, test validation script, and documentation describing the export process.

Changes

Cohort / File(s) Summary
Workflow Configuration
.github/workflows/export-omnilingual-asr-to-onnx.yaml
New GitHub Actions workflow triggered on push to export-omnilingual-asr branch and manual dispatch. Configures Python 3.10 environment with dependencies, executes export script, runs tests, packages artifacts as tar.bz2 files, publishes to HuggingFace Hub via Git LFS, and uploads release tarballs to GitHub Releases.
Export and Quantization
scripts/omnilingual-asr/export-onnx.py
New export script that wraps a pretrained omnilingual ASR model with ModelWrapper, exports to ONNX format with dynamic batch and time axes, embeds metadata via add_meta_data(), and post-quantizes to int8 using MatMul quantization. Generates tokens.txt vocabulary mapping, model.onnx, and model.int8.onnx.
Test Utilities
scripts/omnilingual-asr/test.py
New test script introducing OnnxModel class for CPU-based ONNX Runtime inference, with helper functions to load tokens, normalize audio (resampling to 16kHz, mono extraction, mean-variance normalization), and validate exported models across multiple audio files and iterations, reporting text output and real-time factor.
Documentation
scripts/omnilingual-asr/README.md
New documentation describing the omnilingual ASR export workflow with parameter formulas relating num_samples, num_frames, and frame timing (20ms per frame).

Sequence Diagram(s)

sequenceDiagram
    actor User
    participant Workflow as GitHub Actions<br/>Workflow
    participant Export as export-onnx.py
    participant Model as ONNX Model
    participant Quantize as Quantizer
    participant Test as test.py
    participant HF as HuggingFace Hub
    participant Release as GitHub Releases

    User->>Workflow: Push to export-omnilingual-asr
    Workflow->>Export: Run export script
    Export->>Model: Wrap & export inference model
    Model->>Export: ONNX file (model.onnx)
    Export->>Quantize: Apply int8 quantization
    Quantize->>Export: Quantized model (model.int8.onnx)
    
    Note over Export: Generate tokens.txt<br/>Add metadata
    
    Workflow->>Test: Run validation tests
    Test->>Model: Load & run inference
    Model-->>Test: Logits output
    Test->>Test: Decode & measure RTF
    
    Workflow->>Workflow: Package artifacts
    Workflow->>HF: Push to HuggingFace<br/>(with Git LFS)
    Workflow->>Release: Upload tarballs<br/>to GitHub Releases
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Areas requiring extra attention:

  • ONNX export logic (export-onnx.py): Verify ModelWrapper correctly adapts forward pass, metadata embedding, and dynamic axes configuration for variable batch/sequence lengths
  • Quantization pipeline (export-onnx.py): Ensure int8 quantization strategy (MatMul operator targeting) is appropriate for the model architecture and does not degrade accuracy
  • ONNX Runtime inference (test.py): Validate audio preprocessing (resampling, normalization) matches training expectations; confirm token decoding and real-time factor calculations are correct
  • CI/CD workflow orchestration (.yaml): Review artifact collection, Git LFS tracking configuration, HuggingFace Hub authentication/cloning, and release publishing steps for correctness and security

Possibly related PRs

Poem

🐰 A rabbit hops through ONNX gates,
Wrapping models, quantizing states,
From torch to bytes, from speech to text,
Ten thousand tongues—what's coming next?
With metadata sewn and tests all green,
The finest omnilingual ASR we've seen! ✨

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

📜 Recent review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between f8d6fe5 and 834ad93.

📒 Files selected for processing (4)
  • .github/workflows/export-omnilingual-asr-to-onnx.yaml (1 hunks)
  • scripts/omnilingual-asr/README.md (1 hunks)
  • scripts/omnilingual-asr/export-onnx.py (1 hunks)
  • scripts/omnilingual-asr/test.py (1 hunks)

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@csukuangfj
csukuangfj merged commit db77fe6 into k2-fsa:master Nov 12, 2025
1 of 2 checks passed

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces scripts for exporting the omnilingual-asr model to ONNX and for testing the exported models. The implementation is mostly solid, but I've identified a few areas for improvement. A key issue is in export-onnx.py where batch processing is handled incorrectly, which would cause errors for batch sizes greater than one. In test.py, a hardcoded CTC blank ID could lead to incorrect decoding results. I've also made some suggestions to improve code clarity and maintainability in the export script and the documentation. Please see my detailed comments for suggestions on how to address these points.

Args:
x: (N, num_samples), float32
"""
batch_layout = BatchLayout(shape=x.shape, seq_lens=[x.shape[1]])

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The seq_lens argument for BatchLayout is incorrectly constructed for batch sizes greater than 1. It is currently [x.shape[1]], which means it's a list with a single element. This will fail if the batch size x.shape[0] is greater than 1. To support batching correctly, it should be a list of sequence lengths for each item in the batch.

        batch_layout = BatchLayout(shape=x.shape, seq_lens=[x.shape[1]] * x.shape[0])

ids = logits[0].argmax(axis=-1)
ans = []
prev = -1
blank = 0

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The CTC blank token ID is hardcoded as 0. This is a 'magic number' and might be incorrect, leading to decoding errors. It's better to define it as a named constant. Ideally, the blank ID should be saved as part of the model metadata during export and read from there in the test script to make it more robust.

Comment on lines +11 to +20
```
num_frames = round(num_samples / 318 - 1.5)
num_samples = round(318 * num_frames + 477)

or
num_frames = round(num_samples / 320)

```

20ms per frame

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The formulas provided for num_frames and num_samples are a bit confusing. The formula num_frames = round(num_samples / 320) is consistent with a 20ms frame duration at a 16kHz sampling rate (since 16000 * 0.020 = 320). However, the other set of formulas using 318 and 477 seems to contradict this. Could you please clarify the relationship between these different formulas and explain when each should be used? Providing context on how these numbers are derived from the model architecture would be very helpful for users.

Comment on lines +23 to +24
while len(model.metadata_props):
model.metadata_props.pop()

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This while loop to clear the metadata properties is inefficient. The RepeatedCompositeContainer returned by model.metadata_props has a clear() method that is more efficient and idiomatic for this purpose.

    model.metadata_props.clear()

Comment on lines +55 to +59
vocab_size = pipeline.tokenizer._model.vocabulary_size

with open("tokens.txt", "w") as f:
for i in range(pipeline.tokenizer._model.vocabulary_size):
f.write(f"{pipeline.tokenizer._model.index_to_token(i)} {i}\n")

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Accessing the protected member _model of pipeline.tokenizer is fragile and can lead to issues if the omnilingual-asr library is updated. If there is a public API to get the vocabulary size and map indices to tokens, it would be much safer to use that. If not, it would be good to add a comment here acknowledging the risk.

Comment on lines +54 to +55
if len(fields) == 1:
id2token[int(fields[0])] = " "

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The special handling for lines where len(fields) == 1 is not immediately clear. It would be helpful to add a comment explaining this case (e.g., if it's for handling the space token) to improve code readability and maintainability.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant