Repository navigation
Update Python APIs for Moonshine v2 models - #3235
Conversation
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces support for Moonshine v2 models within the Python API, providing users with a new example to facilitate their integration. It also enhances the CI pipeline to automatically test this new functionality, ensuring stability and correctness. Additionally, a minor but important correction was made to the Python bindings for model configuration, aligning default argument types with their expected string values. Highlights
Changelog
Activity
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
|
Caution Review failedThe pull request is closed. ℹ️ Recent review infoConfiguration used: defaults Review profile: CHILL Plan: Pro 📒 Files selected for processing (3)
📝 WalkthroughWalkthroughAdds operational steps to the CI test script for Moonshine v2 and FireRedASR models, introduces a new Python example script demonstrating offline Moonshine v2 decoding with audio processing, and updates default parameter values in the Python binding for OfflineMoonshineModelConfig from braces to empty strings. Changes
Estimated code review effort🎯 2 (Simple) | ⏱️ ~12 minutes Poem
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Code Review
This pull request updates the Python APIs for Moonshine v2 models, introducing a new example script for offline decoding and a corresponding test case in the CI script. The changes are generally good, but I have a few suggestions to improve code quality and maintainability. In both the test script and the new Python example, hardcoded model paths are repeated; refactoring these into variables would make future updates easier. Additionally, the Python example uses datetime.now() for performance measurement, which could be replaced with the more accurate time.perf_counter() for timing operations.
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27.tar.bz2 | ||
| tar xvf sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27.tar.bz2 | ||
| rm sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27.tar.bz2 | ||
|
|
||
| ls -lh sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27 | ||
|
|
||
| python3 ./python-api-examples/offline-moonshine-decode-files-v2.py | ||
|
|
||
| rm -rf sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27 |
There was a problem hiding this comment.
To improve maintainability and reduce redundancy, consider using variables for the model archive and directory names. This makes it easier to update the model version in the future. The ls command also appears to be for debugging and could be removed from the script.
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27.tar.bz2 | |
| tar xvf sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27.tar.bz2 | |
| rm sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27.tar.bz2 | |
| ls -lh sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27 | |
| python3 ./python-api-examples/offline-moonshine-decode-files-v2.py | |
| rm -rf sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27 | |
| MODEL_ARCHIVE="sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27.tar.bz2" | |
| MODEL_DIR="sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27" | |
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/${MODEL_ARCHIVE} | |
| tar xvf "${MODEL_ARCHIVE}" | |
| rm "${MODEL_ARCHIVE}" | |
| python3 ./python-api-examples/offline-moonshine-decode-files-v2.py | |
| rm -rf "${MODEL_DIR}" |
| import datetime as dt | ||
| from pathlib import Path |
| encoder = "./sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27/encoder_model.ort" | ||
| decoder = ( | ||
| "./sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27/decoder_model_merged.ort" | ||
| ) | ||
| tokens = "./sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27/tokens.txt" | ||
| test_wav = "./sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27/test_wavs/0.wav" |
There was a problem hiding this comment.
To improve readability and maintainability, you can define the model directory path once and reuse it to construct the full paths for the model files. This avoids repeating the long directory name and makes the code cleaner.
| encoder = "./sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27/encoder_model.ort" | |
| decoder = ( | |
| "./sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27/decoder_model_merged.ort" | |
| ) | |
| tokens = "./sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27/tokens.txt" | |
| test_wav = "./sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27/test_wavs/0.wav" | |
| model_dir = Path("./sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27") | |
| encoder = model_dir / "encoder_model.ort" | |
| decoder = model_dir / "decoder_model_merged.ort" | |
| tokens = model_dir / "tokens.txt" | |
| test_wav = model_dir / "test_wavs/0.wav" |
| start_t = dt.datetime.now() | ||
|
|
||
| stream = recognizer.create_stream() | ||
| stream.accept_waveform(sample_rate, audio) | ||
| recognizer.decode_stream(stream) | ||
|
|
||
| end_t = dt.datetime.now() | ||
| elapsed_seconds = (end_t - start_t).total_seconds() |
There was a problem hiding this comment.
For measuring performance, time.perf_counter() is generally more suitable than datetime.datetime.now(). perf_counter() provides a high-resolution monotonic clock that is not affected by system time changes, making it ideal for timing short-duration intervals.
| start_t = dt.datetime.now() | |
| stream = recognizer.create_stream() | |
| stream.accept_waveform(sample_rate, audio) | |
| recognizer.decode_stream(stream) | |
| end_t = dt.datetime.now() | |
| elapsed_seconds = (end_t - start_t).total_seconds() | |
| start_t = time.perf_counter() | |
| stream = recognizer.create_stream() | |
| stream.accept_waveform(sample_rate, audio) | |
| recognizer.decode_stream(stream) | |
| end_t = time.perf_counter() | |
| elapsed_seconds = end_t - start_t |
Usage
(py312) fangjuns-MacBook-Pro:sherpa-onnx fangjun$ python3 ./python-api-examples/offline-moonshine-decode-files-v2.py /Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/offline-stream.cc:AcceptWaveformImpl:133 Creating a resampler: in_sample_rate: 24000 output_sample_rate: 16000 {"lang": "", "emotion": "", "event": "", "text": " Ask not what your country can do for you. Ask what you can do for your country.", "timestamps": [], "durations": [], "tokens":[" Ask", " not", " what", " your", " country", " can", " do", " for", " you", ".", " Ask", " what", " you", " can", " do", " for", " your", " country", "."], "ys_log_probs": [], "words": []} ./sherpa-onnx-moonshine-tiny-en-quantized-2026-02-27/test_wavs/0.wav Text: Ask not what your country can do for you. Ask what you can do for your country. Audio duration: 3.845 s Elapsed: 0.056 s RTF = 0.056/3.845 = 0.015Summary by CodeRabbit
New Features
Bug Fixes