Skip to content

Export Moonshine v2 models to sherpa-onnx - #3234

Merged
csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:export-moonshine-v2
Feb 27, 2026
Merged

csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:export-moonshine-v2

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Feb 27, 2026 •

Copy link
Copy Markdown
Collaborator

You can download the models from
https://github.com/k2-fsa/sherpa-onnx/releases/tag/asr-models

See also

Usage:

./build/bin/sherpa-onnx-offline \
  --moonshine-encoder=./sherpa-onnx-moonshine-base-zh-quantized-2026-02-27/encoder_model.ort \
  --moonshine-merged-decoder=./sherpa-onnx-moonshine-base-zh-quantized-2026-02-27/decoder_model_merged.ort \
  --tokens=./sherpa-onnx-moonshine-base-zh-quantized-2026-02-27/tokens.txt \
  --debug=0 \
  ./sherpa-onnx-moonshine-base-zh-quantized-2026-02-27/test_wavs/0.wav

Output logs are:

/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/parse-options.cc:Read:373 ./build/bin/sherpa-onnx-offline --moonshine-encoder=./sherpa-onnx-moonshine-base-zh-quantized-2026-02-27/encoder_model.ort --moonshine-merged-decoder=./sherpa-onnx-moonshine-base-zh-quantized-2026-02-27/decoder_model_merged.ort --tokens=./sherpa-onnx-moonshine-base-zh-quantized-2026-02-27/tokens.txt --debug=0 ./sherpa-onnx-moonshine-base-zh-quantized-2026-02-27/test_wavs/0.wav 

OfflineRecognizerConfig(feat_config=FeatureExtractorConfig(sampling_rate=16000, feature_dim=80, low_freq=20, high_freq=-400, dither=0, normalize_samples=True, snip_edges=False), model_config=OfflineModelConfig(transducer=OfflineTransducerModelConfig(encoder_filename="", decoder_filename="", joiner_filename=""), paraformer=OfflineParaformerModelConfig(model=""), nemo_ctc=OfflineNemoEncDecCtcModelConfig(model=""), whisper=OfflineWhisperModelConfig(encoder="", decoder="", language="", task="transcribe", tail_paddings=-1, enable_token_timestamps=False, enable_segment_timestamps=False), fire_red_asr=OfflineFireRedAsrModelConfig(encoder="", decoder=""), tdnn=OfflineTdnnModelConfig(model=""), zipformer_ctc=OfflineZipformerCtcModelConfig(model=""), wenet_ctc=OfflineWenetCtcModelConfig(model=""), sense_voice=OfflineSenseVoiceModelConfig(model="", language="auto", use_itn=False), moonshine=OfflineMoonshineModelConfig(preprocessor="", encoder="./sherpa-onnx-moonshine-base-zh-quantized-2026-02-27/encoder_model.ort", uncached_decoder="", cached_decoder="", merged_decoder="./sherpa-onnx-moonshine-base-zh-quantized-2026-02-27/decoder_model_merged.ort"), dolphin=OfflineDolphinModelConfig(model=""), canary=OfflineCanaryModelConfig(encoder="", decoder="", src_lang="", tgt_lang="", use_pnc=True), omnilingual=OfflineOmnilingualAsrCtcModelConfig(model=""), funasr_nano=OfflineFunASRNanoModelConfig(encoder_adaptor="", llm="", embedding="", tokenizer="", system_prompt="You are a helpful assistant.", user_prompt="语音转写:", max_new_tokens=512, temperature=1e-06, top_p=0.8, seed=42, language="", itn=True, hotwords=""), medasr=OfflineMedAsrCtcModelConfig(model=""), fire_red_asr_ctc=OfflineFireRedAsrCtcModelConfig(model=""), telespeech_ctc="", tokens="./sherpa-onnx-moonshine-base-zh-quantized-2026-02-27/tokens.txt", num_threads=2, debug=False, provider="cpu", model_type="", modeling_unit="cjkchar", bpe_vocab=""), lm_config=OfflineLMConfig(model="", scale=0.5, lodr_scale=0.01, lodr_fst="", lodr_backoff_id=-1), ctc_fst_decoder_config=OfflineCtcFstDecoderConfig(graph="", max_active=3000), decoding_method="greedy_search", max_active_paths=4, hotwords_file="", hotwords_score=1.5, blank_penalty=0, rule_fsts="", rule_fars="", hr=HomophoneReplacerConfig(lexicon="", rule_fsts=""))
Creating recognizer ...
recognizer created in 0.351 s
Started
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/offline-stream.cc:AcceptWaveformImpl:133 Creating a resampler:
   in_sample_rate: 24000
   output_sample_rate: 16000

Done!

./sherpa-onnx-moonshine-base-zh-quantized-2026-02-27/test_wavs/0.wav
{"lang": "", "emotion": "", "event": "", "text": " 不要问你的国家能为你做什么,而要问你能为你的国家做什么。", "timestamps": [], "durations": [], "tokens":[" ", "不", "要", "问", "你", "的", "国", "家", "能", "为", "你", "<0xE5>", "<0x81>", "<0x9A>", "<0xE4>", "<0xBB>", "<0x80>", "么", ",", "而", "要", "问", "你", "能", "为", "你", "的", "国", "家", "<0xE5>", "<0x81>", "<0x9A>", "<0xE4>", "<0xBB>", "<0x80>", "么", "。"], "ys_log_probs": [], "words": []}
----
num threads: 2
decoding method: greedy_search
Elapsed seconds: 0.079 s
Real time factor (RTF): 0.079 / 4.759 = 0.017

cc @nshmyrev @nabil6391 @zhitao-zeng @jiangzhuo


CI logs

models/download.moonshine.ai/model/base-en/quantized/base-en:
total 292776
-rw-r--r--  1 runner  staff   104M Feb 27 09:27 decoder_model_merged.ort
-rw-r--r--  1 runner  staff    30M Feb 27 09:26 encoder_model.ort
-rw-r--r--  1 runner  staff   244K Feb 27 09:27 tokenizer.bin
models/download.moonshine.ai/model/base-es/quantized/base-es:
total 126608
-rw-r--r--  1 runner  staff    42M Feb 27 09:26 decoder_model_merged.ort
-rw-r--r--  1 runner  staff    20M Feb 27 09:26 encoder_model.ort
-rw-r--r--  1 runner  staff   236K Feb 27 09:26 tokenizer.bin
models/download.moonshine.ai/model/base-ja/quantized/base-ja:
total 275408
-rw-r--r--  1 runner  staff   104M Feb 27 09:27 decoder_model_merged.ort
-rw-r--r--  1 runner  staff    30M Feb 27 09:27 encoder_model.ort
-rw-r--r--  1 runner  staff   244K Feb 27 09:27 tokenizer.bin
models/download.moonshine.ai/model/base-uk/quantized/base-uk:
total 275408
-rw-r--r--  1 runner  staff   104M Feb 27 09:27 decoder_model_merged.ort
-rw-r--r--  1 runner  staff    30M Feb 27 09:27 encoder_model.ort
-rw-r--r--  1 runner  staff   244K Feb 27 09:27 tokenizer.bin
models/download.moonshine.ai/model/base-vi/quantized/base-vi:
total 275408
-rw-r--r--  1 runner  staff   104M Feb 27 09:27 decoder_model_merged.ort
-rw-r--r--  1 runner  staff    30M Feb 27 09:27 encoder_model.ort
-rw-r--r--  1 runner  staff   244K Feb 27 09:27 tokenizer.bin
models/download.moonshine.ai/model/base-zh/quantized/base-zh:
total 291336
-rw-r--r--  1 runner  staff   104M Feb 27 09:26 decoder_model_merged.ort
-rw-r--r--  1 runner  staff    30M Feb 27 09:26 encoder_model.ort
-rw-r--r--  1 runner  staff   244K Feb 27 09:26 tokenizer.bin
models/download.moonshine.ai/model/tiny-en/quantized/tiny-en:
total 85840
-rw-r--r--  1 runner  staff    29M Feb 27 09:26 decoder_model_merged.ort
-rw-r--r--  1 runner  staff    13M Feb 27 09:26 encoder_model.ort
-rw-r--r--  1 runner  staff   244K Feb 27 09:26 tokenizer.bin
models/download.moonshine.ai/model/tiny-ja/quantized/tiny-ja:
total 140280
-rw-r--r--  1 runner  staff    56M Feb 27 09:27 decoder_model_merged.ort
-rw-r--r--  1 runner  staff    13M Feb 27 09:27 encoder_model.ort
-rw-r--r--  1 runner  staff   244K Feb 27 09:27 tokenizer.bin
models/download.moonshine.ai/model/tiny-ko/quantized/tiny-ko:
total 141024
-rw-r--r--  1 runner  staff    56M Feb 27 09:27 decoder_model_merged.ort
-rw-r--r--  1 runner  staff    13M Feb 27 09:27 encoder_model.ort
-rw-r--r--  1 runner  staff   244K Feb 27 09:27 tokenizer.bin

Summary by CodeRabbit

  • New Features

    • Added Moonshine v2 model support with multiple quantized variants
    • Expanded model distribution to additional platforms
  • Chores

    • Updated model export and publishing workflows

@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Feb 27, 2026
@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the integration of Moonshine v2 Automatic Speech Recognition (ASR) models into the sherpa-onnx framework. It introduces a robust script for exporting and packaging these models, alongside updates to the Python API that simplify their configuration and loading. These changes aim to make it easier for users to utilize Moonshine v2 models within sherpa-onnx for various languages.

Highlights

  • Moonshine v2 Model Export Script: A new shell script (scripts/moonshine/v2/run.sh) was added to automate the process of downloading, packaging, and preparing Moonshine v2 ASR models for various languages and architectures (e.g., base-zh, tiny-en) for use with sherpa-onnx. This script handles model downloads, token generation, and creating tarball archives with test audio.
  • Optional Parameters for Moonshine Model Configuration: The Python binding for OfflineMoonshineModelConfig was updated to make all its constructor parameters optional, providing more flexibility when initializing Moonshine model configurations.
  • Simplified Moonshine v2 Model Loading: A new class method, from_moonshine_v2, was added to OfflineRecognizer in Python. This method streamlines the creation of an OfflineRecognizer instance specifically configured for Moonshine v2 models, requiring only the encoder, merged decoder, and tokens file paths, along with other optional parameters.
Changelog
  • scripts/moonshine/v2/run.sh
    • Added a new script to download, process, and package Moonshine v2 ASR models (encoder, merged decoder, tokens) into sherpa-onnx compatible tarballs.
    • Included logic to download test audio files for each language-specific model.
  • sherpa-onnx/python/csrc/offline-moonshine-model-config.cc
    • Modified the OfflineMoonshineModelConfig constructor in its Python binding to accept all parameters as optional with default empty string values.
  • sherpa-onnx/python/sherpa_onnx/offline_recognizer.py
    • Introduced from_moonshine_v2 as a class method to OfflineRecognizer for convenient initialization of Moonshine v2 models.
    • The new method constructs OfflineModelConfig and OfflineRecognizerConfig using provided Moonshine v2 model paths and other recognition parameters.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/export-moonshine-to-onnx.yaml
Activity
  • The pull request was created by csukuangfj.
  • CI logs indicate successful generation and listing of various Moonshine v2 models (base-vi, base-zh, tiny-en, tiny-ja, tiny-ko) with their respective encoder, decoder, and tokenizer files.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Feb 27, 2026 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info

Configuration used: defaults

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 6762232 and f2ad12d.

📒 Files selected for processing (4)
  • .github/workflows/export-moonshine-to-onnx.yaml
  • scripts/moonshine/v2/run.sh
  • sherpa-onnx/python/csrc/offline-moonshine-model-config.cc
  • sherpa-onnx/python/sherpa_onnx/offline_recognizer.py

📝 Walkthrough

Walkthrough

This PR adds Moonshine v2 model support to sherpa-onnx. Changes include a new GitHub Actions workflow for exporting Moonshine v2 models, a model preparation script that downloads language models and generates quantized variants, updates to make the Python binding constructor arguments optional with defaults, and a new from_moonshine_v2 classmethod for the OfflineRecognizer API.

Changes

Cohort / File(s) Summary
CI/CD Workflow
.github/workflows/export-moonshine-to-onnx.yaml
Adds push trigger for export-moonshine-v2-2 branch, introduces matrix versioning (v2), expands Python dependencies (adds moonshine-voice), executes v2 workflow script, and refactors publishing steps to use loops for multi-model HuggingFace and ModelsScope distribution.
Model Preparation Script
scripts/moonshine/v2/run.sh
New bash script that automates Moonshine v2 model preparation: downloads language models (zh, ar, es, en, ja, ko, vi, uk), generates quantized variants, collects artifacts, downloads corresponding WAV files for testing, and creates tar.bz2 bundles per model.
Python Bindings
sherpa-onnx/python/csrc/offline-moonshine-model-config.cc
Makes all five constructor string arguments (preprocessor, encoder, uncached_decoder, cached_decoder, merged_decoder) optional by adding default empty-string values.
Python API
sherpa-onnx/python/sherpa_onnx/offline_recognizer.py
Adds new from_moonshine_v2 classmethod to OfflineRecognizer that configures and instantiates recognizers for Moonshine v2 models with encoder, merged_decoder, and tokens parameters.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related issues

Suggested labels

size:M

Poem

🐰 A whisper of new models, v2 takes flight,
With scripts that bundle dreams in archive format tight,
Optional defaults smooth the binding's way,
Moonshine's second voice comes out to play! ✨

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@csukuangfj
csukuangfj merged commit fa61676 into k2-fsa:master Feb 27, 2026
1 check was pending
@csukuangfj
csukuangfj deleted the export-moonshine-v2 branch February 27, 2026 08:49

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds support for Moonshine v2 models. It includes a script to download and package models, updates Python bindings, and adds a from_moonshine_v2 factory method. The changes are well-structured. I've provided a few suggestions on the packaging script to remove unnecessary sleep commands and to address a potentially confusing hardcoded date in artifact names.

mv -v models/download.moonshine.ai/model/$name/quantized/$name/* .
python3 ./generate_tokens.py
rm tokenizer.bin
d=sherpa-onnx-moonshine-$name-quantized-2026-02-27

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The hardcoded year 2026 in the directory name is unconventional for versioning and can be misleading as it points to a future date. If this is meant to be the release date, it appears to be a typo for 2024. Using a standard versioning scheme or the correct date would improve clarity and avoid confusion.

python3 -m moonshine_voice.download --language vi --model-arch 1
python3 -m moonshine_voice.download --language uk --model-arch 1

sleep 2

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This sleep 2 command appears to be unnecessary. The preceding python3 download commands are synchronous, so the script will wait for them to finish before proceeding. This sleep introduces an artificial delay and can be safely removed.


ls -lh models/download.moonshine.ai/model/*/*/*

sleep 2

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Similar to the previous sleep, this one also seems unnecessary as the ls command is synchronous. It can be safely removed to avoid an artificial delay.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant