Skip to content

Add C++ runtime and Python API for Google MedASR models - #2935

Merged
csukuangfj merged 3 commits into
k2-fsa:masterfrom
csukuangfj:cpp-medasr
Dec 25, 2025
Merged

csukuangfj merged 3 commits into
k2-fsa:masterfrom
csukuangfj:cpp-medasr

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Dec 25, 2025 •

Copy link
Copy Markdown
Collaborator

Usage

Try it with our Huggingface space

https://huggingface.co/spaces/k2-fsa/automatic-speech-recognition

Screenshot 2025-12-25 at 17 21 09

Build sherpa-onnx

Please refer to our doc: https://k2-fsa.github.io/sherpa/onnx/install/index.html

Download a model

Please download them from
https://github.com/k2-fsa/sherpa-onnx/releases/tag/asr-models

wget https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-medasr-ctc-en-int8-2025-12-25.tar.bz2
tar xvf sherpa-onnx-medasr-ctc-en-int8-2025-12-25.tar.bz2

Run it

./build/bin/sherpa-onnx-offline \
  --tokens=./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/tokens.txt \
  --medasr=./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/model.int8.onnx \
  --debug=0 \
  ./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/0.wav \
  ./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/1.wav \
  ./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/2.wav

The output is

/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/parse-options.cc:Read:373 ./build/bin/sherpa-onnx-offline --tokens=./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/tokens.txt --medasr=./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/model.int8.onnx --debug=0 ./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/0.wav ./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/1.wav ./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/2.wav 

OfflineRecognizerConfig(feat_config=FeatureExtractorConfig(sampling_rate=16000, feature_dim=80, low_freq=20, high_freq=-400, dither=0, normalize_samples=True, snip_edges=False), model_config=OfflineModelConfig(transducer=OfflineTransducerModelConfig(encoder_filename="", decoder_filename="", joiner_filename=""), paraformer=OfflineParaformerModelConfig(model=""), nemo_ctc=OfflineNemoEncDecCtcModelConfig(model=""), whisper=OfflineWhisperModelConfig(encoder="", decoder="", language="", task="transcribe", tail_paddings=-1), fire_red_asr=OfflineFireRedAsrModelConfig(encoder="", decoder=""), tdnn=OfflineTdnnModelConfig(model=""), zipformer_ctc=OfflineZipformerCtcModelConfig(model=""), wenet_ctc=OfflineWenetCtcModelConfig(model=""), sense_voice=OfflineSenseVoiceModelConfig(model="", language="auto", use_itn=False), moonshine=OfflineMoonshineModelConfig(preprocessor="", encoder="", uncached_decoder="", cached_decoder=""), dolphin=OfflineDolphinModelConfig(model=""), canary=OfflineCanaryModelConfig(encoder="", decoder="", src_lang="", tgt_lang="", use_pnc=True), omnilingual=OfflineOmnilingualAsrCtcModelConfig(model=""), medasr=OfflineMedAsrCtcModelConfig(model="./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/model.int8.onnx"), telespeech_ctc="", tokens="./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/tokens.txt", num_threads=2, debug=False, provider="cpu", model_type="", modeling_unit="cjkchar", bpe_vocab=""), lm_config=OfflineLMConfig(model="", scale=0.5, lodr_scale=0.01, lodr_fst="", lodr_backoff_id=-1), ctc_fst_decoder_config=OfflineCtcFstDecoderConfig(graph="", max_active=3000), decoding_method="greedy_search", max_active_paths=4, hotwords_file="", hotwords_score=1.5, blank_penalty=0, rule_fsts="", rule_fars="", hr=HomophoneReplacerConfig(lexicon="", rule_fsts=""))
Creating recognizer ...
recognizer created in 0.377 s
Started
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/offline-stream.cc:AcceptWaveformImpl:133 Creating a resampler:
   in_sample_rate: 24000
   output_sample_rate: 16000

/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/offline-stream.cc:AcceptWaveformImpl:133 Creating a resampler:
   in_sample_rate: 24000
   output_sample_rate: 16000

Done!

./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/0.wav
{"lang": "", "emotion": "", "event": "", "text": "[EXAM TYPE] CT chest PE protocol {period} [INDICATION] 54-year-old female, shortness of breath, evaluate for PE {period} [TECHNIQUE] Standard protocol {period} [FINDINGS] {colon} Pulmonary vasculature {colon} The main PA is patent {period} There are filling defects in the segmental branches of the right lower lobe {comma} compatible with acute PE {period} No saddle embolus {period} Lungs {colon} No pneumothorax {period} Small bilateral effusions {comma} right greater than left {period} {new paragraph} [IMPRESSION] {colon} Acute segmental PE, right lower lobe {period}", "timestamps": [2.20, 2.28, 2.36, 2.44, 2.52, 2.68, 2.72, 2.80, 2.84, 2.88, 2.96, 3.12, 3.36, 3.44, 3.52, 3.76, 3.92, 4.20, 4.28, 4.36, 4.40, 4.48, 4.60, 4.68, 4.76, 5.08, 5.16, 5.24, 5.32, 5.36, 5.40, 5.48, 5.52, 5.56, 5.60, 5.68, 5.72, 5.80, 5.84, 6.04, 6.12, 6.20, 6.24, 6.32, 6.40, 6.48, 6.56, 6.64, 6.72, 6.80, 6.84, 6.96, 7.08, 7.12, 7.20, 7.28, 7.36, 7.44, 7.52, 7.60, 7.68, 7.76, 7.92, 8.00, 8.08, 8.16, 8.24, 8.36, 8.52, 8.68, 8.84, 8.92, 9.00, 9.44, 9.52, 9.56, 9.64, 9.72, 9.76, 9.80, 9.88, 9.92, 9.96, 10.04, 10.08, 10.12, 10.16, 10.24, 10.28, 10.44, 10.48, 10.56, 10.64, 10.68, 10.84, 10.92, 11.00, 11.48, 11.56, 11.64, 11.72, 11.80, 11.84, 11.92, 11.96, 12.04, 12.12, 12.16, 12.24, 12.32, 13.00, 13.08, 13.12, 13.20, 13.28, 13.36, 13.40, 13.48, 13.52, 13.60, 13.64, 13.72, 13.80, 14.00, 14.08, 14.20, 14.72, 14.84, 14.92, 15.08, 15.24, 15.56, 16.08, 16.16, 16.28, 16.56, 16.64, 16.72, 17.52, 17.64, 17.84, 17.92, 17.96, 18.08, 18.20, 18.28, 18.40, 18.56, 18.68, 18.96, 19.04, 19.12, 19.28, 19.40, 19.44, 19.52, 19.56, 19.64, 19.72, 19.88, 20.12, 20.36, 20.44, 20.52, 20.72, 20.80, 20.88, 21.04, 21.12, 21.24, 22.00, 22.20, 22.32, 22.48, 23.24, 23.36, 23.44, 23.64, 23.76, 23.96, 24.04, 24.12, 24.76, 24.96, 25.04, 25.12, 25.20, 25.56, 25.64, 25.76, 25.80, 26.00, 26.08, 26.16, 26.84, 26.92, 27.00, 27.08, 27.24, 27.32, 27.44, 28.00, 28.16, 28.24, 28.32, 28.36, 28.44, 28.52, 28.60, 28.68, 28.80, 29.00, 29.08, 29.20, 29.72, 29.80, 29.88, 29.92, 30.08, 30.16, 30.24, 30.32, 30.40, 30.56, 30.64, 30.76, 30.84, 30.96, 31.20, 31.28, 31.40, 32.36, 32.60, 32.76, 32.88, 33.08, 33.32, 33.40, 33.52, 34.84, 34.92, 35.04, 35.20, 35.28, 36.12, 36.20, 36.28, 36.36, 36.40, 36.44, 36.52, 36.60, 36.64, 36.72, 36.80, 36.84, 37.08, 37.16, 37.24, 38.48, 38.60, 38.68, 38.84, 38.92, 39.00, 39.16, 39.32, 39.44, 39.56, 39.68, 39.88, 39.96, 40.04, 40.16, 40.24, 40.32, 40.56, 40.64, 40.72, 43.64], "durations": [], "tokens":[" [", "E", "X", "A", "M", " T", "Y", "P", "E", "]", " C", "T", " ", "ch", "est", " P", "E", " pro", "t", "o", "c", "ol", " {", "period", "}", " [", "I", "N", "D", "I", "C", "A", "T", "I", "O", "N", "]", " ", "5", "4", "-", "y", "e", "ar", "-", "old", " f", "e", "m", "al", "e", ",", " sh", "or", "t", "ness", " of", " b", "re", "a", "th", ",", " e", "v", "al", "u", "ate", " for", " P", "E", " {", "period", "}", " [", "T", "E", "C", "H", "N", "I", "Q", "U", "E", "]", " S", "t", "an", "d", "ard", " pro", "t", "o", "c", "ol", " {", "period", "}", " [", "F", "I", "N", "D", "I", "N", "G", "S", "]", " {", "colon", "}", " P", "ul", "m", "on", "ar", "y", " ", "v", "a", "s", "cu", "la", "ture", " {", "colon", "}", " The", " ma", "in", " P", "A", " is", " pa", "t", "ent", " {", "period", "}", " There", " are", " f", "ill", "ing", " de", "f", "ect", "s", " in", " the", " se", "g", "ment", "al", " b", "r", "an", "ch", "es", " of", " the", " right", " ", "low", "er", " lo", "b", "e", " {", "comma", "}", " comp", "at", "ible", " with", " a", "cu", "te", " P", "E", " {", "period", "}", " No", " sa", "d", "d", "le", " e", "mb", "ol", "us", " {", "period", "}", " L", "u", "ng", "s", " {", "colon", "}", " No", " p", "ne", "u", "m", "o", "th", "or", "a", "x", " {", "period", "}", " S", "m", "al", "l", " b", "il", "a", "ter", "al", " e", "ff", "us", "ion", "s", " {", "comma", "}", " right", " great", "er", " than", " left", " {", "period", "}", " {", "new", " ", "paragraph", "}", " [", "I", "M", "P", "R", "E", "S", "S", "I", "O", "N", "]", " {", "colon", "}", " A", "cu", "te", " se", "g", "ment", "al", " P", "E", ",", " right", " ", "low", "er", " lo", "b", "e", " {", "period", "}"], "ys_log_probs": [], "words": []}
----
./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/1.wav
{"lang": "", "emotion": "", "event": "", "text": "Biopsy is a medical procedure in which a sample of tissue is removed from the body for examination. Osteoporosis is a condition in which bones become weak and brittle, increasing the risk of fractures.", "timestamps": [0.08, 0.16, 0.24, 0.32, 0.36, 0.44, 0.64, 0.76, 0.88, 0.96, 1.04, 1.12, 1.28, 1.40, 1.48, 1.60, 1.68, 1.88, 2.04, 2.16, 2.36, 2.48, 2.56, 2.68, 2.76, 2.84, 2.92, 2.96, 3.04, 3.12, 3.32, 3.48, 3.60, 3.68, 3.76, 3.84, 4.00, 4.16, 4.28, 4.44, 4.68, 4.92, 5.00, 5.12, 5.32, 5.56, 5.96, 6.08, 6.20, 6.28, 6.40, 6.48, 6.60, 6.68, 6.80, 7.20, 7.32, 7.44, 7.56, 7.68, 7.96, 8.08, 8.36, 8.48, 8.56, 8.80, 8.92, 9.00, 9.20, 9.28, 9.32, 9.52, 9.68, 9.72, 9.76, 9.84, 9.88, 10.00, 10.12, 10.24, 10.32, 10.40, 10.44, 10.52, 10.64, 10.72, 10.80, 10.88, 10.92, 11.04, 11.20, 11.28, 11.32, 11.40, 11.52, 11.60, 11.80], "durations": [], "tokens":[" B", "i", "o", "p", "s", "y", " is", " a", " me", "d", "ic", "al", " pro", "c", "ed", "ur", "e", " in", " which", " a", " sa", "mp", "le", " of", " ", "t", "is", "s", "u", "e", " is", " re", "m", "o", "v", "ed", " from", " the", " ", "body", " for", " ex", "am", "in", "ation", ".", " O", "st", "e", "o", "p", "or", "o", "s", "is", " is", " a", " con", "d", "ition", " in", " which", " bo", "ne", "s", " be", "c", "ome", " we", "a", "k", " and", " b", "ri", "t", "t", "le", ",", " in", "c", "re", "a", "s", "ing", " the", " ", "ri", "s", "k", " of", " f", "ra", "c", "ture", "s", "."], "ys_log_probs": [], "words": []}
----
./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/2.wav
{"lang": "", "emotion": "", "event": "", "text": "[EXAM TYPE] CT chest PE protocol {period} [INDICATION] 54-year-old female, shortness of breath. Evaluate for PE {period} [TECHNIQUE] Standard protocol {period} [FINDINGS] {colon} Pulmonary vasculature {colon} The main PA is patent {period} There are filling defects in the segmental branches of the right lower lobe {comma} compatible with acute PE {period} No saddle embolus {period} Lungs {colon} No pneumothorax {period} Small bilateral effusions {comma} right greater than left {period}", "timestamps": [0.00, 0.04, 0.12, 0.20, 0.24, 0.44, 0.48, 0.56, 0.60, 0.68, 0.76, 0.88, 1.08, 1.16, 1.24, 1.48, 1.60, 1.84, 1.92, 2.00, 2.08, 2.16, 2.36, 2.44, 2.52, 2.84, 2.92, 3.00, 3.08, 3.12, 3.20, 3.24, 3.32, 3.36, 3.44, 3.48, 3.56, 3.64, 3.72, 3.96, 4.08, 4.16, 4.20, 4.28, 4.36, 4.44, 4.56, 4.64, 4.72, 4.84, 4.92, 5.00, 5.20, 5.28, 5.36, 5.44, 5.56, 5.68, 5.76, 5.84, 5.92, 6.00, 6.12, 6.20, 6.32, 6.40, 6.48, 6.68, 6.88, 7.00, 7.20, 7.28, 7.40, 7.60, 7.68, 7.76, 7.80, 7.88, 7.92, 8.00, 8.04, 8.08, 8.12, 8.20, 8.24, 8.28, 8.36, 8.40, 8.48, 8.68, 8.76, 8.84, 8.92, 9.00, 9.20, 9.28, 9.36, 9.68, 9.76, 9.84, 9.92, 10.00, 10.08, 10.12, 10.20, 10.28, 10.32, 10.40, 10.52, 10.60, 11.00, 11.08, 11.16, 11.24, 11.32, 11.44, 11.60, 11.68, 11.72, 11.80, 11.88, 12.00, 12.12, 12.40, 12.48, 12.56, 12.92, 13.08, 13.16, 13.36, 13.52, 13.76, 13.96, 14.04, 14.12, 14.32, 14.40, 14.48, 14.88, 15.00, 15.20, 15.28, 15.36, 15.60, 15.72, 15.84, 15.92, 16.08, 16.20, 16.36, 16.44, 16.56, 16.72, 16.92, 16.96, 17.04, 17.12, 17.20, 17.36, 17.48, 17.64, 17.92, 18.00, 18.08, 18.28, 18.40, 18.48, 18.64, 18.72, 18.84, 19.08, 19.32, 19.44, 19.76, 19.92, 20.04, 20.12, 20.36, 20.52, 20.72, 20.80, 20.88, 21.36, 21.60, 21.68, 21.80, 21.84, 22.00, 22.08, 22.20, 22.32, 22.52, 22.60, 22.68, 23.12, 23.20, 23.28, 23.36, 23.60, 23.68, 23.76, 24.24, 24.36, 24.44, 24.52, 24.56, 24.64, 24.72, 24.80, 24.88, 24.96, 25.16, 25.24, 25.32, 25.80, 25.88, 25.96, 26.00, 26.16, 26.24, 26.32, 26.40, 26.52, 26.68, 26.76, 26.88, 26.96, 27.08, 27.28, 27.36, 27.44, 27.64, 27.96, 28.12, 28.28, 28.48, 28.76, 28.84, 28.96, 29.28], "durations": [], "tokens":[" [", "E", "X", "A", "M", " T", "Y", "P", "E", "]", " C", "T", " ", "ch", "est", " P", "E", " pro", "t", "o", "c", "ol", " {", "period", "}", " [", "I", "N", "D", "I", "C", "A", "T", "I", "O", "N", "]", " ", "5", "4", "-", "y", "e", "ar", "-", "old", " f", "e", "m", "al", "e", ",", " sh", "or", "t", "ness", " of", " b", "re", "a", "th", ".", " E", "v", "al", "u", "ate", " for", " P", "E", " {", "period", "}", " [", "T", "E", "C", "H", "N", "I", "Q", "U", "E", "]", " S", "t", "an", "d", "ard", " pro", "t", "o", "c", "ol", " {", "period", "}", " [", "F", "I", "N", "D", "I", "N", "G", "S", "]", " {", "colon", "}", " P", "ul", "m", "on", "ar", "y", " ", "v", "a", "s", "cu", "la", "ture", " {", "colon", "}", " The", " ma", "in", " P", "A", " is", " pa", "t", "ent", " {", "period", "}", " There", " are", " f", "ill", "ing", " de", "f", "ect", "s", " in", " the", " se", "g", "ment", "al", " b", "r", "an", "ch", "es", " of", " the", " right", " ", "low", "er", " lo", "b", "e", " {", "comma", "}", " comp", "at", "ible", " with", " a", "cu", "te", " P", "E", " {", "period", "}", " No", " sa", "d", "d", "le", " e", "mb", "ol", "us", " {", "period", "}", " L", "u", "ng", "s", " {", "colon", "}", " No", " p", "ne", "u", "m", "o", "th", "or", "a", "x", " {", "period", "}", " S", "m", "al", "l", " b", "il", "a", "ter", "al", " e", "ff", "us", "ion", "s", " {", "comma", "}", " right", " great", "er", " than", " left", " {", "period", "}"], "ys_log_probs": [], "words": []}
----
num threads: 2
decoding method: greedy_search
Elapsed seconds: 5.489 s
Real time factor (RTF): 5.489 / 85.202 = 0.064

Test audio files are

0.mov
1.mov
2.mov

Summary by CodeRabbit

Release Notes

  • New Features

    • Added support for MedASR CTC models for offline speech recognition
    • Introduced Python example demonstrating offline MedASR decoding with single and batch file processing
    • Added Python API to configure and create MedASR CTC recognizers
  • Tests

    • Expanded automated test workflows to validate MedASR model functionality
    • Enhanced test coverage with multiple audio file scenarios

✏️ Tip: You can customize this high-level summary in your review settings.

@dosubot dosubot Bot added the size:XL This PR changes 500-999 lines, ignoring generated files. label Dec 25, 2025
@coderabbitai

coderabbitai Bot commented Dec 25, 2025 •

Copy link
Copy Markdown

Note

Other AI code review bot(s) detected

CodeRabbit has detected other AI code review bot(s) in this pull request and will avoid duplicating their findings in the review comments. This may lead to a less comprehensive review.

📝 Walkthrough

Walkthrough

Adds first-class support for Google MedASR CTC: new C++ model implementation, config, factory integration, Python bindings and API constructor, example decoding script, CI/test updates, and exporter metadata for subsampling_factor.

Changes

Cohort / File(s) Summary
MedASR model sources
sherpa-onnx/csrc/offline-medasr-ctc-model.h, sherpa-onnx/csrc/offline-medasr-ctc-model.cc, sherpa-onnx/csrc/offline-medasr-ctc-model-config.h, sherpa-onnx/csrc/offline-medasr-ctc-model-config.cc
New OfflineMedAsrCtcModel implementation and its config type (Register/Validate/ToString). Exposes Forward, VocabSize, SubsamplingFactor, and Allocator; reads model metadata including subsampling_factor.
Factory & recognizer integration
sherpa-onnx/csrc/offline-ctc-model.cc, sherpa-onnx/csrc/offline-model-config.h, sherpa-onnx/csrc/offline-model-config.cc, sherpa-onnx/csrc/offline-recognizer-impl.cc, sherpa-onnx/csrc/offline-recognizer-ctc-impl.h
OfflineModelConfig extended with medasr field; OfflineCtcModel factory and OfflineRecognizer selection updated to instantiate MedASR paths; MedASR-specific feature extraction settings and CTC post-processing tweaks (skip EOS token when appropriate, trim leading space).
Build inputs & Python bindings
sherpa-onnx/csrc/CMakeLists.txt, sherpa-onnx/python/csrc/CMakeLists.txt, sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.h, sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.cc, sherpa-onnx/python/csrc/offline-model-config.cc
Add new source to core/Python builds; add pybind11 bindings exposing OfflineMedAsrCtcModelConfig and integrate it into Python OfflineModelConfig binding.
Python API & example
sherpa-onnx/python/sherpa_onnx/offline_recognizer.py, python-api-examples/offline-medasr-ctc-decode-files.py
Add OfflineRecognizer.from_medasr_ctc constructor; new example script demonstrating single-file and batched decoding with timing/RTF reporting.
Tests / CI / scripts
.github/scripts/test-python.sh, .github/workflows/export-medasr-ctc-to-onnx.yaml, scripts/medasr/export_onnx.py
CI/test script additions to download and run MedASR and omnilingual model checks; workflow expanded to loop over multiple test WAVs; exporter now writes "subsampling_factor": 4 into model metadata.
Python header comment tweaks
sherpa-onnx/python/csrc/offline-wenet-ctc-model-config.h, sherpa-onnx/python/csrc/offline-wenet-ctc-model-config.cc
Minor comment/path corrections in header comments.
Build list update
sherpa-onnx/csrc/CMakeLists.txt, sherpa-onnx/python/csrc/CMakeLists.txt
Added new source files to the respective CMake src lists.

Sequence Diagram(s)

sequenceDiagram
    autonumber
    participant Py as Python client
    participant OR as OfflineRecognizer (C++)
    participant RC as OfflineRecognizerCtcImpl
    participant MM as OfflineMedAsrCtcModel
    participant ORT as ONNXRuntime

    Note over Py,OR: Instantiate recognizer
    Py->>OR: OfflineRecognizer.from_medasr_ctc(model, tokens...)
    OR->>RC: create CTC impl with medasr config
    RC->>MM: construct OfflineMedAsrCtcModel (load model, read metadata)
    MM->>ORT: create Session, prepare inputs

    Note over Py,MM: Decode request flow
    Py->>OR: decode(audio)
    OR->>RC: forward to CTC impl
    RC->>MM: MM.Forward(features, features_length)
    MM->>ORT: Run session (features, mask) -> outputs
    ORT-->>MM: raw logits / tokens
    MM-->>RC: outputs
    RC-->>OR: decoded tokens/text (apply MedASR post-processing)
    OR-->>Py: final transcription
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~50 minutes

Possibly related PRs

Poem

🐰 I hopped into C++ and Python trees,
Built masks and metadata with nimble ease.
From WAV to tokens, I scurried fast,
MedASR joins the decoding cast.
Little rabbit cheers — new models at last! 🥕

Pre-merge checks and finishing touches

❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 8.89% which is insufficient. The required threshold is 80.00%. You can run @coderabbitai generate docstrings to improve docstring coverage.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The pull request title accurately describes the main change: adding C++ runtime and Python API support for Google MedASR models, which aligns with the substantial additions across multiple files.
✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello @csukuangfj, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances sherpa-onnx by integrating comprehensive support for Google MedASR CTC models. The changes encompass the foundational C++ model implementation, a streamlined Python API for ease of use, and specialized audio feature processing and text post-processing logic to ensure accurate medical speech recognition. This integration also includes updates to the build system and continuous integration tests, alongside a practical Python example, making MedASR models readily available and verifiable within the framework.

Highlights

  • MedASR Model Integration: Added core C++ implementation for Google MedASR CTC models, including model configuration and the model class itself, enabling the framework to process medical speech recognition tasks.
  • Python API Extension: Introduced a new from_medasr_ctc class method in the Python OfflineRecognizer for convenient instantiation and usage of MedASR models, simplifying their adoption for Python developers.
  • Specialized Feature Processing: Implemented specific feature extraction parameters and text post-processing rules (e.g., skipping </s> token, removing leading spaces) tailored for optimal performance and accuracy with MedASR models.
  • CI/CD and Examples: Integrated MedASR testing into the CI/CD pipeline and provided a new Python example script demonstrating how to decode audio files using the newly added MedASR support, ensuring functionality and providing practical usage guidance.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/export-medasr-ctc-to-onnx.yaml
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🧹 Nitpick comments (3)
python-api-examples/offline-medasr-ctc-decode-files.py (1)

99-130: Consider adding strict=True to zip() for robustness.

The batch decoding logic is correct. For added safety, consider using strict=True with zip() (available in Python 3.10+) to ensure filenames and streams have matching lengths.

Optional enhancement
-    for name, stream in zip(filenames, streams):
+    for name, stream in zip(filenames, streams, strict=True):
sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.cc (1)

8-8: Unused include.

<vector> is included but not used in this file.

🔎 Proposed fix
 #include <string>
-#include <vector>

 #include "sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.h"
sherpa-onnx/csrc/offline-medasr-ctc-model.cc (1)

49-54: Potential signed/unsigned or width mismatch in loop.

i is declared as int32_t while batch_size is int64_t. For very large batch sizes (unlikely but possible), this could cause issues. Consider using a consistent type.

🔎 Proposed fix
-  for (int32_t i = 0; i < batch_size; ++i) {
+  for (int64_t i = 0; i < batch_size; ++i) {
📜 Review details

Configuration used: defaults

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between ee3611d and 1fe5949.

📒 Files selected for processing (21)
  • .github/scripts/test-python.sh
  • .github/workflows/export-medasr-ctc-to-onnx.yaml
  • python-api-examples/offline-medasr-ctc-decode-files.py
  • scripts/medasr/export_onnx.py
  • sherpa-onnx/csrc/CMakeLists.txt
  • sherpa-onnx/csrc/offline-ctc-model.cc
  • sherpa-onnx/csrc/offline-medasr-ctc-model-config.cc
  • sherpa-onnx/csrc/offline-medasr-ctc-model-config.h
  • sherpa-onnx/csrc/offline-medasr-ctc-model.cc
  • sherpa-onnx/csrc/offline-medasr-ctc-model.h
  • sherpa-onnx/csrc/offline-model-config.cc
  • sherpa-onnx/csrc/offline-model-config.h
  • sherpa-onnx/csrc/offline-recognizer-ctc-impl.h
  • sherpa-onnx/csrc/offline-recognizer-impl.cc
  • sherpa-onnx/python/csrc/CMakeLists.txt
  • sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.cc
  • sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.h
  • sherpa-onnx/python/csrc/offline-model-config.cc
  • sherpa-onnx/python/csrc/offline-wenet-ctc-model-config.cc
  • sherpa-onnx/python/csrc/offline-wenet-ctc-model-config.h
  • sherpa-onnx/python/sherpa_onnx/offline_recognizer.py
🧰 Additional context used
🧠 Learnings (2)
📚 Learning: 2025-08-06T04:23:50.237Z
Learnt from: litongjava
Repo: k2-fsa/sherpa-onnx PR: 2440
File: sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/core/Core.java:4-6
Timestamp: 2025-08-06T04:23:50.237Z
Learning: The sherpa-onnx JNI library files are stored in Hugging Face repository at https://huggingface.co/csukuangfj/sherpa-onnx-libs under versioned directories like jni/1.12.7/, and the actual Windows JNI library filename is "sherpa-onnx-jni.dll" as defined in Core.java constants.

Applied to files:

  • sherpa-onnx/csrc/offline-ctc-model.cc
📚 Learning: 2025-08-06T04:18:47.981Z
Learnt from: litongjava
Repo: k2-fsa/sherpa-onnx PR: 2440
File: sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/core/Core.java:4-6
Timestamp: 2025-08-06T04:18:47.981Z
Learning: In sherpa-onnx Java API, the native library names in Core.java (WIN_NATIVE_LIBRARY_NAME = "sherpa-onnx-jni.dll", UNIX_NATIVE_LIBRARY_NAME = "libsherpa-onnx-jni.so", MACOS_NATIVE_LIBRARY_NAME = "libsherpa-onnx-jni.dylib") are copied directly from the compiled binary filenames and should not be changed to match other libraries' naming conventions.

Applied to files:

  • sherpa-onnx/csrc/offline-ctc-model.cc
🧬 Code graph analysis (8)
sherpa-onnx/python/csrc/offline-model-config.cc (1)
sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.cc (2)
  • PybindOfflineMedAsrCtcModelConfig (14-20)
  • PybindOfflineMedAsrCtcModelConfig (14-14)
sherpa-onnx/csrc/offline-medasr-ctc-model-config.h (3)
sherpa-onnx/csrc/offline-model-config.h (1)
  • sherpa_onnx (24-111)
sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.h (1)
  • sherpa_onnx (10-14)
sherpa-onnx/csrc/offline-medasr-ctc-model-config.cc (2)
  • Register (14-19)
  • Register (14-14)
sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.h (2)
sherpa-onnx/csrc/offline-medasr-ctc-model-config.h (1)
  • sherpa_onnx (12-27)
sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.cc (2)
  • PybindOfflineMedAsrCtcModelConfig (14-20)
  • PybindOfflineMedAsrCtcModelConfig (14-14)
.github/scripts/test-python.sh (12)
.github/scripts/test-offline-ctc.sh (1)
  • log (5-9)
.github/scripts/test-offline-tts.sh (1)
  • log (5-9)
scripts/apk/build-apk-kws.sh (1)
  • log (11-15)
.github/scripts/test-offline-transducer.sh (1)
  • log (5-9)
.github/scripts/test-online-transducer.sh (1)
  • log (5-9)
.github/scripts/test-online-ctc.sh (1)
  • log (5-9)
.github/scripts/test-offline-moonshine.sh (1)
  • log (5-9)
.github/scripts/test-online-punctuation.sh (1)
  • log (8-12)
.github/scripts/test-offline-whisper.sh (1)
  • log (5-9)
.github/scripts/test-online-paraformer.sh (1)
  • log (5-9)
.github/scripts/test-spoken-language-identification.sh (1)
  • log (5-9)
.github/scripts/test-offline-punctuation.sh (1)
  • log (5-9)
sherpa-onnx/csrc/offline-medasr-ctc-model.h (3)
sherpa-onnx/csrc/offline-medasr-ctc-model-config.h (1)
  • sherpa_onnx (12-27)
sherpa-onnx/csrc/offline-model-config.h (1)
  • sherpa_onnx (24-111)
sherpa-onnx/csrc/offline-ctc-model.h (1)
  • OfflineCtcModel (16-59)
sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.cc (2)
sherpa-onnx/csrc/offline-medasr-ctc-model-config.cc (2)
  • ToString (30-37)
  • ToString (30-30)
sherpa-onnx/csrc/offline-model-config.cc (2)
  • ToString (177-205)
  • ToString (177-177)
python-api-examples/offline-medasr-ctc-decode-files.py (1)
sherpa-onnx/python/sherpa_onnx/offline_recognizer.py (1)
  • from_medasr_ctc (541-588)
sherpa-onnx/csrc/offline-medasr-ctc-model.cc (1)
sherpa-onnx/csrc/offline-medasr-ctc-model.h (1)
  • OfflineMedAsrCtcModel (25-61)
🪛 Ruff (0.14.10)
python-api-examples/offline-medasr-ctc-decode-files.py

49-53: Avoid specifying long messages outside the exception class

(TRY003)


120-120: zip() without an explicit strict= parameter

Add explicit value for parameter strict=

(B905)

⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (20)
  • GitHub Check: Debug shared-OFF tts-ON
  • GitHub Check: Debug shared-OFF tts-OFF
  • GitHub Check: Debug shared-ON tts-ON
  • GitHub Check: Release shared-ON tts-ON
  • GitHub Check: Release shared-OFF tts-OFF
  • GitHub Check: Release shared-OFF tts-ON
  • GitHub Check: Debug shared-ON tts-OFF
  • GitHub Check: Release shared-ON tts-OFF
  • GitHub Check: Debug shared tts-OFF
  • GitHub Check: Release shared tts-OFF
  • GitHub Check: Release shared tts-ON
  • GitHub Check: Release static tts-OFF
  • GitHub Check: Release static tts-ON
  • GitHub Check: rknn shared ON
  • GitHub Check: ubuntu-24.04 3.10
  • GitHub Check: ubuntu-24.04 3.11
  • GitHub Check: ubuntu-24.04 3.8
  • GitHub Check: ubuntu-24.04 3.9
  • GitHub Check: ubuntu-24.04 3.13
  • GitHub Check: ubuntu-24.04 3.12
🔇 Additional comments (35)
sherpa-onnx/python/csrc/offline-wenet-ctc-model-config.h (1)

1-1: LGTM—Comment updated to match filename.

The header comment now correctly references the actual filename.

scripts/medasr/export_onnx.py (1)

110-110: LGTM—Subsampling factor metadata added.

The subsampling factor metadata is correctly added to support the MedASR CTC model implementation.

sherpa-onnx/python/csrc/offline-wenet-ctc-model-config.cc (1)

1-1: LGTM—Comment updated to match filename.

The header comment now correctly references the actual filename.

python-api-examples/offline-medasr-ctc-decode-files.py (4)

26-66: LGTM—Clear validation and recognizer construction.

The file validation and recognizer creation logic is well-structured with helpful error messages.


69-73: LGTM—Audio loading is correct.

The audio loading with librosa at 16kHz sample rate is appropriate and includes a validation assertion.


76-96: LGTM—Single-file decoding with proper timing.

The timing and real-time factor calculation is correctly implemented.


133-138: LGTM—Main function orchestrates decoding correctly.

The example demonstrates both single-file and batch decoding workflows.

sherpa-onnx/csrc/offline-model-config.h (3)

12-12: LGTM—MedASR header included.


40-40: LGTM—MedASR config member added.

The new member follows the existing pattern for other model configurations.


76-76: LGTM—Constructor updated consistently.

The constructor parameter and initializer list are properly extended to include the MedASR configuration.

Also applies to: 95-95

sherpa-onnx/python/csrc/CMakeLists.txt (1)

17-17: LGTM—MedASR Python binding source added to build.

The source file is correctly added to the CMake build configuration.

sherpa-onnx/csrc/offline-recognizer-impl.cc (2)

215-223: LGTM—MedASR integrated into CTC model routing.

The condition correctly routes MedASR models to the CTC implementation, following the pattern of other CTC models.


546-554: LGTM—Managed Create overload updated consistently.

The template Manager version of Create is correctly updated to include MedASR routing.

sherpa-onnx/csrc/offline-ctc-model.cc (3)

24-24: LGTM—MedASR model header included.


130-131: LGTM—MedASR model instantiation added to factory.

The factory method correctly creates an OfflineMedAsrCtcModel when the MedASR configuration is present.


198-199: LGTM—Managed factory overload updated consistently.

The template Manager version of the factory is correctly extended to support MedASR models.

.github/scripts/test-python.sh (1)

11-20: LGTM! MedASR test workflow follows established patterns.

The test block correctly downloads, extracts, validates, and cleans up the MedASR model following the same structure as other model tests in this script.

sherpa-onnx/python/csrc/offline-model-config.cc (3)

14-14: LGTM! MedASR binding integration follows conventions.

The include and binding registration follow the same pattern as other model configurations.

Also applies to: 42-42


59-62: LGTM! Constructor binding correctly extended for MedASR.

The medasr parameter is properly added to the constructor signature with an appropriate default value.

Also applies to: 76-76


94-94: LGTM! Read/write binding correctly exposed.

The medasr field binding follows the established pattern.

.github/workflows/export-medasr-ctc-to-onnx.yaml (2)

6-6: Verify the branch trigger is correct.

The workflow triggers on branch cpp-medasr-2, which appears to be a development/PR branch. Ensure this is intentional or update it to trigger on the appropriate permanent branch (e.g., master or a feature branch) before merging.


45-47: LGTM! Extended test coverage with multiple audio files.

The loop expansion from a single test file to 6 files (0-5.wav) improves test coverage and follows a consistent pattern across download and test steps.

Also applies to: 58-60, 69-71

sherpa-onnx/csrc/CMakeLists.txt (1)

43-44: LGTM! MedASR source files added to build.

The two source files are correctly added to the CMake sources list with proper placement in the alphabetical order.

sherpa-onnx/csrc/offline-model-config.cc (1)

28-28: LGTM! MedASR configuration properly integrated.

The medasr configuration is correctly integrated into Register, Validate, and ToString methods, following the same pattern as other model configurations.

Also applies to: 160-162, 194-194

sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.h (1)

1-16: LGTM! Standard Python binding header.

The header correctly declares the Python binding function with appropriate include guards and namespace.

sherpa-onnx/csrc/offline-recognizer-ctc-impl.h (3)

42-45: LGTM! End-of-sentence token properly filtered.

The token skipping follows the same pattern as the existing SIL token handling and will only apply when is present in the symbol table.


66-68: Verify if leading space trimming should apply to all models.

This code unconditionally trims a leading space from the decoded text for all CTC models. If this behavior is only needed for MedASR, consider adding a conditional check (e.g., if (!config_.model_config.medasr.model.empty())). If it's intentionally applied to all models, please add a comment explaining why.


156-165: Most MedASR feature extraction parameters are confirmed; frequency range lacks explicit documentation.

The feature configuration has been verified against test_onnx.py:

  • ✓ Confirmed: remove_dc_offset=false, dither=0, preemph_coeff=0, window_type="hanning", feature_dim=128, snip_edges=true
  • Not explicitly documented: low_freq=125 and high_freq=7500 (test script uses kaldi_native_fbank without setting frequency bounds, suggesting they may be defaults or handled implicitly)

The parameters that are documented match the test implementation exactly.

sherpa-onnx/csrc/offline-medasr-ctc-model-config.h (1)

1-29: LGTM! Well-structured configuration header.

The OfflineMedAsrCtcModelConfig struct follows the established pattern for model configuration with appropriate constructors and public methods (Register, Validate, ToString).

sherpa-onnx/python/sherpa_onnx/offline_recognizer.py (1)

540-588: LGTM! The new from_medasr_ctc factory method is well-implemented.

The implementation correctly follows the existing pattern established by from_omnilingual_asr_ctc (lines 590-638), which also omits feat_config for models that handle feature extraction internally. The method signature, documentation, and configuration setup are consistent with the codebase conventions.

sherpa-onnx/csrc/offline-medasr-ctc-model-config.cc (1)

14-37: Implementation looks correct.

The Register(), Validate(), and ToString() methods follow the established pattern used by other model config classes in the codebase.

sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.cc (1)

14-20: Python binding implementation is correct.

The binding correctly exposes the constructor, model attribute, and __str__ method, following the established pattern for other model config bindings.

sherpa-onnx/csrc/offline-medasr-ctc-model.h (1)

1-65: Well-structured header file.

The header correctly implements the PIMPL pattern with std::unique_ptr<Impl>, properly inherits from OfflineCtcModel, and provides comprehensive documentation for the Forward method. The include guard and namespace usage are correct.

sherpa-onnx/csrc/offline-medasr-ctc-model.cc (2)

61-101: Model initialization and forward pass implementation look correct.

The Impl class properly initializes the ONNX Runtime session, extracts input/output names, validates model metadata, and implements the forward pass with mask tensor creation. The memory management and tensor creation follow the established patterns in the codebase.


163-196: Public API implementation and template instantiations are correct.

The forwarding methods properly delegate to the Impl class, and the template instantiations for Android and OHOS platforms follow the established pattern for cross-platform support.

Comment thread sherpa-onnx/csrc/offline-medasr-ctc-model-config.cc
Comment on lines +103 to +107
int32_t VocabSize() const { return vocab_size_; }

int32_t SubsamplingFactor() const { return 4; }

OrtAllocator *Allocator() { return allocator_; }

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

SubsamplingFactor() returns hardcoded value instead of using the metadata field.

The method returns a hardcoded 4 at line 105, but subsampling_factor_ is read from model metadata at lines 141-142 and stored in the member variable. This appears to be a bug or oversight.

Additionally, Allocator() in Impl is non-const (line 107), but the public OfflineMedAsrCtcModel::Allocator() is declared const in the header. This works because impl_ is a unique_ptr, but for consistency the Impl method should also be const.

🔎 Proposed fix
-  int32_t SubsamplingFactor() const { return 4; }
+  int32_t SubsamplingFactor() const { return subsampling_factor_; }

-  OrtAllocator *Allocator() { return allocator_; }
+  OrtAllocator *Allocator() const { return allocator_; }
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
int32_t VocabSize() const { return vocab_size_; }
int32_t SubsamplingFactor() const { return 4; }
OrtAllocator *Allocator() { return allocator_; }
int32_t VocabSize() const { return vocab_size_; }
int32_t SubsamplingFactor() const { return subsampling_factor_; }
OrtAllocator *Allocator() const { return allocator_; }
🤖 Prompt for AI Agents
In sherpa-onnx/csrc/offline-medasr-ctc-model.cc around lines 103-107,
SubsamplingFactor() currently returns a hardcoded 4 instead of the member
subsampling_factor_ (which is read from metadata at ~lines 141-142); change
SubsamplingFactor() to return subsampling_factor_. Also make Impl::Allocator() a
const method to match the public const OfflineMedAsrCtcModel::Allocator()
declaration (i.e., change the signature to be const and return allocator_),
preserving behavior but fixing const-correctness.

Comment thread sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.cc Outdated

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds support for the Google MedASR CTC model to sherpa-onnx. The changes are comprehensive, covering the C++ implementation, Python bindings, a new example script, and updates to the CI test scripts. The code is well-structured and follows the existing patterns in the repository. I have a couple of suggestions to improve code quality and maintainability, mainly regarding code duplication in the Python example and an inconsistency in the C++ model implementation. Overall, this is a solid contribution.

Comment on lines +26 to +66
def create_recognizer():
model = "./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/model.int8.onnx"
tokens = "./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/tokens.txt"
test_wav_0 = "./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/0.wav"
test_wav_1 = "./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/1.wav"
test_wav_2 = "./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/2.wav"
test_wav_3 = "./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/3.wav"
test_wav_4 = "./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/4.wav"
test_wav_5 = "./sherpa-onnx-medasr-ctc-en-int8-2025-12-25/test_wavs/5.wav"

for f in [
model,
tokens,
test_wav_0,
test_wav_1,
test_wav_2,
test_wav_3,
test_wav_4,
test_wav_5,
]:
if not Path(f).is_file():
print(f"{f} does not exist")

raise ValueError(
"""Please download model files from
https://github.com/k2-fsa/sherpa-onnx/releases/tag/asr-models
"""
)
return (
sherpa_onnx.OfflineRecognizer.from_medasr_ctc(
model=model,
tokens=tokens,
num_threads=2,
),
test_wav_0,
test_wav_1,
test_wav_2,
test_wav_3,
test_wav_4,
test_wav_5,
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The create_recognizer function contains a lot of repeated strings for file paths, which makes it harder to read and maintain. You can refactor it by defining a base path and constructing the file paths programmatically. This will make the code cleaner and less error-prone.

def create_recognizer():
    model_dir = Path("./sherpa-onnx-medasr-ctc-en-int8-2025-12-25")
    model = model_dir / "model.int8.onnx"
    tokens = model_dir / "tokens.txt"

    test_wavs = [model_dir / "test_wavs" / f"{i}.wav" for i in range(6)]

    all_files = [model, tokens] + test_wavs
    for f in all_files:
        if not f.is_file():
            print(f"{f} does not exist")

            raise ValueError(
                """Please download model files from
                https://github.com/k2-fsa/sherpa-onnx/releases/tag/asr-models
                """
            )

    recognizer = sherpa_onnx.OfflineRecognizer.from_medasr_ctc(
        model=str(model),
        tokens=str(tokens),
        num_threads=2,
    )
    return (recognizer,) + tuple(str(w) for w in test_wavs)


int32_t VocabSize() const { return vocab_size_; }

int32_t SubsamplingFactor() const { return 4; }

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The SubsamplingFactor() method is hardcoded to return 4, while the Init() function reads the subsampling_factor from the model's metadata into the subsampling_factor_ member variable, which is then unused. This is inconsistent. To ensure the value from the model metadata is used, SubsamplingFactor() should return the subsampling_factor_ member variable.

  int32_t SubsamplingFactor() const { return subsampling_factor_; }

@csukuangfj csukuangfj changed the title Cpp medasr Add C++ runtime and Python API for Google MedASR models Dec 25, 2025
@csukuangfj
csukuangfj requested a review from Copilot December 25, 2025 09:10

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 0

♻️ Duplicate comments (1)
sherpa-onnx/csrc/offline-medasr-ctc-model.cc (1)

103-107: SubsamplingFactor fixed; Allocator const-correctness still needs attention.

Good work fixing SubsamplingFactor() to return subsampling_factor_ instead of the hardcoded value!

However, Allocator() at line 107 should be marked const to match the public interface declaration at line 184. The method only returns a member pointer and doesn't modify state.

🔎 Proposed fix
-  OrtAllocator *Allocator() { return allocator_; }
+  OrtAllocator *Allocator() const { return allocator_; }
🧹 Nitpick comments (1)
sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.cc (1)

7-8: Remove unused <vector> include.

The <vector> header is included but not used in this file. Only <string> is needed for the std::string parameter in the constructor binding.

🔎 Proposed fix
 #include <string>
-#include <vector>
 
 #include "sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.h"
📜 Review details

Configuration used: defaults

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 1fe5949 and 07d2a5c.

📒 Files selected for processing (3)
  • sherpa-onnx/csrc/offline-medasr-ctc-model-config.cc
  • sherpa-onnx/csrc/offline-medasr-ctc-model.cc
  • sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.cc
🧰 Additional context used
🧬 Code graph analysis (3)
sherpa-onnx/csrc/offline-medasr-ctc-model-config.cc (1)
sherpa-onnx/csrc/offline-model-config.cc (6)
  • Register (14-61)
  • Register (14-14)
  • Validate (63-175)
  • Validate (63-63)
  • ToString (177-205)
  • ToString (177-177)
sherpa-onnx/csrc/offline-medasr-ctc-model.cc (1)
sherpa-onnx/csrc/offline-medasr-ctc-model.h (1)
  • OfflineMedAsrCtcModel (25-61)
sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.cc (1)
sherpa-onnx/csrc/offline-medasr-ctc-model-config.cc (2)
  • ToString (31-38)
  • ToString (31-31)
🔇 Additional comments (7)
sherpa-onnx/csrc/offline-medasr-ctc-model-config.cc (1)

7-38: LGTM! Previous issue resolved.

The missing <sstream> include from the previous review has been added. The implementation correctly follows the established pattern for model configurations in this codebase.

Note: Line 19 references PR #2934, while this is PR #2935. Verify if the reference should point to #2935 or if #2934 is a related/prerequisite PR.

sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.cc (1)

14-20: LGTM! Previous copyright issue resolved.

The copyright year has been updated to 2025 as suggested in the previous review. The Python binding implementation correctly exposes the constructor, model attribute, and string representation, following established patterns in the codebase.

sherpa-onnx/csrc/offline-medasr-ctc-model.cc (5)

32-57: LGTM!

The GetMask helper correctly constructs a padding mask from the features_length tensor, following the standard pattern of creating a flat (batch_size × max_len) mask with 1s for valid positions and 0s for padding.


63-80: LGTM!

The constructors correctly initialize the model from either the filesystem or a platform-specific asset manager, following the established pattern used by other offline models in the codebase.


82-101: LGTM!

The Forward method correctly constructs a mask tensor from the features_length input and invokes the ONNX session with both features and mask inputs. The mask shape is properly derived from the features tensor dimensions.


110-143: LGTM!

The Init method properly initializes the ONNX session, validates the model type as "medasr_ctc", and reads metadata including vocab_size and subsampling_factor (with a sensible default of 4). Error handling and debug logging are appropriate.


163-196: LGTM!

The public interface correctly delegates to the Impl class, and platform-specific template instantiations for Android and OHOS are properly provided. This follows the established pattern used throughout the codebase.

@csukuangfj
csukuangfj merged commit 1f5f963 into k2-fsa:master Dec 25, 2025
27 checks passed
@csukuangfj
csukuangfj deleted the cpp-medasr branch December 25, 2025 09:26

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR adds comprehensive support for Google MedASR (Medical Automatic Speech Recognition) CTC models to sherpa-onnx, enabling offline speech recognition for medical domain applications. The implementation includes both C++ runtime support and Python API bindings.

Key Changes

  • Added C++ implementation for MedASR CTC model with custom feature extraction configuration (128-dim features, 125-7500 Hz frequency range, Hanning window)
  • Introduced Python API with from_medasr_ctc() factory method for creating MedASR recognizers
  • Enhanced text post-processing to skip </s> tokens and strip leading spaces specific to MedASR models
  • Fixed file header comments in offline-wenet-ctc-model-config files (corrected from "wenet-model" to "wenet-ctc-model")

Reviewed changes

Copilot reviewed 21 out of 21 changed files in this pull request and generated 4 comments.

Show a summary per file
File Description
sherpa-onnx/python/sherpa_onnx/offline_recognizer.py Added from_medasr_ctc() factory method for creating MedASR recognizers
sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.{h,cc} Python bindings for MedASR model configuration
sherpa-onnx/python/csrc/offline-model-config.cc Integrated MedASR config into Python offline model bindings
sherpa-onnx/csrc/offline-medasr-ctc-model.{h,cc} Core C++ implementation of MedASR CTC model with mask generation
sherpa-onnx/csrc/offline-medasr-ctc-model-config.{h,cc} Configuration structure for MedASR models
sherpa-onnx/csrc/offline-recognizer-ctc-impl.h Added MedASR-specific feature config and text post-processing logic
sherpa-onnx/csrc/offline-recognizer-impl.cc Integrated MedASR into CTC recognizer creation paths
sherpa-onnx/csrc/offline-model-config.{h,cc} Added MedASR config to main model configuration
sherpa-onnx/csrc/offline-ctc-model.cc Added MedASR model instantiation in CTC model factory
sherpa-onnx/csrc/CMakeLists.txt Added MedASR source files to build
sherpa-onnx/python/csrc/CMakeLists.txt Added MedASR Python binding sources to build
sherpa-onnx/python/csrc/offline-wenet-ctc-model-config.{h,cc} Fixed header comment from "wenet-model" to "wenet-ctc-model"
scripts/medasr/export_onnx.py Added subsampling_factor metadata to exported ONNX model
python-api-examples/offline-medasr-ctc-decode-files.py Example demonstrating single and batch file decoding with MedASR
.github/workflows/export-medasr-ctc-to-onnx.yaml Workflow to export and test MedASR models with 6 test audio files
.github/scripts/test-python.sh Added MedASR Python API test to CI pipeline

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

@@ -0,0 +1,22 @@
// sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.cc
//
// Copyright (c) 2025 Xiaomi Corporation

Copilot AI Dec 25, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The copyright year is 2023, but this is a new file created in 2025. The copyright year should be updated to 2025 to match the actual creation date of this file, consistent with the corresponding header file and other new files in this PR.

Copilot uses AI. Check for mistakes.
@@ -0,0 +1,16 @@
// sherpa-onnx/python/csrc/offline-medasr-ctc-model-config.h
//
// Copyright (c) 2023 Xiaomi Corporation

Copilot AI Dec 25, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The copyright year is 2023, but this is a new file created in 2025. The copyright year should be updated to 2025 to match the actual creation date of this file.

Suggested change
// Copyright (c) 2023 Xiaomi Corporation
// Copyright (c) 2025 Xiaomi Corporation

Copilot uses AI. Check for mistakes.
#include "sherpa-onnx/csrc/offline-medasr-ctc-model-config.h"

#include <string>
#include <vector>

Copilot AI Dec 25, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The include directive for vector is not used in this file. Consider removing it to keep the includes clean and minimal.

Suggested change
#include <vector>

Copilot uses AI. Check for mistakes.

int32_t VocabSize() const { return vocab_size_; }

int32_t SubsamplingFactor() const { return subsampling_factor_; }

Copilot AI Dec 25, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The SubsamplingFactor method returns a hardcoded value of 4 instead of using the subsampling_factor_ member variable that is read from the model metadata at line 141-142. This should return subsampling_factor_ to properly respect the value from the model metadata.

Copilot uses AI. Check for mistakes.
@stqc

stqc commented Apr 7, 2026

Copy link
Copy Markdown

is it possible to add support for ngram LM? (NOTE: I am asking if I can do it on my own, is so, how? not asking you to add it)

@csukuangfj

Copy link
Copy Markdown
Collaborator Author

@stqc Already supported.

You need to build HLG by yourself.

HINT: Search for hlg in the sherpa-onnx repository and in the icefall repository.

@stqc

stqc commented Apr 8, 2026

Copy link
Copy Markdown

@stqc Already supported.

You need to build HLG by yourself.

HINT: Search for hlg in the sherpa-onnx repository and in the icefall repository.

if I am not mistaken it isn't available for C api (i am currently deploying it on device iOS)

@csukuangfj

Copy link
Copy Markdown
Collaborator Author

We support 12 programming languages, including c api.

As said before, we suggest you first have a look at how we use hlg for ctc models.

@coderabbitai coderabbitai Bot mentioned this pull request May 20, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XL This PR changes 500-999 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants