Skip to content

Export https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3 to sherpa-onnx - #2500

Merged
csukuangfj merged 16 commits into
k2-fsa:masterfrom
csukuangfj:parakeet-tdt-v3
Aug 16, 2025
Merged

csukuangfj merged 16 commits into
k2-fsa:masterfrom
csukuangfj:parakeet-tdt-v3

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Aug 16, 2025 •

Copy link
Copy Markdown
Collaborator

See
https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3

It supports 25 languages!

Its usage is the same as v2. See doc at
https://k2-fsa.github.io/sherpa/onnx/pretrained_models/offline-transducer/nemo-transducer-models.html#sherpa-onnx-nemo-parakeet-tdt-0-6b-v3-int8-25-european-languages

Just replace v2 with v3.

Download the model

The exported model has been uploaded to
https://github.com/k2-fsa/sherpa-onnx/releases/tag/asr-models
Screenshot 2025-08-16 at 18 02 52

Try it with our Huggingface space

You can also try it at
https://huggingface.co/spaces/k2-fsa/automatic-speech-recognition

Screenshot 2025-08-16 at 18 08 23

Try it in Colab

We have created two colab notebooks to show you how to use
parakeet-tdt-0.6b-v3. One uses CPU, while the other uses NVIDIA GPU.

CPU NVIDIA GPU
URL URL

Summary by CodeRabbit

  • New Features
    • Added Parakeet TDT 0.6b v3 export pipeline with INT8 artifacts and tokens generation.
    • Enabled offline export for v2 using a local model file.
    • Added a new APK model option for v3 (type 40) with preconfigured paths.
  • Tests
    • Introduced automated multi-language ONNX validation for v3.
  • Chores
    • CI now supports v2/v3 matrix, adds disk space cleanup on macOS, and tracks weights for publishing.
  • Style
    • Minor grammar fixes in logs/comments.

@coderabbitai

coderabbitai Bot commented Aug 16, 2025 •

Copy link
Copy Markdown

Walkthrough

Introduces v3 export and testing for NeMo Parakeet TDT 0.6b alongside v2, updates CI to run both versions with per-version artifact handling, disables HF transfer, adds macOS runner cleanup, adjusts v2 export to remove FP16 and use external encoder weights, adds APK/kotlin entries for v3, and makes minor C++ message/comment tweaks.

Changes

Cohort / File(s) Summary
CI workflow: multi-version export and publishing
.github/workflows/export-nemo-parakeet-tdt-0.6b-v2.yaml
Adds matrix for v2/v3, sets HF_HUB_ENABLE_HF_TRANSFER=0, adds macOS disk cleanup, splits run/collect per version, tracks *.weights in LFS, publishes versioned artifacts without tar for v3.
Nemo v2 export refactor
scripts/nemo/parakeet-tdt-0.6b-v2/export_onnx.py, scripts/nemo/parakeet-tdt-0.6b-v2/run.sh
Removes FP16 export and onnxmltools, adds offline .nemo restore, saves encoder.onnx with external data (encoder.weights), updates metadata target, downloads model and test WAV, keeps FP32/INT8 tests.
Nemo v3 export pipeline (new)
scripts/nemo/parakeet-tdt-0.6b-v3/export_onnx.py, scripts/nemo/parakeet-tdt-0.6b-v3/run.sh, scripts/nemo/parakeet-tdt-0.6b-v3/test_onnx.py
Adds v3 export script with tokens.txt creation, ONNX export, dynamic quantization, metadata, and external encoder weights. Adds run.sh to download model/samples, install deps, run tests (FP32 vs INT8). test_onnx.py references v2 tester.
APK model list update
scripts/apk/generate-vad-asr-apk-script.py
Adds v3 int8 model entry; fixes directory stack with a trailing popd for the prior entry.
Kotlin recognizer mapping
sherpa-onnx/kotlin-api/OfflineRecognizer.kt
Adds type=40 mapping for nemo parakeet-tdt-0.6b-v3-int8 directory using nemo_transducer config.
C++ minor text/comment fixes
sherpa-onnx/csrc/offline-tts-model-config.cc, sherpa-onnx/csrc/offline-tts.h
Corrects one error message string and adjusts a comment grammar; no functional changes.

Sequence Diagram(s)

sequenceDiagram
  actor Dev as GitHub Actions
  participant Job as CI Job (Matrix v2/v3)
  participant Script as run.sh (v2/v3)
  participant Export as export_onnx.py
  participant HF as HuggingFace Repo

  Dev->>Job: Start matrix (version=v2,v3)
  Job->>Job: macOS cleanup, set env
  Job->>Script: Execute per-version run.sh
  Script->>Export: Export encoder/decoder/joiner ONNX
  Export-->>Script: ONNX + encoder.weights + tokens.txt
  Script->>Script: Run ONNX tests (FP32 / INT8)
  Job->>HF: Track LFS (*.onnx, *.weights) and push artifacts
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

Poem

A bunny hops through v2 and v3,
Packing ONNX with encoder weights free.
Tokens in tow, quantized delight,
CI splits paths by day and night.
Kotlin whispers, APKs agree—
“Parakeet sings, now version three!” 🐇🎶

Tip

🔌 Remote MCP (Model Context Protocol) integration is now available!

Pro plan users can now connect to remote MCP servers from the Integrations page. Connect with popular remote MCPs such as Notion and Linear to add more context to your reviews and chats.

✨ Finishing Touches
  • 📝 Generate Docstrings
🧪 Generate unit tests
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share
🪧 Tips

Chat

There are 3 ways to chat with CodeRabbit:

  • Review comments: Directly reply to a review comment made by CodeRabbit. Example:
    • I pushed a fix in commit <commit_id>, please review it.
    • Open a follow-up GitHub issue for this discussion.
  • Files and specific lines of code (under the "Files changed" tab): Tag @coderabbitai in a new review comment at the desired location with your query.
  • PR comments: Tag @coderabbitai in a new PR comment to ask questions about the PR branch. For the best results, please provide a very specific query, as very limited context is provided in this mode. Examples:
    • @coderabbitai gather interesting stats about this repository and render them as a table. Additionally, render a pie chart showing the language distribution in the codebase.
    • @coderabbitai read the files in the src/scheduler package and generate a class diagram using mermaid and a README in the markdown format.

Support

Need help? Create a ticket on our support page for assistance with any issues or questions.

CodeRabbit Commands (Invoked using PR/Issue comments)

Type @coderabbitai help to get the list of available commands.

Other keywords and placeholders

  • Add @coderabbitai ignore anywhere in the PR description to prevent this PR from being reviewed.
  • Add @coderabbitai summary to generate the high-level summary at a specific location in the PR description.
  • Add @coderabbitai anywhere in the PR title to generate the title automatically.

CodeRabbit Configuration File (.coderabbit.yaml)

  • You can programmatically configure CodeRabbit by adding a .coderabbit.yaml file to the root of your repository.
  • Please see the configuration documentation for more information.
  • If your editor has YAML language server enabled, you can add the path at the top of this file to enable auto-completion and validation: # yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json

Status, Documentation and Community

  • Visit our Status Page to check the current availability of CodeRabbit.
  • Visit our Documentation for detailed information on how to use CodeRabbit.
  • Join our Discord Community to get help, request features, and share feedback.
  • Follow us on X/Twitter for updates and announcements.

@csukuangfj
csukuangfj requested a review from Copilot August 16, 2025 10:12

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

This PR adds support for the newer version (v3) of NVIDIA's parakeet-tdt-0.6b model by exporting it to sherpa-onnx format. The v3 model supports 25 languages, maintaining the same usage pattern as v2 but with improved language coverage.

Key changes:

  • Added export scripts and configuration for parakeet-tdt-0.6b-v3 model
  • Updated Kotlin API to include the new model configuration (case 40)
  • Fixed minor grammar issues in existing code comments

Reviewed Changes

Copilot reviewed 10 out of 10 changed files in this pull request and generated 2 comments.

Show a summary per file
File Description
sherpa-onnx/kotlin-api/OfflineRecognizer.kt Added model configuration for parakeet-tdt-0.6b-v3 (case 40)
sherpa-onnx/csrc/offline-tts.h Fixed grammar in comment ("stop" -> "stops")
sherpa-onnx/csrc/offline-tts-model-config.cc Fixed grammar in error message ("at exactly" -> "exactly")
scripts/nemo/parakeet-tdt-0.6b-v3/ Added complete export pipeline for v3 model including test script, run script, and ONNX export
scripts/nemo/parakeet-tdt-0.6b-v2/ Cleaned up v2 export scripts by removing fp16 support and unused dependencies
scripts/apk/generate-vad-asr-apk-script.py Added v3 model to APK generation script
.github/workflows/export-nemo-parakeet-tdt-0.6b-v2.yaml Updated CI workflow to support both v2 and v3 model exports

Tip: Customize your code reviews with copilot-instructions.md. Create the file or learn how to get started.

for w in en.wav de.wav fr.wav es.wav; do
echo "---fp32----"
python3 ./test_onnx.py \
--encoder ./encoder.int8.onnx \

Copilot AI Aug 16, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The encoder path is inconsistent between fp32 and int8 tests. The fp32 test uses encoder.int8.onnx instead of encoder.onnx, which means it's not actually testing the fp32 model.

Suggested change
--encoder ./encoder.int8.onnx \
--encoder ./encoder.onnx \

Copilot uses AI. Check for mistakes.
@@ -0,0 +1 @@
../parakeet-tdt-0.6b-v2/test_onnx.py No newline at end of file

Copilot AI Aug 16, 2025

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

[nitpick] Using a symbolic link to share test code creates a dependency between v2 and v3 directories. Consider copying the test file or creating a shared test module to avoid this coupling.

Suggested change
../parakeet-tdt-0.6b-v2/test_onnx.py
# Contents of test_onnx.py copied from ../parakeet-tdt-0.6b-v2/test_onnx.py
# (Paste the actual code from the v2 test_onnx.py here)

Copilot uses AI. Check for mistakes.
@csukuangfj
csukuangfj merged commit 091e6ff into k2-fsa:master Aug 16, 2025
@csukuangfj
csukuangfj deleted the parakeet-tdt-v3 branch August 16, 2025 10:15

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (18)
sherpa-onnx/csrc/offline-tts-model-config.cc (1)

27-47: Optional: Align validation logic with “exactly one” wording.

Current logic accepts the first provided model and ignores additional ones, while the message says “exactly one.” Consider enforcing exclusivity to avoid silent misconfiguration.

Example refactor:

 bool OfflineTtsModelConfig::Validate() const {
   if (num_threads < 1) {
     SHERPA_ONNX_LOGE("num_threads should be > 0. Given %d", num_threads);
     return false;
   }

-  if (!vits.model.empty()) {
-    return vits.Validate();
-  }
-
-  if (!matcha.acoustic_model.empty()) {
-    return matcha.Validate();
-  }
-
-  if (!kokoro.model.empty()) {
-    return kokoro.Validate();
-  }
-
-  if (!kitten.model.empty()) {
-    return kitten.Validate();
-  }
-
-  SHERPA_ONNX_LOGE("Please provide exactly one tts model.");
-  return false;
+  int provided = 0;
+  provided += !vits.model.empty();
+  provided += !matcha.acoustic_model.empty();
+  provided += !kokoro.model.empty();
+  provided += !kitten.model.empty();
+
+  if (provided == 1) {
+    if (!vits.model.empty()) return vits.Validate();
+    if (!matcha.acoustic_model.empty()) return matcha.Validate();
+    if (!kokoro.model.empty()) return kokoro.Validate();
+    return kitten.Validate();
+  }
+
+  if (provided == 0) {
+    SHERPA_ONNX_LOGE("Please provide exactly one tts model.");
+  } else {
+    SHERPA_ONNX_LOGE("Multiple TTS models are provided. Please provide exactly one.");
+  }
+  return false;
 }

Also applies to: 49-51

scripts/nemo/parakeet-tdt-0.6b-v2/export_onnx.py (4)

32-42: Make external-data check robust to path variations

Relying on an exact filename string makes this brittle if callers pass a relative/absolute path (e.g., "./encoder.onnx"). Use basename to decide when to externalize weights.

Apply:

-    if filename == "encoder.onnx":
+    if os.path.basename(filename) == "encoder.onnx":

14-31: Broaden type hints for meta_data to match actual usage

You pass ints for several metadata fields and cast to str when saving. The current annotation Dict[str, str] is misleading and may trigger static type checker noise. Prefer Mapping[str, Any].

Apply within this function signature:

-def add_meta_data(filename: str, meta_data: Dict[str, str]):
+def add_meta_data(filename: str, meta_data: "Mapping[str, Any]"):

Additionally apply outside the selected range to fix imports:

# at line 6 (imports)
from typing import Dict
# change to:
from typing import Mapping, Any

58-63: Avoid depending on loop index ‘i’ after the loop

Using the last value of ‘i’ to compute the index is fragile. Compute from len(vocabulary) instead.

-    with open("./tokens.txt", "w", encoding="utf-8") as f:
-        for i, s in enumerate(asr_model.joint.vocabulary):
-            f.write(f"{s} {i}\n")
-        f.write(f"<blk> {i+1}\n")
+    with open("./tokens.txt", "w", encoding="utf-8") as f:
+        vocab = list(asr_model.joint.vocabulary)
+        for i, s in enumerate(vocab):
+            f.write(f"{s} {i}\n")
+        f.write(f"<blk> {len(vocab)}\n")

95-97: Optional: also write metadata to decoder/joiner

If downstream tooling ever inspects metadata on decoder/joiner, you’ll want parity. Current runtime may only read encoder metadata, so this is optional.

Example:

for m in ["decoder", "joiner"]:
    add_meta_data(f"{m}.onnx", meta_data)
    add_meta_data(f"{m}.int8.onnx", meta_data)
scripts/nemo/parakeet-tdt-0.6b-v3/export_onnx.py (3)

32-42: Use basename when deciding to externalize encoder weights

Same robustness concern as v2: compare basenames so the behavior is consistent if a path is provided.

-    if filename == "encoder.onnx":
+    if os.path.basename(filename) == "encoder.onnx":

14-31: Adjust meta_data typing to reflect non-string values

You store ints (e.g., vocab_size) and cast on save. Update the type to prevent type-checker complaints.

-def add_meta_data(filename: str, meta_data: Dict[str, str]):
+def add_meta_data(filename: str, meta_data: "Mapping[str, Any]"):

Also update imports outside the selected range:

# change:
from typing import Dict
# to:
from typing import Mapping, Any

58-63: Compute index from vocabulary length

Avoid dependence on the loop index variable outside its natural scope.

-    with open("./tokens.txt", "w", encoding="utf-8") as f:
-        for i, s in enumerate(asr_model.joint.vocabulary):
-            f.write(f"{s} {i}\n")
-        f.write(f"<blk> {i+1}\n")
+    with open("./tokens.txt", "w", encoding="utf-8") as f:
+        vocab = list(asr_model.joint.vocabulary)
+        for i, s in enumerate(vocab):
+            f.write(f"{s} {i}\n")
+        f.write(f"<blk> {len(vocab)}\n")
scripts/nemo/parakeet-tdt-0.6b-v2/run.sh (3)

12-15: Harden downloads to fail fast on HTTP errors

Use curl -f to fail on non-2xx and catch network issues earlier.

-curl -SL -O https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2/resolve/main/parakeet-tdt-0.6b-v2.nemo
+curl -fSL -O https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2/resolve/main/parakeet-tdt-0.6b-v2.nemo
@@
-curl -SL -O https://dldata-public.s3.us-east-2.amazonaws.com/2086-149220-0033.wav
+curl -fSL -O https://dldata-public.s3.us-east-2.amazonaws.com/2086-149220-0033.wav

30-37: Label is misleading: “fp32” test uses an INT8 encoder

The “fp32” section actually uses encoder.int8.onnx. Consider renaming the marker to reduce confusion, e.g., “mixed (int8 encoder + fp32 dec/joiner)”.

-echo "---fp32----"
+echo "---mixed (int8-enc + fp32-dec/joiner)----"

6-10: Remove unused log() helper or use it

The log() function isn’t used. Either remove it or adopt it for the echo statements for consistency.

Optional removal:

-log() {
-  # This function is from espnet
-  local fname=${BASH_SOURCE[1]##*/}
-  echo -e "$(date '+%Y-%m-%d %H:%M:%S') (${fname}:${BASH_LINENO[0]}:${FUNCNAME[1]}) $*"
-}
+
scripts/nemo/parakeet-tdt-0.6b-v3/run.sh (3)

12-18: Fail fast on download errors

Add -f to curl to fail on 4xx/5xx.

-curl -SL -O https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3/resolve/main/parakeet-tdt-0.6b-v3.nemo
+curl -fSL -O https://huggingface.co/nvidia/parakeet-tdt-0.6b-v3/resolve/main/parakeet-tdt-0.6b-v3.nemo
@@
-curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/en.wav
+curl -fSL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/en.wav
-curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/de.wav
+curl -fSL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/de.wav
-curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/fr.wav
+curl -fSL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/fr.wav
-curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/es.wav
+curl -fSL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/es.wav

35-43: Clarify the test label—this isn’t fully fp32

As with v2, the “fp32” section uses an INT8 encoder; rename to avoid confusion.

-  echo "---fp32----"
+  echo "---mixed (int8-enc + fp32-dec/joiner)----"

6-10: Remove unused log() helper

The helper isn’t used—drop it to reduce noise.

-log() {
-  # This function is from espnet
-  local fname=${BASH_SOURCE[1]##*/}
-  echo -e "$(date '+%Y-%m-%d %H:%M:%S') (${fname}:${BASH_LINENO[0]}:${FUNCNAME[1]}) $*"
-}
+
.github/workflows/export-nemo-parakeet-tdt-0.6b-v2.yaml (4)

60-66: Fix glob safety warnings (SC2035) in v2 step

Shellcheck suggests guarding globs so filenames starting with dashes are not interpreted as options. Also safer if globs ever expand to nothing.

-          ls -lh *.onnx
-          ls -lh *.weights
+          ls -lh -- *.onnx
+          ls -lh -- *.weights
@@
-          mv -v *.onnx ../../..
-          mv -v *.weights ../../..
+          mv -v -- ./*.onnx ../../..
+          mv -v -- ./*.weights ../../..

75-80: Fix glob safety warnings (SC2035) in v3 step

Guard glob expansions for mv.

-          mv -v *.onnx ../../..
-          mv -v *.weights ../../..
-          mv -v tokens.txt ../../..
-          mv *.wav ../../../
+          mv -v -- ./*.onnx ../../..
+          mv -v -- ./*.weights ../../..
+          mv -v -- tokens.txt ../../..
+          mv -- ./*.wav ../../../

87-95: Fix glob safety when copying wavs (SC2035)

Use “--” and explicit ./ to avoid option confusion.

-          cp -v *.wav $d/test_wavs
+          cp -v -- ./*.wav "$d/test_wavs"

106-113: Fix glob safety for int8 artifacts (SC2035)

Same shellcheck warning applies here.

-          cp -v encoder.int8.onnx $d
-          cp -v decoder.int8.onnx $d
-          cp -v joiner.int8.onnx $d
-          cp -v tokens.txt $d
+          cp -v -- encoder.int8.onnx "$d"
+          cp -v -- decoder.int8.onnx "$d"
+          cp -v -- joiner.int8.onnx "$d"
+          cp -v -- tokens.txt "$d"
@@
-          cp -v *.wav $d/test_wavs
+          cp -v -- ./*.wav "$d/test_wavs"
📜 Review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

💡 Knowledge Base configuration:

  • MCP integration is disabled by default for public repositories
  • Jira integration is disabled by default for public repositories
  • Linear integration is disabled by default for public repositories

You can enable these sources in your CodeRabbit configuration.

📥 Commits

Reviewing files that changed from the base of the PR and between 4dfb39c and 335b4ec.

📒 Files selected for processing (10)
  • .github/workflows/export-nemo-parakeet-tdt-0.6b-v2.yaml (4 hunks)
  • scripts/apk/generate-vad-asr-apk-script.py (1 hunks)
  • scripts/nemo/parakeet-tdt-0.6b-v2/export_onnx.py (3 hunks)
  • scripts/nemo/parakeet-tdt-0.6b-v2/run.sh (1 hunks)
  • scripts/nemo/parakeet-tdt-0.6b-v3/export_onnx.py (1 hunks)
  • scripts/nemo/parakeet-tdt-0.6b-v3/run.sh (1 hunks)
  • scripts/nemo/parakeet-tdt-0.6b-v3/test_onnx.py (1 hunks)
  • sherpa-onnx/csrc/offline-tts-model-config.cc (1 hunks)
  • sherpa-onnx/csrc/offline-tts.h (1 hunks)
  • sherpa-onnx/kotlin-api/OfflineRecognizer.kt (1 hunks)
🧰 Additional context used
🧬 Code Graph Analysis (4)
sherpa-onnx/kotlin-api/OfflineRecognizer.kt (3)
sherpa-onnx/c-api/cxx-api.h (1)
  • OfflineModelConfig (269-290)
scripts/go/sherpa_onnx.go (2)
  • OfflineModelConfig (461-490)
  • OfflineTransducerModelConfig (385-389)
scripts/dotnet/OfflineModelConfig.cs (1)
  • OfflineModelConfig (11-32)
scripts/nemo/parakeet-tdt-0.6b-v3/export_onnx.py (2)
scripts/nemo/parakeet-tdt-0.6b-v2/export_onnx.py (2)
  • add_meta_data (14-42)
  • main (46-97)
scripts/spleeter/export_onnx.py (1)
  • export (47-74)
scripts/apk/generate-vad-asr-apk-script.py (6)
scripts/apk/generate-asr-apk-script.py (1)
  • Model (27-44)
scripts/lazarus/generate-subtitles.py (1)
  • Model (28-32)
scripts/apk/generate-asr-2pass-apk-script.py (1)
  • Model (27-43)
scripts/hap/generate-vad-asr-hap-script.py (1)
  • Model (28-48)
scripts/mobile-asr-models/generate-asr.py (1)
  • Model (26-31)
scripts/mobile-asr-models/generate-kws.py (1)
  • Model (26-31)
scripts/nemo/parakeet-tdt-0.6b-v2/export_onnx.py (1)
scripts/nemo/parakeet-tdt-0.6b-v3/export_onnx.py (2)
  • main (46-97)
  • add_meta_data (14-42)
🪛 actionlint (1.7.7)
.github/workflows/export-nemo-parakeet-tdt-0.6b-v2.yaml

56-56: shellcheck reported issue in this script: SC2035:info:4:8: Use ./glob or -- glob so names with dashes won't become options

(shellcheck)


56-56: shellcheck reported issue in this script: SC2035:info:5:8: Use ./glob or -- glob so names with dashes won't become options

(shellcheck)


56-56: shellcheck reported issue in this script: SC2035:info:7:7: Use ./glob or -- glob so names with dashes won't become options

(shellcheck)


56-56: shellcheck reported issue in this script: SC2035:info:8:7: Use ./glob or -- glob so names with dashes won't become options

(shellcheck)


71-71: shellcheck reported issue in this script: SC2035:info:4:8: Use ./glob or -- glob so names with dashes won't become options

(shellcheck)


71-71: shellcheck reported issue in this script: SC2035:info:5:7: Use ./glob or -- glob so names with dashes won't become options

(shellcheck)


71-71: shellcheck reported issue in this script: SC2035:info:6:7: Use ./glob or -- glob so names with dashes won't become options

(shellcheck)


71-71: shellcheck reported issue in this script: SC2035:info:8:4: Use ./glob or -- glob so names with dashes won't become options

(shellcheck)


83-83: shellcheck reported issue in this script: SC2035:info:11:7: Use ./glob or -- glob so names with dashes won't become options

(shellcheck)


102-102: shellcheck reported issue in this script: SC2035:info:10:7: Use ./glob or -- glob so names with dashes won't become options

(shellcheck)


102-102: shellcheck reported issue in this script: SC2035:info:14:8: Use ./glob or -- glob so names with dashes won't become options

(shellcheck)

🔇 Additional comments (8)
sherpa-onnx/csrc/offline-tts-model-config.cc (1)

49-49: Grammar fix LGTM.

Message reads better now; no behavior change.

sherpa-onnx/csrc/offline-tts.h (1)

68-70: Comment tweak LGTM.

Clearer phrasing; interface unchanged.

sherpa-onnx/kotlin-api/OfflineRecognizer.kt (1)

681-692: Type=40 mapping for Parakeet TDT 0.6b v3 is correct and verified.

  • GitHub release includes sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8.tar.bz2 (size 487 170 055 bytes).
  • Paths align with the APK’s modelDir and int8 artifact names.
  • modelType = "nemo_transducer" matches the other transducer cases.

Optional: update the comment block listing example types if you’d like Type 40 to appear there for discoverability.

scripts/nemo/parakeet-tdt-0.6b-v3/export_onnx.py (1)

73-85: LGTM: metadata content is consistent with v2 and runtime expectations

Fields cover what sherpa-onnx expects (version 2 format, subsampling factor, etc.). URL correctly points to v3.

.github/workflows/export-nemo-parakeet-tdt-0.6b-v2.yaml (3)

13-15: LGTM: disabling HF transfer is appropriate for CI

HF_HUB_ENABLE_HF_TRANSFER=0 avoids the optional transfer backend which can be flaky on CI runners.


31-47: Nice addition: macOS runner cleanup

Freeing disk space on macOS runners often prevents failures. Good call.


118-153: HF publish flow looks solid; remember to LFS-track weights (done)

You already track *.onnx, *.wav, and *.weights. This aligns with adding external data for encoder.onnx. No changes needed.

scripts/apk/generate-vad-asr-apk-script.py (1)

681-683: Verified: All cmd blocks have balanced pushd/popd
Ran a sanity‐check script against scripts/apk/generate-vad-asr-apk-script.py and found no imbalanced blocks.

Comment on lines +684 to 699
Model(
model_name="sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8",
idx=40,
lang="multi",
lang2="25_languages",
short_name="parakeet_tdt_0.6b_v3",
cmd="""
pushd $model_name

rm -rfv test_wavs

ls -lh

popd
""",
),

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Verification agent

🧩 Analysis chain

Add v3 model entry (idx=40): looks correct; align lang fields and verify downstream mappings/artifact

The new entry is consistent with existing Nemo Parakeet entries. Two small nits and a couple of verifications:

  • Nit: lang/lang2 values are inconsistent with other multi-language entries (e.g., idx=25 uses multi_lang). Consider aligning for consistency and potential template logic.
  • Verify Android/Kotlin mapping includes type=40 -> modelDir "sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8".
  • Verify the release artifact exists at the expected URL.
  • If Harmony/HAP packaging should also expose this model, ensure scripts/hap/generate-vad-asr-hap-script.py contains a matching entry.

Proposed tiny consistency tweak:

-            lang="multi",
-            lang2="25_languages",
+            lang="multi_lang",
+            lang2="multi_lang",

Check Kotlin mapping and HAP script:

Verify the release artifact exists (HTTP 200 expected, after redirects):


🏁 Script executed:

#!/bin/bash
set -euo pipefail

echo "Searching Kotlin for type=40 mapping and modelDir..."
fd -t f 'OfflineRecognizer.kt' | while read -r f; do
  echo "==> $f"
  rg -n -C3 -P '(case\s+40\b|parakeet-tdt-0.6b-v3-int8|nemo_transducer)' "$f" || true
done

echo
echo "Searching HAP generator for a v3 entry..."
fd -t f 'generate-vad-asr-hap-script.py' | while read -r f; do
  echo "==> $f"
  rg -n -C2 -P '(parakeet-tdt-0.6b-v3-int8|idx\s*=\s*40)' "$f" || true
done

Length of output: 2226


🏁 Script executed:

#!/bin/bash
set -euo pipefail
url="https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8.tar.bz2"
echo "Checking: $url"
curl -sI -L "$url" | sed -n '1p;/^location:/Ip'

Length of output: 1573


Add v3 model entry (idx=40): mapping & artifact verified; align lang fields; extend HAP script

  • Verified Android/Kotlin mapping in sherpa-onnx/kotlin-api/OfflineRecognizer.kt (case 40 → modelDir = "sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8").
  • Release artifact URL returns a valid HTTP 302 redirect to the expected .tar.bz2.
  • No matching block found in scripts/hap/generate-vad-asr-hap-script.py – please add an entry for idx=40 / parakeet-tdt-0.6b-v3-int8.
  • Optional nit: align lang/lang2 with other multi-language models:
-            lang="multi",
-            lang2="25_languages",
+            lang="multi_lang",
+            lang2="multi_lang",
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
Model(
model_name="sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8",
idx=40,
lang="multi",
lang2="25_languages",
short_name="parakeet_tdt_0.6b_v3",
cmd="""
pushd $model_name
rm -rfv test_wavs
ls -lh
popd
""",
),
Model(
model_name="sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8",
idx=40,
lang="multi_lang",
lang2="multi_lang",
short_name="parakeet_tdt_0.6b_v3",
cmd="""
pushd $model_name
rm -rfv test_wavs
ls -lh
popd
""",
),
🤖 Prompt for AI Agents
In scripts/apk/generate-vad-asr-apk-script.py around lines 684 to 699, you added
a new Model entry for sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8 (idx=40) but
did not add the corresponding entry in
scripts/hap/generate-vad-asr-hap-script.py; also ensure lang/lang2 values match
other multi-language models. Add a matching block for idx=40 in
scripts/hap/generate-vad-asr-hap-script.py that mirrors existing multi-language
model entries (use modelDir = "sherpa-onnx-nemo-parakeet-tdt-0.6b-v3-int8"), and
align the lang and lang2 fields in the new APK model entry to the same values
used by other multi models (e.g., lang="multi" and lang2="25_languages") so
mappings and release artifact handling remain consistent.

@@ -0,0 +1 @@
../parakeet-tdt-0.6b-v2/test_onnx.py No newline at end of file

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue

Top-of-file path literal will raise a SyntaxError.

A bare path at Line 1 is invalid Python and will prevent this file from running/importing. If it’s only a breadcrumb, make it a comment; otherwise remove it.

Apply one of the fixes:

-../parakeet-tdt-0.6b-v2/test_onnx.py
+# This file mirrors the v2 test logic:
+# ../parakeet-tdt-0.6b-v2/test_onnx.py

If the intention is to reuse v2 without duplicating code, replace the file with a thin wrapper that executes v2’s test:

# Reuse v2's test_onnx.py
import importlib.util
from pathlib import Path

_v2 = (Path(__file__).resolve().parent.parent / "parakeet-tdt-0.6b-v2" / "test_onnx.py")
spec = importlib.util.spec_from_file_location("test_onnx_v2", _v2)
mod = importlib.util.module_from_spec(spec)
assert spec.loader is not None
spec.loader.exec_module(mod)

if __name__ == "__main__":
    mod.main()
🤖 Prompt for AI Agents
In scripts/nemo/parakeet-tdt-0.6b-v3/test_onnx.py at line 1 there is a bare path
literal which is invalid Python and will raise a SyntaxError; fix by either
converting the path breadcrumb into a comment (prefix with #) or, if you intend
to reuse the v2 test, replace this file with a thin wrapper that imports and
executes the parakeet-tdt-0.6b-v2/test_onnx.py module at runtime (use
importlib.util.spec_from_file_location, module_from_spec,
spec.loader.exec_module and call its main() when __name__ == "__main__") so no
duplicate code remains.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants