Skip to content

Refactor exporting NeMo models - #2362

Merged
csukuangfj merged 8 commits into
k2-fsa:masterfrom
csukuangfj:refactor-export-nemo
Jul 9, 2025
Merged

csukuangfj merged 8 commits into
k2-fsa:masterfrom
csukuangfj:refactor-export-nemo

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Jul 9, 2025 •

Copy link
Copy Markdown
Collaborator

See also

English Japanese
Original huggingface repo parakeet-tdt_ctc-110m parakeet-tdt_ctc-0.6b-ja
Doc in sherpa-onnx sherpa-onnx-nemo-parakeet_tdt_ctc_110m-en-36000 sherpa-onnx-nemo-parakeet-tdt_ctc-0.6b-ja-35000-int8
Real-time ASR Android APK Download Download

Summary by CodeRabbit

  • New Features

    • Added support for exporting, quantizing (int8), testing, and publishing new NeMo Fast Conformer and Parakeet TDT ASR models, including Japanese and English variants.
    • Introduced automation for model export, packaging, and publishing to Hugging Face and GitHub releases.
    • Enhanced Android model selection with new int8 quantized model options for Japanese and English.
  • Documentation

    • Updated model export documentation to include new Parakeet TDT model links.
  • Chores

    • Improved workflow automation to handle and publish additional int8 model variants across multiple pipelines.

@coderabbitai

coderabbitai Bot commented Jul 9, 2025 •

Copy link
Copy Markdown

Walkthrough

The changes introduce support for int8 quantized ONNX models across multiple NeMo ASR model export workflows, scripts, and documentation. New workflow steps automate quantization, packaging, publishing, and testing of these int8 models. Additional model variants and support for the Parakeet TDT CTC model (Japanese) are also added, with corresponding scripts and integration into the Kotlin API.

Changes

File(s) / Path(s) Change Summary
.github/workflows/export-nemo-fast-conformer-hybrid-transducer-ctc-non-streaming.yaml
.github/workflows/export-nemo-fast-conformer-hybrid-transducer-transducer-non-streaming.yaml
Added int8 model variants to arrays for publishing and compression; appended new quantized models to workflow steps.
.github/workflows/export-nemo-fast-conformer-hybrid-transducer-ctc.yaml
.github/workflows/export-nemo-fast-conformer-hybrid-transducer-transducer.yaml
Refactored to loop over expanded list of model directories (including int8 variants) for copying, archiving, and publishing to Hugging Face; added new publishing steps with retry and Git LFS setup.
.github/workflows/export-nemo-parakeet-tdt.yaml New workflow for exporting, quantizing, packaging, publishing, and releasing the Parakeet TDT CTC Japanese model.
scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-ctc.py
scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-ctc-non-streaming.py
scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-transducer.py
scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-transducer-non-streaming.py
Added dynamic quantization step: imports ONNX quantization modules, quantizes exported ONNX models to int8, saves as new files.
scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-ctc.sh
scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-ctc-non-streaming.sh
scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-transducer.sh
scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-transducer-non-streaming.sh
Updated to handle, organize, and test both original and int8 quantized ONNX models for all model variants; added logic for int8 directories and test execution.
scripts/nemo/fast-conformer-hybrid-transducer-ctc/README.md Added Hugging Face URL for the new Parakeet TDT CTC Japanese model.
scripts/nemo/parakeet-tdt_ctc-0.6b-ja/export-onnx-ctc.py New script: exports Parakeet TDT CTC Japanese model to ONNX, writes tokens, adds metadata, and quantizes to int8.
scripts/nemo/parakeet-tdt_ctc-0.6b-ja/run-ctc.sh New script: automates export, packaging, test data download, and model testing for the Parakeet TDT CTC Japanese model.
scripts/nemo/parakeet-tdt_ctc-0.6b-ja/test-onnx-ctc-non-streaming.py New file: references shared test script for ONNX CTC non-streaming models.
scripts/apk/generate-vad-asr-apk-script.py Added two new Model definitions for English and Japanese Parakeet TDT int8 models to the returned model list.
sherpa-onnx/kotlin-api/OfflineRecognizer.kt Added two new cases to support Parakeet TDT int8 models in the Kotlin API's offline model configuration function.

Sequence Diagram(s)

sequenceDiagram
    participant Workflow
    participant ExportScript
    participant ONNX
    participant Quantizer
    participant Archive
    participant HuggingFace

    Workflow->>ExportScript: Trigger export (model, tokens)
    ExportScript->>ONNX: Export model to model.onnx
    ExportScript->>Quantizer: Quantize model.onnx to model.int8.onnx
    ExportScript->>Archive: Package models and tokens
    Workflow->>HuggingFace: Publish original and int8 models
Loading
sequenceDiagram
    participant User
    participant RunScript
    participant ModelDir
    participant Tester

    User->>RunScript: Start run-ctc.sh / run-transducer.sh
    RunScript->>ModelDir: Organize original and int8 models
    RunScript->>Tester: Test original model
    RunScript->>Tester: Test int8 model
Loading

Poem

🐇
New models hop into the light,
Now quantized, int8 and bright!
Scripts and workflows leap ahead,
Parakeet sings in Japanese thread.
Hugging Face gets a fresh new pack—
All thanks to this update track!
🥕✨


📜 Recent review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 242e8d3 and 67f2587.

📒 Files selected for processing (1)
  • scripts/apk/generate-vad-asr-apk-script.py (1 hunks)
🚧 Files skipped from review as they are similar to previous changes (1)
  • scripts/apk/generate-vad-asr-apk-script.py
✨ Finishing Touches
  • 📝 Generate Docstrings

🪧 Tips

Chat

There are 3 ways to chat with CodeRabbit:

  • Review comments: Directly reply to a review comment made by CodeRabbit. Example:
    • I pushed a fix in commit <commit_id>, please review it.
    • Explain this complex logic.
    • Open a follow-up GitHub issue for this discussion.
  • Files and specific lines of code (under the "Files changed" tab): Tag @coderabbitai in a new review comment at the desired location with your query. Examples:
    • @coderabbitai explain this code block.
    • @coderabbitai modularize this function.
  • PR comments: Tag @coderabbitai in a new PR comment to ask questions about the PR branch. For the best results, please provide a very specific query, as very limited context is provided in this mode. Examples:
    • @coderabbitai gather interesting stats about this repository and render them as a table. Additionally, render a pie chart showing the language distribution in the codebase.
    • @coderabbitai read src/utils.ts and explain its main purpose.
    • @coderabbitai read the files in the src/scheduler package and generate a class diagram using mermaid and a README in the markdown format.
    • @coderabbitai help me debug CodeRabbit configuration file.

Support

Need help? Create a ticket on our support page for assistance with any issues or questions.

Note: Be mindful of the bot's finite context window. It's strongly recommended to break down tasks such as reading entire modules into smaller chunks. For a focused discussion, use review comments to chat about specific files and their changes, instead of using the PR comments.

CodeRabbit Commands (Invoked using PR comments)

  • @coderabbitai pause to pause the reviews on a PR.
  • @coderabbitai resume to resume the paused reviews.
  • @coderabbitai review to trigger an incremental review. This is useful when automatic reviews are disabled for the repository.
  • @coderabbitai full review to do a full review from scratch and review all the files again.
  • @coderabbitai summary to regenerate the summary of the PR.
  • @coderabbitai generate docstrings to generate docstrings for this PR.
  • @coderabbitai generate sequence diagram to generate a sequence diagram of the changes in this PR.
  • @coderabbitai resolve resolve all the CodeRabbit review comments.
  • @coderabbitai configuration to show the current CodeRabbit configuration for the repository.
  • @coderabbitai help to get help.

Other keywords and placeholders

  • Add @coderabbitai ignore anywhere in the PR description to prevent this PR from being reviewed.
  • Add @coderabbitai summary to generate the high-level summary at a specific location in the PR description.
  • Add @coderabbitai anywhere in the PR title to generate the title automatically.

CodeRabbit Configuration File (.coderabbit.yaml)

  • You can programmatically configure CodeRabbit by adding a .coderabbit.yaml file to the root of your repository.
  • Please see the configuration documentation for more information.
  • If your editor has YAML language server enabled, you can add the path at the top of this file to enable auto-completion and validation: # yaml-language-server: $schema=https://coderabbit.ai/integrations/schema.v2.json

Documentation and Community

  • Visit our Documentation for detailed information on how to use CodeRabbit.
  • Join our Discord Community to get help, request features, and share feedback.
  • Follow us on X/Twitter for updates and announcements.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

♻️ Duplicate comments (2)
.github/workflows/export-nemo-fast-conformer-hybrid-transducer-ctc.yaml (2)

65-68: Same quoting issue as transducer workflow

Please apply the ${names[@]} → "${names[@]}" quoting fix here as well.


101-116: Replicate the idempotency / commit-message improvements here

The comments made for the transducer workflow’s publish step equally apply to this CTC workflow.

🧹 Nitpick comments (11)
scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-ctc-non-streaming.py (1)

90-94: Consider adding error handling for quantization step.

The dynamic quantization logic is correctly implemented and placed after metadata addition. However, consider adding error handling around the quantization step to provide better feedback if the quantization fails.

+    try:
         quantize_dynamic(
             model_input="./model.onnx",
             model_output="./model.int8.onnx",
             weight_type=QuantType.QUInt8,
         )
+        print("Quantized model saved to model.int8.onnx")
+    except Exception as e:
+        print(f"Warning: Quantization failed: {e}")
+        print("Continuing with original model only")
scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-transducer-non-streaming.py (1)

94-99: Consider adding error handling and improving variable naming.

The dynamic quantization logic is correctly implemented for all three model components. However, consider adding error handling and using a more descriptive variable name.

-    for m in ["encoder", "decoder", "joiner"]:
+    for model_name in ["encoder", "decoder", "joiner"]:
+        try:
             quantize_dynamic(
-                model_input=f"{m}.onnx",
-                model_output=f"{m}.int8.onnx",
+                model_input=f"{model_name}.onnx",
+                model_output=f"{model_name}.int8.onnx",
                 weight_type=QuantType.QUInt8,
             )
+            print(f"Quantized {model_name} model saved to {model_name}.int8.onnx")
+        except Exception as e:
+            print(f"Warning: Quantization failed for {model_name}: {e}")
scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-transducer.py (1)

126-131: Consider adding error handling and improving variable naming.

The dynamic quantization logic is correctly implemented for all three model components. However, consider adding error handling and using a more descriptive variable name.

-    for m in ["encoder", "decoder", "joiner"]:
+    for model_name in ["encoder", "decoder", "joiner"]:
+        try:
             quantize_dynamic(
-                model_input=f"{m}.onnx",
-                model_output=f"{m}.int8.onnx",
+                model_input=f"{model_name}.onnx",
+                model_output=f"{model_name}.int8.onnx",
                 weight_type=QuantType.QUInt8,
             )
+            print(f"Quantized {model_name} model saved to {model_name}.int8.onnx")
+        except Exception as e:
+            print(f"Warning: Quantization failed for {model_name}: {e}")
scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-transducer.sh (1)

26-31: Consider adding error handling for file operations.

The file organization logic is correct, but consider adding error handling to ensure robustness when moving files.

     cp -v tokens.txt $d/

+    # Check if int8 models exist before moving
+    if [ -f encoder.int8.onnx ] && [ -f decoder.int8.onnx ] && [ -f joiner.int8.onnx ]; then
         mv -v encoder.int8.onnx $d_int8/
         mv -v decoder.int8.onnx $d_int8/
         mv -v joiner.int8.onnx $d_int8/
         mv -v tokens.txt $d_int8/
+    else
+        echo "Warning: Some int8 models not found, skipping int8 directory setup"
+    fi
scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-ctc.sh (2)

17-19: Quote array expansions to silence ShellCheck SC2068

ShellCheck flags the un-quoted array expansions:

for m in ${ms[@]}; do        # SC2068
...
for m in ${ms[@]}; do        # SC2068

Use double-quotes to avoid word-splitting and future-proof the loop:

-for m in ${ms[@]}; do
+for m in "${ms[@]}"; do

Also applies to: 41-42


24-30: Streamline tokens.txt handling

tokens.txt is first copied to the FP dir (cp) and then moved to the int8 dir (mv).
A simpler, clearer flow is to copy into both targets and leave the working copy untouched:

-    cp -v tokens.txt $d/
-    ...
-    mv -v tokens.txt $d_int8/
+    cp -v tokens.txt $d/
+    cp -v tokens.txt $d_int8/

This avoids mixing cp/mv, keeps a local copy for diagnostics, and avoids re-creating the file on every loop iteration.

scripts/nemo/parakeet-tdt_ctc-0.6b-ja/run-ctc.sh (1)

1-4: Adopt set -euo pipefail for safer bash scripts

Replacing set -ex with the stricter variant catches unset variables (-u) and pipe errors (-o pipefail):

-set -ex
+set -euo pipefail
+set -x   # keep trace if desired

This small hardening step prevents subtle failures from passing silently.

.github/workflows/export-nemo-fast-conformer-hybrid-transducer-transducer.yaml (1)

65-68: Quote the ${models[@]} expansion inside the loop

Same SC2068 issue as in the shell scripts – quoting avoids accidental word-splitting if a future model name contains spaces:

-for m in ${models[@]}; do
+for m in "${models[@]}"; do
scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-ctc-non-streaming.sh (1)

171-187: Quote list variables inside the for w in … loops

Although the file names are plain ASCII, quoting is a good habit and prevents edge-cases:

-for w in en-english.wav de-german.wav ...; do
+for w in en-english.wav de-german.wav ...; do   # <<< quote "$w" below
...
-  --wav $data/$w
+  --wav "$data/$w"

This pattern repeats for the int8 loop; fix both.

scripts/nemo/parakeet-tdt_ctc-0.6b-ja/export-onnx-ctc.py (2)

74-74: Replace os.system with subprocess for better error handling.

Using os.system for file operations doesn't provide good error handling and is generally discouraged.

-    os.system("ls -lh *.onnx")
+    import subprocess
+    subprocess.run(["ls", "-lh", "*.onnx"], check=True)

Apply the same change to line 84.

Also applies to: 84-84


35-37: Consider parameterizing the hardcoded model name.

The model name is hardcoded, which limits reusability. Consider making it configurable through command-line arguments.

+import argparse
+
+def get_args():
+    parser = argparse.ArgumentParser()
+    parser.add_argument("--model", default="nvidia/parakeet-tdt_ctc-0.6b-ja")
+    return parser.parse_args()
+
 @torch.no_grad()
 def main():
+    args = get_args()
     asr_model = nemo_asr.models.ASRModel.from_pretrained(
-        model_name="nvidia/parakeet-tdt_ctc-0.6b-ja"
+        model_name=args.model
     )
📜 Review details

Configuration used: CodeRabbit UI
Review profile: CHILL
Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between f140577 and 242e8d3.

📒 Files selected for processing (19)
  • .github/workflows/export-nemo-fast-conformer-hybrid-transducer-ctc-non-streaming.yaml (2 hunks)
  • .github/workflows/export-nemo-fast-conformer-hybrid-transducer-ctc.yaml (2 hunks)
  • .github/workflows/export-nemo-fast-conformer-hybrid-transducer-transducer-non-streaming.yaml (2 hunks)
  • .github/workflows/export-nemo-fast-conformer-hybrid-transducer-transducer.yaml (2 hunks)
  • .github/workflows/export-nemo-parakeet-tdt.yaml (1 hunks)
  • scripts/apk/generate-vad-asr-apk-script.py (1 hunks)
  • scripts/nemo/fast-conformer-hybrid-transducer-ctc/README.md (1 hunks)
  • scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-ctc-non-streaming.py (2 hunks)
  • scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-ctc.py (2 hunks)
  • scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-transducer-non-streaming.py (2 hunks)
  • scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-transducer.py (2 hunks)
  • scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-ctc-non-streaming.sh (10 hunks)
  • scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-ctc.sh (1 hunks)
  • scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-transducer-non-streaming.sh (10 hunks)
  • scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-transducer.sh (2 hunks)
  • scripts/nemo/parakeet-tdt_ctc-0.6b-ja/export-onnx-ctc.py (1 hunks)
  • scripts/nemo/parakeet-tdt_ctc-0.6b-ja/run-ctc.sh (1 hunks)
  • scripts/nemo/parakeet-tdt_ctc-0.6b-ja/test-onnx-ctc-non-streaming.py (1 hunks)
  • sherpa-onnx/kotlin-api/OfflineRecognizer.kt (1 hunks)
🧰 Additional context used
🧬 Code Graph Analysis (1)
scripts/nemo/parakeet-tdt_ctc-0.6b-ja/export-onnx-ctc.py (2)
scripts/nemo/parakeet-tdt_ctc-0.6b-ja/test-onnx-ctc-non-streaming.py (1)
  • main (107-165)
scripts/spleeter/export_onnx.py (1)
  • export (47-74)
🪛 LanguageTool
scripts/nemo/fast-conformer-hybrid-transducer-ctc/README.md

[grammar] ~26-~26: Use correct spacing
Context: ...s/nemo/models/parakeet-tdt_ctc-110m - https://huggingface.co/nvidia/parakeet-tdt_ctc-0.6b-ja to sherpa-onnx.

(QB_NEW_EN_OTHER_ERROR_IDS_5)

🪛 markdownlint-cli2 (0.17.2)
scripts/nemo/fast-conformer-hybrid-transducer-ctc/README.md

26-26: Unordered list indentation
Expected: 0; Actual: 2

(MD007, ul-indent)


26-26: Bare URL used

(MD034, no-bare-urls)

🪛 Shellcheck (0.10.0)
scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-ctc.sh

[error] 41-41: Double quote array expansions to avoid re-splitting elements.

(SC2068)

🪛 actionlint (1.7.7)
.github/workflows/export-nemo-parakeet-tdt.yaml

54-54: shellcheck reported issue in this script: SC2068:error:4:10: Double quote array expansions to avoid re-splitting elements

(shellcheck)


54-54: shellcheck reported issue in this script: SC2086:info:5:47: Double quote to prevent globbing and word splitting

(shellcheck)


54-54: shellcheck reported issue in this script: SC2086:info:6:12: Double quote to prevent globbing and word splitting

(shellcheck)


54-54: shellcheck reported issue in this script: SC2086:info:6:23: Double quote to prevent globbing and word splitting

(shellcheck)

🔇 Additional comments (21)
scripts/nemo/parakeet-tdt_ctc-0.6b-ja/test-onnx-ctc-non-streaming.py (2)

2-169: LGTM! Well-structured test implementation.

The Python implementation is well-organized with proper error handling, feature extraction, and CTC decoding logic. The code follows good practices for ONNX model testing.


1-1: Remove the confusing script reference.

Line 1 contains a reference to another script file, but the rest of the file contains a complete Python implementation. This creates confusion about the file's purpose. If this is meant to be a standalone test script, remove the reference line.

Apply this diff to remove the problematic reference:

-../fast-conformer-hybrid-transducer-ctc/test-onnx-ctc-non-streaming.py

Likely an incorrect or invalid review comment.

.github/workflows/export-nemo-fast-conformer-hybrid-transducer-transducer-non-streaming.yaml (2)

64-68: LGTM! Consistent int8 model variant addition.

The int8 model variants are properly added to the publishing pipeline with clear naming conventions.


96-100: LGTM! Matching compression configuration.

The int8 model variants are consistently added to the compression step, maintaining alignment with the publishing configuration.

scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-ctc.py (2)

9-9: LGTM! Proper quantization imports.

The necessary imports for dynamic quantization are correctly added.


118-122: LGTM! Well-placed quantization step.

The dynamic quantization is properly implemented with appropriate weight type and positioned correctly in the workflow after metadata addition.

.github/workflows/export-nemo-fast-conformer-hybrid-transducer-ctc-non-streaming.yaml (2)

64-68: LGTM! Consistent int8 model integration.

The int8 model variants are properly integrated into the publishing workflow with consistent naming patterns.


97-101: LGTM! Aligned compression configuration.

The int8 model variants are consistently added to the compression step, maintaining proper alignment with the publishing workflow.

scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-ctc-non-streaming.py (1)

9-9: LGTM - Quantization import added correctly.

The import of QuantType and quantize_dynamic from onnxruntime.quantization is properly added to support dynamic quantization functionality.

scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-transducer-non-streaming.py (1)

9-9: LGTM - Quantization import added correctly.

The import of QuantType and quantize_dynamic from onnxruntime.quantization is properly added to support dynamic quantization functionality.

scripts/nemo/fast-conformer-hybrid-transducer-ctc/export-onnx-transducer.py (1)

9-9: LGTM - Quantization import added correctly.

The import of QuantType and quantize_dynamic from onnxruntime.quantization is properly added to support dynamic quantization functionality.

scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-transducer.sh (2)

20-22: LGTM - Directory structure for int8 models.

The addition of separate directories for original and int8 quantized models is well-organized and maintains clear separation between model variants.


52-58: LGTM - Testing both model variants.

The testing logic correctly validates both original and int8 quantized models, ensuring parity between the two variants.

sherpa-onnx/kotlin-api/OfflineRecognizer.kt (2)

605-613: LGTM - English Parakeet TDT CTC model configuration.

The new model configuration for the English Parakeet TDT CTC model is properly structured and follows the established pattern for NeMo CTC models.


615-623: LGTM - Japanese Parakeet TDT CTC model configuration.

The new model configuration for the Japanese Parakeet TDT CTC model is properly structured and follows the established pattern for NeMo CTC models.

scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-ctc-non-streaming.sh (1)

21-29: Duplicate mv tokens.txt likely unintended

tokens.txt is first copied to the FP directory (line 22) and then moved there again at line 28 – the second mv is redundant and can fail if the file was already moved elsewhere earlier in the script.

-cp -v tokens.txt $d/
-...
-mv -v tokens.txt $d/

Remove the second line (or switch both to cp) for clarity.

Likely an incorrect or invalid review comment.

scripts/nemo/fast-conformer-hybrid-transducer-ctc/run-transducer-non-streaming.sh (2)

21-33: Well-structured int8 model organization.

The systematic approach of creating separate directories for regular and int8 models with explicit file moves ensures proper organization and avoids potential conflicts.


154-172: Comprehensive testing coverage for int8 models.

The addition of int8 model testing maintains parity with the regular model testing, ensuring both variants are properly validated.

scripts/nemo/parakeet-tdt_ctc-0.6b-ja/export-onnx-ctc.py (1)

12-30: Metadata handling looks correct.

The function properly clears existing metadata before adding new entries, ensuring clean metadata state.

.github/workflows/export-nemo-parakeet-tdt.yaml (2)

64-95: Well-implemented publishing workflow with retry mechanism.

The Hugging Face publishing step includes proper authentication, retry logic, and Git LFS handling for large files.


97-105: Proper release asset handling.

The release step correctly uses file globbing and proper authentication for uploading to the sherpa-onnx repository.

- https://catalog.ngc.nvidia.com/orgs/nvidia/teams/nemo/models/stt_multilingual_fastconformer_hybrid_large_pc

- https://catalog.ngc.nvidia.com/orgs/nvidia/teams/nemo/models/parakeet-tdt_ctc-110m
- https://huggingface.co/nvidia/parakeet-tdt_ctc-0.6b-ja

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue

Fix markdown formatting issues.

The line has incorrect indentation and should format the URL properly according to markdown standards.

Apply this diff to fix the formatting:

-  - https://huggingface.co/nvidia/parakeet-tdt_ctc-0.6b-ja
+- https://huggingface.co/nvidia/parakeet-tdt_ctc-0.6b-ja

Committable suggestion skipped: line range outside the PR's diff.

🧰 Tools
🪛 LanguageTool

[grammar] ~26-~26: Use correct spacing
Context: ...s/nemo/models/parakeet-tdt_ctc-110m - https://huggingface.co/nvidia/parakeet-tdt_ctc-0.6b-ja to sherpa-onnx.

(QB_NEW_EN_OTHER_ERROR_IDS_5)

🪛 markdownlint-cli2 (0.17.2)

26-26: Unordered list indentation
Expected: 0; Actual: 2

(MD007, ul-indent)


26-26: Bare URL used

(MD034, no-bare-urls)

🤖 Prompt for AI Agents
In scripts/nemo/fast-conformer-hybrid-transducer-ctc/README.md at line 26, the
URL is not properly formatted in markdown due to incorrect indentation. Fix this
by removing any leading spaces before the URL and ensure it is either on its own
line or formatted as a proper markdown link so it renders correctly.

Comment on lines +101 to +116
for m in ${models[@]}; do
rm -rf huggingface
export GIT_LFS_SKIP_SMUDGE=1
export GIT_CLONE_PROTECTION_ACTIVE=false
git clone https://csukuangfj:$HF_TOKEN@huggingface.co/csukuangfj/$m huggingface
cp -av $m/* huggingface
cd huggingface
git lfs track "*.onnx"
git lfs track "*.wav"
git status
git add .
git status
git commit -m "first commit"
git push https://csukuangfj:$HF_TOKEN@huggingface.co/csukuangfj/$m main
cd ..
done

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Make the Hugging Face publish loop idempotent & cheaper

  1. git lfs track "*.onnx" and "*.wav" run on every iteration – do it once outside the loop or inside the repo after the first clone.
  2. The commit message "first commit" becomes inaccurate after the first push – consider a generic message ("update models") or include date +"%F" to avoid duplicate-headline warnings.
  3. Add set -euo pipefail at the top of the script block to fail fast on any publishing error.

No diff provided since it spans multiple script blocks.

🤖 Prompt for AI Agents
In
.github/workflows/export-nemo-fast-conformer-hybrid-transducer-transducer.yaml
around lines 101 to 116, the git lfs track commands for "*.onnx" and "*.wav" are
redundantly run inside the loop for each model; move these commands outside the
loop or run them only once after the initial clone to improve efficiency. Also,
update the commit message from the static "first commit" to a more generic
message like "update models" or include a timestamp using date +"%F" to prevent
duplicate commit message warnings. Finally, add set -euo pipefail at the top of
the script block to ensure the script fails immediately on any error during
publishing.

Comment thread scripts/apk/generate-vad-asr-apk-script.py Outdated
Comment on lines +54 to +61
run: |
models=(
sherpa-onnx-nemo-parakeet-tdt_ctc-0.6b-ja-35000-int8
)
for m in ${models[@]}; do
mv -v scripts/nemo/parakeet-tdt_ctc-0.6b-ja/$m .
tar cjfv $m.tar.bz2 $m
done

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🛠️ Refactor suggestion

Fix shell script quoting issues.

Shellcheck identifies several quoting issues that should be addressed for robustness.

          models=(
            sherpa-onnx-nemo-parakeet-tdt_ctc-0.6b-ja-35000-int8
          )
-          for m in ${models[@]}; do
-            mv -v scripts/nemo/parakeet-tdt_ctc-0.6b-ja/$m .
-            tar cjfv $m.tar.bz2 $m
+          for m in "${models[@]}"; do
+            mv -v "scripts/nemo/parakeet-tdt_ctc-0.6b-ja/$m" .
+            tar cjfv "$m.tar.bz2" "$m"
           done
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
run: |
models=(
sherpa-onnx-nemo-parakeet-tdt_ctc-0.6b-ja-35000-int8
)
for m in ${models[@]}; do
mv -v scripts/nemo/parakeet-tdt_ctc-0.6b-ja/$m .
tar cjfv $m.tar.bz2 $m
done
run: |
models=(
sherpa-onnx-nemo-parakeet-tdt_ctc-0.6b-ja-35000-int8
)
for m in "${models[@]}"; do
mv -v "scripts/nemo/parakeet-tdt_ctc-0.6b-ja/$m" .
tar cjfv "$m.tar.bz2" "$m"
done
🧰 Tools
🪛 actionlint (1.7.7)

54-54: shellcheck reported issue in this script: SC2068:error:4:10: Double quote array expansions to avoid re-splitting elements

(shellcheck)


54-54: shellcheck reported issue in this script: SC2086:info:5:47: Double quote to prevent globbing and word splitting

(shellcheck)


54-54: shellcheck reported issue in this script: SC2086:info:6:12: Double quote to prevent globbing and word splitting

(shellcheck)


54-54: shellcheck reported issue in this script: SC2086:info:6:23: Double quote to prevent globbing and word splitting

(shellcheck)

🤖 Prompt for AI Agents
In .github/workflows/export-nemo-parakeet-tdt-tdt.yaml around lines 54 to 61,
the shell script has quoting issues that can cause word splitting or globbing
problems. Fix this by quoting the array expansion as "${models[@]}" in the for
loop and quoting variable expansions like "$m" in the mv and tar commands to
ensure proper handling of filenames and prevent errors.

@csukuangfj
csukuangfj requested a review from Copilot July 9, 2025 07:56

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull Request Overview

Refactors and extends model export support to include new NeMo Parakeet TDT int8 variants for English and Japanese, updating the Kotlin API, export scripts, test runners, and CI workflows.

  • Added support for two new int8 model types in OfflineRecognizer.kt.
  • Enhanced Python export scripts to perform dynamic quantization and metadata injection.
  • Updated shell scripts and GitHub workflows to package, test, and publish int8 model artifacts.

Reviewed Changes

Copilot reviewed 19 out of 19 changed files in this pull request and generated 1 comment.

Show a summary per file
File Description
sherpa-onnx/kotlin-api/OfflineRecognizer.kt Added cases 33/34 for new int8 Parakeet TDT ASR models
scripts/nemo/parakeet-tdt_ctc-0.6b-ja/test-onnx-ctc-non-streaming.py Introduced a placeholder script referencing the test runner
scripts/nemo/parakeet-tdt_ctc-0.6b-ja/export-onnx-ctc.py Imported quantize_dynamic, added int8 quantization step
scripts/nemo/fast-conformer-hybrid-transducer-ctc/*.sh Extended CTC/transducer run scripts to handle int8 models
scripts/nemo/fast-conformer-hybrid-transducer-ctc/*.py Imported quantization, looped over encoder/decoder/joiner
.github/workflows/*.yaml Updated workflows to package and publish int8 variants
Comments suppressed due to low confidence (2)

sherpa-onnx/kotlin-api/OfflineRecognizer.kt:616

  • The directory name for the Japanese int8 model uses a different separator pattern (parakeet-tdt_ctc) compared to the English int8 model (parakeet_tdt_ctc), which may lead to inconsistency or path resolution errors. Consider aligning the naming convention across modelDir values.
            val modelDir = "sherpa-onnx-nemo-parakeet-tdt_ctc-0.6b-ja-35000-int8"

sherpa-onnx/kotlin-api/OfflineRecognizer.kt:605

  • [nitpick] New model type indices (33 and 34) have been added, but there is no corresponding unit test to verify that getOfflineModelConfig returns the expected OfflineModelConfig for these indices. Consider adding tests to cover these new cases.
        33 -> {

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants