Skip to content

Upload DPDFNet models - #3322

Merged
csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:upload-models
Mar 16, 2026
Merged

csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:upload-models

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Mar 16, 2026 •

Copy link
Copy Markdown
Collaborator

See also
https://github.com/ceva-ip/DPDFNet

You can download models from
https://github.com/k2-fsa/sherpa-onnx/releases/tag/speech-enhancement-models

Screenshot 2026-03-16 at 15 10 11

See also #3276

cc @danielr-ceva

16 kHz models

Model Params [M] MACs [G] TFLite Size [MB] ONNX Size [MB] Intended Use
dpdfnet_baseline 2.31 0.36 8.5 8.5 Fastest / lowest resource usage
dpdfnet2 2.49 1.35 10.7 9.9 Real-time / embedded devices
dpdfnet4 2.84 2.36 12.9 11.2 Balanced performance
dpdfnet8 3.54 4.37 17.2 14.1 Best enhancement quality

48 kHz model

Model Params [M] MACs [G] TFLite Size [MB] ONNX Size [MB] Intended Use
dpdfnet2_48khz_hr 2.58 2.42 11.6 10.3 High-resolution 48 kHz audio

Summary by CodeRabbit

  • Chores
    • Enhanced the model upload workflow to automatically download and prepare multiple ONNX model variants for future distribution
    • Introduced ONNX model upload functionality to the release pipeline, currently available for activation when needed
    • Maintained compatibility with the existing tarball-based release upload process, ensuring no disruption to current workflows

@dosubot dosubot Bot added the size:S This PR changes 10-29 lines, ignoring generated files. label Mar 16, 2026
@coderabbitai

coderabbitai Bot commented Mar 16, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Added a new workflow step to download DPDFNet ONNX models from Hugging Face, rename baseline files, and introduced a gated (disabled) Release block for uploading ONNX files to the sherpa-onnx repository while preserving the existing tarball upload workflow.

Changes

Cohort / File(s) Summary
DPDFNet Model Upload Workflow
.github/workflows/upload-models.yaml
Added "Upload DPDFNet" step to download five ONNX model files from Hugging Face and rename baseline.onnx; added disabled Release block for ONNX file uploads to k2-fsa/sherpa-onnx repository with speech-enhancement-models tag.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

  • #2656: Modifies .github/workflows/upload-models.yaml to add model packaging and upload steps for different model sets
  • #2890: Modifies upload-models workflow to add ONNX model upload steps and conditional Release gating logic
  • #3220: Modifies .github/workflows/upload-models.yaml to add gated ONNX and model upload release steps

Suggested labels

size:M

Poem

🐰 Hop, skip, and a bound,
DPDFNet models found!
From Hugging Face they leap with grace,
A workflow to upload and embrace,
With gates of false, we wait and pause,
For workflows to enhance the cause! ✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'Upload DPDFNet models' is directly related to the main change in the pull request, which adds workflow steps to upload DPDFNet ONNX models to a GitHub release.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (2)
.github/workflows/upload-models.yaml (2)

1163-1172: Remove or document the disabled Release block.

Static analysis flagged this if: false condition. If this block is for future use, consider either:

  1. Removing it until needed (can be added back via PR when required)
  2. Adding a comment explaining why it's kept disabled

Additionally, having two steps named "Release" will cause confusion in workflow logs.

♻️ Option: Remove the disabled block entirely
-
-      - name: Release
-        if: false
-        uses: svenstaro/upload-release-action@v2
-        with:
-          file_glob: true
-          file: ./*.onnx
-          overwrite: true
-          repo_name: k2-fsa/sherpa-onnx
-          repo_token: ${{ secrets.UPLOAD_GH_SHERPA_ONNX_TOKEN }}
-          tag: speech-enhancement-models
♻️ Option: Document and rename if keeping for reference
+      # TODO: Enable this step to upload ONNX models to speech-enhancement-models release
+      # when ready for production use
-      - name: Release
+      - name: Release ONNX models (disabled)
         if: false
         uses: svenstaro/upload-release-action@v2
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/workflows/upload-models.yaml around lines 1163 - 1172, The workflow
contains a disabled release step named "Release" using
svenstaro/upload-release-action@v2 with if: false and tag:
speech-enhancement-models; either remove this entire block to avoid dead code
and duplicate step names, or document why it is kept disabled (add an inline
comment near if: false) and rename the step to a unique identifier (e.g.,
"Release (disabled - reference)") to avoid confusion in logs while preserving
the action and inputs for future use.

27-42: Add error handling to wget downloads.

The wget commands will not fail the workflow on HTTP errors (e.g., 404, 500) without explicit flags. If a model file is unavailable, wget may save an HTML error page instead, leading to corrupted uploads or confusing downstream failures.

Consider adding set -e and using wget --fail or wget -q --show-progress with explicit exit code checks.

♻️ Proposed fix for robustness
       - name: Upload DPDFNet
         shell: bash
         run: |
+          set -e

           models=(
             baseline.onnx
             dpdfnet2.onnx
             dpdfnet2_48khz_hr.onnx
             dpdfnet4.onnx
             dpdfnet8.onnx
           )
           for m in ${models[@]}; do
-            wget https://huggingface.co/Ceva-IP/DPDFNet/resolve/main/onnx/$m
+            wget --retry-connrefused --waitretry=1 --tries=3 -q --show-progress \
+              "https://huggingface.co/Ceva-IP/DPDFNet/resolve/main/onnx/$m" || exit 1
           done

           mv baseline.onnx dpdfnet_baseline.onnx
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/workflows/upload-models.yaml around lines 27 - 42, The current
download loop using the models array and wget in the upload workflow can
silently succeed with HTML error pages; update the script to enable strict
failure and make wget fail on HTTP errors by adding a failing shell option
(e.g., set -e) at the top of the run block and using wget --fail (or -q
--show-progress --fail) inside the for m in ${models[@]} loop, and after each
download check the exit status (or rely on set -e) so any missing file causes
the job to fail early; ensure mv baseline.onnx dpdfnet_baseline.onnx only runs
after successful download of baseline.onnx.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In @.github/workflows/upload-models.yaml:
- Around line 1163-1172: The workflow contains a disabled release step named
"Release" using svenstaro/upload-release-action@v2 with if: false and tag:
speech-enhancement-models; either remove this entire block to avoid dead code
and duplicate step names, or document why it is kept disabled (add an inline
comment near if: false) and rename the step to a unique identifier (e.g.,
"Release (disabled - reference)") to avoid confusion in logs while preserving
the action and inputs for future use.
- Around line 27-42: The current download loop using the models array and wget
in the upload workflow can silently succeed with HTML error pages; update the
script to enable strict failure and make wget fail on HTTP errors by adding a
failing shell option (e.g., set -e) at the top of the run block and using wget
--fail (or -q --show-progress --fail) inside the for m in ${models[@]} loop, and
after each download check the exit status (or rely on set -e) so any missing
file causes the job to fail early; ensure mv baseline.onnx dpdfnet_baseline.onnx
only runs after successful download of baseline.onnx.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 194aebbc-fe68-4cb8-83de-402217c12733

📥 Commits

Reviewing files that changed from the base of the PR and between 807c496 and bbc09de.

📒 Files selected for processing (1)
  • .github/workflows/upload-models.yaml

@csukuangfj
csukuangfj merged commit 7bc23b8 into k2-fsa:master Mar 16, 2026
1 check passed
@csukuangfj
csukuangfj deleted the upload-models branch March 16, 2026 07:16
@danielr-ceva

Copy link
Copy Markdown
Contributor

@csukuangfj Thanks for the quick support!

I have one more question: when I click on Speech Enhancement, it takes me to
https://k2-fsa.github.io/sherpa/onnx/speech-enhancement/index.html - but the DPDFNet model
doesn't appear there (only the GTCRN appears at the moment). It only shows up on the GitHub releases page:
https://github.com/k2-fsa/sherpa-onnx/releases/tag/speech-enhancement-models

Is there anything I need to do to get it listed on the documentation page as well,
or is that something you'll be updating in the future?

Thanks a lot!

@csukuangfj

Copy link
Copy Markdown
Collaborator Author
  1. Visit https://k2-fsa.github.io/sherpa/onnx/speech-enhancement/index.html -
  2. Click "Edit on GitHub"
  3. Please create a PR to add the doc for the model

@danielr-ceva

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S This PR changes 10-29 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants