Skip to content

Upload models for https://huggingface.co/CohereLabs/cohere-transcribe-03-2026 - #3453

Merged
csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:upload-models
Apr 1, 2026
Merged

csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:upload-models

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Apr 1, 2026 •

Copy link
Copy Markdown
Collaborator

See also #3442

Download it from

sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01.tar.bz2

Summary by CodeRabbit

  • New Features

    • Added sherpa-onnx-cohere-transcribe-14-lang-int8 model to published releases with multilingual support (Arabic, German, English, Spanish, French, Japanese, Korean, Vietnamese, Chinese) and corresponding test audio files.
  • Chores

    • Updated model publishing workflow configuration and disabled publishing to additional repository sources.

@dosubot dosubot Bot added the size:M This PR changes 30-99 lines, ignoring generated files. label Apr 1, 2026
@gemini-code-assist

Copy link
Copy Markdown

Note

Gemini is unable to generate a review for this pull request due to the file types involved not being currently supported.

@coderabbitai

coderabbitai Bot commented Apr 1, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

Updated .github/workflows/upload-models.yaml to change the trigger branch, add a new step for packaging the sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01 model with test WAV files into a tar.bz2 archive, disable FunASR Nano collection and modelscope publishing, and add the Cohere model to Hugging Face publishing.

Changes

Cohort / File(s) Summary
GitHub Actions Workflow
.github/workflows/upload-models.yaml
Updated trigger branch from upload-models-2 to upload-models; added new step to download and package Cohere transcription model artifacts with multilingual test WAV files (ar, de, en, es, fr, ja, ko, vi, zh) into tar.bz2 archive; disabled "Collect funasr-nano with LLM int8" step; disabled "Publish to modelscope" job; added Cohere model to Hugging Face publishing list.

Estimated code review effort

🎯 2 (Simple) | ⏱️ ~10 minutes

Possibly related PRs

Suggested labels

size:S

Poem

🐰 A new model hops into place,
Cohere's transcription, multi-language grace!
Test WAVs bundled, languages packed tight,
Tar archives ready, uploaded just right! ✨

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title directly references the specific Hugging Face model being uploaded and clearly indicates the primary purpose of the changes.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In @.github/workflows/upload-models.yaml:
- Around line 27-52: Remove the redundant "if: true" from the "Download
cohere-transcribe" GitHub Actions step and verify/add download of tokens.txt:
locate the step titled "Download cohere-transcribe" in the workflow, delete the
if: true line, then check the model repo for a tokens.txt file and if present
add a curl -SL -O call to fetch tokens.txt alongside README.md,
decoder.int8.onnx, encoder.int8.onnx.data and encoder.int8.onnx so the step
mirrors other ASR model downloads.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 30ee06be-eee5-4e6f-96f3-54d475b5754a

📥 Commits

Reviewing files that changed from the base of the PR and between c94cdda and 621cd23.

📒 Files selected for processing (1)
  • .github/workflows/upload-models.yaml

Comment on lines +27 to +52
- name: Download cohere-transcribe
if: true
shell: bash
run: |
d=sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01
mkdir -p $d
pushd $d

mkdir ./test_wavs
cd ./test_wavs
for w in ar de en es fr ja ko vi zh; do
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/test_wavs/$w.wav
done
ls -lh
cd ..

curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/README.md
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/decoder.int8.onnx
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/encoder.int8.onnx.data
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/encoder.int8.onnx

ls -lh
popd

tar cjvf $d.tar.bz2 $d
ls -lh $d $d.tar.bz2

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Check if tokens.txt exists in the ModelScope repository
curl -sI "https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/tokens.txt" | head -20

Repository: k2-fsa/sherpa-onnx

Length of output: 954


🏁 Script executed:

#!/bin/bash
# Search for cohere-transcribe references in the repository
rg "cohere-transcribe" -i --max-count=20

Repository: k2-fsa/sherpa-onnx

Length of output: 44


🏁 Script executed:

#!/bin/bash
# Look for tokens.txt references in the workflow and tests
rg "tokens\.txt" --max-count=20 -A 2 -B 2

Repository: k2-fsa/sherpa-onnx

Length of output: 50375


🏁 Script executed:

#!/bin/bash
# Check the upload-models workflow file structure around this section
head -60 .github/workflows/upload-models.yaml | tail -35

Repository: k2-fsa/sherpa-onnx

Length of output: 1515


Remove redundant if: true condition.

Line 28 has if: true which is redundant—the step runs unconditionally by default and will be flagged by actionlint.

Additionally, this step downloads only README.md, decoder.int8.onnx, encoder.int8.onnx.data, and encoder.int8.onnx, but does not download a tokens.txt file. All other ASR models in this workflow (FireRedASR2, Whisper, SenseVoice, Moonshine, Paraformer, etc.) include tokens.txt in their downloads. Please verify whether tokens.txt is available and required for the cohere-transcribe model.

Suggested fix for redundant condition
       - name: Download cohere-transcribe
-        if: true
         shell: bash
         run: |
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
- name: Download cohere-transcribe
if: true
shell: bash
run: |
d=sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01
mkdir -p $d
pushd $d
mkdir ./test_wavs
cd ./test_wavs
for w in ar de en es fr ja ko vi zh; do
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/test_wavs/$w.wav
done
ls -lh
cd ..
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/README.md
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/decoder.int8.onnx
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/encoder.int8.onnx.data
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/encoder.int8.onnx
ls -lh
popd
tar cjvf $d.tar.bz2 $d
ls -lh $d $d.tar.bz2
- name: Download cohere-transcribe
shell: bash
run: |
d=sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01
mkdir -p $d
pushd $d
mkdir ./test_wavs
cd ./test_wavs
for w in ar de en es fr ja ko vi zh; do
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/test_wavs/$w.wav
done
ls -lh
cd ..
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/README.md
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/decoder.int8.onnx
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/encoder.int8.onnx.data
curl -SL -O https://modelscope.cn/models/csukuangfj/sherpa-onnx-cohere-transcribe-14-lang-int8-2026-04-01/resolve/master/encoder.int8.onnx
ls -lh
popd
tar cjvf $d.tar.bz2 $d
ls -lh $d $d.tar.bz2
🧰 Tools
🪛 actionlint (1.7.11)

[error] 28-28: constant expression "true" in condition. remove the if: section

(if-cond)

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In @.github/workflows/upload-models.yaml around lines 27 - 52, Remove the
redundant "if: true" from the "Download cohere-transcribe" GitHub Actions step
and verify/add download of tokens.txt: locate the step titled "Download
cohere-transcribe" in the workflow, delete the if: true line, then check the
model repo for a tokens.txt file and if present add a curl -SL -O call to fetch
tokens.txt alongside README.md, decoder.int8.onnx, encoder.int8.onnx.data and
encoder.int8.onnx so the step mirrors other ASR model downloads.

@csukuangfj
csukuangfj merged commit aeeb891 into k2-fsa:master Apr 1, 2026
1 check passed
@csukuangfj
csukuangfj deleted the upload-models branch April 1, 2026 11:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M This PR changes 30-99 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant