Skip to content

Upload matcha tts zh-en model - #2865

Merged
csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:matcha-zh-en
Dec 5, 2025
Merged

csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:matcha-zh-en

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Dec 5, 2025 •

Copy link
Copy Markdown
Collaborator

You can find the doc about this model at https://k2-fsa.github.io/sherpa/onnx/tts/all/Chinese-English/matcha-icefall-zh-en.html

Screenshot 2025-12-05 at 11 44 07

Summary by CodeRabbit

  • New Features

    • GitHub Actions workflow for automated model export and packaging to ONNX format with artifact publishing to HuggingFace
  • Documentation

    • Minor README formatting improvements for download links and presentation
  • Refactor

    • Internal token handling and lexicon generation updates

✏️ Tip: You can customize this high-level summary in your review settings.

@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Dec 5, 2025
@coderabbitai

coderabbitai Bot commented Dec 5, 2025 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

Walkthrough

This pull request introduces a GitHub Actions CI/CD workflow that automates the export and packaging of the Matcha Chinese-English TTS model. It includes sample generation using sherpa_onnx, refactors token handling in lexicon generation, and publishes processed artifacts to HuggingFace and GitHub releases.

Changes

Cohort / File(s) Summary
CI/CD Automation
.github/workflows/export-matcha-zh-en.yaml
Introduces GitHub Actions workflow triggered on push to matcha-zh-en branch or manual dispatch. Runs Ubuntu job with Python 3.10 to install dependencies, generate TTS samples, download ONNX models, process tokens/lexicon, create tar.bz2 archive, publish artifacts to HuggingFace with LFS tracking, and conditionally release to GitHub for specified owners.
TTS Sample Generation
scripts/matcha-tts/zh-en/generate_samples.py
New script that configures sherpa_onnx OfflineTts with ONNX models, lexicon, and token paths; generates Chinese speech from text; and saves output to MP3. Includes config validation and error handling.
Lexicon Token Processing
scripts/matcha-tts/zh-en/generate_lexicon.py
Refactors token handling from flat string list to list of single-element lists. Updates "shei2" → "shui2" check and tone digit appending to work with nested token structure (token[0]). Adjusts flattening logic to extract first element from sublists before writing output.
Documentation
scripts/matcha-tts/zh-en/README.md
Minor formatting: removes trailing whitespace and restructures download URL presentation for improved readability.

Sequence Diagram

sequenceDiagram
    participant GHA as GitHub Actions
    participant Runner as Ubuntu Runner
    participant HF as HuggingFace API
    participant GHRel as GitHub Releases
    participant Scripts as Python Scripts

    GHA->>Runner: Trigger on matcha-zh-en push
    Runner->>Scripts: Checkout & Setup Python 3.10
    Scripts->>Scripts: Install dependencies (numpy, sherpa-onnx, etc.)
    Scripts->>Scripts: Run generate_samples.py
    Scripts->>Scripts: Download ONNX model components
    Scripts->>Scripts: Run token generation & lexicon scripts
    Scripts->>Scripts: Create tar.bz2 archive
    Scripts->>HF: Publish artifacts with LFS tracking
    HF-->>Scripts: Confirm upload
    Scripts->>GHRel: Upload release (for owners csukuangfj, k2-fsa)
    GHRel-->>Scripts: Release created
    Scripts->>Runner: Cleanup & Complete
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

  • Workflow file complexity: Multiple conditional steps, artifact handling, and HuggingFace/GitHub API interactions require careful validation
  • Token refactoring logic: The nested list structure changes in generate_lexicon.py need verification of token indexing and flattening correctness across all affected lines
  • New script soundness: generate_samples.py error handling and sherpa_onnx configuration should be verified against API expectations

Possibly related PRs

Suggested labels

size:M

Poem

🐰 Hops through workflows, fast as can be,
Matcha speech flows—zh-EN harmony!
Tokens restack, lexicons align,
Samples sing out—automation divine!
To HuggingFace and releases they go, ✨🎵

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

📜 Recent review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between ab479d8 and 56e1aaf.

📒 Files selected for processing (4)
  • .github/workflows/export-matcha-zh-en.yaml (1 hunks)
  • scripts/matcha-tts/zh-en/README.md (1 hunks)
  • scripts/matcha-tts/zh-en/generate_lexicon.py (1 hunks)
  • scripts/matcha-tts/zh-en/generate_samples.py (1 hunks)

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello @csukuangfj, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request integrates the Matcha TTS Chinese-English model, enhancing the sherpa-onnx project's text-to-speech capabilities. It includes essential updates to the lexicon generation script for improved accuracy, provides a new utility for generating audio samples, and updates documentation to offer alternative download options for model dependencies.

Highlights

  • New Model Support: The pull request introduces the Matcha TTS Chinese-English model, enabling text-to-speech synthesis for both languages within the project.
  • Lexicon Generation Fix: A bug in generate_lexicon.py was corrected to properly handle pinyin token modifications and tone assignments, ensuring accurate lexicon generation for the model.
  • Sample Generation Script: A new script, generate_samples.py, has been added to provide an example of how to use the sherpa-onnx library to generate audio samples with the new Matcha TTS model.
  • Documentation Update: The README.md file has been updated to include an additional download source for the required vocos-16khz-univ.onnx vocoder model, offering more flexibility for users.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/export-matcha-zh-en.yaml
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@csukuangfj
csukuangfj merged commit e6a6599 into k2-fsa:master Dec 5, 2025
1 check was pending
@csukuangfj
csukuangfj deleted the matcha-zh-en branch December 5, 2025 03:45

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a new Matcha TTS Chinese-English model. The README.md has been updated to include an alternative download link for the vocoder, which is a helpful addition. The generate_lexicon.py script includes critical fixes to correctly process the output of the pypinyin library, ensuring accurate lexicon generation. A new script, generate_samples.py, has been added to demonstrate sample generation, but it contains several hardcoded values that could be made configurable for better maintainability and flexibility.

tokens="matcha-icefall-zh-en/tokens.txt",
data_dir="matcha-icefall-zh-en/espeak-ng-data",
),
num_threads=2,

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The num_threads parameter is hardcoded to 2. While this might be sufficient for sample generation, it's generally good practice to make such configuration parameters adjustable, perhaps via command-line arguments, to allow for flexibility in different environments or for performance tuning.

Comment on lines +23 to +24
max_num_sentences=1,
rule_fsts="./matcha-icefall-zh-en/phone-zh.fst,./matcha-icefall-zh-en/date-zh.fst,./matcha-icefall-zh-en/number-zh.fst",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The rule_fsts string is quite long and can impact readability. Consider breaking it into multiple lines or constructing it from a list of strings for better clarity.

    rule_fsts=(
        "./matcha-icefall-zh-en/phone-zh.fst,"
        "./matcha-icefall-zh-en/date-zh.fst,"
        "./matcha-icefall-zh-en/number-zh.fst"
    ),

raise ValueError("Please check your config")

tts = sherpa_onnx.OfflineTts(config)
text = "我最近在学习machine learning,希望能够在未来的artificial intelligence领域有所建树。在这次vocation中,我们计划去Paris欣赏埃菲尔铁塔和卢浮宫的美景。某某银行的副行长和一些行政领导表示,他们去过长江和长白山; 经济不断增长。开始数字测试。2025年12月4号,拨打110或者189202512043。123456块钱。在这个快速发展的时代,人工智能技术正在改变我们的生活方式。语音合成作为人工智能的重要应用之一,让机器能够用自然流畅的语音与人类进行交流。"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The text variable contains a very long string literal. For improved readability and easier modification, especially with long sentences, it's better to use Python's triple-quoted string literals to split it across multiple lines.

text = (
    "我最近在学习machine learning,希望能够在未来的artificial intelligence领域有所建树。"
    "在这次vocation中,我们计划去Paris欣赏埃菲尔铁塔和卢浮宫的美景。"
    "某某银行的副行长和一些行政领导表示,他们去过长江和长白山; 经济不断增长。"
    "开始数字测试。2025年12月4号,拨打110或者189202512043。123456块钱。"
    "在这个快速发展的时代,人工智能技术正在改变我们的生活方式。"
    "语音合成作为人工智能的重要应用之一,让机器能够用自然流畅的语音与人类进行交流。"
)

Comment on lines +37 to +40
"./hf/matcha/icefall-zh-en/mp3/0.mp3",
audio.samples,
samplerate=audio.sample_rate,
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The output path for the generated audio file is hardcoded. To make this script more flexible and reusable, consider allowing the output path to be specified as a command-line argument.

@ZhangWei125521

Copy link
Copy Markdown

@csukuangfj 非常感谢你的提交,但是碰到个问题,用该tts生成时,会碰到一些说话不清楚的地方, 比如“好的,请将草稿放置在下方扫描区域”中的草稿很大概率会出现问题,比如稿字读不出来,请问该如何规避这种问题

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants