Repository navigation
Upload matcha tts zh-en model - #2865
Conversation
|
Caution Review failedThe pull request is closed. WalkthroughThis pull request introduces a GitHub Actions CI/CD workflow that automates the export and packaging of the Matcha Chinese-English TTS model. It includes sample generation using sherpa_onnx, refactors token handling in lexicon generation, and publishes processed artifacts to HuggingFace and GitHub releases. Changes
Sequence DiagramsequenceDiagram
participant GHA as GitHub Actions
participant Runner as Ubuntu Runner
participant HF as HuggingFace API
participant GHRel as GitHub Releases
participant Scripts as Python Scripts
GHA->>Runner: Trigger on matcha-zh-en push
Runner->>Scripts: Checkout & Setup Python 3.10
Scripts->>Scripts: Install dependencies (numpy, sherpa-onnx, etc.)
Scripts->>Scripts: Run generate_samples.py
Scripts->>Scripts: Download ONNX model components
Scripts->>Scripts: Run token generation & lexicon scripts
Scripts->>Scripts: Create tar.bz2 archive
Scripts->>HF: Publish artifacts with LFS tracking
HF-->>Scripts: Confirm upload
Scripts->>GHRel: Upload release (for owners csukuangfj, k2-fsa)
GHRel-->>Scripts: Release created
Scripts->>Runner: Cleanup & Complete
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes
Possibly related PRs
Suggested labels
Poem
✨ Finishing touches
🧪 Generate unit tests (beta)
📜 Recent review detailsConfiguration used: CodeRabbit UI Review profile: CHILL Plan: Pro 📒 Files selected for processing (4)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Summary of ChangesHello @csukuangfj, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request integrates the Matcha TTS Chinese-English model, enhancing the Highlights
Ignored Files
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request introduces a new Matcha TTS Chinese-English model. The README.md has been updated to include an alternative download link for the vocoder, which is a helpful addition. The generate_lexicon.py script includes critical fixes to correctly process the output of the pypinyin library, ensuring accurate lexicon generation. A new script, generate_samples.py, has been added to demonstrate sample generation, but it contains several hardcoded values that could be made configurable for better maintainability and flexibility.
| tokens="matcha-icefall-zh-en/tokens.txt", | ||
| data_dir="matcha-icefall-zh-en/espeak-ng-data", | ||
| ), | ||
| num_threads=2, |
There was a problem hiding this comment.
| max_num_sentences=1, | ||
| rule_fsts="./matcha-icefall-zh-en/phone-zh.fst,./matcha-icefall-zh-en/date-zh.fst,./matcha-icefall-zh-en/number-zh.fst", |
There was a problem hiding this comment.
| raise ValueError("Please check your config") | ||
|
|
||
| tts = sherpa_onnx.OfflineTts(config) | ||
| text = "我最近在学习machine learning,希望能够在未来的artificial intelligence领域有所建树。在这次vocation中,我们计划去Paris欣赏埃菲尔铁塔和卢浮宫的美景。某某银行的副行长和一些行政领导表示,他们去过长江和长白山; 经济不断增长。开始数字测试。2025年12月4号,拨打110或者189202512043。123456块钱。在这个快速发展的时代,人工智能技术正在改变我们的生活方式。语音合成作为人工智能的重要应用之一,让机器能够用自然流畅的语音与人类进行交流。" |
There was a problem hiding this comment.
The text variable contains a very long string literal. For improved readability and easier modification, especially with long sentences, it's better to use Python's triple-quoted string literals to split it across multiple lines.
text = (
"我最近在学习machine learning,希望能够在未来的artificial intelligence领域有所建树。"
"在这次vocation中,我们计划去Paris欣赏埃菲尔铁塔和卢浮宫的美景。"
"某某银行的副行长和一些行政领导表示,他们去过长江和长白山; 经济不断增长。"
"开始数字测试。2025年12月4号,拨打110或者189202512043。123456块钱。"
"在这个快速发展的时代,人工智能技术正在改变我们的生活方式。"
"语音合成作为人工智能的重要应用之一,让机器能够用自然流畅的语音与人类进行交流。"
)| "./hf/matcha/icefall-zh-en/mp3/0.mp3", | ||
| audio.samples, | ||
| samplerate=audio.sample_rate, | ||
| ) |
|
@csukuangfj 非常感谢你的提交,但是碰到个问题,用该tts生成时,会碰到一些说话不清楚的地方, 比如“好的,请将草稿放置在下方扫描区域”中的草稿很大概率会出现问题,比如稿字读不出来,请问该如何规避这种问题 |
You can find the doc about this model at https://k2-fsa.github.io/sherpa/onnx/tts/all/Chinese-English/matcha-icefall-zh-en.html
Summary by CodeRabbit
New Features
Documentation
Refactor
✏️ Tip: You can customize this high-level summary in your review settings.