Skip to content

fix matcha tts zh-en model - #2851

Merged
csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:fix-matcha-tts-zh-en
Dec 3, 2025
Merged

csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:fix-matcha-tts-zh-en

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Dec 3, 2025 •

Copy link
Copy Markdown
Collaborator

Please regenerate the lexicon.txt

or download the lexicon.txt below.

lexicon.txt

Fixes #2847

CC @lkocok


Test script

test_onnx_20251203.py

Summary by CodeRabbit

  • Bug Fixes

    • Improved Chinese character tokenization and tone handling in text-to-speech processing.
    • Removed unintended space token insertion in lexicon generation.
  • Chores

    • Optimized debug logging to activate only when enabled.

✏️ Tip: You can customize this high-level summary in your review settings.

@dosubot dosubot Bot added the size:S This PR changes 10-29 lines, ignoring generated files. label Dec 3, 2025
@coderabbitai

coderabbitai Bot commented Dec 3, 2025 •

Copy link
Copy Markdown

Note

Other AI code review bot(s) detected

CodeRabbit has detected other AI code review bot(s) in this pull request and will avoid duplicating their findings in the review comments. This may lead to a less comprehensive review.

Walkthrough

This PR modifies lexicon generation and TTS processing for Matcha TTS. Changes include switching from lazy_pinyin to pinyin with adjusted tone handling, removing automatic space insertion before punctuation/alphabetic characters, and making debug logging conditional on a configuration flag.

Changes

Cohort / File(s) Summary
Lexicon generation script
scripts/matcha-tts/zh-en/generate_lexicon.py
Replaces lazy_pinyin with pinyin, updates tone handling from tone_sandhi to neutral_tone_with_five, and adjusts result extraction from nested lists (e.g., [0][0] for characters, flattened lists for phrases).
C++ lexicon tokenization
sherpa-onnx/csrc/matcha-tts-lexicon.cc
Removes conditional block that appended space token when input word started with alphabetic or punctuation characters.
Matcha TTS inference
sherpa-onnx/csrc/offline-tts-matcha-impl.h
Makes token aggregation into vector x and debug logging conditional on config_.model.debug; when debug is disabled, x remains empty, affecting downstream tensor creation and inference.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

  • scripts/matcha-tts/zh-en/generate_lexicon.py: Verify pinyin library API changes and nested list flattening logic are correct across both character and phrase processing paths.
  • sherpa-onnx/csrc/matcha-tts-lexicon.cc: Confirm removal of space token insertion does not reintroduce the overlapping audio artifact and validate tokenization output for punctuation-prefixed words.
  • sherpa-onnx/csrc/offline-tts-matcha-impl.h: Ensure conditional debug logic doesn't inadvertently disable model inference when debug flag is off; verify that empty x vector is handled safely downstream.

Possibly related PRs

Suggested labels

size:M

Poem

🐰 Whiskers twitching at punctuation's pace,
We've smoothed the overlaps, erased the trace,
No extra spaces where they shouldn't be,
Tone handling now flows naturally,
Speech blends seamlessly, debug can rest,
The lexicon's dance is truly blessed!

Pre-merge checks and finishing touches

❌ Failed checks (1 warning, 2 inconclusive)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. You can run @coderabbitai generate docstrings to improve docstring coverage.
Title check ❓ Inconclusive The title 'fix matcha tts zh-en model' is vague and generic without describing the specific nature of the fix. Use a more specific title describing the actual fix, e.g., 'Fix overlapping audio at punctuation in matcha TTS zh-en model' or 'Fix matcha TTS zh-en lexicon tokenization issue'.
Out of Scope Changes check ❓ Inconclusive The PR includes conditional debug-only changes in offline-tts-matcha-impl.h that disable token population when debug is off, which appears unrelated to fixing punctuation overlap. Clarify the purpose of the debug conditional changes and how they relate to fixing the punctuation overlap issue, or remove them if unrelated.
✅ Passed checks (2 passed)
Check name Status Explanation
Linked Issues check ✅ Passed The PR changes tokenization logic (lazy_pinyin to pinyin) and removes space-token insertion for punctuation, which may address overlapping audio at punctuation marks in issue #2847.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello @csukuangfj, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request addresses an issue within the Matcha TTS Chinese-English model by refining its lexicon generation and token processing mechanisms. The core changes involve migrating the Python script responsible for lexicon creation to a more precise pinyin conversion method, adjusting C++ code to remove a potentially redundant space token insertion, and integrating new debug logging capabilities for token sequences to facilitate future troubleshooting. The author explicitly requests the regeneration of the lexicon.txt file, which is a direct outcome of these updates.

Highlights

  • Lexicon Generation Script Update: The generate_lexicon.py script was updated to use the pinyin function from the pypinyin library instead of lazy_pinyin. This change includes adjusting function parameters to handle neutral tones correctly and flattening the output of the pinyin function, which returns a list of lists, into a single list of pinyin strings.
  • C++ Lexicon Processing Refinement: A conditional block in matcha-tts-lexicon.cc that previously added a space token if a word started with an alphanumeric character or punctuation was removed, streamlining token handling.
  • Debug Logging for Token Processing: The token processing logic in offline-tts-matcha-impl.h was moved into a debug-only block. This change also introduces detailed logging of token IDs to SHERPA_ONNX_LOGE when debug mode is enabled, which will aid in diagnosing issues.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request aims to fix the Matcha TTS zh-en model by updating the lexicon generation script and C++ implementation. However, I've identified two critical issues that need to be addressed. First, the Python script for lexicon generation contains a bug that will cause it to crash due to incorrect handling of the data structure returned by the pinyin library. Second, in the C++ code, essential logic for preparing model inputs has been mistakenly moved into a debug-only conditional block, which will cause the TTS generation to fail in non-debug builds. Both of these issues are critical and will prevent the code from functioning as intended.

Comment on lines 48 to +57
for i in range(len(tokens)):
if tokens[i] == "shei2":
tokens[i] = "shui2"

if tokens[i][-1] not in ("1", "2", "3", "4", "5"):
tokens[i] += "1"

tokens = " ".join(tokens)
flatten = [t[0] for t in tokens]

tokens = " ".join(flatten)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

The logic for processing pinyin tokens is incorrect after switching from lazy_pinyin to pinyin. The pinyin function returns a list of lists (e.g., [['pin1'], ['yin1']]), so tokens[i] in your loop is a list, not a string. The current loop attempts to perform string operations on this list, which will raise a TypeError and cause the script to fail.

The logic needs to be refactored to correctly handle the nested list structure returned by pinyin.

Suggested change
for i in range(len(tokens)):
if tokens[i] == "shei2":
tokens[i] = "shui2"
if tokens[i][-1] not in ("1", "2", "3", "4", "5"):
tokens[i] += "1"
tokens = " ".join(tokens)
flatten = [t[0] for t in tokens]
tokens = " ".join(flatten)
processed_tokens = []
for t_list in tokens:
t = t_list[0]
if t == "shei2":
t = "shui2"
if t[-1] not in ("1", "2", "3", "4", "5"):
t += "1"
processed_tokens.append(t)
tokens = " ".join(processed_tokens)

Comment on lines +434 to 444
if (config_.model.debug) {
for (const auto &k : tokens) {
x.insert(x.end(), k.begin(), k.end());
}
std::ostringstream oss;
for (int32_t i : x) {
oss << i << ", ";
}
oss << "\n";
SHERPA_ONNX_LOGE("%s\n", oss.str().c_str());
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

critical

The logic for populating the x vector, which is the input to the model, has been mistakenly moved inside a debug-only block. This means that when config_.model.debug is false (i.e., in a release build), x will remain empty. Running the model with an empty input tensor will lead to incorrect audio generation or a crash.

The x vector must be populated regardless of whether debug mode is enabled.

Suggested change
if (config_.model.debug) {
for (const auto &k : tokens) {
x.insert(x.end(), k.begin(), k.end());
}
std::ostringstream oss;
for (int32_t i : x) {
oss << i << ", ";
}
oss << "\n";
SHERPA_ONNX_LOGE("%s\n", oss.str().c_str());
}
for (const auto &k : tokens) {
x.insert(x.end(), k.begin(), k.end());
}
if (config_.model.debug) {
std::ostringstream oss;
for (int32_t i : x) {
oss << i << ", ";
}
oss << "\n";
SHERPA_ONNX_LOGE("%s\n", oss.str().c_str());
}

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 172c906 and 109e4fe.

📒 Files selected for processing (3)
  • scripts/matcha-tts/zh-en/generate_lexicon.py (3 hunks)
  • sherpa-onnx/csrc/matcha-tts-lexicon.cc (0 hunks)
  • sherpa-onnx/csrc/offline-tts-matcha-impl.h (1 hunks)
💤 Files with no reviewable changes (1)
  • sherpa-onnx/csrc/matcha-tts-lexicon.cc
🔇 Additional comments (3)
scripts/matcha-tts/zh-en/generate_lexicon.py (3)

31-31: LGTM!

The change to pinyin() with [0][0] access correctly extracts the pinyin string from the nested list structure. The neutral_tone_with_five=True parameter ensures neutral tones are explicitly marked, which aligns with the tone normalization logic at lines 36-37.


55-57: LGTM with dependency on loop fix!

The flattening logic correctly extracts the first element from each sublist before joining. This is necessary because pinyin() returns nested lists. However, this logic depends on the loop at lines 48-53 being fixed to properly modify the nested structure.


3-3: Confirm the pypinyin API change from lazy_pinyin to pinyin.

The import update is correct. The pinyin() function returns a list of lists (e.g., [['zhōng'], ['xīn']]), requiring nested access patterns like [0][0] to get the first pinyin of the first character. It supports the neutral_tone_with_five parameter to mark neutral tones with "5", and flattening with [t[0] for t in result] extracts one pinyin per character as intended for phrase processing.

Comment on lines +47 to 54
tokens = pinyin(key, style=Style.TONE3, neutral_tone_with_five=True)
for i in range(len(tokens)):
if tokens[i] == "shei2":
tokens[i] = "shui2"

if tokens[i][-1] not in ("1", "2", "3", "4", "5"):
tokens[i] += "1"

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

Verify the loop logic with nested list structure.

After changing to pinyin(), the return value is a nested list (e.g., [['ni3'], ['hao3']]), but the loop at lines 48-53 accesses tokens[i] as if it were a string (checking tokens[i][-1] and comparing with strings). This suggests the loop should access tokens[i][0] instead of tokens[i].

Apply this diff to fix the nested list access:

 tokens = pinyin(key, style=Style.TONE3, neutral_tone_with_five=True)
 for i in range(len(tokens)):
-    if tokens[i] == "shei2":
-        tokens[i] = "shui2"
+    if tokens[i][0] == "shei2":
+        tokens[i][0] = "shui2"
 
-    if tokens[i][-1] not in ("1", "2", "3", "4", "5"):
-        tokens[i] += "1"
+    if tokens[i][0][-1] not in ("1", "2", "3", "4", "5"):
+        tokens[i][0] += "1"
🤖 Prompt for AI Agents
In scripts/matcha-tts/zh-en/generate_lexicon.py around lines 47 to 54, tokens
returned by pinyin() are nested lists (e.g., [['ni3'], ['hao3']]) but the code
treats tokens[i] as a string; update the logic to operate on the inner string
tokens[i][0] (read, modify, then assign back into tokens[i][0]) so checks like
equality to "shei2" and the final-tone check/append use tokens[i][0] and write
modifications back into the nested list.

Comment on lines +434 to 444
if (config_.model.debug) {
for (const auto &k : tokens) {
x.insert(x.end(), k.begin(), k.end());
}
std::ostringstream oss;
for (int32_t i : x) {
oss << i << ", ";
}
oss << "\n";
SHERPA_ONNX_LOGE("%s\n", oss.str().c_str());
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

Critical bug: Token vector x is empty when debug mode is off.

The conditional at line 434 makes the population of vector x dependent on config_.model.debug. When debug is false (typical in production), x remains empty, causing an empty tensor to be passed to the model at line 455. This breaks audio generation entirely.

Apply this diff to fix the issue by separating token aggregation from debug logging:

 std::vector<int64_t> x;
 x.reserve(num_tokens);
-if (config_.model.debug) {
-  for (const auto &k : tokens) {
-    x.insert(x.end(), k.begin(), k.end());
-  }
+for (const auto &k : tokens) {
+  x.insert(x.end(), k.begin(), k.end());
+}
+
+if (config_.model.debug) {
   std::ostringstream oss;
   for (int32_t i : x) {
     oss << i << ", ";
   }
   oss << "\n";
   SHERPA_ONNX_LOGE("%s\n", oss.str().c_str());
 }
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
if (config_.model.debug) {
for (const auto &k : tokens) {
x.insert(x.end(), k.begin(), k.end());
}
std::ostringstream oss;
for (int32_t i : x) {
oss << i << ", ";
}
oss << "\n";
SHERPA_ONNX_LOGE("%s\n", oss.str().c_str());
}
for (const auto &k : tokens) {
x.insert(x.end(), k.begin(), k.end());
}
if (config_.model.debug) {
std::ostringstream oss;
for (int32_t i : x) {
oss << i << ", ";
}
oss << "\n";
SHERPA_ONNX_LOGE("%s\n", oss.str().c_str());
}
🤖 Prompt for AI Agents
In sherpa-onnx/csrc/offline-tts-matcha-impl.h around lines 434 to 444, the token
aggregation into vector x is currently inside the debug-only conditional which
leaves x empty when debug is false and results in an empty tensor being passed
later; move the loop that appends tokens into x so it always runs (populate x
unconditionally before any debug checks), and keep only the ostringstream
construction and SHERPA_ONNX_LOGE call inside the config_.model.debug guard;
optionally reserve x size before the append to avoid reallocations.

@csukuangfj
csukuangfj merged commit a60d6d0 into k2-fsa:master Dec 3, 2025
10 of 27 checks passed
@csukuangfj
csukuangfj deleted the fix-matcha-tts-zh-en branch December 3, 2025 11:20
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:S This PR changes 10-29 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

matcha_tts_zh_en_20251010生成的语音在标点处有重叠音

1 participant