Skip to content

Support passing multiple lexicon files for matcha tts models. - #2765

Merged
csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:fix-matcha
Nov 11, 2025
Merged

csukuangfj merged 2 commits into
k2-fsa:masterfrom
csukuangfj:fix-matcha

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Nov 11, 2025 •

Copy link
Copy Markdown
Collaborator

See also #2763

./build/bin/sherpa-onnx-offline-tts \
  --matcha-acoustic-model=./scripts/matcha-tts/zh-en/model-steps-6.onnx \
  --matcha-vocoder=./scripts/matcha-tts/zh-en/vocos-16khz-univ.onnx \
  --matcha-tokens=./scripts/matcha-tts/zh-en/tokens.txt \
  --matcha-lexicon=./scripts/matcha-tts/zh-en/lexicon.txt,/Users/fangjun/open-source/icefall-models/kokoro-multi-lang-v1_0/lexicon-gb-en.txt \
  --matcha-data-dir=/Users/fangjun/open-source/icefall-models/espeak-ng-data \
  --tts-rule-fsts=./matcha-icefall-zh-baker/phone.fst,./matcha-icefall-zh-baker/date.fst,./matcha-icefall-zh-baker/number.fst \
  --debug=1 \
  --tts-silence-scale=1 \
  --num-threads=4 \
  "中文测试:架起一座桥梁,繁华似锦,岁月如梭,这里是温柔女声电台,安抚你的情绪。下面是英文播报:hello world!once upon a time, a beautiful princess was locked in a tower by an evil witch.下面测试数字,开始:1 2 3 4 567890"

Summary by CodeRabbit

  • New Features

    • Lexicon configuration now supports multiple comma-separated file paths.
  • Improvements

    • Enhanced punctuation handling and tokenization in text-to-speech processing.
    • Improved phrase matching for non-ASCII characters.
    • Configuration validation now checks each lexicon path.
    • Added a character-classification utility to improve text processing.

@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Nov 11, 2025
@coderabbitai

coderabbitai Bot commented Nov 11, 2025 •

Copy link
Copy Markdown

Note

Other AI code review bot(s) detected

CodeRabbit has detected other AI code review bot(s) in this pull request and will avoid duplicating their findings in the review comments. This may lead to a less comprehensive review.

Walkthrough

Updated Matcha TTS lexicon loading to accept comma-separated files and load multiple lexicons; added IsAlphaOrPunct helper and ASCII-aware backward scanning in phrase matching; refined tokenization and punctuation handling; changed OnlinePunctuationModelConfig::debug type from int32_t to bool.

Changes

Cohort / File(s) Summary
API Type Update
sherpa-onnx/c-api/cxx-api.h
OnlinePunctuationModelConfig::debug changed from int32_t to bool
Matcha TTS Lexicon (multi-file)
sherpa-onnx/csrc/matcha-tts-lexicon.cc
New InitLexicon(const std::string &lexicon) overload to split comma-separated paths and load each file; constructors switched to call it; tokenization changes include appending space tokens for words starting with alpha/punct; removed prior single-file ifstream path and adjusted punctuation mappings and sentence-boundary debug logging
Config validation (lexicon paths)
sherpa-onnx/csrc/offline-tts-matcha-model-config.cc
Lexicon option help updated; validation now splits comma-separated paths and checks existence for each file; added text-utils.h include
Phrase matching logic
sherpa-onnx/csrc/phrase-matcher.cc
Added non-ASCII-start path with bounded backward scan (within max_search_len_) to find candidate words; adjusted inner-loop control flow and debug logging to reflect new search behavior
Text utilities
sherpa-onnx/csrc/text-utils.h, sherpa-onnx/csrc/text-utils.cc
Added bool IsAlphaOrPunct(int ch) declaration and definition using std::isalpha/std::ispunct

Sequence Diagram(s)

sequenceDiagram
    participant Caller as Calling Code
    participant Lex as MatchaTtsLexicon
    participant Init as InitLexicon(string)
    participant File as File I/O

    Caller->>Lex: Construct with lexicon string
    Lex->>Init: InitLexicon(comma_separated_paths)
    Init->>Init: Split string by ','
    loop for each path
        Init->>File: open(path)
        File-->>Init: stream/data
        Init->>Init: parse tokens & merge lexicon
    end
    Init-->>Lex: lexicons loaded
    Lex-->>Caller: ready
Loading
sequenceDiagram
    participant PM as PhraseMatcher
    participant Check as IsAlphaOrPunct
    participant Lex as Lexicon Lookup

    PM->>Check: Is current word start ASCII?
    alt Non-ASCII start
        PM->>PM: Set window (i..i+max_search_len_)
        PM->>PM: Backward scan to find candidate end
    else ASCII start
        PM->>PM: Use current position as start
    end
    PM->>Lex: Lookup candidate substring
    alt Match found
        PM->>PM: accept match, advance index, log if debug
    else No match
        PM->>PM: fallback single-word
    end
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

  • Attention points:
    • matcha-tts-lexicon.cc: multi-file parsing, tokenization changes, and sentence-boundary logging
    • phrase-matcher.cc: backward-scan logic correctness and boundary conditions for non-ASCII text
    • offline-tts-matcha-model-config.cc: validation of comma-separated paths
    • cxx-api.h: ensure ABI/public header change (type narrowing) is intentional

Possibly related PRs

Poem

🐰 I nibble commas, files multiplied,
I hop through tokens, spaces at my side,
Backward I scan where non-ASCII hides,
Booleans now whisper where integers once cried,
A rabbit cheers: lexicons unified! 🥕

Pre-merge checks and finishing touches

❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 6.25% which is insufficient. The required threshold is 80.00%. You can run @coderabbitai generate docstrings to improve docstring coverage.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title directly describes the main change: support for multiple lexicon files in Matcha TTS models, which aligns with the core changes across all modified files.
✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

📜 Recent review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 35f1a7f and 53fd8b6.

📒 Files selected for processing (4)
  • sherpa-onnx/csrc/matcha-tts-lexicon.cc (6 hunks)
  • sherpa-onnx/csrc/phrase-matcher.cc (1 hunks)
  • sherpa-onnx/csrc/text-utils.cc (1 hunks)
  • sherpa-onnx/csrc/text-utils.h (1 hunks)
🧰 Additional context used
🧬 Code graph analysis (3)
sherpa-onnx/csrc/text-utils.h (1)
sherpa-onnx/csrc/text-utils.cc (2)
  • IsAlphaOrPunct (811-811)
  • IsAlphaOrPunct (811-811)
sherpa-onnx/csrc/phrase-matcher.cc (1)
sherpa-onnx/csrc/text-utils.cc (5)
  • IsAlphaOrPunct (811-811)
  • IsAlphaOrPunct (811-811)
  • i (114-114)
  • GetWord (794-809)
  • GetWord (794-795)
sherpa-onnx/csrc/matcha-tts-lexicon.cc (2)
sherpa-onnx/csrc/kokoro-multi-lang-lexicon.cc (12)
  • InitLexicon (484-497)
  • InitLexicon (484-484)
  • lexicon (470-481)
  • lexicon (470-470)
  • w (178-211)
  • w (178-178)
  • is (433-433)
  • is (445-468)
  • is (445-445)
  • is (478-478)
  • is (499-543)
  • is (499-499)
sherpa-onnx/csrc/text-utils.cc (5)
  • SplitStringToVector (130-142)
  • SplitStringToVector (130-132)
  • i (114-114)
  • IsAlphaOrPunct (811-811)
  • IsAlphaOrPunct (811-811)
🔇 Additional comments (8)
sherpa-onnx/csrc/text-utils.cc (1)

811-812: LGTM! Clean helper function.

The implementation is straightforward and uses standard C++ library functions. The int parameter type correctly matches the signature of std::isalpha and std::ispunct.

sherpa-onnx/csrc/text-utils.h (1)

168-169: LGTM! Declaration is correct.

The function signature matches the implementation and is appropriately placed among other character utility functions.

sherpa-onnx/csrc/matcha-tts-lexicon.cc (6)

60-60: LGTM: Multi-file lexicon support correctly integrated.

The change from direct file loading to calling InitLexicon(lexicon) enables comma-separated lexicon file paths, matching the pattern used in similar implementations.


86-93: LGTM: Multi-file loading correctly implemented for managed resources.

The implementation properly splits comma-separated lexicon paths and loads each file through the asset manager, consistent with the reference pattern in kokoro-multi-lang-lexicon.cc.


106-106: Punctuation mapping correctly simplified.

The removal of {":", ","} is appropriate—the mapping now consistently converts Chinese punctuation to English equivalents without altering English punctuation. This addresses the past review concern about the duplicate Chinese colon mapping.


208-219: LGTM: Debug logging enhances observability.

The sentence boundary logging provides useful visibility into tokenization when debugging is enabled, with no performance impact in production.


277-277: LGTM: Safer map access pattern.

Using .at() instead of [] is appropriate for debug logging—it will throw an exception if an invalid token ID is encountered, helping catch bugs rather than silently inserting default values.


322-334: LGTM: Multi-file lexicon loading correctly implemented.

The new InitLexicon(const std::string &lexicon) overload cleanly handles comma-separated file paths and delegates stream parsing to the existing InitLexicon(std::istream &is) method, following the established pattern in kokoro-multi-lang-lexicon.cc.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello @csukuangfj, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the flexibility and robustness of Matcha TTS models by enabling the use of multiple lexicon files. This change allows for more comprehensive phonetic mappings, especially beneficial for multilingual applications, while also refining text processing and phrase matching logic to improve overall accuracy and debugging capabilities.

Highlights

  • Multiple Lexicon Files Support: The Matcha TTS models now support loading multiple lexicon files by providing a comma-separated list of paths to the --matcha-lexicon argument, enhancing flexibility for diverse phonetic mappings.
  • Enhanced Lexicon Loading and Validation: The system has been updated to iterate through and load each specified lexicon file individually, and the configuration validation now rigorously checks for the existence of all provided lexicon files, ensuring robust model initialization.
  • Improved Multi-Language Text Processing: Adjustments were made to punctuation handling, including the removal of some implicit punctuation mappings and the addition of space tokens before alphabetic words. Furthermore, specific phrase matching logic was introduced for non-ASCII characters, which is particularly beneficial for multi-language text-to-speech scenarios.
  • Debugging and Type Safety Improvements: The debug member type in OnlinePunctuationModelConfig was corrected from int32_t to bool for better type safety. Additionally, new debug logging was added to display token IDs for generated sentences and to refine the output of token IDs during word conversion, offering clearer insights into the tokenization process.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request successfully adds support for passing multiple lexicon files for Matcha TTS models by parsing a comma-separated string of file paths. The changes are logical and cover both configuration and lexicon loading. My review focuses on improving code robustness and correctness. I've pointed out a potential bug in a punctuation mapping, several places where operations on potentially empty strings could lead to undefined behavior, and an unused variable. I've also included suggestions to adhere to modern C++ practices by using C++ standard headers.

Comment thread sherpa-onnx/csrc/matcha-tts-lexicon.cc Outdated
std::string text = _text;
std::vector<std::pair<std::string, std::string>> replace_str_pairs = {
{",", ","}, {":", ","}, {"、", ","}, {";", ";"}, {":", ":"},
{",", ","}, {":", ":"}, {"、", ","}, {";", ";"}, {":", ":"},

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

There seems to be a copy-paste error here. The pair {":", ":"} is duplicated, and the original {":", ","} has been removed. This is likely unintentional and could affect text normalization. Please restore the original mapping and remove the duplicate.

        {",", ","}, {":", ","},  {"、", ","}, {";", ";"},   {":", ":"},

Comment thread sherpa-onnx/csrc/matcha-tts-lexicon.cc Outdated
}
}

if (isalpha(w.front())) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Calling w.front() on an empty string w would lead to undefined behavior. It's safer to add a check to ensure the string is not empty before accessing its first character.

    if (!w.empty() && isalpha(w.front())) {

Comment thread sherpa-onnx/csrc/phrase-matcher.cc Outdated
auto this_word = GetWord(words, start, end);
if (debug_) {

if (!isascii(words[i].front())) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Calling words[i].front() on an empty string would lead to undefined behavior. It's safer to add a check to ensure the string is not empty before accessing its first character.

Suggested change
if (!isascii(words[i].front())) {
if (!words[i].empty() && !isascii(words[i].front())) {

Comment thread sherpa-onnx/csrc/phrase-matcher.cc Outdated

while (end > start) {
auto this_word = GetWord(words, start, end);
if (isascii(this_word.back())) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The GetWord function can return an empty string. Calling .back() on an empty string results in undefined behavior. Please add a check to ensure this_word is not empty before accessing its last character.

Suggested change
if (isascii(this_word.back())) {
if (!this_word.empty() && isascii(this_word.back())) {

Comment thread sherpa-onnx/csrc/matcha-tts-lexicon.cc Outdated

#include "sherpa-onnx/csrc/matcha-tts-lexicon.h"

#include <ctype.h>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

For better C++ practice, it's recommended to use C++ headers instead of C headers. Please include <cctype> and use functions from the std namespace, like std::isalpha (used on line 273).

#include <cctype>

Comment thread sherpa-onnx/csrc/matcha-tts-lexicon.cc Outdated
private:
std::vector<int32_t> ConvertWordToIds(const std::string &w) const {
std::vector<int32_t> ans;
int32_t space_id = token2id_.at(" ");

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This space_id variable is initialized but never used. Later, at line 274, token2id_.at(" ") is called again. To avoid the redundant map lookup, please use this space_id variable at line 274.

Comment thread sherpa-onnx/csrc/phrase-matcher.cc Outdated
// Copyright (c) 2025 Xiaomi Corporation
#include "sherpa-onnx/csrc/phrase-matcher.h"

#include <ctype.h>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

For better C++ practice, it's recommended to use C++ headers instead of C headers. Please include <cctype> and use functions from the std namespace, like std::isascii (used on lines 64 and 69).

Suggested change
#include <ctype.h>
#include <cctype>

@csukuangfj
csukuangfj requested a review from Copilot November 11, 2025 04:05

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🧹 Nitpick comments (2)
sherpa-onnx/csrc/matcha-tts-lexicon.cc (1)

7-7: Prefer <cctype> for C++ code.

While <ctype.h> works, the C++ standard recommends using <cctype> for C++ code.

Apply this diff:

-#include <ctype.h>
+#include <cctype>
sherpa-onnx/csrc/phrase-matcher.cc (1)

64-98: Add documentation for the non-ASCII phrase matching algorithm.

The new logic introduces complex conditional behavior (ASCII domain checks, backward scanning, candidate filtering) without explanatory comments. This makes the code difficult to understand and maintain.

Consider adding a comment block explaining:

  • Why non-ASCII words trigger multi-word matching
  • Why candidates ending with ASCII characters are skipped (line 69-72)
  • The algorithm's purpose (e.g., correctly handling mixed-language phrases like Chinese text with embedded English words)

Example:

       std::string w;
+
+      // For words starting with non-ASCII (e.g., Chinese, Japanese),
+      // attempt to match multi-word phrases from the lexicon.
+      // Skip candidates ending with ASCII to avoid breaking up
+      // mixed-language phrases incorrectly (e.g., "你好world" should
+      // not match "你好" if "你好world" is in the lexicon).
       if (!isascii(words[i].front())) {
📜 Review details

Configuration used: CodeRabbit UI

Review profile: CHILL

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between 5509c1b and 35f1a7f.

📒 Files selected for processing (4)
  • sherpa-onnx/c-api/cxx-api.h (1 hunks)
  • sherpa-onnx/csrc/matcha-tts-lexicon.cc (8 hunks)
  • sherpa-onnx/csrc/offline-tts-matcha-model-config.cc (2 hunks)
  • sherpa-onnx/csrc/phrase-matcher.cc (2 hunks)
🧰 Additional context used
🧬 Code graph analysis (2)
sherpa-onnx/csrc/offline-tts-matcha-model-config.cc (2)
sherpa-onnx/csrc/matcha-tts-lexicon.cc (2)
  • lexicon (326-338)
  • lexicon (326-326)
sherpa-onnx/csrc/text-utils.cc (2)
  • SplitStringToVector (130-142)
  • SplitStringToVector (130-132)
sherpa-onnx/csrc/matcha-tts-lexicon.cc (3)
sherpa-onnx/csrc/kokoro-multi-lang-lexicon.cc (12)
  • InitLexicon (484-497)
  • InitLexicon (484-484)
  • lexicon (470-481)
  • lexicon (470-470)
  • w (178-211)
  • w (178-178)
  • is (433-433)
  • is (445-468)
  • is (445-445)
  • is (478-478)
  • is (499-543)
  • is (499-499)
sherpa-onnx/csrc/text-utils.cc (3)
  • SplitStringToVector (130-142)
  • SplitStringToVector (130-132)
  • i (114-114)
sherpa-onnx/csrc/character-lexicon.cc (8)
  • w (194-225)
  • w (194-194)
  • is (47-47)
  • is (52-52)
  • is (227-258)
  • is (227-227)
  • is (260-312)
  • is (260-260)
⏰ Context from checks skipped due to timeout of 90000ms. You can increase the timeout in your CodeRabbit configuration to a maximum of 15 minutes (900000ms). (19)
  • GitHub Check: rknn shared OFF
  • GitHub Check: rknn shared ON
  • GitHub Check: Debug shared tts-ON
  • GitHub Check: Debug shared-ON tts-ON
  • GitHub Check: Release shared tts-ON
  • GitHub Check: swift (macos-latest)
  • GitHub Check: Debug shared-OFF tts-OFF
  • GitHub Check: Release shared-OFF tts-OFF
  • GitHub Check: Debug shared-ON tts-OFF
  • GitHub Check: swift (macos-13)
  • GitHub Check: Release shared-OFF tts-ON
  • GitHub Check: ubuntu-24.04 3.12
  • GitHub Check: ubuntu-24.04 3.13
  • GitHub Check: ubuntu-24.04 3.11
  • GitHub Check: Release shared-ON tts-ON
  • GitHub Check: ubuntu-24.04 3.10
  • GitHub Check: ubuntu-24.04 3.9
  • GitHub Check: ubuntu-24.04 3.8
  • GitHub Check: ascend (gpustack/devel-ascendai-cann:8.0.rc3.beta1-310p-ubuntu20.04-v2, 8.0.0-310p)
🔇 Additional comments (10)
sherpa-onnx/c-api/cxx-api.h (1)

745-745: LGTM! Type correctness improvement.

The change from int32_t to bool for the debug field is appropriate and improves type safety.

sherpa-onnx/csrc/offline-tts-matcha-model-config.cc (3)

11-11: LGTM!

The include is necessary for the SplitStringToVector function used in the validation logic.


19-22: LGTM! Clear documentation of multi-file support.

The help text clearly explains that multiple lexicon files can be passed as a comma-separated list, which will help users understand the new capability.


88-99: LGTM! Robust validation logic.

The validation correctly splits the comma-separated lexicon string and verifies the existence of each file individually, providing clear error messages if any file is missing.

sherpa-onnx/csrc/matcha-tts-lexicon.cc (6)

62-62: LGTM! Cleaner multi-file lexicon loading.

The change from directly opening a file stream to calling the new InitLexicon(const std::string &) method enables multi-file lexicon support consistently.


88-95: LGTM! Consistent multi-file loading implementation.

The template constructor correctly splits the comma-separated lexicon string and loads each file individually, consistent with the non-template constructor implementation.


108-108: LGTM! Removed duplicate mapping.

The duplicate {":", ":"} entry has been correctly removed from the punctuation mapping list.


210-220: LGTM! Helpful debug logging.

The debug logging for sentence boundaries provides useful visibility into the tokenization process.


273-275: Verify the space appending behavior for alphabetic words.

The code now appends a space token to the end of the token ID sequence when a word starts with an alphabetic character. This changes the tokenization behavior and could affect TTS output quality.

Ensure this behavior is intentional and that:

  1. It correctly handles mixed Chinese/English text
  2. It doesn't create unexpected pauses or artifacts in the generated speech
  3. The space is appended at the correct position relative to other processing

326-338: LGTM! Well-structured multi-file lexicon loading.

The new InitLexicon(const std::string &) overload correctly implements multi-file lexicon loading by splitting the comma-separated string and loading each file individually. This approach is consistent with similar implementations in the codebase (e.g., kokoro-multi-lang-lexicon.cc).

Comment thread sherpa-onnx/csrc/matcha-tts-lexicon.cc Outdated
Comment thread sherpa-onnx/csrc/phrase-matcher.cc Outdated
Comment on lines +6 to +7
#include <ctype.h>

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

Consider portability of isascii().

The isascii() function is a POSIX extension, not standard C++. If the codebase targets non-POSIX platforms (e.g., certain embedded systems, Windows without POSIX compatibility layer), this may cause compilation failures.

Consider using a standard alternative:

-#include <ctype.h>
-
+// Helper for ASCII check (standard C++ compatible)
+inline bool IsAsciiChar(char c) {
+  return static_cast<unsigned char>(c) <= 127;
+}

Then replace isascii() calls with IsAsciiChar().

Alternatively, verify that all target platforms support POSIX isascii():


🏁 Script executed:

#!/bin/bash
# Check if the codebase already uses other POSIX-specific functions
rg -n 'isascii|isalpha|isdigit|isalnum' --type cpp -g '!*test*' -C 2

Length of output: 2662



---

🏁 Script executed:

```shell
#!/bin/bash
# Check phrase-matcher.cc includes and look for platform/build configuration
head -20 sherpa-onnx/csrc/phrase-matcher.cc

Length of output: 493


🏁 Script executed:

#!/bin/bash
# Check if codebase has platform-specific configurations
fd -t f '(CMakeLists|Makefile|BUILD|bazel)' --type f | head -20

Length of output: 907


🏁 Script executed:

#!/bin/bash
# Check how other files handle character checking for consistency
rg -n 'std::isalpha|std::isdigit' sherpa-onnx/csrc/phrase-matcher.cc

Length of output: 0



---

🏁 Script executed:

```shell
#!/bin/bash
# Find GetWord() implementation to check if it can return empty strings
rg -n 'GetWord\s*\(' sherpa-onnx/csrc/phrase-matcher.cc -A 10

Length of output: 354


🏁 Script executed:

#!/bin/bash
# Check CMakeLists.txt for platform support and POSIX requirement statements
grep -i 'posix\|platform\|windows\|embed' CMakeLists.txt | head -20

Length of output: 122



---

🏁 Script executed:

```shell
#!/bin/bash
# Search for GetWord() function definition in the file
rg -n 'std::string.*GetWord|GetWord.*{' sherpa-onnx/csrc/phrase-matcher.cc -A 15

Length of output: 125


🏁 Script executed:

#!/bin/bash
# Also check phrase-matcher.h for GetWord signature
cat sherpa-onnx/csrc/phrase-matcher.h | grep -A 10 'GetWord'

Length of output: 0



---

🏁 Script executed:

```shell
#!/bin/bash
# Search for GetWord method definition with simpler pattern
rg -n 'GetWord' sherpa-onnx/csrc/phrase-matcher.cc -B 2 -A 8

Length of output: 330


🏁 Script executed:

#!/bin/bash
# View the complete phrase-matcher.cc to understand full structure
cat sherpa-onnx/csrc/phrase-matcher.cc

Length of output: 3618



---

🏁 Script executed:

```shell
#!/bin/bash
# Search for GetWord definition in header and related files
rg -n 'GetWord' sherpa-onnx/csrc/phrase-matcher.h

Length of output: 0


🏁 Script executed:

#!/bin/bash
# Search for GetWord in all cpp/h files in csrc
fd -e cpp -e h . sherpa-onnx/csrc | xargs grep -l 'GetWord' 2>/dev/null

Length of output: 29



---

🏁 Script executed:

```shell
#!/bin/bash
# Search for GetWord definition in text-utils
rg -n 'GetWord' sherpa-onnx/csrc/text-utils.h -B 2 -A 12

Length of output: 277


🏁 Script executed:

#!/bin/bash
# Also check if text-utils.h is included in phrase-matcher.cc
grep -n 'text-utils' sherpa-onnx/csrc/phrase-matcher.cc

Length of output: 43



---

🏁 Script executed:

```shell
#!/bin/bash
# Find GetWord implementation in text-utils.cc
rg -n 'std::string GetWord' sherpa-onnx/csrc/text-utils.cc -A 20

Length of output: 594


Replace isascii() with standard C++ alternative for cross-platform compatibility.

The isascii() function is a POSIX extension, not part of the C++ standard library. While other files use std::isdigit() and std::isalpha(), this file uses bare isascii(), creating an inconsistency. On platforms without POSIX support (e.g., certain embedded systems, Windows without POSIX compatibility), this will cause compilation failures.

Recommended fix:

Replace #include <ctype.h> with a standard-compatible ASCII check:

-#include <ctype.h>
+// Helper for standard C++ compatible ASCII checking
+inline bool IsAsciiChar(char c) {
+  return static_cast<unsigned char>(c) <= 127;
+}

Then replace isascii() calls at lines 64 and 69 with IsAsciiChar(). Additionally, add defensive empty-string checks before .front() and .back() calls to avoid undefined behavior on edge cases.

Committable suggestion skipped: line range outside the PR's diff.

Comment thread sherpa-onnx/csrc/phrase-matcher.cc Outdated
auto this_word = GetWord(words, start, end);
if (debug_) {

if (!isascii(words[i].front())) {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

Add empty string checks to prevent undefined behavior.

Calling .front() on line 64 and .back() on line 69 without checking if the strings are empty invokes undefined behavior and can cause crashes.

Apply this diff to add safety checks:

-      if (!isascii(words[i].front())) {
+      if (!words[i].empty() && !isascii(words[i].front())) {
         int32_t end = std::min(i + max_search_len_ - 1, num_words - 1);
 
         while (end > start) {
           auto this_word = GetWord(words, start, end);
-          if (isascii(this_word.back())) {
+          if (this_word.empty() || isascii(this_word.back())) {
             --end;
             continue;
           }

Also applies to: 69-69

🤖 Prompt for AI Agents
In sherpa-onnx/csrc/phrase-matcher.cc around lines 64 and 69, the code calls
words[i].front() and words[i].back() without verifying the string is non-empty,
which is undefined behavior; add checks that words[i].empty() is false (or
!words[i].empty()) before calling .front() or .back(), and skip or handle empty
strings appropriately (e.g., continue the loop or treat them as
non-ascii/mismatch) so the front()/back() calls are only executed on non-empty
strings.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot encountered an error and was unable to review this pull request. You can try again by re-requesting a review.

@csukuangfj
csukuangfj merged commit 2b81e4d into k2-fsa:master Nov 11, 2025
43 of 44 checks passed
@csukuangfj
csukuangfj deleted the fix-matcha branch November 11, 2025 06:32
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants