Repository navigation
Add CXX API for supertonic TTS - #3279
Wasser1462 wants to merge 24 commits into
Conversation
Made-with: Cursor
|
Caution Review failedPull request was closed or merged during review 📝 WalkthroughWalkthroughThis PR extends the sherpa-onnx C/C++ API to support the Supertonic TTS model by introducing new configuration structures, populating these fields through the API layers, adding a CMake build target, and providing a complete C++ example demonstrating Supertonic English TTS synthesis with progress callbacks. Changes
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~20 minutes Possibly related issues
Possibly related PRs
Suggested labels
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Tip Try Coding Plans. Let us write the prompt for your AI agent so you can ship faster (with fewer bugs). Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request integrates the Supertonic text-to-speech model into the existing framework, enhancing the system's speech synthesis capabilities. It introduces the necessary configuration structures and API mappings to support the new model, alongside a practical C++ example to demonstrate its usage and features like progress callbacks and generation parameter adjustments. Highlights
Changelog
Activity
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
There was a problem hiding this comment.
Code Review
This pull request adds CXX API support for the Supertonic TTS model. The changes include adding new configuration structures for Supertonic in both the C and CXX APIs, implementing the logic to handle these configurations, and providing a new example file to demonstrate its usage. The implementation looks solid and consistent with the existing structure for other TTS models. I've included one suggestion to improve the maintainability of the new example file by refactoring hardcoded model paths.
| config.model.supertonic.duration_predictor = | ||
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/" | ||
| "duration_predictor.int8.onnx"; | ||
| config.model.supertonic.text_encoder = | ||
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/text_encoder.int8.onnx"; | ||
| config.model.supertonic.vector_estimator = | ||
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/vector_estimator.int8.onnx"; | ||
| config.model.supertonic.vocoder = | ||
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/vocoder.int8.onnx"; | ||
| config.model.supertonic.tts_json = | ||
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/tts.json"; | ||
| config.model.supertonic.unicode_indexer = | ||
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/unicode_indexer.bin"; | ||
| config.model.supertonic.voice_style = | ||
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/voice.bin"; |
There was a problem hiding this comment.
To improve maintainability and reduce redundancy, consider defining the base path for the model files in a constant string and then constructing the full paths from it. This makes it easier to update the model path in the future.
| config.model.supertonic.duration_predictor = | |
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/" | |
| "duration_predictor.int8.onnx"; | |
| config.model.supertonic.text_encoder = | |
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/text_encoder.int8.onnx"; | |
| config.model.supertonic.vector_estimator = | |
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/vector_estimator.int8.onnx"; | |
| config.model.supertonic.vocoder = | |
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/vocoder.int8.onnx"; | |
| config.model.supertonic.tts_json = | |
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/tts.json"; | |
| config.model.supertonic.unicode_indexer = | |
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/unicode_indexer.bin"; | |
| config.model.supertonic.voice_style = | |
| "./sherpa-onnx-supertonic-tts-int8-2026-03-06/voice.bin"; | |
| const std::string model_dir = "./sherpa-onnx-supertonic-tts-int8-2026-03-06/"; | |
| config.model.supertonic.duration_predictor = | |
| model_dir + "duration_predictor.int8.onnx"; | |
| config.model.supertonic.text_encoder = model_dir + "text_encoder.int8.onnx"; | |
| config.model.supertonic.vector_estimator = | |
| model_dir + "vector_estimator.int8.onnx"; | |
| config.model.supertonic.vocoder = model_dir + "vocoder.int8.onnx"; | |
| config.model.supertonic.tts_json = model_dir + "tts.json"; | |
| config.model.supertonic.unicode_indexer = | |
| model_dir + "unicode_indexer.bin"; | |
| config.model.supertonic.voice_style = model_dir + "voice.bin"; |
Summary by CodeRabbit