Add KittenTTS v0.8 support - #3591
Conversation
📝 WalkthroughWalkthroughThis PR adds comprehensive Kitten TTS v0.8 model support to sherpa-onnx. It expands the C++ runtime metadata contract, implements dynamic style row selection with speaker speed priors, introduces Kitten-specific text phonemization, provides asset generation tools, updates build systems across platforms, and integrates support into Flutter, Swift, WASM, and CI/CD workflows for v0.8 model variants. ChangesKitten TTS v0.8 Runtime and Build Integration
Sequence Diagram(s)sequenceDiagram
participant Dev as Developer
participant Run as v0_8/run.sh
participant HF as HuggingFace
participant GenV as generate_voices_bin.py
participant GenT as generate_tokens.py
participant AddM as add_meta_data.py
participant Pack as package_dir
Dev->>Run: invoke run.sh (variant)
Run->>HF: download ONNX + voices.npz (if missing)
Run->>GenV: generate voices.bin
Run->>GenT: generate tokens.txt
Run->>AddM: inject metadata into ONNX
AddM-->>Pack: produce model.onnx (with metadata)
GenV-->>Pack: produce voices.bin
GenT-->>Pack: produce tokens.txt
Run-->>Dev: list package_dir
Estimated code review effort🎯 4 (Complex) | ⏱️ ~45 minutes Possibly related PRs
Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Tip 💬 Introducing Slack Agent: The best way for teams to turn conversations into code.Slack Agent is built on CodeRabbit's deep understanding of your code, so your team can collaborate across the entire SDLC without losing context.
Built for teams:
One agent for your entire SDLC. Right inside Slack. Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Code Review
This pull request introduces support for KittenTTS v0.8 models across multiple platforms, including Flutter, Android, and WASM. It adds scripts for model preparation, metadata handling, and token generation, while updating the C++ core to support multi-row style selection and speaker-specific speed priors. A critical issue was identified regarding the lack of bounds checking for the speaker ID (sid), which could lead to out-of-bounds memory access if an invalid ID is provided.
| } | ||
|
|
||
| int32_t num_speakers = meta_data_.num_speakers; | ||
| int32_t sid_int = static_cast<int32_t>(sid); |
There was a problem hiding this comment.
The speaker ID sid is used to index into the styles_ vector and the speaker_speed_priors vector without any bounds checking. If a user provides an invalid sid (e.g., via a client API), it will result in out-of-bounds memory access, leading to undefined behavior or a crash. It is recommended to validate that 0 <= sid < meta_data_.num_speakers and handle invalid values gracefully, for example by logging an error and falling back to speaker 0.
int32_t sid_int = static_cast<int32_t>(sid);
if (sid_int < 0 || sid_int >= meta_data_.num_speakers) {
SHERPA_ONNX_LOGE(
"sid (%d) is out of range. num_speakers: %d. Use 0 by default",
sid_int, meta_data_.num_speakers);
sid_int = 0;
}There was a problem hiding this comment.
Actionable comments posted: 5
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
sherpa-onnx/csrc/offline-tts-kitten-model.cc (1)
71-90:⚠️ Potential issue | 🔴 Critical | ⚡ Quick winValidate
sidbefore indexing speaker-specific buffers.
sidis now used to index bothstyles_andspeaker_speed_priorswithout any bounds check. A negative or oversized speaker ID will read outside those buffers and can crash or corrupt inference state.Suggested fix
- int32_t sid_int = static_cast<int32_t>(sid); + if (sid < 0 || sid >= meta_data_.num_speakers) { + SHERPA_ONNX_LOGE("Invalid speaker id: %lld. Valid range: [0, %d)", + static_cast<long long>(sid), meta_data_.num_speakers); + SHERPA_ONNX_EXIT(-1); + } + + int32_t sid_int = static_cast<int32_t>(sid); int32_t dim1 = style_dim_[1]; int32_t style_rows = style_dim_[0]; int32_t row = SelectStyleRow(x, style_rows);🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@sherpa-onnx/csrc/offline-tts-kitten-model.cc` around lines 71 - 90, Validate sid before using it to index styles_ and meta_data_.speaker_speed_priors: compute a safe num_speakers (e.g., styles_.size() / (style_rows * dim1)) and check that sid_int >= 0 and sid_int < num_speakers, and also ensure sid_int < meta_data_.speaker_speed_priors.size() before accessing those containers; if the checks fail, return/throw an error (or handle gracefully) rather than proceeding to compute p or multiply speed. Update the code around sid_int, styles_, style_dim_, and meta_data_.speaker_speed_priors (and where SelectStyleRow is called) to perform these bounds checks and early exit on invalid sid.
🧹 Nitpick comments (2)
scripts/kitten-tts/v0_8/add_meta_data.py (1)
89-90: 💤 Low valuePrefer
del model.metadata_props[:]over the pop loop.The
while/poppattern is O(n²); clearing a protobufRepeatedCompositeContainerwith a slice delete is O(1) and more idiomatic.♻️ Proposed refactor
- while len(model.metadata_props): - model.metadata_props.pop() + del model.metadata_props[:]🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@scripts/kitten-tts/v0_8/add_meta_data.py` around lines 89 - 90, The loop that repeatedly pops from model.metadata_props is O(n^2); replace the pop loop with a slice deletion to clear the RepeatedCompositeContainer in O(1) by using del model.metadata_props[:] (or the equivalent clear operation) so the metadata_props field on the model is emptied efficiently.scripts/kitten-tts/v0_8/generate_tokens.py (1)
12-12: 💤 Low valueConsider iterable unpacking over list concatenation (Ruff RUF005).
♻️ Proposed refactor
- symbols = [_pad] + list(_punctuation) + list(_letters) + list(_letters_ipa) + symbols = [_pad, *_punctuation, *_letters, *_letters_ipa]🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@scripts/kitten-tts/v0_8/generate_tokens.py` at line 12, Replace the list concatenation used to build symbols with iterable unpacking: instead of creating intermediate lists by adding list(_punctuation), list(_letters), list(_letters_ipa) to [_pad], build the final sequence by unpacking the iterables directly into a single list; update the assignment to symbols to use _pad, _punctuation, _letters, and _letters_ipa with unpacking so there are no unnecessary temporary lists.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@scripts/kitten-tts/v0_8/run.sh`:
- Around line 56-67: ShellCheck SC2086 is triggered by unquoted variable
expansions; update the commands that reference ${onnx_name}, ${base_url},
voices.npz, ${output_name}, and ${model_name} to use quoted expansions to
prevent word-splitting/globbing. Specifically, change the curl invocations to
use "$base_url/$onnx_name" and "$base_url/voices.npz", quote the file tests and
cp like [ ! -f "$onnx_name" ] and cp "$onnx_name" "$output_name", and quote the
add_meta_data.py args such as --model "./$output_name" and --model-name
"$model_name" (also quote any other bare ${...} in this snippet).
In `@sherpa-onnx/csrc/offline-tts-kitten-model.cc`:
- Around line 258-269: The num_content_tokens logic fails to always exclude a
terminal end_id because it only decrements when add_pad_after_end is true;
update the logic in the block that computes num_content_tokens (look for
variables num_content_tokens, p, num_tokens and meta_data_.end_id) to always
decrement num_content_tokens when the last token equals meta_data_.end_id (i.e.,
check p[num_tokens-1] == meta_data_.end_id and --num_content_tokens), keeping
the existing checks for start_id and pad_id intact so SelectStyleRow() receives
the correct content token count.
In `@sherpa-onnx/csrc/piper-phonemize-lexicon.cc`:
- Around line 252-272: The boundary check currently uses
max_current_size_before_token which doesn't account for the tokens that will be
appended for the current phoneme (and the extra space token when p == '.') nor
the trailing end_id/pad_id, so on exact-boundary inputs chunks can exceed
meta_data.max_token_len; change the check to compute the required space before
appending by calculating tokens_to_append = 1 + (p == '.' ? 1 : 0) and
trailing_space = 1 + (meta_data.add_pad_after_end ? 1 : 0) (or equivalent), and
replace the condition using max_current_size_before_token with a check that
current.size() + tokens_to_append + trailing_space >
static_cast<size_t>(meta_data.max_token_len) so you close the current chunk
(push end_id/pad_id, push current into ans, reserve and start a new chunk with
start_id) before adding token2id.at(p) (and the space token if p == '.'); update
uses of max_current_size_before_token to this on-the-fly calculation to ensure
no emitted sequence exceeds meta_data.max_token_len.
In `@swift-api-examples/run-tts-kitten-en.sh`:
- Around line 13-21: The preflight check in run-tts-kitten-en.sh currently only
verifies ./kitten-mini-en-v0_8/model.onnx; update the script to also verify the
presence of ./kitten-mini-en-v0_8/voices.bin and
./kitten-mini-en-v0_8/tokens.txt before proceeding, and if any are missing print
a clear message (similar to the existing echo block) listing the missing files
and exit 1; locate the existing if block that checks for model.onnx and add
tests for voices.bin and tokens.txt (or loop over an array of required files) so
all required assets are validated up front.
- Around line 16-20: Update the setup hint so the copy command targets the
generated model directory instead of the entire v0_8 folder: change the echoed
cp line that currently shows "cp -R ../scripts/kitten-tts/v0_8
./kitten-mini-en-v0_8" to point to the generated subdirectory (e.g., "cp -R
../scripts/kitten-tts/v0_8/kitten-mini-en-v0_8 ./kitten-mini-en-v0_8") so the
script matches the expected runtime layout.
---
Outside diff comments:
In `@sherpa-onnx/csrc/offline-tts-kitten-model.cc`:
- Around line 71-90: Validate sid before using it to index styles_ and
meta_data_.speaker_speed_priors: compute a safe num_speakers (e.g.,
styles_.size() / (style_rows * dim1)) and check that sid_int >= 0 and sid_int <
num_speakers, and also ensure sid_int < meta_data_.speaker_speed_priors.size()
before accessing those containers; if the checks fail, return/throw an error (or
handle gracefully) rather than proceeding to compute p or multiply speed. Update
the code around sid_int, styles_, style_dim_, and
meta_data_.speaker_speed_priors (and where SelectStyleRow is called) to perform
these bounds checks and early exit on invalid sid.
---
Nitpick comments:
In `@scripts/kitten-tts/v0_8/add_meta_data.py`:
- Around line 89-90: The loop that repeatedly pops from model.metadata_props is
O(n^2); replace the pop loop with a slice deletion to clear the
RepeatedCompositeContainer in O(1) by using del model.metadata_props[:] (or the
equivalent clear operation) so the metadata_props field on the model is emptied
efficiently.
In `@scripts/kitten-tts/v0_8/generate_tokens.py`:
- Line 12: Replace the list concatenation used to build symbols with iterable
unpacking: instead of creating intermediate lists by adding list(_punctuation),
list(_letters), list(_letters_ipa) to [_pad], build the final sequence by
unpacking the iterables directly into a single list; update the assignment to
symbols to use _pad, _punctuation, _letters, and _letters_ipa with unpacking so
there are no unnecessary temporary lists.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: 3e66c774-12d4-4c47-9342-471faf8eae32
📒 Files selected for processing (19)
.github/workflows/export-kitten.yamlflutter-examples/tts/lib/model.dartscripts/apk/generate-tts-apk-script.pyscripts/flutter/build-android-tts.sh.inscripts/flutter/build-linux-tts.sh.inscripts/flutter/build-macos-tts.sh.inscripts/flutter/build-windows-tts.sh.inscripts/flutter/generate-tts.pyscripts/kitten-tts/README.mdscripts/kitten-tts/v0_8/add_meta_data.pyscripts/kitten-tts/v0_8/generate_tokens.pyscripts/kitten-tts/v0_8/generate_voices_bin.pyscripts/kitten-tts/v0_8/run.shsherpa-onnx/csrc/offline-tts-kitten-model-meta-data.hsherpa-onnx/csrc/offline-tts-kitten-model.ccsherpa-onnx/csrc/piper-phonemize-lexicon.ccswift-api-examples/run-tts-kitten-en.shswift-api-examples/tts-kitten-en.swiftwasm/tts/sherpa-onnx-tts.js
|
Thank you for your contribution! Will review it soon. |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@swift-api-examples/run-tts-kitten-en.sh`:
- Around line 41-44: The download/extract sequence for the
espeak-ng-data.tar.bz2 archive lacks integrity checks; add a SHA256 checksum
verification step before running tar: define an expected checksum string for
"espeak-ng-data.tar.bz2", compute the actual checksum with a portable tool
(sha256sum or shasum -a 256), compare them, and if they differ print an error
and exit non‑zero (and optionally delete the corrupt file). Only proceed to tar
xf and mv into the "${model_dir}/" when the checksum matches. Ensure the check
runs immediately after the curl -O step and before extraction so tampered or
corrupted archives are rejected.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: bf502df3-7924-47a9-ace8-aa4c74481042
📒 Files selected for processing (2)
scripts/kitten-tts/v0_8/run.shswift-api-examples/run-tts-kitten-en.sh
🚧 Files skipped from review as they are similar to previous changes (1)
- scripts/kitten-tts/v0_8/run.sh
| curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/espeak-ng-data.tar.bz2 | ||
| tar xf espeak-ng-data.tar.bz2 | ||
| rm espeak-ng-data.tar.bz2 | ||
| mv espeak-ng-data "${model_dir}/" |
There was a problem hiding this comment.
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Read-only verification: confirm integrity check is present for downloaded archive.
rg -n 'espeak-ng-data\.tar\.bz2|sha256|shasum|openssl dgst' swift-api-examples/run-tts-kitten-en.shRepository: k2-fsa/sherpa-onnx
Length of output: 235
Add checksum verification for downloaded archive.
The tarball download (line 41) and extraction (line 42) lack integrity verification. This creates a supply-chain security gap where tampering or corruption could go undetected.
Add SHA256 checksum validation before extraction:
Suggested fix
if [ ! -d "${model_dir}/espeak-ng-data" ]; then
- curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/espeak-ng-data.tar.bz2
+ espeak_url=https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/espeak-ng-data.tar.bz2
+ espeak_sha256="<pin-release-sha256-here>"
+ curl -fSL -o espeak-ng-data.tar.bz2 "${espeak_url}"
+ echo "${espeak_sha256} espeak-ng-data.tar.bz2" | shasum -a 256 -c -
tar xf espeak-ng-data.tar.bz2
rm espeak-ng-data.tar.bz2
mv espeak-ng-data "${model_dir}/"
fi🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@swift-api-examples/run-tts-kitten-en.sh` around lines 41 - 44, The
download/extract sequence for the espeak-ng-data.tar.bz2 archive lacks integrity
checks; add a SHA256 checksum verification step before running tar: define an
expected checksum string for "espeak-ng-data.tar.bz2", compute the actual
checksum with a portable tool (sha256sum or shasum -a 256), compare them, and if
they differ print an error and exit non‑zero (and optionally delete the corrupt
file). Only proceed to tar xf and mv into the "${model_dir}/" when the checksum
matches. Ensure the check runs immediately after the curl -O step and before
extraction so tampered or corrupted archives are rejected.
|
@csukuangfj The Swift failures were from our change: CI starts from a clean checkout, so |
csukuangfj
left a comment
There was a problem hiding this comment.
Thank you for your contribution!
…pport
v0.1.5 had the right model file on disk (verifyLayout finally happy)
but synthesis crashed mid-play. Cause: our vendored sherpa-onnx AAR
was 1.12.32 — the Kitten v0.8 model format wasn't supported until
1.13.2 ("Add KittenTTS v0.8 support" — k2-fsa/sherpa-onnx#3591,
released 2026-05-13). v0.8 ships extra ONNX metadata flags
(version=8, end_id=10, add_pad_after_end=1, max_token_len=400) that
1.12.32 doesn't read; the runtime tried v0.1's pipeline against
v0.8's model and JNI-faulted on first synthesize.
Upgraded to 1.13.2 — also the current latest tag — so we're not
shipping an artificially-floored dep. AAR grew slightly (33 MB →
36 MB on disk; APK grows ~2 MB).
Version 0.1.6 (versionCode 7).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Summary
Add KittenTTS v0.8 support across sherpa-onnx.
This includes:
Testing
python3 -m py_compile scripts/kitten-tts/v0_8/add_meta_data.py scripts/kitten-tts/v0_8/generate_tokens.py scripts/kitten-tts/v0_8/generate_voices_bin.py scripts/apk/generate-tts-apk-script.py scripts/flutter/generate-tts.pybash -n scripts/kitten-tts/v0_8/run.shgit diff --check./build-swift-macos.shtest-kitten-en.wavSummary by CodeRabbit