Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 7 additions & 0 deletions .github/scripts/test-dot-net.sh
Original file line number Diff line number Diff line change
Expand Up @@ -74,6 +74,12 @@ cd ../pocket-tts-zero-shot
ls -lh
rm -rf sherpa-onnx-pocket-*

cd ../zipvoice-tts
./run.sh
ls -lh
rm -rf sherpa-onnx-zipvoice-*
rm -f vocos_24khz.onnx

cd ../vad-non-streaming-funasr-nano
./run-ten-vad.sh
rm -fv *.onnx
Expand Down Expand Up @@ -155,6 +161,7 @@ mkdir tts
cp -v dotnet-examples/kokoro-tts/*.wav ./tts
cp -v dotnet-examples/offline-tts/*.wav ./tts
cp -v dotnet-examples/supertonic-tts/*.wav ./tts
cp -v dotnet-examples/zipvoice-tts/*.wav ./tts
popd

cd ../offline-speaker-diarization
Expand Down
1 change: 1 addition & 0 deletions .github/workflows/test-dot-net.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,7 @@ on:
- master
paths:
- '.github/workflows/test-dot-net.yaml'
- '.github/scripts/test-dot-net.sh'
- 'cmake/**'
- 'sherpa-onnx/csrc/*'
- 'dotnet-examples/**'
Expand Down
4 changes: 4 additions & 0 deletions dotnet-examples/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,10 @@ for details.
It shows how to use the online speech denoiser API with GTCRN models.
- [./streaming-speech-enhancement-dpdfnet](./streaming-speech-enhancement-dpdfnet)
It shows how to use the online speech denoiser API with DPDFNet models.
- [./zipvoice-tts](./zipvoice-tts) It shows how to use ZipVoice for
Chinese/English zero-shot text-to-speech.
- [./zipvoice-tts-play](./zipvoice-tts-play) It shows how to use ZipVoice for
Chinese/English zero-shot text-to-speech with playback.
Comment on lines +20 to +23

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The formatting of the new list items is inconsistent with other entries in the list. For better readability and consistency, the description for each item should start on a new line and be indented with two spaces.

Suggested change
- [./zipvoice-tts](./zipvoice-tts) It shows how to use ZipVoice for
Chinese/English zero-shot text-to-speech.
- [./zipvoice-tts-play](./zipvoice-tts-play) It shows how to use ZipVoice for
Chinese/English zero-shot text-to-speech with playback.
- [./zipvoice-tts](./zipvoice-tts)
It shows how to use ZipVoice for Chinese/English zero-shot text-to-speech.
- [./zipvoice-tts-play](./zipvoice-tts-play)
It shows how to use ZipVoice for Chinese/English zero-shot text-to-speech with playback.


```bash
dotnet new console -n offline-tts-play
Expand Down
12 changes: 12 additions & 0 deletions dotnet-examples/sherpa-onnx.sln
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,10 @@ Project("{FAE04EC0-301F-11D3-BF4B-00C04F79EFBC}") = "streaming-speech-enhancemen
EndProject
Project("{FAE04EC0-301F-11D3-BF4B-00C04F79EFBC}") = "streaming-speech-enhancement-dpdfnet", "streaming-speech-enhancement-dpdfnet\streaming-speech-enhancement-dpdfnet.csproj", "{8CD66C3E-3AE3-43AA-8FDA-DD5BA456F2EC}"
EndProject
Project("{FAE04EC0-301F-11D3-BF4B-00C04F79EFBC}") = "zipvoice-tts", "zipvoice-tts\zipvoice-tts.csproj", "{BBC69A08-01A7-4F89-938F-F0D551AD3F6C}"
EndProject
Project("{FAE04EC0-301F-11D3-BF4B-00C04F79EFBC}") = "zipvoice-tts-play", "zipvoice-tts-play\zipvoice-tts-play.csproj", "{84A37E18-095E-42A6-93CC-C27CD90B8478}"
EndProject
Global
GlobalSection(SolutionConfigurationPlatforms) = preSolution
Debug|Any CPU = Debug|Any CPU
Expand Down Expand Up @@ -195,6 +199,14 @@ Global
{8CD66C3E-3AE3-43AA-8FDA-DD5BA456F2EC}.Debug|Any CPU.Build.0 = Debug|Any CPU
{8CD66C3E-3AE3-43AA-8FDA-DD5BA456F2EC}.Release|Any CPU.ActiveCfg = Release|Any CPU
{8CD66C3E-3AE3-43AA-8FDA-DD5BA456F2EC}.Release|Any CPU.Build.0 = Release|Any CPU
{BBC69A08-01A7-4F89-938F-F0D551AD3F6C}.Debug|Any CPU.ActiveCfg = Debug|Any CPU
{BBC69A08-01A7-4F89-938F-F0D551AD3F6C}.Debug|Any CPU.Build.0 = Debug|Any CPU
{BBC69A08-01A7-4F89-938F-F0D551AD3F6C}.Release|Any CPU.ActiveCfg = Release|Any CPU
{BBC69A08-01A7-4F89-938F-F0D551AD3F6C}.Release|Any CPU.Build.0 = Release|Any CPU
{84A37E18-095E-42A6-93CC-C27CD90B8478}.Debug|Any CPU.ActiveCfg = Debug|Any CPU
{84A37E18-095E-42A6-93CC-C27CD90B8478}.Debug|Any CPU.Build.0 = Debug|Any CPU
{84A37E18-095E-42A6-93CC-C27CD90B8478}.Release|Any CPU.ActiveCfg = Release|Any CPU
{84A37E18-095E-42A6-93CC-C27CD90B8478}.Release|Any CPU.Build.0 = Release|Any CPU
EndGlobalSection
GlobalSection(SolutionProperties) = preSolution
HideSolutionNode = FALSE
Expand Down
193 changes: 193 additions & 0 deletions dotnet-examples/zipvoice-tts-play/Program.cs
Original file line number Diff line number Diff line change
@@ -0,0 +1,193 @@
// Copyright (c) 2026 Xiaomi Corporation
//
// This file shows how to use a non-streaming ZipVoice model
// for zero-shot text-to-speech with playback.
// Please refer to
// https://k2-fsa.github.io/sherpa/onnx/tts/zipvoice.html
// and
// https://github.com/k2-fsa/sherpa-onnx/releases/tag/tts-models
// to download pre-trained models
using PortAudioSharp;
using SherpaOnnx;
using System.Collections.Concurrent;
using System.Runtime.InteropServices;

class ZipVoiceTtsDemo
{
static void Main(string[] args)
{
TestZhEn();
}

static void TestZhEn()
{
var config = new OfflineTtsConfig();
config.Model.ZipVoice.Tokens = "./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/tokens.txt";
config.Model.ZipVoice.Encoder = "./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/encoder.int8.onnx";
config.Model.ZipVoice.Decoder = "./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/decoder.int8.onnx";
config.Model.ZipVoice.Vocoder = "./vocos_24khz.onnx";
config.Model.ZipVoice.DataDir = "./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/espeak-ng-data";
config.Model.ZipVoice.Lexicon = "./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/lexicon.txt";

config.Model.NumThreads = 2;
config.Model.Debug = 1;
config.Model.Provider = "cpu";

var referenceWaveFilename = "./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/test_wavs/leijun-1.wav";
var reader = new WaveReader(referenceWaveFilename);

OfflineTtsGenerationConfig genConfig = new OfflineTtsGenerationConfig();
genConfig.ReferenceAudio = reader.Samples;
genConfig.ReferenceSampleRate = reader.SampleRate;
genConfig.ReferenceText = "那还是三十六年前, 一九八七年. 我呢考上了武汉大学的计算机系.";
genConfig.NumSteps = 4;
genConfig.Extra["min_char_in_sentence"] = "10";

var tts = new OfflineTts(config);
var text = "小米的价值观是真诚, 热爱. 真诚,就是不欺人也不自欺. 热爱, 就是全心投入并享受其中.";

Console.WriteLine(PortAudio.VersionInfo.versionText);
PortAudio.Initialize();
Console.WriteLine($"Number of devices: {PortAudio.DeviceCount}");

for (int i = 0; i != PortAudio.DeviceCount; ++i)
{
Console.WriteLine($" Device {i}");
DeviceInfo deviceInfo = PortAudio.GetDeviceInfo(i);
Console.WriteLine($" Name: {deviceInfo.name}");
Console.WriteLine($" Max output channels: {deviceInfo.maxOutputChannels}");
Console.WriteLine($" Default sample rate: {deviceInfo.defaultSampleRate}");
}
int deviceIndex = PortAudio.DefaultOutputDevice;
if (deviceIndex == PortAudio.NoDevice)
{
Console.WriteLine("No default output device found. Please use ../zipvoice-tts instead");
Environment.Exit(1);
}

var info = PortAudio.GetDeviceInfo(deviceIndex);
Console.WriteLine();
Console.WriteLine($"Use output default device {deviceIndex} ({info.name})");

var param = new StreamParameters();
param.device = deviceIndex;
param.channelCount = 1;
param.sampleFormat = SampleFormat.Float32;
param.suggestedLatency = info.defaultLowOutputLatency;
param.hostApiSpecificStreamInfo = IntPtr.Zero;

var dataItems = new BlockingCollection<float[]>();

var myCallback = (IntPtr samples, int n, float progress, IntPtr arg) =>
{
Console.WriteLine($"Progress {progress * 100}%");

float[] data = new float[n];
Marshal.Copy(samples, data, 0, n);
dataItems.Add(data);

// 1 means to keep generating
// 0 means to stop generating
return 1;
};

var playFinished = false;

Comment on lines +94 to +95
float[]? lastSampleArray = null;
int lastIndex = 0;

PortAudioSharp.Stream.Callback playCallback = (IntPtr input, IntPtr output,
UInt32 frameCount,
ref StreamCallbackTimeInfo timeInfo,
StreamCallbackFlags statusFlags,
IntPtr userData
) =>
{
if (dataItems.IsCompleted && lastSampleArray == null && lastIndex == 0)
{
Console.WriteLine("Finished playing");
playFinished = true;
return StreamCallbackResult.Complete;
Comment on lines +94 to +110

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Verify shared flag is accessed from different execution contexts without synchronization.
rg -n -C3 '\bplayFinished\b|Callback|while \(!playFinished\)' dotnet-examples/zipvoice-tts-play/Program.cs

Repository: k2-fsa/sherpa-onnx

Length of output: 2127


🏁 Script executed:

#!/bin/bash
# Examine complete context of playFinished usage and verify threading model
cat -n dotnet-examples/zipvoice-tts-play/Program.cs | sed -n '88,195p'

Repository: k2-fsa/sherpa-onnx

Length of output: 3702


Synchronize playFinished across threads to avoid data race and potential hangs.

playFinished is written from the PortAudio callback thread (line 109) and polled on the main thread (line 188) without synchronization. The missing memory barrier allows the compiler to optimize away the poll loop or cache stale values, potentially causing indefinite hangs.

Use ManualResetEventSlim for thread-safe signaling:

Proposed synchronization fix
-    var playFinished = false;
+    using var playbackDone = new ManualResetEventSlim(false);
     
     PortAudioSharp.Stream.Callback playCallback = ...
     {
       if (dataItems.IsCompleted && lastSampleArray == null && lastIndex == 0)
       {
         Console.WriteLine("Finished playing");
-        playFinished = true;
+        playbackDone.Set();
         return StreamCallbackResult.Complete;
       }
       ...
     };
     
-    while (!playFinished)
-    {
-      Thread.Sleep(100);
-    }
+    playbackDone.Wait();
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dotnet-examples/zipvoice-tts-play/Program.cs` around lines 94 - 110, The
playFinished flag is written from the PortAudio callback thread
(PortAudioSharp.Stream.Callback playCallback) and polled on the main thread,
which causes a data race; replace the plain bool with a thread-safe waiter such
as ManualResetEventSlim (e.g., playFinishedEvent), call playFinishedEvent.Set()
inside the playCallback where you currently set playFinished = true, and replace
the main-thread busy-poll of playFinished with playFinishedEvent.Wait(timeout)
(or Wait without timeout if appropriate) so the signal uses proper memory
barriers and avoids hangs.

}

int expected = Convert.ToInt32(frameCount);
int i = 0;

while ((lastSampleArray != null || dataItems.Count != 0) && (i < expected))
{
int needed = expected - i;

if (lastSampleArray != null)
{
int remaining = lastSampleArray.Length - lastIndex;
if (remaining >= needed)
{
float[] thisBlock = lastSampleArray.Skip(lastIndex).Take(needed).ToArray();
lastIndex += needed;
Comment on lines +120 to +126
if (lastIndex == lastSampleArray.Length)
{
lastSampleArray = null;
lastIndex = 0;
}

Marshal.Copy(thisBlock, 0, IntPtr.Add(output, i * sizeof(float)), needed);
return StreamCallbackResult.Continue;
Comment on lines +125 to +134

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The use of LINQ (Skip, Take, ToArray) inside this audio callback is inefficient as it allocates a new array (thisBlock) on each call. This can lead to performance issues and audio glitches, especially in a real-time audio thread. You can directly use Marshal.Copy with an offset on lastSampleArray to avoid this allocation.

            Marshal.Copy(lastSampleArray, lastIndex, IntPtr.Add(output, i * sizeof(float)), needed);
            lastIndex += needed;
            if (lastIndex == lastSampleArray.Length)
            {
              lastSampleArray = null;
              lastIndex = 0;
            }

            return StreamCallbackResult.Continue;

}

float[] thisBlock2 = lastSampleArray.Skip(lastIndex).Take(remaining).ToArray();
lastIndex = 0;
lastSampleArray = null;

Marshal.Copy(thisBlock2, 0, IntPtr.Add(output, i * sizeof(float)), remaining);
Comment on lines +137 to +141

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Similar to the previous comment, using Skip().Take().ToArray() here creates an unnecessary intermediate array. This is inefficient for an audio callback. You can directly copy from lastSampleArray to the output buffer.

          Marshal.Copy(lastSampleArray, lastIndex, IntPtr.Add(output, i * sizeof(float)), remaining);

i += remaining;
continue;
}

if (dataItems.Count != 0)
{
lastSampleArray = dataItems.Take();
lastIndex = 0;
}
}

if (i < expected)
{
int sizeInBytes = (expected - i) * 4;
Marshal.Copy(new byte[sizeInBytes], 0, IntPtr.Add(output, i * sizeof(float)), sizeInBytes);
}
Comment on lines +153 to +157

return StreamCallbackResult.Continue;
};

PortAudioSharp.Stream stream = new PortAudioSharp.Stream(inParams: null, outParams: param, sampleRate: tts.SampleRate,
framesPerBuffer: 0,
streamFlags: StreamFlags.ClipOff,
callback: playCallback,
userData: IntPtr.Zero
);

stream.Start();

var callback = new OfflineTtsCallbackProgressWithArg(myCallback);
var audio = tts.GenerateWithConfig(text, genConfig, callback);

var outputFilename = "./generated-zipvoice-zh-en-play.wav";
var ok = audio.SaveToWaveFile(outputFilename);

if (ok)
{
Console.WriteLine($"Wrote to {outputFilename} succeeded!");
}
else
{
Console.WriteLine($"Failed to write {outputFilename}");
}

dataItems.CompleteAdding();

Comment on lines +169 to +187
while (!playFinished)
{
Thread.Sleep(100);
}
Comment on lines +188 to +191

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The while (!playFinished) loop with Thread.Sleep(100) is a form of busy-waiting, which consumes CPU resources unnecessarily. A more efficient approach is to use a synchronization primitive like ManualResetEvent to wait for the playback to finish.

For example, you could use a static ManualResetEvent, call .Set() in the audio callback when playback is finished, and call .WaitOne() here to block the main thread until the event is signaled.

}
}
14 changes: 14 additions & 0 deletions dotnet-examples/zipvoice-tts-play/run.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
#!/usr/bin/env bash
set -ex

if [ ! -f ./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/encoder.int8.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
tar xvf sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
rm sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
fi

if [ ! -f ./vocos_24khz.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
fi
Comment on lines +5 to +12

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Use fail-fast curl flags here as well.

Line 5 and Line 11 should fail on HTTP errors to avoid continuing with bad artifacts.

🛠️ Proposed change
-  curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
+  curl -fSL --retry 3 --retry-delay 2 -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
@@
-  curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
+  curl -fSL --retry 3 --retry-delay 2 -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
tar xvf sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
rm sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
fi
if [ ! -f ./vocos_24khz.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
fi
curl -fSL --retry 3 --retry-delay 2 -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
tar xvf sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
rm sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
fi
if [ ! -f ./vocos_24khz.onnx ]; then
curl -fSL --retry 3 --retry-delay 2 -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
fi
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dotnet-examples/zipvoice-tts-play/run.sh` around lines 5 - 12, The curl
invocations in run.sh (the commands that download
sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2 and vocos_24khz.onnx)
don’t fail on HTTP errors; update those curl commands to use fail-fast flags
(e.g., add -f or --fail alongside -SL -O) so the script exits on HTTP failures
and doesn’t continue with missing or incomplete artifacts.


dotnet run
19 changes: 19 additions & 0 deletions dotnet-examples/zipvoice-tts-play/zipvoice-tts-play.csproj
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
<Project Sdk="Microsoft.NET.Sdk">

<PropertyGroup>
<OutputType>Exe</OutputType>
<TargetFramework>net8.0</TargetFramework>
<RootNamespace>zipvoice_tts_play</RootNamespace>
<ImplicitUsings>enable</ImplicitUsings>
<Nullable>enable</Nullable>
</PropertyGroup>

<ItemGroup>
<PackageReference Include="PortAudioSharp2" Version="*" />

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/usr/bin/env bash
# Verify wildcard PackageReference usage across the repo
rg -n --glob '*.csproj' '<PackageReference[^>]*Version="\*"'

Repository: k2-fsa/sherpa-onnx

Length of output: 1232


🌐 Web query:

What is the latest stable NuGet version of PortAudioSharp2, and is using Version="*" in SDK-style .csproj recommended for reproducible builds?

💡 Result:

The latest stable NuGet version of PortAudioSharp2 is 1.0.6, released on October 16, 2025. Using Version="" (a floating version) in SDK-style .csproj files is not recommended for reproducible builds. Floating versions like "" resolve to the highest available version matching the pattern at restore time, which can lead to different package versions across restores, environments, or after new versions are published. This breaks reproducibility, as the same project may build with varying dependencies. For reproducible builds, pin to specific versions (e.g., Version="1.0.6") or use lock files (RestorePackagesWithLockFile) with --locked-mode in CI. Official guidance emphasizes lowest applicable versions without upper bounds but avoids floating versions for consistency. Enable Central Package Management or Dependabot for updates instead.

Citations:


Pin PortAudioSharp2 to a specific version instead of using a wildcard.

Line 12 uses Version="*", which resolves to the latest available version at restore time and makes builds non-reproducible. This can cause different package versions across environments and unexpected breakage. The latest stable version is 1.0.6—pin to this version explicitly.

🔒 Suggested fix
-    <PackageReference Include="PortAudioSharp2" Version="*" />
+    <PackageReference Include="PortAudioSharp2" Version="1.0.6" />

Note: This issue appears in multiple project files across the repository (kokoro-tts-play.csproj, kitten-tts-play.csproj, speech-recognition-from-microphone.csproj, etc.). Consider applying the fix consistently across all affected files.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
<PackageReference Include="PortAudioSharp2" Version="*" />
<PackageReference Include="PortAudioSharp2" Version="1.0.6" />
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dotnet-examples/zipvoice-tts-play/zipvoice-tts-play.csproj` at line 12, The
PackageReference for PortAudioSharp2 currently uses a wildcard Version="*",
making builds non-reproducible; update the PortAudioSharp2 PackageReference (the
XML element with Include="PortAudioSharp2") to pin Version="1.0.6" instead of
"*" and apply the same change to other project files that reference
PortAudioSharp2 (e.g., kokoro-tts-play.csproj, kitten-tts-play.csproj,
speech-recognition-from-microphone.csproj) so all projects use the explicit
1.0.6 version.

</ItemGroup>

<ItemGroup>
<ProjectReference Include="..\Common\Common.csproj" />
</ItemGroup>

</Project>
74 changes: 74 additions & 0 deletions dotnet-examples/zipvoice-tts/Program.cs
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
// Copyright (c) 2026 Xiaomi Corporation
//
// This file shows how to use a non-streaming ZipVoice model
// for zero-shot text-to-speech.
// Please refer to
// https://k2-fsa.github.io/sherpa/onnx/tts/zipvoice.html
// and
// https://github.com/k2-fsa/sherpa-onnx/releases/tag/tts-models
// to download pre-trained models
using SherpaOnnx;
using System.Runtime.InteropServices;

class ZipVoiceTtsDemo
{
static void Main(string[] args)
{
TestZhEn();
}

static void TestZhEn()
{
var config = new OfflineTtsConfig();
config.Model.ZipVoice.Tokens = "./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/tokens.txt";
config.Model.ZipVoice.Encoder = "./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/encoder.int8.onnx";
config.Model.ZipVoice.Decoder = "./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/decoder.int8.onnx";
config.Model.ZipVoice.Vocoder = "./vocos_24khz.onnx";
config.Model.ZipVoice.DataDir = "./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/espeak-ng-data";
config.Model.ZipVoice.Lexicon = "./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/lexicon.txt";

config.Model.NumThreads = 2;
config.Model.Debug = 1;
config.Model.Provider = "cpu";

var referenceWaveFilename = "./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/test_wavs/leijun-1.wav";
var reader = new WaveReader(referenceWaveFilename);

OfflineTtsGenerationConfig genConfig = new OfflineTtsGenerationConfig();
genConfig.ReferenceAudio = reader.Samples;
genConfig.ReferenceSampleRate = reader.SampleRate;
genConfig.ReferenceText = "那还是三十六年前, 一九八七年. 我呢考上了武汉大学的计算机系.";
genConfig.NumSteps = 4;
genConfig.Extra["min_char_in_sentence"] = "10";

var tts = new OfflineTts(config);
var text = "小米的价值观是真诚, 热爱. 真诚,就是不欺人也不自欺. 热爱, 就是全心投入并享受其中.";

var myCallback = (IntPtr samples, int n, float progress, IntPtr arg) =>
{
float[] data = new float[n];
Marshal.Copy(samples, data, 0, n);
Comment on lines +49 to +50

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The data array is created and populated from the samples pointer, but it's a local variable that is never used. This is unnecessary work and memory allocation inside the callback. These lines can be removed.

Comment on lines +49 to +50
Console.WriteLine($"Progress {progress * 100}%");

// 1 means to keep generating
// 0 means to stop generating
return 1;
};

var callback = new OfflineTtsCallbackProgressWithArg(myCallback);

var audio = tts.GenerateWithConfig(text, genConfig, callback);

var outputFilename = "./generated-zipvoice-zh-en.wav";
var ok = audio.SaveToWaveFile(outputFilename);

if (ok)
{
Console.WriteLine($"Wrote to {outputFilename} succeeded!");
}
else
{
Console.WriteLine($"Failed to write {outputFilename}");
}
}
}
14 changes: 14 additions & 0 deletions dotnet-examples/zipvoice-tts/run.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,14 @@
#!/usr/bin/env bash
set -ex

if [ ! -f ./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/encoder.int8.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
tar xvf sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
rm sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
fi

if [ ! -f ./vocos_24khz.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
fi
Comment on lines +5 to +12

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Use fail-fast curl flags for reliable downloads.

Line 5 and Line 11 can succeed on HTTP errors without --fail, producing confusing downstream failures.

🛠️ Proposed change
-  curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
+  curl -fSL --retry 3 --retry-delay 2 -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
@@
-  curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
+  curl -fSL --retry 3 --retry-delay 2 -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dotnet-examples/zipvoice-tts/run.sh` around lines 5 - 12, The two curl
invocations that download the model files (the lines calling "curl -SL -O
https://...sherpa-onnx-zipvoice-...tar.bz2" and "curl -SL -O
https://...vocos_24khz.onnx") should be made fail-fast; update those commands to
include curl's fail and error reporting flags (e.g., add --fail/-f and
--show-error, and keep -S -L -O) so HTTP errors cause the script to exit
immediately and surface the underlying error when "curl -SL -O ..." is executed.


dotnet run
Comment on lines +1 to +14

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This script is identical to dotnet-examples/zipvoice-tts-play/run.sh. To avoid code duplication and improve maintainability, consider extracting the model download logic into a common script (e.g., in a shared parent directory) and sourcing it from both run.sh files. This aligns with the Don't Repeat Yourself (DRY) principle.

Loading
Loading