Skip to content

Add C# API examples for ZipVoice models - #3345

Merged
csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:dotnet-zipvoice
Mar 18, 2026
Merged

csukuangfj merged 1 commit into
k2-fsa:masterfrom
csukuangfj:dotnet-zipvoice

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Mar 18, 2026 •

Copy link
Copy Markdown
Collaborator

Summary by CodeRabbit

  • New Features

    • Added ZipVoice text-to-speech examples demonstrating zero-shot synthesis for Chinese and English
    • Added ZipVoice TTS with real-time audio playback and output device selection
    • Integrated ZipVoice TTS into the automated test workflow
  • Documentation

    • Updated README with links to new ZipVoice TTS examples

@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Mar 18, 2026
@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request expands the .NET examples by adding C# API demonstrations for the ZipVoice text-to-speech models. It provides two distinct examples: one for generating speech and saving it to a file, and another for generating speech and playing it back in real-time. These additions aim to provide clear, functional examples for developers looking to integrate ZipVoice TTS capabilities into their C# applications, enhancing the accessibility and usability of the sherpa-onnx library within the .NET ecosystem.

Highlights

  • New C# API Examples: Introduced two new C# API examples for ZipVoice text-to-speech (TTS) models: one for basic TTS and another for TTS with real-time audio playback.
  • Build System Integration: Integrated the new ZipVoice TTS examples into the .NET build system by updating the solution file and the automated testing script.
  • Documentation Update: Updated the dotnet-examples/README.md to include clear documentation and links for the newly added ZipVoice TTS examples.
Ignored Files
  • Ignored by pattern: .github/workflows/** (1)
    • .github/workflows/test-dot-net.yaml
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 18, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This pull request introduces ZipVoice TTS integration to the .NET examples codebase. Two new example projects are added—one demonstrating non-streaming text-to-speech conversion and another with real-time audio playback. Supporting infrastructure includes run scripts for model downloading, project configurations, documentation updates, and CI workflow modifications.

Changes

Cohort / File(s) Summary
CI/Test Infrastructure
.github/scripts/test-dot-net.sh, .github/workflows/test-dot-net.yaml
Adds ZipVoice TTS workflow execution block with model cleanup (sherpa-onnx-zipvoice-*, vocos_24khz.onnx) and extends workflow trigger to watch the test script path.
Documentation
dotnet-examples/README.md
Adds two documentation links for ZipVoice TTS examples: standard TTS and TTS with audio playback.
Solution Configuration
dotnet-examples/sherpa-onnx.sln
Registers two new ZipVoice projects (zipvoice-tts, zipvoice-tts-play) with GUIDs and Debug/Release build configurations.
ZipVoice TTS (Basic)
dotnet-examples/zipvoice-tts/Program.cs, dotnet-examples/zipvoice-tts/run.sh, dotnet-examples/zipvoice-tts/zipvoice-tts.csproj
Demonstrates non-streaming ZipVoice TTS with Chinese text input, automatic model download/extraction logic, and .NET project structure targeting net8.0.
ZipVoice TTS (Playback)
dotnet-examples/zipvoice-tts-play/Program.cs, dotnet-examples/zipvoice-tts-play/run.sh, dotnet-examples/zipvoice-tts-play/zipvoice-tts-play.csproj
Demonstrates non-streaming ZipVoice TTS with real-time PortAudio-based playback via streaming audio callbacks, includes model download logic and PortAudioSharp2 dependency.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~20 minutes

Possibly related PRs

Suggested labels

size:M

Poem

🐰 Whisker-twitching with joy, a rabbit hops by,
Two TTS friends now speak—no silence, oh my!
One whispers to files, one sings through the air,
ZipVoice echoes sweetly in Chinese and fair,
Models download, callbacks stream, playback takes flight—
Our .NET voice toolkit shines ever more bright! 🎵

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately describes the main change: adding C# API examples for ZipVoice models, which is evident from the new example programs, project files, and documentation additions across the changeset.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
📝 Coding Plan
  • Generate coding plan for human review comments

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds C# examples for ZipVoice models, including both a simple generation-to-file example and a more complex example with real-time playback. The changes are generally good, but I've identified several areas for improvement, particularly concerning performance in the audio playback callback, code duplication in shell scripts, and minor issues in the C# code and documentation formatting. Addressing these points will enhance the quality and maintainability of the new examples.

Comment on lines +125 to +134
float[] thisBlock = lastSampleArray.Skip(lastIndex).Take(needed).ToArray();
lastIndex += needed;
if (lastIndex == lastSampleArray.Length)
{
lastSampleArray = null;
lastIndex = 0;
}

Marshal.Copy(thisBlock, 0, IntPtr.Add(output, i * sizeof(float)), needed);
return StreamCallbackResult.Continue;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The use of LINQ (Skip, Take, ToArray) inside this audio callback is inefficient as it allocates a new array (thisBlock) on each call. This can lead to performance issues and audio glitches, especially in a real-time audio thread. You can directly use Marshal.Copy with an offset on lastSampleArray to avoid this allocation.

            Marshal.Copy(lastSampleArray, lastIndex, IntPtr.Add(output, i * sizeof(float)), needed);
            lastIndex += needed;
            if (lastIndex == lastSampleArray.Length)
            {
              lastSampleArray = null;
              lastIndex = 0;
            }

            return StreamCallbackResult.Continue;

Comment on lines +137 to +141
float[] thisBlock2 = lastSampleArray.Skip(lastIndex).Take(remaining).ToArray();
lastIndex = 0;
lastSampleArray = null;

Marshal.Copy(thisBlock2, 0, IntPtr.Add(output, i * sizeof(float)), remaining);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Similar to the previous comment, using Skip().Take().ToArray() here creates an unnecessary intermediate array. This is inefficient for an audio callback. You can directly copy from lastSampleArray to the output buffer.

          Marshal.Copy(lastSampleArray, lastIndex, IntPtr.Add(output, i * sizeof(float)), remaining);

Comment thread dotnet-examples/README.md
Comment on lines +20 to +23
- [./zipvoice-tts](./zipvoice-tts) It shows how to use ZipVoice for
Chinese/English zero-shot text-to-speech.
- [./zipvoice-tts-play](./zipvoice-tts-play) It shows how to use ZipVoice for
Chinese/English zero-shot text-to-speech with playback.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The formatting of the new list items is inconsistent with other entries in the list. For better readability and consistency, the description for each item should start on a new line and be indented with two spaces.

Suggested change
- [./zipvoice-tts](./zipvoice-tts) It shows how to use ZipVoice for
Chinese/English zero-shot text-to-speech.
- [./zipvoice-tts-play](./zipvoice-tts-play) It shows how to use ZipVoice for
Chinese/English zero-shot text-to-speech with playback.
- [./zipvoice-tts](./zipvoice-tts)
It shows how to use ZipVoice for Chinese/English zero-shot text-to-speech.
- [./zipvoice-tts-play](./zipvoice-tts-play)
It shows how to use ZipVoice for Chinese/English zero-shot text-to-speech with playback.

Comment on lines +188 to +191
while (!playFinished)
{
Thread.Sleep(100);
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The while (!playFinished) loop with Thread.Sleep(100) is a form of busy-waiting, which consumes CPU resources unnecessarily. A more efficient approach is to use a synchronization primitive like ManualResetEvent to wait for the playback to finish.

For example, you could use a static ManualResetEvent, call .Set() in the audio callback when playback is finished, and call .WaitOne() here to block the main thread until the event is signaled.

Comment on lines +49 to +50
float[] data = new float[n];
Marshal.Copy(samples, data, 0, n);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The data array is created and populated from the samples pointer, but it's a local variable that is never used. This is unnecessary work and memory allocation inside the callback. These lines can be removed.

Comment on lines +1 to +14
#!/usr/bin/env bash
set -ex

if [ ! -f ./sherpa-onnx-zipvoice-distill-int8-zh-en-emilia/encoder.int8.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
tar xvf sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
rm sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
fi

if [ ! -f ./vocos_24khz.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
fi

dotnet run

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This script is identical to dotnet-examples/zipvoice-tts-play/run.sh. To avoid code duplication and improve maintainability, consider extracting the model download logic into a common script (e.g., in a shared parent directory) and sourcing it from both run.sh files. This aligns with the Don't Repeat Yourself (DRY) principle.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds new C#/.NET 8 ZipVoice zero-shot TTS examples (with and without PortAudio playback) to the dotnet-examples solution, and wires the non-playback demo into the GitHub Actions .NET test script so it runs in CI.

Changes:

  • Added zipvoice-tts (generate WAV) and zipvoice-tts-play (generate + playback) example projects and run scripts.
  • Registered the new projects in dotnet-examples/sherpa-onnx.sln and linked them from dotnet-examples/README.md.
  • Updated CI to trigger on .github/scripts/test-dot-net.sh changes and to run the new zipvoice-tts example, archiving its generated WAV.

Reviewed changes

Copilot reviewed 10 out of 10 changed files in this pull request and generated 6 comments.

Show a summary per file
File Description
dotnet-examples/zipvoice-tts/zipvoice-tts.csproj New .NET 8 console project for ZipVoice TTS WAV generation.
dotnet-examples/zipvoice-tts/run.sh Downloads ZipVoice + vocoder models and runs the demo.
dotnet-examples/zipvoice-tts/Program.cs Implements ZipVoice zero-shot TTS generation demo.
dotnet-examples/zipvoice-tts-play/zipvoice-tts-play.csproj New .NET 8 console project for ZipVoice TTS with PortAudio playback.
dotnet-examples/zipvoice-tts-play/run.sh Downloads models and runs the playback demo.
dotnet-examples/zipvoice-tts-play/Program.cs Implements streaming playback via PortAudio callback while generating.
dotnet-examples/sherpa-onnx.sln Adds the two new example projects to the solution.
dotnet-examples/README.md Documents the two new ZipVoice examples.
.github/workflows/test-dot-net.yaml Ensures workflow triggers on changes to the test script.
.github/scripts/test-dot-net.sh Runs zipvoice-tts in CI and copies its output WAV to artifacts.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment on lines +49 to +50
float[] data = new float[n];
Marshal.Copy(samples, data, 0, n);
param.suggestedLatency = info.defaultLowOutputLatency;
param.hostApiSpecificStreamInfo = IntPtr.Zero;

var dataItems = new BlockingCollection<float[]>();
Comment on lines +94 to +95
var playFinished = false;

Comment on lines +120 to +126
if (lastSampleArray != null)
{
int remaining = lastSampleArray.Length - lastIndex;
if (remaining >= needed)
{
float[] thisBlock = lastSampleArray.Skip(lastIndex).Take(needed).ToArray();
lastIndex += needed;
Comment on lines +153 to +157
if (i < expected)
{
int sizeInBytes = (expected - i) * 4;
Marshal.Copy(new byte[sizeInBytes], 0, IntPtr.Add(output, i * sizeof(float)), sizeInBytes);
}
Comment on lines +169 to +187
stream.Start();

var callback = new OfflineTtsCallbackProgressWithArg(myCallback);
var audio = tts.GenerateWithConfig(text, genConfig, callback);

var outputFilename = "./generated-zipvoice-zh-en-play.wav";
var ok = audio.SaveToWaveFile(outputFilename);

if (ok)
{
Console.WriteLine($"Wrote to {outputFilename} succeeded!");
}
else
{
Console.WriteLine($"Failed to write {outputFilename}");
}

dataItems.CompleteAdding();

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🧹 Nitpick comments (5)
dotnet-examples/zipvoice-tts/Program.cs (1)

47-51: Remove unused audio copy in the progress callback.

Line 49 and Line 50 allocate/copy samples that are never consumed, adding avoidable overhead during generation.

♻️ Proposed change
     var myCallback = (IntPtr samples, int n, float progress, IntPtr arg) =>
     {
-      float[] data = new float[n];
-      Marshal.Copy(samples, data, 0, n);
       Console.WriteLine($"Progress {progress * 100}%");

       // 1 means to keep generating
       // 0 means to stop generating
       return 1;
     };
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dotnet-examples/zipvoice-tts/Program.cs` around lines 47 - 51, The progress
callback myCallback currently allocates and copies audio data with "float[] data
= new float[n]; Marshal.Copy(samples, data, 0, n);" but never uses it, creating
unnecessary overhead; remove those two lines from the myCallback lambda and keep
only the progress logging (Console.WriteLine($"Progress {progress * 100}%")) so
the callback still receives the parameters but doesn't allocate or copy unused
audio buffers.
dotnet-examples/zipvoice-tts-play/run.sh (1)

1-3: Harden script execution path resolution.

Relative model paths and dotnet run should not depend on where the script is launched from.

♻️ Proposed change
 #!/usr/bin/env bash
 set -ex
+
+script_dir="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
+cd "${script_dir}"
@@
 dotnet run

Also applies to: 14-14

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dotnet-examples/zipvoice-tts-play/run.sh` around lines 1 - 3, The script
run.sh currently relies on the caller's CWD for relative model paths and the
dotnet run invocation; update it to determine its own directory (use BASH_SOURCE
logic) and convert relative model paths to absolute ones based on that
directory, then invoke dotnet run with an explicit --project or by cd'ing to the
script directory so the project and models are resolved consistently; update any
references to model files and the dotnet run call in the script to use the
computed script directory variable.
dotnet-examples/zipvoice-tts/run.sh (1)

1-3: Make execution independent of caller working directory.

Current relative paths and dotnet run assume invocation from this folder.

♻️ Proposed change
 #!/usr/bin/env bash
 set -ex
+
+script_dir="$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)"
+cd "${script_dir}"
@@
 dotnet run

Also applies to: 14-14

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dotnet-examples/zipvoice-tts/run.sh` around lines 1 - 3, The script assumes
the caller's CWD; make it directory-independent by resolving the script
directory via BASH_SOURCE and switching there before running dotnet commands. In
run.sh compute SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" and
either cd "$SCRIPT_DIR" or use "$SCRIPT_DIR" to build absolute paths for any
subsequent dotnet run invocations (or --project paths), so references in run.sh
work regardless of the caller's working directory.
dotnet-examples/zipvoice-tts-play/Program.cs (2)

99-159: Avoid allocations and blocking work inside the PortAudio callback.

The callback currently allocates via LINQ (Skip/Take/ToArray on lines 125 and 137) and can block on queue access (dataItems.Take() on line 148). Additionally, Console.WriteLine on line 108 performs I/O in a real-time audio callback. These operations can cause audio underruns/glitches.

Replace with non-blocking queue access and direct buffer copying:

  • Use TryTake() instead of Take() to avoid blocking when the queue is empty
  • Use direct Marshal.Copy() with pointer arithmetic instead of LINQ allocations
  • Remove I/O operations from the callback path
♻️ Proposed callback-safe copy path
-      while ((lastSampleArray != null || dataItems.Count != 0) && (i < expected))
+      while (i < expected)
       {
-        int needed = expected - i;
-
-        if (lastSampleArray != null)
-        {
-          int remaining = lastSampleArray.Length - lastIndex;
-          if (remaining >= needed)
-          {
-            float[] thisBlock = lastSampleArray.Skip(lastIndex).Take(needed).ToArray();
-            lastIndex += needed;
-            if (lastIndex == lastSampleArray.Length)
-            {
-              lastSampleArray = null;
-              lastIndex = 0;
-            }
-
-            Marshal.Copy(thisBlock, 0, IntPtr.Add(output, i * sizeof(float)), needed);
-            return StreamCallbackResult.Continue;
-          }
-
-          float[] thisBlock2 = lastSampleArray.Skip(lastIndex).Take(remaining).ToArray();
-          lastIndex = 0;
-          lastSampleArray = null;
-
-          Marshal.Copy(thisBlock2, 0, IntPtr.Add(output, i * sizeof(float)), remaining);
-          i += remaining;
-          continue;
-        }
-
-        if (dataItems.Count != 0)
-        {
-          lastSampleArray = dataItems.Take();
-          lastIndex = 0;
-        }
+        if (lastSampleArray is null)
+        {
+          if (!dataItems.TryTake(out lastSampleArray))
+          {
+            break;
+          }
+          lastIndex = 0;
+        }
+
+        int remaining = lastSampleArray.Length - lastIndex;
+        int toCopy = Math.Min(remaining, expected - i);
+        Marshal.Copy(lastSampleArray, lastIndex, IntPtr.Add(output, i * sizeof(float)), toCopy);
+        lastIndex += toCopy;
+        i += toCopy;
+
+        if (lastIndex == lastSampleArray.Length)
+        {
+          lastSampleArray = null;
+          lastIndex = 0;
+        }
       }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dotnet-examples/zipvoice-tts-play/Program.cs` around lines 99 - 159, The
PortAudio callback playCallback performs blocking and allocating work
(Console.WriteLine, LINQ Skip/Take/ToArray and dataItems.Take()) which can cause
glitches; replace these with real-time safe operations: remove the
Console.WriteLine and the playFinished I/O, use dataItems.TryTake(out var
buffer) instead of Take() to avoid blocking, and stop using LINQ—copy directly
from lastSampleArray using offsets (e.g., call Marshal.Copy(lastSampleArray,
lastIndex, IntPtr.Add(output, i * sizeof(float)), count)) updating lastIndex and
nulling lastSampleArray when consumed, and when no data is available fill the
remainder with zeros via a preallocated zero buffer (avoid new byte[]
allocations in the callback); ensure all branches return StreamCallbackResult
appropriately without any blocking or heap allocations.

50-51: Add explicit PortAudio/stream teardown for deterministic resource cleanup.

The code initializes PortAudio at line 50 and starts a stream at line 169 but does not explicitly stop/dispose the stream or terminate PortAudio. In repeated runs this can leak unmanaged audio resources. This pattern affects all PortAudio examples in the codebase.

Wrap the playback logic in a try/finally block to ensure cleanup:

🧹 Proposed lifecycle cleanup via try/finally
-    PortAudio.Initialize();
+    PortAudio.Initialize();
@@
-    PortAudioSharp.Stream stream = new PortAudioSharp.Stream(inParams: null, outParams: param, sampleRate: tts.SampleRate,
+    PortAudioSharp.Stream stream = new PortAudioSharp.Stream(inParams: null, outParams: param, sampleRate: tts.SampleRate,
         framesPerBuffer: 0,
         streamFlags: StreamFlags.ClipOff,
         callback: playCallback,
         userData: IntPtr.Zero
         );
-
-    stream.Start();
-
-    var callback = new OfflineTtsCallbackProgressWithArg(myCallback);
-    var audio = tts.GenerateWithConfig(text, genConfig, callback);
+    try
+    {
+      stream.Start();
+      var callback = new OfflineTtsCallbackProgressWithArg(myCallback);
+      var audio = tts.GenerateWithConfig(text, genConfig, callback);
@@
-    dataItems.CompleteAdding();
-
-    while (!playFinished)
-    {
-      Thread.Sleep(100);
-    }
+      dataItems.CompleteAdding();
+      while (!playFinished)
+      {
+        Thread.Sleep(100);
+      }
+    }
+    finally
+    {
+      stream.Stop();
+      stream.Dispose();
+      PortAudio.Terminate();
+    }

Also applies to: 162-170, 186-192

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dotnet-examples/zipvoice-tts-play/Program.cs` around lines 50 - 51, The
PortAudio initialization (PortAudio.Initialize()) currently lacks deterministic
teardown; wrap the playback logic that creates/opens/starts the PortAudio stream
(the variable named stream and any related playback code around where the stream
is started) in a try/finally so that in the finally you explicitly call
stream.Stop() (if running), stream.Dispose()/Close() (or the stream's proper
dispose method) and PortAudio.Terminate(); ensure these cleanup calls run even
on exceptions to avoid leaking unmanaged audio resources.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@dotnet-examples/zipvoice-tts-play/Program.cs`:
- Around line 94-110: The playFinished flag is written from the PortAudio
callback thread (PortAudioSharp.Stream.Callback playCallback) and polled on the
main thread, which causes a data race; replace the plain bool with a thread-safe
waiter such as ManualResetEventSlim (e.g., playFinishedEvent), call
playFinishedEvent.Set() inside the playCallback where you currently set
playFinished = true, and replace the main-thread busy-poll of playFinished with
playFinishedEvent.Wait(timeout) (or Wait without timeout if appropriate) so the
signal uses proper memory barriers and avoids hangs.

In `@dotnet-examples/zipvoice-tts-play/run.sh`:
- Around line 5-12: The curl invocations in run.sh (the commands that download
sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2 and vocos_24khz.onnx)
don’t fail on HTTP errors; update those curl commands to use fail-fast flags
(e.g., add -f or --fail alongside -SL -O) so the script exits on HTTP failures
and doesn’t continue with missing or incomplete artifacts.

In `@dotnet-examples/zipvoice-tts-play/zipvoice-tts-play.csproj`:
- Line 12: The PackageReference for PortAudioSharp2 currently uses a wildcard
Version="*", making builds non-reproducible; update the PortAudioSharp2
PackageReference (the XML element with Include="PortAudioSharp2") to pin
Version="1.0.6" instead of "*" and apply the same change to other project files
that reference PortAudioSharp2 (e.g., kokoro-tts-play.csproj,
kitten-tts-play.csproj, speech-recognition-from-microphone.csproj) so all
projects use the explicit 1.0.6 version.

In `@dotnet-examples/zipvoice-tts/run.sh`:
- Around line 5-12: The two curl invocations that download the model files (the
lines calling "curl -SL -O https://...sherpa-onnx-zipvoice-...tar.bz2" and "curl
-SL -O https://...vocos_24khz.onnx") should be made fail-fast; update those
commands to include curl's fail and error reporting flags (e.g., add --fail/-f
and --show-error, and keep -S -L -O) so HTTP errors cause the script to exit
immediately and surface the underlying error when "curl -SL -O ..." is executed.

---

Nitpick comments:
In `@dotnet-examples/zipvoice-tts-play/Program.cs`:
- Around line 99-159: The PortAudio callback playCallback performs blocking and
allocating work (Console.WriteLine, LINQ Skip/Take/ToArray and dataItems.Take())
which can cause glitches; replace these with real-time safe operations: remove
the Console.WriteLine and the playFinished I/O, use dataItems.TryTake(out var
buffer) instead of Take() to avoid blocking, and stop using LINQ—copy directly
from lastSampleArray using offsets (e.g., call Marshal.Copy(lastSampleArray,
lastIndex, IntPtr.Add(output, i * sizeof(float)), count)) updating lastIndex and
nulling lastSampleArray when consumed, and when no data is available fill the
remainder with zeros via a preallocated zero buffer (avoid new byte[]
allocations in the callback); ensure all branches return StreamCallbackResult
appropriately without any blocking or heap allocations.
- Around line 50-51: The PortAudio initialization (PortAudio.Initialize())
currently lacks deterministic teardown; wrap the playback logic that
creates/opens/starts the PortAudio stream (the variable named stream and any
related playback code around where the stream is started) in a try/finally so
that in the finally you explicitly call stream.Stop() (if running),
stream.Dispose()/Close() (or the stream's proper dispose method) and
PortAudio.Terminate(); ensure these cleanup calls run even on exceptions to
avoid leaking unmanaged audio resources.

In `@dotnet-examples/zipvoice-tts-play/run.sh`:
- Around line 1-3: The script run.sh currently relies on the caller's CWD for
relative model paths and the dotnet run invocation; update it to determine its
own directory (use BASH_SOURCE logic) and convert relative model paths to
absolute ones based on that directory, then invoke dotnet run with an explicit
--project or by cd'ing to the script directory so the project and models are
resolved consistently; update any references to model files and the dotnet run
call in the script to use the computed script directory variable.

In `@dotnet-examples/zipvoice-tts/Program.cs`:
- Around line 47-51: The progress callback myCallback currently allocates and
copies audio data with "float[] data = new float[n]; Marshal.Copy(samples, data,
0, n);" but never uses it, creating unnecessary overhead; remove those two lines
from the myCallback lambda and keep only the progress logging
(Console.WriteLine($"Progress {progress * 100}%")) so the callback still
receives the parameters but doesn't allocate or copy unused audio buffers.

In `@dotnet-examples/zipvoice-tts/run.sh`:
- Around line 1-3: The script assumes the caller's CWD; make it
directory-independent by resolving the script directory via BASH_SOURCE and
switching there before running dotnet commands. In run.sh compute
SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)" and either cd
"$SCRIPT_DIR" or use "$SCRIPT_DIR" to build absolute paths for any subsequent
dotnet run invocations (or --project paths), so references in run.sh work
regardless of the caller's working directory.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 79cbc52c-d13f-40d1-b538-ddfe4deccea2

📥 Commits

Reviewing files that changed from the base of the PR and between 286d693 and 618a2f6.

📒 Files selected for processing (10)
  • .github/scripts/test-dot-net.sh
  • .github/workflows/test-dot-net.yaml
  • dotnet-examples/README.md
  • dotnet-examples/sherpa-onnx.sln
  • dotnet-examples/zipvoice-tts-play/Program.cs
  • dotnet-examples/zipvoice-tts-play/run.sh
  • dotnet-examples/zipvoice-tts-play/zipvoice-tts-play.csproj
  • dotnet-examples/zipvoice-tts/Program.cs
  • dotnet-examples/zipvoice-tts/run.sh
  • dotnet-examples/zipvoice-tts/zipvoice-tts.csproj

Comment on lines +94 to +110
var playFinished = false;

float[]? lastSampleArray = null;
int lastIndex = 0;

PortAudioSharp.Stream.Callback playCallback = (IntPtr input, IntPtr output,
UInt32 frameCount,
ref StreamCallbackTimeInfo timeInfo,
StreamCallbackFlags statusFlags,
IntPtr userData
) =>
{
if (dataItems.IsCompleted && lastSampleArray == null && lastIndex == 0)
{
Console.WriteLine("Finished playing");
playFinished = true;
return StreamCallbackResult.Complete;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/bin/bash
# Verify shared flag is accessed from different execution contexts without synchronization.
rg -n -C3 '\bplayFinished\b|Callback|while \(!playFinished\)' dotnet-examples/zipvoice-tts-play/Program.cs

Repository: k2-fsa/sherpa-onnx

Length of output: 2127


🏁 Script executed:

#!/bin/bash
# Examine complete context of playFinished usage and verify threading model
cat -n dotnet-examples/zipvoice-tts-play/Program.cs | sed -n '88,195p'

Repository: k2-fsa/sherpa-onnx

Length of output: 3702


Synchronize playFinished across threads to avoid data race and potential hangs.

playFinished is written from the PortAudio callback thread (line 109) and polled on the main thread (line 188) without synchronization. The missing memory barrier allows the compiler to optimize away the poll loop or cache stale values, potentially causing indefinite hangs.

Use ManualResetEventSlim for thread-safe signaling:

Proposed synchronization fix
-    var playFinished = false;
+    using var playbackDone = new ManualResetEventSlim(false);
     
     PortAudioSharp.Stream.Callback playCallback = ...
     {
       if (dataItems.IsCompleted && lastSampleArray == null && lastIndex == 0)
       {
         Console.WriteLine("Finished playing");
-        playFinished = true;
+        playbackDone.Set();
         return StreamCallbackResult.Complete;
       }
       ...
     };
     
-    while (!playFinished)
-    {
-      Thread.Sleep(100);
-    }
+    playbackDone.Wait();
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dotnet-examples/zipvoice-tts-play/Program.cs` around lines 94 - 110, The
playFinished flag is written from the PortAudio callback thread
(PortAudioSharp.Stream.Callback playCallback) and polled on the main thread,
which causes a data race; replace the plain bool with a thread-safe waiter such
as ManualResetEventSlim (e.g., playFinishedEvent), call playFinishedEvent.Set()
inside the playCallback where you currently set playFinished = true, and replace
the main-thread busy-poll of playFinished with playFinishedEvent.Wait(timeout)
(or Wait without timeout if appropriate) so the signal uses proper memory
barriers and avoids hangs.

Comment on lines +5 to +12
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
tar xvf sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
rm sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
fi

if [ ! -f ./vocos_24khz.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
fi

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Use fail-fast curl flags here as well.

Line 5 and Line 11 should fail on HTTP errors to avoid continuing with bad artifacts.

🛠️ Proposed change
-  curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
+  curl -fSL --retry 3 --retry-delay 2 -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
@@
-  curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
+  curl -fSL --retry 3 --retry-delay 2 -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
tar xvf sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
rm sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
fi
if [ ! -f ./vocos_24khz.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
fi
curl -fSL --retry 3 --retry-delay 2 -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
tar xvf sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
rm sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
fi
if [ ! -f ./vocos_24khz.onnx ]; then
curl -fSL --retry 3 --retry-delay 2 -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
fi
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dotnet-examples/zipvoice-tts-play/run.sh` around lines 5 - 12, The curl
invocations in run.sh (the commands that download
sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2 and vocos_24khz.onnx)
don’t fail on HTTP errors; update those curl commands to use fail-fast flags
(e.g., add -f or --fail alongside -SL -O) so the script exits on HTTP failures
and doesn’t continue with missing or incomplete artifacts.

</PropertyGroup>

<ItemGroup>
<PackageReference Include="PortAudioSharp2" Version="*" />

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

#!/usr/bin/env bash
# Verify wildcard PackageReference usage across the repo
rg -n --glob '*.csproj' '<PackageReference[^>]*Version="\*"'

Repository: k2-fsa/sherpa-onnx

Length of output: 1232


🌐 Web query:

What is the latest stable NuGet version of PortAudioSharp2, and is using Version="*" in SDK-style .csproj recommended for reproducible builds?

💡 Result:

The latest stable NuGet version of PortAudioSharp2 is 1.0.6, released on October 16, 2025. Using Version="" (a floating version) in SDK-style .csproj files is not recommended for reproducible builds. Floating versions like "" resolve to the highest available version matching the pattern at restore time, which can lead to different package versions across restores, environments, or after new versions are published. This breaks reproducibility, as the same project may build with varying dependencies. For reproducible builds, pin to specific versions (e.g., Version="1.0.6") or use lock files (RestorePackagesWithLockFile) with --locked-mode in CI. Official guidance emphasizes lowest applicable versions without upper bounds but avoids floating versions for consistency. Enable Central Package Management or Dependabot for updates instead.

Citations:


Pin PortAudioSharp2 to a specific version instead of using a wildcard.

Line 12 uses Version="*", which resolves to the latest available version at restore time and makes builds non-reproducible. This can cause different package versions across environments and unexpected breakage. The latest stable version is 1.0.6—pin to this version explicitly.

🔒 Suggested fix
-    <PackageReference Include="PortAudioSharp2" Version="*" />
+    <PackageReference Include="PortAudioSharp2" Version="1.0.6" />

Note: This issue appears in multiple project files across the repository (kokoro-tts-play.csproj, kitten-tts-play.csproj, speech-recognition-from-microphone.csproj, etc.). Consider applying the fix consistently across all affected files.

📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
<PackageReference Include="PortAudioSharp2" Version="*" />
<PackageReference Include="PortAudioSharp2" Version="1.0.6" />
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dotnet-examples/zipvoice-tts-play/zipvoice-tts-play.csproj` at line 12, The
PackageReference for PortAudioSharp2 currently uses a wildcard Version="*",
making builds non-reproducible; update the PortAudioSharp2 PackageReference (the
XML element with Include="PortAudioSharp2") to pin Version="1.0.6" instead of
"*" and apply the same change to other project files that reference
PortAudioSharp2 (e.g., kokoro-tts-play.csproj, kitten-tts-play.csproj,
speech-recognition-from-microphone.csproj) so all projects use the explicit
1.0.6 version.

Comment on lines +5 to +12
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
tar xvf sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
rm sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
fi

if [ ! -f ./vocos_24khz.onnx ]; then
curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
fi

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Use fail-fast curl flags for reliable downloads.

Line 5 and Line 11 can succeed on HTTP errors without --fail, producing confusing downstream failures.

🛠️ Proposed change
-  curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
+  curl -fSL --retry 3 --retry-delay 2 -O https://github.com/k2-fsa/sherpa-onnx/releases/download/tts-models/sherpa-onnx-zipvoice-distill-int8-zh-en-emilia.tar.bz2
@@
-  curl -SL -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
+  curl -fSL --retry 3 --retry-delay 2 -O https://github.com/k2-fsa/sherpa-onnx/releases/download/vocoder-models/vocos_24khz.onnx
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dotnet-examples/zipvoice-tts/run.sh` around lines 5 - 12, The two curl
invocations that download the model files (the lines calling "curl -SL -O
https://...sherpa-onnx-zipvoice-...tar.bz2" and "curl -SL -O
https://...vocos_24khz.onnx") should be made fail-fast; update those commands to
include curl's fail and error reporting flags (e.g., add --fail/-f and
--show-error, and keep -S -L -O) so HTTP errors cause the script to exit
immediately and surface the underlying error when "curl -SL -O ..." is executed.

@csukuangfj
csukuangfj merged commit a59a5e7 into k2-fsa:master Mar 18, 2026
5 checks passed
@csukuangfj
csukuangfj deleted the dotnet-zipvoice branch March 18, 2026 06:37
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants