Skip to content

demo with UI and Web Worker to avoid main-thread blocking - #3120

Merged
csukuangfj merged 1 commit into
k2-fsa:masterfrom
yuiyideyui:master
Feb 5, 2026
Merged

csukuangfj merged 1 commit into
k2-fsa:masterfrom
yuiyideyui:master

Conversation

@yuiyideyui

@yuiyideyui yuiyideyui commented Feb 2, 2026 •

Copy link
Copy Markdown
Contributor

When TTS is executed on the main thread (without a Web Worker), it blocks rendering and causes UI stuttering。

Summary by CodeRabbit

  • Refactor

    • Restructured text-to-speech synthesis processing architecture for enhanced performance and maintainability.
    • Enhanced error handling and validation for generation requests.
    • Streamlined UI update logic for generation status and speaker range display.
  • Improvements

    • Better progress reporting and status messaging during synthesis operations.
    • Optimized audio playback and buffer initialization.

@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Feb 2, 2026
@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello @yuiyideyui, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the user experience of the Text-to-Speech demo by refactoring the TTS logic to run within a Web Worker. This change ensures that the main browser thread, responsible for rendering the user interface, remains unblocked and responsive, thereby eliminating UI stuttering that previously occurred during TTS execution. The application now leverages asynchronous communication with the worker to manage TTS initialization, progress updates, and audio generation, leading to a smoother and more fluid interaction.

Highlights

  • Web Worker Integration: Introduced a Web Worker (sherpa-onnx-tts.worker.js) to offload Text-to-Speech (TTS) processing from the main thread, preventing UI blocking during audio generation.
  • Asynchronous TTS Processing: The main application now communicates with the Web Worker to initialize the TTS engine and generate audio, receiving results asynchronously via message passing.
  • Improved UI Responsiveness: By moving heavy computation to a background thread, the user interface remains responsive even during intensive TTS operations, addressing the issue of UI stuttering.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Feb 2, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR refactors the TTS (Text-to-Speech) implementation to offload synthesis processing from the main thread to a dedicated Web Worker. The main app now communicates with the worker via messages, while HTML dependencies on direct TTS scripts are removed.

Changes

Cohort / File(s) Summary
Web Worker Implementation
wasm/tts/sherpa-onnx-tts.worker.js
New worker file that handles TTS synthesis asynchronously, manages WASM module initialization via Module callbacks, and processes generate requests by invoking tts.generate with text, speaker ID, and speed parameters.
Main App Refactoring
wasm/tts/app-tts.js
Replaced in-page TTS processing with worker-based communication; removed direct tts object usage and runtime initialization hook; added worker instance and ttsInstanceInfo state for managing progress and readiness; updated generate button handler to send messages to worker and receive audio playback via Web Audio API.
HTML Script Cleanup
wasm/tts/index.html
Removed two TTS script includes (sherpa-onnx-tts.js and sherpa-onnx-wasm-main-tts.js) as TTS logic now executes within the dedicated worker context.

Sequence Diagram(s)

sequenceDiagram
    participant User
    participant MainApp as Main App<br/>(app-tts.js)
    participant Worker as Web Worker<br/>(sherpa-onnx-tts.worker.js)
    participant WasmModule as WASM Module
    participant WebAudio as Web Audio API

    Worker->>WasmModule: Load WASM scripts via importScripts
    WasmModule-->>Worker: onRuntimeInitialized callback
    Worker->>MainApp: Post ready message with numSpeakers
    MainApp->>MainApp: Update UI (enable Generate button)

    User->>MainApp: Click Generate with text & settings
    MainApp->>Worker: Post "generate" message<br/>(text, sid, speed)
    Worker->>WasmModule: Invoke tts.generate()
    WasmModule-->>Worker: Return samples & sampleRate
    Worker->>MainApp: Post "sherpa-onnx-tts-result"<br/>(samples, sampleRate)
    
    MainApp->>WebAudio: Create AudioContext & buffer
    MainApp->>WebAudio: Start audio playback
    WebAudio-->>User: Play audio
    MainApp->>MainApp: Render clip element
Loading

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • k2-fsa/sherpa-onnx#2885: Modifies createOfflineTts function that the new worker imports and invokes for TTS synthesis.

Suggested labels

size:L

Poem

🐰 A Worker takes the stage so bright,
Threading synthesis left and right,
Off the main thread, smooth and fleet,
Making voices sound so sweet! 🎵

🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: refactoring TTS demo to use a Web Worker to prevent main-thread blocking, which is the core objective of the PR.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request effectively refactors the TTS generation to a Web Worker, which is an excellent change to prevent blocking the main UI thread and improve user experience. The implementation is solid, using transferable objects for efficient data passing. My review includes a few suggestions to enhance the message handling logic, particularly by adding error handling for worker-side failures and improving the structure of the message processing.

Comment thread wasm/tts/app-tts.js
Comment on lines +19 to 53
worker.onmessage = (e) => {
if (e.data.type === "sherpa-onnx-tts-progress") {
Module.setStatus(e.data.status);
}
if (e.data.type === "sherpa-onnx-tts-ready") {
ttsInstanceInfo.numSpeakers = e.data.numSpeakers;
ttsInstanceInfo.isReady = true;
generateBtn.disabled = false;
speakerIdLabel.innerHTML = `Speaker ID (0 - ${e.data.numSpeakers - 1}):`;
return;
}
if (e.data.type === "sherpa-onnx-tts-result") {
let audio = e.data;

Module = {};
console.log(audio.samples.length, audio.sampleRate);

// https://emscripten.org/docs/api_reference/module.html#Module.locateFile
Module.locateFile = function(path, scriptDirectory = '') {
console.log(`path: ${path}, scriptDirectory: ${scriptDirectory}`);
return scriptDirectory + path;
if (!audioCtx) {
audioCtx = new AudioContext({ sampleRate: audio.sampleRate });
}

const buffer = audioCtx.createBuffer(
1,
audio.samples.length,
audio.sampleRate,
);

buffer.getChannelData(0).set(audio.samples); // 使用 .set() 比 for 循环快得多
const source = audioCtx.createBufferSource();
source.buffer = buffer;
source.connect(audioCtx.destination);
source.start();

createAudioTag(audio);
}
};

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The web worker can post messages with type: 'error', for instance, if TTS initialization or audio generation fails. The current onmessage handler does not account for this, which would cause errors to fail silently from the user's perspective. It's important to handle these errors to provide feedback to the user.

Additionally, the if/if/if structure can be improved by using an if...else if chain for better readability and to avoid unnecessary condition checks. I've combined both improvements in the suggestion below.

worker.onmessage = (e) => {
  if (e.data.type === "sherpa-onnx-tts-progress") {
    Module.setStatus(e.data.status);
  } else if (e.data.type === "sherpa-onnx-tts-ready") {
    ttsInstanceInfo.numSpeakers = e.data.numSpeakers;
    ttsInstanceInfo.isReady = true;
    generateBtn.disabled = false;
    speakerIdLabel.innerHTML = `Speaker ID (0 - ${e.data.numSpeakers - 1}):`;
  } else if (e.data.type === "sherpa-onnx-tts-result") {
    let audio = e.data;

    console.log(audio.samples.length, audio.sampleRate);

    if (!audioCtx) {
      audioCtx = new AudioContext({ sampleRate: audio.sampleRate });
    }

    const buffer = audioCtx.createBuffer(
      1,
      audio.samples.length,
      audio.sampleRate,
    );

    buffer.getChannelData(0).set(audio.samples); // 使用 .set() 比 for 循环快得多
    const source = audioCtx.createBufferSource();
    source.buffer = buffer;
    source.connect(audioCtx.destination);
    source.start();

    createAudioTag(audio);
  } else if (e.data.type === "error") {
    console.error(e.data.message);
    alert(e.data.message);
  }
};

};
importScripts("/sherpa-onnx-wasm-main-tts.js");
importScripts("/sherpa-onnx-tts.js");
self.onmessage = async (e) => {

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The onmessage handler is marked as async, but it doesn't use the await keyword. The async keyword is unnecessary here and can be removed for code clarity.

Suggested change
self.onmessage = async (e) => {
self.onmessage = (e) => {

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Fix all issues with AI agents
In `@wasm/tts/app-tts.js`:
- Around line 109-110: The alert message shown when validating speakerId has a
typo; update the user-facing string in the alert call inside the validation
branch that checks speakerId against ttsInstanceInfo.numSpeakers (the block with
`if (speakerId > ttsInstanceInfo.numSpeakers - 1)`) to read "Please enter a
number between 0 and X" instead of "Pleaser enter..."; keep the interpolation of
ttsInstanceInfo.numSpeakers - 1 as-is so the numeric range is correct.
- Around line 19-52: Extend the existing worker.onmessage handler to catch
e.data.type === "error" (the worker posts { type: "error" }) and surface the
failure: call Module.setStatus(e.data.message || "TTS worker error"),
console.error(e.data), set ttsInstanceInfo.isReady = false, disable generateBtn,
and update a visible UI element (e.g., speakerIdLabel.innerHTML or a dedicated
error label) with the error text so the failure is not silent; implement this
alongside the existing handlers in the worker.onmessage block that currently
handles "sherpa-onnx-tts-progress", "sherpa-onnx-tts-ready", and
"sherpa-onnx-tts-result".
🧹 Nitpick comments (1)
wasm/tts/sherpa-onnx-tts.worker.js (1)

28-29: Prefer relative URLs for worker imports to keep sub-path deployments working.

Absolute root paths can 404 when the demo is hosted under a subdirectory. Consider relative paths so they resolve next to the worker script.

✅ Suggested update
-importScripts("/sherpa-onnx-wasm-main-tts.js");
-importScripts("/sherpa-onnx-tts.js");
+importScripts("sherpa-onnx-wasm-main-tts.js");
+importScripts("sherpa-onnx-tts.js");

Comment thread wasm/tts/app-tts.js
Comment on lines +19 to +52
worker.onmessage = (e) => {
if (e.data.type === "sherpa-onnx-tts-progress") {
Module.setStatus(e.data.status);
}
if (e.data.type === "sherpa-onnx-tts-ready") {
ttsInstanceInfo.numSpeakers = e.data.numSpeakers;
ttsInstanceInfo.isReady = true;
generateBtn.disabled = false;
speakerIdLabel.innerHTML = `Speaker ID (0 - ${e.data.numSpeakers - 1}):`;
return;
}
if (e.data.type === "sherpa-onnx-tts-result") {
let audio = e.data;

Module = {};
console.log(audio.samples.length, audio.sampleRate);

// https://emscripten.org/docs/api_reference/module.html#Module.locateFile
Module.locateFile = function(path, scriptDirectory = '') {
console.log(`path: ${path}, scriptDirectory: ${scriptDirectory}`);
return scriptDirectory + path;
if (!audioCtx) {
audioCtx = new AudioContext({ sampleRate: audio.sampleRate });
}

const buffer = audioCtx.createBuffer(
1,
audio.samples.length,
audio.sampleRate,
);

buffer.getChannelData(0).set(audio.samples); // 使用 .set() 比 for 循环快得多
const source = audioCtx.createBufferSource();
source.buffer = buffer;
source.connect(audioCtx.destination);
source.start();

createAudioTag(audio);
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Handle worker error messages so failures aren’t silent.

The worker posts { type: "error" }, but the UI ignores it. Add a handler to surface the failure and update status.

✅ Suggested update
 worker.onmessage = (e) => {
   if (e.data.type === "sherpa-onnx-tts-progress") {
     Module.setStatus(e.data.status);
   }
   if (e.data.type === "sherpa-onnx-tts-ready") {
     ttsInstanceInfo.numSpeakers = e.data.numSpeakers;
     ttsInstanceInfo.isReady = true;
     generateBtn.disabled = false;
     speakerIdLabel.innerHTML = `Speaker ID (0 - ${e.data.numSpeakers - 1}):`;
     return;
   }
+  if (e.data.type === "error") {
+    console.error(e.data.message);
+    Module.setStatus(`Error: ${e.data.message}`);
+    return;
+  }
   if (e.data.type === "sherpa-onnx-tts-result") {
     let audio = e.data;
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
worker.onmessage = (e) => {
if (e.data.type === "sherpa-onnx-tts-progress") {
Module.setStatus(e.data.status);
}
if (e.data.type === "sherpa-onnx-tts-ready") {
ttsInstanceInfo.numSpeakers = e.data.numSpeakers;
ttsInstanceInfo.isReady = true;
generateBtn.disabled = false;
speakerIdLabel.innerHTML = `Speaker ID (0 - ${e.data.numSpeakers - 1}):`;
return;
}
if (e.data.type === "sherpa-onnx-tts-result") {
let audio = e.data;
Module = {};
console.log(audio.samples.length, audio.sampleRate);
// https://emscripten.org/docs/api_reference/module.html#Module.locateFile
Module.locateFile = function(path, scriptDirectory = '') {
console.log(`path: ${path}, scriptDirectory: ${scriptDirectory}`);
return scriptDirectory + path;
if (!audioCtx) {
audioCtx = new AudioContext({ sampleRate: audio.sampleRate });
}
const buffer = audioCtx.createBuffer(
1,
audio.samples.length,
audio.sampleRate,
);
buffer.getChannelData(0).set(audio.samples); // 使用 .set() 比 for 循环快得多
const source = audioCtx.createBufferSource();
source.buffer = buffer;
source.connect(audioCtx.destination);
source.start();
createAudioTag(audio);
}
worker.onmessage = (e) => {
if (e.data.type === "sherpa-onnx-tts-progress") {
Module.setStatus(e.data.status);
}
if (e.data.type === "sherpa-onnx-tts-ready") {
ttsInstanceInfo.numSpeakers = e.data.numSpeakers;
ttsInstanceInfo.isReady = true;
generateBtn.disabled = false;
speakerIdLabel.innerHTML = `Speaker ID (0 - ${e.data.numSpeakers - 1}):`;
return;
}
if (e.data.type === "error") {
console.error(e.data.message);
Module.setStatus(`Error: ${e.data.message}`);
return;
}
if (e.data.type === "sherpa-onnx-tts-result") {
let audio = e.data;
console.log(audio.samples.length, audio.sampleRate);
if (!audioCtx) {
audioCtx = new AudioContext({ sampleRate: audio.sampleRate });
}
const buffer = audioCtx.createBuffer(
1,
audio.samples.length,
audio.sampleRate,
);
buffer.getChannelData(0).set(audio.samples); // 使用 .set() 比 for 循环快得多
const source = audioCtx.createBufferSource();
source.buffer = buffer;
source.connect(audioCtx.destination);
source.start();
createAudioTag(audio);
}
🤖 Prompt for AI Agents
In `@wasm/tts/app-tts.js` around lines 19 - 52, Extend the existing
worker.onmessage handler to catch e.data.type === "error" (the worker posts {
type: "error" }) and surface the failure: call Module.setStatus(e.data.message
|| "TTS worker error"), console.error(e.data), set ttsInstanceInfo.isReady =
false, disable generateBtn, and update a visible UI element (e.g.,
speakerIdLabel.innerHTML or a dedicated error label) with the error text so the
failure is not silent; implement this alongside the existing handlers in the
worker.onmessage block that currently handles "sherpa-onnx-tts-progress",
"sherpa-onnx-tts-ready", and "sherpa-onnx-tts-result".

Comment thread wasm/tts/app-tts.js
Comment on lines +109 to +110
if (speakerId > ttsInstanceInfo.numSpeakers - 1) {
alert(`Pleaser enter a number between 0 and ${ttsInstanceInfo.numSpeakers - 1}`);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Fix typo in user-facing alert text.

✏️ Suggested fix
-    alert(`Pleaser enter a number between 0 and ${ttsInstanceInfo.numSpeakers - 1}`);
+    alert(`Please enter a number between 0 and ${ttsInstanceInfo.numSpeakers - 1}`);
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
if (speakerId > ttsInstanceInfo.numSpeakers - 1) {
alert(`Pleaser enter a number between 0 and ${ttsInstanceInfo.numSpeakers - 1}`);
if (speakerId > ttsInstanceInfo.numSpeakers - 1) {
alert(`Please enter a number between 0 and ${ttsInstanceInfo.numSpeakers - 1}`);
🤖 Prompt for AI Agents
In `@wasm/tts/app-tts.js` around lines 109 - 110, The alert message shown when
validating speakerId has a typo; update the user-facing string in the alert call
inside the validation branch that checks speakerId against
ttsInstanceInfo.numSpeakers (the block with `if (speakerId >
ttsInstanceInfo.numSpeakers - 1)`) to read "Please enter a number between 0 and
X" instead of "Pleaser enter..."; keep the interpolation of
ttsInstanceInfo.numSpeakers - 1 as-is so the numeric range is correct.

@csukuangfj

Copy link
Copy Markdown
Collaborator

Is there a website like https://huggingface.co/spaces/k2-fsa/web-assembly-tts-sherpa-onnx-en
where we can try out the changes?

@yuiyideyui

Copy link
Copy Markdown
Contributor Author

You can now test this directly by accessing this address in your browser: https://huggingface.co/spaces/yuiyide/web-assembly-tts-sherpa-onnx

Is there a website like https://huggingface.co/spaces/k2-fsa/web-assembly-tts-sherpa-onnx-en where we can try out the changes?

You can now test this directly by accessing this address in your browser: https://huggingface.co/spaces/yuiyide/web-assembly-tts-sherpa-onnx

@csukuangfj csukuangfj left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for your contribution!

@csukuangfj
csukuangfj merged commit a89db67 into k2-fsa:master Feb 5, 2026
1 check passed
csukuangfj pushed a commit to csukuangfj/sherpa-onnx that referenced this pull request Feb 5, 2026
…based audio generation using sherpa-onnx. (k2-fsa#3120)

This pull request significantly enhances the user experience of the Text-to-Speech demo by refactoring the TTS logic to run within a Web Worker. This change ensures that the main browser thread, responsible for rendering the user interface, remains unblocked and responsive, thereby eliminating UI stuttering that previously occurred during TTS execution. The application now leverages asynchronous communication with the worker to manage TTS initialization, progress updates, and audio generation, leading to a smoother and more fluid interaction.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants