demo with UI and Web Worker to avoid main-thread blocking - #3120
Conversation
…based audio generation using sherpa-onnx.
Summary of ChangesHello @yuiyideyui, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly enhances the user experience of the Text-to-Speech demo by refactoring the TTS logic to run within a Web Worker. This change ensures that the main browser thread, responsible for rendering the user interface, remains unblocked and responsive, thereby eliminating UI stuttering that previously occurred during TTS execution. The application now leverages asynchronous communication with the worker to manage TTS initialization, progress updates, and audio generation, leading to a smoother and more fluid interaction. Highlights
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
📝 WalkthroughWalkthroughThis PR refactors the TTS (Text-to-Speech) implementation to offload synthesis processing from the main thread to a dedicated Web Worker. The main app now communicates with the worker via messages, while HTML dependencies on direct TTS scripts are removed. Changes
Sequence Diagram(s)sequenceDiagram
participant User
participant MainApp as Main App<br/>(app-tts.js)
participant Worker as Web Worker<br/>(sherpa-onnx-tts.worker.js)
participant WasmModule as WASM Module
participant WebAudio as Web Audio API
Worker->>WasmModule: Load WASM scripts via importScripts
WasmModule-->>Worker: onRuntimeInitialized callback
Worker->>MainApp: Post ready message with numSpeakers
MainApp->>MainApp: Update UI (enable Generate button)
User->>MainApp: Click Generate with text & settings
MainApp->>Worker: Post "generate" message<br/>(text, sid, speed)
Worker->>WasmModule: Invoke tts.generate()
WasmModule-->>Worker: Return samples & sampleRate
Worker->>MainApp: Post "sherpa-onnx-tts-result"<br/>(samples, sampleRate)
MainApp->>WebAudio: Create AudioContext & buffer
MainApp->>WebAudio: Start audio playback
WebAudio-->>User: Play audio
MainApp->>MainApp: Render clip element
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly related PRs
Suggested labels
Poem
🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing touches
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Code Review
This pull request effectively refactors the TTS generation to a Web Worker, which is an excellent change to prevent blocking the main UI thread and improve user experience. The implementation is solid, using transferable objects for efficient data passing. My review includes a few suggestions to enhance the message handling logic, particularly by adding error handling for worker-side failures and improving the structure of the message processing.
| worker.onmessage = (e) => { | ||
| if (e.data.type === "sherpa-onnx-tts-progress") { | ||
| Module.setStatus(e.data.status); | ||
| } | ||
| if (e.data.type === "sherpa-onnx-tts-ready") { | ||
| ttsInstanceInfo.numSpeakers = e.data.numSpeakers; | ||
| ttsInstanceInfo.isReady = true; | ||
| generateBtn.disabled = false; | ||
| speakerIdLabel.innerHTML = `Speaker ID (0 - ${e.data.numSpeakers - 1}):`; | ||
| return; | ||
| } | ||
| if (e.data.type === "sherpa-onnx-tts-result") { | ||
| let audio = e.data; | ||
|
|
||
| Module = {}; | ||
| console.log(audio.samples.length, audio.sampleRate); | ||
|
|
||
| // https://emscripten.org/docs/api_reference/module.html#Module.locateFile | ||
| Module.locateFile = function(path, scriptDirectory = '') { | ||
| console.log(`path: ${path}, scriptDirectory: ${scriptDirectory}`); | ||
| return scriptDirectory + path; | ||
| if (!audioCtx) { | ||
| audioCtx = new AudioContext({ sampleRate: audio.sampleRate }); | ||
| } | ||
|
|
||
| const buffer = audioCtx.createBuffer( | ||
| 1, | ||
| audio.samples.length, | ||
| audio.sampleRate, | ||
| ); | ||
|
|
||
| buffer.getChannelData(0).set(audio.samples); // 使用 .set() 比 for 循环快得多 | ||
| const source = audioCtx.createBufferSource(); | ||
| source.buffer = buffer; | ||
| source.connect(audioCtx.destination); | ||
| source.start(); | ||
|
|
||
| createAudioTag(audio); | ||
| } | ||
| }; |
There was a problem hiding this comment.
The web worker can post messages with type: 'error', for instance, if TTS initialization or audio generation fails. The current onmessage handler does not account for this, which would cause errors to fail silently from the user's perspective. It's important to handle these errors to provide feedback to the user.
Additionally, the if/if/if structure can be improved by using an if...else if chain for better readability and to avoid unnecessary condition checks. I've combined both improvements in the suggestion below.
worker.onmessage = (e) => {
if (e.data.type === "sherpa-onnx-tts-progress") {
Module.setStatus(e.data.status);
} else if (e.data.type === "sherpa-onnx-tts-ready") {
ttsInstanceInfo.numSpeakers = e.data.numSpeakers;
ttsInstanceInfo.isReady = true;
generateBtn.disabled = false;
speakerIdLabel.innerHTML = `Speaker ID (0 - ${e.data.numSpeakers - 1}):`;
} else if (e.data.type === "sherpa-onnx-tts-result") {
let audio = e.data;
console.log(audio.samples.length, audio.sampleRate);
if (!audioCtx) {
audioCtx = new AudioContext({ sampleRate: audio.sampleRate });
}
const buffer = audioCtx.createBuffer(
1,
audio.samples.length,
audio.sampleRate,
);
buffer.getChannelData(0).set(audio.samples); // 使用 .set() 比 for 循环快得多
const source = audioCtx.createBufferSource();
source.buffer = buffer;
source.connect(audioCtx.destination);
source.start();
createAudioTag(audio);
} else if (e.data.type === "error") {
console.error(e.data.message);
alert(e.data.message);
}
};| }; | ||
| importScripts("/sherpa-onnx-wasm-main-tts.js"); | ||
| importScripts("/sherpa-onnx-tts.js"); | ||
| self.onmessage = async (e) => { |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Fix all issues with AI agents
In `@wasm/tts/app-tts.js`:
- Around line 109-110: The alert message shown when validating speakerId has a
typo; update the user-facing string in the alert call inside the validation
branch that checks speakerId against ttsInstanceInfo.numSpeakers (the block with
`if (speakerId > ttsInstanceInfo.numSpeakers - 1)`) to read "Please enter a
number between 0 and X" instead of "Pleaser enter..."; keep the interpolation of
ttsInstanceInfo.numSpeakers - 1 as-is so the numeric range is correct.
- Around line 19-52: Extend the existing worker.onmessage handler to catch
e.data.type === "error" (the worker posts { type: "error" }) and surface the
failure: call Module.setStatus(e.data.message || "TTS worker error"),
console.error(e.data), set ttsInstanceInfo.isReady = false, disable generateBtn,
and update a visible UI element (e.g., speakerIdLabel.innerHTML or a dedicated
error label) with the error text so the failure is not silent; implement this
alongside the existing handlers in the worker.onmessage block that currently
handles "sherpa-onnx-tts-progress", "sherpa-onnx-tts-ready", and
"sherpa-onnx-tts-result".
🧹 Nitpick comments (1)
wasm/tts/sherpa-onnx-tts.worker.js (1)
28-29: Prefer relative URLs for worker imports to keep sub-path deployments working.Absolute root paths can 404 when the demo is hosted under a subdirectory. Consider relative paths so they resolve next to the worker script.
✅ Suggested update
-importScripts("/sherpa-onnx-wasm-main-tts.js"); -importScripts("/sherpa-onnx-tts.js"); +importScripts("sherpa-onnx-wasm-main-tts.js"); +importScripts("sherpa-onnx-tts.js");
| worker.onmessage = (e) => { | ||
| if (e.data.type === "sherpa-onnx-tts-progress") { | ||
| Module.setStatus(e.data.status); | ||
| } | ||
| if (e.data.type === "sherpa-onnx-tts-ready") { | ||
| ttsInstanceInfo.numSpeakers = e.data.numSpeakers; | ||
| ttsInstanceInfo.isReady = true; | ||
| generateBtn.disabled = false; | ||
| speakerIdLabel.innerHTML = `Speaker ID (0 - ${e.data.numSpeakers - 1}):`; | ||
| return; | ||
| } | ||
| if (e.data.type === "sherpa-onnx-tts-result") { | ||
| let audio = e.data; | ||
|
|
||
| Module = {}; | ||
| console.log(audio.samples.length, audio.sampleRate); | ||
|
|
||
| // https://emscripten.org/docs/api_reference/module.html#Module.locateFile | ||
| Module.locateFile = function(path, scriptDirectory = '') { | ||
| console.log(`path: ${path}, scriptDirectory: ${scriptDirectory}`); | ||
| return scriptDirectory + path; | ||
| if (!audioCtx) { | ||
| audioCtx = new AudioContext({ sampleRate: audio.sampleRate }); | ||
| } | ||
|
|
||
| const buffer = audioCtx.createBuffer( | ||
| 1, | ||
| audio.samples.length, | ||
| audio.sampleRate, | ||
| ); | ||
|
|
||
| buffer.getChannelData(0).set(audio.samples); // 使用 .set() 比 for 循环快得多 | ||
| const source = audioCtx.createBufferSource(); | ||
| source.buffer = buffer; | ||
| source.connect(audioCtx.destination); | ||
| source.start(); | ||
|
|
||
| createAudioTag(audio); | ||
| } |
There was a problem hiding this comment.
Handle worker error messages so failures aren’t silent.
The worker posts { type: "error" }, but the UI ignores it. Add a handler to surface the failure and update status.
✅ Suggested update
worker.onmessage = (e) => {
if (e.data.type === "sherpa-onnx-tts-progress") {
Module.setStatus(e.data.status);
}
if (e.data.type === "sherpa-onnx-tts-ready") {
ttsInstanceInfo.numSpeakers = e.data.numSpeakers;
ttsInstanceInfo.isReady = true;
generateBtn.disabled = false;
speakerIdLabel.innerHTML = `Speaker ID (0 - ${e.data.numSpeakers - 1}):`;
return;
}
+ if (e.data.type === "error") {
+ console.error(e.data.message);
+ Module.setStatus(`Error: ${e.data.message}`);
+ return;
+ }
if (e.data.type === "sherpa-onnx-tts-result") {
let audio = e.data;📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| worker.onmessage = (e) => { | |
| if (e.data.type === "sherpa-onnx-tts-progress") { | |
| Module.setStatus(e.data.status); | |
| } | |
| if (e.data.type === "sherpa-onnx-tts-ready") { | |
| ttsInstanceInfo.numSpeakers = e.data.numSpeakers; | |
| ttsInstanceInfo.isReady = true; | |
| generateBtn.disabled = false; | |
| speakerIdLabel.innerHTML = `Speaker ID (0 - ${e.data.numSpeakers - 1}):`; | |
| return; | |
| } | |
| if (e.data.type === "sherpa-onnx-tts-result") { | |
| let audio = e.data; | |
| Module = {}; | |
| console.log(audio.samples.length, audio.sampleRate); | |
| // https://emscripten.org/docs/api_reference/module.html#Module.locateFile | |
| Module.locateFile = function(path, scriptDirectory = '') { | |
| console.log(`path: ${path}, scriptDirectory: ${scriptDirectory}`); | |
| return scriptDirectory + path; | |
| if (!audioCtx) { | |
| audioCtx = new AudioContext({ sampleRate: audio.sampleRate }); | |
| } | |
| const buffer = audioCtx.createBuffer( | |
| 1, | |
| audio.samples.length, | |
| audio.sampleRate, | |
| ); | |
| buffer.getChannelData(0).set(audio.samples); // 使用 .set() 比 for 循环快得多 | |
| const source = audioCtx.createBufferSource(); | |
| source.buffer = buffer; | |
| source.connect(audioCtx.destination); | |
| source.start(); | |
| createAudioTag(audio); | |
| } | |
| worker.onmessage = (e) => { | |
| if (e.data.type === "sherpa-onnx-tts-progress") { | |
| Module.setStatus(e.data.status); | |
| } | |
| if (e.data.type === "sherpa-onnx-tts-ready") { | |
| ttsInstanceInfo.numSpeakers = e.data.numSpeakers; | |
| ttsInstanceInfo.isReady = true; | |
| generateBtn.disabled = false; | |
| speakerIdLabel.innerHTML = `Speaker ID (0 - ${e.data.numSpeakers - 1}):`; | |
| return; | |
| } | |
| if (e.data.type === "error") { | |
| console.error(e.data.message); | |
| Module.setStatus(`Error: ${e.data.message}`); | |
| return; | |
| } | |
| if (e.data.type === "sherpa-onnx-tts-result") { | |
| let audio = e.data; | |
| console.log(audio.samples.length, audio.sampleRate); | |
| if (!audioCtx) { | |
| audioCtx = new AudioContext({ sampleRate: audio.sampleRate }); | |
| } | |
| const buffer = audioCtx.createBuffer( | |
| 1, | |
| audio.samples.length, | |
| audio.sampleRate, | |
| ); | |
| buffer.getChannelData(0).set(audio.samples); // 使用 .set() 比 for 循环快得多 | |
| const source = audioCtx.createBufferSource(); | |
| source.buffer = buffer; | |
| source.connect(audioCtx.destination); | |
| source.start(); | |
| createAudioTag(audio); | |
| } |
🤖 Prompt for AI Agents
In `@wasm/tts/app-tts.js` around lines 19 - 52, Extend the existing
worker.onmessage handler to catch e.data.type === "error" (the worker posts {
type: "error" }) and surface the failure: call Module.setStatus(e.data.message
|| "TTS worker error"), console.error(e.data), set ttsInstanceInfo.isReady =
false, disable generateBtn, and update a visible UI element (e.g.,
speakerIdLabel.innerHTML or a dedicated error label) with the error text so the
failure is not silent; implement this alongside the existing handlers in the
worker.onmessage block that currently handles "sherpa-onnx-tts-progress",
"sherpa-onnx-tts-ready", and "sherpa-onnx-tts-result".
| if (speakerId > ttsInstanceInfo.numSpeakers - 1) { | ||
| alert(`Pleaser enter a number between 0 and ${ttsInstanceInfo.numSpeakers - 1}`); |
There was a problem hiding this comment.
Fix typo in user-facing alert text.
✏️ Suggested fix
- alert(`Pleaser enter a number between 0 and ${ttsInstanceInfo.numSpeakers - 1}`);
+ alert(`Please enter a number between 0 and ${ttsInstanceInfo.numSpeakers - 1}`);📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| if (speakerId > ttsInstanceInfo.numSpeakers - 1) { | |
| alert(`Pleaser enter a number between 0 and ${ttsInstanceInfo.numSpeakers - 1}`); | |
| if (speakerId > ttsInstanceInfo.numSpeakers - 1) { | |
| alert(`Please enter a number between 0 and ${ttsInstanceInfo.numSpeakers - 1}`); |
🤖 Prompt for AI Agents
In `@wasm/tts/app-tts.js` around lines 109 - 110, The alert message shown when
validating speakerId has a typo; update the user-facing string in the alert call
inside the validation branch that checks speakerId against
ttsInstanceInfo.numSpeakers (the block with `if (speakerId >
ttsInstanceInfo.numSpeakers - 1)`) to read "Please enter a number between 0 and
X" instead of "Pleaser enter..."; keep the interpolation of
ttsInstanceInfo.numSpeakers - 1 as-is so the numeric range is correct.
|
Is there a website like https://huggingface.co/spaces/k2-fsa/web-assembly-tts-sherpa-onnx-en |
You can now test this directly by accessing this address in your browser: https://huggingface.co/spaces/yuiyide/web-assembly-tts-sherpa-onnx |
csukuangfj
left a comment
There was a problem hiding this comment.
Thank you for your contribution!
…based audio generation using sherpa-onnx. (k2-fsa#3120) This pull request significantly enhances the user experience of the Text-to-Speech demo by refactoring the TTS logic to run within a Web Worker. This change ensures that the main browser thread, responsible for rendering the user interface, remains unblocked and responsive, thereby eliminating UI stuttering that previously occurred during TTS execution. The application now leverages asynchronous communication with the worker to manage TTS initialization, progress updates, and audio generation, leading to a smoother and more fluid interaction.
When TTS is executed on the main thread (without a Web Worker), it blocks rendering and causes UI stuttering。
Summary by CodeRabbit
Refactor
Improvements