feat: add is_final support for streaming Paraformer - #3282
ZhaoChaoqun wants to merge 2 commits into
Conversation
The CIF decoder can occasionally produce sos (token id 1) or eos (token id 2) tokens in its output, particularly when tail frames generate extra CIF fires. Previously only blank (token id 0) was filtered. Add sos and eos to the filter to prevent spurious tokens in the recognition result.
When audio ends with fewer than chunk_size (61) frames remaining, the last chunk is silently dropped and CIF residual alpha is never flushed, causing tail tokens to be lost. This commit adds SetParaformerFinalChunk() to OnlineStream. When set, it enables two behaviors: 1. Short chunk acceptance: IsReady() allows chunks shorter than chunk_size frames. DecodeStream() zero-pads them to chunk_size before feeding to the encoder. 2. CIF tail flush: after the CIF integration loop, if the residual alpha >= 0.6 (empirically tuned), the partially accumulated token is force-fired. This recovers tail tokens that were >60% accumulated but never reached the fire threshold. The flag is automatically cleared on Reset(). API surface: - C++: OnlineStream::SetParaformerFinalChunk() / IsParaformerFinalChunk() - C: SherpaOnnxOnlineStreamSetFinalChunk() - Python: OnlineStream.set_paraformer_final_chunk() Other language bindings (Java, Kotlin, Go, Swift, C#) are not yet updated and can be added in follow-up commits. Test results on 73-entry zh corpus: - Average CER: 0.0762 (baseline) -> 0.0656 (is_final), 14% improvement - 15 improved, 58 unchanged, 0 regressed - 0 mid-chunk differences (intermediate results 100% identical)
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request introduces crucial enhancements to the streaming Paraformer model within Sherpa-Onnx, primarily by adding explicit support for handling the final audio chunk in a stream. This resolves previous limitations where short final segments and residual acoustic information were discarded, leading to reduced accuracy. The changes ensure that all audio data, especially at the end of an utterance, is properly processed and integrated into the transcription, resulting in a notable improvement in average character error rate (CER) for the final decode cycle. Highlights
Changelog
Activity
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension. Footnotes
|
📝 WalkthroughWalkthroughA new C API function Changes
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Poem
🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Code Review
This pull request introduces is_final support for streaming Paraformer, effectively addressing issues with short final audio chunks being dropped and ensuring residual CIF alpha is flushed. The changes are well-implemented across the C-API, C++ core, and Python bindings. The logic for handling the final chunk, including padding, CIF tail flushing, and filtering of SOS/EOS tokens, is sound and clearly commented. I have a couple of suggestions to enhance API consistency and code readability.
| if (t == 0 || t == 1 || t == 2) { | ||
| // skip blank(0), sos(1), eos(2) | ||
| continue; | ||
| } |
There was a problem hiding this comment.
Using magic numbers for special token IDs makes the code less readable and harder to maintain. It's better to define them as named constants.
You could add these constants to the OnlineRecognizerParaformerImpl class, for instance, near kCifTailFlushMinAlpha:
static constexpr int32_t kBlankId = 0;
static constexpr int32_t kSosId = 1;
static constexpr int32_t kEosId = 2;Then, you can use them here to make the condition more explicit.
| if (t == 0 || t == 1 || t == 2) { | |
| // skip blank(0), sos(1), eos(2) | |
| continue; | |
| } | |
| if (t == kBlankId || t == kSosId || t == kEosId) { | |
| // skip blank(0), sos(1), eos(2) | |
| continue; | |
| } |
| void SherpaOnnxOnlineStreamSetFinalChunk( | ||
| const SherpaOnnxOnlineStream *stream) { | ||
| stream->impl->SetParaformerFinalChunk(true); | ||
| } |
There was a problem hiding this comment.
For consistency with the C++ (SetParaformerFinalChunk) and Python (set_paraformer_final_chunk) APIs, and to make it clear this function is specific to Paraformer models, consider renaming this function to SherpaOnnxOnlineStreamSetParaformerFinalChunk.
This change would also need to be applied to:
- The function declaration in
sherpa-onnx/c-api/c-api.h - The exported symbol in
sherpa-onnx/c-api/sherpa-onnx-symbols-c.exp
void SherpaOnnxOnlineStreamSetParaformerFinalChunk(
const SherpaOnnxOnlineStream *stream) {
stream->impl->SetParaformerFinalChunk(true);
}|
Can you add tail padding instead of invoking |
|
Thanks for the review! I actually started with tail padding — that was our baseline approach. Unfortunately, padding alone cannot solve the two core problems: Problem 1: Short final chunk is droppedWhen the remaining audio is shorter than Problem 2: CIF residual alpha is never flushedSilence frames produce near-zero encoder output → near-zero CIF alpha contributions. The residual alpha accumulated from real speech never reaches the fire threshold (1.0), so the final token is permanently lost. Even 1 second of silence padding cannot push the residual over the threshold because silence adds essentially nothing to the alpha accumulator. Evidence from our testing
On the official bundled test wavs:
The
This is analogous to the non-streaming Paraformer's tail handling in Happy to discuss alternative API designs if the concern is about API surface! |
i mean,can you replace Whenever you call
We try to avoid adding a specific API for a specific model. Instead, we want to fix #3101 |
|
Hi @csukuangfj, I've implemented the
Once these are merged, follow-up PRs will migrate Could you take a look and see if this is what you had in mind for #3101? Thanks! |
|
Thanks! Please first fix the comments in the |
|
Can you fix the comments? |
|
Sure, I'll fix the comments in PR #3307 first today, thanks! |
|
Thanks! Closing this PR since all other PRs are merged. |
Summary
Streaming Paraformer silently drops the last audio chunk when it is shorter than 61 frames, and never flushes residual CIF alpha — causing tail tokens to be lost. This PR adds
SetParaformerFinalChunk()to fix both issues, following the same approach as FunASR'sis_final=True.Two commits:
fix: filter sos/eos tokens in streaming Paraformer decoder outputfeat: add is_final support for streaming ParaformerIsReady()allows chunks < 61 frames whenis_finalis set.DecodeStream()zero-pads them to chunk_size.paraformer_is_final_is reset tofalseinOnlineStream::Reset().SherpaOnnxOnlineStreamSetFinalChunk), C++ (OnlineStream::SetParaformerFinalChunk), Python (set_paraformer_final_chunk).Test results with official test_wavs
Using
sherpa-onnx-streaming-paraformer-bilingual-zh-en/test_wavs/(5 files from model package):0.wav1.wav2.wav3.wav8k.wav3 unchanged, 2 improved, 0 regressed.
The improvement is most visible when the audio's tail tokens fall in a short final chunk (<61 frames).
1.wavrecovers 3 truncated characters ("什" → "什么意思啊").Usage
Files changed
sherpa-onnx/csrc/online-stream.h/.cc—SetParaformerFinalChunk/IsParaformerFinalChunk, reset inReset()sherpa-onnx/csrc/online-recognizer-paraformer-impl.h—IsReady,DecodeStream(short chunk + tail flush), sos/eos filtersherpa-onnx/c-api/c-api.h/.cc—SherpaOnnxOnlineStreamSetFinalChunksherpa-onnx/c-api/sherpa-onnx-symbols-c.exp— macOS symbol exportsherpa-onnx/python/csrc/online-stream.cc— Python binding