Skip to content

feat: add is_final support for streaming Paraformer - #3282

Closed
ZhaoChaoqun wants to merge 2 commits into
k2-fsa:masterfrom
ZhaoChaoqun:main
Closed

ZhaoChaoqun wants to merge 2 commits into
k2-fsa:masterfrom
ZhaoChaoqun:main

Conversation

@ZhaoChaoqun

@ZhaoChaoqun ZhaoChaoqun commented Mar 11, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Streaming Paraformer silently drops the last audio chunk when it is shorter than 61 frames, and never flushes residual CIF alpha — causing tail tokens to be lost. This PR adds SetParaformerFinalChunk() to fix both issues, following the same approach as FunASR's is_final=True.

Two commits:

  1. fix: filter sos/eos tokens in streaming Paraformer decoder output

    • The CIF decoder can produce sos (id=1) or eos (id=2) tokens. Previously only blank (id=0) was filtered.
  2. feat: add is_final support for streaming Paraformer

    • Short chunk acceptance: IsReady() allows chunks < 61 frames when is_final is set. DecodeStream() zero-pads them to chunk_size.
    • CIF tail flush: After the CIF integration loop, if residual alpha ≥ 0.6, the partially accumulated token is force-fired.
    • Reset clears flag: paraformer_is_final_ is reset to false in OnlineStream::Reset().
    • API: C (SherpaOnnxOnlineStreamSetFinalChunk), C++ (OnlineStream::SetParaformerFinalChunk), Python (set_paraformer_final_chunk).
    • Other language bindings (Java, Kotlin, Go, Swift, C#) can be added in follow-up.

Test results with official test_wavs

Using sherpa-onnx-streaming-paraformer-bilingual-zh-en/test_wavs/ (5 files from model package):

File Baseline is_final Diff
0.wav 昨天是monday today day is礼拜二the day after tomorrow是星期三 (same) SAME
1.wav ...always o s 什 ...always o s 什么意思啊 ✓ IMPROVED
2.wav ...接下来frequently苹繁的 (same) SAME
3.wav ...后后它是参一些商品 (same) SAME
8k.wav ...他现在正在教他 ...他现在正在教他的 ✓ IMPROVED

3 unchanged, 2 improved, 0 regressed.

The improvement is most visible when the audio's tail tokens fall in a short final chunk (<61 frames). 1.wav recovers 3 truncated characters ("什" → "什么意思啊").

Usage

// C API
SherpaOnnxOnlineStreamSetFinalChunk(stream);  // call before final decode
SherpaOnnxOnlineStreamInputFinished(stream);
while (SherpaOnnxIsOnlineStreamReady(recognizer, stream)) {
    SherpaOnnxDecodeOnlineStream(recognizer, stream);
}
# Python
stream.set_paraformer_final_chunk()
stream.input_finished()
while recognizer.is_ready(stream):
    recognizer.decode_stream(stream)

Files changed

  • sherpa-onnx/csrc/online-stream.h/.cc — SetParaformerFinalChunk / IsParaformerFinalChunk, reset in Reset()
  • sherpa-onnx/csrc/online-recognizer-paraformer-impl.h — IsReady, DecodeStream (short chunk + tail flush), sos/eos filter
  • sherpa-onnx/c-api/c-api.h/.cc — SherpaOnnxOnlineStreamSetFinalChunk
  • sherpa-onnx/c-api/sherpa-onnx-symbols-c.exp — macOS symbol export
  • sherpa-onnx/python/csrc/online-stream.cc — Python binding

The CIF decoder can occasionally produce sos (token id 1) or eos
(token id 2) tokens in its output, particularly when tail frames
generate extra CIF fires. Previously only blank (token id 0) was
filtered. Add sos and eos to the filter to prevent spurious tokens
in the recognition result.
When audio ends with fewer than chunk_size (61) frames remaining, the
last chunk is silently dropped and CIF residual alpha is never flushed,
causing tail tokens to be lost.

This commit adds SetParaformerFinalChunk() to OnlineStream. When set,
it enables two behaviors:

1. Short chunk acceptance: IsReady() allows chunks shorter than
   chunk_size frames. DecodeStream() zero-pads them to chunk_size
   before feeding to the encoder.

2. CIF tail flush: after the CIF integration loop, if the residual
   alpha >= 0.6 (empirically tuned), the partially accumulated token
   is force-fired. This recovers tail tokens that were >60%
   accumulated but never reached the fire threshold.

The flag is automatically cleared on Reset().

API surface:
- C++: OnlineStream::SetParaformerFinalChunk() / IsParaformerFinalChunk()
- C:   SherpaOnnxOnlineStreamSetFinalChunk()
- Python: OnlineStream.set_paraformer_final_chunk()

Other language bindings (Java, Kotlin, Go, Swift, C#) are not yet
updated and can be added in follow-up commits.

Test results on 73-entry zh corpus:
- Average CER: 0.0762 (baseline) -> 0.0656 (is_final), 14% improvement
- 15 improved, 58 unchanged, 0 regressed
- 0 mid-chunk differences (intermediate results 100% identical)
@dosubot dosubot Bot added the size:M This PR changes 30-99 lines, ignoring generated files. label Mar 11, 2026
@gemini-code-assist

Copy link
Copy Markdown

Summary of Changes

Hello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request introduces crucial enhancements to the streaming Paraformer model within Sherpa-Onnx, primarily by adding explicit support for handling the final audio chunk in a stream. This resolves previous limitations where short final segments and residual acoustic information were discarded, leading to reduced accuracy. The changes ensure that all audio data, especially at the end of an utterance, is properly processed and integrated into the transcription, resulting in a notable improvement in average character error rate (CER) for the final decode cycle.

Highlights

  • Streaming Paraformer Accuracy Improvement: Addressed issues in streaming Paraformer where the last audio chunk was dropped if shorter than 61 frames, and residual CIF alpha was not flushed, leading to lost tail tokens. This significantly improves accuracy for final chunks.
  • Token Filtering Enhancement: Expanded the token filtering in the streaming Paraformer decoder to include sos (start-of-sentence) and eos (end-of-sentence) tokens, in addition to the existing blank token filtering.
  • Short Chunk Acceptance: Implemented logic to allow IsReady() to accept audio chunks shorter than the standard chunk_size_ when the is_final flag is set, ensuring all audio is processed.
  • CIF Tail Token Flushing: Introduced a mechanism to force-fire partially accumulated tokens (residual alpha) if the integration value is above a certain threshold (0.6) when processing the final chunk, preventing loss of tail tokens.
  • New API for Final Chunk Signaling: Provided new C and Python APIs (SherpaOnnxOnlineStreamSetFinalChunk and stream.set_paraformer_final_chunk) to explicitly signal that a given chunk is the final one, enabling the new behaviors.
Changelog
  • sherpa-onnx/c-api/c-api.cc
    • Implemented SherpaOnnxOnlineStreamSetFinalChunk to set the final chunk flag.
  • sherpa-onnx/c-api/c-api.h
    • Declared SherpaOnnxOnlineStreamSetFinalChunk with documentation for its functionality.
  • sherpa-onnx/c-api/sherpa-onnx-symbols-c.exp
    • Exported the new C API function _SherpaOnnxOnlineStreamSetFinalChunk.
  • sherpa-onnx/csrc/online-recognizer-paraformer-impl.h
    • Modified IsReady to allow processing of short final audio chunks.
    • Updated DecodeStream to handle short final chunks by zero-padding and consuming all frames.
    • Added logic to DecodeStream for force-firing residual CIF tail tokens when is_final is true.
    • Extended token filtering in the decoder to skip sos (1) and eos (2) tokens in addition to blank (0).
    • Defined kCifTailFlushMinAlpha constant for tail flushing.
  • sherpa-onnx/csrc/online-stream.cc
    • Added paraformer_is_final_ member to track the final chunk status.
    • Implemented SetParaformerFinalChunk and IsParaformerFinalChunk methods.
    • Ensured paraformer_is_final_ is reset to false when the stream is reset.
  • sherpa-onnx/csrc/online-stream.h
    • Declared SetParaformerFinalChunk and IsParaformerFinalChunk methods for OnlineStream.
  • sherpa-onnx/python/csrc/online-stream.cc
    • Added Python binding for the new set_paraformer_final_chunk method.
Activity
  • The pull request introduces new functionality and includes test results demonstrating a significant improvement in Average CER (from 0.0762 to 0.0656) with zero regressions or mid-chunk differences, indicating a successful implementation.
  • Usage examples for both C and Python APIs are provided, guiding users on how to integrate the new is_final support.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@coderabbitai

coderabbitai Bot commented Mar 11, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

A new C API function SherpaOnnxOnlineStreamSetFinalChunk is introduced to signal the final chunk in streaming Paraformer processing. The implementation spans C API declarations, stream state tracking, enhanced Paraformer decoding logic for final chunks, symbol exports, and Python language bindings.

Changes

Cohort / File(s) Summary
C API Surface
sherpa-onnx/c-api/c-api.h, sherpa-onnx/c-api/c-api.cc, sherpa-onnx/c-api/sherpa-onnx-symbols-c.exp
Introduces new public C API function SherpaOnnxOnlineStreamSetFinalChunk() with declaration, implementation, and symbol export. Enables signaling final chunks from C clients.
Stream State Management
sherpa-onnx/csrc/online-stream.h, sherpa-onnx/csrc/online-stream.cc
Adds new public methods SetParaformerFinalChunk() and IsParaformerFinalChunk() to track final chunk state; introduces internal paraformer_is_final_ flag cleared during reset.
Paraformer Decoding Logic
sherpa-onnx/csrc/online-recognizer-paraformer-impl.h
Enhances final chunk handling: IsReady() logic allows short final chunks; DecodeStream() uses reduced chunk size for final chunks and pads with zeros; CIF tail-flush emits residual token when thresholds met; token filtering expands to skip sos and eos tokens. Adds kCifTailFlushMinAlpha constant.
Python Bindings
sherpa-onnx/python/csrc/online-stream.cc
Adds Python method set_paraformer_final_chunk() binding to the C\+\+ SetParaformerFinalChunk() with default is_final=true.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Poem

🐰 A final chunk arrives, the stream declares,
With CIF tails flushed and alpha-flush cares,
Short frames padded, tokens pruned just right,
Paraformer streams finish with delight! ✨

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Title check ✅ Passed The title 'feat: add is_final support for streaming Paraformer' directly and clearly describes the main change—adding is_final support for streaming Paraformer to fix tail token loss.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces is_final support for streaming Paraformer, effectively addressing issues with short final audio chunks being dropped and ensuring residual CIF alpha is flushed. The changes are well-implemented across the C-API, C++ core, and Python bindings. The logic for handling the final chunk, including padding, CIF tail flushing, and filtering of SOS/EOS tokens, is sound and clearly commented. I have a couple of suggestions to enhance API consistency and code readability.

Comment on lines +415 to 418
if (t == 0 || t == 1 || t == 2) {
// skip blank(0), sos(1), eos(2)
continue;
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Using magic numbers for special token IDs makes the code less readable and harder to maintain. It's better to define them as named constants.

You could add these constants to the OnlineRecognizerParaformerImpl class, for instance, near kCifTailFlushMinAlpha:

  static constexpr int32_t kBlankId = 0;
  static constexpr int32_t kSosId = 1;
  static constexpr int32_t kEosId = 2;

Then, you can use them here to make the condition more explicit.

Suggested change
if (t == 0 || t == 1 || t == 2) {
// skip blank(0), sos(1), eos(2)
continue;
}
if (t == kBlankId || t == kSosId || t == kEosId) {
// skip blank(0), sos(1), eos(2)
continue;
}

Comment on lines +354 to +357
void SherpaOnnxOnlineStreamSetFinalChunk(
const SherpaOnnxOnlineStream *stream) {
stream->impl->SetParaformerFinalChunk(true);
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

For consistency with the C++ (SetParaformerFinalChunk) and Python (set_paraformer_final_chunk) APIs, and to make it clear this function is specific to Paraformer models, consider renaming this function to SherpaOnnxOnlineStreamSetParaformerFinalChunk.

This change would also need to be applied to:

  • The function declaration in sherpa-onnx/c-api/c-api.h
  • The exported symbol in sherpa-onnx/c-api/sherpa-onnx-symbols-c.exp
void SherpaOnnxOnlineStreamSetParaformerFinalChunk(
    const SherpaOnnxOnlineStream *stream) {
  stream->impl->SetParaformerFinalChunk(true);
}

@csukuangfj

Copy link
Copy Markdown
Collaborator

Can you add tail padding instead of invoking SetParaformerFinalChunk?

@ZhaoChaoqun

Copy link
Copy Markdown
Contributor Author

Thanks for the review! I actually started with tail padding — that was our baseline approach. Unfortunately, padding alone cannot solve the two core problems:

Problem 1: Short final chunk is dropped

When the remaining audio is shorter than chunk_size (e.g., 12 frames left but chunk_size=32), IsReady() returns false and those frames are never processed. No amount of appended silence can help because the frames are already accepted — they just don't meet the chunk_size threshold.

Problem 2: CIF residual alpha is never flushed

Silence frames produce near-zero encoder output → near-zero CIF alpha contributions. The residual alpha accumulated from real speech never reaches the fire threshold (1.0), so the final token is permanently lost. Even 1 second of silence padding cannot push the residual over the threshold because silence adds essentially nothing to the alpha accumulator.

Evidence from our testing

Approach CER (67-entry zh corpus) Notes
Baseline (1s silence padding) 0.0856 Old approach — appended 16000 zero samples before InputFinished()
SetParaformerFinalChunk 0.0756 This PR — 15 items improved, 0 regressed

On the official bundled test wavs:

File Baseline (with implicit padding) With SetParaformerFinalChunk
1.wav 什 什么意思啊 (3 characters recovered)
8k.wav 教他 教他的 (1 character recovered)

The SetParaformerFinalChunk approach works because it:

  1. Allows short chunks to be processed (zero-padded to chunk_size internally)
  2. Force-fires a tail token when residual CIF alpha ≥ 0.6 (empirically tuned threshold)

This is analogous to the non-streaming Paraformer's tail handling in paraformer.cc, where the CIF module always flushes residual alpha at the end of the utterance. The streaming version currently lacks this mechanism — this PR adds it.

Happy to discuss alternative API designs if the concern is about API surface!

@csukuangfj

Copy link
Copy Markdown
Collaborator

IsReady() returns false and those frames are never processed

i mean,can you replace SetParaformerFinalChunk() with stream.AcceptWaveform(with some tail padding) ?

Whenever you call SetParaformerFinalChunk(), you replace it with stream.AcceptWaveform().


Happy to discuss alternative API designs if the concern is about API surface!

We try to avoid adding a specific API for a specific model. Instead, we want to fix #3101

@ZhaoChaoqun

Copy link
Copy Markdown
Contributor Author

Hi @csukuangfj,

I've implemented the SetOption/GetOption mechanism for OnlineStream/OfflineStream as discussed in #3101. The core changes are split into 3 small PRs for easier review:

  1. feat: add generic SetOption/GetOption to OnlineStream/OfflineStream (C++ core) #3307 — C++ core: add SetOption/HasOption/GetOption/GetOptionInt/GetOptionFloat to OnlineStream and OfflineStream (4 files, +138 lines)
  2. feat: add SetOption/GetOption C API and export symbols #3308 — C API: add SherpaOnnxOnline/OfflineStreamSetOption/GetOption + export symbols (3 files, +65 lines)
  3. feat: add SetOption/GetOption CXX wrapper #3309 — CXX wrapper: add SetOption/GetOption to cxx-api (2 files, +22 lines)

Once these are merged, follow-up PRs will migrate SetParaformerFinalChunk to use SetOption("is_final", "true") (#3310) and add language bindings (Python, Java, Kotlin, Go, C#, WASM).

Could you take a look and see if this is what you had in mind for #3101? Thanks!

@csukuangfj

Copy link
Copy Markdown
Collaborator

@ZhaoChaoqun

Thanks!

Please first fix the comments in the core PR (#3307). After it is merged, please rebase all of your remaining PRs to avoid adding the same changes.

@csukuangfj

Copy link
Copy Markdown
Collaborator

@ZhaoChaoqun

Can you fix the comments?

@ZhaoChaoqun

Copy link
Copy Markdown
Contributor Author

Sure, I'll fix the comments in PR #3307 first today, thanks!

@csukuangfj

Copy link
Copy Markdown
Collaborator

Thanks! Closing this PR since all other PRs are merged.

@csukuangfj csukuangfj closed this Mar 19, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:M This PR changes 30-99 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants