Fix streaming paraformer feature frontend drift from its training config - #3821
Conversation
The offline paraformer sets hamming window, high_freq=0 and snip_edges=true (InitFeatConfig, added in k2-fsa#1148) to match the FunASR training frontend, but the online paraformer still used the fbank defaults (povey window, high_freq=-400, snip_edges=false), degrading streaming accuracy. Related: k2-fsa#3511
|
Caution The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased. |
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (2)
📝 WalkthroughWalkthroughParaformer feature extraction settings are centralized in ChangesParaformer streaming recognition
Estimated code review effort: 2 (Simple) | ~10 minutes Suggested labels: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
csukuangfj
left a comment
There was a problem hiding this comment.
Thank you for your contribution!
Problem
The offline paraformer configures the feature frontend the model was trained with (offline-recognizer-paraformer-impl.h:210-217):
The online paraformer sets only
normalize_samples, in both constructors, so streaming decoding runs on the fbank defaults (povey window, high_freq=-400, snip_edges=false) while the model's upstream config.yaml (modelscope speech_paraformer, online variant) specifieswindow: hamming. We noticed this while running the streaming and offline paraformer on the same audio: the streaming output was consistently worse, and the frontend difference turned out to be the reason. A user in #3511 found the same drift and worked around it with--high-freq=0 --snip-edges=true --window-type=hamming. Related: #3511.Fix
Give the online impl the same private
InitFeatConfig()and call it from both constructors. The gap dates to 6038e2a (#263), which copied the offline paraformer's then-current frontend handling; 25f0a10 (#1148) later upgraded only the offline impl.Tested
New
test_streaming_paraformerin test_online_recognizer.py decodes test_wavs/2.wav of sherpa-onnx-streaming-paraformer-bilingual-zh-en (greedy search, deterministic):The audio says 频繁 twice; the default frontend misrecognizes it both times. Individual tokens on other utterances can flip either way, but aggregate accuracy improves (about 10 errors fixed on a 4-minute zh recording, converging toward the offline paraformer's output on the same audio). The test fails on master without this change and passes with it; the full test_online_recognizer.py suite passes with no new skips.
./scripts/check_style_cpplint.shpasses.Only touches online-recognizer-paraformer-impl.h and its Python test; the offline paraformer and all other online recognizers are untouched.
Summary by CodeRabbit
Bug Fixes
Tests