Add multilingual Nemotron-3.5 streaming ASR support - #3671
csukuangfj merged 5 commits into
Conversation
📝 WalkthroughWalkthroughEnd-to-end multilingual Nemotron streaming ASR: CI exporter + script produce ONNX artifacts with prompt metadata; C++ model detects multilingual encoders and accepts per-stream prompt IDs; stream-setOption bindings across SDKs; language-tag filtering and tests validate behavior. ChangesMultilingual Nemotron Streaming ASR
Sequence DiagramsequenceDiagram
participant Stream as OnlineStream
participant Recognizer as OnlineRecognizer
participant Model as OnlineTransducerNeMoModel
participant Encoder as ONNX Encoder
Stream->>Recognizer: setOption("language","ja")
Recognizer->>Model: GetLanguagePromptId("ja")
Model-->>Recognizer: prompt_id (e.g., 105)
Recognizer->>Model: RunEncoder(features, states, [105])
Model->>Encoder: Execute with prompt_index=[105]
Encoder-->>Model: encoded outputs
Recognizer->>Recognizer: filter language-tag tokens
Recognizer-->>Stream: final transcription (tags removed)
Estimated code review effort🎯 4 (Complex) | ⏱️ ~60 minutes Possibly related PRs
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Code Review
This pull request introduces support for prompt-conditioned multilingual Nemotron-3.5 streaming ASR models across C++, Python, Flutter, Swift, and C APIs. It includes an ONNX export script, metadata parsing for language prompt dictionaries, language tag filtering, and stream-level language configuration options. The code review feedback focuses on critical performance optimizations in hot paths, such as avoiding redundant string normalization, eliminating temporary vector allocations during encoder runs, and skipping language tag filtering when no tags are present in the decoded tokens.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
| OnlineRecognizerResult r; | ||
| if (language_tag_token_ids_.empty()) { | ||
| r = Convert(s->GetResult(), symbol_table_, frame_shift_ms, | ||
| subsampling_factor, s->GetCurrentSegment(), | ||
| s->GetNumFramesSinceStart()); | ||
| } else { | ||
| auto filtered = FilterLanguageTags(s->GetResult()); | ||
| r = Convert(filtered, symbol_table_, frame_shift_ms, subsampling_factor, | ||
| s->GetCurrentSegment(), s->GetNumFramesSinceStart()); | ||
| } |
There was a problem hiding this comment.
We can optimize this hot path by checking if there are actually any language tags in s->GetResult().tokens before calling FilterLanguageTags. Since language tags are only emitted occasionally, this check will succeed in almost all frames, completely avoiding the overhead of vector allocations and copying inside FilterLanguageTags.
Additionally, we can use the ternary operator directly to initialize OnlineRecognizerResult r, which allows the compiler to perform copy elision (NRVO) and avoids default-constructing and then copy-assigning r.
| OnlineRecognizerResult r; | |
| if (language_tag_token_ids_.empty()) { | |
| r = Convert(s->GetResult(), symbol_table_, frame_shift_ms, | |
| subsampling_factor, s->GetCurrentSegment(), | |
| s->GetNumFramesSinceStart()); | |
| } else { | |
| auto filtered = FilterLanguageTags(s->GetResult()); | |
| r = Convert(filtered, symbol_table_, frame_shift_ms, subsampling_factor, | |
| s->GetCurrentSegment(), s->GetNumFramesSinceStart()); | |
| } | |
| bool has_language_tag = false; | |
| if (!language_tag_token_ids_.empty()) { | |
| for (int32_t token : s->GetResult().tokens) { | |
| if (language_tag_token_ids_.count(token)) { | |
| has_language_tag = true; | |
| break; | |
| } | |
| } | |
| } | |
| OnlineRecognizerResult r = (!has_language_tag) | |
| ? Convert(s->GetResult(), symbol_table_, frame_shift_ms, | |
| subsampling_factor, s->GetCurrentSegment(), | |
| s->GetNumFramesSinceStart()) | |
| : Convert(FilterLanguageTags(s->GetResult()), symbol_table_, | |
| frame_shift_ms, subsampling_factor, | |
| s->GetCurrentSegment(), s->GetNumFramesSinceStart()); |
| std::vector<int64_t> prompt_id_buf; | ||
| Ort::Value prompt_id_tensor{nullptr}; | ||
| if (is_multilingual_) { | ||
| if (static_cast<int32_t>(language_prompt_ids.size()) == batch_size) { | ||
| prompt_id_buf = language_prompt_ids; | ||
| } else { | ||
| prompt_id_buf.assign(batch_size, default_prompt_id_); | ||
| } | ||
|
|
||
| std::array<int64_t, 1> prompt_id_shape{batch_size}; | ||
| prompt_id_tensor = Ort::Value::CreateTensor<int64_t>( | ||
| allocator_, prompt_id_shape.data(), prompt_id_shape.size()); | ||
| std::copy(prompt_id_buf.begin(), prompt_id_buf.end(), | ||
| prompt_id_tensor.GetTensorMutableData<int64_t>()); | ||
| inputs.push_back(std::move(prompt_id_tensor)); | ||
| } |
There was a problem hiding this comment.
We can avoid allocating a temporary std::vector<int64_t> prompt_id_buf on every single encoder run by writing directly to the tensor's mutable data pointer. This is a significant performance improvement since RunEncoder is called on every audio chunk during streaming.
Ort::Value prompt_id_tensor{nullptr};
if (is_multilingual_) {
std::array<int64_t, 1> prompt_id_shape{batch_size};
prompt_id_tensor = Ort::Value::CreateTensor<int64_t>(
allocator_, prompt_id_shape.data(), prompt_id_shape.size());
int64_t *p = prompt_id_tensor.GetTensorMutableData<int64_t>();
if (static_cast<int32_t>(language_prompt_ids.size()) == batch_size) {
std::copy(language_prompt_ids.begin(), language_prompt_ids.end(), p);
} else {
std::fill_n(p, batch_size, default_prompt_id_);
}
inputs.push_back(std::move(prompt_id_tensor));
}| int64_t GetLanguagePromptId(const std::string &language) const { | ||
| if (!is_multilingual_) { | ||
| return default_prompt_id_; | ||
| } | ||
|
|
||
| auto normalized = NormalizeLanguage(language); | ||
| if (normalized.empty() || normalized == "auto") { | ||
| return default_prompt_id_; | ||
| } | ||
|
|
||
| auto it = language_prompt_ids_.find(normalized); | ||
| if (it != language_prompt_ids_.end()) { | ||
| return it->second; | ||
| } | ||
|
|
||
| SHERPA_ONNX_LOGE( | ||
| "Unsupported language '%s' for multilingual NeMo transducer; using " | ||
| "auto", | ||
| language.c_str()); | ||
| return default_prompt_id_; | ||
| } |
There was a problem hiding this comment.
Since GetLanguagePromptId is called on every audio chunk for every stream, performing string normalization (NormalizeLanguage) every time is highly redundant and expensive. We can add a fast-path check to look up the raw language string directly in language_prompt_ids_ first, which will succeed in almost all cases and completely avoid string allocations and copies.
int64_t GetLanguagePromptId(const std::string &language) const {
if (!is_multilingual_) {
return default_prompt_id_;
}
auto it = language_prompt_ids_.find(language);
if (it != language_prompt_ids_.end()) {
return it->second;
}
auto normalized = NormalizeLanguage(language);
if (normalized.empty() || normalized == "auto") {
return default_prompt_id_;
}
it = language_prompt_ids_.find(normalized);
if (it != language_prompt_ids_.end()) {
return it->second;
}
SHERPA_ONNX_LOGE(
"Unsupported language '%s' for multilingual NeMo transducer; using "
"auto",
language.c_str());
return default_prompt_id_;
}There was a problem hiding this comment.
Actionable comments posted: 2
🧹 Nitpick comments (1)
sherpa-onnx/csrc/online-transducer-nemo-model.cc (1)
110-121: ⚡ Quick winOnly synthesize bare-language aliases when the base code is unambiguous.
AddBaseLanguageAliases()keeps the first locale it sees for each base code. Combined with the exporter’s sortedprompt_dictionary, bare values like"pt"or"zh"will map to whichever regional variant sorts first, not to an explicit default. That makesset_option("language", "...")silently choose the wrong prompt once a base language has multiple locales.Suggested fix
void AddBaseLanguageAliases( std::unordered_map<std::string, int64_t> *language_prompt_ids, const std::vector<std::pair<std::string, int64_t>> &ordered_prompt_ids) { + std::unordered_map<std::string, int32_t> counts; + for (const auto &p : ordered_prompt_ids) { + auto pos = p.first.find('-'); + if (pos != std::string::npos && pos != 0) { + ++counts[p.first.substr(0, pos)]; + } + } + for (const auto &p : ordered_prompt_ids) { auto pos = p.first.find('-'); if (pos == std::string::npos || pos == 0) { continue; } auto base = p.first.substr(0, pos); - language_prompt_ids->emplace(std::move(base), p.second); + if (counts[base] == 1) { + language_prompt_ids->emplace(std::move(base), p.second); + } } }🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@sherpa-onnx/csrc/online-transducer-nemo-model.cc` around lines 110 - 121, AddBaseLanguageAliases currently maps the first seen locale for a base language (e.g., "pt-BR") to the bare base ("pt"), which can silently pick the wrong variant; instead only synthesize a bare-language alias when that base is unambiguous. Change AddBaseLanguageAliases to first scan ordered_prompt_ids to count variants per base (extract base via p.first.substr(0,pos)), then on a second pass only emplace the base->id into language_prompt_ids when the count for that base equals 1; reference the function AddBaseLanguageAliases, the ordered_prompt_ids input vector and the language_prompt_ids output map when making this change.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.github/workflows/export-nemotron-3.5-asr-streaming-0.6b.yaml:
- Line 25: Replace all floating action tags with full commit SHAs: locate the
uses entries such as "uses: actions/checkout@v4" and the other three floating
tags referenced in the comment, and change them to pinned SHAs (e.g.,
actions/checkout@<full-commit-sha>) by finding the corresponding upstream
repository commit you want to pin to; update each "uses:" line so it uses the
full commit SHA instead of a tag or floating ref and commit the updated
workflow.
- Around line 178-181: The publish step currently always runs git commit -m
"first commit" which fails on no-op runs; change the workflow to check for
staged changes (e.g., use git diff --cached --quiet or git status --porcelain)
and only run git add ., git commit -m "first commit" and the git push command
when there are staged changes to commit; ensure the guard wraps the commit and
push (the git commit -m "first commit" and git push
https://csukuangfj2:$HF_TOKEN@huggingface.co/csukuangfj2/$m main) so repeated
runs with identical content are idempotent.
---
Nitpick comments:
In `@sherpa-onnx/csrc/online-transducer-nemo-model.cc`:
- Around line 110-121: AddBaseLanguageAliases currently maps the first seen
locale for a base language (e.g., "pt-BR") to the bare base ("pt"), which can
silently pick the wrong variant; instead only synthesize a bare-language alias
when that base is unambiguous. Change AddBaseLanguageAliases to first scan
ordered_prompt_ids to count variants per base (extract base via
p.first.substr(0,pos)), then on a second pass only emplace the base->id into
language_prompt_ids when the count for that base equals 1; reference the
function AddBaseLanguageAliases, the ordered_prompt_ids input vector and the
language_prompt_ids output map when making this change.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: 4c99930e-f029-43ab-82d0-295e4124be03
📒 Files selected for processing (15)
.github/workflows/export-nemotron-3.5-asr-streaming-0.6b.yamlc-api-examples/streaming-nemotron-c-api.ccxx-api-examples/streaming-nemotron-cxx-api.ccflutter-examples/streaming_asr/lib/streaming_asr.dartflutter/sherpa_onnx/lib/src/online_stream.dartflutter/sherpa_onnx/lib/src/sherpa_onnx_bindings.dartscripts/nemo/nemotron-3.5-asr-streaming-0.6b/README.mdscripts/nemo/nemotron-3.5-asr-streaming-0.6b/export_onnx.pysherpa-onnx/csrc/online-recognizer-transducer-nemo-impl.hsherpa-onnx/csrc/online-transducer-greedy-search-nemo-decoder.ccsherpa-onnx/csrc/online-transducer-nemo-model.ccsherpa-onnx/csrc/online-transducer-nemo-model.hsherpa-onnx/csrc/sherpa-onnx.ccsherpa-onnx/python/tests/test_online_recognizer.pyswift-api-examples/SherpaOnnx.swift
csukuangfj
left a comment
There was a problem hiding this comment.
Thank you for your contribution!
Left some minor comments. Otherwise, it looks great to me.
| @@ -0,0 +1,544 @@ | |||
| #!/usr/bin/env python3 | |||
| # Copyright 2026 Xiaomi Corp. (authors: Fangjun Kuang) | |||
There was a problem hiding this comment.
Can you update it to use your own information?
| return s; | ||
| } | ||
|
|
||
| bool ParseLanguagePromptEntry(const std::string &entry, std::string *language, |
There was a problem hiding this comment.
We are using https://github.com/nlohmann/json
Please see also
Please use json to parse it. No need to parse it manually.
There was a problem hiding this comment.
Done in fc071e3. Switched to nlohmann/json with the non-throwing parse (is_discarded), same as the qwen tokenizer uses, and removed the manual parser. Net 25 lines lighter.
| bool IsMultilingual() const { return is_multilingual_; } | ||
|
|
||
| int64_t GetLanguagePromptId(const std::string &language) const { | ||
| if (!is_multilingual_) { |
There was a problem hiding this comment.
| if (!is_multilingual_) { | |
| if (!is_multilingual_ || language.empty()) { |
| // std::vector<float> left_paddings(static_cast<int>(0.3 * sampling_rate)); | ||
| // s->AcceptWaveform(sampling_rate, left_paddings.data(), |
There was a problem hiding this comment.
Please uncomment these two lines.
I just tested that if you uncomment these two lines,
./build/bin/sherpa-onnx \
--encoder=./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/encoder.int8.onnx \
--decoder=./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/decoder.int8.onnx \
--joiner=./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/joiner.int8.onnx \
--tokens=./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/tokens.txt \
./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/test_wavs/en.wavgives the output:
Start to create recognizer
Recognizer created in 1.26584 s
./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/test_wavs/en.wav
Number of threads: 1, Elapsed seconds: 1.4, Audio duration (s): 7.2, Real time factor (RTF) = 1.4/7.2 = 0.19
The Triba tief then called for the boy and presented him with fifty pieces of gold
{ "text": " The Triba tief then called for the boy and presented him with fifty pieces of gold", "tokens": [" The", " T", "ri", "ba", " t", "ie", "f", " the", "n", " call", "ed", " for", " the", " bo", "y", " and", " pre", "s", "ent", "ed", " ", "h", "im", " with", " ", "fi", "f", "ty", " pie", "ce", "s", " of", " ", "g", "ol", "d"], "timestamps": [1.28, 2.00, 2.00, 2.16, 2.80, 2.80, 2.88, 2.96, 2.96, 3.12, 3.20, 3.28, 3.36, 3.52, 3.76, 4.48, 4.64, 4.72, 4.72, 4.80, 5.04, 5.04, 5.04, 5.20, 5.60, 5.60, 5.68, 5.68, 5.92, 6.08, 6.08, 6.32, 6.88, 6.88, 6.88, 6.96], "ys_probs": [], "lm_probs": [], "context_scores": [], "segment": 0, "words": [], "start_time": 0.00, "is_final": false, "is_eof": false}
If you don't, the output is
Start to create recognizer
Recognizer created in 2.21577 s
./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/test_wavs/en.wav
Number of threads: 1, Elapsed seconds: 1.6, Audio duration (s): 7.2, Real time factor (RTF) = 1.6/7.2 = 0.22
The tribal chief then called for the boy and present him with fifty pieces of
{ "text": " The tribal chief then called for the boy and present him with fifty pieces of", "tokens": [" The", " t", "ri", "ba", "l", " ", "chi", "e", "f", " the", "n", " call", "ed", " for", " the", " bo", "y", " and", " pre", "s", "ent", " ", "h", "im", " with", " ", "fi", "f", "ty", " pie", "ce", "s", " of"], "timestamps": [1.12, 1.76, 1.76, 1.84, 2.00, 2.32, 2.32, 2.32, 2.48, 2.64, 2.64, 2.80, 2.96, 3.04, 3.12, 3.36, 3.44, 4.00, 4.16, 4.32, 4.32, 4.72, 4.72, 4.72, 4.88, 5.20, 5.20, 5.28, 5.28, 5.60, 5.76, 5.76, 6.00], "ys_probs": [], "lm_probs": [], "context_scores": [], "segment": 0, "words": [], "start_time": 0.00, "is_final": false, "is_eof": false}
You can see the last part gold is dropped.
There was a problem hiding this comment.
Done in fc071e3. Reproduced your before/after on the 560ms package and I get your exact with-padding output, trailing "gold" included. Re-ran the ja forced/auto and English-only package checks too, all still good.
|
Models exported by the CI in this PR can be found at
|
- Parse the prompt_dictionary metadata (a JSON object) with nlohmann/json instead of the manual string parser - Return the default prompt id directly when the requested language is empty - Feed 0.3s of left padding in the sherpa-onnx CLI so trailing words are not dropped - Update the export script copyright header
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
sherpa-onnx/csrc/online-transducer-nemo-model.cc (1)
459-461:⚠️ Potential issue | 🟠 Major | ⚡ Quick winValidate the encoder input contract before enabling multilingual mode.
Line 459 only checks whether
prompt_indexexists anywhere inencoder_input_names_, butRunEncoder()always appends the prompt tensor as the last value. If a model exposesprompt_indexin any other slot, the names/value arrays no longer match and ORT will bind the wrong tensor to the wrong input. This should gate on the exact contract (size == 6and input 6 isprompt_index) and fail fast otherwise.Proposed fix
- is_multilingual_ = - std::find(encoder_input_names_.begin(), encoder_input_names_.end(), - "prompt_index") != encoder_input_names_.end(); + auto prompt_index_it = + std::find(encoder_input_names_.begin(), encoder_input_names_.end(), + "prompt_index"); + + is_multilingual_ = + encoder_input_names_.size() == 6 && + prompt_index_it == encoder_input_names_.begin() + 5; + + if (prompt_index_it != encoder_input_names_.end() && !is_multilingual_) { + SHERPA_ONNX_LOGE( + "Expected multilingual NeMo encoder input 6 to be prompt_index; " + "got an unexpected encoder input order."); + SHERPA_ONNX_EXIT(-1); + }🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@sherpa-onnx/csrc/online-transducer-nemo-model.cc` around lines 459 - 461, The current multilingual detection only checks for the presence of "prompt_index" in encoder_input_names_ (used to set is_multilingual_), which is unsafe because RunEncoder() always appends the prompt tensor as the last input; change the gate to validate the exact encoder input contract: require encoder_input_names_.size() == 6 and encoder_input_names_[5] == "prompt_index" before setting is_multilingual_; if the contract is not met, fail fast (log/error/throw) rather than enabling multilingual mode so RunEncoder() and ORT bindings remain consistent.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@sherpa-onnx/csrc/online-transducer-nemo-model.cc`:
- Around line 459-461: The current multilingual detection only checks for the
presence of "prompt_index" in encoder_input_names_ (used to set
is_multilingual_), which is unsafe because RunEncoder() always appends the
prompt tensor as the last input; change the gate to validate the exact encoder
input contract: require encoder_input_names_.size() == 6 and
encoder_input_names_[5] == "prompt_index" before setting is_multilingual_; if
the contract is not met, fail fast (log/error/throw) rather than enabling
multilingual mode so RunEncoder() and ORT bindings remain consistent.
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: 7fb2b6da-2493-404b-ac78-90cdd86e0d4a
📒 Files selected for processing (3)
scripts/nemo/nemotron-3.5-asr-streaming-0.6b/export_onnx.pysherpa-onnx/csrc/online-transducer-nemo-model.ccsherpa-onnx/csrc/sherpa-onnx.cc
🚧 Files skipped from review as they are similar to previous changes (1)
- scripts/nemo/nemotron-3.5-asr-streaming-0.6b/export_onnx.py

Fixes #3664
Adds support for nvidia/nemotron-3.5-asr-streaming-0.6b (40 locales, OpenMDW-1.1).
API shape is what you described in the issue: language is a plain string set per stream via the existing SherpaOnnxOnlineStreamSetOption(stream, "language", "ja"), the numeric prompt index stays internal, and detection is automatic from the prompt_index encoder input. English-only Nemotron models take exactly the same code path as before. tokenizer.model is converted to tokens.txt at export.
The language-to-index mapping lives in the encoder ONNX metadata (prompt_dictionary + auto_prompt_id), written at export time. Moving it to a separate file next to tokens.txt is a small change if you prefer that.
The model also emits language-tag tokens like
<en-US>in its output stream, which I noticed while testing the exported packages. The runtime filters those from text/tokens/timestamps. Tag ids are resolved once from the symbol table at init, no per-token string matching in the decode loop.Contents: the C++ runtime change, export script + CI workflow (modeled on the English Nemotron one), Python tests, a README, and setOption on the Flutter and Swift online stream wrappers (Java/Kotlin/Rust already had it).
Two gotchas worth knowing:
The multilingual Python test skips when the model package isn't present, like the other model tests, and activates once the package is published.
Validation: I ran the workflow end to end on my fork (NeMo install, export of 80/160/560/1120ms in fp32 and int8) and decoded the packages locally. Japanese with forced language, auto, and unset all transcribe correctly, English forced too. Roughly RTF 0.12 for int8 on an M-series Mac. The published English-only package (sherpa-onnx-nemotron-speech-streaming-en-0.6b-560ms-int8) decodes identically to master.
Summary by CodeRabbit
New Features
--languagehint; stream-option setter available in Flutter, Swift, C, and C++ examples.Documentation
Tests