Skip to content

Add multilingual Nemotron-3.5 streaming ASR support - #3671

Merged
csukuangfj merged 5 commits into
k2-fsa:masterfrom
JulianPscheid:feature/nemotron-multilingual-streaming
Jun 12, 2026
Merged

csukuangfj merged 5 commits into
k2-fsa:masterfrom
JulianPscheid:feature/nemotron-multilingual-streaming

Conversation

@JulianPscheid

@JulianPscheid JulianPscheid commented Jun 12, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #3664

Adds support for nvidia/nemotron-3.5-asr-streaming-0.6b (40 locales, OpenMDW-1.1).

API shape is what you described in the issue: language is a plain string set per stream via the existing SherpaOnnxOnlineStreamSetOption(stream, "language", "ja"), the numeric prompt index stays internal, and detection is automatic from the prompt_index encoder input. English-only Nemotron models take exactly the same code path as before. tokenizer.model is converted to tokens.txt at export.

The language-to-index mapping lives in the encoder ONNX metadata (prompt_dictionary + auto_prompt_id), written at export time. Moving it to a separate file next to tokens.txt is a small change if you prefer that.

The model also emits language-tag tokens like <en-US> in its output stream, which I noticed while testing the exported packages. The runtime filters those from text/tokens/timestamps. Tag ids are resolved once from the symbol table at init, no per-token string matching in the decode loop.

Contents: the C++ runtime change, export script + CI workflow (modeled on the English Nemotron one), Python tests, a README, and setOption on the Flutter and Swift online stream wrappers (Java/Kotlin/Rust already had it).

Two gotchas worth knowing:

  • The export needs NeMo from git main; the released nemo-toolkit doesn't have EncDecRNNTBPEModelWithPrompt yet. NVIDIA's model card also still shows the old pip #egg syntax, which current pip rejects, so the workflow uses the name @ URL form.
  • The export script asserts the encoder graph contract (input/output names including prompt_index) after the fp32 export, the metadata rewrite, and the int8 quantization. An earlier version of the script silently produced an unconditioned encoder that loaded fine and decoded to empty text. That can't slip through anymore.

The multilingual Python test skips when the model package isn't present, like the other model tests, and activates once the package is published.

Validation: I ran the workflow end to end on my fork (NeMo install, export of 80/160/560/1120ms in fp32 and int8) and decoded the packages locally. Japanese with forced language, auto, and unset all transcribe correctly, English forced too. Roughly RTF 0.12 for int8 on an M-series Mac. The published English-only package (sherpa-onnx-nemotron-speech-streaming-en-0.6b-560ms-int8) decodes identically to master.

Summary by CodeRabbit

  • New Features

    • Multilingual Nemotron streaming ASR: per-stream language hints (including auto-detect) and a new CLI --language hint; stream-option setter available in Flutter, Swift, C, and C++ examples.
    • Automated export & publishing pipeline: exports ONNX (fp32/int8) for multiple chunk sizes, bundles test audio, and uploads archives and model repos.
  • Documentation

    • Added Nemotron export & usage guide with language configuration examples.
  • Tests

    • Added English and multilingual streaming tests validating outputs and language-tag filtering.

@dosubot dosubot Bot added the size:XXL This PR changes 1000+ lines, ignoring generated files. label Jun 12, 2026
@coderabbitai

coderabbitai Bot commented Jun 12, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

End-to-end multilingual Nemotron streaming ASR: CI exporter + script produce ONNX artifacts with prompt metadata; C++ model detects multilingual encoders and accepts per-stream prompt IDs; stream-setOption bindings across SDKs; language-tag filtering and tests validate behavior.

Changes

Multilingual Nemotron Streaming ASR

Layer / File(s) Summary
NeMo export workflow and exporter script
.github/workflows/export-nemotron-3.5-asr-streaming-0.6b.yaml, scripts/nemo/nemotron-3.5-asr-streaming-0.6b/export_onnx.py, scripts/nemo/nemotron-3.5-asr-streaming-0.6b/README.md
CI workflow and Python exporter produce per-chunk FP32/INT8 ONNX artifacts, inject prompt_dictionary/auto_prompt_id metadata, extract SentencePiece tokens to tokens.txt, quantize to INT8, organize outputs by chunk size, archive test waves, and publish artifacts to GitHub releases and Hugging Face.
Stream option FFI & SDK bindings and CLI/examples
flutter/sherpa_onnx/lib/src/sherpa_onnx_bindings.dart, flutter/sherpa_onnx/lib/src/online_stream.dart, swift-api-examples/SherpaOnnx.swift, sherpa-onnx/csrc/sherpa-onnx.cc, c-api-examples/*, cxx-api-examples/*, flutter-examples/*
Adds FFI typedefs and binding lookup for SherpaOnnxOnlineStreamSetOption, implements OnlineStream.setOption (Flutter), Swift setOption(key:value:), CLI --language handling and example usages showing per-stream language hints.
C++ model: multilingual detection & prompt routing
sherpa-onnx/csrc/online-transducer-nemo-model.h, sherpa-onnx/csrc/online-transducer-nemo-model.cc
Detects multilingual encoder by prompt_index input, parses prompt_dictionary and auto_prompt_id from encoder metadata (normalizes keys, adds base-language aliases), exposes IsMultilingual() and GetLanguagePromptId(), and extends RunEncoder to accept language_prompt_ids and emit prompt_index tensor per batch.
Recognizer: stream option collection & language-tag filtering
sherpa-onnx/csrc/online-recognizer-transducer-nemo-impl.h
Collects per-stream language prompt IDs from stream options, passes them into model encoder runs, and filters out language-tag tokens (e.g., <ja>) from decoder outputs while preserving aligned score/timestamp vectors.
Decoder helper, tests, and small refactor
sherpa-onnx/csrc/online-transducer-greedy-search-nemo-decoder.cc, sherpa-onnx/python/tests/test_online_recognizer.py
Introduces BuildStateViews helper for decoder state wrapping; adds Python tests covering English-only and multilingual Nemotron streaming decodes and assertions that final text has no language-tag tokens.

Sequence Diagram

sequenceDiagram
  participant Stream as OnlineStream
  participant Recognizer as OnlineRecognizer
  participant Model as OnlineTransducerNeMoModel
  participant Encoder as ONNX Encoder
  Stream->>Recognizer: setOption("language","ja")
  Recognizer->>Model: GetLanguagePromptId("ja")
  Model-->>Recognizer: prompt_id (e.g., 105)
  Recognizer->>Model: RunEncoder(features, states, [105])
  Model->>Encoder: Execute with prompt_index=[105]
  Encoder-->>Model: encoded outputs
  Recognizer->>Recognizer: filter language-tag tokens
  Recognizer-->>Stream: final transcription (tags removed)
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • k2-fsa/sherpa-onnx#3309: Introduced foundational SherpaOnnxOnlineStreamSetOption/GetOption C API used by this PR's bindings.
  • k2-fsa/sherpa-onnx#3461: Adds JNI-exposed OnlineStream option methods that connect to the same option-setting codepath.
  • k2-fsa/sherpa-onnx#3044: Related changes to initial audio padding behavior in the stream input pipeline.

Suggested reviewers

  • csukuangfj

Poem

🐰 I hopped through prompts and ONNX night,

Packaged tokens and languages right.
Streams now speak ja, en, or auto,
C++ routes prompts in a neat flow.
Hooray — models sing in many tongues!

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 6.35% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'Add multilingual Nemotron-3.5 streaming ASR support' accurately and concisely describes the main feature addition: multilingual support for the Nemotron-3.5 streaming ASR model. It is clear, specific, and directly reflects the primary objective of the changeset.
Linked Issues check ✅ Passed The PR comprehensively addresses all coding requirements from issue #3664: multilingual detection via prompt_index input, language-to-prompt-id mapping via metadata, per-stream language configuration, runtime support, API exposure, language-tag token filtering, export tooling with validation, and example implementations across multiple language bindings.
Out of Scope Changes check ✅ Passed All changes are within scope of supporting multilingual Nemotron: C++ runtime updates, export script/workflow, Python tests, API wrappers (Flutter/Swift/C/C++), documentation, and CLI enhancements. No unrelated refactoring or feature drift detected.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces support for prompt-conditioned multilingual Nemotron-3.5 streaming ASR models across C++, Python, Flutter, Swift, and C APIs. It includes an ONNX export script, metadata parsing for language prompt dictionaries, language tag filtering, and stream-level language configuration options. The code review feedback focuses on critical performance optimizations in hot paths, such as avoiding redundant string normalization, eliminating temporary vector allocations during encoder runs, and skipping language tag filtering when no tags are present in the decoded tokens.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +105 to +114
OnlineRecognizerResult r;
if (language_tag_token_ids_.empty()) {
r = Convert(s->GetResult(), symbol_table_, frame_shift_ms,
subsampling_factor, s->GetCurrentSegment(),
s->GetNumFramesSinceStart());
} else {
auto filtered = FilterLanguageTags(s->GetResult());
r = Convert(filtered, symbol_table_, frame_shift_ms, subsampling_factor,
s->GetCurrentSegment(), s->GetNumFramesSinceStart());
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

We can optimize this hot path by checking if there are actually any language tags in s->GetResult().tokens before calling FilterLanguageTags. Since language tags are only emitted occasionally, this check will succeed in almost all frames, completely avoiding the overhead of vector allocations and copying inside FilterLanguageTags.

Additionally, we can use the ternary operator directly to initialize OnlineRecognizerResult r, which allows the compiler to perform copy elision (NRVO) and avoids default-constructing and then copy-assigning r.

Suggested change
OnlineRecognizerResult r;
if (language_tag_token_ids_.empty()) {
r = Convert(s->GetResult(), symbol_table_, frame_shift_ms,
subsampling_factor, s->GetCurrentSegment(),
s->GetNumFramesSinceStart());
} else {
auto filtered = FilterLanguageTags(s->GetResult());
r = Convert(filtered, symbol_table_, frame_shift_ms, subsampling_factor,
s->GetCurrentSegment(), s->GetNumFramesSinceStart());
}
bool has_language_tag = false;
if (!language_tag_token_ids_.empty()) {
for (int32_t token : s->GetResult().tokens) {
if (language_tag_token_ids_.count(token)) {
has_language_tag = true;
break;
}
}
}
OnlineRecognizerResult r = (!has_language_tag)
? Convert(s->GetResult(), symbol_table_, frame_shift_ms,
subsampling_factor, s->GetCurrentSegment(),
s->GetNumFramesSinceStart())
: Convert(FilterLanguageTags(s->GetResult()), symbol_table_,
frame_shift_ms, subsampling_factor,
s->GetCurrentSegment(), s->GetNumFramesSinceStart());

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in cd500b0.

Comment on lines +236 to +251
std::vector<int64_t> prompt_id_buf;
Ort::Value prompt_id_tensor{nullptr};
if (is_multilingual_) {
if (static_cast<int32_t>(language_prompt_ids.size()) == batch_size) {
prompt_id_buf = language_prompt_ids;
} else {
prompt_id_buf.assign(batch_size, default_prompt_id_);
}

std::array<int64_t, 1> prompt_id_shape{batch_size};
prompt_id_tensor = Ort::Value::CreateTensor<int64_t>(
allocator_, prompt_id_shape.data(), prompt_id_shape.size());
std::copy(prompt_id_buf.begin(), prompt_id_buf.end(),
prompt_id_tensor.GetTensorMutableData<int64_t>());
inputs.push_back(std::move(prompt_id_tensor));
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

We can avoid allocating a temporary std::vector<int64_t> prompt_id_buf on every single encoder run by writing directly to the tensor's mutable data pointer. This is a significant performance improvement since RunEncoder is called on every audio chunk during streaming.

    Ort::Value prompt_id_tensor{nullptr};
    if (is_multilingual_) {
      std::array<int64_t, 1> prompt_id_shape{batch_size};
      prompt_id_tensor = Ort::Value::CreateTensor<int64_t>(
          allocator_, prompt_id_shape.data(), prompt_id_shape.size());
      int64_t *p = prompt_id_tensor.GetTensorMutableData<int64_t>();
      if (static_cast<int32_t>(language_prompt_ids.size()) == batch_size) {
        std::copy(language_prompt_ids.begin(), language_prompt_ids.end(), p);
      } else {
        std::fill_n(p, batch_size, default_prompt_id_);
      }
      inputs.push_back(std::move(prompt_id_tensor));
    }

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in cd500b0.

Comment on lines +351 to +371
int64_t GetLanguagePromptId(const std::string &language) const {
if (!is_multilingual_) {
return default_prompt_id_;
}

auto normalized = NormalizeLanguage(language);
if (normalized.empty() || normalized == "auto") {
return default_prompt_id_;
}

auto it = language_prompt_ids_.find(normalized);
if (it != language_prompt_ids_.end()) {
return it->second;
}

SHERPA_ONNX_LOGE(
"Unsupported language '%s' for multilingual NeMo transducer; using "
"auto",
language.c_str());
return default_prompt_id_;
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Since GetLanguagePromptId is called on every audio chunk for every stream, performing string normalization (NormalizeLanguage) every time is highly redundant and expensive. We can add a fast-path check to look up the raw language string directly in language_prompt_ids_ first, which will succeed in almost all cases and completely avoid string allocations and copies.

  int64_t GetLanguagePromptId(const std::string &language) const {
    if (!is_multilingual_) {
      return default_prompt_id_;
    }

    auto it = language_prompt_ids_.find(language);
    if (it != language_prompt_ids_.end()) {
      return it->second;
    }

    auto normalized = NormalizeLanguage(language);
    if (normalized.empty() || normalized == "auto") {
      return default_prompt_id_;
    }

    it = language_prompt_ids_.find(normalized);
    if (it != language_prompt_ids_.end()) {
      return it->second;
    }

    SHERPA_ONNX_LOGE(
        "Unsupported language '%s' for multilingual NeMo transducer; using "
        "auto",
        language.c_str());
    return default_prompt_id_;
  }

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in cd500b0.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🧹 Nitpick comments (1)
sherpa-onnx/csrc/online-transducer-nemo-model.cc (1)

110-121: ⚡ Quick win

Only synthesize bare-language aliases when the base code is unambiguous.

AddBaseLanguageAliases() keeps the first locale it sees for each base code. Combined with the exporter’s sorted prompt_dictionary, bare values like "pt" or "zh" will map to whichever regional variant sorts first, not to an explicit default. That makes set_option("language", "...") silently choose the wrong prompt once a base language has multiple locales.

Suggested fix
 void AddBaseLanguageAliases(
     std::unordered_map<std::string, int64_t> *language_prompt_ids,
     const std::vector<std::pair<std::string, int64_t>> &ordered_prompt_ids) {
+  std::unordered_map<std::string, int32_t> counts;
+  for (const auto &p : ordered_prompt_ids) {
+    auto pos = p.first.find('-');
+    if (pos != std::string::npos && pos != 0) {
+      ++counts[p.first.substr(0, pos)];
+    }
+  }
+
   for (const auto &p : ordered_prompt_ids) {
     auto pos = p.first.find('-');
     if (pos == std::string::npos || pos == 0) {
       continue;
     }
 
     auto base = p.first.substr(0, pos);
-    language_prompt_ids->emplace(std::move(base), p.second);
+    if (counts[base] == 1) {
+      language_prompt_ids->emplace(std::move(base), p.second);
+    }
   }
 }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@sherpa-onnx/csrc/online-transducer-nemo-model.cc` around lines 110 - 121,
AddBaseLanguageAliases currently maps the first seen locale for a base language
(e.g., "pt-BR") to the bare base ("pt"), which can silently pick the wrong
variant; instead only synthesize a bare-language alias when that base is
unambiguous. Change AddBaseLanguageAliases to first scan ordered_prompt_ids to
count variants per base (extract base via p.first.substr(0,pos)), then on a
second pass only emplace the base->id into language_prompt_ids when the count
for that base equals 1; reference the function AddBaseLanguageAliases, the
ordered_prompt_ids input vector and the language_prompt_ids output map when
making this change.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In @.github/workflows/export-nemotron-3.5-asr-streaming-0.6b.yaml:
- Line 25: Replace all floating action tags with full commit SHAs: locate the
uses entries such as "uses: actions/checkout@v4" and the other three floating
tags referenced in the comment, and change them to pinned SHAs (e.g.,
actions/checkout@<full-commit-sha>) by finding the corresponding upstream
repository commit you want to pin to; update each "uses:" line so it uses the
full commit SHA instead of a tag or floating ref and commit the updated
workflow.
- Around line 178-181: The publish step currently always runs git commit -m
"first commit" which fails on no-op runs; change the workflow to check for
staged changes (e.g., use git diff --cached --quiet or git status --porcelain)
and only run git add ., git commit -m "first commit" and the git push command
when there are staged changes to commit; ensure the guard wraps the commit and
push (the git commit -m "first commit" and git push
https://csukuangfj2:$HF_TOKEN@huggingface.co/csukuangfj2/$m main) so repeated
runs with identical content are idempotent.

---

Nitpick comments:
In `@sherpa-onnx/csrc/online-transducer-nemo-model.cc`:
- Around line 110-121: AddBaseLanguageAliases currently maps the first seen
locale for a base language (e.g., "pt-BR") to the bare base ("pt"), which can
silently pick the wrong variant; instead only synthesize a bare-language alias
when that base is unambiguous. Change AddBaseLanguageAliases to first scan
ordered_prompt_ids to count variants per base (extract base via
p.first.substr(0,pos)), then on a second pass only emplace the base->id into
language_prompt_ids when the count for that base equals 1; reference the
function AddBaseLanguageAliases, the ordered_prompt_ids input vector and the
language_prompt_ids output map when making this change.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 4c99930e-f029-43ab-82d0-295e4124be03

📥 Commits

Reviewing files that changed from the base of the PR and between a2b7ab9 and 6b5abf8.

📒 Files selected for processing (15)
  • .github/workflows/export-nemotron-3.5-asr-streaming-0.6b.yaml
  • c-api-examples/streaming-nemotron-c-api.c
  • cxx-api-examples/streaming-nemotron-cxx-api.cc
  • flutter-examples/streaming_asr/lib/streaming_asr.dart
  • flutter/sherpa_onnx/lib/src/online_stream.dart
  • flutter/sherpa_onnx/lib/src/sherpa_onnx_bindings.dart
  • scripts/nemo/nemotron-3.5-asr-streaming-0.6b/README.md
  • scripts/nemo/nemotron-3.5-asr-streaming-0.6b/export_onnx.py
  • sherpa-onnx/csrc/online-recognizer-transducer-nemo-impl.h
  • sherpa-onnx/csrc/online-transducer-greedy-search-nemo-decoder.cc
  • sherpa-onnx/csrc/online-transducer-nemo-model.cc
  • sherpa-onnx/csrc/online-transducer-nemo-model.h
  • sherpa-onnx/csrc/sherpa-onnx.cc
  • sherpa-onnx/python/tests/test_online_recognizer.py
  • swift-api-examples/SherpaOnnx.swift

Comment thread .github/workflows/export-nemotron-3.5-asr-streaming-0.6b.yaml
Comment thread .github/workflows/export-nemotron-3.5-asr-streaming-0.6b.yaml

@csukuangfj csukuangfj left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for your contribution!
Left some minor comments. Otherwise, it looks great to me.

@@ -0,0 +1,544 @@
#!/usr/bin/env python3
# Copyright 2026 Xiaomi Corp. (authors: Fangjun Kuang)

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you update it to use your own information?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in fc071e3.

return s;
}

bool ParseLanguagePromptEntry(const std::string &entry, std::string *language,

@csukuangfj csukuangfj Jun 12, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We are using https://github.com/nlohmann/json

Please see also

Please use json to parse it. No need to parse it manually.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in fc071e3. Switched to nlohmann/json with the non-throwing parse (is_discarded), same as the qwen tokenizer uses, and removed the manual parser. Net 25 lines lighter.

bool IsMultilingual() const { return is_multilingual_; }

int64_t GetLanguagePromptId(const std::string &language) const {
if (!is_multilingual_) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
if (!is_multilingual_) {
if (!is_multilingual_ || language.empty()) {

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in fc071e3.

Comment thread sherpa-onnx/csrc/sherpa-onnx.cc Outdated
Comment on lines 135 to 136
// std::vector<float> left_paddings(static_cast<int>(0.3 * sampling_rate));
// s->AcceptWaveform(sampling_rate, left_paddings.data(),

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please uncomment these two lines.

I just tested that if you uncomment these two lines,

  ./build/bin/sherpa-onnx \
    --encoder=./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/encoder.int8.onnx \
    --decoder=./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/decoder.int8.onnx \
    --joiner=./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/joiner.int8.onnx \
    --tokens=./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/tokens.txt \
    ./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/test_wavs/en.wav

gives the output:

Start to create recognizer
Recognizer created in 1.26584 s
./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/test_wavs/en.wav
Number of threads: 1, Elapsed seconds: 1.4, Audio duration (s): 7.2, Real time factor (RTF) = 1.4/7.2 = 0.19
 The Triba tief then called for the boy and presented him with fifty pieces of gold
{ "text": " The Triba tief then called for the boy and presented him with fifty pieces of gold", "tokens": [" The", " T", "ri", "ba", " t", "ie", "f", " the", "n", " call", "ed", " for", " the", " bo", "y", " and", " pre", "s", "ent", "ed", " ", "h", "im", " with", " ", "fi", "f", "ty", " pie", "ce", "s", " of", " ", "g", "ol", "d"], "timestamps": [1.28, 2.00, 2.00, 2.16, 2.80, 2.80, 2.88, 2.96, 2.96, 3.12, 3.20, 3.28, 3.36, 3.52, 3.76, 4.48, 4.64, 4.72, 4.72, 4.80, 5.04, 5.04, 5.04, 5.20, 5.60, 5.60, 5.68, 5.68, 5.92, 6.08, 6.08, 6.32, 6.88, 6.88, 6.88, 6.96], "ys_probs": [], "lm_probs": [], "context_scores": [], "segment": 0, "words": [], "start_time": 0.00, "is_final": false, "is_eof": false}

If you don't, the output is

Start to create recognizer
Recognizer created in 2.21577 s
./sherpa-onnx-nemotron-3.5-asr-streaming-0.6b-560ms-int8-2026-06-11/test_wavs/en.wav
Number of threads: 1, Elapsed seconds: 1.6, Audio duration (s): 7.2, Real time factor (RTF) = 1.6/7.2 = 0.22
 The tribal chief then called for the boy and present him with fifty pieces of
{ "text": " The tribal chief then called for the boy and present him with fifty pieces of", "tokens": [" The", " t", "ri", "ba", "l", " ", "chi", "e", "f", " the", "n", " call", "ed", " for", " the", " bo", "y", " and", " pre", "s", "ent", " ", "h", "im", " with", " ", "fi", "f", "ty", " pie", "ce", "s", " of"], "timestamps": [1.12, 1.76, 1.76, 1.84, 2.00, 2.32, 2.32, 2.32, 2.48, 2.64, 2.64, 2.80, 2.96, 3.04, 3.12, 3.36, 3.44, 4.00, 4.16, 4.32, 4.32, 4.72, 4.72, 4.72, 4.88, 5.20, 5.20, 5.28, 5.28, 5.60, 5.76, 5.76, 6.00], "ys_probs": [], "lm_probs": [], "context_scores": [], "segment": 0, "words": [], "start_time": 0.00, "is_final": false, "is_eof": false}

You can see the last part gold is dropped.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Done in fc071e3. Reproduced your before/after on the 560ms package and I get your exact with-padding output, trailing "gold" included. Re-ran the ja forced/auto and English-only package checks too, all still good.

@csukuangfj

Copy link
Copy Markdown
Collaborator

Models exported by the CI in this PR can be found at
https://github.com/k2-fsa/sherpa-onnx/releases/tag/asr-models

Screenshot 2026-06-12 at 12 26 45

- Parse the prompt_dictionary metadata (a JSON object) with nlohmann/json
  instead of the manual string parser
- Return the default prompt id directly when the requested language is empty
- Feed 0.3s of left padding in the sherpa-onnx CLI so trailing words are
  not dropped
- Update the export script copyright header

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
sherpa-onnx/csrc/online-transducer-nemo-model.cc (1)

459-461: ⚠️ Potential issue | 🟠 Major | ⚡ Quick win

Validate the encoder input contract before enabling multilingual mode.

Line 459 only checks whether prompt_index exists anywhere in encoder_input_names_, but RunEncoder() always appends the prompt tensor as the last value. If a model exposes prompt_index in any other slot, the names/value arrays no longer match and ORT will bind the wrong tensor to the wrong input. This should gate on the exact contract (size == 6 and input 6 is prompt_index) and fail fast otherwise.

Proposed fix
-    is_multilingual_ =
-        std::find(encoder_input_names_.begin(), encoder_input_names_.end(),
-                  "prompt_index") != encoder_input_names_.end();
+    auto prompt_index_it =
+        std::find(encoder_input_names_.begin(), encoder_input_names_.end(),
+                  "prompt_index");
+
+    is_multilingual_ =
+        encoder_input_names_.size() == 6 &&
+        prompt_index_it == encoder_input_names_.begin() + 5;
+
+    if (prompt_index_it != encoder_input_names_.end() && !is_multilingual_) {
+      SHERPA_ONNX_LOGE(
+          "Expected multilingual NeMo encoder input 6 to be prompt_index; "
+          "got an unexpected encoder input order.");
+      SHERPA_ONNX_EXIT(-1);
+    }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@sherpa-onnx/csrc/online-transducer-nemo-model.cc` around lines 459 - 461, The
current multilingual detection only checks for the presence of "prompt_index" in
encoder_input_names_ (used to set is_multilingual_), which is unsafe because
RunEncoder() always appends the prompt tensor as the last input; change the gate
to validate the exact encoder input contract: require
encoder_input_names_.size() == 6 and encoder_input_names_[5] == "prompt_index"
before setting is_multilingual_; if the contract is not met, fail fast
(log/error/throw) rather than enabling multilingual mode so RunEncoder() and ORT
bindings remain consistent.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@sherpa-onnx/csrc/online-transducer-nemo-model.cc`:
- Around line 459-461: The current multilingual detection only checks for the
presence of "prompt_index" in encoder_input_names_ (used to set
is_multilingual_), which is unsafe because RunEncoder() always appends the
prompt tensor as the last input; change the gate to validate the exact encoder
input contract: require encoder_input_names_.size() == 6 and
encoder_input_names_[5] == "prompt_index" before setting is_multilingual_; if
the contract is not met, fail fast (log/error/throw) rather than enabling
multilingual mode so RunEncoder() and ORT bindings remain consistent.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 7fb2b6da-2493-404b-ac78-90cdd86e0d4a

📥 Commits

Reviewing files that changed from the base of the PR and between cd500b0 and fc071e3.

📒 Files selected for processing (3)
  • scripts/nemo/nemotron-3.5-asr-streaming-0.6b/export_onnx.py
  • sherpa-onnx/csrc/online-transducer-nemo-model.cc
  • sherpa-onnx/csrc/sherpa-onnx.cc
🚧 Files skipped from review as they are similar to previous changes (1)
  • scripts/nemo/nemotron-3.5-asr-streaming-0.6b/export_onnx.py

@csukuangfj csukuangfj left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:XXL This PR changes 1000+ lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Feature request: Support multilingual Nemotron-3.5 streaming ASR (prompt_index input)

2 participants