Skip to content

Add hotwords support for Qwen3-ASR - #3434

Merged
csukuangfj merged 10 commits into
k2-fsa:masterfrom
Wasser1462:feature/qwen3-asr-hotword
Apr 1, 2026
Merged

csukuangfj merged 10 commits into
k2-fsa:masterfrom
Wasser1462:feature/qwen3-asr-hotword

Conversation

@Wasser1462

@Wasser1462 Wasser1462 commented Mar 27, 2026 •

Copy link
Copy Markdown
Collaborator

Add hotwords support for Qwen3-ASR

Adds optional hotwords to the offline Qwen3-ASR path so callers can bias recognition toward domain phrases (comma-separated text), wired through the C/C++ core and language bindings/examples.

Example (qwen3-asr-cxx-api, same audio, 4 threads)

With hotwords 骨质疏松症,打败咬死:

Loading hotwords: 骨质疏松症,打败咬死
...
text: 广西壮族自治区爱吃红鲤鱼与绿鲤鱼与驴的出租车司机,拉着苗族土家族自制粥,爱喝自制的刘奶奶榴莲牛奶的骨质疏松症患者,遇见别着喇叭的哑巴,打败咬死山前四十四棵死色柿子树的四十四个只石狮子之后,碰到年年练刘娘的牛郎,念着灰飞灰化肥发黑,会挥发走出山岗。官方网站设置组到广西壮族自治区首府南宁市民总医院就医。
Duration: 20.760s
Elapsed seconds: 6.038s
RTF = 0.291

Without hotwords (empty string):

Loading hotwords:
...
text: 广西壮族自治区爱吃红鲤鱼与绿鲤鱼与驴的出租车司机,拉着苗族土家族自治州爱喝自制的刘奶奶榴莲牛奶的古痴,输中症患者,遇见别着喇叭的哑巴,打败遥死山前四十四棵死色柿子树的四十自知十狮子之后,碰到年年练牛娘的牛郎,捏着灰黑灰化肥发黑灰发走出香港官方网站,设置组到广西壮族自治区首府南宁市民动医院就医。
Duration: 20.760s
Elapsed seconds: 5.826s
RTF = 0.281

Summary by CodeRabbit

  • New Features

    • Hotwords support added across languages and APIs: users can supply optional comma-separated hotwords (defaults to empty) and create recognition streams with hotwords.
    • New CLI option to pass hotwords for non-streaming ASR.
  • Bug Fixes / Improvements

    • Prompt construction now incorporates hotwords and provides diagnostics when context limits are tight.
    • Token post-processing refined for cleaner, more consistent final transcriptions.

@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Mar 27, 2026
@coderabbitai

coderabbitai Bot commented Mar 27, 2026 •

Copy link
Copy Markdown

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review
📝 Walkthrough

Walkthrough

Adds an optional Qwen3-ASR hotwords string across core, headers, bindings, FFI, language APIs, examples, and WASM; wires hotwords into config structs, stream creation, prompt/token construction, token postprocessing, allocation, and free paths.

Changes

Cohort / File(s) Summary
Core ASR implementation
sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.h, sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc, sherpa-onnx/csrc/offline-qwen3-asr-model-config.h, sherpa-onnx/csrc/offline-qwen3-asr-model-config.cc
Added hotwords member and ctor arg; BuildSourceIds now accepts hotwords; prompt construction, token accounting, hotword formatting, GenerateText behavior, and token postprocessing updated.
C/C++ public API
sherpa-onnx/c-api/c-api.h, sherpa-onnx/c-api/c-api.cc, sherpa-onnx/c-api/cxx-api.h, sherpa-onnx/c-api/cxx-api.cc
Extended C/C++ model config structs with hotwords; conversion/assignment logic added for C/C++ APIs.
Bindings & FFI (JNI, Python, Java, Kotlin, Rust, Go, .NET, Pascal, Swift, Flutter)
sherpa-onnx/jni/..., sherpa-onnx/java-api/..., sherpa-onnx/kotlin-api/..., sherpa-onnx/python/..., sherpa-onnx/rust/..., scripts/go/sherpa_onnx.go, scripts/dotnet/OfflineQwen3AsrModelConfig.cs, sherpa-onnx/pascal-api/sherpa_onnx.pas, swift-api-examples/*, flutter/...
Added hotwords fields/getters/setters, FFI marshaling (allocate/free), pybind exposure, new JNI createStreamWithHotwords binding, and updated defaults.
Examples & tests
*/*-api-examples/*, nodejs-addon-examples/*, nodejs-examples/*, python-api-examples/*
Updated many examples/tests to include hotwords (usually empty string ""); Python example adds CLI --hotwords.
WASM / NodeJS runtime
wasm/asr/sherpa-onnx-asr.js, wasm/nodejs/sherpa-onnx-wasm-nodejs.cc
Serialized hotwords into contiguous UTF‑8 buffer in JS init; adjusted struct layout and compile-time size assertions.
HarmonyOS / UI exports & native bridge
harmony-os/.../NonStreamingAsr.ets, harmony-os/.../Index.ets, harmony-os/.../non-streaming-asr.cc
Exported OfflineQwen3AsrModelConfig; wired hotwords parsing from JS/NAPI and added cleanup/free.
Glue & serialization
sherpa-onnx/csrc/offline-qwen3-asr-model-config.cc, various binding conversion helpers
Registered qwen3-asr-hotwords option, included hotwords in ToString()/serialization and cross-language struct layouts.

Sequence Diagram

sequenceDiagram
    participant App
    participant Config
    participant Recognizer
    participant Stream
    participant PromptBuilder
    participant Model

    App->>Config: build OfflineQwen3AsrModelConfig(hotwords)
    App->>Recognizer: createStream() or createStream(hotwords)
    Recognizer->>Stream: allocate stream (attach hotwords)
    Stream->>PromptBuilder: BuildSourceIds(hotwords, audio_len)
    PromptBuilder->>PromptBuilder: format hotwords, encode prefix+hotwords+suffix
    PromptBuilder-->>Stream: return token IDs
    App->>Recognizer: send audio -> Recognizer.generate
    Recognizer->>Model: infer tokens (uses stream/config hotwords)
    Model-->>Recognizer: token pieces
    Recognizer->>Stream: postprocess tokens (buffer, trim <asr_text>, emit)
    Recognizer-->>App: recognition result
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • csukuangfj

"🐰 I hopped through code, nibbling strings and streams,
Hotwords tucked in prompts, stitched into dreams.
From C to Swift, from Dart to Rust,
I sewed a prompt with gentle thrust.
Now Qwen3 listens — soft as beams."

🚥 Pre-merge checks | ✅ 2 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 6.06% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title directly and clearly summarizes the main change: adding optional hotwords support for the Qwen3-ASR model. It is concise, specific, and accurately reflects the primary purpose of the changeset across all modified files.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request adds hotword support to the Qwen3-ASR model across multiple language bindings and platforms, including C++, Go, Python, Java, and WASM. Key changes include a new "hotwords" configuration field, prompt formatting logic, and an API for per-utterance hotwords. Review feedback identifies a critical double-free bug in the HarmonyOS C++ bindings, suggests refactoring a string utility for better reuse, recommends replacing magic numbers with named constants, and highlights a performance optimization for string concatenation.

Comment thread harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/cpp/non-streaming-asr.cc Outdated
Comment on lines +63 to +70
static inline void Qwen3TrimInplace(std::string *s) {
if (!s) return;
auto &str = *s;
auto not_space = [](unsigned char c) { return !std::isspace(c); };
str.erase(str.begin(), std::find_if(str.begin(), str.end(), not_space));
str.erase(std::find_if(str.rbegin(), str.rend(), not_space).base(),
str.end());
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

This Qwen3TrimInplace function is a general-purpose string trimming utility. To promote code reuse and avoid future duplication, it should be moved to a common utility file (e.g., text-utils.h and text-utils.cc) for reuse across the codebase. The function could also be renamed to something more generic like TrimInplace.

References
  1. Move duplicated utility functions, such as Trim, to a common utility file (e.g., text-utils.h and text-utils.cc) for reuse across the codebase.

Comment on lines +793 to +794
const bool tight = hotword_tokens >= 48 ||
(one_audio_len > 0 && room < one_audio_len * 32);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The numbers 48 and 32 are used here without explanation. To improve code readability and maintainability, please define them as named constants with descriptive names that clarify their purpose. For example: kHotwordTokensWarningThreshold and kMinAudioChunksForWarning.

Comment on lines +1051 to +1058
for (size_t i = 0; i < all_tokens.size(); ++i) {
concat += all_tokens[i];
if (concat.find("<asr_text>") != std::string::npos) {
skip = i + 1;
stripped = true;
break;
}
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The string concatenation concat += all_tokens[i] inside the loop is inefficient as it may cause multiple reallocations, leading to quadratic complexity. This can be optimized. For instance, you could calculate the total size of all tokens first, reserve capacity for concat, and then append. A better approach might be to avoid building the full concat string just to find a marker, if possible.

@dosubot dosubot Bot added size:XL This PR changes 500-999 lines, ignoring generated files. and removed size:L This PR changes 100-499 lines, ignoring generated files. labels Mar 27, 2026
@dosubot dosubot Bot added size:L This PR changes 100-499 lines, ignoring generated files. and removed size:XL This PR changes 500-999 lines, ignoring generated files. labels Mar 27, 2026

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/cpp/non-streaming-asr.cc (1)

462-473: ⚠️ Potential issue | 🔴 Critical

Double-free bug: qwen3_asr fields are freed twice.

Lines 462-465 and 470-473 both delete the same qwen3_asr strings (conv_frontend, encoder, decoder, tokenizer). This causes undefined behavior and will likely crash at runtime.

Remove the duplicate deletion block at lines 470-473.

🐛 Proposed fix
  SHERPA_ONNX_DELETE_C_STR(c.model_config.qwen3_asr.conv_frontend);
  SHERPA_ONNX_DELETE_C_STR(c.model_config.qwen3_asr.encoder);
  SHERPA_ONNX_DELETE_C_STR(c.model_config.qwen3_asr.decoder);
  SHERPA_ONNX_DELETE_C_STR(c.model_config.qwen3_asr.tokenizer);
  SHERPA_ONNX_DELETE_C_STR(c.model_config.qwen3_asr.hotwords);

  SHERPA_ONNX_DELETE_C_STR(c.model_config.fire_red_asr_ctc.model);

-  SHERPA_ONNX_DELETE_C_STR(c.model_config.qwen3_asr.conv_frontend);
-  SHERPA_ONNX_DELETE_C_STR(c.model_config.qwen3_asr.encoder);
-  SHERPA_ONNX_DELETE_C_STR(c.model_config.qwen3_asr.decoder);
-  SHERPA_ONNX_DELETE_C_STR(c.model_config.qwen3_asr.tokenizer);
-
   SHERPA_ONNX_DELETE_C_STR(c.model_config.tokens);
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/cpp/non-streaming-asr.cc`
around lines 462 - 473, The diff shows the qwen3_asr C-strings being freed twice
(symbols: c.model_config.qwen3_asr.conv_frontend, .encoder, .decoder,
.tokenizer) which can cause a double-free; remove the duplicate deletion block
(the second group that repeats SHERPA_ONNX_DELETE_C_STR for those qwen3_asr
fields) so each qwen3_asr string is freed exactly once, leaving the single,
original SHERPA_ONNX_DELETE_C_STR calls and any other distinct frees (e.g.,
fire_red_asr_ctc.model) intact.
sherpa-onnx/python/csrc/offline-qwen3-asr-model-config.cc (1)

16-23: ⚠️ Potential issue | 🔴 Critical

Add missing hotwords parameter to constructor binding.

The C++ constructor accepts hotwords as the 10th parameter with a default value, but the Python binding (lines 16-23) exposes only 9 parameters. The Python factory (offline_recognizer.py:471) passes hotwords as a constructor kwarg, causing a TypeError at runtime.

Fix
   py::class_<PyClass>(*m, "OfflineQwen3ASRModelConfig")
       .def(py::init<const std::string &, const std::string &,
                     const std::string &, const std::string &, int32_t, int32_t,
-                    float, float, int32_t>(),
+                    float, float, int32_t, const std::string &>(),
            py::arg("conv_frontend") = "", py::arg("encoder") = "",
            py::arg("decoder") = "", py::arg("tokenizer") = "",
            py::arg("max_total_len") = 512, py::arg("max_new_tokens") = 128,
            py::arg("temperature") = 1e-6f, py::arg("top_p") = 0.8f,
-           py::arg("seed") = 42)
+           py::arg("seed") = 42, py::arg("hotwords") = "")
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@sherpa-onnx/python/csrc/offline-qwen3-asr-model-config.cc` around lines 16 -
23, The Python binding for the constructor in offline-qwen3-asr-model-config.cc
is missing the 10th parameter hotwords; update the py::init<> binding for the
class (the constructor binding shown with py::init<const std::string &, ...
int32_t>()) to include py::arg("hotwords") with the correct default (e.g. an
empty std::vector<std::string>()) as the 10th argument before the final seed arg
so the Python kwarg hotwords passed from offline_recognizer.py maps to the
native constructor.
🧹 Nitpick comments (1)
dart-api-examples/non-streaming-asr/bin/qwen3-asr.dart (1)

40-40: Expose hotwords via CLI instead of hard-coding an empty value.

On Line 40, hotwords is always '', so this example cannot actually demonstrate or use hotword biasing from callers. Consider wiring an optional --hotwords argument and passing it through.

Suggested diff
   final parser = ArgParser()
     ..addOption('conv-frontend', help: 'Path to the conv frontend model')
     ..addOption('encoder', help: 'Path to the encoder model')
     ..addOption('decoder', help: 'Path to the decoder model')
     ..addOption('tokenizer', help: 'Path to the tokenizer directory')
+    ..addOption(
+      'hotwords',
+      help: 'Comma-separated hotword phrases to bias recognition',
+    )
     ..addOption('input-wav', help: 'Path to input.wav to transcribe');
@@
   final tokenizer = res['tokenizer'] as String;
+  final hotwords = res['hotwords'] as String? ?? '';
   final inputWav = res['input-wav'] as String;
@@
-    hotwords: '',
+    hotwords: hotwords,
   );
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@dart-api-examples/non-streaming-asr/bin/qwen3-asr.dart` at line 40, The
example hard-codes hotwords: '' so callers cannot exercise hotword biasing; add
an optional CLI flag (e.g., --hotwords) in main() argument parsing, read its
value into a local variable (hotwords) and pass that variable into the ASR
request where hotwords: '' is currently set (referencing the hotwords field in
the request object and the main() function/argument parsing code) so the example
forwards user-provided hotwords instead of the empty string.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc`:
- Around line 408-415: CreateStream currently only calls
OfflineStream::SetOption("hotwords", hotwords) when hotwords is non-empty, so
passing an empty string cannot override recognizer-level qwen3_config.hotwords;
modify OfflineRecognizerQwen3ASRImpl::CreateStream to set the "hotwords" option
unconditionally (or add and set a separate explicit override flag option) so
GenerateText will see the empty-string override for that stream; make the same
change in the other CreateStream occurrence referenced (the block around the
second instance) to ensure per-utterance hotword disabling works.
- Around line 112-119: The Qwen3FormatHotwordsForPrompt function currently
rejoins parsed hotwords with a single space, which merges comma-separated
multi-word phrases; update Qwen3FormatHotwordsForPrompt to rejoin using a
delimiter that preserves phrase boundaries (e.g., ", " or "\n") instead of ' '
so that outputs from Qwen3ParseHotwordsCsv and NormalizeQwen3AsrHotwordSlashes
remain distinct when later consumed by BuildSourceIds; locate
Qwen3FormatHotwordsForPrompt and replace the join logic to use the chosen
delimiter consistently.

In `@sherpa-onnx/jni/offline-recognizer.cc`:
- Around line 540-562: Wrap
Java_com_k2fsa_sherpa_onnx_OfflineRecognizer_createStreamWithHotwords in
SafeJNI, call ValidatePointer on the casted sherpa_onnx::OfflineRecognizer* at
the top (same pattern as decode()), and if env->GetStringUTFChars(j_hotwords,
nullptr) returns nullptr, release any resources if needed and immediately return
0 instead of creating a stream; only call recognizer->CreateStream() or
recognizer->CreateStream(hotwords) when ValidatePointer succeeds and
GetStringUTFChars returns a valid pointer, and remember to ReleaseStringUTFChars
after copying to std::string.

---

Outside diff comments:
In `@harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/cpp/non-streaming-asr.cc`:
- Around line 462-473: The diff shows the qwen3_asr C-strings being freed twice
(symbols: c.model_config.qwen3_asr.conv_frontend, .encoder, .decoder,
.tokenizer) which can cause a double-free; remove the duplicate deletion block
(the second group that repeats SHERPA_ONNX_DELETE_C_STR for those qwen3_asr
fields) so each qwen3_asr string is freed exactly once, leaving the single,
original SHERPA_ONNX_DELETE_C_STR calls and any other distinct frees (e.g.,
fire_red_asr_ctc.model) intact.

In `@sherpa-onnx/python/csrc/offline-qwen3-asr-model-config.cc`:
- Around line 16-23: The Python binding for the constructor in
offline-qwen3-asr-model-config.cc is missing the 10th parameter hotwords; update
the py::init<> binding for the class (the constructor binding shown with
py::init<const std::string &, ... int32_t>()) to include py::arg("hotwords")
with the correct default (e.g. an empty std::vector<std::string>()) as the 10th
argument before the final seed arg so the Python kwarg hotwords passed from
offline_recognizer.py maps to the native constructor.

---

Nitpick comments:
In `@dart-api-examples/non-streaming-asr/bin/qwen3-asr.dart`:
- Line 40: The example hard-codes hotwords: '' so callers cannot exercise
hotword biasing; add an optional CLI flag (e.g., --hotwords) in main() argument
parsing, read its value into a local variable (hotwords) and pass that variable
into the ASR request where hotwords: '' is currently set (referencing the
hotwords field in the request object and the main() function/argument parsing
code) so the example forwards user-provided hotwords instead of the empty
string.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: fdc05231-1151-4721-bacc-bf2fc666a965

📥 Commits

Reviewing files that changed from the base of the PR and between 7d61865 and 59cce1f.

📒 Files selected for processing (42)
  • c-api-examples/qwen3-asr-c-api.c
  • cxx-api-examples/qwen3-asr-cxx-api.cc
  • dart-api-examples/non-streaming-asr/bin/qwen3-asr.dart
  • dotnet-examples/non-streaming-qwen3-asr-decode-files/Program.cs
  • dotnet-examples/vad-non-streaming-qwen3-asr/Program.cs
  • flutter/sherpa_onnx/lib/src/offline_recognizer.dart
  • flutter/sherpa_onnx/lib/src/sherpa_onnx_bindings.dart
  • go-api-examples/non-streaming-qwen3-asr-decode-files/main.go
  • harmony-os/SherpaOnnxHar/sherpa_onnx/Index.ets
  • harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/cpp/non-streaming-asr.cc
  • harmony-os/SherpaOnnxHar/sherpa_onnx/src/main/ets/components/NonStreamingAsr.ets
  • java-api-examples/NonStreamingDecodeFileQwen3Asr.java
  • nodejs-addon-examples/test_asr_non_streaming_qwen3_asr.js
  • nodejs-addon-examples/test_asr_non_streaming_qwen3_asr_async.js
  • nodejs-examples/test-offline-qwen3-asr.js
  • pascal-api-examples/non-streaming-asr/qwen3_asr.pas
  • python-api-examples/offline-qwen3-asr-decode-files.py
  • rust-api-examples/examples/qwen3_asr.rs
  • scripts/dotnet/OfflineQwen3AsrModelConfig.cs
  • scripts/go/sherpa_onnx.go
  • sherpa-onnx/c-api/c-api.cc
  • sherpa-onnx/c-api/c-api.h
  • sherpa-onnx/c-api/cxx-api.cc
  • sherpa-onnx/c-api/cxx-api.h
  • sherpa-onnx/csrc/offline-qwen3-asr-model-config.cc
  • sherpa-onnx/csrc/offline-qwen3-asr-model-config.h
  • sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc
  • sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.h
  • sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/OfflineQwen3AsrModelConfig.java
  • sherpa-onnx/java-api/src/main/java/com/k2fsa/sherpa/onnx/OfflineRecognizer.java
  • sherpa-onnx/jni/offline-recognizer.cc
  • sherpa-onnx/jni/sherpa-onnx-symbols.exp
  • sherpa-onnx/kotlin-api/OfflineRecognizer.kt
  • sherpa-onnx/pascal-api/sherpa_onnx.pas
  • sherpa-onnx/python/csrc/offline-qwen3-asr-model-config.cc
  • sherpa-onnx/python/sherpa_onnx/offline_recognizer.py
  • sherpa-onnx/rust/sherpa-onnx-sys/src/offline_asr.rs
  • sherpa-onnx/rust/sherpa-onnx/src/offline_asr.rs
  • swift-api-examples/SherpaOnnx.swift
  • swift-api-examples/qwen3-asr.swift
  • wasm/asr/sherpa-onnx-asr.js
  • wasm/nodejs/sherpa-onnx-wasm-nodejs.cc

Comment on lines +112 to +119
static std::string Qwen3FormatHotwordsForPrompt(const std::string &csv) {
std::vector<std::string> parts = Qwen3ParseHotwordsCsv(csv);
std::string s;
for (size_t i = 0; i < parts.size(); ++i) {
if (i) s += ' ';
s += NormalizeQwen3AsrHotwordSlashes(std::move(parts[i]));
}
return s;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Preserve phrase boundaries when rejoining parsed hotwords.

This formatter turns "foo,bar" into "foo bar", so multiple hotwords are merged before BuildSourceIds() inserts them into the system prompt. That breaks the advertised comma-separated hotword behavior for multi-word phrases. Rejoin with a delimiter that survives prompt construction, such as ", " or a newline.

Suggested fix
 static std::string Qwen3FormatHotwordsForPrompt(const std::string &csv) {
   std::vector<std::string> parts = Qwen3ParseHotwordsCsv(csv);
   std::string s;
   for (size_t i = 0; i < parts.size(); ++i) {
-    if (i) s += ' ';
+    if (i) s += ", ";
     s += NormalizeQwen3AsrHotwordSlashes(std::move(parts[i]));
   }
   return s;
 }
📝 Committable suggestion

‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.

Suggested change
static std::string Qwen3FormatHotwordsForPrompt(const std::string &csv) {
std::vector<std::string> parts = Qwen3ParseHotwordsCsv(csv);
std::string s;
for (size_t i = 0; i < parts.size(); ++i) {
if (i) s += ' ';
s += NormalizeQwen3AsrHotwordSlashes(std::move(parts[i]));
}
return s;
static std::string Qwen3FormatHotwordsForPrompt(const std::string &csv) {
std::vector<std::string> parts = Qwen3ParseHotwordsCsv(csv);
std::string s;
for (size_t i = 0; i < parts.size(); ++i) {
if (i) s += ", ";
s += NormalizeQwen3AsrHotwordSlashes(std::move(parts[i]));
}
return s;
}
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc` around lines 112 -
119, The Qwen3FormatHotwordsForPrompt function currently rejoins parsed hotwords
with a single space, which merges comma-separated multi-word phrases; update
Qwen3FormatHotwordsForPrompt to rejoin using a delimiter that preserves phrase
boundaries (e.g., ", " or "\n") instead of ' ' so that outputs from
Qwen3ParseHotwordsCsv and NormalizeQwen3AsrHotwordSlashes remain distinct when
later consumed by BuildSourceIds; locate Qwen3FormatHotwordsForPrompt and
replace the join logic to use the chosen delimiter consistently.

Comment thread sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc Outdated
Comment on lines +540 to +562
SHERPA_ONNX_EXTERN_C
JNIEXPORT jlong JNICALL
Java_com_k2fsa_sherpa_onnx_OfflineRecognizer_createStreamWithHotwords(
JNIEnv *env, jobject /*obj*/, jlong ptr, jstring j_hotwords) {
auto recognizer = reinterpret_cast<sherpa_onnx::OfflineRecognizer *>(ptr);
if (!j_hotwords) {
std::unique_ptr<sherpa_onnx::OfflineStream> s = recognizer->CreateStream();
sherpa_onnx::OfflineStream *p = s.release();
return (jlong)p;
}
const char *utf = env->GetStringUTFChars(j_hotwords, nullptr);
if (!utf) {
std::unique_ptr<sherpa_onnx::OfflineStream> s = recognizer->CreateStream();
sherpa_onnx::OfflineStream *p = s.release();
return (jlong)p;
}
std::string hotwords(utf);
env->ReleaseStringUTFChars(j_hotwords, utf);
std::unique_ptr<sherpa_onnx::OfflineStream> s =
recognizer->CreateStream(hotwords);
sherpa_onnx::OfflineStream *p = s.release();
return (jlong)p;
}

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

🧩 Analysis chain

🏁 Script executed:

# First, let's examine the entire file to find decode() and other JNI entry points
wc -l sherpa-onnx/jni/offline-recognizer.cc

Repository: k2-fsa/sherpa-onnx

Length of output: 103


🏁 Script executed:

# Search for SafeJNI and ValidatePointer definitions
rg "SafeJNI|ValidatePointer" sherpa-onnx/jni/

Repository: k2-fsa/sherpa-onnx

Length of output: 1530


🏁 Script executed:

# Find the decode() function to see how it handles error cases
rg -A 20 "Java_com_k2fsa_sherpa_onnx_OfflineRecognizer_decode\(" sherpa-onnx/jni/offline-recognizer.cc | head -50

Repository: k2-fsa/sherpa-onnx

Length of output: 969


🏁 Script executed:

# Search for GetStringUTFChars usage in the codebase to see patterns
rg "GetStringUTFChars" sherpa-onnx/jni/ -B 2 -A 5

Repository: k2-fsa/sherpa-onnx

Length of output: 16064


🏁 Script executed:

# Read SafeJNI and ValidatePointer definitions from common.h
sed -n '1,150p' sherpa-onnx/jni/common.h | grep -A 30 "SafeJNI\|ValidatePointer"

Repository: k2-fsa/sherpa-onnx

Length of output: 203


🏁 Script executed:

# Get more complete view of common.h, focusing on template/function definitions
cat sherpa-onnx/jni/common.h | head -200

Repository: k2-fsa/sherpa-onnx

Length of output: 10541


🏁 Script executed:

# Get the rest of ValidatePointer implementation
sed -n '200,250p' sherpa-onnx/jni/common.h

Repository: k2-fsa/sherpa-onnx

Length of output: 317


🏁 Script executed:

# Check keyword-spotter hotwords function (called createStreamWithKeywords)
rg -B 5 -A 25 "Java_com_k2fsa_sherpa_onnx_KeywordSpotter_createStreamWithKeywords" sherpa-onnx/jni/keyword-spotter.cc

Repository: k2-fsa/sherpa-onnx

Length of output: 44


🏁 Script executed:

# Check online-recognizer hotwords function
rg -B 5 -A 25 "Java_com_k2fsa_sherpa_onnx_OnlineRecognizer_createStreamWithHotwords" sherpa-onnx/jni/online-recognizer.cc

Repository: k2-fsa/sherpa-onnx

Length of output: 44


🏁 Script executed:

# Search more broadly in keyword-spotter.cc
rg "createStream" sherpa-onnx/jni/keyword-spotter.cc -B 5 -A 20

Repository: k2-fsa/sherpa-onnx

Length of output: 816


🏁 Script executed:

# Search in online-recognizer.cc
rg "createStream" sherpa-onnx/jni/online-recognizer.cc -B 5 -A 20

Repository: k2-fsa/sherpa-onnx

Length of output: 970


🏁 Script executed:

# Now let's see the actual hotwords function in offline-recognizer more carefully
sed -n '540,562p' sherpa-onnx/jni/offline-recognizer.cc

Repository: k2-fsa/sherpa-onnx

Length of output: 965


Wrap this function with SafeJNI and ValidatePointer, and return early on GetStringUTFChars() failure.

When GetStringUTFChars() returns nullptr, JNI has already set a pending exception (typically OutOfMemoryError). Creating and returning a stream object in this case is problematic—Java will receive the exception rather than the stream pointer, leaving the native stream unreleased. Follow the pattern used by decode(): wrap the entire function in SafeJNI, validate the recognizer pointer with ValidatePointer, and return 0 if string conversion fails instead of falling back to CreateStream().

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@sherpa-onnx/jni/offline-recognizer.cc` around lines 540 - 562, Wrap
Java_com_k2fsa_sherpa_onnx_OfflineRecognizer_createStreamWithHotwords in
SafeJNI, call ValidatePointer on the casted sherpa_onnx::OfflineRecognizer* at
the top (same pattern as decode()), and if env->GetStringUTFChars(j_hotwords,
nullptr) returns nullptr, release any resources if needed and immediately return
0 instead of creating a stream; only call recognizer->CreateStream() or
recognizer->CreateStream(hotwords) when ValidatePointer succeeds and
GetStringUTFChars returns a valid pointer, and remember to ReleaseStringUTFChars
after copying to std::string.

Comment thread scripts/go/sherpa_onnx.go Outdated
Comment thread sherpa-onnx/c-api/c-api.h Outdated
Comment thread sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc Outdated
Comment thread sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc Outdated
@Wasser1462
Wasser1462 requested a review from csukuangfj March 31, 2026 00:51
Comment thread c-api-examples/qwen3-asr-c-api.c Outdated
Comment thread python-api-examples/offline-qwen3-asr-decode-files.py Outdated
Comment thread sherpa-onnx/c-api/c-api.h Outdated
Comment thread sherpa-onnx/csrc/offline-qwen3-asr-model-config.cc Outdated
Comment thread sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc Outdated
Comment thread sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc Outdated
Comment thread sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc Outdated
Comment thread sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.h Outdated

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
sherpa-onnx/python/sherpa_onnx/offline_recognizer.py (1)

466-477: ⚠️ Potential issue | 🔴 Critical

Constructor kwarg mismatch will crash from_qwen3_asr at runtime.

Line 471 passes hotwords into OfflineQwen3ASRModelConfig(...), but the pybind constructor only accepts 9 parameters (conv_frontend, encoder, decoder, tokenizer, max_total_len, max_new_tokens, temperature, top_p, seed). This will raise a TypeError at runtime when from_qwen3_asr is called. While the C++ constructor supports hotwords and it is exposed as a read-write property in Python, it cannot be passed during initialization. Set it after construction instead:

Fix
         qwen3 = OfflineQwen3ASRModelConfig(
             conv_frontend=conv_frontend,
             encoder=encoder,
             decoder=decoder,
             tokenizer=tokenizer,
-            hotwords=hotwords,
             max_total_len=max_total_len,
             max_new_tokens=max_new_tokens,
             temperature=temperature,
             top_p=top_p,
             seed=seed,
         )
+        qwen3.hotwords = hotwords
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@sherpa-onnx/python/sherpa_onnx/offline_recognizer.py` around lines 466 - 477,
The call to OfflineQwen3ASRModelConfig in from_qwen3_asr passes an unsupported
keyword hotwords which will raise a TypeError because the pybind constructor
only accepts (conv_frontend, encoder, decoder, tokenizer, max_total_len,
max_new_tokens, temperature, top_p, seed); remove hotwords from the constructor
call and after creating qwen3 assign qwen3.hotwords = hotwords (since hotwords
is exposed as a read-write property) so the value is set post-construction.
♻️ Duplicate comments (1)
sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc (1)

56-58: ⚠️ Potential issue | 🟠 Major

Preserve hotword boundaries when rebuilding the prompt text.

Joining the parsed CSV with " " turns foo,bar into foo bar, so separate hotwords collapse before they reach the system prompt. That breaks multi-phrase hotwords like New York,Los Angeles. Use a delimiter that survives prompt construction, e.g. ", ".

💡 Suggested fix
 static std::string Qwen3FormatHotwordsForPrompt(const std::string &csv) {
   const std::vector<std::string> parts = SplitStringAndTrim(csv, ',');
-  return Join(parts, " ");
+  return Join(parts, ", ");
 }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc` around lines 56 - 58,
Qwen3FormatHotwordsForPrompt currently joins SplitStringAndTrim(csv, ',') with a
single space which collapses adjacent hotwords into one phrase; change the join
delimiter to a sequence that preserves CSV boundaries (for example ", ") so
multi-phrase hotwords like "New York,Los Angeles" remain distinct when passed to
the prompt; update the call in Qwen3FormatHotwordsForPrompt (which uses
SplitStringAndTrim and Join) to use ", " as the separator.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc`:
- Around line 976-993: The loop currently sets skip = i + 1 which drops any
transcript bytes that reside in the same token that completes the "<asr_text>"
marker; change the logic so when concat.find("<asr_text>") returns a position
you compute the marker end offset inside the concatenated stream and then
preserve the remainder of the triggering token. Concretely: when you detect the
marker (use auto pos = concat.find("<asr_text>"); auto marker_end = pos +
strlen("<asr_text>")), compute how many bytes of all_tokens[i] lie after that
marker (using concat.length() and all_tokens[i].length()), push the suffix of
all_tokens[i] that follows the marker into result.tokens (instead of dropping
it), set skip to i+1 for the remaining whole tokens and then push the rest as
before; keep symbols result.tokens, all_tokens, concat, skip, stripped and the
marker string "<asr_text>" to locate the change.

---

Outside diff comments:
In `@sherpa-onnx/python/sherpa_onnx/offline_recognizer.py`:
- Around line 466-477: The call to OfflineQwen3ASRModelConfig in from_qwen3_asr
passes an unsupported keyword hotwords which will raise a TypeError because the
pybind constructor only accepts (conv_frontend, encoder, decoder, tokenizer,
max_total_len, max_new_tokens, temperature, top_p, seed); remove hotwords from
the constructor call and after creating qwen3 assign qwen3.hotwords = hotwords
(since hotwords is exposed as a read-write property) so the value is set
post-construction.

---

Duplicate comments:
In `@sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc`:
- Around line 56-58: Qwen3FormatHotwordsForPrompt currently joins
SplitStringAndTrim(csv, ',') with a single space which collapses adjacent
hotwords into one phrase; change the join delimiter to a sequence that preserves
CSV boundaries (for example ", ") so multi-phrase hotwords like "New York,Los
Angeles" remain distinct when passed to the prompt; update the call in
Qwen3FormatHotwordsForPrompt (which uses SplitStringAndTrim and Join) to use ",
" as the separator.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: 2a9d7048-d5be-497e-bda4-2f1235447fa9

📥 Commits

Reviewing files that changed from the base of the PR and between 3807bec and 7c63577.

📒 Files selected for processing (10)
  • c-api-examples/qwen3-asr-c-api.c
  • cxx-api-examples/qwen3-asr-cxx-api.cc
  • python-api-examples/offline-qwen3-asr-decode-files.py
  • sherpa-onnx/c-api/c-api.h
  • sherpa-onnx/c-api/cxx-api.h
  • sherpa-onnx/csrc/offline-qwen3-asr-model-config.cc
  • sherpa-onnx/csrc/offline-qwen3-asr-model-config.h
  • sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc
  • sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.h
  • sherpa-onnx/python/sherpa_onnx/offline_recognizer.py
✅ Files skipped from review due to trivial changes (3)
  • sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.h
  • c-api-examples/qwen3-asr-c-api.c
  • sherpa-onnx/c-api/c-api.h
🚧 Files skipped from review as they are similar to previous changes (3)
  • python-api-examples/offline-qwen3-asr-decode-files.py
  • cxx-api-examples/qwen3-asr-cxx-api.cc
  • sherpa-onnx/csrc/offline-qwen3-asr-model-config.cc

Comment on lines +976 to +993
std::string concat;
size_t skip = 0;
bool stripped = false;
for (size_t i = 0; i < all_tokens.size(); ++i) {
concat += all_tokens[i];
if (concat.find("<asr_text>") != std::string::npos) {
skip = i + 1;
stripped = true;
break;
}
}
if (stripped && skip <= all_tokens.size()) {
result.tokens.reserve(all_tokens.size() - skip);
for (size_t i = skip; i < all_tokens.size(); ++i) {
result.tokens.push_back(std::move(all_tokens[i]));
}
} else {
result.tokens = std::move(all_tokens);

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟠 Major

Don’t drop transcript bytes that share the <asr_text> chunk.

skip = i + 1 discards the whole chunk that first completes <asr_text>. Since the marker search already spans multiple chunks, the triggering entry can also contain the first transcript bytes, e.g. "<asr_text>Hel", and those bytes never reach result.tokens.

💡 Suggested fix
+    static constexpr char kAsrTextMarker[] = "<asr_text>";
     std::string concat;
     size_t skip = 0;
     bool stripped = false;
     for (size_t i = 0; i < all_tokens.size(); ++i) {
+      const size_t prev_len = concat.size();
       concat += all_tokens[i];
-      if (concat.find("<asr_text>") != std::string::npos) {
+      size_t pos = concat.find(kAsrTextMarker);
+      if (pos != std::string::npos) {
         skip = i + 1;
         stripped = true;
+        const size_t marker_end = pos + std::strlen(kAsrTextMarker);
+        if (marker_end > prev_len) {
+          const size_t suffix_offset = marker_end - prev_len;
+          if (suffix_offset < all_tokens[i].size()) {
+            result.tokens.push_back(all_tokens[i].substr(suffix_offset));
+          }
+        }
         break;
       }
     }
     if (stripped && skip <= all_tokens.size()) {
-      result.tokens.reserve(all_tokens.size() - skip);
+      result.tokens.reserve(result.tokens.size() + all_tokens.size() - skip);
       for (size_t i = skip; i < all_tokens.size(); ++i) {
         result.tokens.push_back(std::move(all_tokens[i]));
       }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc` around lines 976 -
993, The loop currently sets skip = i + 1 which drops any transcript bytes that
reside in the same token that completes the "<asr_text>" marker; change the
logic so when concat.find("<asr_text>") returns a position you compute the
marker end offset inside the concatenated stream and then preserve the remainder
of the triggering token. Concretely: when you detect the marker (use auto pos =
concat.find("<asr_text>"); auto marker_end = pos + strlen("<asr_text>")),
compute how many bytes of all_tokens[i] lie after that marker (using
concat.length() and all_tokens[i].length()), push the suffix of all_tokens[i]
that follows the marker into result.tokens (instead of dropping it), set skip to
i+1 for the remaining whole tokens and then push the rest as before; keep
symbols result.tokens, all_tokens, concat, skip, stripped and the marker string
"<asr_text>" to locate the change.

@csukuangfj

Copy link
Copy Markdown
Collaborator

我更倾向于只修改 impl.cc 这一处,通过 stream 的 set/get/has option 机制来支持热词功能。这样无需改动 config,也不需要新增专门的字段,更不用调整其他编程语言的 binding。

之所以之前是在创建 stream 时通过参数传入 hotwords,是因为当时还没有 set/get/has option 这套机制。

如果在 model config 中指定 hotwords,会带来一个问题:每次识别时都会增加额外的计算开销(因为输入 token 数量变长)。

相比之下,只改动 impl.cc,实现更简单,代码量也更少。

@Wasser1462

Copy link
Copy Markdown
Collaborator Author

热词已经改为在 GenerateText 中通过 stream 的 SetOption / GetOption / HasOption("hotwords")读取

Comment thread wasm/asr/sherpa-onnx-asr.js Outdated
Comment thread swift-api-examples/SherpaOnnx.swift Outdated

@csukuangfj csukuangfj left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thank you for your contribution!

@csukuangfj
csukuangfj merged commit 36a2923 into k2-fsa:master Apr 1, 2026
26 of 27 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants