Add Qwen3-ASR support - #3399
Conversation
📝 WalkthroughWalkthroughThis PR introduces comprehensive Qwen3-ASR offline speech recognition support to sherpa-onnx across C, C++, and Python APIs. It adds model configuration structures, a tokenizer implementation, LLM-based recognizer pipeline with KV-cache management, and example programs in three languages. Supporting changes include FunASR-nano tokenizer template refactoring and WebAssembly size assertion updates. Changes
Sequence DiagramsequenceDiagram
participant User
participant App as Application
participant Recognizer as OfflineRecognizer
participant Model as OfflineQwen3ASRModel
participant Tokenizer as QwenAsrTokenizer
participant ONNX as ONNX Runtime
User->>App: Provide WAV file + config
App->>Recognizer: Create with Qwen3-ASR config
Recognizer->>Model: Load model (conv/encoder/decoder)
Recognizer->>Tokenizer: Initialize tokenizer
Recognizer->>ONNX: Create sessions
App->>Recognizer: Read WAV audio
Recognizer->>Model: ForwardConvFrontend(audio_features)
Model->>ONNX: Run conv_frontend session
ONNX-->>Model: Conv output
Recognizer->>Model: ForwardEncoder(conv_output, attention_mask)
Model->>ONNX: Run encoder session
ONNX-->>Model: Audio features
Recognizer->>Model: CreateEmptyKVCache()
Model-->>Recognizer: KV cache tensors
Recognizer->>Model: ForwardLLM(input_ids, audio_features, cache_kv)
Model->>ONNX: Run decoder session (autoregressive loop)
ONNX-->>Model: Logits + KV deltas
Model->>Recognizer: ApplyKvDeltaInplace(cache_kv, kv_delta)
Recognizer->>Tokenizer: Decode token_ids to text
Tokenizer-->>Recognizer: Recognized text
Recognizer-->>App: Recognition result
Estimated code review effort🎯 4 (Complex) | ⏱️ ~75 minutes Possibly related PRs
Suggested reviewers
Poem
🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Summary of ChangesHello, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed! This pull request significantly expands the capabilities of sherpa-onnx by integrating offline Qwen3-ASR models. It provides the necessary infrastructure, including model loading, configuration, and a specialized tokenizer, to enable accurate speech recognition for these models. The addition of C, C++, and Python examples ensures ease of adoption for developers looking to leverage Qwen3-ASR in their applications. Highlights
Using Gemini Code AssistThe full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips. Invoking Gemini You can request assistance from Gemini at any point by creating a comment using either
Customization To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a Limitations & Feedback Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for GitHub and other Google products, sign up here. Footnotes
|
|
Thanks! I am reviewing it. |
|
Warning Gemini encountered an error creating the review. You can try again by commenting |
There was a problem hiding this comment.
Actionable comments posted: 8
🧹 Nitpick comments (4)
sherpa-onnx/csrc/CMakeLists.txt (1)
260-263: Please add a small core smoke test for the new Qwen3 tokenizer/config path.These sources extend
sherpa-onnx-core, but the gtest list later in this file is unchanged. A round-trip tokenizer test or config-validation smoke test would keep this path from regressing silently.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@sherpa-onnx/csrc/CMakeLists.txt` around lines 260 - 263, Add a small gtest-based smoke test that exercises the new Qwen3 tokenizer/config code and wire it into the existing test target: create a test file (e.g., tests/test_qwen3_tokenizer.cc) that performs a simple round-trip or config-validation using the tokenizer/config types implemented in qwen-asr-tokenizer.cc and offline-qwen3-asr-model-config.cc (for example, serialize/deserialize a config or tokenize->detokenize a sample). Then add that test source to the gtest sources list in CMakeLists.txt (the same place where other test .cc files are listed) so the test is built and run with the existing sherpa-onnx-core tests. Ensure the test target name matches the project’s convention and the test includes the relevant headers for qwen-asr-tokenizer and offline-qwen3-asr-model-config.wasm/nodejs/sherpa-onnx-wasm-nodejs.cc (1)
48-49: Add a dedicated size invariant forSherpaOnnxOfflineQwen3ASRModelConfig.This file already keeps per-struct layout checks for the other offline model configs. Relying only on the aggregate
SherpaOnnxOfflineModelConfigassert makes future drift in the new Qwen3 layout harder to diagnose.Possible follow-up
static_assert(sizeof(SherpaOnnxOfflineFunASRNanoModelConfig) == 13 * 4, ""); +static_assert(sizeof(SherpaOnnxOfflineQwen3ASRModelConfig) == 9 * 4, ""); static_assert(sizeof(SherpaOnnxOfflineDolphinModelConfig) == 4, "");🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@wasm/nodejs/sherpa-onnx-wasm-nodejs.cc` around lines 48 - 49, Add a dedicated size invariant for SherpaOnnxOfflineQwen3ASRModelConfig similar to the existing per-struct layout checks: add a static_assert (or the same ASSERT used for other structs) that verifies sizeof(SherpaOnnxOfflineQwen3ASRModelConfig) equals the expected byte size so layout drift is caught independently of SherpaOnnxOfflineModelConfig; place this check alongside the other struct sizeof assertions and update the expected size constant to the correct value for Qwen3 if necessary.sherpa-onnx/c-api/cxx-api.h (1)
582-592: Optional: add field-level doc comments for API consistency.This new public config struct is the only nearby one without per-field comments; adding brief docs would keep
cxx-api.hstyle uniform.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@sherpa-onnx/c-api/cxx-api.h` around lines 582 - 592, Add brief per-field doc comments to the public struct OfflineQwen3ASRModelConfig to match the style of nearby config structs: document conv_frontend, encoder, decoder, tokenizer as path/identifier strings for model components; max_total_len and max_new_tokens as token-length limits with their default semantics; temperature and top_p as sampling/hyperparameter controls; and seed as RNG seed. Place single-line comments above each member (referencing OfflineQwen3ASRModelConfig and the specific members conv_frontend, encoder, decoder, tokenizer, max_total_len, max_new_tokens, temperature, top_p, seed) following the file’s existing comment style.sherpa-onnx/csrc/offline-stream.cc (1)
53-71: Consider extracting shared frame/mel option initialization.Lines 54-66 largely duplicate Lines 73-87. A small helper for common option assignment would reduce drift risk between whisper/non-whisper branches.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@sherpa-onnx/csrc/offline-stream.cc` around lines 53 - 71, Extract the duplicated frame/mel option assignment into a small helper (e.g., PopulateFrameAndMelOptions) that takes the source config and a reference to the target options (or the object owning them) and sets frame_opts.dither, snip_edges, samp_freq, frame_shift_ms, frame_length_ms, remove_dc_offset, window_type and mel_opts.num_bins, high_freq, low_freq, is_librosa; then replace the duplicated blocks with calls to this helper before constructing knf::WhisperFeatureOptions and creating whisper_fbank_ (use the same helper from the non-whisper branch as well so opts_.frame_opts and opts_.mel_opts are populated consistently).
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@sherpa-onnx/c-api/c-api.h`:
- Line 1068: You changed the public C ABI by adding the qwen3_asr field to
SherpaOnnxOfflineModelConfig (which is embedded in
SherpaOnnxOfflineRecognizerConfig); instead of extending the existing structs,
introduce a versioned config and constructor: create
SherpaOnnxOfflineModelConfigV2 that includes qwen3_asr, add a corresponding
factory/creation function (e.g., sherpa_onnx_offline_model_config_v2_create or a
recognizer create that accepts the V2 type), and keep the original
SherpaOnnxOfflineModelConfig and its create/recognizer APIs unchanged so older
binaries remain compatible; ensure the library checks which struct version is
passed (or exposes distinct symbols) and document that callers must opt into V2
to use qwen3_asr.
In `@sherpa-onnx/csrc/offline-qwen3-asr-model-config.cc`:
- Around line 83-95: The current validation checks only vocab.json and
merges.txt but not tokenizer_config.json, causing later initialization failures
in the tokenizer reader; update the same validation block that uses the
tokenizer variable and FileExists to also verify tokenizer_config.json exists,
and on failure call SHERPA_ONNX_LOGE with a similar message referencing
tokenizer.c_str() and return false so partially copied tokenizers are rejected
early.
In `@sherpa-onnx/csrc/offline-recognizer-impl.cc`:
- Around line 223-225: The manager-backed factory overload
OfflineRecognizerImpl::Create(Manager *mgr, const OfflineRecognizerConfig&
config, ...) is missing the Qwen3 dispatch; add the same conditional that checks
config.model_config.qwen3_asr.conv_frontend.empty() (or non-empty as done in the
other overload) and return
std::make_unique<OfflineRecognizerQwen3ASRImpl>(config) for that case so
Manager-backed creation can instantiate Qwen3-ASR; ensure you follow the exact
symbol names OfflineRecognizerImpl::Create, OfflineRecognizerConfig, and
OfflineRecognizerQwen3ASRImpl when locating the code to modify.
In `@sherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc`:
- Around line 821-846: The loop is passing the wrong cache position (cur_len
already includes the newly pushed token) causing KV cache misalignment; change
the cache position passed to next_cache_position to be context_len + (step - 1)
(or equivalently use cur_len - 1) so the first generated token is written at
position context_len; update the creation of next_cache_position (the Ort::Value
created for cache positions used in the decoder call) to use that corrected
value instead of cur_len, keeping other tensors (one_tensor,
next_attention_mask) unchanged.
In `@sherpa-onnx/csrc/qwen-asr-tokenizer.cc`:
- Around line 1144-1187: The tokenizer initialization currently doesn't validate
presence of required Qwen control tokens, so missing tokens like "<|audio_pad|>"
or "<|im_end|>" let construction succeed with invalid eos/pad IDs and later
break InitPromptTemplateIds; update the initialization (after ParseAddedTokens
and BuildSpecialTokens and before returning) to explicitly check token2id_ for
required markers ("<|audio_pad|>", "<|im_end|>", and any other Qwen control
tokens used by OfflineRecognizerQwen3ASRImpl::InitPromptTemplateIds) and if any
are missing, log an error via the same logging mechanism and return/fail
initialization (or throw) so the caller cannot proceed with invalid
eos_token_id_, pad_token_id_, or im_end_token_id_; reference token2id_,
eos_token_id_, pad_token_id_, im_end_token_id_ to locate where to add these
validations.
- Around line 1269-1297: Decode() currently calls DecodeBytes(buffer)
unconditionally and can emit a partial UTF-8 codepoint at sequence end; change
QwenAsrTokenizer::Decode to reuse the pending-byte handling from
GetTokenStringStreaming: before calling DecodeBytes(buffer) (both inside the
special-token branch and the final flush), inspect buffer bytes for the last
valid UTF-8 boundary, call DecodeBytes only on the complete-prefix portion, and
save the trailing incomplete bytes into a member (e.g., pending_bytes_ or the
existing streaming buffer) instead of appending them; ensure
IsSpecialToken/IsSkippableSpecialToken handling also flushes only complete bytes
and moves leftovers to the pending storage.
In `@sherpa-onnx/python/csrc/offline-model-config.cc`:
- Around line 68-71: The new binding added for OfflineModelConfig places
qwen3_asr as an optional positional argument that shifts existing positional
parameters; fix by making qwen3_asr keyword-only or preserve positional order
via a factory lambda. Locate the pybind11 init/binding for OfflineModelConfig in
offline-model-config.cc (the init signature that includes qwen3_asr,
telespeech_ctc, tokens, num_threads, etc.) and either insert a py::kw_only()
before the qwen3_asr py::arg(...) so callers must use qwen3_asr=... or replace
the init<> with a small lambda/factory that accepts the old positional
parameters in the original order and handles qwen3_asr as an optional keyword,
applying the same change for the other occurrence noted around the later block
(lines ~88-92).
In `@sherpa-onnx/python/sherpa_onnx/offline_recognizer.py`:
- Around line 415-416: The public constructor/factory in offline_recognizer.py
exposes a configurable feature_dim but
OfflineRecognizerQwen3ASRImpl::InitFeatConfig() always uses 128; fix this by
making the API truthful: either remove/hardcode feature_dim to 128 in the
factory/constructor or validate the passed feature_dim and raise a ValueError if
it is not 128. Update the spots that define the signature (the
constructor/factory that currently takes feature_dim and the other duplicated
location) to enforce or hardcode 128 so the Python config cannot diverge from
InitFeatConfig().
---
Nitpick comments:
In `@sherpa-onnx/c-api/cxx-api.h`:
- Around line 582-592: Add brief per-field doc comments to the public struct
OfflineQwen3ASRModelConfig to match the style of nearby config structs: document
conv_frontend, encoder, decoder, tokenizer as path/identifier strings for model
components; max_total_len and max_new_tokens as token-length limits with their
default semantics; temperature and top_p as sampling/hyperparameter controls;
and seed as RNG seed. Place single-line comments above each member (referencing
OfflineQwen3ASRModelConfig and the specific members conv_frontend, encoder,
decoder, tokenizer, max_total_len, max_new_tokens, temperature, top_p, seed)
following the file’s existing comment style.
In `@sherpa-onnx/csrc/CMakeLists.txt`:
- Around line 260-263: Add a small gtest-based smoke test that exercises the new
Qwen3 tokenizer/config code and wire it into the existing test target: create a
test file (e.g., tests/test_qwen3_tokenizer.cc) that performs a simple
round-trip or config-validation using the tokenizer/config types implemented in
qwen-asr-tokenizer.cc and offline-qwen3-asr-model-config.cc (for example,
serialize/deserialize a config or tokenize->detokenize a sample). Then add that
test source to the gtest sources list in CMakeLists.txt (the same place where
other test .cc files are listed) so the test is built and run with the existing
sherpa-onnx-core tests. Ensure the test target name matches the project’s
convention and the test includes the relevant headers for qwen-asr-tokenizer and
offline-qwen3-asr-model-config.
In `@sherpa-onnx/csrc/offline-stream.cc`:
- Around line 53-71: Extract the duplicated frame/mel option assignment into a
small helper (e.g., PopulateFrameAndMelOptions) that takes the source config and
a reference to the target options (or the object owning them) and sets
frame_opts.dither, snip_edges, samp_freq, frame_shift_ms, frame_length_ms,
remove_dc_offset, window_type and mel_opts.num_bins, high_freq, low_freq,
is_librosa; then replace the duplicated blocks with calls to this helper before
constructing knf::WhisperFeatureOptions and creating whisper_fbank_ (use the
same helper from the non-whisper branch as well so opts_.frame_opts and
opts_.mel_opts are populated consistently).
In `@wasm/nodejs/sherpa-onnx-wasm-nodejs.cc`:
- Around line 48-49: Add a dedicated size invariant for
SherpaOnnxOfflineQwen3ASRModelConfig similar to the existing per-struct layout
checks: add a static_assert (or the same ASSERT used for other structs) that
verifies sizeof(SherpaOnnxOfflineQwen3ASRModelConfig) equals the expected byte
size so layout drift is caught independently of SherpaOnnxOfflineModelConfig;
place this check alongside the other struct sizeof assertions and update the
expected size constant to the correct value for Qwen3 if necessary.
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: 74682a7b-5f3a-4267-9d1e-0088f5643bd6
📒 Files selected for processing (29)
c-api-examples/CMakeLists.txtc-api-examples/qwen3-asr-c-api.ccxx-api-examples/CMakeLists.txtcxx-api-examples/qwen3-asr-cxx-api.ccpython-api-examples/offline-qwen3-asr-decode-files.pysherpa-onnx/c-api/c-api.ccsherpa-onnx/c-api/c-api.hsherpa-onnx/c-api/cxx-api.ccsherpa-onnx/c-api/cxx-api.hsherpa-onnx/csrc/CMakeLists.txtsherpa-onnx/csrc/offline-model-config.ccsherpa-onnx/csrc/offline-model-config.hsherpa-onnx/csrc/offline-qwen3-asr-model-config.ccsherpa-onnx/csrc/offline-qwen3-asr-model-config.hsherpa-onnx/csrc/offline-qwen3-asr-model.ccsherpa-onnx/csrc/offline-qwen3-asr-model.hsherpa-onnx/csrc/offline-recognizer-impl.ccsherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.ccsherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.hsherpa-onnx/csrc/offline-stream.ccsherpa-onnx/csrc/qwen-asr-tokenizer.ccsherpa-onnx/csrc/qwen-asr-tokenizer.hsherpa-onnx/python/csrc/CMakeLists.txtsherpa-onnx/python/csrc/offline-model-config.ccsherpa-onnx/python/csrc/offline-qwen3-asr-model-config.ccsherpa-onnx/python/csrc/offline-qwen3-asr-model-config.hsherpa-onnx/python/sherpa_onnx/__init__.pysherpa-onnx/python/sherpa_onnx/offline_recognizer.pywasm/nodejs/sherpa-onnx-wasm-nodejs.cc
csukuangfj
left a comment
There was a problem hiding this comment.
Thanks! Looks great to me.
Left only some minor comments.
There was a problem hiding this comment.
♻️ Duplicate comments (2)
sherpa-onnx/csrc/qwen-asr-tokenizer.cc (2)
1246-1272:⚠️ Potential issue | 🟠 MajorReuse the streaming UTF-8 boundary handling in
Decode().
DecodeBytes(buffer)is appended directly here, so a truncated multibyte sequence can leak invalid UTF-8 intoresult.text. This also skips the boundary handling thatGetTokenStringStreaming()already applies around special tokens, so<|im_end|>/<|im_start|>can merge byte fragments across token boundaries.🔧 Suggested fix
std::string QwenAsrTokenizer::Decode(const std::vector<int64_t> &token_ids) { std::string ans; std::string buffer; + std::string pending_bytes; for (int64_t id : token_ids) { @@ const std::string &token = id2token_[static_cast<size_t>(id)]; if (IsSpecialToken(token)) { - if (IsSkippableSpecialToken(token)) { - continue; - } if (!buffer.empty()) { - ans.append(DecodeBytes(buffer)); + pending_bytes.append(DecodeBytes(buffer)); + ans.append( + ConsumeAvailableUtf8(&pending_bytes, /*flush_incomplete=*/true)); buffer.clear(); } + if (IsSkippableSpecialToken(token)) { + continue; + } ans.append(token); } else { buffer.append(token); @@ if (!buffer.empty()) { - ans.append(DecodeBytes(buffer)); + pending_bytes.append(DecodeBytes(buffer)); } + + ans.append( + ConsumeAvailableUtf8(&pending_bytes, /*flush_incomplete=*/true)); return ans; }🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@sherpa-onnx/csrc/qwen-asr-tokenizer.cc` around lines 1246 - 1272, QwenAsrTokenizer::Decode currently appends DecodeBytes(buffer) directly which can leak truncated UTF-8 and merge byte fragments across token boundaries; replace direct DecodeBytes calls with the streaming-safe helper GetTokenStringStreaming so boundary handling is reused: when hitting a special token (after flushing buffer) call GetTokenStringStreaming on the special token and on the flushed buffer instead of DecodeBytes, and likewise use GetTokenStringStreaming for the final buffer flush at the end of Decode; update references around id2token_, IsSpecialToken, and IsSkippableSpecialToken to ensure every appended piece goes through GetTokenStringStreaming.
1136-1179:⚠️ Potential issue | 🟠 MajorFail fast when the required Qwen control tokens are missing.
OfflineRecognizerQwen3ASRImpl::InitPromptTemplateIds()builds prompts with<|im_start|>,<|audio_start|>,<|audio_pad|>,<|audio_end|>, and<|im_end|>insherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.cc, Lines 318-332. If any of those tokens are absent here,Encode()silently byte-encodes the literal marker text instead, andeos_token_id_can remain-1, so generation proceeds with a broken prompt rather than failing early.🔧 Suggested guard
ParseAddedTokens(config_content, &token2id_, &id2token_); BuildSpecialTokens(token2id_, &special_tokens_); @@ - auto it = token2id_.find("<|im_end|>"); - if (it != token2id_.end()) { - eos_token_id_ = it->second; - im_end_token_id_ = it->second; - } + auto require_token = [&](const char *token, int64_t *out = nullptr) { + auto it2 = token2id_.find(token); + if (it2 == token2id_.end()) { + SHERPA_ONNX_LOGE("Missing required Qwen control token: %s", token); + SHERPA_ONNX_EXIT(-1); + } + if (out) { + *out = it2->second; + } + }; + + require_token("<|im_start|>"); + require_token("<|audio_start|>"); + require_token("<|audio_pad|>"); + require_token("<|audio_end|>"); + require_token("<|im_end|>", &eos_token_id_); + im_end_token_id_ = eos_token_id_; + + auto it = token2id_.find("<|padding|>");🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@sherpa-onnx/csrc/qwen-asr-tokenizer.cc` around lines 1136 - 1179, The tokenizer currently may leave control token IDs like eos_token_id_/im_end_token_id_ unset and allow generation to proceed; after ParseAddedTokens/BuildSpecialTokens, validate that token2id_ contains the required Qwen control tokens ("<|im_start|>", "<|audio_start|>", "<|audio_pad|>", "<|audio_end|>", "<|im_end|>") and set the corresponding IDs (e.g., eos_token_id_, im_end_token_id_, pad_token_id_ as appropriate); if any required token is missing, fail fast by returning an error or throwing (or calling processLogger/error handling used in your codebase) with a clear message identifying the missing token(s) so InitPromptTemplateIds()/Encode() cannot silently proceed with byte-encoding the literal markers.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Duplicate comments:
In `@sherpa-onnx/csrc/qwen-asr-tokenizer.cc`:
- Around line 1246-1272: QwenAsrTokenizer::Decode currently appends
DecodeBytes(buffer) directly which can leak truncated UTF-8 and merge byte
fragments across token boundaries; replace direct DecodeBytes calls with the
streaming-safe helper GetTokenStringStreaming so boundary handling is reused:
when hitting a special token (after flushing buffer) call
GetTokenStringStreaming on the special token and on the flushed buffer instead
of DecodeBytes, and likewise use GetTokenStringStreaming for the final buffer
flush at the end of Decode; update references around id2token_, IsSpecialToken,
and IsSkippableSpecialToken to ensure every appended piece goes through
GetTokenStringStreaming.
- Around line 1136-1179: The tokenizer currently may leave control token IDs
like eos_token_id_/im_end_token_id_ unset and allow generation to proceed; after
ParseAddedTokens/BuildSpecialTokens, validate that token2id_ contains the
required Qwen control tokens ("<|im_start|>", "<|audio_start|>",
"<|audio_pad|>", "<|audio_end|>", "<|im_end|>") and set the corresponding IDs
(e.g., eos_token_id_, im_end_token_id_, pad_token_id_ as appropriate); if any
required token is missing, fail fast by returning an error or throwing (or
calling processLogger/error handling used in your codebase) with a clear message
identifying the missing token(s) so InitPromptTemplateIds()/Encode() cannot
silently proceed with byte-encoding the literal markers.
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: 24a0f63f-3912-43cc-be70-79ddacccc9d7
📒 Files selected for processing (14)
sherpa-onnx/csrc/funasr-nano-tokenizer.ccsherpa-onnx/csrc/funasr-nano-tokenizer.hsherpa-onnx/csrc/offline-qwen3-asr-model-config.ccsherpa-onnx/csrc/offline-qwen3-asr-model-config.hsherpa-onnx/csrc/offline-qwen3-asr-model.ccsherpa-onnx/csrc/offline-recognizer-impl.ccsherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.ccsherpa-onnx/csrc/offline-recognizer-qwen3-asr-impl.hsherpa-onnx/csrc/qwen-asr-tokenizer.ccsherpa-onnx/csrc/qwen-asr-tokenizer.hsherpa-onnx/python/csrc/offline-model-config.ccsherpa-onnx/python/csrc/offline-qwen3-asr-model-config.ccsherpa-onnx/python/sherpa_onnx/offline_recognizer.pyswift-api-examples/SherpaOnnx.swift
🚧 Files skipped from review as they are similar to previous changes (6)
- sherpa-onnx/csrc/offline-recognizer-impl.cc
- sherpa-onnx/python/csrc/offline-qwen3-asr-model-config.cc
- sherpa-onnx/python/sherpa_onnx/offline_recognizer.py
- sherpa-onnx/python/csrc/offline-model-config.cc
- sherpa-onnx/csrc/offline-qwen3-asr-model-config.cc
- sherpa-onnx/csrc/offline-qwen3-asr-model-config.h
|
Warning Gemini encountered an error creating the review. You can try again by commenting |
csukuangfj
left a comment
There was a problem hiding this comment.
Thank you for your contribution!
Summary
Adds offline Qwen3-ASR to sherpa-onnx: model path, tokenizer wiring, and hooks in
sherpa-onnx-offlineandsherpa-onnx-vad-with-offline-asr. Tested with Qwen3-ASR-0.6B and Qwen3-ASR-1.7B ONNX builds.Models & export
Scope & caveats
provider=cpu.Example:
sherpa-onnx-offline(0.6B)Example output (
textmay include model-style prefixes such aslanguage Chinese<asr_text>):{ "lang": "", "emotion": "", "event": "", "text": "language Chinese<asr_text>开放时间:早上九点至下午五点。", "timestamps": [], "durations": [], "tokens": [ "language", " Chinese", "<asr_text>", "开放", "时间", ":", "早上", "九", "点", "至", "下午", "五", "点", "。" ], "ys_log_probs": [], "words": [] }Example:
sherpa-onnx-vad-with-offline-asr(1.7B)Example output (time range → transcript per VAD segment):
CPU benchmark (reference only)
Machine: Intel Xeon Gold 6448Y, x86_64, 32 cores — 4 threads, CPU provider.
zh.wavsong.wavfar_4.wavsong.wavSummary by CodeRabbit
Release Notes
New Features
Chores