Skip to content

Export X-ASR models to sherpa-onnx - #3662

Merged
csukuangfj merged 4 commits into
k2-fsa:masterfrom
csukuangfj:export-x-asr
Jun 5, 2026
Merged

csukuangfj merged 4 commits into
k2-fsa:masterfrom
csukuangfj:export-x-asr

Conversation

@csukuangfj

@csukuangfj csukuangfj commented Jun 5, 2026 •

Copy link
Copy Markdown
Collaborator

See also https://github.com/Gilgamesh-J/X-ASR


https://github.com/k2-fsa/sherpa-onnx/releases/tag/asr-models

Screenshot 2026-06-05 at 17 43 21

Usage

Non-streaming

wget https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03.tar.bz2
tar xvf sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03.tar.bz2
rm sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03.tar.bz2
ls -lh sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/
total 356128
-rw-r--r--@ 1 fangjun  staff   116K  5 Jun 17:39 bpe.model
-rw-r--r--@ 1 fangjun  staff    11M  5 Jun 17:39 decoder-epoch-99-avg-1.onnx
-rw-r--r--@ 1 fangjun  staff   153M  5 Jun 17:39 encoder-epoch-99-avg-1.int8.onnx
-rw-r--r--@ 1 fangjun  staff   2.5M  5 Jun 17:39 joiner-epoch-99-avg-1.int8.onnx
-rw-r--r--@ 1 fangjun  staff    96B  5 Jun 17:39 README.md
-rwxr-xr-x@ 1 fangjun  staff   6.7K  5 Jun 17:39 test_onnx.py
drwxr-xr-x@ 7 fangjun  staff   224B  5 Jun 17:44 test_wavs
-rw-r--r--@ 1 fangjun  staff    57K  5 Jun 17:39 tokens.txt
./build/bin/sherpa-onnx-offline \
  --encoder=./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/encoder-epoch-99-avg-1.int8.onnx \
  --decoder=./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/decoder-epoch-99-avg-1.onnx \
  --joiner=./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/joiner-epoch-99-avg-1.int8.onnx \
  --tokens=./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/tokens.txt \
  ./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/test_wavs/0.wav
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/parse-options.cc:Read:373 ./build/bin/sherpa-onnx-offline --encoder=./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/encoder-epoch-99-avg-1.int8.onnx --decoder=./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/decoder-epoch-99-avg-1.onnx --joiner=./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/joiner-epoch-99-avg-1.int8.onnx --tokens=./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/tokens.txt ./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/test_wavs/0.wav 

OfflineRecognizerConfig(feat_config=FeatureExtractorConfig(sampling_rate=16000, feature_dim=80, low_freq=20, high_freq=-400, dither=0, normalize_samples=True, snip_edges=False), model_config=OfflineModelConfig(transducer=OfflineTransducerModelConfig(encoder_filename="./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/encoder-epoch-99-avg-1.int8.onnx", decoder_filename="./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/decoder-epoch-99-avg-1.onnx", joiner_filename="./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/joiner-epoch-99-avg-1.int8.onnx"), paraformer=OfflineParaformerModelConfig(model=""), nemo_ctc=OfflineNemoEncDecCtcModelConfig(model=""), whisper=OfflineWhisperModelConfig(encoder="", decoder="", language="", task="transcribe", tail_paddings=-1, enable_token_timestamps=False, enable_segment_timestamps=False), fire_red_asr=OfflineFireRedAsrModelConfig(encoder="", decoder=""), tdnn=OfflineTdnnModelConfig(model=""), zipformer_ctc=OfflineZipformerCtcModelConfig(model=""), wenet_ctc=OfflineWenetCtcModelConfig(model=""), sense_voice=OfflineSenseVoiceModelConfig(model="", language="auto", use_itn=False), moonshine=OfflineMoonshineModelConfig(preprocessor="", encoder="", uncached_decoder="", cached_decoder="", merged_decoder=""), dolphin=OfflineDolphinModelConfig(model=""), canary=OfflineCanaryModelConfig(encoder="", decoder="", src_lang="", tgt_lang="", use_pnc=True), cohere_transcribe=OfflineCohereTranscribeModelConfig(encoder="", decoder="", language="", use_punct=True, use_itn=True), omnilingual=OfflineOmnilingualAsrCtcModelConfig(model=""), funasr_nano=OfflineFunASRNanoModelConfig(encoder_adaptor="", llm="", embedding="", tokenizer="", system_prompt="You are a helpful assistant.", user_prompt="语音转写:", max_new_tokens=512, temperature=1e-06, top_p=0.8, seed=42, language="", itn=True, hotwords=""), medasr=OfflineMedAsrCtcModelConfig(model=""), fire_red_asr_ctc=OfflineFireRedAsrCtcModelConfig(model=""), qwen3_asr=OfflineQwen3ASRModelConfig(conv_frontend="", encoder="", decoder="", tokenizer="", hotwords="", max_total_len=512, max_new_tokens=128, temperature=1e-06, top_p=0.8, seed=42), telespeech_ctc="", tokens="./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/tokens.txt", num_threads=2, debug=False, provider="cpu", model_type="", modeling_unit="cjkchar", bpe_vocab=""), lm_config=OfflineLMConfig(model="", scale=0.5, lodr_scale=0.01, lodr_fst="", lodr_backoff_id=-1), ctc_fst_decoder_config=OfflineCtcFstDecoderConfig(graph="", max_active=3000), decoding_method="greedy_search", max_active_paths=4, hotwords_file="", hotwords_score=1.5, blank_penalty=0, rule_fsts="", rule_fars="", hr=HomophoneReplacerConfig(lexicon="", rule_fsts=""))
Creating recognizer ...
recognizer created in 1.159 s
Started
Done!

./sherpa-onnx-x-asr-zipformer-transducer-zh-en-punct-int8-2026-06-03/test_wavs/0.wav
----
num threads: 2
decoding method: greedy_search
Elapsed seconds: 0.121 s
Real time factor (RTF): 0.121 / 10.053 = 0.012
{"lang": "", "emotion": "", "event": "", "text": " 昨天是 Monday , today is 礼拜二 , the day after tomorrow 是星期三。", "timestamps": [0.96, 1.28, 1.72, 2.20, 2.24, 2.32, 2.56, 2.84, 3.64, 4.12, 5.00, 5.52, 5.76, 6.04, 6.60, 6.96, 7.24, 7.64, 8.00, 8.28, 8.40, 8.52, 8.60, 8.96, 9.28, 9.48, 9.72, 9.96], "durations": [], "tokens":[" 昨", " 天", " 是", " ", "M", "on", "da", "y", " ,", " today", " is", " 礼", " 拜", " 二", " ,", " the", " day", " after", " to", "m", "or", "ro", "w", " 是", " 星", " 期", " 三", " 。"], "ys_log_probs": [-0.004713, -0.031719, -0.001996, -0.552704, -0.003760, -0.276808, -0.642796, -0.014487, -0.726585, -0.458149, -0.060577, -0.034548, -0.713135, -0.342164, -0.489986, -0.420195, -0.061359, -0.002247, -0.066935, -0.022328, -0.067989, -0.205115, -0.022569, -0.001078, -0.001451, -0.002481, -0.010420, -0.550426], "words": []}

Streaming

wget https://github.com/k2-fsa/sherpa-onnx/releases/download/asr-models/sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05.tar.bz2
tar xvf sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05.tar.bz2
rm sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05.tar.bz2
ls -lh sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/
total 356160
-rw-r--r--@ 1 fangjun  staff   116K  5 Jun 17:38 bpe.model
-rw-r--r--@ 1 fangjun  staff    11M  5 Jun 17:38 decoder.onnx
-rw-r--r--@ 1 fangjun  staff   148M  5 Jun 17:38 encoder.int8.onnx
-rw-r--r--@ 1 fangjun  staff   2.5M  5 Jun 17:38 joiner.int8.onnx
-rw-r--r--@ 1 fangjun  staff    96B  5 Jun 17:38 README.md
-rw-r--r--@ 1 fangjun  staff   6.7K  5 Jun 17:38 test_onnx.py
drwxr-xr-x@ 6 fangjun  staff   192B  5 Jun 17:35 test_wavs
-rw-r--r--@ 1 fangjun  staff    57K  5 Jun 17:38 tokens.txt
./build/bin/sherpa-onnx \
  --encoder=./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/encoder.int8.onnx \
  --decoder=./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/decoder.onnx \
  --joiner=./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/joiner.int8.onnx \
  --tokens=./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/tokens.txt \
  --model-type=zipformer2 \
  ./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/test_wavs/0.wav
/Users/fangjun/open-source/sherpa-onnx/sherpa-onnx/csrc/parse-options.cc:Read:373 ./build/bin/sherpa-onnx --encoder=./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/encoder.int8.onnx --decoder=./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/decoder.onnx --joiner=./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/joiner.int8.onnx --tokens=./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/tokens.txt --model-type=zipformer2 ./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/test_wavs/0.wav 

OnlineRecognizerConfig(feat_config=FeatureExtractorConfig(sampling_rate=16000, feature_dim=80, low_freq=20, high_freq=-400, dither=0, normalize_samples=True, snip_edges=False), model_config=OnlineModelConfig(transducer=OnlineTransducerModelConfig(encoder="./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/encoder.int8.onnx", decoder="./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/decoder.onnx", joiner="./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/joiner.int8.onnx"), paraformer=OnlineParaformerModelConfig(encoder="", decoder=""), wenet_ctc=OnlineWenetCtcModelConfig(model="", chunk_size=16, num_left_chunks=4), zipformer2_ctc=OnlineZipformer2CtcModelConfig(model=""), nemo_ctc=OnlineNeMoCtcModelConfig(model=""), t_one_ctc=OnlineToneCtcModelConfig(model=""), provider_config=ProviderConfig(device=0, provider="cpu", cuda_config=CudaConfig(cudnn_conv_algo_search=1), trt_config=TensorrtConfig(trt_max_workspace_size=2147483647, trt_max_partition_iterations=10, trt_min_subgraph_size=5, trt_fp16_enable="True", trt_detailed_build_log="False", trt_engine_cache_enable="True", trt_engine_cache_path=".", trt_timing_cache_enable="True", trt_timing_cache_path=".",trt_dump_subgraphs="False" )), tokens="./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/tokens.txt", num_threads=1, warm_up=0, debug=False, model_type="zipformer2", modeling_unit="cjkchar", bpe_vocab=""), lm_config=OnlineLMConfig(model="", scale=0.5, lodr_scale=0.01, lodr_fst="", lodr_backoff_id=-1, shallow_fusion=True), endpoint_config=EndpointConfig(rule1=EndpointRule(must_contain_nonsilence=False, min_trailing_silence=2.4, min_utterance_length=0), rule2=EndpointRule(must_contain_nonsilence=True, min_trailing_silence=1.2, min_utterance_length=0), rule3=EndpointRule(must_contain_nonsilence=False, min_trailing_silence=0, min_utterance_length=20)), ctc_fst_decoder_config=OnlineCtcFstDecoderConfig(graph="", max_active=3000), enable_endpoint=True, max_active_paths=4, hotwords_score=1.5, hotwords_file="", decoding_method="greedy_search", blank_penalty=0, temperature_scale=2, rule_fsts="", rule_fars="", reset_encoder=False, hr=HomophoneReplacerConfig(lexicon="", rule_fsts=""))
Start to create recognizer
Recognizer created in 0.69465 s
./sherpa-onnx-x-asr-480ms-streaming-zipformer-transducer-zh-en-punct-int8-2026-06-05/test_wavs/0.wav
Number of threads: 1, Elapsed seconds: 0.35, Audio duration (s): 10, Real time factor (RTF) = 0.35/10 = 0.035
 昨天是 Monday , today is 礼拜二 , the day after tomorrow 是星期三
{ "text": " 昨天是 Monday , today is 礼拜二 , the day after tomorrow 是星期三", "tokens": [" 昨", " 天", " 是", " ", "M", "on", "da", "y", " ,", " today", " is", " 礼", " 拜", " 二", " ,", " the", " day", " after", " to", "m", "or", "ro", "w", " 是", " 星", " 期", " 三"], "timestamps": [1.12, 1.48, 1.80, 2.24, 2.28, 2.32, 2.60, 2.84, 3.96, 4.44, 5.28, 5.60, 5.84, 6.16, 6.76, 7.20, 7.40, 7.76, 8.16, 8.40, 8.52, 8.60, 8.76, 9.20, 9.60, 9.88, 10.28], "ys_probs": [-0.415120, -0.121638, -0.627824, -1.397655, -0.059075, -0.764276, -0.293221, -0.142320, -1.110069, -0.713361, -0.469155, -0.472615, -1.096575, -0.836425, -0.949083, -0.625053, -0.470173, -0.520002, -0.741348, -0.723526, -0.268818, -0.635884, -0.287995, -0.415925, -0.170761, -0.162517, -0.213999], "lm_probs": [], "context_scores": [], "segment": 0, "words": [], "start_time": 0.00, "is_final": false, "is_eof": false}

Summary by CodeRabbit

  • New Features

    • Added X-ASR model export pipelines for non-streaming and streaming variants.
    • New testing scripts for validating ONNX model inference.
    • Support for int8 quantization alongside fp32 models.
  • Improvements

    • Enhanced CJK text spacing normalization in recognition output.

@dosubot dosubot Bot added the size:L This PR changes 100-499 lines, ignoring generated files. label Jun 5, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces export and test scripts for streaming and non-streaming Zipformer transducer models, and refactors C++ source files to use a shared RemoveSpaceBetweenCjk utility. Feedback from the reviewer highlights potential runtime crashes in the test scripts due to hardcoded input data types (np.int32 and np.int64), recommending dynamic detection of these types from the ONNX models. Additionally, the reviewer suggests a more robust implementation for the load_tokens function in both test scripts to prevent ValueError crashes when parsing empty lines or space tokens.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

model_type = encoder_meta["model_type"]
assert model_type == "zipformer2", model_type

self.context_size = self.decoder.get_inputs()[0].shape[1]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The input data types for the encoder length and decoder inputs are currently hardcoded to np.int32. However, depending on whether --use-int32-inputs is set to 1 or 0 during export, these inputs can be either int32 or int64. Hardcoding them to np.int32 will cause a runtime type mismatch crash in ONNX Runtime when testing models exported with int64 inputs (such as the no-punct model in this PR). We should dynamically detect the expected input types from the ONNX model.

Suggested change
self.context_size = self.decoder.get_inputs()[0].shape[1]
self.context_size = self.decoder.get_inputs()[0].shape[1]
encoder_len_input = self.encoder.get_inputs()[1]
self.encoder_len_dtype = np.int64 if "int64" in encoder_len_input.type else np.int32
decoder_input = self.decoder.get_inputs()[0]
self.decoder_dtype = np.int64 if "int64" in decoder_input.type else np.int32

"""
Args: x: (1, T, C]
"""
x_len = np.array([x.shape[1]], dtype=np.int32)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Use the dynamically detected encoder length data type instead of hardcoded np.int32 to prevent runtime crashes.

Suggested change
x_len = np.array([x.shape[1]], dtype=np.int32)
x_len = np.array([x.shape[1]], dtype=self.encoder_len_dtype)

return out[0]

def run_decoder(self, hyp):
hyp = np.array([hyp], dtype=np.int32)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Use the dynamically detected decoder input data type instead of hardcoded np.int32 to prevent runtime crashes.

Suggested change
hyp = np.array([hyp], dtype=np.int32)
hyp = np.array([hyp], dtype=self.decoder_dtype)


decoder_meta = self.decoder.get_modelmeta().custom_metadata_map
self.context_size = int(decoder_meta["context_size"])
self.vocab_size = int(decoder_meta["vocab_size"])

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The input data type for the decoder is currently hardcoded to np.int64. To make the script robust to models exported with either int32 or int64 inputs, we should dynamically detect the expected input type from the ONNX model.

Suggested change
self.vocab_size = int(decoder_meta["vocab_size"])
self.vocab_size = int(decoder_meta["vocab_size"])
decoder_input = self.decoder.get_inputs()[0]
self.decoder_dtype = np.int64 if "int64" in decoder_input.type else np.int32

return out[0], out[1:]

def run_decoder(self, hyp):
hyp = np.array([hyp], dtype=np.int64)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

Use the dynamically detected decoder input data type instead of hardcoded np.int64 to prevent runtime crashes.

Suggested change
hyp = np.array([hyp], dtype=np.int64)
hyp = np.array([hyp], dtype=self.decoder_dtype)

Comment on lines +46 to +54
def load_tokens(filename):
ans = dict()
with open(filename, encoding="utf-8") as f:
for line in f:
t, i = line.strip().split()
ans[int(i)] = t
pass

return ans

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The current implementation of load_tokens can fail with a ValueError if a line is empty or if the token is a space character (which is common in BPE/character vocabularies, where a space token might be represented as <id>). Stripping and splitting such lines results in a single-element list, causing unpacking to fail. We should handle these cases robustly.

def load_tokens(filename):
    ans = dict()
    with open(filename, encoding="utf-8") as f:
        for line in f:
            line = line.rstrip("\r\n")
            parts = line.split()
            if len(parts) == 1:
                t = " "
                i = parts[0]
            elif len(parts) == 2:
                t, i = parts
            else:
                continue
            ans[int(i)] = t
    return ans

Comment on lines +44 to +52
def load_tokens(filename):
ans = dict()
with open(filename, encoding="utf-8") as f:
for line in f:
t, i = line.strip().split()
ans[int(i)] = t
pass

return ans

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The current implementation of load_tokens can fail with a ValueError if a line is empty or if the token is a space character. We should handle these cases robustly.

def load_tokens(filename):
    ans = dict()
    with open(filename, encoding="utf-8") as f:
        for line in f:
            line = line.rstrip("\r\n")
            parts = line.split()
            if len(parts) == 1:
                t = " "
                i = parts[0]
            elif len(parts) == 2:
                t, i = parts
            else:
                continue
            ans[int(i)] = t
    return ans

@coderabbitai

coderabbitai Bot commented Jun 5, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

This PR adds a complete CI/CD pipeline to automatically export X-ASR Zipformer transducer models in both non-streaming and streaming modes, creates Python test scripts for ONNX inference validation, disables conflicting legacy upload steps, and refactors CJK text spacing normalization into a shared utility module.

Changes

X-ASR Model Export & Test Infrastructure

Layer / File(s) Summary
Non-streaming Export Workflow and Scripts
.github/workflows/export-x-asr.yaml, scripts/zipformer-transducer/x-asr/export-non-streaming.sh, scripts/zipformer-transducer/x-asr/test_onnx_non_streaming.py
The non-streaming workflow job clones icefall, executes a Docker export script to generate encoder/decoder/joiner ONNX models in fp32 and int8 for punct/no-punct variants, and uploads tarballs to the release. The shell script downloads test audio and model checkpoints, runs the ONNX export twice for both variants, and validates inference. The Python test module loads audio, computes fbank features, and performs greedy decoding using the three ONNX session models.
Streaming Export Workflow and Scripts
.github/workflows/export-x-asr.yaml, scripts/zipformer-transducer/x-asr/export-streaming.sh, scripts/zipformer-transducer/x-asr/test_onnx_streaming.py
The streaming workflow job adds a matrix over four chunk sizes (160/480/960/1920 ms), runs streaming export and packaging per chunk size variant, and uploads all bundles to the same release tag. The shell script exports streaming ONNX models with normalized artifact naming (int8 encoder/joiner, fp32 decoder). The Python test module loads audio, computes online fbank features with padding, initializes encoder state tensors, and iterates over feature chunks performing streaming transducer decoding with frame-level joiner calls.
Disable Legacy Upload Steps
.github/workflows/upload-models.yaml
Disables the two existing non-streaming X-ASR zipformer transducer upload steps (punct and int8 variants) to avoid duplication now that exports are handled by the new workflow.

CJK Text Post-Processing Refactor

Layer / File(s) Summary
Consolidate RemoveSpaceBetweenCjk to Shared Utility
sherpa-onnx/csrc/online-recognizer-transducer-impl.h, sherpa-onnx/csrc/qnn/offline-recognizer-transducer-qnn-impl.h
The online transducer recognizer includes text-utils.h and applies RemoveSpaceBetweenCjk normalization to decoded text during the Convert result assembly. The QNN offline recognizer removes its duplicate local function definition, relying on the shared implementation via included headers.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • k2-fsa/sherpa-onnx#3660: Both PRs wire RemoveSpaceBetweenCjk from text-utils.h into the transducer text post-processing path during Convert for online/QNN recognition.
  • k2-fsa/sherpa-onnx#3656: Both PRs modify the same upload-models.yaml workflow steps for uploading X-ASR non-streaming zipformer transducer model artifacts.
  • k2-fsa/sherpa-onnx#3655: Both PRs modify the RemoveSpaceBetweenCjk(...) logic in sherpa-onnx/csrc/qnn/offline-recognizer-transducer-qnn-impl.h to consolidate the CJK text post-processing.

Poem

🐰 Workflows now export the X-ASR way,
With streaming chunks and tests in play,
Two variants dance, int8 and float,
While CJK spacing stays afloat,
Models packed in boxes, bundled tight—
Sherpa hops into the night! 🌙

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 4.17% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title 'Export X-ASR models to sherpa-onnx' clearly and concisely summarizes the main objective of the changeset, which adds workflows and scripts to export X-ASR models and assemble release bundles for the sherpa-onnx repository.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (6)
.github/workflows/export-x-asr.yaml (4)

1-11: ⚖️ Poor tradeoff

Consider workflow security hardening.

The workflow has several security posture gaps flagged by static analysis:

  • No permissions: block (defaults to broad GITHUB_TOKEN permissions)
  • Actions not pinned to commit SHAs (tags can be moved)
  • actions/checkout does not set persist-credentials: false (credentials remain accessible)

While not actively exploitable, these reduce defense-in-depth. Consider:

  1. Adding permissions: {} at workflow level and granting minimal permissions per job
  2. Pinning actions to commit SHAs for supply-chain integrity
  3. Setting persist-credentials: false on checkout steps
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/export-x-asr.yaml around lines 1 - 11, The workflow
"export-x-asr" lacks explicit permissions and has unpinned actions and an
insecure checkout; add a top-level permissions: {} to deny everything by
default, then grant minimal required permissions per job (e.g., permissions:
contents: read only where needed), replace action references (e.g.,
actions/checkout) with commit SHAs instead of tags to pin them, and set
persist-credentials: false on the actions/checkout step to avoid leaking
GITHUB_TOKEN to subsequent steps.

280-280: 💤 Low value

Remove redundant if: true condition.

Same as the non-streaming release step.

🧹 Proposed fix
       - name: Release
-        if: true
         uses: svenstaro/upload-release-action@v2
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/export-x-asr.yaml at line 280, Remove the redundant "if:
true" condition from the workflow step so it behaves like the non-streaming
release step; locate the literal line "if: true" in the export-x-asr GitHub
Actions YAML and delete that key from the step definition so the step runs by
default without the unnecessary conditional.

63-134: ⚖️ Poor tradeoff

Hardcoded release dates in directory names.

The directory names include 2026-06-03 (non-streaming) and 2026-06-05 (streaming), which will require manual updates for future releases. Consider using a variable like ${{ env.RELEASE_DATE }} set at the workflow level for easier maintenance.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/export-x-asr.yaml around lines 63 - 134, Replace the
hardcoded dates in the directory-name assignments (the lines that set d=... like
in the "Collect model files (punct fp32)", "Collect model files (punct int8)",
"Collect model files (no-punct fp32)", and "Collect model files (no-punct int8)"
steps) with a workflow-level variable (e.g. use env.RELEASE_DATE) and construct
d using that variable instead of literal "2026-06-03"/"2026-06-05"; also add
RELEASE_DATE to the workflow env section so all steps reference the same date
variable for future releases.

137-137: 💤 Low value

Remove redundant if: true condition.

The step will run by default without an explicit if: true. Per actionlint, constant conditions should be removed.

🧹 Proposed fix
       - name: Release
-        if: true
         uses: svenstaro/upload-release-action@v2
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In @.github/workflows/export-x-asr.yaml at line 137, Remove the redundant
constant condition "if: true" from the workflow step (the literal "if: true"
entry) so the step runs by default; locate the step block containing that "if:
true" and delete that line entirely to satisfy actionlint and simplify the
workflow.
scripts/zipformer-transducer/x-asr/export-non-streaming.sh (1)

43-84: ⚡ Quick win

Quote variable expansions to prevent word splitting.

Shellcheck flags unquoted $dir variables (lines 52-53, 63, 74-75, 85). While unlikely to cause issues in the controlled CI environment, quoting prevents potential word-splitting if the path ever contains spaces.

🛡️ Proposed fix
   --exp-dir $dir/punct \
   --tokens $dir/punct/tokens.txt \
+  --exp-dir "$dir/punct" \
+  --tokens "$dir/punct/tokens.txt" \

Apply the same fix to all instances of $dir in lines 52-53, 63, 74-75, and 85.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/zipformer-transducer/x-asr/export-non-streaming.sh` around lines 43 -
84, The $dir variable expansions in the shell invocations and the ls command
(e.g., the two calls to ./zipformer/export-streaming-as-non-streaming-onnx.py
and ls -lh $dir/punct) are unquoted and can suffer word-splitting; update every
occurrence of $dir (including $dir/punct, $dir/no-punct, and any uses in
--exp-dir and --tokens flags and the ls command) to be quoted (e.g., "$dir",
"$dir/punct", "$dir/no-punct") so paths with spaces are handled safely while
leaving the rest of the arguments unchanged.
scripts/zipformer-transducer/x-asr/export-streaming.sh (1)

48-112: ⚡ Quick win

Quote variable expansions to prevent word splitting.

Shellcheck flags unquoted $dir variables throughout the export commands and path references. While unlikely to cause issues in CI, quoting prevents potential word-splitting if paths contain spaces.

🛡️ Proposed fix
-  --exp-dir $dir/punct \
-  --tokens $dir/punct/tokens.txt \
+  --exp-dir "$dir/punct" \
+  --tokens "$dir/punct/tokens.txt" \

Apply the same fix to all instances of $dir in lines 55-56, 66, 68, 78, 89-90, 100, 102, and 112.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@scripts/zipformer-transducer/x-asr/export-streaming.sh` around lines 48 -
112, The script uses unquoted variable expansions of $dir (e.g. the --exp-dir
and --tokens args in the export-onnx-streaming.py calls, and commands like ls
-lh $dir/punct, pushd $dir/punct, mv ... $dir/punct, rm ... $dir/no-punct, popd)
which can cause word-splitting; update every occurrence to use quoted expansions
(\"$dir\", \"$dir/punct\", \"$dir/no-punct\", and \"$dir/punct/tokens.txt\") so
paths with spaces are handled safely, ensuring you quote the --exp-dir and
--tokens arguments and all ls/pushd/mv/rm uses referencing $dir.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In @.github/workflows/export-x-asr.yaml:
- Around line 1-11: The workflow "export-x-asr" lacks explicit permissions and
has unpinned actions and an insecure checkout; add a top-level permissions: {}
to deny everything by default, then grant minimal required permissions per job
(e.g., permissions: contents: read only where needed), replace action references
(e.g., actions/checkout) with commit SHAs instead of tags to pin them, and set
persist-credentials: false on the actions/checkout step to avoid leaking
GITHUB_TOKEN to subsequent steps.
- Line 280: Remove the redundant "if: true" condition from the workflow step so
it behaves like the non-streaming release step; locate the literal line "if:
true" in the export-x-asr GitHub Actions YAML and delete that key from the step
definition so the step runs by default without the unnecessary conditional.
- Around line 63-134: Replace the hardcoded dates in the directory-name
assignments (the lines that set d=... like in the "Collect model files (punct
fp32)", "Collect model files (punct int8)", "Collect model files (no-punct
fp32)", and "Collect model files (no-punct int8)" steps) with a workflow-level
variable (e.g. use env.RELEASE_DATE) and construct d using that variable instead
of literal "2026-06-03"/"2026-06-05"; also add RELEASE_DATE to the workflow env
section so all steps reference the same date variable for future releases.
- Line 137: Remove the redundant constant condition "if: true" from the workflow
step (the literal "if: true" entry) so the step runs by default; locate the step
block containing that "if: true" and delete that line entirely to satisfy
actionlint and simplify the workflow.

In `@scripts/zipformer-transducer/x-asr/export-non-streaming.sh`:
- Around line 43-84: The $dir variable expansions in the shell invocations and
the ls command (e.g., the two calls to
./zipformer/export-streaming-as-non-streaming-onnx.py and ls -lh $dir/punct) are
unquoted and can suffer word-splitting; update every occurrence of $dir
(including $dir/punct, $dir/no-punct, and any uses in --exp-dir and --tokens
flags and the ls command) to be quoted (e.g., "$dir", "$dir/punct",
"$dir/no-punct") so paths with spaces are handled safely while leaving the rest
of the arguments unchanged.

In `@scripts/zipformer-transducer/x-asr/export-streaming.sh`:
- Around line 48-112: The script uses unquoted variable expansions of $dir (e.g.
the --exp-dir and --tokens args in the export-onnx-streaming.py calls, and
commands like ls -lh $dir/punct, pushd $dir/punct, mv ... $dir/punct, rm ...
$dir/no-punct, popd) which can cause word-splitting; update every occurrence to
use quoted expansions (\"$dir\", \"$dir/punct\", \"$dir/no-punct\", and
\"$dir/punct/tokens.txt\") so paths with spaces are handled safely, ensuring you
quote the --exp-dir and --tokens arguments and all ls/pushd/mv/rm uses
referencing $dir.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: d20d041e-a459-4704-8742-e42cb53e6373

📥 Commits

Reviewing files that changed from the base of the PR and between 8799533 and de125ca.

📒 Files selected for processing (8)
  • .github/workflows/export-x-asr.yaml
  • .github/workflows/upload-models.yaml
  • scripts/zipformer-transducer/x-asr/export-non-streaming.sh
  • scripts/zipformer-transducer/x-asr/export-streaming.sh
  • scripts/zipformer-transducer/x-asr/test_onnx_non_streaming.py
  • scripts/zipformer-transducer/x-asr/test_onnx_streaming.py
  • sherpa-onnx/csrc/online-recognizer-transducer-impl.h
  • sherpa-onnx/csrc/qnn/offline-recognizer-transducer-qnn-impl.h
💤 Files with no reviewable changes (1)
  • sherpa-onnx/csrc/qnn/offline-recognizer-transducer-qnn-impl.h

@csukuangfj
csukuangfj merged commit f2991e2 into k2-fsa:master Jun 5, 2026
22 of 27 checks passed
@csukuangfj
csukuangfj deleted the export-x-asr branch June 5, 2026 14:11
jimmy1984xu pushed a commit to jimmy1984xu/sherpa-onnx that referenced this pull request Jul 29, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L This PR changes 100-499 lines, ignoring generated files.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant