Export nemotron-speech-streaming-en-0.6b to QNN - #3725
Conversation
📝 WalkthroughWalkthroughThis PR adds an end-to-end pipeline for exporting the nemotron-speech-streaming-en-0.6b NeMo model to ONNX and quantizing it to QNN. It includes a patched ONNX export wrapper, a streaming ONNX inference test script, a shell launcher, a build matrix generator, and a new GitHub Actions workflow. ChangesNemotron speech streaming QNN export
Estimated code review effort: 4 (Complex) | ~60 minutes Sequence Diagram(s)sequenceDiagram
participant Push
participant OnnxJob
participant GenerateMatrixJob
participant QnnJob
participant Release
Push->>OnnxJob: trigger workflow, run.sh exports ONNX
OnnxJob->>OnnxJob: test_onnx.py validates encoder/decoder/joiner
OnnxJob->>Release: upload per-chunk-size ONNX artifact
Push->>GenerateMatrixJob: run generate_nemotron_speech_streaming.py
GenerateMatrixJob->>QnnJob: expose SoC/chunk-size build matrix
QnnJob->>OnnxJob: download matching ONNX artifact
QnnJob->>QnnJob: convert/quantize via qnn-onnx-converter, build context binaries
QnnJob->>Release: package tarballs, publish release artifacts
Related PRs: None identified. Suggested labels: ci, export, qnn, nemo Suggested reviewers: csukuangfj 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@scripts/nemo/qnn/nemotron-speech-streaming-en-0.6b/run.sh`:
- Line 29: Quote the shell variables in the wrapper.py invocation to prevent
word splitting and globbing; update the run.sh command that passes chunk_size_ms
and model_id so both arguments are wrapped in quotes when used in the python3
./wrapper.py call.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro
Run ID: cb410d9b-6b5e-42af-b360-6178033bef2e
📒 Files selected for processing (5)
.github/scripts/export-qnn/generate_nemotron_speech_streaming.py.github/workflows/export-nemotron-speech-streaming-en-0.6b-qnn.yamlscripts/nemo/qnn/nemotron-speech-streaming-en-0.6b/run.shscripts/nemo/qnn/nemotron-speech-streaming-en-0.6b/test_onnx.pyscripts/nemo/qnn/nemotron-speech-streaming-en-0.6b/wrapper.py
| model_id="$2" | ||
| fi | ||
|
|
||
| python3 ./wrapper.py --chunk-size-ms $chunk_size_ms --model-id $model_id |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Quote $chunk_size_ms and $model_id to avoid globbing/word-splitting.
🔧 Proposed fix
-python3 ./wrapper.py --chunk-size-ms $chunk_size_ms --model-id $model_id
+python3 ./wrapper.py --chunk-size-ms "$chunk_size_ms" --model-id "$model_id"📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| python3 ./wrapper.py --chunk-size-ms $chunk_size_ms --model-id $model_id | |
| python3 ./wrapper.py --chunk-size-ms "$chunk_size_ms" --model-id "$model_id" |
🧰 Tools
🪛 Shellcheck (0.11.0)
[info] 29-29: Double quote to prevent globbing and word splitting.
(SC2086)
[info] 29-29: Double quote to prevent globbing and word splitting.
(SC2086)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@scripts/nemo/qnn/nemotron-speech-streaming-en-0.6b/run.sh` at line 29, Quote
the shell variables in the wrapper.py invocation to prevent word splitting and
globbing; update the run.sh command that passes chunk_size_ms and model_id so
both arguments are wrapped in quotes when used in the python3 ./wrapper.py call.
Source: Linters/SAST tools
There was a problem hiding this comment.
Code Review
This pull request introduces scripts and wrappers to export and test the nemotron-speech-streaming-en-0.6b model for Qualcomm QNN. Key changes include a configuration generator for GitHub Actions, a test script using ONNX Runtime, and a wrapper that monkey-patches NeMo's ConformerEncoder to ensure QNN compatibility. The review feedback highlights several critical issues: a potential ModuleNotFoundError due to missing import paths, a crash when chunk_size_ms is 80, potential parsing failures with whitespace tokens, and inefficient array upcasting to float64. Additionally, minor shell scripting improvements are suggested to prevent globbing and word-splitting issues.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
|
|
||
| import json | ||
|
|
||
| from device_info import soc_info_dict |
There was a problem hiding this comment.
The script imports device_info directly, but device_info.py is located in scripts/qnn/device_info.py, which is not in the Python search path when running this script. This will cause a ModuleNotFoundError when executed. Please add the scripts/qnn directory to sys.path before importing.
| from device_info import soc_info_dict | |
| import sys | |
| from pathlib import Path | |
| # Add scripts/qnn to sys.path so we can import device_info | |
| sys.path.append(str(Path(__file__).resolve().parents[3] / "scripts" / "qnn")) | |
| from device_info import soc_info_dict |
| for line in f: | ||
| t, idx = line.split() | ||
| id2token[int(idx)] = t |
There was a problem hiding this comment.
Using line.split() to parse tokens.txt will fail with a ValueError if any token is a space or empty string (which is common in SentencePiece/BPE vocabularies), because split() ignores leading/trailing whitespace and splits on any whitespace. Using rsplit(" ", 1) is much more robust.
| for line in f: | |
| t, idx = line.split() | |
| id2token[int(idx)] = t | |
| for line in f: | |
| parts = line.rstrip("\r\n").rsplit(" ", 1) | |
| if len(parts) == 2: | |
| t, idx = parts | |
| id2token[int(idx)] = t |
|
|
||
| window_size = chunk_size + pre_encode_cache_size | ||
|
|
||
| window_shift = chunk_size |
There was a problem hiding this comment.
| ) | ||
| sample_rate = 16000 | ||
|
|
||
| tail_padding = np.zeros(sample_rate * 1) |
There was a problem hiding this comment.
Creating tail_padding without specifying dtype defaults to float64. When concatenated with audio (which is float32), it upcasts the entire array to float64. This is inefficient and can cause type mismatch issues with the ONNX model which expects float32 inputs.
| tail_padding = np.zeros(sample_rate * 1) | |
| tail_padding = np.zeros(sample_rate * 1, dtype=np.float32) |
| set -ex | ||
|
|
||
| pip install \ | ||
| nemo_toolkit['asr'] \ |
| model_id="$2" | ||
| fi | ||
|
|
||
| python3 ./wrapper.py --chunk-size-ms $chunk_size_ms --model-id $model_id |
There was a problem hiding this comment.
C++ runtime will be added in a separate pull request.
You can find the exported models at
(search for streaming in the above two pages)
Summary by CodeRabbit
New Features
Bug Fixes