Repository navigation
refactor(umi): remove redundant allocations, use byte comparison - #137
Conversation
- IdentityUmiAssigner::assign: uppercase each UMI once instead of twice, eliminate redundant unique_canonicals Vec by using HashSet directly, and use sort_unstable for deterministic ID assignment - TagSets: replace consensus_reverse()/consensus_revcomp() methods that allocated Vec<String> on every call with CONSENSUS_REVERSE/CONSENSUS_REVCOMP static slice constants - count_mismatches: use byte comparison instead of char iteration since UMI sequences are always ASCII
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #137 +/- ##
==========================================
- Coverage 83.50% 83.50% -0.01%
==========================================
Files 126 126
Lines 51375 51375
==========================================
- Hits 42901 42899 -2
- Misses 8474 8476 +2 ☔ View full report in Codecov by Sentry. 🚀 New features to boost your workflow:
|
📝 WalkthroughWalkthroughThis PR refactors the UMI assignment logic to compare bytes instead of chars and improves canonical UMI deduplication by using HashSet for explicit deduplication. Two public helper functions are replaced with public constants, reducing allocations and simplifying the API. Tests and internal usage are updated accordingly to match the new constant-based approach. 🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Summary
TagSets::consensus_reverse()andconsensus_revcomp()methods (which allocatedVec<String>on every call) withCONSENSUS_REVERSEandCONSENSUS_REVCOMPstatic slice constantsIdentityUmiAssigner::assignto uppercase each UMI once instead of twicecount_mismatchesfrom.chars().zip()to.as_bytes().iter().zip()for more efficient byte-level comparison on ASCII UMI sequencesTest plan
cargo nextest run -p fgumi-umi— all tests passcargo clippy -p fgumi-umi --all-features -- -D warnings— no warningscargo check(full workspace) — clean