Skip to content

feat(group): add --parallel-group-min-templates for size-based parallel UMI assignment - #118

Closed
nh13 wants to merge 1 commit into
mainfrom
feat/parallel-group-min-templates
Closed

nh13 wants to merge 1 commit into
mainfrom
feat/parallel-group-min-templates

Conversation

@nh13

@nh13 nh13 commented Feb 19, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Add --parallel-group-min-templates N CLI option that enables the parallel UMI assigner (Edit, Adjacency, Paired) for position groups with at least N templates
  • Decouple parallel assigner from --allow-unmapped so amplicon and other workflows with large mapped position groups can opt in
  • Extract create_umi_assigner helper to eliminate duplicated assigner selection logic between pipeline and single-threaded code paths

Stacked on #39 (feat/allow-unmapped-reads).

Details

When --parallel-group-min-templates is set, groups meeting the threshold use rayon-based parallel edge discovery. --allow-unmapped continues to unconditionally use the parallel path (all unmapped reads form one giant group).

The CLI docs include a warning about thread over-subscription: the parallel assigner uses rayon's thread pool, which is separate from the pipeline's worker threads (--threads). For amplicon data (few large groups), pipeline threads are mostly idle so over-subscription is minimal.

Test plan

  • New test: test_parallel_group_min_templates_activates_for_mapped_data — verifies parallel path activates with Some(1) threshold
  • New test: test_parallel_group_min_templates_none_uses_sequential — verifies sequential path when threshold is None
  • All existing allow_unmapped tests pass unchanged
  • Full CI: cargo ci-test && cargo ci-fmt && cargo ci-lint (1840 tests passed)

@nh13
nh13 temporarily deployed to github-actions February 19, 2026 04:53 — with GitHub Actions Inactive
@codecov

codecov Bot commented Feb 19, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.70115% with 2 lines in your changes missing coverage. Please review.
✅ Project coverage is 88.99%. Comparing base (7c1cc80) to head (11c57e7).

Files with missing lines Patch % Lines
src/commands/group.rs 97.70% 2 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #118      +/-   ##
==========================================
+ Coverage   88.97%   88.99%   +0.02%     
==========================================
  Files         113      113              
  Lines       55038    55111      +73     
==========================================
+ Hits        48970    49047      +77     
+ Misses       6068     6064       -4     

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@nh13
nh13 force-pushed the feat/allow-unmapped-reads branch 3 times, most recently from 3691c0c to de39ec6 Compare February 19, 2026 16:23
@nh13
nh13 force-pushed the feat/parallel-group-min-templates branch from b00b34c to 1089920 Compare February 19, 2026 16:27
@nh13
nh13 temporarily deployed to github-actions February 19, 2026 16:27 — with GitHub Actions Inactive
@nh13
nh13 force-pushed the feat/allow-unmapped-reads branch from de39ec6 to 35d4c03 Compare February 19, 2026 16:33
@nh13
nh13 force-pushed the feat/parallel-group-min-templates branch from 1089920 to 9c38dcb Compare February 27, 2026 20:06
@nh13
nh13 force-pushed the feat/allow-unmapped-reads branch from 35d4c03 to 24bb0a2 Compare February 27, 2026 20:18
@nh13
nh13 force-pushed the feat/parallel-group-min-templates branch from 9c38dcb to 5efac96 Compare February 27, 2026 20:19
@nh13
nh13 temporarily deployed to github-actions February 27, 2026 20:19 — with GitHub Actions Inactive
@nh13
nh13 force-pushed the feat/allow-unmapped-reads branch from 24bb0a2 to da20643 Compare February 27, 2026 20:24
@nh13
nh13 force-pushed the feat/parallel-group-min-templates branch from 5efac96 to 5c87430 Compare February 27, 2026 20:26
@nh13
nh13 temporarily deployed to github-actions February 27, 2026 20:26 — with GitHub Actions Inactive
@nh13
nh13 force-pushed the feat/allow-unmapped-reads branch from da20643 to 349829c Compare February 27, 2026 20:34
@nh13
nh13 force-pushed the feat/parallel-group-min-templates branch from 5c87430 to 960f9e0 Compare February 27, 2026 20:34
@nh13
nh13 temporarily deployed to github-actions February 27, 2026 20:34 — with GitHub Actions Inactive
@nh13 nh13 changed the title feat: add --parallel-group-min-templates for size-based parallel UMI assignment feat(group): add --parallel-group-min-templates for size-based parallel UMI assignment Mar 4, 2026
@nh13
nh13 force-pushed the feat/parallel-group-min-templates branch from 960f9e0 to 9a3f6de Compare March 4, 2026 05:50
@nh13
nh13 temporarily deployed to github-actions March 4, 2026 05:50 — with GitHub Actions Inactive
@nh13
nh13 force-pushed the feat/allow-unmapped-reads branch 4 times, most recently from 30d5cfb to b20b3cd Compare March 6, 2026 05:43
Base automatically changed from feat/allow-unmapped-reads to main March 6, 2026 05:49
@nh13

nh13 commented Mar 29, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Mar 29, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai

coderabbitai Bot commented Mar 29, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

A new CLI option --parallel-group-min-templates was added to control when parallel UMI assigners activate for position groups based on template count thresholds. The UMI assigner selection logic was refactored through a new helper function create_umi_assigner(...) that chooses between sequential and parallel variants based on assignment strategy and threshold conditions, decoupling it from the allow_unmapped flag. Both the unified-pipeline and single-threaded processing paths were updated to pass this parameter and compute parallelization decisions. Tests were added to verify the threshold-based activation behavior for mapped data.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title directly and concisely describes the main feature added: a new CLI option for size-based parallel UMI assignment, which aligns with the core changeset.
Description check ✅ Passed The description clearly explains the feature, its purpose, behavior, and includes test details—all directly relevant to the changeset.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/parallel-group-min-templates

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@nh13

nh13 commented Apr 4, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 4, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

🧹 Nitpick comments (1)
src/commands/group.rs (1)

5054-5125: These tests don't actually prove the selector changed paths.

Both cases only assert final grouping, so they'd still pass if parallel_group_min_templates were ignored and mapped groups always stayed on the sequential edit assigner. A small unit test around create_umi_assigner() / the use_parallel decision, or a paired-strategy regression case that fails when the parallel branch is taken incorrectly, would lock down the new behavior directly.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/commands/group.rs` around lines 5054 - 5125, The tests only assert final
grouping and don't verify that create_umi_assigner() selects the parallel path
when GroupReadsByUmi.parallel_group_min_templates is set; add a focused unit
test that constructs the command (or directly calls create_umi_assigner()) with
parallel_group_min_templates = Some(1) and with None, then asserts the
selector/returned assigner type or a behavior unique to the parallel assigner
(e.g., type name, enum variant, or a mocked method call) to prove use_parallel
is true for the former and false for the latter; reference GroupReadsByUmi,
parallel_group_min_templates, and create_umi_assigner() to locate the code under
test.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@crates/fgumi-dna/src/bitenc.rs`:
- Around line 130-148: The public methods base_at and with_base_at currently use
debug_assert! so bounds and base-range checks are disabled in release builds;
replace the debug_assert! calls in base_at and with_base_at with unconditional
assert! (or otherwise make these checked public wrappers that call private
unchecked helpers) so pos < self.len and base < 4 are always enforced, and keep
the existing bit-masking logic (mask, bit_pos, new_bits) unchanged.

In `@src/commands/group.rs`:
- Around line 422-425: The early return when detecting ParallelPairedAssigner
bypasses paired-UMI preprocessing (the parts.len()!=2 check and the
is_r1_earlier logic) and causes same-string paired UMIs to lose their /A vs /B
distinction and malformed paired UMIs to hit asserts in
ParallelPairedAssigner::assign(); remove this shortcut and ensure the same
preprocessing runs for all assigners: parse the UMI into parts, validate
parts.len() == 2, compute/retain is_r1_earlier, normalize/canonicalize the UMI
as the rest of the code expects, then call assigner.assign() (or
assigner.as_any().downcast_ref::<ParallelPairedAssigner>().unwrap().assign()) so
ParallelPairedAssigner still gets canonicalized input and downstream asserts are
avoided while preserving correct paired-suffix behavior (see
ParallelPairedAssigner and assign()).
- Around line 1555-1576: The create_umi_assigner helper currently constructs
ParallelIdentityAssigner when use_parallel is true, which violates the intended
behavior for Strategy::Identity; change the logic so that if strategy ==
Strategy::Identity you always return the sequential identity assigner (use
strategy.new_assigner_full(effective_edits, 1, index_threshold) or the existing
sequential creation path) regardless of use_parallel, and only construct
ParallelEditAssigner / ParallelAdjacencyAssigner / ParallelPairedAssigner for
non-identity strategies when use_parallel is true; update the conditional in
create_umi_assigner to check Strategy::Identity first and return the sequential
assigner instead of ParallelIdentityAssigner::new.

In `@src/lib/umi/parallel_assigner.rs`:
- Around line 365-369: The parallel assigner currently treats
BitEnc::from_umi_str() returning None as a simple invalid UMI, which causes long
UMIs (>32 bases) to be skipped in parallel clustering; change the logic in the
parallel assigner so that if any UMI in a group fails BitEnc::from_umi_str() (or
its length exceeds BitEnc capacity) the entire group is handled by the
sequential path instead of attempting parallel clustering. Concretely: before
populating umi_counts/umi_to_original (where BitEnc::from_umi_str is called),
scan the group for UMIs that cannot be encoded (or check length >32) and, if any
exist, mark the group to use the sequential fallback (respecting the existing
--parallel-group-min-templates behavior) rather than continuing with parallel
processing; update the code paths around BitEnc::from_umi_str, umi_counts, and
umi_to_original to implement this gate or per-group sequential fallback.
- Around line 1-23: Add the module-level unsafe guard by inserting the crate
attribute to deny unsafe code at the top of the file (before any use statements
or module docs); ensure you add #![deny(unsafe_code)] as the first
non-comment/non-doc line in src/lib/umi/parallel_assigner.rs so the compiler
enforces the repo guideline for unsafe code while keeping the rest of the file
(types like Umi, UmiAssigner and functions using BitEnc, MoleculeId, rayon,
ahash, etc.) unchanged.

---

Nitpick comments:
In `@src/commands/group.rs`:
- Around line 5054-5125: The tests only assert final grouping and don't verify
that create_umi_assigner() selects the parallel path when
GroupReadsByUmi.parallel_group_min_templates is set; add a focused unit test
that constructs the command (or directly calls create_umi_assigner()) with
parallel_group_min_templates = Some(1) and with None, then asserts the
selector/returned assigner type or a behavior unique to the parallel assigner
(e.g., type name, enum variant, or a mocked method call) to prove use_parallel
is true for the former and false for the latter; reference GroupReadsByUmi,
parallel_group_min_templates, and create_umi_assigner() to locate the code under
test.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 0340362f-03b5-4fb7-855f-55ca76004138

📥 Commits

Reviewing files that changed from the base of the PR and between c98dbbf and 9a3f6de.

⛔ Files ignored due to path filters (1)
  • Cargo.lock is excluded by !**/*.lock
📒 Files selected for processing (5)
  • crates/fgumi-dna/Cargo.toml
  • crates/fgumi-dna/src/bitenc.rs
  • src/commands/group.rs
  • src/lib/umi/mod.rs
  • src/lib/umi/parallel_assigner.rs

Comment thread crates/fgumi-dna/src/bitenc.rs
Comment thread src/commands/group.rs
Comment thread src/commands/group.rs
Comment thread src/lib/umi/parallel_assigner.rs
Comment thread src/lib/umi/parallel_assigner.rs
@nh13
nh13 force-pushed the feat/parallel-group-min-templates branch from 9a3f6de to 1f33928 Compare April 5, 2026 08:42
@nh13
nh13 temporarily deployed to github-actions April 5, 2026 08:42 — with GitHub Actions Inactive
…el UMI assignment

Add --parallel-group-min-templates N CLI option that enables the parallel
UMI assigner (Edit, Adjacency, Paired) for position groups with at least
N templates. Decouples parallel assigner from --allow-unmapped so amplicon
and other workflows with large mapped position groups can opt in.

Extract create_umi_assigner helper to eliminate duplicated assigner
selection logic between pipeline and single-threaded code paths.
@nh13
nh13 force-pushed the feat/parallel-group-min-templates branch from 1f33928 to 11c57e7 Compare April 5, 2026 08:45
@nh13
nh13 temporarily deployed to github-actions April 5, 2026 08:45 — with GitHub Actions Inactive
@nh13

nh13 commented Apr 5, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 5, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Apr 5, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 5, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

♻️ Duplicate comments (1)
src/commands/group.rs (1)

1560-1574: ⚠️ Potential issue | 🔴 Critical

Keep Strategy::Paired on the sequential assigner for now.

Line 411 still returns the raw uppercase UMI for ParallelPairedAssigner, so is_r1_earlier never reaches paired canonicalization. With this helper, thresholded mapped groups can now lose their /A vs /B distinction for same-string paired UMIs.

Suggested safe fallback
-    if matches!(strategy, Strategy::Identity) {
+    if matches!(strategy, Strategy::Identity | Strategy::Paired) {
         return strategy.new_assigner_full(effective_edits, 1, index_threshold);
     }

     if use_parallel {
         match strategy {
-            Strategy::Identity => unreachable!("handled above"),
+            Strategy::Identity | Strategy::Paired => unreachable!("handled above"),
             Strategy::Edit => Box::new(ParallelEditAssigner::new(effective_edits, num_threads)),
             Strategy::Adjacency => {
                 Box::new(ParallelAdjacencyAssigner::new(effective_edits, num_threads))
             }
-            Strategy::Paired => Box::new(ParallelPairedAssigner::new(effective_edits, num_threads)),
         }
     } else {
         strategy.new_assigner_full(effective_edits, 1, index_threshold)
     }
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/commands/group.rs` around lines 1560 - 1574, The ParallelPairedAssigner
must not be used; keep paired canonicalization on the sequential assigner. In
the match inside the use_parallel branch, replace the Strategy::Paired arm
(currently Box::new(ParallelPairedAssigner::new(effective_edits, num_threads)))
with a call to the sequential assigner via
strategy.new_assigner_full(effective_edits, 1, index_threshold) so Paired falls
back to the non-parallel path; leave the other arms unchanged and keep the
earlier Identity handling as-is.
🧹 Nitpick comments (1)
src/commands/group.rs (1)

5361-5397: These tests don't prove activation.

Both use test_group_cmd(), so they stay on the fast path; Lines 1283-1292 are still uncovered, and the MI counts here are identical on the sequential and parallel edit paths. A broken branch selection would still pass.

Suggested test shape
+    #[test]
+    fn test_create_umi_assigner_selects_parallel_edit() {
+        let assigner = create_umi_assigner(Strategy::Edit, 1, 100, 4, true);
+        assert!(assigner.as_any().is::<ParallelEditAssigner>());
+    }
+
+    #[test]
+    fn test_create_umi_assigner_selects_sequential_edit() {
+        let assigner = create_umi_assigner(Strategy::Edit, 1, 100, 4, false);
+        assert!(!assigner.as_any().is::<ParallelEditAssigner>());
+    }

Also applies to: 5400-5435

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/commands/group.rs` around lines 5361 - 5397, The test is not actually
exercising the parallel branch because it still uses test_group_cmd(); fix by
constructing and invoking GroupReadsByUmi directly with
parallel_group_min_templates set (as done in the test but without using
test_group_cmd), or by adding a targeted test that calls the parallel-specific
routine (the parallel grouping path) instead of the shared helper; ensure you
create input that crosses the parallel threshold and then assert a behavior
unique to the parallel path (e.g., run both the sequential and parallel paths
and compare outputs or detect the parallel helper invocation) so the code paths
for the parallel implementation are truly covered.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Duplicate comments:
In `@src/commands/group.rs`:
- Around line 1560-1574: The ParallelPairedAssigner must not be used; keep
paired canonicalization on the sequential assigner. In the match inside the
use_parallel branch, replace the Strategy::Paired arm (currently
Box::new(ParallelPairedAssigner::new(effective_edits, num_threads))) with a call
to the sequential assigner via strategy.new_assigner_full(effective_edits, 1,
index_threshold) so Paired falls back to the non-parallel path; leave the other
arms unchanged and keep the earlier Identity handling as-is.

---

Nitpick comments:
In `@src/commands/group.rs`:
- Around line 5361-5397: The test is not actually exercising the parallel branch
because it still uses test_group_cmd(); fix by constructing and invoking
GroupReadsByUmi directly with parallel_group_min_templates set (as done in the
test but without using test_group_cmd), or by adding a targeted test that calls
the parallel-specific routine (the parallel grouping path) instead of the shared
helper; ensure you create input that crosses the parallel threshold and then
assert a behavior unique to the parallel path (e.g., run both the sequential and
parallel paths and compare outputs or detect the parallel helper invocation) so
the code paths for the parallel implementation are truly covered.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 9cf6271b-12a6-485b-ae12-67a874ebf9ab

📥 Commits

Reviewing files that changed from the base of the PR and between 9a3f6de and 11c57e7.

📒 Files selected for processing (1)
  • src/commands/group.rs

@nh13
nh13 marked this pull request as draft May 16, 2026 02:24
@nh13

nh13 commented May 26, 2026

Copy link
Copy Markdown
Member Author

Superseded by #371, which adopts the same design (decouple from --allow-unmapped, extract create_umi_assigner, over-subscription warning) but replaces the single user-supplied integer threshold with . The new auto mode uses per-strategy thresholds (Identity: never, Edit: 1536, Adjacency: 3072, Paired: 128) justified by a new microbench (benches/umi_assigner_threshold.rs). Bench data showed a flat threshold was the wrong shape for three of four strategies.

@nh13 nh13 closed this May 26, 2026
@nh13

nh13 commented May 26, 2026

Copy link
Copy Markdown
Member Author

(Correction to my previous comment — shell quoting ate the flag name.)

Superseded by #371, which keeps the same design from this PR (decouple from --allow-unmapped, extract create_umi_assigner, over-subscription warning) but takes the flag value as N|auto instead of a bare integer.

The new auto mode uses per-strategy thresholds (Identity: usize::MAX, Edit: 1536, Adjacency: 3072, Paired: 128) justified by a new microbench (benches/umi_assigner_threshold.rs). Bench data showed a flat threshold was the wrong shape for three of four strategies — Paired wins parallel at ~128 templates while Edit/Adjacency need 1.5k–3k and Identity never.

This branch was previously deployed

1 inactive deployment
github-actions — 11c57e71 Deployed Apr 5, 2026 by nh13 via coverage #906
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant