Skip to content

fix(group): emit consecutive MI integers 0..N-1 - #273

Merged
nh13 merged 1 commit into
mainfrom
nh/fix-group-mi-consecutive
Apr 15, 2026
Merged

nh13 merged 1 commit into
mainfrom
nh/fix-group-mi-consecutive

Conversation

@nh13

@nh13 nh13 commented Apr 14, 2026

Copy link
Copy Markdown
Member

Summary

  • fgumi group and fgumi dedup reserved MI counter blocks sized by template count, but each position group only emits one MI per distinct UMI family (and paired A/B share a numeric id), so the block was almost always larger than needed. Unused slots became gaps in the global MI space (~22% wasted IDs on a 1 M-pair test) and propagated to downstream consensus read names, breaking cross-tool comparison with fgbio.
  • Track a distinct_mi_count on ProcessedPositionGroup / ProcessedDedupGroup computed as max(MoleculeId::id()) + 1 across assigned templates, and advance the global MI counter by that value instead of templates.len(). Fix applies to both the unified parallel pipeline and the single-threaded streaming path.
  • Adds two integration tests (single-threaded and multi-threaded pipeline paths) asserting max(MI) == distinct_count - 1, matching fgbio's GroupReadsByUmi.

Fixes #269

Test plan

  • New failing test test_group_command_assigns_consecutive_mi_values reproduces the bug: MIs {0, 4, 5, 9}, max 9 vs expected 3.
  • Same test passes after the fix.
  • New test_group_command_assigns_consecutive_mi_values_multi_threaded covers the parallel pipeline path with --threads 4 and 20 position groups.
  • cargo ci-test: 2495 passed, 19 skipped.
  • cargo ci-fmt clean.
  • cargo ci-lint clean.

Scope note

This removes the gaps in the MI integer space (the literal issue). It does not guarantee MI values appear in monotonically increasing order in the output BAM — parallel serialize workers may reserve their (now correctly-sized) blocks in any interleaved order. If byte-for-byte output-order monotonicity to match fgbio is also required, that is a follow-up (option (a) from the issue: end-of-pipeline renumber, or serial MI assignment during output write).

The `group` and `dedup` commands reserved MI counter blocks sized by
template count, but each position group only emits a distinct MI per
UMI family. Multiple templates in the same family share a `MoleculeId`,
and `PairedA(id)` / `PairedB(id)` share the same numeric id, so the
block size was almost always larger than the number of IDs actually
used. The unused slots became gaps in the global MI space (observed as
~22% wasted IDs on a 1 M-pair test), which propagated to downstream
consensus read names (e.g. `fgumi codec`'s `smoke:N`) and broke
cross-tool read-name comparison with fgbio's `GroupReadsByUmi`.

Track a `distinct_mi_count` on `ProcessedPositionGroup` and
`ProcessedDedupGroup` — computed as `max(MoleculeId::id()) + 1` across
assigned templates — and advance the global MI counter by that value
instead of `templates.len()`. This applies to both the unified parallel
pipeline and the single-threaded streaming path.

Adds two integration tests covering both paths that assert
`max(MI) == distinct_count - 1`, matching fgbio's output.

Fixes #269
@nh13
nh13 temporarily deployed to github-actions April 14, 2026 23:50 — with GitHub Actions Inactive
@nh13
nh13 marked this pull request as ready for review April 14, 2026 23:52
@coderabbitai

coderabbitai Bot commented Apr 14, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

The pull request introduces a distinct_mi_count: u64 field to track distinct numeric MoleculeId values per group, replacing template count for MI counter block allocation. This ensures emitted MI tag integers remain consecutive (0..N-1) even when templates share the same numeric MoleculeId. Changes propagate across dedup.rs, group.rs, and grouper.rs to compute and pass distinct_mi_count through both single-threaded and parallel serialization pipelines. Integration tests validate that the group command now assigns consecutive MI values across single and multi-threaded execution.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title directly matches the main change: fixing MI integer emission to be consecutive 0..N-1 instead of having gaps.
Description check ✅ Passed The description thoroughly explains the problem, solution, test coverage, and scope—clearly related to the changeset.
Linked Issues check ✅ Passed The PR fully addresses issue #269: tracks distinct_mi_count, advances MI counter accordingly, and adds tests asserting consecutive 0..N-1 integers as required.
Out of Scope Changes check ✅ Passed All changes directly support fixing MI consecutiveness: ProcessedPositionGroup/ProcessedDedupGroup field additions, MI counter logic updates, and integration tests. No extraneous modifications present.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch nh/fix-group-mi-consecutive

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@codecov

codecov Bot commented Apr 14, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.30769% with 2 lines in your changes missing coverage. Please review.
✅ Project coverage is 89.85%. Comparing base (25ede93) to head (d02426a).
⚠️ Report is 5 commits behind head on main.

Files with missing lines Patch % Lines
src/lib/commands/dedup.rs 92.30% 1 Missing ⚠️
src/lib/commands/group.rs 92.30% 1 Missing ⚠️
Additional details and impacted files
@@           Coverage Diff           @@
##             main     #273   +/-   ##
=======================================
  Coverage   89.84%   89.85%           
=======================================
  Files         120      120           
  Lines       61359    61379   +20     
=======================================
+ Hits        55127    55151   +24     
+ Misses       6232     6228    -4     

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@nh13

nh13 commented Apr 15, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 15, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Apr 15, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 15, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
src/lib/commands/group.rs (1)

1287-1297: Extract the distinct_mi_count calculation.

This max(id) + 1 rule now exists here, again in the single-threaded path below, and in src/lib/commands/dedup.rs. It’s worth centralizing so the contiguity invariant only has one definition.

♻️ Helper sketch
+fn distinct_mi_count(templates: &[Template]) -> u64 {
+    templates.iter().filter_map(|t| t.mi.id()).max().map_or(0, |max_id| max_id + 1)
+}
+
 // ...
-                let distinct_mi_count: u64 = templates
-                    .iter()
-                    .filter_map(|t| t.mi.id())
-                    .max()
-                    .map(|max_id| max_id + 1)
-                    .unwrap_or(0);
+                let distinct_mi_count = distinct_mi_count(&templates);

Apply the same helper in process_and_write_position_group and the matching dedup path.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/commands/group.rs` around lines 1287 - 1297, The calculation of
distinct_mi_count (using templates.iter().filter_map(|t|
t.mi.id()).max().map(|m| m+1).unwrap_or(0)) should be extracted into a single
helper (e.g., compute_distinct_mi_count or distinct_molecule_id_count) so the
contiguity invariant lives in one place; implement the helper to accept an
iterator or slice of templates (or a closure to get mi ids) and return u64,
replace the inline logic in this file (where distinct_mi_count is computed), in
process_and_write_position_group, and in the dedup path
(src/lib/commands/dedup.rs) to call the new helper, and add tests or a small doc
comment on the helper to codify the max+1 semantics.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@src/lib/commands/group.rs`:
- Around line 1287-1297: The calculation of distinct_mi_count (using
templates.iter().filter_map(|t| t.mi.id()).max().map(|m| m+1).unwrap_or(0))
should be extracted into a single helper (e.g., compute_distinct_mi_count or
distinct_molecule_id_count) so the contiguity invariant lives in one place;
implement the helper to accept an iterator or slice of templates (or a closure
to get mi ids) and return u64, replace the inline logic in this file (where
distinct_mi_count is computed), in process_and_write_position_group, and in the
dedup path (src/lib/commands/dedup.rs) to call the new helper, and add tests or
a small doc comment on the helper to codify the max+1 semantics.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 9403bdf0-210b-4e89-bec5-fd887fe87524

📥 Commits

Reviewing files that changed from the base of the PR and between 25ede93 and d02426a.

📒 Files selected for processing (4)
  • src/lib/commands/dedup.rs
  • src/lib/commands/group.rs
  • src/lib/grouper.rs
  • tests/integration/test_group_command.rs

@nh13
nh13 merged commit f893802 into main Apr 15, 2026
9 checks passed
@nh13
nh13 deleted the nh/fix-group-mi-consecutive branch April 15, 2026 19:26
@nh13 nh13 mentioned this pull request Apr 15, 2026
@nh13 nh13 mentioned this pull request Apr 22, 2026

This branch was previously deployed

1 inactive deployment
github-actions — d02426ac Deployed Apr 14, 2026 by nh13 via coverage #1103
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

group assigns non-consecutive MI integers; downstream consensus read names diverge from fgbio

1 participant