Skip to content

fix(group): reject UMIs of differing length in all assigners (GRP-01) - #510

Merged
nh13 merged 1 commit into
mainfrom
nh/fix-group-unequal-umi-length
Jul 14, 2026
Merged

nh13 merged 1 commit into
mainfrom
nh/fix-group-unequal-umi-length

Conversation

@nh13

@nh13 nh13 commented Jul 9, 2026 •

Copy link
Copy Markdown
Member

What

fgbio's GroupReadsByUmi throws on UMIs of differing length: its adjacency assigner has an explicit require(orderedNodes.forall(_.umi.length == umiLength), "Multiple UMI lengths"), and Sequences.countMismatches requires equal lengths (GroupReadsByUmiTest.scala:513, "fail when umis have different length").

fgumi's count_mismatches instead returns usize::MAX for differing lengths, which every mismatch-based matcher reads as "no match", so a differing-length UMI silently forms its own molecule — under-grouping with no error. Only the sequential adjacency assigner carried the guard (an inline assert!); the sequential paired/edit assigners and all three parallel assigners (edit/adjacency/paired) did not, so the default threaded path under-grouped silently for every strategy.

Change

  • Add a shared assert_uniform_umi_length helper (plus base_umi_length) in fgumi-umi::assigner, and wire it into every mismatch-based assigner path — sequential edit/adjacency/paired and parallel edit/adjacency/paired — replacing the ad-hoc inline adjacency check.
  • Identity grouping is exact-match and needs no length guard, matching fgbio's identity assigner.
  • --min-umi-length still works: it truncates every UMI to a common length before assignment, so the guard never fires on the variable-length input the flag supports. A new truncate → assign test (test_variable_length_umis_group_after_truncation) locks that interaction end-to-end.

Parity evidence (§0 step 2/4)

Fixture: reports/fgbio-parity-fixtures/grp-01-unequal-umi-length.sh — two read pairs at the same coordinates/orientation, RX = ACT-ACT (7 chars) vs ACT-AC (6 chars), run under Strategy.Paired/Adjacency/Edit, edits=1 (mirrors GroupReadsByUmiTest.scala:513).

before:
  fgbio  Paired    exit: 1   (rejects)      fgumi  Paired    exit: 0    (silent under-group)
  fgbio  Adjacency exit: 1   (rejects)      fgumi  Adjacency exit: 101  (already guarded)
  fgbio  Edit      exit: 1   (rejects)      fgumi  Edit      exit: 0    (silent under-group)

after:
  fgbio  Paired    exit: 1   (rejects)      fgumi  Paired    exit: 101  (rejects)
  fgbio  Adjacency exit: 1   (rejects)      fgumi  Adjacency exit: 101  (rejects)
  fgbio  Edit      exit: 1   (rejects)      fgumi  Edit      exit: 101  (rejects)

Both tools now reject differing-length UMIs (non-zero exit) across all three strategies; before, Paired and Edit silently under-grouped.

Notes

  • The guard panics (assert!), consistent with the existing sequential-adjacency guard and the sibling "is not a paired UMI" guards in the same file. Exit is non-zero (101) rather than fgbio's 1, which satisfies the missing-guard (S3/BS6) parity standard used throughout the tracker; converting the assign trait to return Result would be an invasive refactor rippling into dedup and is out of scope for this parity fix.

Tests

  • New should_panic cases for all six assigner paths (sequential + parallel × edit/adjacency/paired) asserting "Multiple UMI lengths", plus direct unit tests for assert_uniform_umi_length / base_umi_length, and the --min-umi-length regression test. cargo ci-fmt && cargo ci-lint && cargo ci-test all green (2219 passed).

Tracker: GRP-01 (S3/S1). GRP-02 wontfix.

Summary by CodeRabbit

  • New Features
    • Added assert_uniform_umi_length(...) to enforce consistent UMI lengths, including consistent paired behavior (delimiter ignored).
  • Bug Fixes
    • Edit/Adjacency/Paired assigners now fail fast on mixed-length valid UMIs; differences from invalid/non-encodable UMIs no longer trigger length panics.
    • Invalid/non-encodable UMIs now share or diverge molecule assignments by distinct invalid-string identity, aligning sequential and parallel results.
    • Adjacency profiling reports base UMI length directly.
  • Tests
    • Added truncation regression coverage and expanded parameterized panic/non-panic cases across strategies.

@nh13
nh13 temporarily deployed to github-actions July 9, 2026 05:26 — with GitHub Actions Inactive
@coderabbitai

coderabbitai Bot commented Jul 9, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

Walkthrough

Encodable UMIs with mixed base lengths now fail before clustering; invalid UMIs are excluded from length checks and receive stable per-string molecule assignments. Sequential and parallel paths share this behavior, with regression coverage for paired UMIs and truncation.

Changes

UMI assignment consistency

Layer / File(s) Summary
Sequential assignment behavior
crates/fgumi-umi/src/assigner.rs
Adds the exported length guard and invalid-UMI fallback, applies both across sequential Edit, Adjacency, and Paired assigners, and tests valid versus invalid length handling.
Parallel assignment behavior
src/lib/umi/parallel_assigner.rs
Validates encoded lengths before graph discovery and gives repeated invalid strings shared molecule IDs, including canonical handling for paired UMIs.
Cross-assigner regression coverage
src/lib/umi/parallel_assigner.rs, src/lib/commands/group.rs
Adds parameterized parity and panic tests, invalid-UMI isolation checks, and adjacency grouping coverage after fixed-length truncation.

Estimated code review effort: 4 (Complex) | ~45 minutes

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: rejecting differing-length UMIs across all assigners.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch nh/fix-group-unequal-umi-length

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Jul 9, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 85.41667% with 21 lines in your changes missing coverage. Please review.
✅ Project coverage is 92.59%. Comparing base (f638abc) to head (d03a568).
⚠️ Report is 28 commits behind head on main.

Files with missing lines Patch % Lines
src/lib/umi/parallel_assigner.rs 75.36% 17 Missing ⚠️
crates/fgumi-umi/src/assigner.rs 92.85% 4 Missing ⚠️

❌ Your patch check has failed because the patch coverage (85.41%) is below the target coverage (90.00%). You can increase the patch coverage or adjust the target coverage.

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #510      +/-   ##
==========================================
+ Coverage   91.12%   92.59%   +1.47%     
==========================================
  Files          78      166      +88     
  Lines       51540   100040   +48500     
==========================================
+ Hits        46964    92629   +45665     
- Misses       4576     7411    +2835     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@nh13

nh13 commented Jul 11, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 11, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 11, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 11, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 11, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 11, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 11, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 11, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/fgumi-umi/src/assigner.rs`:
- Around line 989-991: The sequential and parallel assignment paths validate
different UMI sets before assert_uniform_umi_length. Update the logic around
assert_uniform_umi_length at crates/fgumi-umi/src/assigner.rs:989-991 and
1884-1888 so invalid, non-BitEnc-encodable UMIs are filtered before length
validation in both paths, or consistently rejected before filtering; ensure
mixed valid/invalid inputs produce identical behavior across sequential,
parallel edit, and paired paths.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: bc7769ae-4c15-4072-b2b1-d3f27e7a4a69

📥 Commits

Reviewing files that changed from the base of the PR and between f638abc and 3745dc7.

📒 Files selected for processing (3)
  • crates/fgumi-umi/src/assigner.rs
  • src/lib/commands/group.rs
  • src/lib/umi/parallel_assigner.rs

Comment thread crates/fgumi-umi/src/assigner.rs Outdated
@nh13
nh13 force-pushed the nh/fix-group-unequal-umi-length branch from 3745dc7 to 718460d Compare July 11, 2026 21:17
@nh13
nh13 temporarily deployed to github-actions July 11, 2026 21:17 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 12, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 12, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 12, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 12, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/fgumi-umi/src/assigner.rs`:
- Around line 651-688: Make the production-facing base_umi_length function
private, or move its logic into the tests that need it; do not expose it
publicly when there are no production call sites. Keep assert_uniform_umi_length
public and unchanged.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 03412e25-a8db-4ca8-ac7f-395cd31fad96

📥 Commits

Reviewing files that changed from the base of the PR and between 3745dc7 and 718460d.

📒 Files selected for processing (3)
  • crates/fgumi-umi/src/assigner.rs
  • src/lib/commands/group.rs
  • src/lib/umi/parallel_assigner.rs

Comment thread crates/fgumi-umi/src/assigner.rs Outdated
@nh13
nh13 force-pushed the nh/fix-group-unequal-umi-length branch from 718460d to 032beae Compare July 12, 2026 15:11
@nh13
nh13 temporarily deployed to github-actions July 12, 2026 15:11 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 12, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 12, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 12, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 12, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 12, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 12, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai

coderabbitai Bot commented Jul 13, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/fgumi-umi/src/assigner.rs`:
- Around line 2105-2155: Extend
test_sequential_assigner_excludes_invalid_umis_from_length_guard with an
fgbio-derived expected baseline for the mixed valid/invalid Edit and Paired
inputs, and assert fgumi’s panic/grouping behavior matches it. If fgbio
intentionally differs for these cases, document that divergence explicitly in
the test instead of leaving the expected behavior unsupported.

In `@src/lib/umi/parallel_assigner.rs`:
- Around line 1068-1120: Extend the regression tests around
test_parallel_assigner_rejects_differing_umi_lengths and
test_parallel_assigner_excludes_invalid_umis_from_length_guard to
programmatically obtain the corresponding fgbio baseline results, then assert
each parallel assigner matches that baseline. Preserve the existing
differing-length panic expectation and mixed valid/invalid grouping assertions
while adding coverage for both the sequential/local behavior and parallel
parity.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: a7d99c93-839d-48e6-b075-386855874979

📥 Commits

Reviewing files that changed from the base of the PR and between 032beae and 6396833.

📒 Files selected for processing (3)
  • crates/fgumi-umi/src/assigner.rs
  • src/lib/commands/group.rs
  • src/lib/umi/parallel_assigner.rs

Comment thread crates/fgumi-umi/src/assigner.rs
Comment thread src/lib/umi/parallel_assigner.rs
@nh13
nh13 force-pushed the nh/fix-group-unequal-umi-length branch from 6396833 to bd1f81c Compare July 13, 2026 03:21
@nh13
nh13 temporarily deployed to github-actions July 13, 2026 03:21 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 13, 2026

Copy link
Copy Markdown
Member Author

Addressed the two length-guard baseline comments (assigner.rs sequential + parallel_assigner.rs excludes_invalid_umis tests) by documenting the intentional fgbio divergence rather than invoking fgbio from crate unit tests (a JVM tool with no presence in these crates' unit-test CI; fgbio parity in this repo lives in e2e baselines).

Traced to fgbio source: for the mixed valid/invalid input, fgbio never drops non-encodable UMIs, so SimpleErrorUmiAssigner (Edit) throws via Sequences.countMismatches's require(s1.length == s2.length) (Sequences.scala:99) and PairedUmiAssigner extends AdjacencyUmiAssigner throws via the umiLength require (GroupReadsByUmi.scala:283) — both seeing the length-3 NNN / length-5 NN-NN beside the length-4 AAAA / length-7 ACT-ACT. fgumi deliberately tolerates invalid-base UMIs of a differing length by excluding them from the guard (GRP-01). Differing-length valid UMIs still reject in parity with fgbio (the rejects_differing_umi_lengths tests). Both test doc comments now state this explicitly.

@nh13

nh13 commented Jul 13, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 13, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/fgumi-umi/src/assigner.rs`:
- Around line 980-988: Update the sequential edit-assignment flow around
assert_uniform_umi_length to skip non-encodable UMIs during matching, preventing
duplicate invalid values from being deduplicated into one molecule. Allocate
each invalid input its own molecule ID while constructing results, matching
parallel assignment output; add a parity fixture covering duplicate invalid UMIs
such as ["AAAA", "NNN", "NNN"].
- Around line 2132-2165: Extend
test_sequential_assigner_excludes_invalid_umis_from_length_guard with an
Adjacency case using mixed valid and invalid UMIs, then update the sequential
adjacency fast path so invalid UMIs receive separate assignments rather than
sharing the sole encoded UMI’s ID. Add coverage for empty and single-read
families, and verify sequential and parallel adjacency outputs are identical
while preserving valid-UMI grouping.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 78bd6a26-1ced-424e-b74e-fa6df1e3d9b0

📥 Commits

Reviewing files that changed from the base of the PR and between 6396833 and bd1f81c.

📒 Files selected for processing (3)
  • crates/fgumi-umi/src/assigner.rs
  • src/lib/commands/group.rs
  • src/lib/umi/parallel_assigner.rs

Comment thread crates/fgumi-umi/src/assigner.rs
Comment thread crates/fgumi-umi/src/assigner.rs
@nh13
nh13 force-pushed the nh/fix-group-unequal-umi-length branch from bd1f81c to d03a568 Compare July 13, 2026 18:45
@nh13
nh13 temporarily deployed to github-actions July 13, 2026 18:45 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 13, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 13, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
crates/fgumi-umi/src/assigner.rs (1)

1938-1968: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Exclude non-BitEnc paired UMIs before counting/grouping. Invalid canonical strings can still hit the umi_counts.len() == 1 fast path and get PairedA/B, and same-length invalid neighbors can be pulled into the graph and join a valid molecule. Filter umi_counts to encodable entries up front, then route the dropped UMIs through the Single fallback.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/fgumi-umi/src/assigner.rs` around lines 1938 - 1968, The paired
assignment flow must exclude non-BitEnc UMIs before fast-path and adjacency
grouping. Filter `umi_counts` to BitEnc-encodable entries up front, use the
filtered counts for `umi_counts.len()` and `build_adjacency_graph`, and route
dropped invalid UMIs through the `Single` fallback while preserving valid paired
assignment behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@crates/fgumi-umi/src/assigner.rs`:
- Around line 1938-1968: The paired assignment flow must exclude non-BitEnc UMIs
before fast-path and adjacency grouping. Filter `umi_counts` to BitEnc-encodable
entries up front, use the filtered counts for `umi_counts.len()` and
`build_adjacency_graph`, and route dropped invalid UMIs through the `Single`
fallback while preserving valid paired assignment behavior.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 328ccc45-ee8a-4532-96b0-7834f77c5d34

📥 Commits

Reviewing files that changed from the base of the PR and between 6396833 and d03a568.

📒 Files selected for processing (3)
  • crates/fgumi-umi/src/assigner.rs
  • src/lib/commands/group.rs
  • src/lib/umi/parallel_assigner.rs

fgbio's GroupReadsByUmi throws on UMIs of differing length: its adjacency
assigner has an explicit require(orderedNodes.forall(_.umi.length == umiLength),
"Multiple UMI lengths"), and Sequences.countMismatches requires equal lengths
(GroupReadsByUmiTest.scala:513, "fail when umis have different length").

fgumi's count_mismatches instead returns usize::MAX for differing lengths, which
every mismatch-based matcher reads as "no match", so a differing-length UMI
silently forms its own molecule -- under-grouping with no error. Only the
sequential adjacency assigner carried the guard (an inline assert!); the
sequential paired/edit assigners and all three parallel assigners
(edit/adjacency/paired) did not, so the default threaded path under-grouped
silently for every strategy.

Add a shared assert_uniform_umi_length helper in fgumi-umi::assigner and wire it
into every mismatch-based assigner path, replacing the ad-hoc adjacency check.
Identity grouping is exact-match and needs no guard, matching fgbio's identity
assigner. --min-umi-length still works: it truncates every UMI to a common
length before assignment (covered by a new truncate->assign test).

Also make the sequential and parallel assigners agree on invalid
(non-BitEnc-encodable) UMIs, which are production-reachable: validate_umi (like
fgbio, GroupReadsByUmi.scala:673) filters only uppercase-N UMIs upstream, so a
UMI with a lowercase 'n' or a non-ACGT IUPAC base passes the filter and reaches
assign() as non-encodable. Previously the six assigner paths handled these
inconsistently: sequential edit string-matched them (grouping N-neighbours),
parallel edit/adjacency/paired split duplicate invalid UMIs into distinct
molecules, and the sequential adjacency single-UMI fast path tagged an invalid
UMI into the real molecule. Depending on --threads, identical reads grouped
differently.

Unify all six paths on fgbio's map-by-string model: every invalid UMI gets its
own molecule keyed by its uppercased string (identical strings share, distinct
strings never merge, an invalid UMI never joins a valid molecule), via a shared
assign_with_invalid_fallback helper. The one residual divergence from fgbio is
structural: fgbio treats N as a mismatch character and can merge a
within-threshold N-neighbour, but BitEnc cannot represent N, so fgumi isolates
such a UMI; both fgumi assigners agree on the isolation.

Add test_sequential_and_parallel_assigners_induce_same_partition, a
cross-assigner parity harness covering valid grouping plus the invalid-UMI edge
cases per strategy, and run the invalid-UMI behavioural tests on both the
sequential and parallel assigners.

The paired sequential path needed one extra step to actually match that model: the
group command tags each half of a paired UMI with an orientation prefix
(lower_read_umi_prefix/higher_read_umi_prefix, e.g. "aa:ACGT-bb:TTTT") before calling
assign(), so a whole-key BitEnc check reads every production key as non-encodable -- it
would isolate every molecule as a filter and never fire as the length guard. The paired
assigner now strips that prefix per half (underlying_umi_len) before deciding
encodability, so genuinely non-encodable paired UMIs (e.g. a lowercase-n / IUPAC base
that bypassed the uppercase-N filter) are isolated to their own molecule -- matching the
parallel paired assigner and grouping the same reads identically regardless of --threads
-- while the differing-length guard now fires on the prefixed production keys instead of
silently no-op'ing.

Parity (Strategy Paired/Adjacency/Edit, edits=1, RX "ACT-ACT" vs "ACT-AC"):
  before: fgbio rejects (exit 1); fgumi Paired/Edit exit 0 (silent under-group),
          Adjacency exit 101
  after:  fgbio rejects (exit 1); fgumi rejects all three (exit 101)

Tracker: GRP-01 (S3/S1).
@nh13
nh13 force-pushed the nh/fix-group-unequal-umi-length branch from d03a568 to 8b6838b Compare July 13, 2026 22:00
@nh13
nh13 temporarily deployed to github-actions July 13, 2026 22:00 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 13, 2026

Copy link
Copy Markdown
Member Author

Addressed the outside-diff comment on the paired assigner (crates/fgumi-umi/src/assigner.rs, ~1938-1968) in the pushed amend.

The literal suggestion — filter umi_counts to BitEnc-encodable entries — can't be applied as-worded: the sequential PairedUmiAssigner is fed orientation-prefixed keys by the group command (umi_for_read_impl emits aa:ACGT-bb:TTTT), which are never BitEnc-encodable as a whole, so a whole-key filter would isolate every molecule. For the same reason the existing differing-length guard was a silent no-op on those production keys.

Fixed it prefix-aware instead: a new underlying_umi_len strips the aa:/bb: prefix per half before checking ACGT-encodability. Non-encodable paired UMIs (e.g. a lowercase-n/IUPAC base that bypasses the upstream uppercase-N filter) now get their own Single molecule via assign_with_invalid_fallback across all three paired code paths (all-invalid, len()==1 fast path, adjacency graph) — matching the parallel paired assigner and the edit/adjacency policy, so grouping no longer depends on --threads. The length guard now also fires on the prefixed production keys.

Coverage: test_paired_prefixed_invalid_umi_is_isolated (prefixed production-key form), plus two clean-input cases added to the cross-assigner parity harness (paired_invalid_neighbor_fast_path, paired_invalid_neighbor_graph_path) that are RED before the fix. test_paired_backward_capture's XXXX filler became TTTT since non-ACGT filler is now correctly isolated rather than grouped. Full suite green (2234 tests).

@nh13

nh13 commented Jul 13, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 13, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
src/lib/umi/parallel_assigner.rs (2)

686-711: 🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift

Parse prefixed paired keys before BitEnc encoding.

Production keys such as aa:ACGT-bb:TTTT contain : and a B prefix, so whole-string encoding filters valid UMIs out. They then hit the all-invalid branch, lose /A//B, and bypass both adjacency and length validation. Share the sequential prefix-aware parsing and add a parallel production-key regression.

As per path instructions, parallel paired assignment must match the sequential path and fgbio baseline.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/lib/umi/parallel_assigner.rs` around lines 686 - 711, Update the parallel
paired-assignment preprocessing around sorted_umis to parse production-prefixed
paired keys with the same prefix-aware logic used by the sequential assigner
before calling BitEnc::from_umi_str. Preserve the original key for counts and
output while encoding the parsed UMI components, so /A and /B handling,
adjacency assignment, and uniform-length validation remain active; add a
regression covering prefixed production keys.

Source: Path instructions


770-792: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Key invalid paired fallbacks by the raw uppercase string.

Canonical keys merge distinct reversed invalid strings only when a valid UMI is also present. For example, ACGN-TTTT and TTTT-ACGN share an ID here, but remain distinct in the sequential and all-invalid paths.

Proposed fix
                 } else {
-                    *invalid_to_id.entry(canonical).or_insert_with(|| {
+                    *invalid_to_id.entry(umi.to_uppercase()).or_insert_with(|| {
                         let id = MoleculeId::Single(next_mol_id);
                         next_mol_id += 1;
                         id

Add this reversed-invalid case to the parity harness.

As per path instructions, parallel assignment must produce output identical to sequential assigners.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/lib/umi/parallel_assigner.rs` around lines 770 - 792, Update the invalid
fallback in the raw_umis mapping to key invalid UMIs by their raw uppercase
string rather than the canonicalized value, while retaining canonical_to_mol
lookup and strand assignment for valid UMIs. Ensure reversed invalid UMIs
receive distinct MoleculeId::Single values and match sequential/all-invalid
behavior. Extend the parity harness with the reversed-invalid case such as
ACGN-TTTT and TTTT-ACGN.

Source: Path instructions

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Outside diff comments:
In `@src/lib/umi/parallel_assigner.rs`:
- Around line 686-711: Update the parallel paired-assignment preprocessing
around sorted_umis to parse production-prefixed paired keys with the same
prefix-aware logic used by the sequential assigner before calling
BitEnc::from_umi_str. Preserve the original key for counts and output while
encoding the parsed UMI components, so /A and /B handling, adjacency assignment,
and uniform-length validation remain active; add a regression covering prefixed
production keys.
- Around line 770-792: Update the invalid fallback in the raw_umis mapping to
key invalid UMIs by their raw uppercase string rather than the canonicalized
value, while retaining canonical_to_mol lookup and strand assignment for valid
UMIs. Ensure reversed invalid UMIs receive distinct MoleculeId::Single values
and match sequential/all-invalid behavior. Extend the parity harness with the
reversed-invalid case such as ACGN-TTTT and TTTT-ACGN.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: dd5b6bd5-139b-4240-a533-e162963ad85f

📥 Commits

Reviewing files that changed from the base of the PR and between d03a568 and 8b6838b.

📒 Files selected for processing (3)
  • crates/fgumi-umi/src/assigner.rs
  • src/lib/commands/group.rs
  • src/lib/umi/parallel_assigner.rs

@nh13

nh13 commented Jul 14, 2026

Copy link
Copy Markdown
Member Author

Re the two outside-diff findings on src/lib/umi/parallel_assigner.rs in the latest review:

1. "Parse prefixed paired keys before BitEnc encoding" (686-711) — not applicable (false positive). The premise is that the parallel paired assigner receives prefixed production keys like aa:ACGT-bb:TTTT. It does not: the orientation prefix is added only for the sequential PairedUmiAssigner. umi_for_read_impl (src/lib/commands/group.rs:313-317) returns the raw uppercase UMI without a prefix when the assigner is a ParallelPairedAssigner, because that assigner canonicalizes internally. So ParallelPairedAssigner always sees clean ACGT-TTTT keys, its whole-string BitEnc::from_umi_str encodes them correctly, and no valid UMI is filtered into the all-invalid branch. The prefix-aware parsing that underlying_umi_len does is needed on the sequential side precisely because that's the only side that gets prefixed keys. No change.

2. "Key invalid paired fallbacks by the raw uppercase string" (770-792) — real, fixed in the stacked #525. Confirmed: the parallel main-path fallback keyed invalid UMIs by their canonical form, so reverses like ACGN-TTTT / TTTT-ACGN merged into one Single when a valid UMI was present, while the sequential assigner and the all_invalid_molecule_ids branch key by raw uppercase and keep them distinct. Reproduced:

umis = ["ACGT-ACGT", "ACGN-TTTT", "TTTT-ACGN"]
sequential: [PairedB(0), Single(1), Single(2)]
parallel:   [PairedB(0), Single(1), Single(1)]   # reverses merged

The fix (entry(canonical) → entry(umi.to_uppercase())) lands in #525, which reworks this exact ParallelPairedAssigner::assign final-mapping block (a fix in #510 would just be overwritten by #525's rework). Added a paired_reversed_invalid case to the cross-assigner parity harness covering it. #525 is rebased on this branch; full suite green (2237 tests).

@nh13
nh13 merged commit 808930d into main Jul 14, 2026
8 checks passed
@nh13
nh13 deleted the nh/fix-group-unequal-umi-length branch July 14, 2026 02:05
@nh13 nh13 mentioned this pull request Jul 14, 2026
nh13 added a commit that referenced this pull request Jul 21, 2026
The sequential edit, adjacency, and paired assigners each re-ran
BitEnc::from_umi_str several times per UMI and discarded the result,
keeping only .len() or .is_some(). On the paired path this compounded:
the differing-length guard added in #510 called underlying_umi_len (which
encodes both halves of the prefixed key) up to three times per UMI -- in
the encodability filter, the length guard, and the per-record strand
resolution in the single-molecule fast path.

Encode each UMI once into a Vec<Option<BitEnc>> (paired: a Vec of
underlying base lengths) and reuse it for all three sites. BitEnc is a
16-byte Copy struct, so caching it is far cheaper than repeating the
per-byte encode. count_paired now carries the underlying length; this is
sound because canonicalization only swaps the two halves (A-B <-> B-A),
which preserves both BitEnc-encodability and the summed base count, so
every raw UMI folding into a canonical form contributes the same value.

Grouping output is unchanged: family-size histograms and molecule counts
are byte-identical before and after. On a c7g.4xlarge at threads=0 this
recovers ~1% of paired group runtime on agilent-hs2 (93.4s -> 92.4s user),
with the untouched identity strategy flat as a control.

Encode calls per UMI: edit 2->1, adjacency 2->1, paired 6->2.
nh13 added a commit that referenced this pull request Jul 23, 2026
The sequential edit, adjacency, and paired assigners each re-ran
BitEnc::from_umi_str several times per UMI and discarded the result,
keeping only .len() or .is_some(). On the paired path this compounded:
the differing-length guard added in #510 called underlying_umi_len (which
encodes both halves of the prefixed key) up to three times per UMI -- in
the encodability filter, the length guard, and the per-record strand
resolution in the single-molecule fast path.

Encode each UMI once into a Vec<Option<BitEnc>> (paired: a Vec of
underlying base lengths) and reuse it for all three sites. BitEnc is a
16-byte Copy struct, so caching it is far cheaper than repeating the
per-byte encode. count_paired now carries the underlying length; this is
sound because canonicalization only swaps the two halves (A-B <-> B-A),
which preserves both BitEnc-encodability and the summed base count, so
every raw UMI folding into a canonical form contributes the same value.

Grouping output is unchanged: family-size histograms and molecule counts
are byte-identical before and after. On a c7g.4xlarge at threads=0 this
recovers ~1% of paired group runtime on agilent-hs2 (93.4s -> 92.4s user),
with the untouched identity strategy flat as a control.

Encode calls per UMI: edit 2->1, adjacency 2->1, paired 6->2.

This branch was previously deployed

1 inactive deployment
github-actions — 8b6838b8 Deployed Jul 13, 2026 by nh13 via coverage #2535
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant