Skip to content

fix(umi): make sequential and parallel assigners agree on UMI case - #455

Merged
nh13 merged 1 commit into
mainfrom
nh/fix-assigner-case-parity
Jun 24, 2026
Merged

nh13 merged 1 commit into
mainfrom
nh/fix-assigner-case-parity

Conversation

@nh13

@nh13 nh13 commented Jun 23, 2026 •

Copy link
Copy Markdown
Member

Summary

Each UMI assignment strategy has a sequential implementation (crates/fgumi-umi/src/assigner.rs) and a parallel one (src/lib/umi/parallel_assigner.rs), and both are documented to "produce identical results". The parallel implementations fold UMI case (they count, match, and tie-break on the uppercased UMI), but the sequential implementations did not. On mixed-case UMIs the two could therefore diverge:

Strategy Sequential (before) Parallel Reachable via CLI?
Adjacency equal-count tie-break compared the raw first-seen string uppercased string No — group/dedup pre-uppercase non-paired UMIs
Edit matching was case-sensitive uppercased No — same CLI masking as adjacency
Paired counting, matching, and strand A/B assignment all case-sensitive fully case-insensitive (canonicalizes to uppercase) Yes — the sequential paired path keeps raw case (group.rs builds prefix:part0-prefix:part1 from the raw segments)

The paired case is the only one reachable through the CLI today: fgumi group/dedup selects the parallel paired assigner when --threads > 1 and the sequential one otherwise, so on mixed-case paired/duplex UMIs the two thread settings could produce different groupings and different strand assignments. The adjacency and edit divergences are library-only, because the CLI uppercases non-paired UMIs before assign().

Behavior comparison: fgbio main vs fgumi main vs this PR

How each tool/implementation handles UMI case, per strategy. "case-insensitive" = acgt and ACGT are treated as the same UMI; "case-sensitive" = they are treated as different.

Strategy fgbio main fgumi main — sequential (--threads 1) fgumi main — parallel (--threads N) This PR (both)
Identity case-insensitive case-insensitive case-insensitive case-insensitive (unchanged)
Edit case-sensitive case-sensitive case-insensitive case-insensitive
Adjacency case-sensitive counting case-insensitive, tie-break case-sensitive case-insensitive case-insensitive
Paired case-sensitive case-sensitive case-insensitive case-insensitive

Verified against fgbio main (GroupReadsByUmi.scala): its IdentityUmiAssigner.assign uppercases, but the edit/adjacency/paired assigners key on the raw (case-sensitive) UMI and the tool feeds them the raw RX tag — so fgbio is itself internally inconsistent on case across strategies.

Takeaways:

  • On standard (uppercase ACGT) UMIs, all three columns agree — every difference above manifests only on non-standard mixed-case UMIs.
  • fgbio main is internally inconsistent (identity case-insensitive; edit/adjacency/paired case-sensitive).
  • fgumi main is inconsistent between its own sequential and parallel implementations for edit/adjacency/paired (the bug this PR fixes).
  • This PR makes fgumi uniformly case-insensitive across all strategies and both implementations.
  • This does not change fgumi-vs-fgbio behavior on uppercase data. fgumi deliberately normalizes UMIs to uppercase (at the CLI boundary and now uniformly in the assigners), so on mixed-case input fgumi intentionally differs from fgbio's mostly-case-sensitive behavior — and cannot be byte-identical to fgbio there regardless, because fgumi's adjacency counting case-folds via BitEnc by design.

Fix

Fold case in all three sequential assigners so they match their parallel counterparts:

  • Adjacency — tie-break on the UMI bytes uppercased on the fly (Iterator::cmp, no per-comparison allocation).
  • Edit and Paired — uppercase the inputs once at the top of assign() and run the existing logic over the uppercased UMIs (strand-assignment logic is unchanged; it now simply operates on the uppercased forms).

This is a no-op for the uppercase UMIs the CLI emits, so it does not change CLI output on standard data.

Tests

  • Sequential-vs-parallel parity on mixed-case input for all three strategies (test_sequential_and_parallel_{adjacency,edit,paired}_agree_on_mixed_case), comparing grouping structure (and, for paired, strand). Each fails on main and passes here.
  • A sequential-only case-folded tie-break test for the adjacency assigner (test_adjacency_equal_count_tiebreak_is_case_insensitive).

cargo ci-fmt, cargo ci-lint, and cargo ci-test (2210 tests) all pass.

Reading order

  1. crates/fgumi-umi/src/assigner.rs — the three sequential fixes + docs.
  2. src/lib/umi/parallel_assigner.rs — the four new tests (and a corrected stale "matches the sequential assigner exactly" comment).

Summary by CodeRabbit

  • Bug Fixes

    • Updated UMI assignment to treat UMIs case-insensitively across error-correction strategies, including deterministic equal-count tie-breaking for mixed-case inputs.
    • Ensured sequential and parallel assignment produce matching grouping topology for adjacency, edit, and paired/duplex strategies.
  • Tests

    • Added regression coverage for case-insensitive behavior and parity between sequential and parallel implementations.
    • Included additional assertions covering paired strand/base relationships and consistent base linkage across orientations.

@nh13
nh13 temporarily deployed to github-actions June 23, 2026 22:11 — with GitHub Actions Inactive
@coderabbitai

coderabbitai Bot commented Jun 23, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: dacb8b50-995e-42ca-9f73-f3c58f82287a

📥 Commits

Reviewing files that changed from the base of the PR and between bceabf3 and 5215dc8.

📒 Files selected for processing (2)
  • crates/fgumi-umi/src/assigner.rs
  • src/lib/umi/parallel_assigner.rs

Walkthrough

Mixed-case UMI inputs are now folded to uppercase before counting, matching, and result mapping in the sequential assigners, with adjacency tie-breaking updated to use case-folded ordering and parity tests added for sequential versus parallel behavior.

Changes

Case-insensitive UMI assignment

Layer / File(s) Summary
SimpleErrorUmiAssigner case folding
crates/fgumi-umi/src/assigner.rs
Docs note case-insensitive matching; assign() builds upper_umis, resizes seen from folded input, deduplicates over upper_umis, and maps the final result from folded inputs.
AdjacencyUmiAssigner tie-break and test
crates/fgumi-umi/src/assigner.rs
Docs state case-insensitive counting and tie-breaking; equal-count sorting compares ASCII uppercased bytes instead of raw bytes; the regression test checks mixed-case inputs produce the same grouping topology as the case-folded ordering.
PairedUmiAssigner folding and helpers
crates/fgumi-umi/src/assigner.rs
Docs state case-insensitive counting, matching, and strand assignment; assign() folds paired UMIs before counting and fast-path assignment; is_same_umi and canonicalize uppercase before equivalence and canonical-form checks; the helper regression test covers mixed-case behavior.
Sequential and parallel parity tests
src/lib/umi/parallel_assigner.rs
ParallelAdjacencyAssigner docs state case-folded tie-breaking; same_partition compares grouping structure while ignoring MoleculeId; mixed-case parity tests cover adjacency, edit, and paired assigners, including paired strand and base-molecule checks.

Estimated code review effort

🎯 3 (Moderate) | ⏱️ ~25 minutes

Possibly related PRs

  • fulcrumgenomics/fgumi#47: Also changes AdjacencyUmiAssigner::assign() in crates/fgumi-umi/src/assigner.rs, touching the same assignment path as the tie-break update.

Suggested labels

bug

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 68.42% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly matches the main change: aligning sequential and parallel UMI assigners on case handling.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch nh/fix-assigner-case-parity

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Jun 23, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 91.11%. Comparing base (5ac181b) to head (5215dc8).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #455      +/-   ##
==========================================
+ Coverage   91.05%   91.11%   +0.06%     
==========================================
  Files          78       78              
  Lines       51324    51376      +52     
==========================================
+ Hits        46731    46813      +82     
+ Misses       4593     4563      -30     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 3

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/fgumi-umi/src/assigner.rs`:
- Around line 1614-1619: The paired helper methods `is_same_umi()` and
`canonicalize()` do not perform case-folding on their inputs while `assign()`
does through `upper_umis`, creating inconsistency in the case-insensitive
contract. Update both `is_same_umi()` (around line 1832-1836) and
`canonicalize()` (around line 1901-1904) to uppercase their input UMI parameters
before processing, matching the case-folding behavior already implemented in
`assign()`. This ensures mixed-case paired callers get consistent sequential
behavior that agrees with the parallel paired canonicalizer.

In `@src/lib/umi/parallel_assigner.rs`:
- Around line 1236-1311: The test functions
test_sequential_and_parallel_adjacency_agree_on_mixed_case,
test_sequential_and_parallel_edit_agree_on_mixed_case, and
test_sequential_and_parallel_paired_agree_on_mixed_case verify parity between
sequential and parallel implementations, but they do not verify compliance with
the fgbio baseline contract. Add programmatically generated expected outputs for
each test that validate the grouping/assignment results against the known fgbio
baseline behavior, or document and generate fixtures for any intentional
divergence from fgbio. This is especially important for
test_sequential_and_parallel_paired_agree_on_mixed_case since PairedUmiAssigner
is reachable via the CLI. The verification should ensure that the
output-changing UMI assignment behavior for mixed-case handling matches the
expected fgbio behavior.
- Around line 1303-1310: The current test at the assertion block starting with
same_partition() does not adequately verify that the reverse read keeps the same
base molecule ID as the forward reads. The current assert_ne check only verifies
that sequential[2] has a different partition from sequential[0], but this could
pass even if index 2 were assigned to a completely unrelated molecule. Add
explicit base-ID equality assertions after the same_partition() call to verify
that sequential[0], sequential[1], and sequential[2] all share the same base
molecule ID (just with different strand assignments), and add the same base-ID
equality checks for the parallel result to ensure both sequential and parallel
results maintain strand-assignment behavior correctly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 421a8d6e-e4a4-4da1-8468-0adbbd26c432

📥 Commits

Reviewing files that changed from the base of the PR and between 5ac181b and 1814ac2.

📒 Files selected for processing (2)
  • crates/fgumi-umi/src/assigner.rs
  • src/lib/umi/parallel_assigner.rs

Comment thread crates/fgumi-umi/src/assigner.rs
Comment thread src/lib/umi/parallel_assigner.rs
Comment thread src/lib/umi/parallel_assigner.rs
@nh13
nh13 force-pushed the nh/fix-assigner-case-parity branch from 1814ac2 to b451fcf Compare June 24, 2026 01:59
@nh13
nh13 temporarily deployed to github-actions June 24, 2026 01:59 — with GitHub Actions Inactive
@nh13
nh13 force-pushed the nh/fix-assigner-case-parity branch from b451fcf to bceabf3 Compare June 24, 2026 02:09
@nh13
nh13 temporarily deployed to github-actions June 24, 2026 02:09 — with GitHub Actions Inactive
@nh13

nh13 commented Jun 24, 2026

Copy link
Copy Markdown
Member Author

fgbio baseline validation for the case-folding change

Following up on the request for fgbio-baseline coverage: I drove fgbio's actual UmiAssigner classes (SimpleErrorUmiAssigner, AdjacencyUmiAssigner, PairedUmiAssigner) on the exact inputs the new parity tests use, to confirm empirically what fgbio does with mixed-case UMIs rather than asserting it from memory.

What fgbio does

  • Matching is case-insensitive. Sequences.countMismatches uppercases each base (Character.toUpperCase), so the raw edit assigner groups ACGT and acgt together.
  • Adjacency counting and the equal-count tie-break are case-sensitive. Counts come from SimpleCounter over the raw strings, and ordering is sortBy((-count, umi)) on the raw UMI string.
  • The GroupReadsByUmi CLI uppercases every UMI before assignment (canonicalize(rawTag.toUpperCase)), so fgbio's assigner only ever sees uppercase UMIs in practice.

Result

With CLI-style uppercasing (what fgbio actually feeds its assigners), fgbio and fgumi produce identical groupings on mixed-case input:

input fgbio raw class fgbio CLI (uppercases first) fgumi (this PR)
edit [ACGT, acgt] {ACGT, acgt} {ACGT} {ACGT, acgt} → one molecule
adjacency [AAAaAA×2, AAAGAC×2, AAAAAC] {AAAGAC, AAAAAC}, {AAAaAA} {AAAAAA, AAAAAC}, {AAAGAC} same as fgbio CLI
paired [ACGT-TGCA, acgt-tgca, TGCA-ACGT] acgt-tgca split off one molecule, fwd → strand A, rev → strand B same as fgbio CLI

So this PR makes fgumi converge with fgbio's user-facing output on mixed-case input — it does not diverge. The only difference is where the case fold happens: fgbio folds in the CLI before calling the assigner, while fgumi folds inside the assigner itself, which keeps the library API from being misused case-sensitively. On uppercase input (all real data) the fold is a no-op, so byte-for-byte parity with fgbio is unchanged there.

The parity tests now pin these fgbio-derived groupings (not just sequential-vs-parallel agreement), and the # Case sensitivity doc comments describe the relationship accurately.

@nh13

nh13 commented Jun 24, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jun 24, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/fgumi-umi/src/assigner.rs`:
- Around line 2446-2466: The new PairedUmiAssigner case-insensitivity tests
cover local behavior, but the output-changing UMI assignment semantics also need
generated fgbio-baseline identity coverage or an explicit documented divergence.
Add or reference a programmatic baseline test that asserts identity for the
affected edit, adjacency, and paired mixed-case inputs, using the existing
PairedUmiAssigner helpers and canonicalization flow so the expected behavior is
pinned against fgbio. If the new behavior intentionally differs from fgbio,
document that divergence in the same baseline test area instead of leaving only
local assertions.

In `@src/lib/umi/parallel_assigner.rs`:
- Around line 1339-1352: The paired fgbio-baseline test in parallel_assigner
only verifies same/opposite strand behavior, so a global A/B flip could still
pass; tighten the assertions in the sequential/parallel comparison block by
explicitly checking the expected PairedA/PairedB strand variants for indices 0,
1, and 2. Use the existing sequential/parallel test values in parallel_assigner
to anchor the check, and keep the base_id_string and opposite-strand assertions
alongside the new absolute orientation assertions so parity coverage remains
intact.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: a8293fb1-f8c5-4b9a-b064-90a8c008de11

📥 Commits

Reviewing files that changed from the base of the PR and between b451fcf and bceabf3.

📒 Files selected for processing (2)
  • crates/fgumi-umi/src/assigner.rs
  • src/lib/umi/parallel_assigner.rs

Comment thread crates/fgumi-umi/src/assigner.rs
Comment thread src/lib/umi/parallel_assigner.rs
@nh13

nh13 commented Jun 24, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jun 24, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

The adjacency, edit, and paired UMI assigners each have a sequential and a
parallel implementation that are documented to produce identical groupings.
The parallel implementations fold UMI case (count/match/tie-break on the
uppercased UMI), but the sequential implementations did not, so on mixed-case
UMIs the two could diverge:

- Adjacency: the equal-count tie-break compared the raw first-seen string
  while the parallel assigner compared the uppercased string. Reachable only
  via direct library calls; the CLI pre-uppercases non-paired UMIs.
- Edit: sequential matching was case-sensitive. Same CLI masking as adjacency.
- Paired: sequential counting, matching, and strand assignment were all
  case-sensitive. NOT masked by the CLI, since the sequential paired path
  keeps raw case, so `group`/`dedup --threads 1` vs `--threads N` could group
  mixed-case paired UMIs (and assign strands) differently.

Fold case in all three sequential assigners so they match their parallel
counterparts. This is a no-op for the uppercase UMIs the CLI emits and does
not change fgbio parity on standard (uppercase) data. Add sequential-vs-
parallel parity tests on mixed-case input for all three strategies, plus a
sequential-only case-folded tie-break test for the adjacency assigner.
@nh13
nh13 force-pushed the nh/fix-assigner-case-parity branch from bceabf3 to 5215dc8 Compare June 24, 2026 16:14
@nh13
nh13 temporarily deployed to github-actions June 24, 2026 16:14 — with GitHub Actions Inactive
@nh13
nh13 merged commit 93d082d into main Jun 24, 2026
10 checks passed
@nh13
nh13 deleted the nh/fix-assigner-case-parity branch June 24, 2026 17:37
@nh13 nh13 mentioned this pull request Jun 24, 2026

This branch was previously deployed

1 inactive deployment
github-actions — 5215dc81 Deployed Jun 24, 2026 by nh13 via coverage #1744
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant