Repository navigation
fix(group): add reverse-orientation edges to the parallel paired assigner (GRP3-01) - #525
Conversation
|
Note Reviews pausedIt looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the Use the following commands to manage reviews:
Use the checkboxes below for quick actions:
WalkthroughParallel paired assignment now discovers reverse-orientation UMI edges, handles split-point collisions through sequential fallback, tracks cluster roots, computes strand labels, and preserves distinct invalid UMIs. Tests validate fgbio parity, sequential equivalence, and deterministic results across thread counts. ChangesPaired UMI orientation handling
Estimated code review effort: 4 (Complex) | ~60 minutes Sequence Diagram(s)sequenceDiagram
participant Reads
participant ParallelPairedAssigner
participant discover_paired_reverse_edges
participant paired_canonical_strands
Reads->>ParallelPairedAssigner: submit paired UMI reads
ParallelPairedAssigner->>ParallelPairedAssigner: detect split-point collision
ParallelPairedAssigner->>discover_paired_reverse_edges: discover reverse-orientation edges
discover_paired_reverse_edges-->>ParallelPairedAssigner: combined adjacency edges
ParallelPairedAssigner->>paired_canonical_strands: compute root-relative strands
paired_canonical_strands-->>ParallelPairedAssigner: PairedA/PairedB labels
ParallelPairedAssigner-->>Reads: molecule assignments
Possibly related issues
Possibly related PRs
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #525 +/- ##
==========================================
+ Coverage 92.65% 92.84% +0.19%
==========================================
Files 166 166
Lines 100136 102064 +1928
==========================================
+ Hits 92776 94765 +1989
+ Misses 7360 7299 -61 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
e199117 to
742fa85
Compare
742fa85 to
dcbbed5
Compare
|
Real-data confirmation (resolves the deferred agilent check). Ran on the full 53,139,936-record agilent-hs2 paired dataset with the parallel path forced on every ≥128-template position group (the natural GRP3-01 condition):
This complements the GRP3-T2 proptest (parallel ≡ sequential across threads {1,4,16}, 3000 cases). The |
3745dc7 to
718460d
Compare
dcbbed5 to
e84180f
Compare
718460d to
032beae
Compare
e84180f to
89f556d
Compare
6396833 to
bd1f81c
Compare
89f556d to
81c8d8c
Compare
bd1f81c to
d03a568
Compare
81c8d8c to
9555fd6
Compare
d03a568 to
8b6838b
Compare
b2859e1 to
2fc5c0a
Compare
2fc5c0a to
015573d
Compare
✅ Action performedReview finished.
|
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@src/lib/umi/parallel_assigner.rs`:
- Around line 280-281: Update the key construction in the forward-index map used
by the parallel paired assigner so it preserves the canonical pair’s split point
instead of relying on BitEnc alone. Apply the same split-aware key consistently
wherever forward assignments are looked up, and add a regression covering paired
UMIs with identical concatenated bases but different half lengths.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 8f8cf27a-d9b9-448e-ba68-37b1ba7b573e
📒 Files selected for processing (2)
proptest-regressions/lib/umi/parallel_assigner.txtsrc/lib/umi/parallel_assigner.rs
9518d65 to
452ffb9
Compare
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@src/lib/umi/parallel_assigner.rs`:
- Around line 275-304: Update paired assignment around
discover_paired_reverse_edges and its callers so every asymmetric paired pool
falls back to the sequential assigner, not only pools with encoding collisions.
Detect differing forward/reverse half lengths and delegate those pools before
parallel reverse-edge discovery; otherwise make reverse matching split-aware
while preserving the sequential assigner’s threshold behavior and output.
- Around line 1920-2032: Replace the hand-derived expectations in
test_paired_matches_fgbio_oracle with a checked-in fixture generated
programmatically by a pinned fgbio invocation, including the command/version and
reproducible input data. Load the fixture in assert_matches_fgbio_oracle and
validate both PairedUmiAssigner and ParallelPairedAssigner against the generated
fgbio base-molecule and strand results, preserving coverage for the listed
reverse, palindrome, tie-break, and threshold cases.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 87bf8561-f7bb-4e53-84ef-136d74f7b5ab
📒 Files selected for processing (2)
proptest-regressions/lib/umi/parallel_assigner.txtsrc/lib/umi/parallel_assigner.rs
| /// Assert `actual` matches a hand-derived fgbio oracle. | ||
| /// | ||
| /// `expected[i] = (group, strand)`, where `group` is an arbitrary label shared by | ||
| /// every read fgbio places in one *base* molecule and `strand` is fgbio's absolute | ||
| /// duplex suffix (`'A'` for `/A` / [`MoleculeId::PairedA`], `'B'` for `/B` / | ||
| /// [`MoleculeId::PairedB`]). This checks the base-molecule partition and each read's | ||
| /// absolute strand, but not the concrete numeric molecule id (a relabeling). | ||
| /// | ||
| /// The oracle values are derived by hand from fgbio's `PairedUmiAssigner` | ||
| /// (`GroupReadsByUmi.assignIdsToNodes`): the canonical (lexically-smaller-half-first) | ||
| /// spelling of a cluster's root maps to `/A` and its reverse to `/B`; a descendant | ||
| /// takes `/A` for the orientation closer to the root and `/B` for the other; and | ||
| /// fgbio's `Map` last-write-wins gives the palindrome quirks (a palindrome root | ||
| /// canonical spelling resolves to `/B`, a palindrome descendant to `/A`). Because the | ||
| /// oracle is independent of BOTH fgumi assigners, a shared sequential/parallel defect | ||
| /// cannot satisfy it. | ||
| fn assert_matches_fgbio_oracle(actual: &[MoleculeId], expected: &[(u32, char)]) { | ||
| assert_eq!( | ||
| actual.len(), | ||
| expected.len(), | ||
| "length mismatch: actual={actual:?} expected={expected:?}" | ||
| ); | ||
| // Base-molecule partition: reads i and j share a base id IFF the oracle groups match. | ||
| for i in 0..actual.len() { | ||
| for j in 0..actual.len() { | ||
| let actual_same = actual[i].base_id_string() == actual[j].base_id_string(); | ||
| let expected_same = expected[i].0 == expected[j].0; | ||
| assert_eq!( | ||
| actual_same, expected_same, | ||
| "base-partition mismatch at ({i},{j}); actual={actual:?} expected={expected:?}" | ||
| ); | ||
| } | ||
| } | ||
| // Absolute strand per read. | ||
| for (i, (id, &(_, strand))) in actual.iter().zip(expected).enumerate() { | ||
| let actual_strand = match id { | ||
| MoleculeId::PairedA(_) => 'A', | ||
| MoleculeId::PairedB(_) => 'B', | ||
| other => panic!("read {i}: expected a paired strand, got {other:?}"), | ||
| }; | ||
| assert_eq!( | ||
| actual_strand, strand, | ||
| "strand mismatch at read {i}; actual={actual:?} expected={expected:?}" | ||
| ); | ||
| } | ||
| } | ||
|
|
||
| /// GRP3-01 fgbio oracle: pin the fgbio-derived base-molecule partition AND absolute | ||
| /// `/A` `/B` strand for the reverse-aware paired cases, and assert BOTH the sequential | ||
| /// `PairedUmiAssigner` and the parallel `ParallelPairedAssigner` reproduce it. | ||
| /// | ||
| /// The other reverse-orientation suites compare parallel against the sequential | ||
| /// assigner only, so a defect shared by both would pass. This table's expectations | ||
| /// are hand-derived from fgbio's `PairedUmiAssigner` (see | ||
| /// [`assert_matches_fgbio_oracle`]), giving an oracle independent of either fgumi | ||
| /// implementation. Covers reverse edges, palindrome roots, palindrome children, | ||
| /// equal-count tie-breaking, and the child-count threshold boundary (inclusion and | ||
| /// exclusion), per the grouping path's fgbio-parity requirement. | ||
| #[rstest] | ||
| // Reverse-orientation edge that canonicalization alone misses; equal counts, so the | ||
| // root is the lexically-smaller canonical "ATTT-CAAA" (=> /A) and "CAAA-GTTT" the | ||
| // reversed descendant (=> /B). The tie-break choosing the root is what fixes strand | ||
| // here: rooting on the other member would invert both labels. | ||
| #[case::reverse_edge_equal_count_tiebreak( | ||
| &["CAAA-GTTT", "ATTT-CAAA"], | ||
| &[(0, 'B'), (0, 'A')], | ||
| )] | ||
| // Palindrome root (halves equal => umi == reverse(umi)): fgbio maps the root canonical | ||
| // spelling then its (identical) reverse, so last-write-wins lands the palindrome root | ||
| // on /B. Its non-palindrome descendant "ACGT-ACGA", reached in the reverse spelling, | ||
| // takes /A. | ||
| #[case::palindrome_root( | ||
| &["ACGT-ACGT", "ACGT-ACGT", "ACGT-ACGA"], | ||
| &[(0, 'B'), (0, 'B'), (0, 'A')], | ||
| )] | ||
| // Palindrome descendant "AAAA-AAAA" under a non-palindrome root "AAAA-AAAC": fgbio's | ||
| // else-branch maps the palindrome child to /B then its identical reverse to /A, so | ||
| // last-write-wins lands the palindrome child on /A. The reverse spelling of the root | ||
| // ("AAAC-AAAA", index 2) is the strand-/B counterpart, so not every read is /A. | ||
| #[case::palindrome_child( | ||
| &["AAAA-AAAC", "AAAA-AAAC", "AAAC-AAAA", "AAAA-AAAA"], | ||
| &[(0, 'A'), (0, 'A'), (0, 'B'), (0, 'A')], | ||
| )] | ||
| // Threshold boundary, INCLUDED: descendant count 3 == root_count(5)/2 + 1 == 3, so the | ||
| // "AAAA-CCCG" node joins the "AAAA-CCCC" molecule (one group). | ||
| #[case::threshold_boundary_included( | ||
| &[ | ||
| "AAAA-CCCC", "AAAA-CCCC", "AAAA-CCCC", "AAAA-CCCC", "AAAA-CCCC", | ||
| "AAAA-CCCG", "AAAA-CCCG", "AAAA-CCCG", | ||
| ], | ||
| &[(0, 'A'), (0, 'A'), (0, 'A'), (0, 'A'), (0, 'A'), (0, 'A'), (0, 'A'), (0, 'A')], | ||
| )] | ||
| // Threshold boundary, EXCLUDED: descendant count 4 > root_count(5)/2 + 1 == 3, so the | ||
| // "AAAA-CCCG" node splits off into its own molecule (group 1), each a /A root. | ||
| #[case::threshold_boundary_excluded( | ||
| &[ | ||
| "AAAA-CCCC", "AAAA-CCCC", "AAAA-CCCC", "AAAA-CCCC", "AAAA-CCCC", | ||
| "AAAA-CCCG", "AAAA-CCCG", "AAAA-CCCG", "AAAA-CCCG", | ||
| ], | ||
| &[(0, 'A'), (0, 'A'), (0, 'A'), (0, 'A'), (0, 'A'), (1, 'A'), (1, 'A'), (1, 'A'), (1, 'A')], | ||
| )] | ||
| fn test_paired_matches_fgbio_oracle(#[case] umis: &[&str], #[case] expected: &[(u32, char)]) { | ||
| let umis: Vec<Umi> = umis.iter().map(|s| (*s).to_string()).collect(); | ||
|
|
||
| // Both the sequential assigner AND the parallel assigner (at multiple thread | ||
| // counts) must reproduce the independent fgbio oracle. | ||
| let sequential = crate::umi::PairedUmiAssigner::new(1).assign(&umis); | ||
| assert_matches_fgbio_oracle(&sequential, expected); | ||
| for threads in [1usize, 4, 16] { | ||
| let parallel = ParallelPairedAssigner::new(1, threads).assign(&umis); | ||
| assert_matches_fgbio_oracle(¶llel, expected); | ||
| } | ||
| } |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | 🏗️ Heavy lift
Generate the fgbio oracle programmatically.
The expectations are documented as hand-derived, so this still does not satisfy the required reproducible fgbio-baseline coverage. Generate a checked-in fixture from a pinned fgbio invocation and validate both assigners against it.
As per path instructions, output-changing grouping requires parity coverage against “the sequential code path and the fgbio baseline, generated programmatically.”
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@src/lib/umi/parallel_assigner.rs` around lines 1920 - 2032, Replace the
hand-derived expectations in test_paired_matches_fgbio_oracle with a checked-in
fixture generated programmatically by a pinned fgbio invocation, including the
command/version and reproducible input data. Load the fixture in
assert_matches_fgbio_oracle and validate both PairedUmiAssigner and
ParallelPairedAssigner against the generated fgbio base-molecule and strand
results, preserving coverage for the listed reverse, palindrome, tie-break, and
threshold cases.
Source: Path instructions
There was a problem hiding this comment.
Deferring this one, intentionally leaving the thread open for a maintainer call rather than acting on it in this PR.
The test_paired_matches_fgbio_oracle expectations are already independent of both fgumi implementations — they are hand-derived from fgbio's PairedUmiAssigner.assignIdsToNodes algorithm (documented case-by-case in the table: reverse edges, palindrome roots/children via fgbio's last-write-wins Map, equal-count tie-breaking, and the child-count threshold boundary on both sides). Because the oracle is derived independently of both the sequential and parallel fgumi assigners, a defect shared by both cannot satisfy it — which is the coverage the path instruction is after.
Replacing it with a checked-in fixture generated by a pinned fgbio invocation runs against two repo conventions: fgumi's test suite is pure Rust with no JVM/fgbio runtime dependency, and this project generates test data programmatically rather than committing fixture files. A programmatic alternative (a Rust reimplementation of fgbio's paired assignment to generate expectations at test time) would just re-implement the logic under test, adding a second copy to keep in sync without adding independent signal.
Leaving this open so a maintainer can weigh in on whether to invest in a fgbio-provenanced fixture harness (better as its own infra task than folded into this focused reverse-edge fix). Happy to file a tracking issue if that's the preferred direction.
There was a problem hiding this comment.
@nh13 Risk remains: the hand-derived oracle is independent, but it is not reproducible, pinned fgbio-baseline coverage as required by the path instruction.
Keeping this open for maintainer scope judgment is reasonable. If you confirm, I can create a follow-up GitHub issue for a pinned, generated fgbio fixture/harness—explicitly preserving the Rust-only default test path and avoiding a JVM runtime requirement in normal CI.
✏️ Learnings added
Learnt from: nh13
Repo: fulcrumgenomics/fgumi PR: 525
File: src/lib/umi/parallel_assigner.rs:1963-2075
Timestamp: 2026-07-16T15:17:40.377Z
Learning: In `src/lib/umi/parallel_assigner.rs`, fgumi’s normal test suite is pure Rust and does not depend on a JVM or fgbio runtime; the repository convention favors programmatic test-data generation over committed fixture files. Any future fgbio-baseline harness for paired UMI parity should preserve a Rust-only default test path.
Note: Learnings are effective only in the context of similar code segments. To apply general coding guidelines, please configure review instructions. You can manage existing learnings in the app.
452ffb9 to
9cb3ecc
Compare
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@src/lib/umi/parallel_assigner.rs`:
- Around line 2064-2073: Update test_paired_matches_fgbio_oracle to parameterize
max_mismatches and pass that threshold to both PairedUmiAssigner and
ParallelPairedAssigner. Extend the oracle cases with a reverse-only distance-2
input using max_mismatches = 2, while retaining the existing k=1 cases and
fgbio-oracle assertions.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 6fedf953-c547-4029-b5ed-4e8a81e4ef6b
📒 Files selected for processing (2)
proptest-regressions/lib/umi/parallel_assigner.txtsrc/lib/umi/parallel_assigner.rs
…gner (GRP3-01)
The parallel `ParallelPairedAssigner` discovered adjacency edges by forward
Hamming distance on canonical (lexically smaller of A-B / B-A) forms only,
claiming "canonicalization already handles reverse-orientation matching". It does
not: a single mismatch can flip which half is lexically smaller, so two reads of
one molecule canonicalize to forms that are far apart forward yet within threshold
when one is reversed. The multi-threaded paired/duplex path therefore over-split
molecules versus fgbio and versus fgumi's own sequential `PairedUmiAssigner`
(reachable at the Paired Auto parallel threshold of 128 templates/position, and
forced by `--allow-unmapped`).
The sequential assigner matches on `within(l, r) OR within(reverse(l), r)`. This
change unions the forward edges with a reverse-orientation edge pass
(`discover_paired_reverse_edges`), so the parallel base-molecule partition now
matches the sequential one.
Reverse-orientation edges also mean a read reached via such an edge is on the
OPPOSITE strand, so strand (/A /B) can no longer be derived from "raw == its own
canonical". Strand is now assigned relative to each molecule's root (highest-count
member), mirroring the sequential assigner — including its palindrome insert-order
quirk (a UMI whose halves are identical is its own reverse; a palindromic root is
labeled B and a palindromic non-root child A, matching the sequential's last-write
insert order).
Adds `test_parallel_paired_groups_reverse_orientation_edge` (the minimal
"CAAA-GTTT" / "ATTT-CAAA" case) and GRP3-T2, a parallel-vs-sequential PAIRED parity
proptest that builds pools from random molecules in both orientations with
single-base mutations and asserts identical base-molecule AND strand partitions at
threads {1, 4, 16} (which also pins thread-determinism). The proptest surfaced the
palindrome strand cases now covered by the committed regression seeds.
Also key the parallel main-path invalid-UMI fallback by the raw uppercase string
instead of the canonical form. Two invalid (non-encodable) UMIs that are reverses
of each other (e.g. "ACGN-TTTT" / "TTTT-ACGN") canonicalize to the same form, so
the canonical keying merged them into one Single molecule whenever a valid UMI was
also present (the main path, not the all-invalid fast path). The sequential
assigner and the all-invalid branch both key by the raw uppercase string and keep
such reverses distinct, so this was a --threads-dependent divergence. Now all three
paths agree (one molecule per distinct raw UMI string), covered by a new
paired_reversed_invalid case in the cross-assigner parity harness.
Also guard the pre-existing `BitEnc`-drops-dash limitation in the parallel paired
path. `BitEnc::from_umi_str` discards the `-`, so the parallel forward AND reverse
edit distances are dash-blind, whereas fgbio and the sequential `PairedUmiAssigner`
compare the dash-delimited string byte-for-byte in both orientations. The two agree
exactly when every UMI has symmetric halves (the dash sits at a fixed position); with
asymmetric halves (read-1 UMI length != read-2 UMI length) the dash-blind path
diverges in two ways: distinct forms with the same concatenated bases but different
splits collide to one encoding (`AC-GTA` / `ACG-TA`), and the halves-swapped reverse
encoding can forge a within-threshold reverse edge the dash-sensitive reference never
draws (`A-AC` reverse `ACA` is Hamming-1 from `A-CT`'s `ACT`). Until the full
split-aware rework lands (tracked in #586), delegate any pool containing an
asymmetric-halves UMI to the sequential assigner, which is byte-for-byte faithful to
fgbio; symmetric pools take the fast parallel path unchanged. Adds
`test_parallel_paired_mixed_half_length_collision_matches_sequential` (forward
collision) and `test_parallel_paired_asymmetric_noncollision_matches_sequential` (the
reverse-edge case with no forward collision), both of which diverged before the guard.
9cb3ecc to
855db1b
Compare
…halves The parallel paired UMI assigner encoded canonical paired UMIs with `BitEnc::from_umi_str`, which drops the `-`. With asymmetric halves the split point varies, so distinct canonical forms collide (`AC-GTA` / `ACG-TA`) and a halves-swapped reverse encoding can forge within-threshold edges the dash-sensitive reference never draws. PR #525 worked around this by delegating any asymmetric-halves pool to the single-threaded sequential `PairedUmiAssigner`. Replace that fallback with a native dash-aware path. For pools that contain an asymmetric-halves UMI, discover edges by comparing the full dash-delimited canonical strings position-for-position -- exactly the sequential assigner's `matches_paired` relation -- in parallel over ordered `(parent, child)` pairs. The relation is directional (with asymmetric halves `reverse()` moves the dash, so `matches_paired` is not symmetric), and the sequential adjacency graph is a dynamic directed BFS in which a node pulled into a cluster can then absorb a lower-indexed node. So the asymmetric branch feeds a directed adjacency to the existing count-gated BFS, reproducing the sequential traversal exactly. The common symmetric pool keeps the sub-quadratic `BitEnc` neighbour-generation path unchanged. Add an asymmetric-halves parity proptest across threads {1,4,16} and pinned regression cases for the two subtle failure modes (directional reverse edge; a child absorbed by a higher-indexed node). Closes #586.
Stacked on #510 (base branch
nh/fix-group-unequal-umi-length) — both touchcrates/fgumi-umi/src/assigner.rsandsrc/lib/umi/parallel_assigner.rs. Merge #510 first.Fixes the audit's second remaining S1, GRP3-01.
The bug
The parallel
ParallelPairedAssignerdiscovered adjacency edges by forward Hamming distance on canonical (lexically smaller ofA-B/B-A) forms only, on the claim that "canonicalization already handles reverse-orientation matching." It does not: a single mismatch can flip which half is lexically smaller, so two reads of one molecule canonicalize to forms that are far apart forward yet within threshold when one is reversed.Minimal case (
--edits 1):CAAA-GTTTandATTT-CAAAare the same molecule (reverse ofCAAA-GTTTisGTTT-CAAA, Hamming-1 fromATTT-CAAA), but their canonical forms differ at every base forward. The sequentialPairedUmiAssignergroups them via itswithin(l, r) OR within(reverse(l), r)check; the parallel path split them:The multi-threaded paired/duplex path therefore over-split molecules versus fgbio and versus fgumi's own sequential assigner. Reachable at the Paired Auto parallel threshold of 128 templates/position, and forced by
--allow-unmapped.The fix
discover_paired_reverse_edges), mirroring the sequentialmatches_paired. The generic forwarddiscover_edges_parallel_kis shared with the non-paired assigners and is left untouched./A/Bcan no longer be derived from "raw == its own canonical". Strand is now assigned relative to each molecule's root, mirroring the sequential assigner — including its palindrome insert-order quirk (a UMI whose halves are identical is its own reverse; a palindromic root is labeledB, a palindromic non-root childA).Tests
test_parallel_paired_groups_reverse_orientation_edge— the minimalCAAA-GTTT/ATTT-CAAAcase (RED before the fix).cargo ci-fmt && ci-lint && ci-testgreen (2221 tests); no regression in the existing paired / group / determinism suites.Notes
group --strategy paired --threads 16vs the hardenedcompare bams) could not be run — the scratch SSD is not mounted right now. The programmatic RED test + proptest fully reproduce the divergence and verify parity; the agilent comparison can be run for extra confidence when the SSD is back.coderabbit --agentflagged one minor item inSimpleErrorUmiAssigner(the edit assigner,assigner.rs:989-991) about fix(group): reject UMIs of differing length in all assigners (GRP-01) #510's length-guard ordering on invalid UMIs — outside this diff and unrelated to the paired assigner, so left for fix(group): reject UMIs of differing length in all assigners (GRP-01) #510 rather than mixed in here.Summary by CodeRabbit
Bug Fixes
Testing