Skip to content

fix(umi): make consensus read downsampling deterministic - #1166

Merged
nh13 merged 2 commits into
mainfrom
nh_deterministic-consensus-downsampling
Aug 10, 2026
Merged

nh13 merged 2 commits into
mainfrom
nh_deterministic-consensus-downsampling

Conversation

@nh13

@nh13 nh13 commented Aug 7, 2026 •

Copy link
Copy Markdown
Member

Problem

When a cap is set (--max-reads, --max-reads-per-strand, --max-read-pairs) and a tag family exceeds it, consensus output was not reproducible under --threads > 1. The same input BAM produced different consensus bases, qualities and depths on every run.

The downsampler used a Random(42) held on the caller instance:

private val random = new Random(42)
...
val capped = if (reads.size <= this.options.maxReads) reads else this.random.shuffle(reads).take(this.options.maxReads)

The RNG is stateful, so which reads survive depends on how many families that instance already downsampled. ConsensusCallingIterator gives each thread its own caller via emptyClone(), and ParIterator hands chunks to a fork-join pool, so the family-to-thread assignment — and hence each caller's RNG position for a given family — varies run to run.

This affects CallMolecularConsensusReads, CallDuplexConsensusReads and CallCodecConsensusReads, all of which downsample through VanillaUmiConsensusCaller.consensusCall. Single-threaded runs were already deterministic.

Fix

Rank reads by a Murmur3 hash of their read name and keep the lowest-ranking ones. The retained subset becomes a pure function of the family, independent of the caller's history, the thread count, and the order reads arrive in. This mirrors the existing hash-based downsampling in CollectDuplexSeqMetrics and SamOrder.Random.

The rank is computed once per read rather than passed to sortBy, which re-evaluates its key on every comparison (measured at ~25-30x the element count for large families, each evaluation hashing a read name and allocating an Option).

Behaviour changes

Two things change for anyone who sets a cap. Both are intentional and want a release note.

  1. Which reads are retained differs from 4.1.0, single-threaded runs included. There is no way to fix the determinism bug while preserving the old selection, since the old selection was the bug.
  2. Both ends of a template are now retained or discarded together. Ends are downsampled by separate consensusCall invocations; under the old independent shuffles each end sampled templates independently, whereas mates share a read name and so now share a rank. This makes --max-read-pairs actually cap read pairs. It is pinned by a test and stated in the arg docs.

Tests

Four new tests, each watched failing before the implementation existed:

  • the subset a caller picks does not depend on how many families it already downsampled — the original bug;
  • CallDuplexConsensusReads produces byte-identical output at --threads 1 and --threads 8 over 300 molecules (spanning several ConsensusCallingIterator chunks rather than relying on fork-join splitting within one);
  • the subset is invariant to input read order;
  • both ends of a template are retained together.

Plus a golden test pinning which reads are retained, whose expected value was derived independently from htsjdk's Murmur3(42) rather than from fgbio. Without it, a deterministic-but-biased rule (e.g. sorting by read name, which on Illumina names sorts by tile/x/y) passes every test above.

Full suite: 1611 tests, 0 failures.

Second commit

CallDuplexConsensusReads accepted --max-reads-per-strand 0, which emptied every strand and then failed deep in consensus calling with a bare IllegalArgumentException: requirement failed: Too few reads to create a consensus. The sibling tools already validate their equivalent options. Separate commit; reviewable independently.

Known gap, deliberately not addressed here

Reads dropped by downsampling are not routed through rejectRecords: they are counted as used in raw_reads_used / frac_raw_reads_used and never reach --rejects. A family of 10 capped to 3 reports raw_reads_used = 10 while the consensus carries cD:i:3. This is pre-existing — the old shuffle.take lost them the same way — and fixing it means a new RejectionReason, touching the usedByVanilla/usedByDuplex/usedByCodec tables and metrics output. It does not belong in a determinism fix, and is tracked separately in #1167.

@nh13
nh13 requested review from clintval and tfenne as code owners August 7, 2026 19:44
@coderabbitai

coderabbitai Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

You’ve reached a temporary PR review limit under our Fair Usage Limits Policy.

Your recent review volume is higher than typical usage, so adaptive limits are currently applied.

Next review available in: 5 minutes

Enable usage-based reviews in Billing to review now. Otherwise, wait until the next included review is available.
You're only billed for reviews past your plan's rate limits ($0.25/file).

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 41d9a404-caae-4e65-a852-5ed02487deb3

📥 Commits

Reviewing files that changed from the base of the PR and between 798aca4 and e748657.

📒 Files selected for processing (7)
  • src/main/scala/com/fulcrumgenomics/umi/CallCodecConsensusReads.scala
  • src/main/scala/com/fulcrumgenomics/umi/CallDuplexConsensusReads.scala
  • src/main/scala/com/fulcrumgenomics/umi/CallMolecularConsensusReads.scala
  • src/main/scala/com/fulcrumgenomics/umi/CodecConsensusCaller.scala
  • src/main/scala/com/fulcrumgenomics/umi/VanillaUmiConsensusCaller.scala
  • src/test/scala/com/fulcrumgenomics/umi/CallDuplexConsensusReadsTest.scala
  • src/test/scala/com/fulcrumgenomics/umi/VanillaUmiConsensusCallerTest.scala

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Aug 7, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 95.95%. Comparing base (798aca4) to head (e748657).

Additional details and impacted files
@@           Coverage Diff           @@
##             main    #1166   +/-   ##
=======================================
  Coverage   95.95%   95.95%           
=======================================
  Files         132      132           
  Lines        8347     8354    +7     
  Branches      984      943   -41     
=======================================
+ Hits         8009     8016    +7     
  Misses        338      338           
Flag Coverage Δ
unittests 95.95% <100.00%> (+<0.01%) ⬆️

Flags with carried forward coverage won't be shown. Click here to find out more.

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

github-actions Bot commented Aug 7, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.6.1

🚀 View preview at
https://fulcrumgenomics.github.io/fgbio/pr-preview/pr-1166/

Built to branch gh-pages at 2026-08-10 15:56 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

nh13 added a commit to fulcrumgenomics/fgumi that referenced this pull request Aug 9, 2026
Downsampling selected which reads to keep by shuffling with a seeded StdRng held
on the consensus caller. The RNG is stateful, so the surviving subset for a tag
family depended on how many families that caller instance had already processed.

fgumi did not have fgbio's thread-count bug: the unified pipeline builds a fresh
caller per batch at a fixed batch size, so --threads 1 and --threads 8 agreed.
But the no---threads path holds a single caller for the entire run, putting its
RNG at a different position for every family, so the two execution modes
disagreed under a cap. On a 77k-molecule library, duplex output at
--max-reads-per-strand 2 differed on 566 of 77,804 records (0.73%). Uncapped runs
always agreed.

Rank reads by a Murmur3 hash of the read name and keep the lowest-ranking ones.
The retained subset becomes a pure function of the family: independent of caller
history, execution mode, and the order reads arrive in. Because both ends of a
template share a read name they share a rank, so a template is retained or
discarded as a unit -- which also recovers reads the old shuffle lost by
splitting templates across ends (simplex --max-reads 2 on that library goes from
726,854 to 734,536 consensus reads, matching fgbio's count).

The rank is computed once per read and stored, never inside a sort key closure:
Rust's sort_by_key may evaluate its closure more than once per element, so
hashing there would re-hash on every comparison. Ranks compare as signed i32,
matching fgbio's Scala Int, and the sort is stable so tied ranks keep input
order. Selection now lives in one place, caller::select_lowest_ranking, shared
by the raw-record and codec paths.

With no RNG left on the caller there is no per-instance state for downsampling to
depend on, so the seed option and the rand dependency are removed and the
determinism becomes structural rather than incidental.

Which reads are retained differs from previous releases for any run that sets a
cap. There is no way to fix the reproducibility bug while preserving the old
selection, since the old selection was the bug. Runs without a cap are
byte-identical, verified against a pre-change binary.

Ports fulcrumgenomics/fgbio#1166. Once that merges, this also restores byte
parity with fgbio under a cap; a parity oracle test is deliberately deferred
until then rather than pinned to the 4.1.0 shuffle it would disagree with.

Closes #725
nh13 added a commit to fulcrumgenomics/fgumi that referenced this pull request Aug 9, 2026
A cap of zero empties every strand, so no molecule can produce a consensus, and
fgumi exited 0 having written an empty BAM. simplex and codec already validate
their equivalents.

Only a lower bound is checked, matching fgbio: duplex's --min-reads is a vector
with its own ordering rule, so the max >= min comparison the sibling tools make
does not apply, and a cap below it is already handled gracefully -- consensus_call
returns no consensus rather than panicking. A negative value is rejected by clap,
since the option is an Option<usize>, so validate() only has to cover zero; fgbio
needs the wider check because its option is an Option[Int].

Ports the second commit of fulcrumgenomics/fgbio#1166.

Refs #725
nh13 added a commit to fulcrumgenomics/fgumi that referenced this pull request Aug 10, 2026
* fix(consensus)!: make consensus read downsampling deterministic

Downsampling selected which reads to keep by shuffling with a seeded StdRng held
on the consensus caller. The RNG is stateful, so the surviving subset for a tag
family depended on how many families that caller instance had already processed.

fgumi did not have fgbio's thread-count bug: the unified pipeline builds a fresh
caller per batch at a fixed batch size, so --threads 1 and --threads 8 agreed.
But the no---threads path holds a single caller for the entire run, putting its
RNG at a different position for every family, so the two execution modes
disagreed under a cap. On a 77k-molecule library, duplex output at
--max-reads-per-strand 2 differed on 566 of 77,804 records (0.73%). Uncapped runs
always agreed.

Rank reads by a Murmur3 hash of the read name and keep the lowest-ranking ones.
The retained subset becomes a pure function of the family: independent of caller
history, execution mode, and the order reads arrive in. Because both ends of a
template share a read name they share a rank, so a template is retained or
discarded as a unit -- which also recovers reads the old shuffle lost by
splitting templates across ends (simplex --max-reads 2 on that library goes from
726,854 to 734,536 consensus reads, matching fgbio's count).

The rank is computed once per read and stored, never inside a sort key closure:
Rust's sort_by_key may evaluate its closure more than once per element, so
hashing there would re-hash on every comparison. Ranks compare as signed i32,
matching fgbio's Scala Int, and the sort is stable so tied ranks keep input
order. Selection now lives in one place, caller::select_lowest_ranking, shared
by the raw-record and codec paths.

With no RNG left on the caller there is no per-instance state for downsampling to
depend on, so the seed option and the rand dependency are removed and the
determinism becomes structural rather than incidental.

Which reads are retained differs from previous releases for any run that sets a
cap. There is no way to fix the reproducibility bug while preserving the old
selection, since the old selection was the bug. Runs without a cap are
byte-identical, verified against a pre-change binary.

Ports fulcrumgenomics/fgbio#1166. Once that merges, this also restores byte
parity with fgbio under a cap; a parity oracle test is deliberately deferred
until then rather than pinned to the 4.1.0 shuffle it would disagree with.

Closes #725

* fix(duplex): reject --max-reads-per-strand 0

A cap of zero empties every strand, so no molecule can produce a consensus, and
fgumi exited 0 having written an empty BAM. simplex and codec already validate
their equivalents.

Only a lower bound is checked, matching fgbio: duplex's --min-reads is a vector
with its own ordering rule, so the max >= min comparison the sibling tools make
does not apply, and a cap below it is already handled gracefully -- consensus_call
returns no consensus rather than panicking. A negative value is rejected by clap,
since the option is an Option<usize>, so validate() only has to cover zero; fgbio
needs the wider check because its option is an Option[Int].

Ports the second commit of fulcrumgenomics/fgbio#1166.

Refs #725
@nh13
nh13 force-pushed the nh_deterministic-consensus-downsampling branch from 6236030 to 98a7bcd Compare August 10, 2026 15:49
Comment thread src/main/scala/com/fulcrumgenomics/umi/VanillaUmiConsensusCaller.scala Outdated
nh13 added 2 commits August 10, 2026 14:28
When --max-reads/--max-reads-per-strand/--max-read-pairs is set and a tag
family exceeds the cap, VanillaUmiConsensusCaller downsampled the family with
a `Random(42)` held on the caller instance. Because the RNG is stateful, which
reads survived depended on how many families that instance had already
downsampled.

ConsensusCallingIterator gives each thread its own caller via emptyClone(), so
under `--threads > 1` the family-to-thread assignment (and hence each caller's
RNG position) varies from run to run. The same input therefore produced
different consensus bases, qualities and depths on every run. This affects
CallMolecularConsensusReads, CallDuplexConsensusReads and
CallCodecConsensusReads, all of which downsample through this code path.

Rank reads by a Murmur3 hash of their read name and keep the lowest ranking
ones instead. The retained subset is now a pure function of the family, so it
no longer depends on the caller's history, on the number of threads, or on the
order the reads arrive in. This mirrors the existing hash-based downsampling in
CollectDuplexSeqMetrics and SamOrder.Random. The rank is computed once per read
rather than handed to `sortBy`, which would re-hash on every comparison.

Because both ends of a template share a read name they now receive the same
rank, so where both ends survive the upstream alignment filtering a template is
retained or discarded on both ends together. Previously each end drew an
independent sample. A test pins this.

Note this changes which reads are retained relative to previous releases, and
so changes consensus output for runs that set a cap, single-threaded runs
included.
CallDuplexConsensusReads accepted `--max-reads-per-strand 0`, which downsampled
every strand to an empty set and then failed deep inside consensus calling with
a bare `IllegalArgumentException: requirement failed: Too few reads to create a
consensus.` rather than a user-facing validation error.

CallMolecularConsensusReads and CallCodecConsensusReads already validate their
equivalent options; this brings CallDuplexConsensusReads in line.
@nh13
nh13 force-pushed the nh_deterministic-consensus-downsampling branch from 98a7bcd to e748657 Compare August 10, 2026 21:29
@nh13
nh13 merged commit 39b6bbf into main Aug 10, 2026
13 checks passed
@nh13
nh13 deleted the nh_deterministic-consensus-downsampling branch August 10, 2026 21:38
nh13 added a commit to fulcrumgenomics/fgumi that referenced this pull request Aug 18, 2026
…785)

Reads discarded by the `--max-reads` consensus-downsampling cap were counted as used and never reached `--rejects`. Because `raw_reads_used = total_input_reads - filtered_reads` and only a recorded rejection increments `filtered_reads`, a family of 10 capped to 3 reported `raw_reads_used = 10` against a consensus of depth 3, and the 7 discarded reads were dropped silently.

Add a `Downsampled` rejection reason and record the discarded reads at each drop site so `raw_reads_used` excludes them:

- simplex (`VanillaUmiConsensusCaller::process_group`): `downsample_reads` now returns the discarded records alongside the survivors; they are counted and, when rejects tracking is enabled, routed to the `--rejects` output byte-for-byte in input order.
- codec (`CodecConsensusCaller`): the per-strand cap counts the discarded reads, mirroring how the codec caller already counts its other per-read rejections (e.g. MinorityAlignment). Per-read routing to `--rejects` for the codec waits on its reject-mask machinery (#751), so this only fixes the count there.

The duplex/codec per-strand `consensus_call` path is unaffected: since #727 it retains the full uncapped source reads (the cap shapes only the consensus bases/quals/depths), so no reads are dropped there.

`Downsampled` is fgumi-specific: fgbio has no equivalent metric yet (fulcrumgenomics/fgbio#1166, #1167), so `raw_reads_used` and the new `raw_reads_rejected_for_downsampled` row diverge from fgbio under a cap. The new row is emitted only when non-zero.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants