Repository navigation
fix(consensus): handle NaN likelihoods like fgbio - #1029
Conversation
ln_sum_exp_array starts its minimum search at f64::INFINITY with index 0 and skips NaN lanes. When every lane was NaN or +inf, the sum was seeded with the +inf start value instead of lane 0, so [NaN] and [NaN, +inf] returned +inf rather than NaN. The sum is now seeded from the selected lane itself, so any NaN lane yields NaN. fgbio's LogProbability.or (NumericTypes.scala:167) throws for those inputs (MathUtilTest.scala:57, :93); fgumi returns NaN. This alone changes no consensus output. ln_prob_to_phred now maps NaN to MIN_PHRED (Q2). NaN passed through clamp and the u8 cast as Q0, below the documented [2, 93] range; fgbio computes PhredScore.cap(fromLogProbability(NaN)) = Q2 (ConsensusCaller.scala:173). The NaN arises when an observation's correct term is ln(0) = -inf, which takes a Q0 on either side of the two-trial error: --error-rate-post-umi 0 (simplex only) or a Q0 input base with --min-input-base-quality 0. The Kahan update turns that lane NaN on the next add, as fgbio's kahanAdd (ConsensusCaller.scala:128-133) does. Observations A, C, G leave lanes [NaN, NaN, -inf, finite] and call (T, Q2), as in fgbio. With all four bases observed, three lanes are NaN and the lane of the last-added base is -inf. FgbioCompat selected that lane, so the called base depended on read order; fgbio throws there. A -inf maximum is now a no-call (N, Q2) under FgbioCompat too, as under UlpRelative, in the selection shared by call_full and the multi-base fast path. fgbio_unique_max_index stays an exact port. Over all 38,812 pileups of length 1-5 (pre-UMI 45; post-UMI 40 with Q0/Q30 bases, post-UMI 0 with Q30 bases), every case fgbio calls is unchanged and matches it; the 1,008 where fgbio throws are now N/Q2. Changes output for: positions with a -inf correct term on every observation, reachable only with --error-rate-post-umi 0 or with --min-input-base-quality 0 and Q0 input bases. Three distinct bases: the single-strand quality goes from Q0 to Q2, so simplex with --min-consensus-base-quality 1 or 2 emits the base instead of N, duplex emits (T, Q4) instead of N when both strands are degenerate, and codec emits (T, Q4) instead of (T, Q0) in the duplex region and N instead of (T, Q0) in a single-strand region. All four bases: (N, Q2) instead of the last-added base at Q0, which simplex at its default --min-consensus-base-quality 2 and duplex already masked to N but codec could emit. Adds tests porting the fgbio NaN rows, pinning ln_prob_to_phred(NaN) at Q2, the shared Kahan -inf -> NaN behavior, the three-base pileup over both Q0 routes, three add orders and both tie rules, the four-base no-call over five read orders, and the degenerate position through the simplex, duplex and codec callers.
|
Note Reviews pausedUse the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configuration
📒 Files selected for processing (5)
Included review availability: This review used your included allowance. 3 included reviews remain after this review. Your included PR review attempts over the past 7 days set your current allowance at 4 reviews per hour. WalkthroughDegenerate Q0 likelihoods could produce order-dependent base calls or incorrect qualities. The changes preserve NaN in likelihood sums, map NaN to MIN_PHRED, and reject negative-infinity selections in FgbioCompat mode. ChangesDegenerate likelihood handling
Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~20 minutes Change: Bug fix Merge Risk: ⚪ Minimal · up to The degenerate-likelihood changes are ready to merge after normal checks; no actionable issue remains identified. 🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
Comment |
|
@coderabbitai pause |
✅ Action performedReviews paused. |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #1029 +/- ##
==========================================
- Coverage 96.47% 96.47% -0.01%
==========================================
Files 299 299
Lines 152214 152302 +88
==========================================
+ Hits 146854 146932 +78
- Misses 5360 5370 +10 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
Summary
ln_sum_exp_arrayfinds the minimum lane by starting atf64::INFINITYwith index 0 and skippingNaNlanes. When no lane compared below that start value (every laneNaNor+inf), the sum was seeded with the+infstart value instead of lane 0, so aNaNin lane 0 was never folded in.[NaN]and[NaN, +inf]returned+inf, a plausible-looking log-probability, instead ofNaN.The sum is now seeded from the selected lane itself (
values[min_index]), so anyNaNlane yieldsNaN. This is a one-line change with no extra pass over the lanes on the consensus hot path. The rustdoc now states the contract:-inffor an empty or all--infarray,NaNwhenever any lane isNaN.This PR also fixes a sibling bug in
ln_prob_to_phred: aNaNerror probability passed throughclampand theas u8cast turned it into Q0, below the documented[2, 93]range. It now returnsMIN_PHRED(Q2), as fgbio does.With that guard, a position where all four bases carry a
-infcorrect term (three lanesNaN, the last-added base's lane-inf) would be called under the defaultFgbioCompatrule as whichever base was added last, at Q2, so the base depended on read order. fgbio throws there.unique_max_index_withnow no-calls a-infmaximum underFgbioCompat, asUlpRelativealready did, which covers bothcall_fulland the multi-base fast path;fgbio_unique_max_indexstays an exact port.Changes output for: positions with a
-infcorrect term on every observation, reachable only with--error-rate-post-umi 0or with--min-input-base-quality 0and Q0 input bases.--min-consensus-base-quality1 or 2 emits the base instead of N; duplex emits (T, Q4) instead of N when both strands are degenerate; codec emits (T, Q4) instead of (T, Q0) in the duplex region and N instead of (T, Q0) in a single-strand region. All match fgbio.--min-consensus-base-quality 2and duplex already masked to N but codec could emit. fgbio throws here.Default settings (
--min-input-base-quality 10, nonzero error rates) never produce a-infterm, so default output is unchanged.When the NaN is reachable
The adjusted probability of a correct base is
ln(0) = -infwhenever either side of the two-trial error is Q0 (error probability 1;ln_error_prob_two_trialsreturns the dominantln(1) = 0, as fgbio'sprobabilityOfErrorTwoTrialsdoes):--error-rate-post-umi 0(simplex only; duplex and codec reject it, fgbio accepts it), or a Q0 input base kept by--min-input-base-quality 0(simplex, duplex, codec). The Kahan update inConsensusBaseBuilder::addthen sets that lane's compensation to(-inf - sum) - -inf = NaN, and the next add turns the lane itself intoNaN. fgbio'skahanAdd(ConsensusCaller.scala:128-133, added in fgbio #1120 / e2ccac9) does the same arithmetic, so this lane behavior is shared and is pinned here rather than changed.Example: pre-UMI 45, observations A, C, G, with either post-UMI 0 and Q30 bases or post-UMI 40 and Q0 bases. The lanes end as
[NaN, NaN, -inf, finite],Tis the unique non-NaNmaximum, and the final error probability isNaN.The
ln_sum_exp_arraychange alone alters no consensus output: real likelihood lanes are never+inf, four all-NaNlanes already summed toNaN, andConsensusBaseBuilder::call_fullno-calls an all-NaNposition inunique_max_index_withbefore the sum is used. The output changes are the Q0 -> Q2 quality above, its downstream effects, and the four-base no-call.fgbio parity (e51a661)
NumericTypes.scala:167(LogProbability.or(Array)): returns-infwhen every lane is-inf(including empty), otherwise seeds the sum withMathUtil.minWithIndexand folds the other lanes in index order. fgumi matches the-infguard and the fold.MathUtil.scala:75-100(minWithIndex): skipsNaNand-inflanes and throwsNoSuchElementExceptionwhen none remain. For those inputs (every laneNaNor-inf, at least oneNaN) fgumi returnsNaNinstead of throwing. For any other array with aNaNlane, fgbio'soralso returnsNaN, so the two agree.ConsensusCaller.scala:173: the quality isPhredScore.cap(PhredScore.fromLogProbability(p)). ForNaN,fromLogProbability(NumericTypes.scala:83) returnsMath.floor(NaN).toByte, which is 0, andcap(NumericTypes.scala:69) raises it toMinValue= 2. fgumi'sln_prob_to_phredfolds the cap in, so it now returns Q2 forNaN.ConsensusBaseBuilder: all 38,812 pileups of length 1-5 at pre-UMI 45 (post-UMI 40 with Q0/Q30 bases; post-UMI 0 with Q30 bases). All 37,804 cases fgbio calls match exactly; the 1,008 where fgbio throwsNoSuchElementExceptionare N/Q2 in fgumi.Tests
test_ln_sum_exp_array_with_nan_and_no_finite_lane_is_nan(rstest):[NaN]ports the row atMathUtilTest.scala:93("MathUtil.maxWithIndex should throw exceptions on invalid inputs",:91, whose row callsminWithIndex).[-inf, NaN]ports the row atMathUtilTest.scala:57("MathUtil.minWithIndex should throw exceptions on invalid inputs",:54).[NaN, -inf]and[NaN, NaN]cover the lane-order variants.[NaN, +inf]is fgumi's own: fgbio seeds with the+inflane andor(+inf, NaN)givesNaN.[NaN]and[NaN, +inf]failed before the fix; the other rows are regression guards. The empty and[-inf]rows (MathUtilTest.scala:55-56) never reachminWithIndexthroughor;test_ln_sum_exp_array_all_neg_inf_is_neg_infalready pins them.test_ln_prob_to_phred_nan_is_min_phred(rstest,NaNand-NaN): asserts exactly Q2.test_kahan_neg_inf_term_turns_lane_nan_on_next_add: pins the shared Kahan-inf->NaNlane behavior with post-UMI Q0 (lane A is-infafter one add andNaNafter the next).test_degenerate_pileup_calls_min_phred_like_fgbio(rstest over both Q0 routes x three add orders x both tie rules, 12 cases): assertscall()andcall_full()both return exactly(T, 2). Every case fails without the guard with(T, 0). fgbio has no test for this case; the expected value was derived from the fgbio e51a661 source above.test_all_four_bases_with_neg_inf_correct_term_no_calls(rstest over both Q0 routes x five read orders x both tie rules, 20 cases): pins that the last-added base's lane is-infand the restNaN, and thatcall()andcall_full()both return (N, Q2). The 10FgbioCompatcases fail without the-infno-call.test_degenerate_position_emits_min_phred_base(vanilla caller, rstest over both Q0 routes): with--min-consensus-base-quality 2, the degenerate position is emitted asTat Q2; without the guard it was masked to N.test_degenerate_position_on_both_strands_calls_duplex_t_at_q4(duplex caller): AB and BA families with A/C/G at Q0 at one position give (T, Q4), derived from fgbio (DuplexConsensusCaller.scala:142,:424,:431;VanillaUmiConsensusCaller.scala:354). Without the guard it was (N, Q2).test_degenerate_position_matches_fgbio(codec caller, rstest): (T, Q4) in the duplex region and (N, Q2) in a single-strand region, derived from fgbio (CodecConsensusCaller.scala:251-252pads with Q0, thenDuplexConsensusCaller.scala:424-431). Without the guard both were (T, Q0).Mutation-checked: seeding from the start value again fails
[NaN]and[NaN, +inf]. An all-NaNguard in its place (no other change) still fails[NaN, +inf]. Removing theln_prob_to_phredNaNguard fails 19 tests (bothNaNrows, all 12 degenerate-pileup cases, both vanilla cases, the duplex case, both codec cases). Removing the-infno-call fails all 10FgbioCompatfour-base cases.Performance
The
NaNguard is one never-taken, perfectly predictable branch perln_prob_to_phredcall (noNaNarises at default settings), next to a division and afloor; the seed change replaces one load with another. Neither adds a pass over the lanes.Checks
cargo ci-fmt,cargo ci-lint,cargo ci-doc,cargo ci-tag-literals,cargo nextest run --workspace(10323 passed, 31 skipped).Risk: Consensus output changes for degenerate Q0 cases, pinned by regression tests; no
unsafechanges, so theCLAUDE.mdallowlist needs no update; no memory-bound, queue-capacity, or thread/backpressure policy changes.NaN likelihoods now propagate through
ln_sum_exp_array, andln_prob_to_phredmaps NaN to Q2. InFgbioCompat, an all--infmaximum now produces a no-call instead of an order-dependent base. These changes affect consensus output in edge cases, including positions reached with zero post-UMI error rate or Q0 minimum input quality. Tests pin the base and quality expectations in the vanilla, duplex, and CODEC callers.Reported checks include
cargo ci-fmt,cargo ci-lint,cargo ci-doc,cargo ci-tag-literals, andcargo nextest run --workspace(10,323 passed; 31 skipped). These results are author-reported; this inspection did not run tests.