The downsample command samples UMI families by the full MI tag value, including the /A / /B strand suffix that the paired (duplex) grouping strategy assigns. Because N/A and N/B are distinct MI values, they are treated as two separate families and each gets an independent keep/reject draw.
Impact
For a duplex molecule to survive as a duplex, both its strand draws must hit — probability fraction². At a realistic keep fraction (e.g. 0.05) that is ~0.25%, so nearly every duplex family loses one strand and collapses to simplex. Downsampling a duplex library therefore silently destroys its duplex structure: duplex-metrics on the result reports ~0 duplex families even when the input was duplex-rich. Simplex data is unaffected (one MI tag == one molecule).
Root cause
src/lib/commands/downsample.rs, FamilyIterator::next_family():
let record_mi = get_mi_tag(record)?;
if record_mi != mi { break; } // "1/A" != "1/B" -> two families
get_mi_tag returns the raw MI string ("1/A", "1/B"), and the family break compares it verbatim. The two strands are already adjacent in template-coordinate order, so no re-sort is needed — only the grouping key is wrong for duplex.
Proposed fix
- Add an opt-in by-molecule mode that strips the trailing
/[AB] before the family-break comparison, so both strands of a molecule group into one family and share a single keep/reject draw. This makes duplex downsampling correct while leaving current (per-tag) behavior the default for backward compatibility.
- Warn when MI values carry a
/A / /B suffix (i.e. paired/duplex grouping) but by-molecule mode is not enabled — this is almost always a mistake, and today it fails silently.
Repro
Group a duplex library (group --strategy paired), run downsample -f 0.05, then duplex-metrics on the output: duplex family count collapses to ~0 versus the input.
Authored by AI, reviewed by a Human.
The
downsamplecommand samples UMI families by the fullMItag value, including the/A//Bstrand suffix that thepaired(duplex) grouping strategy assigns. BecauseN/AandN/Bare distinct MI values, they are treated as two separate families and each gets an independent keep/reject draw.Impact
For a duplex molecule to survive as a duplex, both its strand draws must hit — probability
fraction². At a realistic keep fraction (e.g.0.05) that is ~0.25%, so nearly every duplex family loses one strand and collapses to simplex. Downsampling a duplex library therefore silently destroys its duplex structure:duplex-metricson the result reports ~0 duplex families even when the input was duplex-rich. Simplex data is unaffected (one MI tag == one molecule).Root cause
src/lib/commands/downsample.rs,FamilyIterator::next_family():get_mi_tagreturns the raw MI string ("1/A","1/B"), and the family break compares it verbatim. The two strands are already adjacent in template-coordinate order, so no re-sort is needed — only the grouping key is wrong for duplex.Proposed fix
/[AB]before the family-break comparison, so both strands of a molecule group into one family and share a single keep/reject draw. This makes duplex downsampling correct while leaving current (per-tag) behavior the default for backward compatibility./A//Bsuffix (i.e. paired/duplex grouping) but by-molecule mode is not enabled — this is almost always a mistake, and today it fails silently.Repro
Group a duplex library (
group --strategy paired), rundownsample -f 0.05, thenduplex-metricson the output: duplex family count collapses to ~0 versus the input.Authored by AI, reviewed by a Human.