feat(dedup): sample column and marginal duplication-ladder columns (#802) - #841
Conversation
#802) Toward Picard DuplicationMetrics / dupblaster --stats parity, building on the per-library metrics + saturation ladder in #799: - `--metrics` gains a leading `sample` column, resolved from the header's unique `@RG SM:` values (comma-joined) with an optional `--sample` override. - `--duplication-ladder` gains marginal per-window columns (`window_templates`, `window_duplicate_fraction`) alongside the cumulative ones — the marginal duplication rate isolates each depth band instead of averaging over all prior ones (dupblaster's complexity.rs view). The per-reason pair/orphan duplicate breakdown (the third dupblaster metrics gap) is a distinct algorithmic change and is tracked separately. Column order is pinned by a test and updated accordingly (sample first).
|
Note Reviews pausedUse the following commands to manage reviews:
Use the checkboxes below for quick actions:
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
Comment |
|
@coderabbitai pause |
✅ Action performedReviews paused. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #841 +/- ##
=======================================
Coverage 94.46% 94.47%
=======================================
Files 187 187
Lines 115301 115361 +60
=======================================
+ Hits 108921 108982 +61
+ Misses 6380 6379 -1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
Closes #802. Supersedes #803 (auto-closed when its base branch
786/nhomer/feat-dedup-complexity-metricswas deleted by the #799 squash-merge; GitHub cannot reopen a PR whose base branch was deleted). Same head branch and commit, now rebased ontomain.Toward Picard
DuplicationMetrics/ dupblaster--statsparity, two dupblaster-derived additions tofgumi dedup's QC output:samplecolumn--metricsgains a leadingsamplecolumn, resolved from the header's unique@RG SM:values (comma-joined), with an optional--sampleoverride. Constant across every row of a run; leads the schema, matching dupblaster/Picard.Marginal duplication-ladder columns
--duplication-laddergains per-window columns (window_templates,window_duplicate_fraction) alongside the cumulative ones. The marginal rate isolates each depth band instead of averaging over all prior templates — "usually the more legible view" of the saturation curve (dupblaster'scomplexity.rs). The windows partition the cumulative total exactly.Not included
The pair/orphan duplicate breakdown (Picard
READ_PAIRS_EXAMINED/READ_PAIR_DUPLICATES/UNPAIRED_*) is a distinct algorithmic change — it requires classifying each template as pair vs orphan during marking — so it's tracked as its own follow-up rather than bundled here.Verification
New tests:
sampleresolved from@RG SM, via--sampleoverride, and the empty case (no@RG SM:and no--sample→ empty column); ladder window columns pinned exactly (eachwindow_templatesequals the increment since the previous snapshot, marginal fractions pinned tonum/den, windows summing to the per-library total). The pinned column-order test is updated (samplefirst). The distinct-sample and override tests now pin exact row identities (["libA","libB","All Reads"]/["libA","All Reads"]), not just non-empty.