Repository navigation
docs(dedup): stop calling the duplication ladder a saturation curve - #980
Conversation
The ladder counts templates in coordinate order, and a template's duplicate status is decided by the other templates at its own position. So the cumulative fraction after N templates is the duplicate rate of the genome covered so far, not the rate a library sequenced to N templates would show. A flattening ladder says nothing about library saturation. Reword the --duplication-ladder help, the DuplicationLadderMetrics docs and the internal comments to describe what the ladder measures.
|
Navigate logical layers of code changes, visualize relationships, and explore their blast radius. Note Reviews pausedUse the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Repository: fulcrumgenomics/fgumi/.coderabbit.yaml Review profile: ASSERTIVE Plan: Essentials Run ID: 📒 Files selected for processing (4)
Included review availability: 1 review is currently available. Your included PR review attempts over the past 7 days set your current allowance at 2 reviews per hour. WalkthroughDocumentation and comments now describe the duplication ladder as coordinate-ordered cumulative duplicate fractions that show variation across genomic regions, not sequencing-depth saturation. Runtime behavior is unchanged. ChangesDuplication ladder documentation
Estimated code review effort: 2 (Simple) | ~5 minutes Suggested labels: Merge Risk: ⚪ Minimal · up to This change clarifies the duplication ladder’s meaning without changing runtime behavior; no actionable merge risk is identified. 🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
Comment |
|
@coderabbitai pause |
✅ Action performedReviews paused. |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #980 +/- ##
==========================================
- Coverage 96.15% 96.13% -0.03%
==========================================
Files 293 293
Lines 147588 147588
==========================================
- Hits 141919 141879 -40
- Misses 5669 5709 +40 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
The
--duplication-ladderhelp text and theDuplicationLadderMetricsdocs called the ladder a saturation curve, which it isn't.dedupcounts templates in coordinate order, and whether a template is a duplicate depends only on the other templates at its position. So the cumulative duplicate fraction after N templates is the duplicate rate of the genome covered so far. It is not the rate you'd see if the library were sequenced to N templates. A saturation curve needs templates in random order, as downsampling gives. A flattening ladder therefore doesn't mean the library is saturated, and its bumps come from regions with different duplicate rates.This rewords the CLI help, the metric docs and the internal comments. No code changes.
Risk: Command output changes: none;
unsafechanges: none, and the CLAUDE.md allowlist is unchanged; memory bounds, queue capacity, and thread/backpressure policy changes: none. Fix: clarify that the duplication ladder shows duplicate rates across templates in coordinate order, not library saturation.Updated the
--duplication-ladderhelp text,DuplicationLadderMetricsdocumentation, and internal comments to use “duplication ladder” and describe its cumulative and window fractions. No runtime behavior changes are described.