Skip to content

fix: dedup --no-umi OOM on production WES data - #231

Merged
nh13 merged 3 commits into
mainfrom
nh/dedup-fix-oom
Apr 4, 2026
Merged

nh13 merged 3 commits into
mainfrom
nh/dedup-fix-oom

Conversation

@nh13

@nh13 nh13 commented Apr 4, 2026

Copy link
Copy Markdown
Member

Summary

  • perf(dedup): Replace unbounded SegQueue<CollectedDedupMetrics> with Mutex<CollectedDedupMetrics> and merge metrics incrementally during the serialize step, eliminating unbounded memory growth from per-position-group metric accumulation
  • fix(pipeline): Remove the is_draining() bypass on Q5 memory backpressure in the Process step (bam, fastq, and base pipelines). Previously, once the Read step completed, the Process step could push unlimited data to Q5 with no memory governor, causing OOM on large inputs
  • test(dedup): Add regression test with 5,000 templates at the same position in --no-umi mode

Context

fgumi dedup --no-umi was OOM-killed on a production WES BAM (1000 Genomes HG00100, 219M records, 13 GB) on AWS c7g.4xlarge (32 GB RAM) at ~40M records. Two root causes:

  1. The SegQueue<CollectedDedupMetrics> accumulated one entry per position group for the entire pipeline run, only draining after completion — hundreds of MB on large inputs.

  2. The Q5 (processed queue) memory backpressure threshold (256 MB) was completely bypassed during draining mode. WES data has extreme depth pileups at capture targets creating massive position groups. With 128 queue slots and no memory limit, Q5 could accumulate many GB.

Validation

After fix, the same 219M-record BAM completes in 70 seconds with 11.8 GB peak RSS (8 threads), well within the 32 GB instance. Previously it was OOM-killed and never completed.

For comparison on the same data: samtools markdup completes in 170s / 3.4 GB RSS, GATK MarkDuplicates in 1687s.

Test plan

  • All 1886 tests pass (cargo nextest run)
  • cargo ci-fmt && cargo ci-lint clean
  • New regression test: 5,000 templates at same position with --no-umi
  • Production validation: 219M-record WES BAM completes in 70s / 11.8 GB peak RSS

@nh13
nh13 temporarily deployed to github-actions April 4, 2026 17:48 — with GitHub Actions Inactive
@codecov

codecov Bot commented Apr 4, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 99.20000% with 1 line in your changes missing coverage. Please review.
✅ Project coverage is 88.31%. Comparing base (0c6b0ae) to head (268d3fc).
⚠️ Report is 2 commits behind head on main.

Files with missing lines Patch % Lines
src/lib/unified_pipeline/base.rs 0.00% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #231      +/-   ##
==========================================
+ Coverage   88.27%   88.31%   +0.03%     
==========================================
  Files         113      113              
  Lines       53215    53322     +107     
==========================================
+ Hits        46977    47090     +113     
+ Misses       6238     6232       -6     

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@nh13

nh13 commented Apr 4, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 4, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai

coderabbitai Bot commented Apr 4, 2026 •

Copy link
Copy Markdown

Warning

Rate limit exceeded

@nh13 has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 5 minutes and 3 seconds before requesting another review.

Your organization is not enrolled in usage-based pricing. Contact your admin to enable usage-based pricing to continue reviews beyond the rate limit, or try again in 5 minutes and 3 seconds.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: bea5a027-1ba2-40c5-8540-ef4da4f8c009

📥 Commits

Reviewing files that changed from the base of the PR and between 5d2f432 and 268d3fc.

📒 Files selected for processing (4)
  • src/commands/dedup.rs
  • src/lib/unified_pipeline/bam.rs
  • src/lib/unified_pipeline/base.rs
  • src/lib/unified_pipeline/fastq.rs
📝 Walkthrough

Walkthrough

This change refactors dedup metrics collection from lock-free concurrent queues to mutex-guarded aggregation, consolidating per-group metrics directly during pipeline execution. It adds a regression test for large position-group deduplication without UMI tags. Across the unified pipeline (BAM, FASTQ, and base modules), the change removes drain-mode exceptions to memory backpressure enforcement, ensuring consistent backpressure behavior throughout pipeline execution. Documentation is updated to clarify that drain mode represents "input exhausted, completing remaining work" without special memory handling rules.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed Title clearly identifies the main fix: OOM issue in dedup with --no-umi flag on production data.
Description check ✅ Passed Description is directly related to the changeset, detailing the two root causes (SegQueue memory growth and Q5 backpressure bypass) and validation results.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch nh/dedup-fix-oom

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@nh13 nh13 added fgumi dedup bug Something isn't working labels Apr 4, 2026
@nh13

nh13 commented Apr 4, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 4, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Apr 4, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 4, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
src/lib/unified_pipeline/base.rs (1)

4723-4734: Tighten the ProcessPipelineState contract.

The trait still advertises drain-specific backpressure semantics, but shared_try_step_process() no longer consults is_draining(). That stale contract will mislead the next impl or refactor.

✏️ Suggested doc cleanup
-    /// Check if backpressure should be applied before processing new work.
-    /// Returns true if queue is full OR memory is high (unless draining).
-    /// Default: just checks queue capacity (backwards compatible).
+    /// Check if backpressure should be applied before processing new work.
+    /// Returns true if queue is full OR memory is high.
+    /// Default: just checks queue capacity for backwards compatibility.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/unified_pipeline/base.rs` around lines 4723 - 4734, The trait
ProcessPipelineState advertises drain-specific backpressure semantics that are
stale because shared_try_step_process() no longer consults is_draining(); update
the trait docs and methods to match actual behavior by removing or revising
drain-related wording and behavior: either delete the unused is_draining()
method from the ProcessPipelineState trait (and any impls), or change its
documentation and the should_apply_process_backpressure() docstring to state
that backpressure is determined only by process_output_is_full() (and
is_draining is not consulted), and update any default implementations/comments
in should_apply_process_backpressure() and is_draining() to avoid misleading
callers (referencing the ProcessPipelineState trait,
should_apply_process_backpressure(), is_draining(), and
shared_try_step_process()).
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@src/lib/unified_pipeline/base.rs`:
- Around line 4723-4734: The trait ProcessPipelineState advertises
drain-specific backpressure semantics that are stale because
shared_try_step_process() no longer consults is_draining(); update the trait
docs and methods to match actual behavior by removing or revising drain-related
wording and behavior: either delete the unused is_draining() method from the
ProcessPipelineState trait (and any impls), or change its documentation and the
should_apply_process_backpressure() docstring to state that backpressure is
determined only by process_output_is_full() (and is_draining is not consulted),
and update any default implementations/comments in
should_apply_process_backpressure() and is_draining() to avoid misleading
callers (referencing the ProcessPipelineState trait,
should_apply_process_backpressure(), is_draining(), and
shared_try_step_process()).

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: b8dd1e28-08be-46a5-9bf3-631393aff8e8

📥 Commits

Reviewing files that changed from the base of the PR and between 0c6b0ae and 5d2f432.

📒 Files selected for processing (4)
  • src/commands/dedup.rs
  • src/lib/unified_pipeline/bam.rs
  • src/lib/unified_pipeline/base.rs
  • src/lib/unified_pipeline/fastq.rs

@nh13
nh13 marked this pull request as ready for review April 4, 2026 20:15
nh13 added 3 commits April 4, 2026 13:18
…Queue

Replace SegQueue<CollectedDedupMetrics> with Mutex<CollectedDedupMetrics> and merge
metrics in-place during the serial Serialize step. This eliminates unbounded memory
growth from accumulating per-position-group metrics for the entire pipeline run.
Also removes unnecessary .clone() calls on dedup_metrics and family_sizes.
Remove the is_draining() bypass on Q5 memory backpressure in the Process
step. Previously, once the Read step completed, the Process step could push
unlimited data to Q5 with no memory governor, causing OOM on large inputs.
The slot-based is_full() check still guarantees forward progress during
draining, so removing the memory bypass is deadlock-safe.
Verify that dedup handles a position group with 5000 templates at the same
position without unbounded memory growth. Exercises the --no-umi code path
that was OOM-ing on production WES data.
@nh13
nh13 force-pushed the nh/dedup-fix-oom branch from 5d2f432 to 268d3fc Compare April 4, 2026 20:18
@nh13
nh13 temporarily deployed to github-actions April 4, 2026 20:18 — with GitHub Actions Inactive
@nh13
nh13 merged commit 0c5b318 into main Apr 4, 2026
7 checks passed
@nh13
nh13 deleted the nh/dedup-fix-oom branch April 4, 2026 20:21
@nh13 nh13 mentioned this pull request Apr 4, 2026

This branch was previously deployed

1 inactive deployment
github-actions — 268d3fc0 Deployed Apr 4, 2026 by nh13 via coverage #873
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working fgumi dedup

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant