Skip to content

fix(pipeline): stream rejects to disk in simplex/duplex/codec/correct - #293

Merged
nh13 merged 4 commits into
mainfrom
nh/pipeline-rejects-streaming
Apr 20, 2026
Merged

nh13 merged 4 commits into
mainfrom
nh/pipeline-rejects-streaming

Conversation

@nh13

@nh13 nh13 commented Apr 17, 2026

Copy link
Copy Markdown
Member

Summary

Stacked on #290. Removes the per-thread `Vec<Vec>` reject-buffering pattern from the four commands that keep it (after #290 removed the metric-buffering pattern from all seven). Rejected raw BAM records are now streamed directly to the rejects BAM during `serialize_fn` via a shared BGZF writer guarded by a `parking_lot::Mutex`. The mutex only serializes the byte append; BGZF compression runs on the writer's own thread pool.

Commits (one per command):

  1. `fix(simplex)` — open rejects writer up front; write bytes from serialize; finalize writer post-pipeline. Drops `rejects: Vec<Vec>` from the per-thread accumulator.
  2. `fix(duplex)` — mirrors simplex.
  3. `fix(codec)` — mirrors simplex/duplex.
  4. `fix(correct)` — mirrors simplex/duplex/codec but on `rejected_raw_records`. Rejects use the input header (correct preserves alignment context).
  5. `docs(changelog)` — note.

Why

With #287/#290 landed, metric memory in the consensus commands and `correct` is capped at O(threads × distinct keys). The remaining unbounded growth vector is reject buffering: every rejected raw BAM record was held in per-thread Vecs for the life of the run, then written after the pipeline completed. On protocols with high rejection rates (e.g. duplex calls that fail `/A`/`/B` partnering, correct runs where most UMIs don't match the whitelist), this can retain multiple GB of raw BAM bytes.

Streaming during the pipeline caps reject memory at roughly the BGZF block size per worker, independent of reject count.

Why not the pipeline's `with_secondary` API

`run_bam_pipeline_from_reader_with_secondary` exists and `filter` uses it, but it writes the secondary output using `output_header`. simplex/duplex/codec build an unmapped-consensus output header that differs from the input header — rejects need the input header since they're original mapped reads. Extending the pipeline API to accept a separate secondary header is a larger surgery; the shared-Mutex pattern here is surgical and keeps the pipeline internals untouched. Worth revisiting once all four commands are in the same shape and a generalization has a clear target.

Test plan

  • `cargo ci-fmt`, `cargo ci-lint`, `cargo ci-test` — 2545 tests pass.
  • Each command's integration tests (simplex/duplex/codec/correct) pass, including the `*_rejects_has_bgzf_eof` BGZF-EOF sanity tests.
  • A/B memory comparison on a run with a high-rejection-rate input — confirm flat RSS across reject count.
  • Byte-identity check on rejects BAM output vs. the old buffered path (ordering may differ since writes are no longer post-sorted; behavior already was not order-guaranteed).

Note

The final changelog commit is unsigned (three 1Password biometric signing attempts failed locally). It should be re-signed before merge.

@nh13
nh13 temporarily deployed to github-actions April 17, 2026 21:19 — with GitHub Actions Inactive
@codecov

codecov Bot commented Apr 17, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 74.35897% with 50 lines in your changes missing coverage. Please review.
✅ Project coverage is 90.26%. Comparing base (54147fd) to head (b7dceb0).
⚠️ Report is 1 commits behind head on main.

Files with missing lines Patch % Lines
src/lib/commands/simplex.rs 74.57% 15 Missing ⚠️
src/lib/commands/codec.rs 73.91% 12 Missing ⚠️
src/lib/commands/correct.rs 73.33% 12 Missing ⚠️
src/lib/commands/duplex.rs 75.55% 11 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #293      +/-   ##
==========================================
+ Coverage   90.19%   90.26%   +0.06%     
==========================================
  Files         124      124              
  Lines       60598    60618      +20     
==========================================
+ Hits        54657    54716      +59     
+ Misses       5941     5902      -39     

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@nh13

nh13 commented Apr 18, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 18, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai

coderabbitai Bot commented Apr 18, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 8ac4369b-8e66-45a5-a123-4be257f218f6

📥 Commits

Reviewing files that changed from the base of the PR and between 8ad2f8f and b7dceb0.

📒 Files selected for processing (5)
  • src/lib/commands/correct.rs
  • tests/integration/helpers/assertions.rs
  • tests/integration/test_bgzf_eof.rs
  • tests/integration/test_correct_command.rs
  • tests/integration/test_duplex_command.rs
🚧 Files skipped from review as they are similar to previous changes (5)
  • tests/integration/helpers/assertions.rs
  • tests/integration/test_bgzf_eof.rs
  • tests/integration/test_duplex_command.rs
  • tests/integration/test_correct_command.rs
  • src/lib/commands/correct.rs

📝 Walkthrough

Walkthrough

Rejects are no longer buffered per-batch; when --rejects is enabled a single shared RawBamWriter is created before the pipeline and wrapped in Arc<Mutex<Option<...>>>. Worker threads stream rejected raw records directly to that writer under the mutex. Rejects headers are rewritten with header_as_unsorted (SO:unsorted, remove GO/SS). Pipelines always attempt to finalize the writer and surface any finish() errors alongside pipeline errors.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed Title clearly and concisely summarizes the main change: streaming rejects to disk across four pipeline commands.
Description check ✅ Passed Description comprehensively covers the motivation, implementation approach, trade-offs, and test status relevant to the changeset.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch nh/pipeline-rejects-streaming

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@nh13 nh13 added the enhancement New feature or request label Apr 18, 2026
@nh13

nh13 commented Apr 18, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 18, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13
nh13 force-pushed the nh/pipeline-metrics-per-thread branch from 3072739 to e7ed89c Compare April 18, 2026 02:27
@nh13

nh13 commented Apr 18, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 18, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13
nh13 force-pushed the nh/pipeline-metrics-per-thread branch from e7ed89c to cadc96c Compare April 18, 2026 03:09
@nh13

nh13 commented Apr 18, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 18, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Apr 18, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 18, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (2)
src/lib/commands/codec.rs (1)

566-669: ⚠️ Potential issue | 🟠 Major

Finalize the rejects writer even when the pipeline fails.

Line 654 returns before Lines 667-669 on any pipeline error, so a partially written rejects BAM can miss finalization/EOF. Capture the pipeline result, finalize rejects, then propagate the pipeline error.

Proposed fix
-        let groups_processed = run_bam_pipeline_from_reader(
+        let pipeline_result = run_bam_pipeline_from_reader(
             pipeline_config,
             reader,
             input_header,
             &self.io.output,
@@
         )
-        .map_err(|e| anyhow::anyhow!("Pipeline error: {e}"))?;
+        .map_err(|e| anyhow::anyhow!("Pipeline error: {e}"));
 
-        // ========== Post-pipeline: Aggregate metrics and close rejects writer ==========
+        let finish_rejects_result: Result<()> = (|| {
+            if let Some(rw_arc) = rejects_writer {
+                if let Some(writer) = rw_arc.lock().take() {
+                    writer.finish().context("Failed to finish rejects file")?;
+                    info!("Rejected reads streamed to rejects file during processing");
+                }
+            }
+            Ok(())
+        })();
+
+        let groups_processed = pipeline_result?;
+        finish_rejects_result?;
+
+        // ========== Post-pipeline: Aggregate metrics ==========
         let mut total_groups = 0u64;
         let mut merged_stats = CodecConsensusStats::default();
@@
-        // Rejects were streamed during serialize; now finalize the writer.
-        if let Some(rw_arc) = rejects_writer {
-            if let Some(writer) = rw_arc.lock().take() {
-                writer.finish().context("Failed to finish rejects file")?;
-                info!("Rejected reads streamed to rejects file during processing");
-            }
-        }
-
         // Log statistics
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/commands/codec.rs` around lines 566 - 669, The pipeline call
currently uses .map_err(...) which returns early on error before finalizing the
rejects writer (rejects_writer / rw_arc), so a partially-written rejects BAM may
never be finished; change to capture the Result from
run_bam_pipeline_from_reader into a variable (e.g., let pipeline_result =
run_bam_pipeline_from_reader(...);), then after aggregating metrics always
finalize the rejects writer (if let Some(rw_arc) = rejects_writer { if let
Some(writer) = rw_arc.lock().take() { writer.finish().context("Failed to finish
rejects file")?; } }), and finally propagate the pipeline_result error (e.g.,
pipeline_result.map_err(|e| anyhow::anyhow!("Pipeline error: {e}"))?), ensuring
rejects_writer finalization runs regardless of pipeline success or failure.
src/lib/commands/duplex.rs (1)

702-740: ⚠️ Potential issue | 🟠 Major

Rejects are still buffered per processed batch.

serialize_fn streams rejects only after process_fn has built all_rejects, so a high-rejection MI batch can still retain all rejected raw records until serialization. Consider flushing rejects per MI group through the shared writer or otherwise keeping DuplexProcessedBatch free of reject payloads.

Also applies to: 755-765

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/commands/duplex.rs` around lines 702 - 740, The batch loop
accumulates per-MI rejects into all_rejects and returns them in
DuplexProcessedBatch, causing high-memory buffering; instead, when track_rejects
is enabled, immediately drain caller.take_rejected_reads() for each RawMiGroup
and write/stream them via the shared reject writer (the same writer used by
serialize_fn) before moving to the next MI, and remove the reject payload from
DuplexProcessedBatch (or keep only lightweight metadata). Update process_fn to
call the shared writer inside the loop (after caller.take_rejected_reads()) and
stop extending all_rejects so DuplexProcessedBatch no longer holds raw reject
payloads; keep merging stats (batch_stats, batch_overlapping) as before.
🧹 Nitpick comments (2)
src/lib/commands/codec.rs (1)

506-508: Update the stale rejects comment.

The comment still says rejects buffering is preserved and tracked as follow-up, but this implementation now streams rejects during serialization.

Proposed fix
-        // Per-thread metrics accumulator: bounded metric memory, no unbounded
-        // queue. Rejects buffering semantics are preserved (see follow-up).
+        // Per-thread metrics accumulator: bounded metric memory, no unbounded
+        // queue. Rejects are streamed separately during serialization.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/commands/codec.rs` around lines 506 - 508, The comment above the
PerThreadAccumulator::<CollectedCodecMetrics>::new(num_threads) instantiation is
outdated: it claims "Rejects buffering semantics are preserved (see follow-up)"
but the current implementation streams rejects during serialization; update the
comment to accurately reflect that rejects are streamed during serialization
(not buffered) and that the accumulator enforces bounded per-thread metric
memory. Locate the comment near the collected_metrics /
PerThreadAccumulator::<CollectedCodecMetrics>::new(...) and change the wording
to state that rejects are streamed during serialization and that the accumulator
enforces bounded metric memory with no unbounded queue.
src/lib/commands/simplex.rs (1)

499-500: Update the stale rejects comment.

Rejects are no longer “preserved” for a follow-up; they are streamed during serialization now.

Proposed comment update
-        // Per-thread metrics accumulator: bounded metric memory, no unbounded
-        // queue. Rejects buffering semantics are preserved (see follow-up).
+        // Per-thread metrics accumulator: bounded metric memory, no unbounded
+        // queue. Rejects are streamed during serialize, not stored here.
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/commands/simplex.rs` around lines 499 - 500, Update the stale inline
comment next to the Per-thread metrics accumulator so it no longer says "Rejects
buffering semantics are preserved (see follow-up)"; change it to state that
rejects are streamed during serialization (e.g., replace that clause with
"Rejects are streamed during serialization.") so the comment accurately reflects
current behavior in the code surrounding the Per-thread metrics accumulator.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@CHANGELOG.md`:
- Around line 10-11: Remove the stale clause "Rejects buffering in the consensus
commands is unchanged and tracked separately." from the release-note bullet that
begins "Replace unbounded `SegQueue<CollectedXxxMetrics>`..." and either delete
the clause entirely or replace it with a corrected statement noting that
rejected records are now streamed to disk (as documented in the next bullet) so
reject buffering is no longer unchanged; ensure the two bullets are consistent
so readers aren't confused about reject-buffering behavior.

In `@src/lib/per_thread_accumulator.rs`:
- Around line 81-91: into_slots currently mutates shared slots via
std::mem::take when Arc::try_unwrap fails; change its behavior to avoid mutating
state held by other Arc owners by altering the signature to return
Result<Vec<A>, Arc<Self>> (keep the receiver as Arc<Self>), use
Arc::try_unwrap(self) => Ok(inner) branch to consume and return the Vec as
before, and in the Err(arc) branch return Err(arc) without performing
std::mem::take or locking/mutating slots — callers can then call the
non-consuming slots() accessor to inspect or reduce values safely; refer to
into_slots, Arc::try_unwrap, and slots() when making the change.

---

Outside diff comments:
In `@src/lib/commands/codec.rs`:
- Around line 566-669: The pipeline call currently uses .map_err(...) which
returns early on error before finalizing the rejects writer (rejects_writer /
rw_arc), so a partially-written rejects BAM may never be finished; change to
capture the Result from run_bam_pipeline_from_reader into a variable (e.g., let
pipeline_result = run_bam_pipeline_from_reader(...);), then after aggregating
metrics always finalize the rejects writer (if let Some(rw_arc) = rejects_writer
{ if let Some(writer) = rw_arc.lock().take() { writer.finish().context("Failed
to finish rejects file")?; } }), and finally propagate the pipeline_result error
(e.g., pipeline_result.map_err(|e| anyhow::anyhow!("Pipeline error: {e}"))?),
ensuring rejects_writer finalization runs regardless of pipeline success or
failure.

In `@src/lib/commands/duplex.rs`:
- Around line 702-740: The batch loop accumulates per-MI rejects into
all_rejects and returns them in DuplexProcessedBatch, causing high-memory
buffering; instead, when track_rejects is enabled, immediately drain
caller.take_rejected_reads() for each RawMiGroup and write/stream them via the
shared reject writer (the same writer used by serialize_fn) before moving to the
next MI, and remove the reject payload from DuplexProcessedBatch (or keep only
lightweight metadata). Update process_fn to call the shared writer inside the
loop (after caller.take_rejected_reads()) and stop extending all_rejects so
DuplexProcessedBatch no longer holds raw reject payloads; keep merging stats
(batch_stats, batch_overlapping) as before.

---

Nitpick comments:
In `@src/lib/commands/codec.rs`:
- Around line 506-508: The comment above the
PerThreadAccumulator::<CollectedCodecMetrics>::new(num_threads) instantiation is
outdated: it claims "Rejects buffering semantics are preserved (see follow-up)"
but the current implementation streams rejects during serialization; update the
comment to accurately reflect that rejects are streamed during serialization
(not buffered) and that the accumulator enforces bounded per-thread metric
memory. Locate the comment near the collected_metrics /
PerThreadAccumulator::<CollectedCodecMetrics>::new(...) and change the wording
to state that rejects are streamed during serialization and that the accumulator
enforces bounded metric memory with no unbounded queue.

In `@src/lib/commands/simplex.rs`:
- Around line 499-500: Update the stale inline comment next to the Per-thread
metrics accumulator so it no longer says "Rejects buffering semantics are
preserved (see follow-up)"; change it to state that rejects are streamed during
serialization (e.g., replace that clause with "Rejects are streamed during
serialization.") so the comment accurately reflects current behavior in the code
surrounding the Per-thread metrics accumulator.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 5504d2e3-1ad7-46e8-ac38-7f82aefbc72c

📥 Commits

Reviewing files that changed from the base of the PR and between cadc96c and 4afa6f6.

📒 Files selected for processing (10)
  • CHANGELOG.md
  • src/lib/commands/clip.rs
  • src/lib/commands/codec.rs
  • src/lib/commands/correct.rs
  • src/lib/commands/duplex.rs
  • src/lib/commands/filter.rs
  • src/lib/commands/group.rs
  • src/lib/commands/simplex.rs
  • src/lib/mod.rs
  • src/lib/per_thread_accumulator.rs

Comment thread CHANGELOG.md Outdated
Comment thread src/lib/per_thread_accumulator.rs
@nh13
nh13 force-pushed the nh/pipeline-metrics-per-thread branch from cadc96c to ee7ac26 Compare April 18, 2026 07:40
Base automatically changed from nh/pipeline-metrics-per-thread to main April 18, 2026 07:42
@nh13
nh13 force-pushed the nh/pipeline-rejects-streaming branch from 4afa6f6 to d907ac6 Compare April 18, 2026 08:00
@nh13
nh13 temporarily deployed to github-actions April 18, 2026 08:00 — with GitHub Actions Inactive
@nh13
nh13 marked this pull request as ready for review April 18, 2026 08:00
@nh13

nh13 commented Apr 18, 2026

Copy link
Copy Markdown
Member Author

Addressed review feedback and rebased:

  • Dropped 9 commits already on main via fix(group): cap metric memory with per-thread accumulators (#285) #287/refactor(pipeline): PerThreadAccumulator for all per-batch metric collection #290; dropped both docs(changelog) commits since release-plz + git-cliff owns CHANGELOG.md (per release-plz.toml).
  • Fixed stale "Rejects buffering semantics are preserved (see follow-up)" comments in simplex/duplex/codec.
  • Finalize the rejects writer even when the pipeline errors out, so a partial rejects BAM still gets a valid BGZF EOF (simplex/duplex/codec/correct).
  • Moved reject streaming from serialize_fn into process_fn, writing per-MI-group (per-template for correct) directly through the shared BGZF writer. ProcessedBatch no longer carries any reject payload, so peak reject memory is now the BGZF block size rather than O(batch_size × rejects_per_group).

Branch is rebased onto current main, 4 commits (one per pipeline command).

The per_thread_accumulator.rs thread is now obsolete: after rebase, that file isn't touched by this PR (it belongs to #290).

@nh13

nh13 commented Apr 18, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 18, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 6

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@src/lib/commands/correct.rs`:
- Around line 859-875: The current flush_rejects closure (used in process_fn and
writing via rejects_writer_for_process) writes rejects under a mutex but outside
the pipeline's reorder/serialize stage, so reject BAMs can be out-of-order while
still carrying a sorted header; fix by either routing reject records through the
same ordered pipeline/serialize stage (i.e., stop calling flush_rejects
pre-serialize and instead enqueue or emit rejects alongside normal records so
the reorder/serialize step handles them) or explicitly mark the rejects BAM
header as unsorted/unknown when creating/initializing rejects_writer_for_process
(update the header HD SO tag to "unsorted" or "unknown") so downstream tools
won’t assume sort order.
- Around line 1032-1042: The current code uses pipeline_result? before checking
rejects_finish_result which hides a rejects.finish() error; change the control
flow to first capture pipeline_result without early-return (e.g., let
pipeline_res = pipeline_result;), then if let Some(result) =
rejects_finish_result { result? } to propagate any finish() error first, and
only after successful rejects finalization return or ? the pipeline result
(e.g., match pipeline_res { Ok(records_written) => records_written, Err(e) =>
return Err(e) }). Ensure you reference and use the existing symbols
rejects_writer, rejects_finish_result, pipeline_result, and writer.finish() so
that a rejects.finish() failure is not dropped when the pipeline also fails.

In `@src/lib/commands/duplex.rs`:
- Around line 703-719: The current flush_rejects closure writes worker-produced
reject records under a mutex but not in input order, which can falsify sort
metadata in the rejects BAM; fix by either (A) routing rejects through an
ordered stage: have worker code send (sequence_id, Vec<Vec<u8>>) to a single
ordered_consumer on the main thread which calls w.write_raw_record in sequence
(introduce a channel and an ordered drain in place of direct flush_rejects
usage), or (B) make the rejects header explicitly unordered before any writes by
mutating input_header (remove or set the HD/SO tag to "unsorted") so the
contract matches nondeterministic ordering; locate changes around flush_rejects,
rejects_writer_for_process, w.write_raw_record, and input_header.
- Around line 786-797: The pipeline error can short-circuit before we observe
rejects_writer.finish() failures: ensure the rejects writer is finalized and any
finish() error is surfaced even when pipeline_result is Err. Change control flow
around rejects_finish_result and pipeline_result so you first compute
rejects_finish_result (already done), then handle both results: if
rejects_finish_result is Some(Err) return that error (or combine it), otherwise
if pipeline_result is Err return the pipeline error; only proceed to set
groups_processed when pipeline_result is Ok. Reference symbols: rejects_writer,
rejects_finish_result, pipeline_result, groups_processed, and the
writer.finish() call — ensure rejects_writer.lock().take() ->
writer.finish().context(...) is awaited/checked before returning
pipeline_result.map_err(...)? so finish() failures are not masked.

In `@src/lib/commands/simplex.rs`:
- Around line 599-615: The current flush_rejects closure writes per-worker
rejects under a mutex but does so in arrival order, allowing output order to
diverge from input order when process_fn runs concurrently; change to preserve
input ordering by tagging each rejected record with its input sequence/index and
funneling all rejects into a single ordered writer that emits records in
increasing sequence (e.g., have workers push (seq, Vec<u8>) into a shared
ordered queue or BTreeMap and have rejects_writer_for_process take/flush entries
only when the next expected sequence is present), or alternatively buffer
per-worker batches and merge/flush them by sequence before calling
write_raw_record; apply the same change to the other reject-flush site (lines
~627-655) so rejects are always emitted in input order.
- Around line 699-710: The code returns early on pipeline_result with `?` before
checking `rejects_finish_result`, so a failed `rejects_writer.finish()` can be
lost; change control flow to evaluate and propagate `rejects_finish_result` (the
result of `rejects_writer.and_then(...).map(|writer| writer.finish()...)`)
before unwrapping `pipeline_result` into `groups_processed`—i.e., first if
`rejects_finish_result` is Some, call `result?` and log, then map/`?` the
`pipeline_result` into `groups_processed` so any rejects finalization error is
surfaced prior to returning the pipeline error. Ensure you reference and update
`rejects_finish_result`, `rejects_writer`, and `pipeline_result` in that order.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 11819411-946d-43dc-a693-5ff0cecc6069

📥 Commits

Reviewing files that changed from the base of the PR and between 4afa6f6 and d907ac6.

📒 Files selected for processing (4)
  • src/lib/commands/codec.rs
  • src/lib/commands/correct.rs
  • src/lib/commands/duplex.rs
  • src/lib/commands/simplex.rs
✅ Files skipped from review due to trivial changes (1)
  • src/lib/commands/codec.rs

Comment thread src/lib/commands/correct.rs Outdated
Comment thread src/lib/commands/correct.rs Outdated
Comment thread src/lib/commands/duplex.rs
Comment thread src/lib/commands/duplex.rs Outdated
Comment thread src/lib/commands/simplex.rs
Comment thread src/lib/commands/simplex.rs Outdated
@nh13
nh13 force-pushed the nh/pipeline-rejects-streaming branch from d907ac6 to e338c91 Compare April 18, 2026 17:11
@nh13
nh13 temporarily deployed to github-actions April 18, 2026 17:11 — with GitHub Actions Inactive
@coderabbitai

coderabbitai Bot commented Apr 19, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
tests/integration/test_correct_command.rs (1)

187-187: Nit: collapse literal.

Diff
-    let expected_rejects = 10 + 10 + 10;
+    let expected_rejects = 30;

Or derive from the family sizes to keep them in sync.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@tests/integration/test_correct_command.rs` at line 187, The test sets
expected_rejects as a sum of three identical literals (let expected_rejects = 10
+ 10 + 10;) — replace this with a single literal (e.g. 30) or, preferably,
compute it from the existing family-size variables used in the test (so
expected_rejects is derived from those family size identifiers) to keep the
values in sync and avoid magic repetition.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@src/lib/commands/duplex.rs`:
- Around line 77-79: Update the stale comments that say rejects are streamed
during serialize to reflect the current behavior: rejects are now streamed
during process_fn. Specifically, edit the doc comment on the per-thread
accumulator (the comment above the accumulator struct/variable in duplex.rs
referencing "Rejected records are streamed directly to the rejects BAM during
serialize") and the later comment near the code referencing serialize (around
the block related to serialize/process flow) to say "streamed during process_fn"
and, if helpful, mention the process_fn function name where the streaming now
occurs so readers can locate the implementation.

---

Nitpick comments:
In `@tests/integration/test_correct_command.rs`:
- Line 187: The test sets expected_rejects as a sum of three identical literals
(let expected_rejects = 10 + 10 + 10;) — replace this with a single literal
(e.g. 30) or, preferably, compute it from the existing family-size variables
used in the test (so expected_rejects is derived from those family size
identifiers) to keep the values in sync and avoid magic repetition.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 00399063-a873-409a-af88-d01e0e1164f2

📥 Commits

Reviewing files that changed from the base of the PR and between af68aa9 and 779a38d.

📒 Files selected for processing (8)
  • crates/fgumi-consensus/src/duplex_caller.rs
  • src/lib/commands/codec.rs
  • src/lib/commands/correct.rs
  • src/lib/commands/duplex.rs
  • tests/integration/helpers/assertions.rs
  • tests/integration/test_bgzf_eof.rs
  • tests/integration/test_correct_command.rs
  • tests/integration/test_duplex_command.rs
🚧 Files skipped from review as they are similar to previous changes (4)
  • tests/integration/helpers/assertions.rs
  • tests/integration/test_bgzf_eof.rs
  • tests/integration/test_duplex_command.rs
  • src/lib/commands/correct.rs

Comment thread src/lib/commands/duplex.rs Outdated
@nh13
nh13 force-pushed the nh/pipeline-rejects-streaming branch from 779a38d to d6a96fe Compare April 19, 2026 06:19
@nh13
nh13 temporarily deployed to github-actions April 19, 2026 06:19 — with GitHub Actions Inactive

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
src/lib/commands/duplex.rs (1)

706-722: Nit: mutex is held across the entire group's record loop.

This is fine (and arguably desirable — keeps a single MI group's rejects contiguous), but worth noting that for very large groups this serializes workers. The BGZF thread pool still compresses concurrently, so it's unlikely to matter. No change required.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/commands/duplex.rs` around lines 706 - 722, The mutex in the closure
flush_rejects is held for the entire loop over a group's records (rw_arc lock is
acquired before iterating and calling write_raw_record), which can serialize
workers for very large groups; to reduce contention, acquire the lock only for
the minimal write operation (e.g., take a short-lived guard per record or
move/copy each raw record out and lock just to call write_raw_record), ensuring
you still call write_raw_record on the same writer (guard.as_mut()) and preserve
group contiguity as needed; update the flush_rejects closure (and uses of
rw_arc, guard, and write_raw_record) accordingly.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@crates/fgumi-consensus/src/duplex_caller.rs`:
- Around line 1748-1757: The closure collect_rejected_raw currently clones
a_records and b_records causing double memory; change it to move the raw records
out instead of cloning: when track is true, drain/consume a_records and
b_records and return the combined Vec<Vec<u8>> (and return an empty Vec when
false). Update the closure/call sites so it takes ownership (make it FnOnce or
otherwise consume the vectors) and adjust each rejection return that uses
collect_rejected_raw (including the single-strand paths via ss_caller and the
other listed rejection sites) to pass tracking=true and return the moved Vec
rather than cloned copies; references you’ll need to modify:
collect_rejected_raw, a_records, b_records, and the rejection return points (and
ss_caller usage).

---

Nitpick comments:
In `@src/lib/commands/duplex.rs`:
- Around line 706-722: The mutex in the closure flush_rejects is held for the
entire loop over a group's records (rw_arc lock is acquired before iterating and
calling write_raw_record), which can serialize workers for very large groups; to
reduce contention, acquire the lock only for the minimal write operation (e.g.,
take a short-lived guard per record or move/copy each raw record out and lock
just to call write_raw_record), ensuring you still call write_raw_record on the
same writer (guard.as_mut()) and preserve group contiguity as needed; update the
flush_rejects closure (and uses of rw_arc, guard, and write_raw_record)
accordingly.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 886e51f0-e887-4d77-8dbf-48cc78696959

📥 Commits

Reviewing files that changed from the base of the PR and between 779a38d and d6a96fe.

📒 Files selected for processing (8)
  • crates/fgumi-consensus/src/duplex_caller.rs
  • src/lib/commands/codec.rs
  • src/lib/commands/correct.rs
  • src/lib/commands/duplex.rs
  • tests/integration/helpers/assertions.rs
  • tests/integration/test_bgzf_eof.rs
  • tests/integration/test_correct_command.rs
  • tests/integration/test_duplex_command.rs
✅ Files skipped from review due to trivial changes (1)
  • tests/integration/test_correct_command.rs
🚧 Files skipped from review as they are similar to previous changes (2)
  • tests/integration/helpers/assertions.rs
  • src/lib/commands/correct.rs

Comment thread crates/fgumi-consensus/src/duplex_caller.rs Outdated
@nh13
nh13 force-pushed the nh/pipeline-rejects-streaming branch from d6a96fe to fa60481 Compare April 19, 2026 08:59
@nh13
nh13 temporarily deployed to github-actions April 19, 2026 08:59 — with GitHub Actions Inactive
@nh13

nh13 commented Apr 19, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 19, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
src/lib/commands/codec.rs (1)

594-606: Mutex held across write_raw_record calls — watch worker contention.

flush_rejects holds the parking_lot::Mutex for the full loop over recs, so write_raw_record back-pressure (BGZF channel full) serializes all workers. Fine for typical reject volumes given the PR's stated goals, but if a high-rejection input regresses throughput vs. the old buffered path, consider a bounded MPSC to a dedicated rejects-writer thread so workers never block each other.

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/commands/codec.rs` around lines 594 - 606, flush_rejects currently
holds the parking_lot::Mutex (rejects_writer_for_process) across the entire loop
calling write_raw_record, which causes workers to block on back-pressure; change
this so the mutex is not held while doing I/O: either (A) implement a bounded
MPSC channel and spawn a dedicated rejects-writer thread (create a sender stored
in rejects_writer_for_process that workers .try_send() / .send() to, and have
the writer read from the channel and call write_raw_record) or (B) if you want a
minimal change, lock only briefly to take ownership of the inner writer (e.g.,
swap out the Option<Writer> or clone/move it out of the mutex) and then release
the lock before iterating recs and calling write_raw_record; reference
flush_rejects, rejects_writer_for_process, and write_raw_record when making the
change.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Nitpick comments:
In `@src/lib/commands/codec.rs`:
- Around line 594-606: flush_rejects currently holds the parking_lot::Mutex
(rejects_writer_for_process) across the entire loop calling write_raw_record,
which causes workers to block on back-pressure; change this so the mutex is not
held while doing I/O: either (A) implement a bounded MPSC channel and spawn a
dedicated rejects-writer thread (create a sender stored in
rejects_writer_for_process that workers .try_send() / .send() to, and have the
writer read from the channel and call write_raw_record) or (B) if you want a
minimal change, lock only briefly to take ownership of the inner writer (e.g.,
swap out the Option<Writer> or clone/move it out of the mutex) and then release
the lock before iterating recs and calling write_raw_record; reference
flush_rejects, rejects_writer_for_process, and write_raw_record when making the
change.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 8729633b-29c6-4e41-ab20-24bf6863103c

📥 Commits

Reviewing files that changed from the base of the PR and between d6a96fe and fa60481.

📒 Files selected for processing (8)
  • crates/fgumi-consensus/src/duplex_caller.rs
  • src/lib/commands/codec.rs
  • src/lib/commands/correct.rs
  • src/lib/commands/duplex.rs
  • tests/integration/helpers/assertions.rs
  • tests/integration/test_bgzf_eof.rs
  • tests/integration/test_correct_command.rs
  • tests/integration/test_duplex_command.rs
✅ Files skipped from review due to trivial changes (1)
  • tests/integration/test_correct_command.rs
🚧 Files skipped from review as they are similar to previous changes (2)
  • tests/integration/helpers/assertions.rs
  • src/lib/commands/correct.rs

nh13 added 3 commits April 20, 2026 11:16
Rejects are now written to the rejects BAM directly from serialize_fn via
a shared BGZF writer guarded by a parking_lot mutex. The mutex only
serializes the byte append; BGZF compression runs on the writer's own
thread pool. This removes the last unbounded Vec<Vec<u8>> growth path
from simplex, complementing the metric-memory fix in #285/#290.

Peak reject memory drops from proportional to the reject count to a
small fixed working set of the BGZF block buffer.
Same pattern as the simplex change: a shared BGZF writer guarded by a
parking_lot mutex receives reject bytes directly from serialize_fn, so
rejected raw BAM records no longer accumulate in per-thread Vecs.
Same pattern as the simplex/duplex changes: shared BGZF writer guarded
by a parking_lot mutex receives reject bytes directly from serialize_fn.
@nh13
nh13 force-pushed the nh/pipeline-rejects-streaming branch from fa60481 to 8ad2f8f Compare April 20, 2026 18:18
@nh13
nh13 temporarily deployed to github-actions April 20, 2026 18:18 — with GitHub Actions Inactive
@nh13

nh13 commented Apr 20, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 20, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/lib/commands/codec.rs (1)

500-502: ⚠️ Potential issue | 🟡 Minor

Stale comment: rejects are no longer buffered.

The "Rejects buffering semantics are preserved (see follow-up)" note contradicts this PR — rejects now stream straight to disk, not via per-thread buffers. Update or drop.

♻️ Suggested tweak
-        // Per-thread metrics accumulator: bounded metric memory, no unbounded
-        // queue. Rejects buffering semantics are preserved (see follow-up).
+        // Per-thread metrics accumulator: bounded metric memory, no unbounded
+        // queue. Rejects are streamed directly to disk in process_fn (no buffering).
         let collected_metrics = PerThreadAccumulator::<CollectedCodecMetrics>::new(num_threads);
🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/commands/codec.rs` around lines 500 - 502, The inline comment on the
PerThreadAccumulator creation is stale: update or remove the phrase "Rejects
buffering semantics are preserved (see follow-up)" because rejects now stream
directly to disk instead of being buffered per-thread; locate the instantiation
of PerThreadAccumulator::<CollectedCodecMetrics> (bound to collected_metrics) in
codec.rs and either remove the misleading clause or replace it with a brief
accurate note (e.g., "Per-thread metrics accumulator; rejects stream directly to
disk, not buffered") so the comment matches current behavior.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@tests/integration/test_correct_command.rs`:
- Around line 179-239: The test currently only creates 63 input records so it
never spans CorrectUmis’ 1000-template batch and asserts any 30 unique rejects;
to fix, increase the uncorrectable family sizes so the total number of templates
exceeds the batch size (e.g. set far_family_size such that far_family_size * 3 +
other templates > 1000), keep expected_rejects = far_family_size * 3, run the
same Command::new("correct") invocation, then in the reader/seen loop assert
that total == expected_rejects and that the set of seen read names exactly
matches the specific uncorrectable reads produced by the three far families (use
the create_umi_family naming pattern for far_a/far_b/far_c to generate the
expected read names) rather than allowing any 30 records.

---

Outside diff comments:
In `@src/lib/commands/codec.rs`:
- Around line 500-502: The inline comment on the PerThreadAccumulator creation
is stale: update or remove the phrase "Rejects buffering semantics are preserved
(see follow-up)" because rejects now stream directly to disk instead of being
buffered per-thread; locate the instantiation of
PerThreadAccumulator::<CollectedCodecMetrics> (bound to collected_metrics) in
codec.rs and either remove the misleading clause or replace it with a brief
accurate note (e.g., "Per-thread metrics accumulator; rejects stream directly to
disk, not buffered") so the comment matches current behavior.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: dfd68707-ba19-4706-9d77-7eba06147a9c

📥 Commits

Reviewing files that changed from the base of the PR and between fa60481 and 8ad2f8f.

📒 Files selected for processing (12)
  • crates/fgumi-consensus/src/duplex_caller.rs
  • crates/fgumi-sam/src/lib.rs
  • docs/src/guide/migration-from-fgbio.md
  • src/lib/commands/codec.rs
  • src/lib/commands/correct.rs
  • src/lib/commands/duplex.rs
  • src/lib/commands/simplex.rs
  • src/lib/sam/mod.rs
  • tests/integration/helpers/assertions.rs
  • tests/integration/test_bgzf_eof.rs
  • tests/integration/test_correct_command.rs
  • tests/integration/test_duplex_command.rs
✅ Files skipped from review due to trivial changes (1)
  • crates/fgumi-consensus/src/duplex_caller.rs
🚧 Files skipped from review as they are similar to previous changes (6)
  • src/lib/sam/mod.rs
  • tests/integration/helpers/assertions.rs
  • docs/src/guide/migration-from-fgbio.md
  • tests/integration/test_bgzf_eof.rs
  • src/lib/commands/simplex.rs
  • tests/integration/test_duplex_command.rs

Comment thread tests/integration/test_correct_command.rs Outdated
Rejected raw records are now written to the rejects BAM directly from
serialize_fn via a shared BGZF writer guarded by a parking_lot mutex,
completing the rejects-streaming work started in simplex/duplex/codec.
Removes the last per-thread Vec<Vec<u8>> accumulator in the pipeline
commands.
@nh13
nh13 force-pushed the nh/pipeline-rejects-streaming branch from 8ad2f8f to b7dceb0 Compare April 20, 2026 19:44
@nh13
nh13 temporarily deployed to github-actions April 20, 2026 19:44 — with GitHub Actions Inactive
@nh13
nh13 merged commit 5742d4f into main Apr 20, 2026
7 of 9 checks passed
@nh13
nh13 deleted the nh/pipeline-rejects-streaming branch April 20, 2026 19:52
This was referenced Apr 20, 2026

This branch was previously deployed

1 inactive deployment
github-actions — b7dceb0c Deployed Apr 20, 2026 by nh13 via coverage #1271
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

enhancement New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant