feat(sort): add per-phase --sort-threads / --merge-threads - #608
Conversation
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (3)
WalkthroughThe sorter now accepts separate Phase 1 and Phase 2 thread counts through the CLI and builder API. Worker pools cap active workers per phase, while sort paths switch caps at the merge boundary and tests verify clamping and byte-identical output. ChangesPer-phase sorter concurrency
Estimated code review effort: 4 (Complex) | ~60 minutes Sequence Diagram(s)sequenceDiagram
participant CLI
participant RawExternalSorter
participant SortWorkerPool
participant Phase1
participant Phase2
CLI->>RawExternalSorter: configure sort and merge thread counts
RawExternalSorter->>SortWorkerPool: create pool and activate Phase 1 workers
RawExternalSorter->>Phase1: accumulate, sort, and spill
RawExternalSorter->>SortWorkerPool: activate Phase 2 workers
RawExternalSorter->>Phase2: merge and write output
🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #608 +/- ##
========================================
Coverage 93.52% 93.52%
========================================
Files 175 175
Lines 105404 105530 +126
========================================
+ Hits 98574 98699 +125
- Misses 6830 6831 +1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
There was a problem hiding this comment.
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
crates/fgumi-sort/src/external.rs (1)
2581-2664: 🚀 Performance & Scalability | 🟠 Major | ⚡ Quick win
--write-indexpath doesn't honormerge_threads.
pool.set_active_workers(self.phase2_threads())is raised here, but the two writers this function actually uses ignore it: the in-memory branch'screate_indexing_bam_writer(..., self.threads)(~line 2603) andmerge_chunks_with_index'slet writer_threads = self.threads;(~line 3493) both hardcode the basethreads, notphase2_threads(). Contrast with the non-indexed merge (merge_chunks_generic), whosePooledBamWritercorrectly inherits the raised cap via the shared pool. Sofgumi sort --order coordinate --write-index --merge-threads Nsilently caps output compression at--threadsinstead ofN— output stays byte-identical (compression thread count doesn't change block content), but the documented scheduling contract ("output write" phase) is broken for this path. No test combineswrite_indexwithsort_threads/merge_threadseither, which is how this slipped through.🐛 Proposed fix
@@ sort_coordinate_with_index (in-memory branch) timer.time_write_output(|| { let mut writer = create_indexing_bam_writer( output, &output_header, self.output_compression, - self.threads, + self.phase2_threads(), )?; @@ merge_chunks_with_index - let writer_threads = self.threads; + let writer_threads = self.phase2_threads();🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@crates/fgumi-sort/src/external.rs` around lines 2581 - 2664, Update the write-index paths to use the Phase 2 merge-thread count rather than the base thread count: pass self.phase2_threads() to create_indexing_bam_writer in the in-memory branch and replace the self.threads-based writer_threads value in merge_chunks_with_index. Preserve the existing output and indexing behavior while ensuring --merge-threads controls compression during indexed output.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Outside diff comments:
In `@crates/fgumi-sort/src/external.rs`:
- Around line 2581-2664: Update the write-index paths to use the Phase 2
merge-thread count rather than the base thread count: pass self.phase2_threads()
to create_indexing_bam_writer in the in-memory branch and replace the
self.threads-based writer_threads value in merge_chunks_with_index. Preserve the
existing output and indexing behavior while ensuring --merge-threads controls
compression during indexed output.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: c5fd5733-f75e-4170-b328-1bc40427ef76
📒 Files selected for processing (3)
crates/fgumi-sort/src/external.rscrates/fgumi-sort/src/worker_pool.rssrc/lib/commands/sort.rs
5c1b587 to
2fdf559
Compare
|
Addressed the outside-diff finding: the Also closed the coverage gap: |
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
`fgumi sort` had a single `-@/--threads` sizing both Phase 1 (accumulate/sort/spill) and Phase 2 (merge/write). Those phases contend with different things: Phase 1 competes with whatever feeds the sort, while Phase 2 runs after the producer is done. A user running `bwa mem -t 32 ... | fgumi sort -@ 8` had no way to cede cores to the aligner during ingest while keeping the merge wide. `SortWorkerPool` gains an `active_worker_limit` and `set_active_workers`. One pool still spans both phases -- sized to the wider of the two counts -- and the driver flips the active cap at the ingest/merge boundary. Workers above the cap idle instead of taking new work, and re-check the limit within `MAX_BACKOFF_US`, so raising it reactivates them without an explicit wake. Crucially, a capped worker still drains items it already holds and only declines to acquire *new* work. Held items are per-worker, so a worker that froze while holding work would strand that output with no other worker able to advance it. `RawExternalSorter` gains `sort_threads`/`merge_threads` builders and `phase1_threads()`/`phase2_threads()`, each falling back to `threads` independently, and the CLI exposes `--sort-threads` / `--merge-threads` with the same defaulting. This is purely a scheduling knob: output is byte-identical for any split. Tests cover byte-identical output against a plain `--threads` run across four asymmetric splits (exercising both the wider-Phase-1 and wider-Phase-2 pool sizings over a spilling multi-chunk merge), independent fallback of each override, clamping of the active-worker count, and -- on a single pool, which is the only way to show it -- that workers idled by a cap come back online when the cap is raised.
2fdf559 to
6585365
Compare
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
fgumi sorthad a single-@/--threadssizing both Phase 1 (accumulate/sort/spill) and Phase 2 (merge/write). Those phases contend with different things: Phase 1 competes with whatever feeds the sort, while Phase 2 runs after the producer is finished. A user runningbwa mem -t 32 ... | fgumi sort -@ 8had no way to cede cores to the aligner during ingest while keeping the merge wide.Engine
SortWorkerPoolgains anactive_worker_limitandset_active_workers. A single pool still spans both phases — sized to the wider of the two counts — and the driver flips the active cap at the ingest/merge boundary. Workers above the cap idle rather than taking new work and re-check the limit withinMAX_BACKOFF_US, so raising it reactivates them without an explicit wake.The invariant that makes this safe: a capped worker still drains items it already holds and only declines to acquire new work. Held items are per-worker, so a worker that froze while holding work would strand that output with no other worker able to advance it.
RawExternalSortergainssort_threads/merge_threadsbuilders plusphase1_threads()/phase2_threads(), each falling back tothreadsindependently. The CLI exposes--sort-threads/--merge-threadswith the same defaulting.This is purely a scheduling knob: output is byte-identical for any split.
On testing a knob that by definition changes nothing observable
An end-to-end run cannot tell a correctly wired flag from one that was parsed and then silently dropped — both produce identical bytes. So the sorter construction is extracted from
executeintoSort::build_sorter, and the wiring is asserted directly on the resolved per-phase counts. I verified that test is non-vacuous by unwiring the flag and confirming it fails; the output-identity test, as expected, does not.Both properties are covered separately:
build_sorterresolves each override, and each falls back to--threadsindependently.--threadsrun over a spilling multi-chunk merge, exercising both the wider-Phase-1 and wider-Phase-2 pool sizings.Runall-only halves of the upstream change —
SortSpillDecompress::with_max_concurrencyand the chains builder — are deliberately not ported; they depend on crates that do not exist onmain.Verification
cargo ci-test(5557 tests),cargo ci-fmt,cargo ci-lint, andRUSTDOCFLAGS="-D warnings" cargo ci-docpass. Patch coverage 96%.Summary by CodeRabbit
--sort-threadsand--merge-threadsforfgumi sort, with safe clamping.