Skip to content

perf(sort): redesign phase 2 work stealing and bound rayon to --threads - #247

Merged
nh13 merged 1 commit into
mainfrom
nh/perf-sort-no-rayon-oversubscribe
Apr 9, 2026
Merged

nh13 merged 1 commit into
mainfrom
nh/perf-sort-no-rayon-oversubscribe

Conversation

@nh13

@nh13 nh13 commented Apr 8, 2026

Copy link
Copy Markdown
Member

Summary

Fixes two independent problems hurting the N+2 sort worker-pool path on large inputs, plus a post-review cleanup sweep.

Phase 2 redesign: per-file work-stealing state

Phase 2 regressed badly on WES (1kg-wes-HG00100.bam, 13 GB, 219M records): 450s+ wall and ~42 GB peak RSS under --threads 8 --max-memory 768M, driven by an over-coupled producer/consumer with global queues that deadlocked under backpressure and left workers idle while OOMing.

Replace it with per-file state:

struct Phase2FileState {
    reader: Mutex<Phase2Reader>,           // raw BGZF block reads
    reader_eof: AtomicBool,                // hot-path fast check
    raw_blocks: Mutex<VecDeque<...>>,      // bounded FIFO
    decompressed: Mutex<ReorderBuffer>,    // in-order decompressed
    decomp_in_flight: AtomicUsize,         // race-safe drain
}

Workers scan files with try_lock everywhere and steal work across spill files. Lock order is strict: reader -> raw_blocks -> decompressed.

  • Deadlock-free admission: a popped raw block is always admitted when it is the gap-filler the reorder buffer is waiting for, even if the decompressed cap is full. Prevents head-of-line deadlock where every worker holds a stale raw block while the buffer waits on a missing serial.
  • Race fix: is_drained() must observe decomp_in_flight > 0 to avoid exiting Phase 2 between a worker popping a raw block and inserting the decompressed bytes.
  • FIFO bug fix: serial assignment and the matching push into raw_blocks happen under both the reader and raw_blocks locks held simultaneously. Releasing the reader lock first allowed another worker to interleave a higher serial ahead of the lower one and break the gap-filler invariant.
  • Hot-path fast check: is_drained() fast-paths on a lock-free reader_eof atomic; the common case (reader not at EOF) exits after a single atomic load, versus three mutex acquisitions per decompressed block.

Rayon oversubscription

Rayon's global pool defaults to num_cpus::get(), which silently violated --threads on machines with more physical cores than the requested thread count. Every par_sort / par_sort_into_chunks / par_sort_unstable_by / par_chunks_mut call in the raw sort path now runs under a rayon::ThreadPool built with num_threads = self.threads, via rayon_pool.install(|| ...). rayon::current_num_threads() inside parallel_radix_sort_* now returns the requested thread count, so chunk sizing is consistent with the worker-pool sizing.

Oversubscription with the SortWorkerPool is not introduced: every call site is preceded by drain_pending_spill, which joins the prior chunk's I/O thread, so the sort worker pool is provably idle when rayon fans out.

Additional fixes

  • Use par_sort in the three in-memory sort fallbacks (coordinate, coordinate-index, template-coordinate). The single-threaded buffer.sort() serialized the sort step onto one core even with --threads set high when the entire input fit in memory.
  • In-memory split path replaces the O(n²) drain(..chunk_size) loop with O(n) split_off from the tail. The old loop memmoved the tail on every drain and became a measurable cost once entries grew into the millions under tight --max-memory budgets.
  • Drop Phase2FileState::source_id (always equal to the file's index in the shared vector), SourceParserState::eof, and the unused _worker parameter on is_phase_complete — all written or threaded through but never read.
  • Refresh stale module, struct, and function docs to describe the per-file work-stealing Phase 2 model.

Result

WES (1kg-wes-HG00100.bam, --threads 8 --max-memory 4g):

wall peak RSS sort violations
before (5df7e17) 450s+ ~42 GB OOM-prone
this branch 139 s 37.6 GB 0

Test plan

  • `cargo ci-fmt`
  • `cargo ci-lint`
  • `cargo ci-test` (2351 tests pass)
  • WES end-to-end sort with coordinate order, verified sorted output
  • Reviewer spot-check of Phase 2 lock order and admission rule

@nh13
nh13 had a problem deploying to github-actions April 8, 2026 22:44 — with GitHub Actions Failure
@nh13
nh13 marked this pull request as ready for review April 8, 2026 22:45
@coderabbitai

coderabbitai Bot commented Apr 8, 2026 •

Copy link
Copy Markdown

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: d22809da-2018-4b75-a11d-1a53d1bdec68

📥 Commits

Reviewing files that changed from the base of the PR and between 50071ce and 789fc66.

📒 Files selected for processing (4)
  • src/lib/sort/inline_buffer.rs
  • src/lib/sort/raw.rs
  • src/lib/sort/read_ahead.rs
  • src/lib/sort/worker_pool.rs
✅ Files skipped from review due to trivial changes (1)
  • src/lib/sort/worker_pool.rs
🚧 Files skipped from review as they are similar to previous changes (1)
  • src/lib/sort/read_ahead.rs

📝 Walkthrough

Walkthrough

Refactors Phase 2 into a single Phase2FileWork step with per-file Phase2FileState (disk reader, raw-block FIFO, decompressed ReorderBuffer, atomic in-flight counter) and uses is_drained() for completion gating. Replaces consumer-side per-source buffering with a main-thread SourceParserState that consumes from each file’s reorder buffer and removes prior pending bookkeeping. Adds Phase2Guard for RAII teardown and pool APIs set_phase2_files / phase2_files / clear_phase2_files. Constrains Rayon via RawExternalSorter::build_sort_rayon_pool and routes sorts through rayon_pool.install(...). SortPhaseTimer::time_sort now returns the closure’s value. PooledInputStream::drain_queue uses a dynamic reorder-cap loop.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed Title accurately summarizes the two main changes: Phase 2 work-stealing redesign and Rayon thread-count binding to --threads flag.
Description check ✅ Passed Description comprehensively explains both Phase 2 restructuring and Rayon oversubscription fixes with technical details, performance results, and test coverage.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch nh/perf-sort-no-rayon-oversubscribe

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@src/lib/sort/raw.rs`:
- Around line 1717-1724: The code calls a non-existent method
par_sort_into_chunks on buffer (used when building memory_chunks), which breaks
compilation; either implement par_sort_into_chunks on the concrete buffer types
(RecordBuffer/TemplateRecordBuffer) or revert these branches to use the existing
par_sort() and then materialize chunked Vec<Vec<(RawCoordinateKey, Vec<u8>)>>
via the buffer's current chunking/flush API (i.e., sort with buffer.par_sort()
inside rayon_pool.install and then split the sorted buffer into chunks using the
existing method you already have), and apply the same fix to the other two
occurrences around the referenced blocks.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: fb5cd1c6-670a-46a4-be8f-5509b71efeac

📥 Commits

Reviewing files that changed from the base of the PR and between 0757238 and ac35c07.

📒 Files selected for processing (3)
  • src/lib/sort/raw.rs
  • src/lib/sort/read_ahead.rs
  • src/lib/sort/worker_pool.rs

Comment thread src/lib/sort/raw.rs
@nh13
nh13 force-pushed the nh/perf-sort-no-rayon-oversubscribe branch from ac35c07 to 22e94bd Compare April 9, 2026 00:31
@nh13
nh13 temporarily deployed to github-actions April 9, 2026 00:31 — with GitHub Actions Inactive
@codecov

codecov Bot commented Apr 9, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 93.93382% with 33 lines in your changes missing coverage. Please review.
✅ Project coverage is 89.42%. Comparing base (0757238) to head (789fc66).
⚠️ Report is 2 commits behind head on main.

Files with missing lines Patch % Lines
src/lib/sort/raw.rs 90.32% 15 Missing ⚠️
src/lib/sort/worker_pool.rs 93.53% 15 Missing ⚠️
src/lib/sort/inline_buffer.rs 98.69% 2 Missing ⚠️
src/lib/sort/read_ahead.rs 75.00% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #247      +/-   ##
==========================================
+ Coverage   89.29%   89.42%   +0.12%     
==========================================
  Files         118      118              
  Lines       57201    57398     +197     
==========================================
+ Hits        51080    51326     +246     
+ Misses       6121     6072      -49     

☔ View full report in Codecov by Sentry.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/lib/sort/raw.rs (1)

2451-2460: ⚠️ Potential issue | 🟠 Major

Make Phase 2 teardown run on every exit path.

From pool.set_phase(phase::PHASE2) until the manual cleanup block, several ? sites can return early. That leaves the pool in PHASE2 and keeps phase2_files published on merge errors. Wrap this in a drop guard so the reset always happens.

Suggested shape
+struct Phase2Cleanup<'a> {
+    pool: &'a Arc<SortWorkerPool>,
+    active: bool,
+}
+
+impl Drop for Phase2Cleanup<'_> {
+    fn drop(&mut self) {
+        if self.active {
+            self.pool.set_phase(phase::LEGACY);
+            self.pool.clear_phase2_files();
+        }
+    }
+}
+
         let mut consumer: Option<MainThreadChunkConsumer<K>> = if num_disk > 0 {
             let files = pool.phase2_files();
             let consumer = MainThreadChunkConsumer::new(
                 files,
                 pool.decompress_error_flag(),
                 pool.chunk_read_error_flag(),
                 pool.worker_panicked_flag(),
             );
             pool.set_phase(phase::PHASE2);
             Some(consumer)
         } else {
             None
         };
+        let _phase2_cleanup = consumer.as_ref().map(|_| Phase2Cleanup { pool, active: true });
...
-        if consumer.is_some() {
-            pool.set_phase(phase::LEGACY);
-            drop(consumer.take());
-            pool.clear_phase2_files();
-        }

Also applies to: 2472-2475, 2531-2558

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/sort/raw.rs` around lines 2451 - 2460, The code sets
pool.set_phase(phase::PHASE2) and then constructs a MainThreadChunkConsumer and
performs operations that may return early via ?; if an early return occurs the
pool remains in PHASE2 and phase2_files stay published. Fix by introducing a
RAII drop guard (e.g., a small struct ResetPhaseOnDrop) created immediately
after calling pool.set_phase(phase::PHASE2) that in its Drop impl resets the
phase and clears/hides phase2_files, then ensure the guard is dropped only after
the manual cleanup block completes (or explicitly into_inner when success) so
the reset runs on every exit path; apply the same pattern around the other
regions mentioned (near lines where MainThreadChunkConsumer is created and where
phase is set, e.g., the blocks covering 2472-2475 and 2531-2558).
♻️ Duplicate comments (1)
src/lib/sort/raw.rs (1)

1721-1722: ⚠️ Potential issue | 🔴 Critical

Replace par_sort_into_chunks or add that API first.

These call sites still target a method that RecordBuffer / TemplateRecordBuffer do not expose, so this is a hard compile break. Either add par_sort_into_chunks on those buffer types or keep par_sort() and materialize merge chunks with the existing iterators/split logic.

Also applies to: 1893-1894, 2298-2299

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/sort/raw.rs` around lines 1721 - 1722, Call sites use
buffer.par_sort_into_chunks which RecordBuffer/TemplateRecordBuffer lack,
causing compile errors; either implement par_sort_into_chunks on those types or
revert the call sites to use the existing par_sort() and then materialize chunks
via the current iterators/split logic. Update occurrences such as the
timer.time_sort(|| rayon_pool.install(||
buffer.par_sort_into_chunks(self.threads))) (and the similar sites around the
other mentioned lines) to either call a newly added
RecordBuffer::par_sort_into_chunks/TemplateRecordBuffer::par_sort_into_chunks
that returns chunked iterators compatible with the merge stage, or change them
back to buffer.par_sort() and follow with the code that splits the sorted buffer
into merge chunks using the existing iterator/split helpers so callers compile
without the new API.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Outside diff comments:
In `@src/lib/sort/raw.rs`:
- Around line 2451-2460: The code sets pool.set_phase(phase::PHASE2) and then
constructs a MainThreadChunkConsumer and performs operations that may return
early via ?; if an early return occurs the pool remains in PHASE2 and
phase2_files stay published. Fix by introducing a RAII drop guard (e.g., a small
struct ResetPhaseOnDrop) created immediately after calling
pool.set_phase(phase::PHASE2) that in its Drop impl resets the phase and
clears/hides phase2_files, then ensure the guard is dropped only after the
manual cleanup block completes (or explicitly into_inner when success) so the
reset runs on every exit path; apply the same pattern around the other regions
mentioned (near lines where MainThreadChunkConsumer is created and where phase
is set, e.g., the blocks covering 2472-2475 and 2531-2558).

---

Duplicate comments:
In `@src/lib/sort/raw.rs`:
- Around line 1721-1722: Call sites use buffer.par_sort_into_chunks which
RecordBuffer/TemplateRecordBuffer lack, causing compile errors; either implement
par_sort_into_chunks on those types or revert the call sites to use the existing
par_sort() and then materialize chunks via the current iterators/split logic.
Update occurrences such as the timer.time_sort(|| rayon_pool.install(||
buffer.par_sort_into_chunks(self.threads))) (and the similar sites around the
other mentioned lines) to either call a newly added
RecordBuffer::par_sort_into_chunks/TemplateRecordBuffer::par_sort_into_chunks
that returns chunked iterators compatible with the merge stage, or change them
back to buffer.par_sort() and follow with the code that splits the sorted buffer
into merge chunks using the existing iterator/split helpers so callers compile
without the new API.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 0978eb85-d5dd-4302-bec0-b19eedce75e5

📥 Commits

Reviewing files that changed from the base of the PR and between ac35c07 and 22e94bd.

📒 Files selected for processing (3)
  • src/lib/sort/raw.rs
  • src/lib/sort/read_ahead.rs
  • src/lib/sort/worker_pool.rs
✅ Files skipped from review due to trivial changes (1)
  • src/lib/sort/worker_pool.rs

@nh13
nh13 force-pushed the nh/perf-sort-no-rayon-oversubscribe branch from 22e94bd to 50071ce Compare April 9, 2026 02:48
@nh13
nh13 temporarily deployed to github-actions April 9, 2026 02:48 — with GitHub Actions Inactive
@nh13

nh13 commented Apr 9, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Apr 9, 2026

Copy link
Copy Markdown
✅ Actions performed

Review triggered.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
src/lib/sort/inline_buffer.rs (1)

248-294: Factor the chunking skeleton once.

These two methods now duplicate the same threshold check, chunk sizing, parallel chunk sort, and materialization flow. A small internal helper would keep future tuning and fixes in one place.

Also applies to: 844-876

🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed.

In `@src/lib/sort/inline_buffer.rs` around lines 248 - 294, Extract the duplicated
chunking/sorting/materialization logic into a private helper (e.g., fn
materialize_sorted_chunks(&mut self, threads: usize) ->
Vec<Vec<(RawCoordinateKey, Vec<u8>)>>) that encapsulates the threshold checks
(using RADIX_THRESHOLD, n = self.refs.len(), and the threads <= 1 / small-n
single-threaded path), chunk_size calculation (n.div_ceil(threads)), parallel
radix-sorting of chunks via self.refs.par_chunks_mut(chunk_size) with
radix_sort_record_refs, and the final materialization using
self.refs.chunks(chunk_size) mapping each Ref to (RawCoordinateKey { sort_key:
r.sort_key }, self.get_record(r).to_vec()). Replace the bodies of
par_sort_into_chunks and the other similar method (the one at lines ~844-876) to
call this new helper so tuning and fixes live in one place; keep all semantics
and use of radix_sort_record_refs, get_record, refs, and RADIX_THRESHOLD
unchanged.
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.

Inline comments:
In `@src/lib/sort/raw.rs`:
- Around line 2136-2158: The emitted chunk boundaries currently don't match the
ranges sorted by par_chunks_mut because you split from the tail with split_off;
change the emission logic so chunks are carved from the head in the same
fixed-size ranges used by par_chunks_mut (i.e., [0..chunk_size),
[chunk_size..2*chunk_size), ...), ensuring each pushed Vec corresponds to a
sorted sub-slice; specifically adjust the code that builds chunks (the
remaining/split_off loop that produces chunks) to iterate over entries from the
start in chunk_size increments (handling the final short chunk) so that
chunk_size, par_chunks_mut, remaining, split_off and chunks align with the k-way
merge assumption that every source chunk is sorted.

---

Nitpick comments:
In `@src/lib/sort/inline_buffer.rs`:
- Around line 248-294: Extract the duplicated chunking/sorting/materialization
logic into a private helper (e.g., fn materialize_sorted_chunks(&mut self,
threads: usize) -> Vec<Vec<(RawCoordinateKey, Vec<u8>)>>) that encapsulates the
threshold checks (using RADIX_THRESHOLD, n = self.refs.len(), and the threads <=
1 / small-n single-threaded path), chunk_size calculation (n.div_ceil(threads)),
parallel radix-sorting of chunks via self.refs.par_chunks_mut(chunk_size) with
radix_sort_record_refs, and the final materialization using
self.refs.chunks(chunk_size) mapping each Ref to (RawCoordinateKey { sort_key:
r.sort_key }, self.get_record(r).to_vec()). Replace the bodies of
par_sort_into_chunks and the other similar method (the one at lines ~844-876) to
call this new helper so tuning and fixes live in one place; keep all semantics
and use of radix_sort_record_refs, get_record, refs, and RADIX_THRESHOLD
unchanged.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 988417c3-a276-4e96-8096-c2d89d4d6942

📥 Commits

Reviewing files that changed from the base of the PR and between 22e94bd and 50071ce.

📒 Files selected for processing (4)
  • src/lib/sort/inline_buffer.rs
  • src/lib/sort/raw.rs
  • src/lib/sort/read_ahead.rs
  • src/lib/sort/worker_pool.rs
🚧 Files skipped from review as they are similar to previous changes (1)
  • src/lib/sort/read_ahead.rs

Comment thread src/lib/sort/raw.rs
Two independent problems were hurting the N+2 sort worker-pool path
on large inputs; this squash fixes both and lands a post-review
cleanup sweep.

## Phase 2 redesign: per-file work-stealing state

Phase 2 regressed badly on WES (1kg-wes-HG00100.bam, 13G, 219M
records): 450s+ wall and ~42 GB peak RSS under --threads 8
--max-memory 768M, driven by an over-coupled producer/consumer with
global queues that deadlocked under backpressure and left workers
idle while OOMing.

Replace it with per-file state:

  struct Phase2FileState {
      reader: Mutex<Phase2Reader>,           // raw BGZF block reads
      reader_eof: AtomicBool,                // hot-path fast check
      raw_blocks: Mutex<VecDeque<...>>,      // bounded FIFO
      decompressed: Mutex<ReorderBuffer>,    // in-order decompressed
      decomp_in_flight: AtomicUsize,         // race-safe drain
  }

Workers scan files with try_lock everywhere and steal work across
spill files. Lock order is strict: reader -> raw_blocks ->
decompressed.

- Deadlock-free admission: a popped raw block is always admitted when
  it is the gap-filler the reorder buffer is waiting for, even if the
  decompressed cap is full, preventing head-of-line deadlock where
  every worker holds a stale raw block while the buffer waits on a
  missing serial.
- Race fix: is_drained() must observe decomp_in_flight > 0 to avoid
  exiting Phase 2 between a worker popping a raw block and inserting
  the decompressed bytes.
- FIFO bug fix: serial assignment and the matching push into
  raw_blocks happen under both the reader and raw_blocks locks held
  simultaneously. Releasing the reader lock first allowed another
  worker to interleave a higher serial ahead of the lower one and
  break the gap-filler invariant.
- is_drained() fast-paths on a lock-free reader_eof atomic instead of
  acquiring the reader mutex; the common case (reader not at EOF)
  now exits after a single atomic load, versus three mutexes per
  decompressed block on the original code.

## Rayon oversubscription

Rayon's global pool defaults to num_cpus::get(), which silently
violated --threads on machines with more physical cores than the
requested thread count. Every par_sort / par_sort_into_chunks /
par_sort_unstable_by / par_chunks_mut call in the raw sort path now
runs under a rayon::ThreadPool built with num_threads = self.threads,
via rayon_pool.install(|| ...). rayon::current_num_threads() inside
parallel_radix_sort_* now returns the requested thread count, so
chunk sizing is consistent with the worker pool sizing.

Oversubscription with the SortWorkerPool is not introduced: every
call site is preceded by drain_pending_spill, which joins the prior
chunk's I/O thread, so the sort worker pool is provably idle when
rayon fans out.

## Additional fixes

- Use par_sort in the three in-memory sort fallbacks (coordinate,
  coordinate-index, template-coordinate). The single-threaded
  buffer.sort() serialized the sort step onto one core even with
  --threads set high when the entire input fit in memory.
- In-memory split path in RawExternalSorter replaces the O(n^2)
  drain(..chunk_size) loop with O(n) split_off from the tail. The
  old loop memmoved the tail on every drain and became a measurable
  cost once entries grew into the millions under tight --max-memory
  budgets. The merge consumer accepts chunks in any order so the
  original layout need not be preserved.
- Drop Phase2FileState::source_id: it was always equal to the file's
  index in the shared vector. Log lines use the cursor index and
  set_phase2_files simplifies to &[PathBuf].
- Drop SourceParserState::eof and the unused _worker parameter on
  is_phase_complete; both were written or threaded through but never
  read.
- Refresh stale module, struct, and function docs to describe the
  per-file work-stealing Phase 2 model. No stale references to
  ReadChunkBlocks, decompressed_chunks, or GenericKeyedChunkReader
  remain.

## Result

WES (1kg-wes-HG00100.bam, --threads 8 --max-memory 4g):
  before (5df7e17): 450s+ wall, ~42 GB peak RSS (OOM-prone)
  after:            139 s wall,  37.6 GB peak RSS, 0 sort violations
@nh13
nh13 force-pushed the nh/perf-sort-no-rayon-oversubscribe branch from 50071ce to 789fc66 Compare April 9, 2026 17:20
@nh13
nh13 temporarily deployed to github-actions April 9, 2026 17:20 — with GitHub Actions Inactive
@nh13
nh13 merged commit 09ce075 into main Apr 9, 2026
9 checks passed
@nh13
nh13 deleted the nh/perf-sort-no-rayon-oversubscribe branch April 9, 2026 18:43
@nh13 nh13 mentioned this pull request Apr 9, 2026

This branch was previously deployed

1 inactive deployment
github-actions — 789fc66e Deployed Apr 9, 2026 by nh13 via coverage #1000
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant