Repository navigation
perf: parallel gather for the in-memory sort fast path - #966
Conversation
The streaming sort's SortMerge fast path (single already-sorted in-memory chunk, no k-way merge) gathered every record into output blocks serially on the lone detached coordination thread. Profiling a 20M-record coordinate sort at 32 threads showed this gather as a flat ~1.2s serial cost that does not scale with thread count and floors the wall clock while the worker pool sits ~80% idle -- its CPU is per-record record_bytes() arena slices plus the extend_from_slice memcpy into block buffers. Fan that gather across a bounded, step-owned rayon pool sized to the phase-2 (merge) thread budget: - Precompute output-block boundaries up front from the per-record lengths alone (new InMemoryChunk/MemoryChunkErased::record_len, off the cold arena), reproducing the serial count/byte-cap split byte-for-byte. A MergeBatchBuilder::FRAME_OVERHEAD_PER_RECORD const keeps the byte-cap plan in lockstep with each output builder (BlockOutput frames +4/record, RecordBatchOutput +0). - Gather a bounded window of blocks in parallel (chunk is Send+Sync; record_bytes is a read-only slice), then push them in strict ascending dense ordinal order through the existing held-slot backpressure idiom -- required, since the Detached output edge has no reorder stage (reassembly is downstream at BgzfCompress). Gated on a min record count so small chunks keep the cheaper serial gather. Output is identical (same records, block boundaries, and dense ordinals), covered by a serial-vs-parallel parity test across the count-cap, byte-cap, single-block, and partial-tail regimes, and the existing oracle parity suite.
|
Note Reviews pausedUse the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Essentials Run ID: 📒 Files selected for processing (6)
Included review availability: 0 reviews are currently available. Your included PR review attempts over the past 7 days set your current allowance at 1 review per hour. WalkthroughChangesParallel in-memory sort gathering
Priority: ➖ Normal Estimated code review effort: 4 (Complex) | ~45 minutes Change: Feature Sequence Diagram(s)sequenceDiagram
participant SortMerge
participant MemoryChunk
participant WorkerPool
participant Output
SortMerge->>MemoryChunk: Read record lengths
SortMerge->>WorkerPool: Submit bounded block ranges
WorkerPool->>Output: Produce ordered blocks or batches
Output-->>SortMerge: Accept or reject output
SortMerge->>Output: Resume pending output
Suggested labels: Merge Risk: ⚪ Minimal · up to The parallel fast path is mergeable based on the available evidence, with no actionable risk identified. 🚥 Pre-merge checks | ✅ 2 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (2 passed)
Comment |
|
@coderabbitai pause |
✅ Action performedReviews paused. |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## main #966 +/- ##
==========================================
- Coverage 96.08% 96.05% -0.03%
==========================================
Files 291 292 +1
Lines 144093 145159 +1066
==========================================
+ Hits 138447 139435 +988
- Misses 5646 5724 +78 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
What
fgumi sort's in-memory fast path gathered the whole sorted chunk into output blocks serially on one detached coordination thread — a flat ~1.2 s cost that did not scale with--threadsand floored the sort's wall clock while the worker pool sat mostly idle. This PR fans that gather across a bounded, step-owned rayon pool sized to the phase-2 (merge) thread budget.Only the fast path (a single already-sorted in-memory chunk, no k-way merge) is touched. The spill / k-way-merge path is deliberately untouched — it is a genuinely serial loser-tree walk whose winner order cannot be produced out of order.
How
record_lenaccessors (InMemoryChunk/TemplateMemChunk/MemoryChunkErased) read only the per-recordlenindex.plan_fast_path_blocksreproduces the serial count/byte-cap split byte-for-byte from lengths alone. AMergeBatchBuilder::FRAME_OVERHEAD_PER_RECORDconst (4for the framedBlockBuilder,0forRecordBatchBuilder) keeps the byte-cap plan in lockstep with each builder'stotal_bytes()growth.SortMergeisDetached, so its output edge has no reorder stage (the by-ordinal reassembly is downstream atBgzfCompressand needs a dense, gap-free stream). Gathered-but-unpushed blocks ride the existing held-slot backpressure idiom plus a smallfast_pendingqueue, drained in order at the top of the next dispatch.fast_path_threads * 4blocks. Gated onFAST_PATH_PARALLEL_MIN_RECORDS = 64Kiso small chunks keep the cheaper serial gather; a degenerate single-block plan falls through to serial rather than spinning up a pool. The pool is step-owned and bounded tonum_phase2_threadsso it never oversubscribes past--threads.Results (EC2 c8g.8xlarge, 32 vCPU Graviton4, coordinate sort, 20 M records)
Total CPU-seconds is roughly flat — the win is converting serial gather time into parallel gather time. 16 threads is the throughput sweet spot (21.9 M rec/s).
Correctness
Output is byte-for-byte identical to the serial gather. Verified with
fgumi compare bams --command sort(record-content + order) across coordinate / template-coordinate / queryname at 16 threads, and against the spill regime (fast path not taken) — all IDENTICAL. Parity unit tests assert the parallel gather emits identical dense ordinals and framed bytes as the serial gather across count-cap, byte-cap, small-multi-block, partial-tail, and a deterministic backpressure regime (small output queue forcing sustained mid-window reject/resume) for both theBlockOutputandRecordBatchOutputbuilders.Gates
cargo ci-fmt,cargo ci-lint,cargo ci-docclean;cargo ci-test— 10219 passed, 31 skipped.Risk: sort output stays byte-identical through serial-equivalent block planning and ordered emission; unsafe code: none added, and the CLAUDE.md allowlist remains unchanged; memory and backpressure policy: bounded parallel gathering uses the merge-thread budget.
Fix: parallelize only large, multi-block in-memory gathers. Keep small inputs serial. Keep spill and k-way merge serial.
record_lenaccessors for in-memory chunks.MergeBatchBuilder.