Skip to content

perf(sort): narrow radix passes, reuse the stored queryname NUL, and borrow record bytes on ingest - #606

Merged
nh13 merged 3 commits into
mainfrom
nh/perf-sort-radix-key-borrow
Jul 21, 2026
Merged

nh13 merged 3 commits into
mainfrom
nh/perf-sort-radix-key-borrow

Conversation

@nh13

@nh13 nh13 commented Jul 21, 2026 •

Copy link
Copy Markdown
Member

Three independent throughput improvements to the sort engine, ported from feat-runall and re-verified against main. Output is byte-identical in all three cases; the only behavior changes are bug fixes described below.

Suggested reading order is commit by commit — they touch different files and are independently reviewable.

perf(sort): build the natural queryname key from the NUL BAM already stores

extract_queryname_key stripped the trailing NUL that BAM stores and re-appended one into a fresh Vec per record. BAM already stores the name NUL-terminated and l_read_name counts the terminator, so the stored bytes are directly usable.

The fast path applies only when the declared name is in bounds and the byte it ends on is genuinely a NUL. That second condition is a soundness requirement rather than a nicety: RawQuerynameKey::cmp passes name.as_ptr() to natural_compare_nul, which scans until it finds a terminator, so a key built from a record whose l_read_name does not land on one would read past the end of its own allocation. Anything failing either condition falls back to the previous strip-and-re-append path, which terminates unconditionally.

Tested with an rstest table checking the new extractor against the pre-optimization implementation as an independent oracle, across well-formed, empty-name, truncated, missing-terminator, non-NUL-terminated, and zero-length records.

perf(sort): size radix passes from the max mapped key, not the unmapped sentinel

Mapped coordinate keys occupy ~5-6 bytes, but unmapped reads carry u64::MAX, which dragged bytes_needed to the full 8 and cost three wasted LSD passes over every record. Coordinate BAMs essentially always carry an unmapped tail, so this was paid on real input almost without exception.

Deriving the bound from the largest non-sentinel key needs no contract from the caller: u64::MAX is the largest u64 and truncates to the all-0xFF maximum at every radix width, so unmapped records still sort to the tail and stay stable among themselves regardless of pass count.

Measured on an otherwise idle c6a.4xlarge over 25 contigs with a 5% unmapped tail, criterion 100 samples, confidence intervals within ±0.5%:

records before after speedup
1M 23.6 M/s 29.0 M/s 1.23x
8M 21.0 M/s 25.0 M/s 1.19x

Radix is roughly a seventh of sort CPU, so this is a low-single-digit percentage of total sort time — worth having, but not a headline number.

An earlier revision of this change also tracked the bound incrementally across pushes so the sort could skip the scan entirely. That measured at a further 1.04x — about half a percent of total sort time — and is deliberately not kept: it cost a running field on RecordBuffer, three reset sites, an ordering hazard where the reset had to precede a macro that returns from inside itself, and a public entry point whose unchecked precondition silently mis-sorts when violated. The scan is three lines and carries no invariant.

Bug fix included: a derived bound of 0 does not imply the keys are all equal, so the sort is never skipped outright. A zero key can coexist with sentinels and the two do not compare equal — PackedCoordinateKey::new packs tid = 0, pos = -1, reverse = false to exactly 0, so a malformed record makes this reachable, and skipping would leave the sentinels ordered ahead of it. A single pass separates them and is a stable no-op when the keys genuinely are identical. Covered by a regression test with zero keys interleaved with sentinels above the radix threshold.

perf(sort): borrow record bytes from the decompressed block on ingest

Pooled ingest copied each record twice after decompression: block into a RawRecord, then RawRecord into the sort arena. PooledInputStream::next_record_borrowed lends a slice straight out of the current decompressed block, falling back to a reusable scratch buffer only when the body or its 4-byte length prefix straddles a block boundary. Copy amplification goes from two to one on the coordinate and template-coordinate paths. The keyed/queryname path keeps owned records, which it requires.

RecordSource grows a matching next_record_borrowed. Both non-pooled variants already yield owned records and have nothing shared to borrow from, so each stores the record it just took and lends into it. That includes the Stream variant, which does not exist upstream — its producer-error handling is shared with the Iterator impl so a failure reaches take_error() identically on both paths. This matters because the ingest loops treat Ok(None) as end-of-input, so an error swallowed there would silently truncate a sort rather than fail it.

Tested with records and length prefixes straddling block boundaries down to 1-byte blocks, parity against the owned read_raw_record path, a proptest over randomized bodies and block sizes, and per-variant agreement plus mid-stream producer-error propagation.

Verification

cargo ci-test (5573 tests), cargo ci-fmt, cargo ci-lint, and RUSTDOCFLAGS="-D warnings" cargo ci-doc all pass. Patch coverage is 100% of changed lines. Each new test was checked for non-vacuity by breaking the corresponding fix and confirming the test fails.

Summary by CodeRabbit

  • Performance
    • Improved coordinate and template sorting efficiency by sorting directly from borrowed BAM bytes.
    • Optimized radix sorting by ignoring unmapped sentinel keys for radix pass sizing; consistent behavior across chunked/parallel sorting.
    • Added coordinate radix sort benchmarks for multiple radix strategies and dataset sizes.
  • Bug Fixes
    • Hardened query name key extraction to properly handle malformed or truncated BAM records while keeping names NUL-terminated.
    • Fixed sorting ingest when decompressed records span decompression block boundaries.
  • Documentation
    • Updated high-performance “unsafe” documentation to clarify the coordinate radix-sorting hot path.

@nh13
nh13 temporarily deployed to github-actions July 21, 2026 00:27 — with GitHub Actions Inactive
@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 4be058c8-a4da-4092-9708-070e7660b266

📥 Commits

Reviewing files that changed from the base of the PR and between eba0243 and f8a371b.

📒 Files selected for processing (9)
  • CLAUDE.md
  • benches/core_functions.rs
  • crates/fgumi-raw-bam/src/lib.rs
  • crates/fgumi-raw-bam/src/raw_bam_record.rs
  • crates/fgumi-sort/src/external.rs
  • crates/fgumi-sort/src/inline.rs
  • crates/fgumi-sort/src/keys.rs
  • crates/fgumi-sort/src/lib.rs
  • crates/fgumi-sort/src/read_ahead.rs

Walkthrough

The PR adds borrow-in-place BAM access, routes coordinate and template ingestion through borrowed bytes, bounds radix passes using mapped keys, hardens queryname extraction, and adds exports, tests, and benchmarks.

Changes

Borrowed ingest and sorting

Layer / File(s) Summary
Borrow-in-place record ingestion
crates/fgumi-raw-bam/..., crates/fgumi-sort/src/read_ahead.rs, crates/fgumi-sort/src/external.rs
Record sources expose borrowed BAM slices, reconstruct block-straddling records through scratch storage, preserve producer errors, and feed coordinate/template buffers without intermediate RawRecord copies.
Sentinel-aware radix sorting
crates/fgumi-sort/src/inline.rs, crates/fgumi-sort/src/lib.rs
Serial, parallel, and chunked radix sorting use mapped-only maximum keys while retaining unmapped u64::MAX records and stable ordering; constructors and sorting functions are re-exported.
Robust queryname key extraction
crates/fgumi-sort/src/keys.rs
Extraction copies valid in-bounds NUL-terminated names directly and falls back safely for malformed records, with malformed-input coverage.
Benchmark and unsafe-region documentation
benches/core_functions.rs, CLAUDE.md
A coordinate radix benchmark compares bound-selection strategies against sort_unstable_by_key, and the documented unsafe region names the bounded radix implementation.

Estimated code review effort: 4 (Complex) | ~60 minutes

Possibly related issues

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title matches the three main performance changes: narrower radix passes, queryname NUL reuse, and borrowed-byte ingest.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch nh/perf-sort-radix-key-borrow

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Jul 21, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.45%. Comparing base (37398d0) to head (f8a371b).
⚠️ Report is 1 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #606      +/-   ##
==========================================
+ Coverage   93.43%   93.45%   +0.02%     
==========================================
  Files         175      175              
  Lines      104981   105405     +424     
==========================================
+ Hits        98089    98511     +422     
- Misses       6892     6894       +2     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@nh13

nh13 commented Jul 21, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 21, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/fgumi-sort/src/inline.rs`:
- Around line 1658-1690: Update bytes_needed in radix_sort_record_refs_with_max
so it increases by one when max_key fills all bits of its currently computed
byte width, preventing mapped all-0xFF keys from tying with the u64::MAX
sentinel. Add a regression covering 0xFFFF-class mapped keys mixed with unmapped
records and verify the sentinel ordering is correct.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 4d37c4da-e0ab-4354-b06c-651b786137e2

📥 Commits

Reviewing files that changed from the base of the PR and between f3b0c78 and eba0243.

📒 Files selected for processing (9)
  • CLAUDE.md
  • benches/core_functions.rs
  • crates/fgumi-raw-bam/src/lib.rs
  • crates/fgumi-raw-bam/src/raw_bam_record.rs
  • crates/fgumi-sort/src/external.rs
  • crates/fgumi-sort/src/inline.rs
  • crates/fgumi-sort/src/keys.rs
  • crates/fgumi-sort/src/lib.rs
  • crates/fgumi-sort/src/read_ahead.rs

Comment thread crates/fgumi-sort/src/inline.rs
nh13 added 3 commits July 21, 2026 02:00
…stores

`extract_queryname_key` stripped the trailing NUL that BAM stores and then
re-appended one into a freshly allocated `Vec` for every record. BAM already
stores the read name NUL-terminated and `l_read_name` counts the terminator, so
`bam[32..32 + l_read_name]` is directly usable as the key bytes.

The fast path only applies when the declared name is in bounds *and* the byte it
ends on is genuinely a NUL. That second condition is a soundness requirement,
not a nicety: `RawQuerynameKey::cmp` passes `name.as_ptr()` to
`natural_compare_nul`, which scans until it finds a terminator, so a key built
from a malformed record whose `l_read_name` does not land on one would read past
the end of its own allocation. Records failing either condition fall back to the
previous strip-and-re-append path, which terminates unconditionally.

Adds an rstest table that checks the extracted key against the pre-optimization
implementation as an independent oracle, covering well-formed, empty-name,
truncated, missing-terminator, non-NUL-terminated, and zero-length records, and
asserts the null-termination invariant on each.
…ed sentinel

Coordinate keys pack `(tid << 34) | ((pos + 1) << 1) | reverse`, so a mapped key
occupies only ~5-6 bytes. Unmapped reads, however, carry `u64::MAX` as their
sort key, which dragged `bytes_needed` to the full 8 and cost three wasted LSD
radix passes over every record. Coordinate-sorted BAMs essentially always have
an unmapped tail, so this was paid on real input almost without exception.

Sizing the passes from the largest *non-sentinel* key fixes it, and needs no
contract from the caller: `u64::MAX` is the largest `u64` and truncates to the
all-`0xFF` maximum at every radix width, so unmapped records still sort to the
tail and stay stable among themselves no matter how many passes run. Both
scanning entry points derive their bound this way, and `par_sort_into_chunks`
derives one bound for the whole buffer and shares it across chunks.

A derived bound of 0 does not imply the keys are all equal, so the sort is never
skipped outright: a zero key can coexist with sentinels, and those two do not
compare equal. `PackedCoordinateKey::new` packs `tid = 0, pos = -1,
reverse = false` to exactly 0, so a malformed record makes this reachable, and
skipping would leave the sentinels ordered ahead of it. A single pass separates
them and is a stable no-op when the keys genuinely are identical.

Measured on an otherwise idle c6a.4xlarge over 25 contigs with a 5% unmapped
tail, criterion 100 samples (confidence intervals within +/-0.5%):

    records   before      after     speedup
    1M        23.6 M/s    29.0 M/s  1.23x
    8M        21.0 M/s    25.0 M/s  1.19x

An earlier revision also tracked the bound incrementally as records were pushed,
letting the sort skip the scan entirely. That was measured at a further 1.04x
(30.2 and 26.0 M/s respectively) -- roughly half a percent of total sort time,
since the radix is about a seventh of it -- and is not kept: it cost a running
field on `RecordBuffer`, three reset sites, an ordering hazard where the reset
had to precede a macro that returns from inside itself, and a public entry point
whose unchecked precondition silently mis-sorts when violated. The scan is three
lines and carries no invariant.

Tests cover sentinel handling through both scanning entry points in serial and
parallel, zero keys interleaved with sentinels above the radix threshold,
`par_sort_into_chunks` on both its single- and multi-threaded drain paths, a
`RecordBuffer` end-to-end mix of mapped and unmapped records, and an
output-identity check that a narrowed sort is byte-for-byte identical to a
full-width one including the stable order among equal keys. The added criterion
benchmark separates pass-count reduction from scan elimination.
Pooled sort ingest copied each record's bytes twice after decompression: once
out of the decompressed block into a `RawRecord`, then again from that
`RawRecord` into the sort arena. The first copy is avoidable whenever the record
body lies wholly within the current block, which is the common case.

Adds `PooledInputStream::next_record_borrowed`, a lending reader that returns a
slice borrowed straight out of `current_buf`, falling back to a reusable scratch
buffer only when the body or its 4-byte length prefix straddles a block
boundary. The coordinate and template-coordinate ingest loops consume borrowed
slices and push them into the buffer directly, taking copy amplification from
two to one. The keyed/queryname path keeps owned records, which it requires.

`RecordSource` grows a matching `next_record_borrowed`. Both non-pooled variants
already yield owned records and so have nothing shared to borrow from: each
stores the record it just took and lends a slice into it. This includes the
`Stream` variant, which is not present upstream -- it gains a held slot
alongside its error slot, and its producer-error handling is shared with the
`Iterator` impl so a failure reaches `take_error()` identically on both paths.
That matters because the ingest loops treat `Ok(None)` as end-of-input, so an
error swallowed there would silently truncate a sort rather than fail it.

Tests cover records and length prefixes straddling block boundaries down to
1-byte blocks, parity against the owned `read_raw_record` path, a proptest over
randomized bodies and block sizes, and -- for the `Stream` variant -- agreement
with owned iteration plus propagation of a mid-stream producer error.
@nh13
nh13 force-pushed the nh/perf-sort-radix-key-borrow branch from eba0243 to f8a371b Compare July 21, 2026 09:03
@nh13
nh13 temporarily deployed to github-actions July 21, 2026 09:03 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 21, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 21, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 21, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 21, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 21, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13
nh13 merged commit 5fc3955 into main Jul 21, 2026
14 checks passed
@nh13
nh13 deleted the nh/perf-sort-radix-key-borrow branch July 21, 2026 14:51
@nh13 nh13 mentioned this pull request Jul 21, 2026

This branch was previously deployed

1 inactive deployment
github-actions — f8a371b4 Deployed Jul 21, 2026 by nh13 via coverage #2843
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant