Skip to content

perf(sort): borrow record bytes from the decompressed block on ingest - #437

Merged
nh13 merged 1 commit into
feat-runallfrom
feat-runall-36-sort-ingest-borrow
Jun 18, 2026
Merged

nh13 merged 1 commit into
feat-runallfrom
feat-runall-36-sort-ingest-borrow

Conversation

@nh13

@nh13 nh13 commented Jun 18, 2026

Copy link
Copy Markdown
Member

Summary

On the pooled sort ingest path each record's bytes were copied twice after decompression:

decompress ──► current_buf        (libdeflate output — unavoidable)
current_buf ──► RawRecord         (copy #1 — removable)
RawRecord  ──► SegmentedBuf arena (copy #2 — sort must own the bytes)

This removes copy #1 on the common path. A new PooledInputStream::next_record_borrowed returns a slice borrowed directly out of current_buf when the record body lies wholly within the current decompressed block, falling back to a reusable scratch buffer only when the record body or its 4-byte length prefix straddles a block boundary. RecordSource grows a matching next_record_borrowed; the coordinate and template-coordinate ingest loops now consume borrowed slices and push them straight into the buffer (copy amplification 2 → 1). The keyed/queryname path keeps owned RawRecords, which it must retain for Vec<(K, RawRecord)>.

Input-read memmove measured ~8% of sort CPU on a Time Profiler of fgumi sort; this drops one of the two record-body copies (PooledInputStream path only — it does not touch the ~63% that is BGZF de/compression).

Correctness

This is a core ingest-path refactor, so correctness was prioritized over speed:

  • Byte-identical output verified by an A/B over the 8.2M-record CODEC BAM (coordinate, -m 8GiB, --threads 8): the sorted record streams compare equal (samtools view | cmp). The only .bam-level difference is BGZF block boundaries, which are already non-deterministic under multithreaded compression on feat-runall itself.
  • Parity unit test against the owned read_raw_record path over identically-chunked streams.
  • Dedicated straddle tests: records and 4-byte length prefixes straddling decompressed-block boundaries across many block sizes (down to 1 byte each), plus exact-boundary, empty-stream, and truncated-body cases.
  • No new unsafe — the fast path is a slice borrow and the slow path is a read_exact into reused scratch.

Validation

  • cargo build --features compare,simulate,profile-adjacency
  • cargo ci-lint (clippy pedantic, -D warnings) — clean
  • cargo ci-fmt — clean
  • cargo nextest -p fgumi-sort — 476 passed (6 new next_record_borrowed tests)
  • cargo nextest --features compare,simulate,profile-adjacency -E 'test(sort) or test(runall) or test(streaming)' — 296 passed

Closes #436

@nh13
nh13 temporarily deployed to github-actions June 18, 2026 07:45 — with GitHub Actions Inactive
@codecov

codecov Bot commented Jun 18, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
⚠️ Please upload report for BASE (feat-runall@b0d7229). Learn more about missing BASE report.

Additional details and impacted files
@@              Coverage Diff               @@
##             feat-runall     #437   +/-   ##
==============================================
  Coverage               ?   93.39%           
==============================================
  Files                  ?      139           
  Lines                  ?    54817           
  Branches               ?        0           
==============================================
  Hits                   ?    51199           
  Misses                 ?     3618           
  Partials               ?        0           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@nh13

nh13 commented Jun 18, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jun 18, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai

coderabbitai Bot commented Jun 18, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 14da2745-8824-457b-aa67-197422676e21

📥 Commits

Reviewing files that changed from the base of the PR and between 41bc769 and 6a241a7.

📒 Files selected for processing (4)
  • crates/fgumi-raw-bam/src/lib.rs
  • crates/fgumi-raw-bam/src/raw_bam_record.rs
  • crates/fgumi-sort/src/external.rs
  • crates/fgumi-sort/src/read_ahead.rs
✅ Files skipped from review due to trivial changes (1)
  • crates/fgumi-raw-bam/src/lib.rs
🚧 Files skipped from review as they are similar to previous changes (3)
  • crates/fgumi-raw-bam/src/raw_bam_record.rs
  • crates/fgumi-sort/src/external.rs
  • crates/fgumi-sort/src/read_ahead.rs

📝 Walkthrough

Walkthrough

read_block_size in fgumi-raw-bam is promoted from private to pub with expanded docs covering EOF, block-boundary straddle, and error conditions. PooledInputStream gains a reusable scratch: RawRecord field and next_record_borrowed method, which reads the 4-byte framing prefix (inline if buffered, else via read_block_size across block boundaries), then either borrows the body directly from the decompressed block or reassembles straddling records into scratch and returns a slice. RecordSource's ReadAhead variant is extended to hold Option<RawRecord> to lend borrowed slices. Both sort phase-1 ingest loops switch from for record in record_source.by_ref() to while let Some(bam_bytes) = record_source.next_record_borrowed()?. Tests validate fast/slow paths, block-boundary straddles, parity against the owned reader, and truncated-body error handling.

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed Title clearly summarizes the main optimization: borrowing record bytes from decompressed blocks to eliminate intermediate copies during ingest.
Description check ✅ Passed Description comprehensively explains the problem (double-copy pattern), solution (lending API), correctness validation, and test coverage, all aligned with the changeset.
Linked Issues check ✅ Passed All requirements from issue #436 are met: lending API implemented via next_record_borrowed, block-straddle fallback to scratch buffer, byte-identical output verified, dedicated straddle tests added, no unsafe code introduced.
Out of Scope Changes check ✅ Passed All changes remain scoped to the optimization: read_block_size publicized for the API, PooledInputStream and RecordSource extended with lending methods, ingest loops refactored to consume borrowed slices, keyed path unchanged.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat-runall-36-sort-ingest-borrow

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🧹 Nitpick comments (1)
crates/fgumi-sort/src/read_ahead.rs (1)

660-814: 🏗️ Heavy lift

Add a proptest parity/property test for borrowed ingest.

Given boundary-heavy logic (prefix/body straddles), add randomized property checks (record sizes/content/block sizes) alongside the fixed examples.

As per coding guidelines: Use proptest for property-based testing in Rust test files.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@crates/fgumi-sort/src/read_ahead.rs` around lines 660 - 814, Add a
`proptest`-based property test that generates randomized record bodies and block
sizes to complement the fixed example tests. Create a new test function (similar
to test_next_record_borrowed_parity_with_read_record) that uses proptest
strategies to randomly generate various record sizes, content patterns, and
block lengths, then verify that the borrowed ingest path produces byte-identical
results to either the owned read_record path or the input records. This should
test a wider range of boundary conditions than the current fixed test cases like
test_next_record_borrowed_prefix_straddle and
test_next_record_borrowed_body_exact_boundary_and_straddle.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/fgumi-sort/src/read_ahead.rs`:
- Around line 733-742: The test function
test_next_record_borrowed_matches_input_across_block_sizes uses a for loop to
iterate over different block_len values instead of using rstest for
parameterization. Refactor this test by adding the #[rstest] attribute macro and
converting the for loop into a parameterized block_len parameter using
#[values(...)] with the array of test values (1usize, 2, 3, 4, 5, 6, 7, 8, 16,
64, 1024, 65_535). Remove the for loop and make block_len a direct parameter of
the test function. Apply the same refactoring pattern to the other test
mentioned at lines 772-794.

---

Nitpick comments:
In `@crates/fgumi-sort/src/read_ahead.rs`:
- Around line 660-814: Add a `proptest`-based property test that generates
randomized record bodies and block sizes to complement the fixed example tests.
Create a new test function (similar to
test_next_record_borrowed_parity_with_read_record) that uses proptest strategies
to randomly generate various record sizes, content patterns, and block lengths,
then verify that the borrowed ingest path produces byte-identical results to
either the owned read_record path or the input records. This should test a wider
range of boundary conditions than the current fixed test cases like
test_next_record_borrowed_prefix_straddle and
test_next_record_borrowed_body_exact_boundary_and_straddle.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro

Run ID: 08b6d256-f867-47dc-8709-a63eb46d4df1

📥 Commits

Reviewing files that changed from the base of the PR and between b0d7229 and 41bc769.

📒 Files selected for processing (4)
  • crates/fgumi-raw-bam/src/lib.rs
  • crates/fgumi-raw-bam/src/raw_bam_record.rs
  • crates/fgumi-sort/src/external.rs
  • crates/fgumi-sort/src/read_ahead.rs

Comment thread crates/fgumi-sort/src/read_ahead.rs Outdated
On the pooled sort ingest path each record's bytes were copied twice after
decompression: once from the decompressed block into a `RawRecord` (via
`read_exact`) and again from the `RawRecord` into the sort arena. The first
copy is removable when the record body lies wholly within the current
decompressed block, which is the common case.

Add `PooledInputStream::next_record_borrowed`, a lending reader that returns a
slice borrowed directly out of `current_buf` on the fast path, falling back to
a reusable scratch buffer only when the record body or its 4-byte length prefix
straddles a decompressed-block boundary. `RecordSource` grows a matching
`next_record_borrowed` (the `ReadAhead` thread variant lends the owned record it
just received). The coordinate and template-coordinate ingest loops now consume
borrowed slices and push them straight into the buffer, dropping the intermediate
`RawRecord` copy (copy amplification 2 -> 1; input-read memmove was ~8% of sort
CPU). The keyed/queryname path keeps owned records, which it must retain.

Byte-identical output verified by an A/B over a 8.2M-record BAM (sorted record
streams compare equal) and by a parity unit test against the owned
`read_raw_record` path. Adds dedicated tests for records and length prefixes
straddling block boundaries across many block sizes (down to 1 byte).

Closes #436
@nh13
nh13 force-pushed the feat-runall-36-sort-ingest-borrow branch from 41bc769 to 6a241a7 Compare June 18, 2026 16:37
@nh13
nh13 temporarily deployed to github-actions June 18, 2026 16:38 — with GitHub Actions Inactive
@nh13

nh13 commented Jun 18, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jun 18, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13
nh13 merged commit de51ec9 into feat-runall Jun 18, 2026
10 checks passed
@nh13
nh13 deleted the feat-runall-36-sort-ingest-borrow branch June 18, 2026 16:56
nh13 added a commit that referenced this pull request Jun 18, 2026
…#437)

On the pooled sort ingest path each record's bytes were copied twice after
decompression: once from the decompressed block into a `RawRecord` (via
`read_exact`) and again from the `RawRecord` into the sort arena. The first
copy is removable when the record body lies wholly within the current
decompressed block, which is the common case.

Add `PooledInputStream::next_record_borrowed`, a lending reader that returns a
slice borrowed directly out of `current_buf` on the fast path, falling back to
a reusable scratch buffer only when the record body or its 4-byte length prefix
straddles a decompressed-block boundary. `RecordSource` grows a matching
`next_record_borrowed` (the `ReadAhead` thread variant lends the owned record it
just received). The coordinate and template-coordinate ingest loops now consume
borrowed slices and push them straight into the buffer, dropping the intermediate
`RawRecord` copy (copy amplification 2 -> 1; input-read memmove was ~8% of sort
CPU). The keyed/queryname path keeps owned records, which it must retain.

Byte-identical output verified by an A/B over a 8.2M-record BAM (sorted record
streams compare equal) and by a parity unit test against the owned
`read_raw_record` path. Adds dedicated tests for records and length prefixes
straddling block boundaries across many block sizes (down to 1 byte).

Closes #436

This branch was previously deployed

1 inactive deployment
github-actions — 6a241a7a Deployed Jun 18, 2026 by nh13 via coverage #1690
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant