Skip to content

feat(bgzf): add slice decompression and caller-driven buffer recycling - #714

Merged
nh13 merged 1 commit into
main-runallfrom
nh/runall-04-bgzf-additive
Aug 5, 2026
Merged

nh13 merged 1 commit into
main-runallfrom
nh/runall-04-bgzf-additive

Conversation

@nh13

@nh13 nh13 commented Aug 5, 2026 •

Copy link
Copy Markdown
Member

Phase 2 of the feat-runall landing series, on top of #697. Three additive entry points in fgumi-bgzf that later phases consume, landed ahead of their callers so the crate arrives complete rather than growing under each subsequent port.

Do not read this as a port of feat-runall's fgumi-bgzf. git diff main-runall feat-runall -- crates/fgumi-bgzf reads +512/−578, but almost all of the deletions are main work that feat-runall predates: header.rs in full (#659's single BGZF header predicate), the Cargo.toml conversion to workspace dependencies, and every CHANGELOG entry from 0.3.1 through 0.5.0. Applying that diff would silently revert #614 and #659. The three additions were hand-ported instead, and BGZF_HEADER_SIZE is taken from header.rs rather than reintroduced as a local constant.

What lands

decompress_into_slice — the fixed-slice analogue of decompress_block_slice_into. It decompresses a block straight into a caller-sized &mut [u8] rather than appending to a Vec, so the sort ingest path can decompress directly into an arena slot instead of into a staging buffer it then copies out of. Same integrity guarantee as its siblings: the payload must exactly fill out and match the footer CRC32, and level-0 input takes the same deflate stored-block fast path.

parse_stored_frame — extracted, not new. copy_stored_and_verify had its framing checks inlined in its body, so the slice variant would have had to duplicate them. Both copy paths now call the one parser, which is what keeps the Vec and slice entry points from drifting apart on what counts as a well-formed stored frame.

uncompressed_size — the accessor that makes decompress_into_slice's contract safe to satisfy. A caller has to size the slot before calling, and ISIZE is a u32 sitting in the file, so sizing straight off the footer means allocating from unvalidated input — up to 4 GiB from a corrupt block. This reads the footer off the same &[u8] (no owned Vec, so no copying the block to read four bytes) and bounds the claim to MAX_UNCOMPRESSED_BLOCK_SIZE. decompress_into_slice resolves the size through it too, so what a caller allocates and what the decompressor accepts cannot diverge.

parse_stored_frame, is_stored_block, deflate_into_slice_and_verify — extracted, not new. copy_stored_and_verify had its framing checks inlined in its body, so the slice variant would have had to duplicate them. The BTYPE dispatch predicate and the inflate-then-verify tail were duplicated for the same reason. All three are now shared with decompress_and_verify, so the Vec and fixed-slice entry points cannot drift on what a well-formed stored frame is or on the exact-fill invariant.

InlineBgzfCompressor::recycle_buffer — lets a take_blocks consumer hand a drained block buffer back to the pool. Only write_blocks_to recycled before, so a consumer driving the compressor with write_all + flush + take_blocks left the pool permanently empty and allocated a fresh output Vec for every block.

pub use libdeflater::Decompressor — so a consumer can name the type every decompress_* entry point takes without declaring its own libdeflater dependency, and so the version it names is necessarily the one this crate decompresses with.

Two things here are not purely additive

Both fall out of the additions rather than being bundled with them, but they are behaviour changes and should be reviewed as such.

write_blocks_to now bounds the pool and survives a write error. Routing it through recycle_buffer means draining into a temporary rather than holding a borrow on completed_blocks, since recycle_buffer takes &mut self. Two consequences. The pool is now capped at MAX_POOLED_BUFFERS; previously this method pushed one buffer per block written with nothing capping the growth, so the "bounded" claim in recycle_buffer's docs would have been false the moment the two paths disagreed. And the write-error path became recoverable, so it is handled: the failing block and its tail are restored to the queue instead of being dropped by the Drain guard. The old for block in self.completed_blocks.drain(..) { output.write_all(&block.data)?; } silently discarded every unwritten block on error, and the error says nothing about how far the write got.

decompress_into_slice validates out.len() against the footer ISIZE up front. The doc stated this precondition; nothing enforced it. Both decompress paths compare against out.len() on the assumption that it is the ISIZE, so a mis-sized slot was reported as a fault in the block — the stored path claimed BGZF stored block ISIZE mismatch: footer=105, LEN=100 for a block whose footer says 100. That names a value the footer does not contain and sends a reader after file corruption that isn't there, when the defect is in the arena that sized the slot.

write_blocks_to's retained blocks are documented as safe to inspect, not to replay. io::Write::write_all loops over write, advancing past each Ok(n), so it can commit part of the failing block before an error surfaces — and the whole block is re-queued, not the unwritten remainder. Writing the queue again would repeat those bytes inside a gzip member and produce a stream no BGZF reader can decode.

Both new functions land without callers

decompress_into_slice gets its first consumer at P3 (arena slots) and P4 (fgumi-pipeline-io's spill_decompress). recycle_buffer's eventual consumer is the pipeline BGZF compress step.

Worth naming explicitly: crates/fgumi-sort/src/worker_pool.rs:2193 is already a take_blocks consumer of exactly the shape recycle_buffer was written for, and it still allocates a fresh Vec per block. It is left alone here because P3 rewrites that file wholesale; wiring it now would be undone in two phases. That gap is sequencing, not an oversight.

Verification

cargo ci-fmt, cargo ci-lint, cargo ci-doc clean; 7554 tests pass, 27 skipped.

Every guard here was mutation-tested rather than assumed — revert the fix, confirm a named test fails:

Mutation Test that fails
Delete the stored-path CRC/size verify case_2_bad_crc_stored
Make the exact-fill check tautological case_3_short_fill
Remove the ISIZE upper bound case_4_isize_above_max
Off-by-one on that bound (> → >=) uncompressed_size_accepts_the_block_maximum
Bypass the shared size accessor case_4_isize_above_max, case_7_too_short_block
Drop the pool capacity bound test_recycle_buffer_refuses_an_oversized_buffer
Pop-and-discard instead of reusing test_recycle_buffer_repopulates_the_pool
Let write_blocks_to bypass the cap test_write_blocks_to_recycles_up_to_the_cap

The first two matter most: both mutations previously left the whole suite green, so neither integrity check this crate advertises was actually pinned.

Two constants were measured, not assumed. MAX_POOLED_BUFFER_BYTES (130560) against real block buffers, which are 65376 bytes at levels 0/6/12 on incompressible input — so no legitimate buffer is refused. And dropping the pop-site clear() is safe because bgzf 0.4's resize_uninit opens with Vec::clear, i.e. the compressor clears unconditionally.

Still outstanding from P1

fgumi-sort does not yet have the test-utils feature. The tracker files it under P1, but it is inert until P3 gates chunk_sorter.rs on it, so it belongs in that phase rather than shipping here as a dead feature flag.

Summary by CodeRabbit

  • New Features

    • Added direct decompression into caller-provided buffers.
    • Added utilities for checking uncompressed block sizes.
    • Exposed decompression and buffer-recycling capabilities for advanced integrations.
    • Added safeguards limiting uncompressed BGZF blocks to 64 KiB.
  • Bug Fixes

    • Improved validation of block sizes, framing, output lengths, and CRC checks.
    • Write failures now preserve queued data for retry.
  • Performance

    • Reuses compression buffers to reduce unnecessary allocations.

@nh13
nh13 temporarily deployed to github-actions August 5, 2026 06:50 — with GitHub Actions Inactive
@coderabbitai

coderabbitai Bot commented Aug 5, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: c8b53269-8ce8-445e-89eb-cec714657a22

📥 Commits

Reviewing files that changed from the base of the PR and between b81ba56 and bf39a45.

📒 Files selected for processing (3)
  • crates/fgumi-bgzf/src/lib.rs
  • crates/fgumi-bgzf/src/reader.rs
  • crates/fgumi-bgzf/src/writer.rs

Walkthrough

The reader adds bounded, fixed-slice BGZF decompression with stored-frame, size, and CRC validation. The writer adds bounded buffer recycling and preserves queued blocks after write failures. Public exports expose the new decompression API and reusable decompressor type.

Changes

BGZF I/O changes

Layer / File(s) Summary
Reader validation and fixed-slice decompression
crates/fgumi-bgzf/src/lib.rs, crates/fgumi-bgzf/src/reader.rs
The reader validates block sizes, stored frames, output lengths, DEFLATE output, and CRC values. It adds decompress_into_slice, uncompressed_size, and public re-exports. Tests cover malformed blocks and boundary conditions.
Writer buffer pooling and write recovery
crates/fgumi-bgzf/src/writer.rs
The writer adds bounded buffer recycling, preserves failed and subsequent queued blocks, retains pooled capacity, and tests reuse, pool limits, and write failures.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related issues

  • fulcrumgenomics/fgumi issue 478 — The reader changes implement its ISIZE bounds and stored-block validation objectives.

Possibly related PRs

Suggested labels: fgumi runall

Sequence Diagram(s)

sequenceDiagram
  participant Caller
  participant decompress_into_slice
  participant Decompressor
  Caller->>decompress_into_slice: provide BGZF block and output slice
  decompress_into_slice->>Decompressor: decompress DEFLATE payload
  Decompressor-->>decompress_into_slice: return decompressed bytes
  decompress_into_slice-->>Caller: return validated byte count
Loading
sequenceDiagram
  participant InlineBgzfCompressor
  participant Output
  participant BufferPool
  InlineBgzfCompressor->>Output: write queued compressed blocks
  Output-->>InlineBgzfCompressor: return write result
  InlineBgzfCompressor->>BufferPool: recycle successful eligible buffers
  InlineBgzfCompressor-->>InlineBgzfCompressor: retain failed and subsequent blocks
Loading
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely identifies the two primary changes: slice decompression and caller-driven buffer recycling.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch nh/runall-04-bgzf-additive

Comment @coderabbitai help to get the list of available commands.

@nh13

nh13 commented Aug 5, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai pause

@coderabbitai

coderabbitai Bot commented Aug 5, 2026

Copy link
Copy Markdown
✅ Action performed

Reviews paused.

@codecov

codecov Bot commented Aug 5, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.97980% with 6 lines in your changes missing coverage. Please review.
⚠️ Please upload report for BASE (main-runall@b81ba56). Learn more about missing BASE report.

Files with missing lines Patch % Lines
crates/fgumi-bgzf/src/writer.rs 96.15% 4 Missing ⚠️
crates/fgumi-bgzf/src/reader.rs 98.96% 2 Missing ⚠️
Additional details and impacted files
@@              Coverage Diff               @@
##             main-runall     #714   +/-   ##
==============================================
  Coverage               ?   93.88%           
==============================================
  Files                  ?      206           
  Lines                  ?   117945           
  Branches               ?        0           
==============================================
  Hits                   ?   110732           
  Misses                 ?     7213           
  Partials               ?        0           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@nh13
nh13 force-pushed the nh/runall-04-bgzf-additive branch from 4000b6a to 86bacc8 Compare August 5, 2026 07:55
@nh13
nh13 temporarily deployed to github-actions August 5, 2026 07:55 — with GitHub Actions Inactive
@nh13

nh13 commented Aug 5, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Aug 5, 2026 •

Copy link
Copy Markdown
⚠️ Action not completed

Review rate limited.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Aug 5, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Aug 5, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/fgumi-bgzf/src/reader.rs`:
- Around line 548-587: Update decompress_into_slice to return Ok(0) immediately
when uncompressed_size is zero, before is_stored_block or DEFLATE dispatch;
preserve the existing output-length validation and nonzero decompression paths.
Add a regression test covering the BGZF_EOF payload with an empty output slice
and asserting successful zero-byte decompression.
🪄 Autofix

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: cd1db323-c182-41c0-bab7-928256e9dd8f

📥 Commits

Reviewing files that changed from the base of the PR and between b81ba56 and 86bacc8.

📒 Files selected for processing (3)
  • crates/fgumi-bgzf/src/lib.rs
  • crates/fgumi-bgzf/src/reader.rs
  • crates/fgumi-bgzf/src/writer.rs

Comment thread crates/fgumi-bgzf/src/reader.rs
Three additive entry points that `fgumi-sort`'s arena engine and the
pipeline BGZF steps need, landed ahead of their consumers so the crate
arrives complete rather than growing under each later port.

`decompress_into_slice` is the fixed-slice analogue of
`decompress_block_slice_into`: it decompresses a block straight into a
caller-sized `&mut [u8]` instead of appending to a `Vec`, which lets the
sort ingest path decompress directly into an arena slot rather than into
a staging buffer it then copies out of. It keeps the integrity guarantee
its siblings provide -- the payload must exactly fill `out` and match the
footer CRC32 -- and takes the same deflate stored-block fast path for
level-0 input.

`uncompressed_size` is the accessor that makes that contract safe to
satisfy. A caller has to size the slot before calling, and ISIZE is a
`u32` sitting in the file, so sizing straight off the footer means
allocating from unvalidated input -- up to 4 GiB from a corrupt block.
This reads the footer off the same `&[u8]` (no owned `Vec`, so no copy of
the block just to read four bytes) and bounds the claim to
`MAX_UNCOMPRESSED_BLOCK_SIZE`. `decompress_into_slice` resolves the size
through it too, so what a caller allocates and what the decompressor
accepts cannot diverge.

The stored fast path is now shared three ways. `copy_stored_and_verify`'s
framing checks were inlined in its body, so the slice variant would have
had to duplicate them; they move to `parse_stored_frame`. The BTYPE
dispatch predicate becomes `is_stored_block`, and the inflate-then-verify
tail becomes `deflate_into_slice_and_verify`, both shared with
`decompress_and_verify`. Extracting only the framing check would have
left the same drift exposure one level up.

`InlineBgzfCompressor::recycle_buffer` lets a `take_blocks` consumer hand
a drained block buffer back to the pool. Only `write_blocks_to` recycled
before, so a consumer driving the compressor with `write_all` + `flush` +
`take_blocks` left the pool permanently empty and allocated a fresh
output `Vec` for every block. `write_blocks_to` now routes through the
same method. The pool is bounded on both axes -- count and buffer
capacity -- because bounding the count alone would let one oversized
`Vec` handed in by a caller sit there for the compressor's lifetime.

Rewriting `write_blocks_to` to call `recycle_buffer` (a `&mut self`
method) means draining into a temporary rather than holding a borrow on
`completed_blocks`. That makes the write-error path recoverable, so it is
handled rather than left as-is: the failing block and its tail are
restored to the queue instead of being dropped by the `Drain` guard. They
are documented as safe to inspect, not to replay -- `write_all` can
commit part of the failing block before erroring, and the whole block is
re-queued, so writing the queue again would repeat those bytes inside a
gzip member.
@nh13
nh13 force-pushed the nh/runall-04-bgzf-additive branch from 86bacc8 to bf39a45 Compare August 5, 2026 18:01
@nh13
nh13 temporarily deployed to github-actions August 5, 2026 18:01 — with GitHub Actions Inactive
@nh13

nh13 commented Aug 5, 2026

Copy link
Copy Markdown
Member Author

Addressed in bf39a451.

decompress_into_slice now returns Ok(0) when the footer's ISIZE is zero, and there are two regression tests: decompress_into_slice_accepts_the_eof_marker and decompress_into_slice_rejects_a_sized_slot_for_an_empty_block.

One correction on the diagnosis, since it changes what the change is for. The finding states that with an empty out, deflate_decompress returns InsufficientSpace, mapped to InvalidData. It does not — it returns Ok(0), and decompress_into_slice(&BGZF_EOF, .., &mut []) already returned Ok(0) before this change. Probed against the built crate:

PROBE isize=Ok(0)
PROBE compressed=[03, 00] is_stored=false
PROBE decompress_into_slice(BGZF_EOF, out=[]) -> Ok(0)
PROBE raw deflate_decompress(payload, &mut []) -> Ok(0)

Removing the new short-circuit leaves all 87 crate tests green, which is the same thing said another way: this is not repairing a failure.

It is still worth taking, for the reason the finding is right about — 03 00 is a fixed-Huffman frame, so is_stored_block is false and the EOF marker does reach libdeflater with a zero-length output slice. Resting on an undocumented edge of the C library for a block every stream ends with is a bad bet across libdeflater versions, and both Vec siblings already short-circuit here, so the three entry points now agree.

One deviation from the suggested placement. The prompt says to return before the stored/DEFLATE dispatch while preserving the output-length validation; those two conflict if taken literally, because the size check sits between them. Placing the short-circuit ahead of the size check makes a caller passing a wrongly-sized slot for a zero-ISIZE block get a silent Ok(0) instead of being told the slot is wrong — the exact misreport that check exists to prevent. It goes after the check instead, and decompress_into_slice_rejects_a_sized_slot_for_an_empty_block fails if it is moved ahead of it (verified by mutation).

@nh13

nh13 commented Aug 5, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Aug 5, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai

coderabbitai Bot commented Aug 5, 2026

Copy link
Copy Markdown

Note

GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer.

@nh13
nh13 merged commit 735bdae into main-runall Aug 5, 2026
14 checks passed
@nh13
nh13 deleted the nh/runall-04-bgzf-additive branch August 5, 2026 18:28
nh13 added a commit that referenced this pull request Aug 6, 2026
#714)

Three additive entry points that `fgumi-sort`'s arena engine and the
pipeline BGZF steps need, landed ahead of their consumers so the crate
arrives complete rather than growing under each later port.

`decompress_into_slice` is the fixed-slice analogue of
`decompress_block_slice_into`: it decompresses a block straight into a
caller-sized `&mut [u8]` instead of appending to a `Vec`, which lets the
sort ingest path decompress directly into an arena slot rather than into
a staging buffer it then copies out of. It keeps the integrity guarantee
its siblings provide -- the payload must exactly fill `out` and match the
footer CRC32 -- and takes the same deflate stored-block fast path for
level-0 input.

`uncompressed_size` is the accessor that makes that contract safe to
satisfy. A caller has to size the slot before calling, and ISIZE is a
`u32` sitting in the file, so sizing straight off the footer means
allocating from unvalidated input -- up to 4 GiB from a corrupt block.
This reads the footer off the same `&[u8]` (no owned `Vec`, so no copy of
the block just to read four bytes) and bounds the claim to
`MAX_UNCOMPRESSED_BLOCK_SIZE`. `decompress_into_slice` resolves the size
through it too, so what a caller allocates and what the decompressor
accepts cannot diverge.

The stored fast path is now shared three ways. `copy_stored_and_verify`'s
framing checks were inlined in its body, so the slice variant would have
had to duplicate them; they move to `parse_stored_frame`. The BTYPE
dispatch predicate becomes `is_stored_block`, and the inflate-then-verify
tail becomes `deflate_into_slice_and_verify`, both shared with
`decompress_and_verify`. Extracting only the framing check would have
left the same drift exposure one level up.

`InlineBgzfCompressor::recycle_buffer` lets a `take_blocks` consumer hand
a drained block buffer back to the pool. Only `write_blocks_to` recycled
before, so a consumer driving the compressor with `write_all` + `flush` +
`take_blocks` left the pool permanently empty and allocated a fresh
output `Vec` for every block. `write_blocks_to` now routes through the
same method. The pool is bounded on both axes -- count and buffer
capacity -- because bounding the count alone would let one oversized
`Vec` handed in by a caller sit there for the compressor's lifetime.

Rewriting `write_blocks_to` to call `recycle_buffer` (a `&mut self`
method) means draining into a temporary rather than holding a borrow on
`completed_blocks`. That makes the write-error path recoverable, so it is
handled rather than left as-is: the failing block and its tail are
restored to the queue instead of being dropped by the `Drain` guard. They
are documented as safe to inspect, not to replay -- `write_all` can
commit part of the failing block before erroring, and the whole block is
re-queued, so writing the queue again would repeat those bytes inside a
gzip member.
nh13 added a commit that referenced this pull request Aug 7, 2026
#714)

Three additive entry points that `fgumi-sort`'s arena engine and the
pipeline BGZF steps need, landed ahead of their consumers so the crate
arrives complete rather than growing under each later port.

`decompress_into_slice` is the fixed-slice analogue of
`decompress_block_slice_into`: it decompresses a block straight into a
caller-sized `&mut [u8]` instead of appending to a `Vec`, which lets the
sort ingest path decompress directly into an arena slot rather than into
a staging buffer it then copies out of. It keeps the integrity guarantee
its siblings provide -- the payload must exactly fill `out` and match the
footer CRC32 -- and takes the same deflate stored-block fast path for
level-0 input.

`uncompressed_size` is the accessor that makes that contract safe to
satisfy. A caller has to size the slot before calling, and ISIZE is a
`u32` sitting in the file, so sizing straight off the footer means
allocating from unvalidated input -- up to 4 GiB from a corrupt block.
This reads the footer off the same `&[u8]` (no owned `Vec`, so no copy of
the block just to read four bytes) and bounds the claim to
`MAX_UNCOMPRESSED_BLOCK_SIZE`. `decompress_into_slice` resolves the size
through it too, so what a caller allocates and what the decompressor
accepts cannot diverge.

The stored fast path is now shared three ways. `copy_stored_and_verify`'s
framing checks were inlined in its body, so the slice variant would have
had to duplicate them; they move to `parse_stored_frame`. The BTYPE
dispatch predicate becomes `is_stored_block`, and the inflate-then-verify
tail becomes `deflate_into_slice_and_verify`, both shared with
`decompress_and_verify`. Extracting only the framing check would have
left the same drift exposure one level up.

`InlineBgzfCompressor::recycle_buffer` lets a `take_blocks` consumer hand
a drained block buffer back to the pool. Only `write_blocks_to` recycled
before, so a consumer driving the compressor with `write_all` + `flush` +
`take_blocks` left the pool permanently empty and allocated a fresh
output `Vec` for every block. `write_blocks_to` now routes through the
same method. The pool is bounded on both axes -- count and buffer
capacity -- because bounding the count alone would let one oversized
`Vec` handed in by a caller sit there for the compressor's lifetime.

Rewriting `write_blocks_to` to call `recycle_buffer` (a `&mut self`
method) means draining into a temporary rather than holding a borrow on
`completed_blocks`. That makes the write-error path recoverable, so it is
handled rather than left as-is: the failing block and its tail are
restored to the queue instead of being dropped by the `Drain` guard. They
are documented as safe to inspect, not to replay -- `write_all` can
commit part of the failing block before erroring, and the whole block is
re-queued, so writing the queue again would repeat those bytes inside a
gzip member.
nh13 added a commit that referenced this pull request Aug 9, 2026
#714)

Three additive entry points that `fgumi-sort`'s arena engine and the
pipeline BGZF steps need, landed ahead of their consumers so the crate
arrives complete rather than growing under each later port.

`decompress_into_slice` is the fixed-slice analogue of
`decompress_block_slice_into`: it decompresses a block straight into a
caller-sized `&mut [u8]` instead of appending to a `Vec`, which lets the
sort ingest path decompress directly into an arena slot rather than into
a staging buffer it then copies out of. It keeps the integrity guarantee
its siblings provide -- the payload must exactly fill `out` and match the
footer CRC32 -- and takes the same deflate stored-block fast path for
level-0 input.

`uncompressed_size` is the accessor that makes that contract safe to
satisfy. A caller has to size the slot before calling, and ISIZE is a
`u32` sitting in the file, so sizing straight off the footer means
allocating from unvalidated input -- up to 4 GiB from a corrupt block.
This reads the footer off the same `&[u8]` (no owned `Vec`, so no copy of
the block just to read four bytes) and bounds the claim to
`MAX_UNCOMPRESSED_BLOCK_SIZE`. `decompress_into_slice` resolves the size
through it too, so what a caller allocates and what the decompressor
accepts cannot diverge.

The stored fast path is now shared three ways. `copy_stored_and_verify`'s
framing checks were inlined in its body, so the slice variant would have
had to duplicate them; they move to `parse_stored_frame`. The BTYPE
dispatch predicate becomes `is_stored_block`, and the inflate-then-verify
tail becomes `deflate_into_slice_and_verify`, both shared with
`decompress_and_verify`. Extracting only the framing check would have
left the same drift exposure one level up.

`InlineBgzfCompressor::recycle_buffer` lets a `take_blocks` consumer hand
a drained block buffer back to the pool. Only `write_blocks_to` recycled
before, so a consumer driving the compressor with `write_all` + `flush` +
`take_blocks` left the pool permanently empty and allocated a fresh
output `Vec` for every block. `write_blocks_to` now routes through the
same method. The pool is bounded on both axes -- count and buffer
capacity -- because bounding the count alone would let one oversized
`Vec` handed in by a caller sit there for the compressor's lifetime.

Rewriting `write_blocks_to` to call `recycle_buffer` (a `&mut self`
method) means draining into a temporary rather than holding a borrow on
`completed_blocks`. That makes the write-error path recoverable, so it is
handled rather than left as-is: the failing block and its tail are
restored to the queue instead of being dropped by the `Drain` guard. They
are documented as safe to inspect, not to replay -- `write_all` can
commit part of the failing block before erroring, and the whole block is
re-queued, so writing the queue again would repeat those bytes inside a
gzip member.
nh13 added a commit that referenced this pull request Aug 10, 2026
#714)

Three additive entry points that `fgumi-sort`'s arena engine and the
pipeline BGZF steps need, landed ahead of their consumers so the crate
arrives complete rather than growing under each later port.

`decompress_into_slice` is the fixed-slice analogue of
`decompress_block_slice_into`: it decompresses a block straight into a
caller-sized `&mut [u8]` instead of appending to a `Vec`, which lets the
sort ingest path decompress directly into an arena slot rather than into
a staging buffer it then copies out of. It keeps the integrity guarantee
its siblings provide -- the payload must exactly fill `out` and match the
footer CRC32 -- and takes the same deflate stored-block fast path for
level-0 input.

`uncompressed_size` is the accessor that makes that contract safe to
satisfy. A caller has to size the slot before calling, and ISIZE is a
`u32` sitting in the file, so sizing straight off the footer means
allocating from unvalidated input -- up to 4 GiB from a corrupt block.
This reads the footer off the same `&[u8]` (no owned `Vec`, so no copy of
the block just to read four bytes) and bounds the claim to
`MAX_UNCOMPRESSED_BLOCK_SIZE`. `decompress_into_slice` resolves the size
through it too, so what a caller allocates and what the decompressor
accepts cannot diverge.

The stored fast path is now shared three ways. `copy_stored_and_verify`'s
framing checks were inlined in its body, so the slice variant would have
had to duplicate them; they move to `parse_stored_frame`. The BTYPE
dispatch predicate becomes `is_stored_block`, and the inflate-then-verify
tail becomes `deflate_into_slice_and_verify`, both shared with
`decompress_and_verify`. Extracting only the framing check would have
left the same drift exposure one level up.

`InlineBgzfCompressor::recycle_buffer` lets a `take_blocks` consumer hand
a drained block buffer back to the pool. Only `write_blocks_to` recycled
before, so a consumer driving the compressor with `write_all` + `flush` +
`take_blocks` left the pool permanently empty and allocated a fresh
output `Vec` for every block. `write_blocks_to` now routes through the
same method. The pool is bounded on both axes -- count and buffer
capacity -- because bounding the count alone would let one oversized
`Vec` handed in by a caller sit there for the compressor's lifetime.

Rewriting `write_blocks_to` to call `recycle_buffer` (a `&mut self`
method) means draining into a temporary rather than holding a borrow on
`completed_blocks`. That makes the write-error path recoverable, so it is
handled rather than left as-is: the failing block and its tail are
restored to the queue instead of being dropped by the `Drain` guard. They
are documented as safe to inspect, not to replay -- `write_all` can
commit part of the failing block before erroring, and the whole block is
re-queued, so writing the queue again would repeat those bytes inside a
gzip member.
nh13 added a commit that referenced this pull request Aug 19, 2026
#714)

Three additive entry points that `fgumi-sort`'s arena engine and the
pipeline BGZF steps need, landed ahead of their consumers so the crate
arrives complete rather than growing under each later port.

`decompress_into_slice` is the fixed-slice analogue of
`decompress_block_slice_into`: it decompresses a block straight into a
caller-sized `&mut [u8]` instead of appending to a `Vec`, which lets the
sort ingest path decompress directly into an arena slot rather than into
a staging buffer it then copies out of. It keeps the integrity guarantee
its siblings provide -- the payload must exactly fill `out` and match the
footer CRC32 -- and takes the same deflate stored-block fast path for
level-0 input.

`uncompressed_size` is the accessor that makes that contract safe to
satisfy. A caller has to size the slot before calling, and ISIZE is a
`u32` sitting in the file, so sizing straight off the footer means
allocating from unvalidated input -- up to 4 GiB from a corrupt block.
This reads the footer off the same `&[u8]` (no owned `Vec`, so no copy of
the block just to read four bytes) and bounds the claim to
`MAX_UNCOMPRESSED_BLOCK_SIZE`. `decompress_into_slice` resolves the size
through it too, so what a caller allocates and what the decompressor
accepts cannot diverge.

The stored fast path is now shared three ways. `copy_stored_and_verify`'s
framing checks were inlined in its body, so the slice variant would have
had to duplicate them; they move to `parse_stored_frame`. The BTYPE
dispatch predicate becomes `is_stored_block`, and the inflate-then-verify
tail becomes `deflate_into_slice_and_verify`, both shared with
`decompress_and_verify`. Extracting only the framing check would have
left the same drift exposure one level up.

`InlineBgzfCompressor::recycle_buffer` lets a `take_blocks` consumer hand
a drained block buffer back to the pool. Only `write_blocks_to` recycled
before, so a consumer driving the compressor with `write_all` + `flush` +
`take_blocks` left the pool permanently empty and allocated a fresh
output `Vec` for every block. `write_blocks_to` now routes through the
same method. The pool is bounded on both axes -- count and buffer
capacity -- because bounding the count alone would let one oversized
`Vec` handed in by a caller sit there for the compressor's lifetime.

Rewriting `write_blocks_to` to call `recycle_buffer` (a `&mut self`
method) means draining into a temporary rather than holding a borrow on
`completed_blocks`. That makes the write-error path recoverable, so it is
handled rather than left as-is: the failing block and its tail are
restored to the queue instead of being dropped by the `Drain` guard. They
are documented as safe to inspect, not to replay -- `write_all` can
commit part of the failing block before erroring, and the whole block is
re-queued, so writing the queue again would repeat those bytes inside a
gzip member.
nh13 added a commit that referenced this pull request Aug 19, 2026
#714)

Three additive entry points that `fgumi-sort`'s arena engine and the
pipeline BGZF steps need, landed ahead of their consumers so the crate
arrives complete rather than growing under each later port.

`decompress_into_slice` is the fixed-slice analogue of
`decompress_block_slice_into`: it decompresses a block straight into a
caller-sized `&mut [u8]` instead of appending to a `Vec`, which lets the
sort ingest path decompress directly into an arena slot rather than into
a staging buffer it then copies out of. It keeps the integrity guarantee
its siblings provide -- the payload must exactly fill `out` and match the
footer CRC32 -- and takes the same deflate stored-block fast path for
level-0 input.

`uncompressed_size` is the accessor that makes that contract safe to
satisfy. A caller has to size the slot before calling, and ISIZE is a
`u32` sitting in the file, so sizing straight off the footer means
allocating from unvalidated input -- up to 4 GiB from a corrupt block.
This reads the footer off the same `&[u8]` (no owned `Vec`, so no copy of
the block just to read four bytes) and bounds the claim to
`MAX_UNCOMPRESSED_BLOCK_SIZE`. `decompress_into_slice` resolves the size
through it too, so what a caller allocates and what the decompressor
accepts cannot diverge.

The stored fast path is now shared three ways. `copy_stored_and_verify`'s
framing checks were inlined in its body, so the slice variant would have
had to duplicate them; they move to `parse_stored_frame`. The BTYPE
dispatch predicate becomes `is_stored_block`, and the inflate-then-verify
tail becomes `deflate_into_slice_and_verify`, both shared with
`decompress_and_verify`. Extracting only the framing check would have
left the same drift exposure one level up.

`InlineBgzfCompressor::recycle_buffer` lets a `take_blocks` consumer hand
a drained block buffer back to the pool. Only `write_blocks_to` recycled
before, so a consumer driving the compressor with `write_all` + `flush` +
`take_blocks` left the pool permanently empty and allocated a fresh
output `Vec` for every block. `write_blocks_to` now routes through the
same method. The pool is bounded on both axes -- count and buffer
capacity -- because bounding the count alone would let one oversized
`Vec` handed in by a caller sit there for the compressor's lifetime.

Rewriting `write_blocks_to` to call `recycle_buffer` (a `&mut self`
method) means draining into a temporary rather than holding a borrow on
`completed_blocks`. That makes the write-error path recoverable, so it is
handled rather than left as-is: the failing block and its tail are
restored to the queue instead of being dropped by the `Drain` guard. They
are documented as safe to inspect, not to replay -- `write_all` can
commit part of the failing block before erroring, and the whole block is
re-queued, so writing the queue again would repeat those bytes inside a
gzip member.
nh13 added a commit that referenced this pull request Aug 23, 2026
#714)

Three additive entry points that `fgumi-sort`'s arena engine and the
pipeline BGZF steps need, landed ahead of their consumers so the crate
arrives complete rather than growing under each later port.

`decompress_into_slice` is the fixed-slice analogue of
`decompress_block_slice_into`: it decompresses a block straight into a
caller-sized `&mut [u8]` instead of appending to a `Vec`, which lets the
sort ingest path decompress directly into an arena slot rather than into
a staging buffer it then copies out of. It keeps the integrity guarantee
its siblings provide -- the payload must exactly fill `out` and match the
footer CRC32 -- and takes the same deflate stored-block fast path for
level-0 input.

`uncompressed_size` is the accessor that makes that contract safe to
satisfy. A caller has to size the slot before calling, and ISIZE is a
`u32` sitting in the file, so sizing straight off the footer means
allocating from unvalidated input -- up to 4 GiB from a corrupt block.
This reads the footer off the same `&[u8]` (no owned `Vec`, so no copy of
the block just to read four bytes) and bounds the claim to
`MAX_UNCOMPRESSED_BLOCK_SIZE`. `decompress_into_slice` resolves the size
through it too, so what a caller allocates and what the decompressor
accepts cannot diverge.

The stored fast path is now shared three ways. `copy_stored_and_verify`'s
framing checks were inlined in its body, so the slice variant would have
had to duplicate them; they move to `parse_stored_frame`. The BTYPE
dispatch predicate becomes `is_stored_block`, and the inflate-then-verify
tail becomes `deflate_into_slice_and_verify`, both shared with
`decompress_and_verify`. Extracting only the framing check would have
left the same drift exposure one level up.

`InlineBgzfCompressor::recycle_buffer` lets a `take_blocks` consumer hand
a drained block buffer back to the pool. Only `write_blocks_to` recycled
before, so a consumer driving the compressor with `write_all` + `flush` +
`take_blocks` left the pool permanently empty and allocated a fresh
output `Vec` for every block. `write_blocks_to` now routes through the
same method. The pool is bounded on both axes -- count and buffer
capacity -- because bounding the count alone would let one oversized
`Vec` handed in by a caller sit there for the compressor's lifetime.

Rewriting `write_blocks_to` to call `recycle_buffer` (a `&mut self`
method) means draining into a temporary rather than holding a borrow on
`completed_blocks`. That makes the write-error path recoverable, so it is
handled rather than left as-is: the failing block and its tail are
restored to the queue instead of being dropped by the `Drain` guard. They
are documented as safe to inspect, not to replay -- `write_all` can
commit part of the failing block before erroring, and the whole block is
re-queued, so writing the queue again would repeat those bytes inside a
gzip member.

This branch was previously deployed

1 inactive deployment
github-actions — bf39a451 Deployed Aug 5, 2026 by nh13 via coverage #3290
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant