Skip to content

fix(sort): honour --compression-level for the tail of a spilling sort - #653

Merged
nh13 merged 1 commit into
mainfrom
nh/fix-sort-spill-tail-compression
Jul 25, 2026
Merged

nh13 merged 1 commit into
mainfrom
nh/fix-sort-spill-tail-compression

Conversation

@nh13

@nh13 nh13 commented Jul 24, 2026 •

Copy link
Copy Markdown
Member

Follow-up to #633. That PR fixed the fully-in-memory sort, which never entered Phase 2 at all and so wrote its entire output at temp_compression. This fixes the remaining leak on the spilling path, where only the tail of the file is affected.

The bug

Phase2Guard::deactivate returned the pool to LEGACY before the output writer was finished, at all four merge sites. But select_priorities only offers CompressOutput in Phase 2 — in LEGACY the sole compress step is CompressSpill, which binds worker.compressor (the temp_compression one). Every output block still sitting in the compress queue at teardown, plus the writer's final flush, was therefore compressed at the temp level, and the tail of the BAM silently ignored --compression-level.

The comments at both merge sites ("workers still need to compress", "CompressOutput is eligible whenever the queue is non-empty") describe the intent correctly, and is_available really does gate CompressOutput on a non-empty queue. What they miss is that in LEGACY the step is never dispatched in the first place, so the eligibility check never gets consulted.

Measured

5,000-pair coordinate sort forced to spill, writing at --compression-level 0:

temp_compression output size
0 2,062,091 bytes
9 2,021,843 bytes

~40 KB (~2%) of the output came out at the temp level. The in-memory path is unaffected — it never builds a Phase2Guard, so there is no early teardown to get wrong, and it produces exactly the 2,062,091-byte level-0 output either way.

The fix

Split the guard's teardown into its two halves:

  • release_sources drops the consumer and unpublishes the spill files — which is what stops workers touching files that are about to be deleted — while leaving the pool in Phase 2, so trailing blocks still route to output_compressor.
  • deactivate returns the pool to LEGACY after the writer is finished, and remains what Drop calls, so error paths are unchanged.

try_phase2_file_work returns immediately on an empty file vector, so output compression is the only Phase 2 work left in that window.

Testing

New integration test test_temp_compression_does_not_reach_the_output_bam sorts twice at output_compression 0, varying only the temp level, and asserts the two outputs are the same size. It fails on all three pooled sort orders before this change:

temp_compression leaked into the output for Coordinate: the same sort wrote 2062091 bytes
at temp_compression=0 and 2021843 bytes at temp_compression=9, but only output_compression
(0 for both) may affect the output

sort_at_level gained a temp_compression parameter so both tests share it; the existing test passes 1 as before.

cargo ci-fmt, cargo ci-lint clean; cargo ci-test 6476 passed / 27 skipped.

Summary by CodeRabbit

  • Bug Fixes
    • Fixed output compression so the final BAM consistently honors the selected output compression level (not temporary/spill settings), including spilling runs.
    • Improved Phase 2 merge finalization to ensure writer/index output is finalized at the correct stage, including when all sources are empty.
  • Refactor
    • Clarified merge teardown so output finalization happens before the worker pool transitions to the earlier phase.
  • Tests
    • Added integration coverage for BGZF compressor routing across spilling and in-memory paths, plus deterministic checks for BAI/index bytes.
    • Added unit tests validating the Phase 2 teardown contract and empty-merge behavior.

@nh13
nh13 temporarily deployed to github-actions July 24, 2026 20:24 — with GitHub Actions Inactive
@coderabbitai

coderabbitai Bot commented Jul 24, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 39a76f6b-a29b-426e-8c7f-3563879d2f41

📥 Commits

Reviewing files that changed from the base of the PR and between ecda4d0 and e0c4f26.

📒 Files selected for processing (3)
  • crates/fgumi-sort/src/external.rs
  • crates/fgumi-sort/src/worker_pool.rs
  • crates/fgumi-sort/tests/integration/test_output_compression.rs

Walkthrough

Risk is incorrect phase ordering during BAM/BAI finalization, which can route output compression through temporary settings. The fix releases merge sources, finalizes output in Phase 2, then deactivates; integration tests compare output and index identity across compression settings.

Changes

Phase 2 merge output lifecycle

Layer / File(s) Summary
Release sources before output finalization
crates/fgumi-sort/src/external.rs
Phase2Guard separates source release from deactivation, and indexed and unindexed merge paths finalize writers while Phase 2 remains active before returning the pool to LEGACY.
Validate teardown invariants
crates/fgumi-sort/src/external.rs, crates/fgumi-sort/src/worker_pool.rs
Tests cover Phase 2 visibility during finalization, cleanup after finalizer failure, empty-source merges, header-only output, complete indexes, and test-only phase inspection.

Compression control and verification

Layer / File(s) Summary
Control and verify compression settings
crates/fgumi-sort/tests/integration/test_output_compression.rs
Integration helpers pass temporary and output compression explicitly, force spilling consistently, and compare BAM sizes, decoded records, BAI bytes, and indexed query results across compression settings.

Estimated code review effort: 3 (Moderate) | ~25 minutes

Possibly related PRs

Sequence Diagram(s)

sequenceDiagram
  participant MergeChunks
  participant Phase2Guard
  participant PooledBamWriter
  participant SortWorkerPool
  MergeChunks->>Phase2Guard: release_sources()
  Phase2Guard->>SortWorkerPool: keep Phase 2 active
  MergeChunks->>PooledBamWriter: finish() or finish_index()
  MergeChunks->>Phase2Guard: deactivate()
  Phase2Guard->>SortWorkerPool: set LEGACY
Loading
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title matches the main change: spilling-sort tail output now honors --compression-level.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch nh/fix-sort-spill-tail-compression

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Jul 24, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 93.68%. Comparing base (1cdbb28) to head (e0c4f26).
⚠️ Report is 2 commits behind head on main.

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #653      +/-   ##
==========================================
+ Coverage   93.63%   93.68%   +0.05%     
==========================================
  Files         175      175              
  Lines      107217   107601     +384     
==========================================
+ Hits       100388   100803     +415     
+ Misses       6829     6798      -31     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@nh13

nh13 commented Jul 24, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 24, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@crates/fgumi-sort/src/external.rs`:
- Around line 3684-3687: The spilling temp-compression integration test does not
cover the indexed finalization path in merge_chunks_with_index. Extend that test
with write_index(true), verify the resulting .bai file exists, and validate that
its virtual offsets resolve correctly after release_sources(), finish_index(),
and deactivate(); otherwise document why the unindexed case is sufficient.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 96615585-b6e3-4d2b-ba78-3deb0fac183c

📥 Commits

Reviewing files that changed from the base of the PR and between b5622b1 and ecda4d0.

📒 Files selected for processing (2)
  • crates/fgumi-sort/src/external.rs
  • crates/fgumi-sort/tests/integration/test_output_compression.rs

Comment thread crates/fgumi-sort/src/external.rs Outdated
@nh13
nh13 force-pushed the nh/fix-sort-spill-tail-compression branch from ecda4d0 to eb0d818 Compare July 25, 2026 00:07
@nh13
nh13 temporarily deployed to github-actions July 25, 2026 00:07 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 25, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 25, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 25, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 25, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 25, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 25, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 25, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 25, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 25, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 25, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

`Phase2Guard::deactivate` returned the pool to `LEGACY` *before* the output
writer was finished, in all four merge sites. But `get_sort_priorities` only
offers `CompressOutput` in Phase 2 — in `LEGACY` the sole compress step is
`CompressSpill`, which binds `worker.compressor` (the `temp_compression` one).
So every output block still sitting in the compress queue at teardown, plus the
writer's final flush, was compressed at the temp level and the tail of the BAM
silently ignored `--compression-level`.

The stale comment at both merge sites ("workers still need to compress",
"CompressOutput is eligible whenever the queue is non-empty") describes the
intent correctly; `is_step_eligible` does gate `CompressOutput` on a non-empty
queue. What it misses is that in `LEGACY` the step is never dispatched in the
first place, so the eligibility check is moot.

Measured on a 5,000-pair coordinate sort forced to spill, writing at
`--compression-level 0`:

    temp_compression 0 -> 2,062,091 bytes
    temp_compression 9 -> 2,021,843 bytes

~40 KB (~2%) of the output came out at the temp level. The in-memory path was
already correct: it never builds a `Phase2Guard`, so it has no early teardown,
and it produces exactly the 2,062,091-byte level-0 output either way. That the
merge path was the odd one out is the argument that this fix, though it does
not touch the underlying coupling, is at a defensible depth.

Give the guard a `finish_output` combinator that owns the whole teardown
order: it releases the drained merge sources, runs the caller's finalizer, and
only then returns the pool to `LEGACY`. Releasing the sources first is what
keeps a slow `finish` from holding the merge's file descriptors and per-file
reorder buffers alive while it drains; finalizing before the phase reset is
what keeps the trailing blocks routed to `output_compressor`. `Drop` still
deactivates, so error paths are unchanged.

The ordering is the whole fix, so it is expressed as one method rather than as
a pair of calls each of the four sites has to sequence correctly — getting it
wrong produces a wrongly-compressed BAM tail, not an error. The two
`initial_keys.is_empty()` branches build their writer inside the closure,
which also closes the window where the writer was constructed outside Phase 2.

The root cause is one level down: which compressor a block gets is decided by
the pool's *global phase* when a worker pops it, rather than by the block's own
origin, even though the single submit site knows which it is. Tagging each
`CompressJob` with its target would retire this ordering contract and the
mirror-image one on `set_phase` (Phase 1 spill jobs still queued when the phase
flips are picked up by `CompressOutput`); that is a separate change, noted in
`finish_output`'s docs.

The new integration test sorts twice at `output_compression 0`, varying only
the temp level, and asserts the two outputs are the same size. It fails on all
three pooled sort orders before this change.

The integration test also gains a `coordinate_indexed` case, so the spilling
merge runs through `merge_chunks_with_index` as well, where the tail block's
level additionally moves the BGZF boundaries the BAI records. That case asserts
the sidecar exists, that an indexed whole-reference query resolves to exactly
the record stream a linear scan of the same file yields, and that the two temp
levels produce byte-identical indexes.

The `initial_keys.is_empty()` early return in each merge finalizes through the
same `finish_output` and so has the same way to go wrong, but Phase 1 checks
its spill trigger only after pushing a record and therefore never writes an
empty chunk — the branch is unreachable from `sort()`. Four unit tests drive
it directly from an empty BGZF chunk: one per merge function, plus two on
`Phase2Guard` itself. The first asserts the phase from *inside* the finalizer,
which is the only place the invariant is observable — asserting either side of
it passes even when the output is finalized in `LEGACY`. The second covers the
error path, pinning the reason `finish_output` binds the finalizer's result
rather than applying `?` to it: the phase reset must not depend on `Drop`.
Reading the phase back needs `SortWorkerPool::current_phase`, a test-only
accessor, since `shared` is private to `worker_pool`.
@nh13
nh13 force-pushed the nh/fix-sort-spill-tail-compression branch from eb0d818 to e0c4f26 Compare July 25, 2026 17:34
@nh13
nh13 temporarily deployed to github-actions July 25, 2026 17:34 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 25, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 25, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13
nh13 merged commit ef0f6f8 into main Jul 25, 2026
14 checks passed
@nh13
nh13 deleted the nh/fix-sort-spill-tail-compression branch July 25, 2026 18:20
nh13 added a commit that referenced this pull request Jul 26, 2026
One `compress_queue` carries both spill and output blocks, and `CompressJob`
recorded only the *codec* (BGZF vs zstd), never which of the two it was. So a
worker popping a job re-derived the compression level from `shared.phase`, a
mutable atomic that advances independently of what is already sitting in the
queue. Every block that outlived a phase transition was compressed at the wrong
level, silently.

That single coupling produced the same bug twice, in opposite directions:

  - Output blocks popped outside Phase 2 went through the spill compressor at
    `temp_compression`, so `--compression-level` was silently discarded. Fixed
    twice at the symptom layer — by entering Phase 2 even for an in-memory sort
    (#633), then by keeping the pool in Phase 2 across the writer's `finish()`
    (#653).
  - Phase 1 spill blocks still queued when `begin_phase2` fired went through the
    output compressor at `output_compression`. This one was never a live bug: it
    was held off only by `drain_pending_spill` happening to sit immediately
    before `enter_output_phase`, and was written down as a caller obligation on
    `begin_phase2` rather than enforced. Its cost is wasted CPU and disk, not
    wrong bytes, since any BGZF level decodes.

The identity was already known for certain at submit time, by the one caller who
cannot be wrong about it: there is a single production submit site,
`StagingBuffer::flush`, and its two constructors are statically distinct — the
output writer and the spill writer. So record it rather than infer it.

`CompressJob` gains a `target: CompressTarget { Spill, Output }` beside the
existing `codec`; `StagingBuffer` carries it and stamps every job it submits;
`try_compress` selects `compressor` or `output_compressor` from `job.target`.
Nothing about the choice reads shared mutable state, so a block is compressed at
the level its writer asked for however long it waited in the queue and whatever
phase the pool is in when a worker gets to it.

This is the second form of that selection. `5df7e176`, which introduced the
worker pool, read it from `shared.phase` at pop time and then narrowed it to the
dispatched `SortStep` to close the window where `set_phase` fired between the pop
and the choice. But a step is itself chosen from the phase, so blocks that
outlived a transition were still mis-compressed — the window was narrowed, not
closed.

With the level travelling on the job, `SortStep::CompressSpill` and
`CompressOutput` did the same thing, so they collapse into one `Compress`. There
was only ever one queue; two steps implied two, and the per-step stats buckets
they fed ("CmpSpl" / "CmpOut") could not be trusted to mean what they said.
`SortStep::COUNT` drops from 5 to 4. `SortStep` is `pub` but `worker_pool` is
`pub(crate)`, so this is not an API change.

Also drops `begin_phase2`'s caller obligation, which no longer exists, and pins
zstd's spill-only invariant now that it is expressible: there is one zstd
compressor per worker, fixed at `temp_compression`, because the output BAM is
always BGZF. That one asserts unconditionally rather than in debug only — an
`Output` job reaching it would be the same silent wrong-level output this
mechanism exists to prevent, and the check is one compare inside the zstd arm,
which the BGZF path never evaluates.

The ordering #653 introduced — release the drained sources, finalize, then
leave Phase 2 — is kept, but for the reason that survives: not holding the
merge's file descriptors and 2 MiB-per-chunk reorder buffers across a slow
`finish`. Its compression rationale is gone, and its docs and tests say so.
Reverting that ordering now leaves `test_temp_compression_does_not_reach_the_
output_bam` passing, which is the check that this commit fixes the cause rather
than a third symptom.

`test_compress_target_decides_level_regardless_of_phase` sweeps the phase across
`LEGACY`/`PHASE1`/`PHASE2` and asserts a `Spill` job stays stored at
`temp_compression = 0` while an `Output` job compresses at `output_compression =
9`. All three cases fail if the compressor is selected from the phase, including
the mirror direction that had no coverage before.

No measurable throughput cost. The job grows by one byte in a struct that
already owns a `Vec<u8>`, and the per-block work is one `match` on a `Copy` enum
replacing one `match` on the step; scheduling is untouched, since the two
collapsed steps shared an eligibility predicate and neither was ever exclusive
(only `ReadInputBlocks` is).

Measured on a spilling coordinate sort of a 3,689,310-record BAM (idt-cfdna,
162 MB, 4 spill chunks, 8 threads, `-m 256m`, output level 6, temp level 1),
10 interleaved reps per binary, baseline being #653's merge commit (`ef0f6f8b`,
identical content in these files):

    CPU (user+sys)  baseline 11.497s ± 0.312   candidate 11.326s ± 0.462   -1.5%
    wall            baseline  5.663s ± 0.941   candidate  5.910s ± 1.044   +4.4%

Welch t on CPU time is -0.97, so the difference is not significant; wall clock
on this host is too noisy (σ ≈ 1 s) to resolve anything under ~15%, which is why
CPU time is the reported metric. All 20 runs exited 0 and wrote exactly
3,689,310 records. Byte-for-byte, the two binaries produce the same output: the
sorted BAMs differ only in the `@PG` `CL:` field (it records the binary's path),
and the record streams have identical checksums.
nh13 added a commit that referenced this pull request Jul 26, 2026
One `compress_queue` carries both spill and output blocks, and `CompressJob`
recorded only the *codec* (BGZF vs zstd), never which of the two it was. So a
worker popping a job re-derived the compression level from `shared.phase`, a
mutable atomic that advances independently of what is already sitting in the
queue. Every block that outlived a phase transition was compressed at the wrong
level, silently.

That single coupling produced the same bug twice, in opposite directions:

  - Output blocks popped outside Phase 2 went through the spill compressor at
    `temp_compression`, so `--compression-level` was silently discarded. Fixed
    twice at the symptom layer — by entering Phase 2 even for an in-memory sort
    (#633), then by keeping the pool in Phase 2 across the writer's `finish()`
    (#653).
  - Phase 1 spill blocks still queued when `begin_phase2` fired went through the
    output compressor at `output_compression`. This one was never a live bug: it
    was held off only by `drain_pending_spill` happening to sit immediately
    before `enter_output_phase`, and was written down as a caller obligation on
    `begin_phase2` rather than enforced. Its cost is wasted CPU and disk, not
    wrong bytes, since any BGZF level decodes.

The identity was already known for certain at submit time, by the one caller who
cannot be wrong about it: there is a single production submit site,
`StagingBuffer::flush`, and its two constructors are statically distinct — the
output writer and the spill writer. So record it rather than infer it.

`CompressJob` gains a `target: CompressTarget { Spill, Output }` beside the
existing `codec`; `StagingBuffer` carries it and stamps every job it submits;
`try_compress` selects `compressor` or `output_compressor` from `job.target`.
Nothing about the choice reads shared mutable state, so a block is compressed at
the level its writer asked for however long it waited in the queue and whatever
phase the pool is in when a worker gets to it.

This is the second form of that selection. `5df7e176`, which introduced the
worker pool, read it from `shared.phase` at pop time and then narrowed it to the
dispatched `SortStep` to close the window where `set_phase` fired between the pop
and the choice. But a step is itself chosen from the phase, so blocks that
outlived a transition were still mis-compressed — the window was narrowed, not
closed.

With the level travelling on the job, `SortStep::CompressSpill` and
`CompressOutput` did the same thing, so they collapse into one `Compress`. There
was only ever one queue; two steps implied two, and the per-step stats buckets
they fed ("CmpSpl" / "CmpOut") could not be trusted to mean what they said.
`SortStep::COUNT` drops from 5 to 4. `SortStep` is `pub` but `worker_pool` is
`pub(crate)`, so this is not an API change.

Also drops `begin_phase2`'s caller obligation, which no longer exists, and pins
zstd's spill-only invariant now that it is expressible: there is one zstd
compressor per worker, fixed at `temp_compression`, because the output BAM is
always BGZF. That one asserts unconditionally rather than in debug only — an
`Output` job reaching it would be the same silent wrong-level output this
mechanism exists to prevent, and the check is one compare inside the zstd arm,
which the BGZF path never evaluates.

The ordering #653 introduced — release the drained sources, finalize, then
leave Phase 2 — is kept, but for the reason that survives: not holding the
merge's file descriptors and 2 MiB-per-chunk reorder buffers across a slow
`finish`. Its compression rationale is gone, and its docs and tests say so.
Reverting that ordering now leaves `test_temp_compression_does_not_reach_the_
output_bam` passing, which is the check that this commit fixes the cause rather
than a third symptom.

`test_compress_target_decides_level_regardless_of_phase` sweeps the phase across
`LEGACY`/`PHASE1`/`PHASE2` and asserts a `Spill` job stays stored at
`temp_compression = 0` while an `Output` job compresses at `output_compression =
9`. All three cases fail if the compressor is selected from the phase, including
the mirror direction that had no coverage before.

No measurable throughput cost. The job grows by one byte in a struct that
already owns a `Vec<u8>`, and the per-block work is one `match` on a `Copy` enum
replacing one `match` on the step; scheduling is untouched, since the two
collapsed steps shared an eligibility predicate and neither was ever exclusive
(only `ReadInputBlocks` is).

Measured on a spilling coordinate sort of a 3,689,310-record BAM (idt-cfdna,
162 MB, 4 spill chunks, 8 threads, `-m 256m`, output level 6, temp level 1),
10 interleaved reps per binary, baseline being #653's merge commit (`ef0f6f8b`,
identical content in these files):

    CPU (user+sys)  baseline 11.497s ± 0.312   candidate 11.326s ± 0.462   -1.5%
    wall            baseline  5.663s ± 0.941   candidate  5.910s ± 1.044   +4.4%

Welch t on CPU time is -0.97, so the difference is not significant; wall clock
on this host is too noisy (σ ≈ 1 s) to resolve anything under ~15%, which is why
CPU time is the reported metric. All 20 runs exited 0 and wrote exactly
3,689,310 records. Byte-for-byte, the two binaries produce the same output: the
sorted BAMs differ only in the `@PG` `CL:` field (it records the binary's path),
and the record streams have identical checksums.
nh13 added a commit that referenced this pull request Jul 26, 2026
…ase (#658)

One `compress_queue` carries both spill and output blocks, and `CompressJob`
recorded only the *codec* (BGZF vs zstd), never which of the two it was. So a
worker popping a job re-derived the compression level from `shared.phase`, a
mutable atomic that advances independently of what is already sitting in the
queue. Every block that outlived a phase transition was compressed at the wrong
level, silently.

That single coupling produced the same bug twice, in opposite directions:

  - Output blocks popped outside Phase 2 went through the spill compressor at
    `temp_compression`, so `--compression-level` was silently discarded. Fixed
    twice at the symptom layer — by entering Phase 2 even for an in-memory sort
    (#633), then by keeping the pool in Phase 2 across the writer's `finish()`
    (#653).
  - Phase 1 spill blocks still queued when `begin_phase2` fired went through the
    output compressor at `output_compression`. This one was never a live bug: it
    was held off only by `drain_pending_spill` happening to sit immediately
    before `enter_output_phase`, and was written down as a caller obligation on
    `begin_phase2` rather than enforced. Its cost is wasted CPU and disk, not
    wrong bytes, since any BGZF level decodes.

The identity was already known for certain at submit time, by the one caller who
cannot be wrong about it: there is a single production submit site,
`StagingBuffer::flush`, and its two constructors are statically distinct — the
output writer and the spill writer. So record it rather than infer it.

`CompressJob` gains a `target: CompressTarget { Spill, Output }` beside the
existing `codec`; `StagingBuffer` carries it and stamps every job it submits;
`try_compress` selects `compressor` or `output_compressor` from `job.target`.
Nothing about the choice reads shared mutable state, so a block is compressed at
the level its writer asked for however long it waited in the queue and whatever
phase the pool is in when a worker gets to it.

This is the second form of that selection. `5df7e176`, which introduced the
worker pool, read it from `shared.phase` at pop time and then narrowed it to the
dispatched `SortStep` to close the window where `set_phase` fired between the pop
and the choice. But a step is itself chosen from the phase, so blocks that
outlived a transition were still mis-compressed — the window was narrowed, not
closed.

With the level travelling on the job, `SortStep::CompressSpill` and
`CompressOutput` did the same thing, so they collapse into one `Compress`. There
was only ever one queue; two steps implied two, and the per-step stats buckets
they fed ("CmpSpl" / "CmpOut") could not be trusted to mean what they said.
`SortStep::COUNT` drops from 5 to 4. `SortStep` is `pub` but `worker_pool` is
`pub(crate)`, so this is not an API change.

Also drops `begin_phase2`'s caller obligation, which no longer exists, and pins
zstd's spill-only invariant now that it is expressible: there is one zstd
compressor per worker, fixed at `temp_compression`, because the output BAM is
always BGZF. That one asserts unconditionally rather than in debug only — an
`Output` job reaching it would be the same silent wrong-level output this
mechanism exists to prevent, and the check is one compare inside the zstd arm,
which the BGZF path never evaluates.

The ordering #653 introduced — release the drained sources, finalize, then
leave Phase 2 — is kept, but for the reason that survives: not holding the
merge's file descriptors and 2 MiB-per-chunk reorder buffers across a slow
`finish`. Its compression rationale is gone, and its docs and tests say so.
Reverting that ordering now leaves `test_temp_compression_does_not_reach_the_
output_bam` passing, which is the check that this commit fixes the cause rather
than a third symptom.

`test_compress_target_decides_level_regardless_of_phase` sweeps the phase across
`LEGACY`/`PHASE1`/`PHASE2` and asserts a `Spill` job stays stored at
`temp_compression = 0` while an `Output` job compresses at `output_compression =
9`. All three cases fail if the compressor is selected from the phase, including
the mirror direction that had no coverage before.

No measurable throughput cost. The job grows by one byte in a struct that
already owns a `Vec<u8>`, and the per-block work is one `match` on a `Copy` enum
replacing one `match` on the step; scheduling is untouched, since the two
collapsed steps shared an eligibility predicate and neither was ever exclusive
(only `ReadInputBlocks` is).

Measured on a spilling coordinate sort of a 3,689,310-record BAM (idt-cfdna,
162 MB, 4 spill chunks, 8 threads, `-m 256m`, output level 6, temp level 1),
10 interleaved reps per binary, baseline being #653's merge commit (`ef0f6f8b`,
identical content in these files):

    CPU (user+sys)  baseline 11.497s ± 0.312   candidate 11.326s ± 0.462   -1.5%
    wall            baseline  5.663s ± 0.941   candidate  5.910s ± 1.044   +4.4%

Welch t on CPU time is -0.97, so the difference is not significant; wall clock
on this host is too noisy (σ ≈ 1 s) to resolve anything under ~15%, which is why
CPU time is the reported metric. All 20 runs exited 0 and wrote exactly
3,689,310 records. Byte-for-byte, the two binaries produce the same output: the
sorted BAMs differ only in the `@PG` `CL:` field (it records the binary's path),
and the record streams have identical checksums.

This branch was previously deployed

1 inactive deployment
github-actions — e0c4f266 Deployed Jul 25, 2026 by nh13 via coverage #3062
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant