Repository navigation
fix(sort): scale the per-thread memory budget by the sort phase - #692
Merged
Merged
Conversation
|
Note Reviews pausedUse the following commands to manage reviews:
Use the checkboxes below for quick actions:
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Member
|
@coderabbitai pause |
✅ Action performedReviews paused. |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #692 +/- ##
==========================================
- Coverage 93.95% 93.93% -0.02%
==========================================
Files 178 178
Lines 108086 108109 +23
==========================================
+ Hits 101552 101553 +1
- Misses 6534 6556 +22 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
nh13
force-pushed
the
tf_sort_thread_logging
branch
from
August 2, 2026 00:02
8d95c13 to
6eec80e
Compare
nh13
force-pushed
the
tf_sort_memory_threads
branch
from
August 2, 2026 00:02
e68abb3 to
53ac577
Compare
nh13
approved these changes
Aug 2, 2026
`--memory-per-thread` multiplied `--max-memory` by `--threads` alone. Since
`--sort-threads` defaults to `--threads` rather than the reverse, a run that
set only the per-phase override got the one-thread budget while sorting with
eight workers:
$ fgumi sort -i in.bam -o out.bam --max-memory 100M --sort-threads 8
INFO Max memory: 95.4 MiB (95.4 MiB/thread x 1 threads, from --threads)
The budget sizes the in-memory accumulation buffer, which the sort phase
fills, so it now scales by that phase:
INFO Max memory: 762.9 MiB (95.4 MiB/thread x 8 threads, from --sort-threads)
Take the larger of `--threads` and `--sort-threads` rather than the sort
phase alone. Lowering only the sort phase (`-@ 32 --sort-threads 4`) is the
documented way to cede cores to an upstream producer while keeping the merge
wide; scaling strictly by the sort phase would cut that run's buffer eightfold
and trade a scheduling hint for a throughput cliff. The larger of the two
raises the budget in the case `--threads` under-counts and never lowers it, so
no existing invocation resolves less memory than before.
That makes `--sort-threads` a memory knob as well as a scheduling one, so the
three docs that said otherwise are corrected: its own help no longer claims it
"only changes scheduling", `--threads` is described as the floor for the
multiplier rather than the multiplier, and `--memory-per-thread` names the pair
it scales by. `--merge-threads` is left alone -- the merge phase never reads
`memory_limit`, so that flag really is scheduling-only.
#691 added a `, from --threads` suffix to the `Max memory` line so it could not
be misread against the per-phase `Threads:` line below it. That attribution is
now conditional, since the count can come from either flag, so it reports
whichever one supplied it.
`--threads 0` still reaches `resolve_memory_budget`'s rejection rather than
being clamped, and `--memory-per-thread false` is unaffected. The `Max memory`
log line and the `--max-memory auto` initial-capacity cap both follow the same
count, so neither can disagree with the budget. The auto cap moves into
`Sort::auto_initial_capacity` so it is unit-testable, including its overflow
guard.
Testing: the `Max memory` log line is printed from the same local that feeds
`resolve_memory_budget`, so an integration test asserting on it cannot tell a
correctly wired budget from one that logs the sort-phase count and resolves
`--threads`. `test_memory_budget_threads_resolves_the_scaled_budget` pins the
resolved byte count independently. The sorter-agreement test takes the expected
phase-1 count from its case table instead of re-applying the implementation's
own `max`, which could not have detected an engine fallback resolving below
`--threads`.
nh13
force-pushed
the
tf_sort_memory_threads
branch
from
August 2, 2026 00:12
53ac577 to
a628610
Compare
Merged
Merged
This branch was previously deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #691 — please review/merge that one first; this PR targets its branch, so the diff here is just the memory change.
--memory-per-threadmultiplied--max-memoryby--threadsalone. Since--sort-threadsdefaults to--threadsand not the reverse, a run that set only the per-phase override got the one-thread budget while sorting with eight workers:The budget sizes the in-memory accumulation buffer, which the sort phase fills, so it now scales by that phase:
Why
max(--threads, --sort-threads)and not just--sort-threadsScaling strictly by the sort phase would be a silent regression for the pattern
--sort-threadsexists to serve.-@ 32 --sort-threads 4is the documented way to cede cores to an upstream producer while keeping the merge wide; under sort-phase-only scaling that run's buffer drops from 32× to 4× — an eightfold cut, trading a scheduling hint for a throughput cliff. Taking the larger of the two raises the budget exactly where--threadsunder-counts and never lowers it, so no existing invocation resolves less memory than it does today.Scope
--threads 0still reachesresolve_memory_budget's rejection rather than being clamped up by the new helper.--memory-per-thread falseis unaffected (still a fixed total).Max memorylog line and the--max-memory autoinitial-capacity cap both follow the same count, so neither can disagree with the resolved budget.fgumi sortis touched.QueueMemoryOptions(filter/clip/correct) has a single thread count with no sort/merge phases, so it needs no equivalent — and the README/performance-tuning memory notes describe that pipeline-queue path, not this one.The helper re-derives the sort-phase fallback because the budget is needed before the sorter exists (so it can't read
phase1_threadsoff it, as theThreads:line in #691 does). A test asserts the two definitions agree, so they can't drift.Tests: 7 cases over the multiplier (including both zero cases), 4 drift-guard cases, and two end-to-end tests asserting the logged budget — one for the bug, one pinning the
-@ 32 --sort-threads 4shape against the regression above. I confirmed the end-to-end tests fail before the fix, at95.4 MiB (... x 1 threads).ci-fmt/ci-lintclean.ci-testis 6905 passed / 3 failed — the sametest_input_source_matrix::declared_{sam,stdin}_support_*trio that fails on a pristinemain(they diff R-generatedsimplex_qc.pdfbyte-for-byte and R stamps/CreationDateinto it), as noted in #690 and #691.