perf(compare): make compare bams faster than the sort it validates - #696
Conversation
|
Note Reviews pausedUse the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (12)
Walkthrough
ChangesBAM comparison pipeline
Estimated code review effort: 4 (Complex) | ~45 minutes Sequence Diagram(s)sequenceDiagram
participant CompareBams
participant OpenedInput
participant CheckedRecords
participant PositionalCompare
participant OrderCheck
CompareBams->>OpenedInput: open_pair(bam1, bam2, threads)
OpenedInput->>CheckedRecords: create checked record streams
CheckedRecords-->>PositionalCompare: yield records or labeled read errors
PositionalCompare->>OrderCheck: observe records
PositionalCompare-->>CompareBams: comparison result and order diagnostics
Possibly related issues
Possibly related PRs
Suggested labels: 🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
Comment |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #696 +/- ##
==========================================
+ Coverage 93.93% 94.06% +0.13%
==========================================
Files 178 178
Lines 108128 108601 +473
==========================================
+ Hits 101565 102152 +587
+ Misses 6563 6449 -114 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
@coderabbitai pause |
✅ Action performedReviews paused. |
|
Ran this branch at production scale on NumbersBoth arms
Three comparisons ran on the after side, all at 1.33B records:
#686's acceptance criteria at this scale
Met, and not narrowly. On the coordinate cell the Caveats, since these are not an interleaved A/BThe two rows are different datasets and different record counts — the before is a 1000 Genomes HG00096 pair from an earlier run, the after is GIAB HG002 60x. I have normalized to items/s, and the instance type, thread count and storage configuration are identical, but this is not the same-inputs comparison your workstation numbers are. Treat 5.84x as "the effect is the same order at production scale on the target hardware", not as a precise replication. I also did not sweep AlsoThe reader change held up on inputs where per-record cost is high: the template-coordinate cells run over a BAM with 77.3M both-ends-unmapped records, and the only thing that stopped one of those four comparisons was the pending-window abort that #700 fixes — not throughput or memory. |
7585aba to
1399b02
Compare
|
@coderabbitai review |
✅ Action performedReview finished.
|
There was a problem hiding this comment.
Actionable comments posted: 3
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@src/lib/commands/compare/engines/mod.rs`:
- Around line 67-80: Update the read-ahead flow used by CheckedRecords and
RawReadAheadReader::next_record so a producer-channel disconnect before the EOF
sentinel is represented as an error rather than None. Track explicit producer
completion or propagate the disconnect failure, and surface it from
CheckedRecords with the existing reader context, matching positional_compare’s
“reader disconnected before EOF” behavior while preserving normal sentinel-based
EOF.
In `@src/lib/commands/compare/engines/sort_verify.rs`:
- Around line 1508-1512: Add a paired rejection test case for
SortOrder::TemplateCoordinate in the sort verification table, using the existing
mapped helper with records whose positions descend so core_cmp is Less and one
violation is expected. Keep the existing template_coordinate_accepts_equal_keys
case and follow the naming and structure of the other order-specific reject
cases.
In `@tests/integration/test_compare_bams.rs`:
- Around line 3189-3192: Update the damaged_bam1_while_pairing and
damaged_bam2_while_pairing cases in the test case definitions to use an intact
record count above the truncated BAM’s 20,000-record failure point, such as
20_000; leave the while_draining cases at 1.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 3f856848-8ee9-4327-93bf-40bde9375ab3
📒 Files selected for processing (12)
.gitignoreCLAUDE.mdcrates/fgumi-sort/src/lib.rscrates/fgumi-sort/src/read_ahead.rscrates/fgumi-sort/src/verify.rssrc/lib/commands/compare/bams.rssrc/lib/commands/compare/engines/mod.rssrc/lib/commands/compare/engines/positional.rssrc/lib/commands/compare/engines/sort_verify.rssrc/lib/commands/compare/molecule.rstests/integration/test_compare_bams.rstests/integration/test_compare_mutation.rs
1399b02 to
f6a7f76
Compare
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
Note GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer. |
There was a problem hiding this comment.
Actionable comments posted: 1
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@src/lib/commands/compare/molecule.rs`:
- Around line 152-161: Update the read-error arm in the iterator implementation
to clear or discard pending before returning Some(Err(e.into())). Ensure
subsequent polling cannot emit the buffered partial molecule run, preserving the
invariant documented by the relevant test while leaving successful end-of-stream
handling unchanged.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: ASSERTIVE
Plan: Pro
Run ID: 2599d700-adf0-45bd-90f3-99caff897ec0
📒 Files selected for processing (12)
.gitignoreCLAUDE.mdcrates/fgumi-sort/src/lib.rscrates/fgumi-sort/src/read_ahead.rscrates/fgumi-sort/src/verify.rssrc/lib/commands/compare/bams.rssrc/lib/commands/compare/engines/mod.rssrc/lib/commands/compare/engines/positional.rssrc/lib/commands/compare/engines/sort_verify.rssrc/lib/commands/compare/molecule.rstests/integration/test_compare_bams.rstests/integration/test_compare_mutation.rs
`compare bams --command sort` and `--command group` decoded both inputs inline on the single main thread: `OpenedInput::open` took no thread count, so `--threads` created no threads at all and the whole comparison ran on one core. Profiling a 20.1M-record pair put ~61% of that one thread in BGZF decode (~55% libdeflate inflate, ~6% the per-block CRC32) while eleven cores sat idle; tricorder confirmed it independently with mean_load=98% and a peak OS thread count of 1. Open both inputs through the multithreaded BGZF reader that content mode has always used, wrapped in a per-input read-ahead thread, so each input's decode overlaps the other's and the comparison itself. On a 12-core host that takes sort from 23.44s to 8.93s and grouping from 24.36s to 9.92s (2.6x and 2.5x, interleaved arms with the page cache warmed identically before every run), raising mean_load to 370%. Split the `--threads` budget between the two inputs rather than giving each the full count: they are symmetric and both must be fully decoded, so handing each the whole budget would request twice the requested parallelism and oversubscribe the host. This differs from zipper, which gives its unmapped input 1 thread and its mapped input the full count because its two inputs differ in cost. The read-ahead reader signals a read failure by ending iteration and parking the error for a later `take_error()`, which is a hazard here: a truncated or CRC-corrupt BAM would read as a *shorter* file, reported as a record-count DIFFER, or as a false IDENTICAL when both inputs are damaged alike. Rather than requiring every engine loop to remember a trailing check, `CheckedRecords` restores the `Result` contract at the boundary so the engines' existing `rec?` propagation stays correct by construction. Generalize `verify_sort_order` and `molecule_runs` from a concrete reader type to any `Iterator<Item = io::Result<RawRecord>>` so both reader stacks satisfy them; the `Result` item type is deliberate, keeping a read failure distinguishable from end-of-stream. Verified against a 121-comparison differential corpus spanning matched, missing, extra, edited, reordered, truncated, header-conflicting and corrupt inputs: all verdicts, exit codes and RESULT lines are byte-identical to the pre-change binary. Error text changed only on damaged inputs, where the new messages name the failing file and diagnose it more precisely. This does not yet close #686. The gain is from overlap, not from `--threads`: scaling stays flat (-t 1 8.64s, -t 8 8.89s) because the comparison thread is now the critical path, balanced against ~8.0s of per-file decode. `fgumi sort` over the same data is 8.82s, so comparison went from 2.7x slower than the sort it validates to roughly par, but not strictly faster. Cutting comparison-thread work is the next step, after which the BGZF parallelism wired up here should begin to matter. Refs #686
… keys Profiling `--command sort` after the reader work put ~32% of the critical-path thread in the canonical content key that `RunCanceller::observe` builds for every record: `hash_one` 10.2%, `RunCanceller::observe` 11.1%, allocator traffic 6.2%, `collect_tag_entries` 3.2%, and the key's own tag sort 1.2%. Each record costs a `Vec` per aux tag, a sort, a concatenation, and an `AHashMap` hash plus insert. In the overwhelmingly common case — two files that agree, record for record — none of that is needed. Compare the raw bytes of the two records first and, when they are identical, count the pair and move on. This is sound rather than merely convenient. The run comparison decides on the symmetric difference of the two sides' record multisets, and removing one element from each side *with the same key* leaves that symmetric difference unchanged, regardless of what else is pending — so the fast path needs no precondition on the canceller's state. Byte equality implies content-key equality, and strictly so: `content_key_exact` excludes `bin`, width-normalizes integer tags and sorts the tag multiset, so byte-identical records necessarily share a key. The converse does not hold, which is exactly why this is a fast path and not a replacement: records that are content-equal but differ in tag order, integer tag width or `bin` fall through to the key path and still cancel there. On a 20.1M-record pair, `--command sort` goes from 8.93s to 4.14s at `--threads 8` (warm cache, 3 reps, minimum), 5.7x against the pre-campaign 23.44s. With the comparison thread no longer the longest leg, the BGZF parallelism wired up previously now pays: 8.92s at `-t 1`, 5.58s at `-t 4`, 4.14s at `-t 8`. Both of #686's acceptance criteria are now met: comparison is 2.0x faster than the `fgumi sort` it validates over the same data (4.14s vs 8.20s), having started 2.7x slower, and throughput scales measurably from `--threads 1` to `--threads 8`. Verdicts are unchanged: all 121 comparisons in the differential corpus — matched, missing, extra, edited, tag-reordered, intra-run reordered, mis-sorted, truncated, header-conflicting and corrupt inputs, in both directions — report identical verdicts, exit codes and RESULT lines. The `tag-reorder` cases are the ones that matter here: they are content-equal but not byte-equal, so they prove the key path still runs and still matches. Grouping mode is untouched by this change; it pairs molecules through `molecule_join` rather than this engine, and remains at ~10s. Closes #686
`content` mode read each input twice. `CompareBams::execute` called
`verify_records_in_order` on both paths — a complete traversal of each file — and
only then did `positional_compare` open both and stream them again: four full BAM
traversals for a comparison that needs two. `sort_verify` had already solved this
for `--command sort` with its `OrderChecked` adapter; `content` never adopted it.
Fold the check into the comparison's own pass. `OrderCheck` extracts each record's
sort key and tracks monotonicity as records flow by, and `positional_compare` feeds
every record it receives through its file's checker — including the ones drained
after pairing stops, which is what keeps the totals equal to a dedicated pass.
The check accumulates rather than failing on the first bad record because the
diagnostic reports *how many* records violated the order, which is only known once
the file has been fully seen. Evaluating bam1's result before bam2's preserves
which file a mis-sorted pair names. The resulting message is byte-identical to the
one the standalone pass produced, count and first-violation position included.
Read errors from the comparison readers now name the offending input rather than
its positional slot ("BAM1"), matching how the sort and grouping engines already
report them; a two-input tool that says only "BAM1" leaves the reader to map that
back to a path.
Total CPU for a 20.1M-record content comparison drops from 47.1s to 34.5s (-27%),
which is the decode work the two removed traversals were doing.
`verify_records_in_order` had no callers left and is removed; `OrderCheck`
subsumes it, and `fgumi sort --verify` continues to use
`fgumi_sort::verify_sort_order` directly.
All 121 comparisons in the differential corpus report identical verdicts, exit
codes and RESULT lines, including the mis-sorted and corrupt inputs that exercise
this path.
…pairs Content mode paired every record through `record_keys_match` — extracting flags, testing secondary/supplementary, and walking both read names — before handing the pair to `content_diffs`, whose very first act is to compare the two records byte for byte and return early when they match. For the common case of two files that agree, all of that key work was performed only to confirm what a `memcmp` was about to establish anyway. Compare the bytes once, up front, and skip both checks when they are equal. A `RecordKey` is a pure function of a record's bytes, so identical bytes yield identical keys; and byte equality implies content equality under every `ContentPredicate`, which is exactly why `content_diffs` opens with that test. Records that differ take the original path unchanged, so a genuine desync is still reported as a key mismatch rather than a content diff. Also split the decompression budget across content mode's two readers, matching the convention the sort and grouping engines already follow. `--threads 8` was starting 8 BGZF workers per reader — 16 in total, plus two reader threads and the main thread — which oversubscribes a smaller host without decoding any faster. Total CPU for a 20.1M-record content comparison falls from 34.5s to 27.4s, and to 27.4s from 47.1s before the traversal work — a 42% reduction overall. Also corrects a comment that still described content mode as making two passes over each input; the order-verification pass it referred to is now folded into the comparison pass. All 121 comparisons in the differential corpus report identical verdicts, exit codes and RESULT lines. The `tag-reorder` and `swap-adjacent` cases matter most here: both are non-byte-equal, so they exercise the slow path and confirm the fast path has not swallowed the checks.
`get_mi_tag_raw` probed `find_int_tag` and then fell back to `find_string_tag`, and each of those walks the record's aux data from the start. `fgumi group` writes MI as `MI:Z:<id>[/A|/B]`, so the integer probe always missed and its full scan was wasted, and `MI` is appended late in the tag block, so both walks covered nearly all of it. This runs once per record on the grouping engine's critical path, where profiling attributed 7.1% of CPU to `find_tag_position` and a further 1.9% to `get_mi_tag_raw` itself. Use the existing `RawTagsView::get`, which resolves a tag to a typed `TagValue` in a single `find_tag_position` call, and branch on the type. No new API: the zero-copy typed accessor was already there. Behaviour is unchanged, including the cases the two-probe form handled implicitly. An integer MI still yields `MiKey::Int`, a `Z` payload is still parsed for the `<id>` and `<id>/A|B` forms, and any other aux type is still treated as "no MI" — previously by both probes failing, now by an explicit match arm. Negative integer ids remain accepted, which is why this does not reuse `fgumi_raw_bam::find_mi_tag`: that helper collapses `MI:i:` and `MI:Z:.../A` into one representation and rejects negatives, both of which `MiKey` deliberately keeps distinct. All 121 comparisons in the differential corpus are unchanged, including the `mi-renumber` case that exercises MI parsing directly.
Captures what the `compare bams` optimization work established about measuring this codebase, so the next campaign does not rediscover it. The profiling ladder, in the order worth reaching for: `tricorder` for core utilization and I/O (its `mean_load` and the `--trace` `n_threads` column answer "is this actually parallel" with no symbolication to get wrong), then in-tree phase timers on the `SortPhaseTimer` pattern, then `perf` on Linux when per-function attribution is genuinely needed. macOS sampling is documented as a last resort because it produced three confidently wrong profiles here: the release profile strips debuginfo; blocked threads are counted as CPU unless samples are weighted by `threadCPUDelta`, which made an 18s run look like 285s across 21 threads; and `atos -o <binary>` mis-resolves addresses belonging to other images, which is where "38% of CPU in `clap_builder::error::Error::print`" on a *successful* run came from. Also records the before/after numbers for all three comparison modes, and two results that outlive this change: content-mode conclusions rest on CPU-seconds because this host's wall time is unreliable under load (identical runs measured 3.56s and 16.56s), and parallel BGZF decode costs ~40% more total CPU than single-threaded decode for the same work — which raising the read-ahead batch size did not recover, so that time is the consumer waiting on decode rather than per-handoff overhead.
…d reader Self-review of this branch found the new order-verification path under-tested and two docs left describing the pre-change flow. `content_mode_rejects_records_not_in_declared_coordinate_order` asserted only that the error message contained "order", which a degraded diagnostic would still satisfy. It now asserts the message names the declared order that was violated and the offending input. Three properties of the folded check had no test at all, and each is one a later refactor could quietly drop: - the reported violation count covers the whole file, which is why the check accumulates rather than failing on the first bad record; - bam1 is reported before bam2 when both inputs are mis-sorted, which is what keeps the diagnostic stable now that both files are checked in a single pass rather than one after the other; - a correctly ordered pair still compares normally at more than one thread count, covering the split decompression budget alongside the check. `damaged_input_yields_err_not_clean_eof` only exercised the single-threaded reader, where decode runs inline on the read-ahead thread. Above one thread it runs on a worker pool, so the failure travels a different route — worker to error slot to `CheckedRecords` — which is the route `open_pair` actually uses. The test is now parameterized over thread count as well as damage mode. Also updates two doc comments the byte-identical fast path invalidated: `compare_run` still described every record as being cancelled through `RunCanceller`, and `RunCanceller` still described every arriving record as being reduced to a content key. Neither holds for the pairs the fast path handles. Records why `positional_compare` takes `verify_order` as a required parameter: a convenience overload defaulting it to `None` would let a caller opt out of order verification by omission, which is the failure the parameter exists to prevent.
f6a7f76 to
2b68b78
Compare
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
Note GitHub couldn't provide a complete incremental comparison for this pull request, so CodeRabbit is performing a full review instead. This review may take a little longer. |
Closes #686.
fgumi compare bams --command sortran ~3x slower than thefgumi sortit validates, and--threadsdid nothing. Profiling found why, and four changes fix it; a fifth removes a redundant traversal incontentmode.Result
20.1M records/file, ~2 GB BAMs, M2 Max (12 core), page cache warmed identically before every run, minimum of 3 reps:
--command sort-t 8--command sortscaling-t 1) → 5.58s (-t 4) → 4.14s (-t 8)--command group-t 8--command simplex(content)Both of #686's acceptance criteria hold: comparison is now 2.0x faster than the sort it validates (4.14s vs 8.20s), having started 2.7x slower, and throughput scales measurably with
--threads.Content-mode results are quoted in CPU-seconds deliberately — see Measurement.
What was wrong
tricorder --traceshowedsortandgrouprunning the entire comparison in exactly one OS thread —n_threadsnever left 1, so--threadscreated nothing at all. Sampling put ~61% of that single thread in BGZF decode (~55%libdeflateinflate, ~6% the per-block CRC32) while eleven cores idled.contentreached only 2.46 cores and additionally traversed each input twice, becauseverify_records_in_orderran over both paths beforepositional_comparereopened and streamed them again — four traversals for a job needing two.The changes
Suggested reading order is commit order; each commit is independently revertable.
f7fbf95Read both inputs concurrently.OpenedInput::open_pairgives each input a share of the--threadsBGZF budget and its own read-ahead thread, so the two decodes overlap each other and the comparison. The budget is split rather than handed to each input in full: they are symmetric and both must be fully decoded, so giving each the whole count would request double the parallelism asked for. (zippersplits asymmetrically because its two inputs differ in cost; a comparison's do not.)352cfa1Cancel byte-identical records without building content keys.RunCancellerbuilt a canonical content key per record — aVecper aux tag, a sort, a concatenation, a hash and a map insert. Comparing the raw bytes first skips all of it for pairs that agree.13d8907Verify sort order during the content pass. Folds the order check into the comparison's own record pass viaOrderCheck, removing the two dedicated traversals.388ff00Skip key and content checks for byte-identical pairs (content mode).record_keys_matchran on every pair beforecontent_diffs, whose first act is the same byte comparison. Also splits content mode's decompression budget, which was still starting 8 BGZF workers per reader at-t 8.f0126b3Read the MI tag in one aux scan.get_mi_tag_rawprobedfind_int_tagthenfind_string_tag, each walking the aux block from the start;fgumi groupwritesMI:Z:, so the integer probe always missed and its scan was wasted.Why the byte-equality fast paths are sound
Both rest on the same argument, and neither replaces the general path. A
RecordKeyis a pure function of a record's bytes, andcontent_key_exactexcludesbin, width-normalizes integer tags and sorts the tag multiset — so byte equality strictly implies key and content equality. The converse fails, which is exactly why these are fast paths: records that are content-equal but differ in tag order, integer width orbinfall through and still match. For the run comparison there is a second step: removing one record from each side with the same key leaves the multiset symmetric difference unchanged whatever else is pending, so the fast path needs no precondition on canceller state.Correctness
The reader change carried a real trap.
RawReadAheadReaderisIterator<Item = RawRecord>— noResult— and signals a read failure by ending iteration and parking the error for a latertake_error(). Adopted directly, a truncated or CRC-corrupt BAM would have read as a shorter file: a record-countDIFFER, or a falseIDENTICALwhen both inputs are damaged alike. Rather than requiring every engine loop to remember a trailing check,CheckedRecordsrestores theResultcontract at the boundary, so the engines' existingrec?propagation is correct by construction and forgetting the check is not an available mistake.Every change was gated against a 121-comparison differential corpus — matched, missing, extra, edited, tag-reordered, intra-run reordered, mis-sorted, truncated, header-conflicting and byte-corrupt inputs, in both directions and at three thread counts — requiring byte-identical verdicts, exit codes and RESULT lines against a binary built from
main. The corpus is generated (fgumi simulate), not committed. Two cases carry particular weight:tag-reorderandswap-adjacentare content-equal but not byte-equal, so they prove the fast paths did not swallow the general checks.Error text changed only for damaged inputs, where the messages now name the failing file and diagnose it more precisely (
block data checksum mismatchrather than a generic open failure).New tests: sort-order violations are pinned by count, by which file is named, and by acceptance of correctly-ordered input at more than one thread count; the damaged-input regression test is parameterized over thread count as well as damage mode, covering the worker-pool error route that
open_pairactually uses. An existing test asserting only that an errorcontains("order")was strengthened to assert the actual contract.Measurement
Content-mode numbers are CPU-seconds, not wall. This host's wall time proved unreliable under interactive load — identical
-t 8runs measured 3.56s and 16.56s — while total CPU stayed within 24–27s across the same sweep. Where wall is quoted, arms were interleaved with the page cache warmed identically before every run and the minimum of 3 reps taken; the outliers are all slower, i.e. contention.All numbers are from this workstation. #686's criteria were written against a 780M-record
c7g.4xlarge, and that run has not been done — the effects here (5.7x) are far outside the noise, but the issue's own scale has not been reproduced.Notes for reviewers
fgumi-sortgains one export and one signature generalization.RawReadAheadReaderis re-exported, andverify_sort_ordernow takes anyIterator<Item = io::Result<RawRecord>>rather than a concrete reader. TheResultitem type is deliberate: it keeps a read failure distinguishable from end-of-stream.positional_comparegained a required parameter, so this is a breaking change for out-of-tree callers of that feature-gated API. A defaulting overload was considered and rejected: it would let a caller opt out of order verification by omission.verify_records_in_orderis removed —OrderChecksubsumes it and it had no callers left.fgumi sort --verifycontinues to usefgumi_sort::verify_sort_order.molecule_join) was otherwise untouched, and its remaining hot spots arefind_tag_positionandget_mi_tag_raw.Not done, deliberately
Migrating the readers to the sort engine's pooled architecture was investigated and rejected for now.
PooledInputStream::next_blockblocks onstd::thread::park()and pool workersunpark()the thread captured at construction — correct for the sort pipeline's single stream and single consumer. Two pools feeding one comparison thread would let stream A's workers wake the consumer while it waits on stream B, producing a cross-stream unpark storm: a CPU-burning busy-wait, the opposite of the intent. Fixing it properly needs per-stream condvars or a per-stream consumer thread, and the latter reintroduces the channel it would be removing.Raising the read-ahead batch from 256 to 4096 records was also measured and reverted: CPU was unchanged (26.5 vs 27.0 CPU-s) while peak RSS tripled. That null result is informative — the time in
crossbeam_channel::recvis the comparison thread waiting on decode, not per-handoff overhead, so the remaining lever is decode throughput rather than handoff granularity.CLAUDE.mdrecords the profiling ladder this work established (tricorderfirst, then in-tree phase timers, thenperfon Linux) along with three macOS sampling traps that produced confidently wrong profiles here — including "38% of CPU inclap_builder::error::Error::print" on a successful run, from symbolicating another image's addresses against the fgumi binary.Risk: comparison verdicts, exit codes,
RESULTlines,--max-diffs, and diagnostics remain pinned;unsafechanges are none and theCLAUDE.mdallowlist is unchanged; thread allocation and read-ahead behavior change, with no reported memory-bound or queue-capacity change.fgumi compare bamsreads both BAM inputs concurrently.--threadsapplies to sort and grouping comparisons through split decompression budgets.CheckedRecordspreserves deferred read errors and identifies the affected input.CLAUDE.md.