Skip to content

feat(compare): harden fgumi compare into a sound, faithful fgbio-parity oracle - #530

Merged
nh13 merged 39 commits into
mainfrom
nh/compare-hardening
Jul 19, 2026
Merged

nh13 merged 39 commits into
mainfrom
nh/compare-hardening

Conversation

@nh13

@nh13 nh13 commented Jul 10, 2026 •

Copy link
Copy Markdown
Member

What

Hardens fgumi compare (the feature-gated developer oracle both fgbio-parity audits lean on) into a sound and faithful fgbio-parity comparator: it should never report MATCH when files genuinely differ (no false negative in the oracle), and should compare everything by default so a real difference is surfaced (no silent blind spot). Both prior parity audits explained every missed divergence via a compare blind spot (BS1–BS7) yet compare itself was never audited — this is that "audit the auditor" pass.

Fully self-contained: everything is under the compare feature (new engines/ module + rewritten dispatch). No production command behavior changes.

Latest revision — streaming molecule-join + semantic header gate

The grouping comparator was reworked from an external-sort key-join into an order-independent streaming molecule-join, and a hard header precondition was added. This supersedes the earlier key-join engine and closes the review findings against it (details under Resolved review findings below). For reviewers who saw the prior revision, the net changes are:

  • Deleted the group key-join engine (keyjoin.rs), its MiBijectionTracker, the whole-file external queryname sort, and the --sort-memory/--sort-tmp-dir flags.
  • Added molecule.rs + engines/molecule_join.rs and the engines/header.rs precondition.

Architecture

Comparison decomposes into two orthogonal axes — how records are paired, and what makes a pair equal — composed per command by an explicit preset table, with no resync (any key/position mismatch is an immediate DIFFER — avoids swap≡drop+add masking):

  • Header precondition (require_compatible_headers) — runs once, up front, and hard-exits the program on an incompatible header rather than cascading into record diffs. @SQ/@RG must match; sort-order identity is decided semantically by mapping both @HD headers to the SortOrder enum, so fgumi's bare SS:template-coordinate and fgbio's SO-prefixed SS:unsorted:template-coordinate collapse to one order (the byte-for-byte SS compare that would false-fail every cross-tool run is gone). @PG/@CO are per-invocation metadata and never compared. Genuinely orderless output (extract/zipper, SO:unsorted no SS) compares fine; a header that claims a recognized order this engine can't verify is rejected.
  • Universal sort-order verification — when both inputs declare a supported order, the records are verified to actually be in it (the tag is not trusted), as a precondition for order-dependent comparison — not just under --command sort.
  • RecordKey — collision-resistant record identity (name + segment + secondary/supplementary + multimap locus).
  • Positional engine — key-lockstep exact content comparison for order-preserving commands (extract, correct, dedup, zipper, filter, consensus), gated by the verified sort order.
  • Streaming molecule-join engine (group / --mode grouping) — exploits that grouped output keeps same-MI reads consecutive: each input is cut into per-molecule runs, each molecule identified by its MI-invariant canonical id (the lexicographically smallest read name in the molecule), and molecules are matched across the two files by a never-bounded two-sided streaming hash-join (no external sort, no spill). Each matched molecule is compared by three local checks — membership (equal RecordKey set), content under ExactMinusMi, and duplex strand-partition (the /A//B split, accepted modulo a global A↔B relabel). A molecule present on only one side is a DIFFER. This replaces the key-join's global MI-bijection tracker with per-molecule membership equivalence, and the greedy equal-key-run pairing disappears entirely.
  • Sort-verify engine — detects the shared sort order from @HD, verifies each file is independently correctly ordered, and compares as a multiset grouped by maximal equal-sort-key run.
  • Metrics row key-join — localized diffs, ε-tolerant floats.

Content predicates: Exact / ExactMinusMi, each a narrow carve-out over exact equality.

Accepted divergences (the only tolerances — each individually signed off)

  1. Metrics float representation (ε) — numeric equality is the contract; integer-to-integer is exact, mixed integer/float is ε-tolerant.
  2. group MI numbering (value + within-group order) — the molecule-join matches molecules by MI-invariant canonical id, so any MI renumbering that preserves the grouping is tolerated.
  3. fgumi-only tc tag — zipper writes it and dedup consumes it; fgbio never persists it. Ignored in every predicate (narrow — only tc).
  4. sort template-coordinate tie residue (name-hash vs lexical within equal-sort-key runs).

Consensus/duplex depth-tag saturation is no longer a tolerance: fgumi now clamps every scalar depth tag to fgbio's Short ceiling at the source, so those tags compare exactly (see the paired consensus-clamp work).

Misuse is rejected, not tolerated: a grouping comparison of two BAMs with no MI tags at all is a hard error ("grouping comparison requires MI-tagged (grouped) input") — consistent with fgbio's own consensus callers, which throw IllegalStateException on a read missing its MI tag. Empty-vs-empty is a vacuous MATCH.

Validation

  • Mutation/metamorphic soundness harness — each catalog case re-runs the real comparator (molecule_join_compare/positional/sort-verify/metrics) and asserts its verdict, so a dimension that stops being compared fails the suite; a completeness guard requires every compared dimension to be named by a case.
  • Real fgbio 4.1.0 parity run — built fgbio 4.1.0 from source, ran GroupReadsByUmi and fgumi group (adjacency, --edits 1) on 20k molecules / 119k records, and compared with the new molecule-join and cross-checked the old key-join engine on the same pair: both EQUIVALENT (exit 0), in agreement. Identical family-size distribution, consistent MI bijection, no crash and no gate false-fail on real fgbio output (SS:unsorted:template-coordinate header, integer MI:Z: tags).
  • Streaming/no-spill assertion — a 50k-molecule comparison creates no spill/scratch directory; peak buffering scales with in-flight molecules, not file size.
  • --max-diffs independence — the grouping verdict is keyed off the matched-molecule counter, so it is correct even at --max-diffs 0 (a diff-cap can no longer mask a real difference in the --quiet CI path).

Resolved review findings

All CodeRabbit findings on this PR are resolved: the three heavy-lift grouping findings (F1 secondary/supplementary locus guard, greedy MI-unaware equal-key pairing, unbounded equal-key run memory) are superseded — the code they flag is deleted; the --sort-memory/--sort-tmp-dir doc finding is moot (flags removed); and the remaining items are fixed — mixed int/float doc contract, rejecting unsupported SO:coordinate SS sub-sorts, requiring cM/cE in the mutation-catalog completeness guard, bounding the @SQ diagnostic allocation, and record-byte-exact same-seed determinism assertions.

Quality

Built via a task-by-task TDD process with an independent review after each task and a final whole-branch review (which caught and fixed a --max-diffs 0 verdict soundness bug). cargo nextest (2,700+ tests), clippy -D warnings, and rustfmt all clean.

Related

Summary by CodeRabbit

  • New Features
    • Added fgumi compare bams presets via --command with a dedicated --command sort mode.
    • Added fgumi compare metrics --key-columns with outer-join style comparison, multi-column keys, and row reordering tolerance.
  • Bug Fixes
    • Strengthened BAM header compatibility gating and harmonized mismatch reporting across content, grouping, molecule-based, and sort verification comparisons.
    • Enforced unique --key-columns and tightened numeric-key/tolerance behavior in metrics comparisons (including float representation handling).
  • Documentation
    • Refreshed compare bams/compare metrics docs, examples, and CI guidance for --command presets and --key-columns.
  • Tests
    • Expanded engine and end-to-end integration coverage, including preset wiring, --quiet behavior, and new failure-mode assertions.

@nh13
nh13 temporarily deployed to github-actions July 10, 2026 04:14 — with GitHub Actions Inactive
@coderabbitai

coderabbitai Bot commented Jul 10, 2026 •

Copy link
Copy Markdown

Review Change Stack

Note

Reviews paused

It looks like this branch is under active development. To avoid overwhelming you with review comments due to an influx of new commits, CodeRabbit has automatically paused this review. You can configure this behavior by changing the reviews.auto_review.auto_pause_after_reviewed_commits setting.

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 2f3e3c7b-9ede-459d-b0ff-e38541711d6d

📥 Commits

Reviewing files that changed from the base of the PR and between 772097e and c00345c.

📒 Files selected for processing (10)
  • docs/compare-cli.md
  • src/lib/commands/compare/bams.rs
  • src/lib/commands/compare/engines/header.rs
  • src/lib/commands/compare/engines/molecule_join.rs
  • src/lib/commands/compare/engines/sort_verify.rs
  • src/lib/commands/compare/metrics.rs
  • src/lib/commands/compare/molecule.rs
  • tests/integration/test_compare_bams.rs
  • tests/integration/test_compare_mutation.rs
  • tests/integration/test_e2e_regression.rs

Walkthrough

Replaces legacy BAM modes with content, grouping, and sort-verification engines. Metrics comparison now uses configurable outer key joins. Header compatibility, typed predicates, structured outcomes, CLI presets, documentation, integration coverage, and determinism checks are expanded.

Changes

Comparison engine overhaul

Layer / File(s) Summary
Comparison contracts and shared semantics
src/lib/commands/compare/{engines,raw_compare,record_key,molecule}.rs
Adds predicates, header compatibility rules, collision-resistant record identity, MI/tag semantics, molecule-run validation, and shared engine utilities.
Positional and grouping comparison flow
src/lib/commands/compare/{bams.rs,engines/positional.rs,engines/molecule_join.rs}
Dispatches presets to lockstep content or streaming molecule-join comparisons with structured mismatch outcomes.
Sort verification engine
src/lib/commands/compare/engines/sort_verify.rs
Detects declared sort orders, verifies monotonicity, and compares equal-key runs as multisets.
Metrics key-join comparison
src/lib/commands/compare/metrics.rs
Adds --key-columns, structured TSV validation, outer joins, duplicate-key errors, typed comparisons, and localized reporting.
CLI wiring and cross-engine validation
src/lib/commands/compare/bams.rs, tests/integration/*
Adds preset dispatch, shared fixtures, mutation coverage, CLI integration tests, no-spill checks, and byte-level determinism assertions.
Comparison documentation
docs/compare-cli.md
Documents command presets, grouping and sort behavior, metrics key joins, comparison rules, and CI invocations.

Estimated code review effort: 5 (Critical) | ~120 minutes

Sequence Diagram(s)

sequenceDiagram
  participant CLI
  participant CompareBams
  participant Engine
  participant BAMFiles
  CLI->>CompareBams: select command preset
  CompareBams->>BAMFiles: read headers and records
  CompareBams->>Engine: run content, grouping, or sort verification
  Engine->>CompareBams: return structured outcome
  CompareBams->>CLI: print IDENTICAL or DIFFER
Loading

Possibly related PRs

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately summarizes the main change: hardening fgumi compare into a sound, parity-focused comparison oracle.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch nh/compare-hardening

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Jul 10, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.09944% with 91 lines in your changes missing coverage. Please review.
✅ Project coverage is 93.38%. Comparing base (602fd60) to head (c00345c).

Files with missing lines Patch % Lines
src/lib/commands/compare/engines/sort_verify.rs 86.72% 45 Missing ⚠️
src/lib/commands/compare/metrics.rs 95.00% 24 Missing ⚠️
src/lib/commands/compare/engines/positional.rs 86.23% 15 Missing ⚠️
src/lib/commands/compare/engines/header.rs 98.95% 4 Missing ⚠️
src/lib/commands/compare/engines/molecule_join.rs 99.70% 1 Missing ⚠️
src/lib/commands/compare/molecule.rs 99.34% 1 Missing ⚠️
src/lib/commands/compare/raw_compare.rs 99.54% 1 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #530      +/-   ##
==========================================
+ Coverage   93.00%   93.38%   +0.37%     
==========================================
  Files         167      175       +8     
  Lines      103262   104539    +1277     
==========================================
+ Hits        96041    97620    +1579     
+ Misses       7221     6919     -302     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@nh13

nh13 commented Jul 16, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 16, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 16, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 16, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai

coderabbitai Bot commented Jul 16, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 8

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
docs/compare-cli.md (1)

82-98: 🎯 Functional Correctness | 🟠 Major | ⚡ Quick win

Do not present non-presets as --command presets.

clip, downsample, and review appear in this preset-settings table but are not accepted by the documented/runtime --command preset list. Remove them from the preset table, label them as manual mode guidance, or add corresponding CLI presets.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/compare-cli.md` around lines 82 - 98, Update the command preset table
near the documented Command Presets section so it only lists stages accepted by
the documented/runtime --command preset list. Remove clip, downsample, and
review from the preset rows, or clearly move them into separate manual-mode
guidance; do not present them as --command presets unless corresponding CLI
presets are added.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/compare-cli.md`:
- Line 48: Update the --ignore-order entry in docs/compare-cli.md to describe it
as applying only to grouping modes, replacing the incorrect reference to
parallel consensus output. Keep the existing flag type and default unchanged.
- Line 211: Update the “Comparing fgumi vs fgbio output” documentation to
distinguish tolerated row-order differences from missing or extra rows. State
that row-set differences are known parity gaps reported as DIFFER, while
preserving the existing note about floating-point representation differences.

In `@src/lib/commands/compare/engines/content.rs`:
- Around line 148-166: The tag comparison currently collapses duplicate non-MI
tags via BTreeMap, causing false matches. Replace tag_typed_map_excluding_mi and
the map-based logic in tags_match_excluding_mi with entry-level multiset
matching that filters both MI and tc tags while preserving duplicate
occurrences; add a regression test covering duplicate NM entries versus a single
NM entry.

In `@src/lib/commands/compare/metrics.rs`:
- Around line 225-269: Replace the fully materialized row and key-tree workflow
in parse and the keyed comparison flow with externally sorted temporary row
storage and a streaming merge of both inputs. Keep only bounded buffers in
memory, compare matching keys during the merge, and preserve duplicate-key
detection so repeated UMI keys remain an error rather than being silently
collapsed. Apply the same approach to the related key-joining logic in the
referenced ranges.
- Around line 505-509: Update the public compare_metrics boundary to validate
every tolerance in MetricsCompareConfig, rejecting negative or non-finite values
before comparison begins while preserving the existing Result-based error
handling for invalid configuration.

In `@tests/integration/test_compare_bams.rs`:
- Around line 60-84: Update run_compare_command and the analogous helpers
covering the referenced test ranges to return the subprocess ExitStatus and
captured stderr alongside stdout instead of reducing status to success(). Adjust
negative-test assertions to verify the expected exit code and diagnostic,
distinguishing comparison mismatches from argument-rejection failures.
- Around line 797-851: Strengthen
test_canonicalize_to_queryname_sorts_and_preserves_records and read_all_names
with an independent semantic fingerprint covering each record’s flags, loci,
sequence, qualities, tags, and complete queryname tie-break lanes. Expand the
fixture records to include duplicate QNAMEs spanning segment, secondary, and
supplementary cases, then compare the output fingerprints against the input
multiset and assert non-decreasing order using the full queryname sort key
rather than names alone.

In `@tests/integration/test_compare_metrics_command.rs`:
- Around line 139-152: Strengthen subprocess_quiet_suppresses_output_on_differ
by asserting stderr is empty, not merely that it lacks “Error:”. Update the
--quiet execution path to suppress all informational and warning logging so only
the exit code communicates the comparison result.

---

Outside diff comments:
In `@docs/compare-cli.md`:
- Around line 82-98: Update the command preset table near the documented Command
Presets section so it only lists stages accepted by the documented/runtime
--command preset list. Remove clip, downsample, and review from the preset rows,
or clearly move them into separate manual-mode guidance; do not present them as
--command presets unless corresponding CLI presets are added.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 84ba94c2-4cb1-4969-a8c9-f23b75846b60

📥 Commits

Reviewing files that changed from the base of the PR and between cbc7d72 and 63d7d30.

📒 Files selected for processing (20)
  • docs/compare-cli.md
  • src/lib/commands/compare/bams.rs
  • src/lib/commands/compare/engines/content.rs
  • src/lib/commands/compare/engines/header.rs
  • src/lib/commands/compare/engines/keyjoin.rs
  • src/lib/commands/compare/engines/mod.rs
  • src/lib/commands/compare/engines/positional.rs
  • src/lib/commands/compare/engines/sort_verify.rs
  • src/lib/commands/compare/metrics.rs
  • src/lib/commands/compare/mod.rs
  • src/lib/commands/compare/raw_compare.rs
  • src/lib/commands/compare/record_key.rs
  • src/lib/commands/sort.rs
  • tests/integration/helpers/bam_generator.rs
  • tests/integration/helpers/mod.rs
  • tests/integration/main.rs
  • tests/integration/test_compare_bams.rs
  • tests/integration/test_compare_metrics_command.rs
  • tests/integration/test_compare_mutation.rs
  • tests/integration/test_e2e_regression.rs

Comment thread docs/compare-cli.md Outdated
Comment thread docs/compare-cli.md Outdated
Comment thread src/lib/commands/compare/engines/content.rs Outdated
Comment thread src/lib/commands/compare/metrics.rs
Comment thread src/lib/commands/compare/metrics.rs
Comment thread tests/integration/test_compare_bams.rs Outdated
Comment thread tests/integration/test_compare_bams.rs Outdated
Comment thread tests/integration/test_compare_metrics_command.rs
@nh13
nh13 force-pushed the nh/compare-hardening branch from 63d7d30 to 13ae4ac Compare July 16, 2026 15:59
@nh13
nh13 temporarily deployed to github-actions July 16, 2026 15:59 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 16, 2026

Copy link
Copy Markdown
Member Author

Addressed this review round (pushed to nh/compare-hardening):

Fixed (threads resolved):

  • --ignore-order docs — documented as valid only with --mode grouping (the CLI rejects it otherwise), dropped the "parallel consensus output" wording.
  • Metrics row-set docs — clarified that only row order is tolerated; a differing row set is a known fgbio-parity gap reported as DIFFER, never a success.
  • Duplicate non-MI tags (content.rs) — ExactMinusMi no longer collapses tags through a name-keyed BTreeMap ([NM=1,NM=2] vs [NM=2] was a false MATCH); it now delegates to an entry-level multiset comparator (raw_tags_equal_order_independent_excluding_mi), matching Exact. Regression tests added at both the raw and content_diffs levels.
  • Tolerance validation (metrics.rs) — compare_metrics now rejects negative/non-finite --rel-tol/--abs-tol at the boundary, with an rstest table.
  • Subprocess exit codes (test_compare_bams.rs) — run_compare/run_compare_command return the exact exit code + stderr; negative tests assert Some(1) and the argument-rejection tests assert the specific diagnostic, so a panic/usage error can no longer satisfy a !success check.
  • Canonicalize test (test_compare_bams.rs + the twin in keyjoin.rs) — independent full-record fingerprint (bin-stripped) for exact multiset preservation, a duplicate-QNAME fixture spanning the segment/secondary/supplementary flag lanes, and a non-decreasing assertion on the full (name, flag-order) sort key.
  • Quiet mode (metrics.rs) — see the thread; narrowed to the fgbio-consistent contract and gated the command's own logging behind !quiet.

Outside-diff comment (preset table, docs 82-98): removed clip/downsample/review from the preset rows (they have no --command preset) and added the missing sort/dedup presets, via a new Preset? column that marks each stage.

Deferred (thread left open): the metrics external-sort/memory-bound suggestion — QC metric TSVs are bounded and small (largest is .umi_counts.txt, one short row per distinct UMI); rationale in the thread.

All green locally: full test suite (2443 passed), cargo ci-fmt, cargo ci-lint (pedantic).

@nh13

nh13 commented Jul 16, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai resume

@nh13

nh13 commented Jul 16, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 16, 2026

Copy link
Copy Markdown
✅ Action performed

Reviews resumed.

@coderabbitai

coderabbitai Bot commented Jul 16, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 10

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
src/lib/commands/compare/raw_compare.rs (1)

190-220: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Match duplicate tags by both name and value.

The first unused same-name entry is selected before value comparison. Thus [NM=1, NM=2] falsely differs from reordered [NM=2, NM=1]. Search all unmatched same-name entries for a semantically equal value, then mark that candidate; add this reordered-duplicate regression.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/lib/commands/compare/raw_compare.rs` around lines 190 - 220, Update the
matching loop around entries1, entries2, and matched so it searches every
unmatched entry with the same tag name and selects one whose value matches
byte-for-byte or via integer semantic decoding, instead of returning on the
first same-name mismatch. Mark only the successfully matched candidate and
return false if no semantically equal unmatched candidate exists; add a
regression test covering reordered duplicate values such as [NM=1, NM=2] versus
[NM=2, NM=1].
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/compare-cli.md`:
- Around line 52-53: Update the scope descriptions for --sort-memory and
--sort-tmp-dir in the CLI documentation to include the explicit --mode grouping
--ignore-order path alongside --command group, making clear these settings
configure the key-join engine in both modes while preserving the existing
behavior details.

In `@src/lib/commands/compare/bams.rs`:
- Around line 605-607: Gate OperationTimer creation/completion and informational
logging in the compare command’s dispatch paths behind !self.quiet, including
execute_sort_verify and the additional referenced branches. Preserve all command
execution and exit-code behavior while ensuring quiet mode emits no timers or
info logs, matching the metrics command.
- Around line 645-660: Update the predicate selection around self.command and
mode so an explicit CompareMode::Content always uses ContentPredicate::Exact,
overriding any command preset such as group; retain preset predicates only when
content mode was not explicitly selected, and add a regression test covering
--command group --mode content with an MI-only difference that must not match.

In `@src/lib/commands/compare/engines/keyjoin.rs`:
- Around line 50-58: Replace the unconditional DEFAULT_DISK_BACKED_TMP_DIR
fallback with a platform-aware default that is valid on Windows and preserves
the disk-backed behavior on Unix-like systems. Update the related
temporary-directory resolution in the command-group comparison flow to verify
actual scratch-directory creation, not merely inspect the resolved path, before
passing it to RawExternalSorter. Add or update tests covering Windows-compatible
fallback selection and creation failure.

In `@src/lib/commands/compare/engines/sort_verify.rs`:
- Around line 122-158: Update detect_sort_order so SS validation checks the
complete value before selecting a comparator: accept only documented bare
subsort aliases or exact SO-prefixed forms matching the current SO, and reject
mismatched prefixes such as coordinate:natural under SO:queryname. Preserve the
existing SortOrder mappings and error behavior for unsupported SS values.
- Around line 187-257: The equal-key comparison currently permits unbounded
memory and quadratic CPU usage. Update pull_run and multiset_equal to compare
canonical record encodings using linear-time counting, and introduce a bounded
threshold that spills or externally sorts records when a run exceeds it;
preserve exact multiset semantics and avoid buffering an entire large run in
memory.

In `@src/lib/commands/compare/metrics.rs`:
- Around line 355-368: Update canonicalize_key_field and the precision-handling
logic in parse_value so non-finite multipliers or scaled float results never
canonicalize distinct finite keys to inf/NaN; preserve the original finite value
or reject unsafe rounding and leave the raw key representation unchanged. Apply
the same safeguard to the related call sites, and add a regression covering
high-precision keys with float-rounding overflow.

In `@src/lib/commands/compare/record_key.rs`:
- Around line 74-79: Update segment_of to handle the (true, true) FIRST|LAST
state explicitly instead of routing it through Segment::Fragment. Represent all
four flag combinations with distinct Segment variants, or reject the invalid
combined state, while preserving the existing First, Last, and Fragment behavior
for the other combinations.

In `@tests/integration/test_compare_mutation.rs`:
- Around line 524-546: Split
positional_consensus_cm_and_ce_differences_are_flagged into independent tests
that vary only cD, only cM, or only cE while keeping the other tags identical,
and assert each mutation produces DIFFER. Add a corresponding catalog entry for
each tag-specific test, including the related cases around the referenced
additional section, and strengthen any assertions that do not explicitly verify
the stated difference contract.
- Around line 1034-1046: Replace the self-maintained
REQUIRED_SUBSTRINGS/MUTATION_CATALOG completeness check with an oracle owned by
production comparison behavior. Export or expose the compared-dimension
inventory from positional_compare, compare_headers, keyjoin_compare,
sort_verify_compare, and CompareMetrics::execute, then have
mutation_catalog_covers_every_compared_dimension validate catalog coverage
against that inventory rather than only catalog fields or test-local closures.

---

Outside diff comments:
In `@src/lib/commands/compare/raw_compare.rs`:
- Around line 190-220: Update the matching loop around entries1, entries2, and
matched so it searches every unmatched entry with the same tag name and selects
one whose value matches byte-for-byte or via integer semantic decoding, instead
of returning on the first same-name mismatch. Mark only the successfully matched
candidate and return false if no semantically equal unmatched candidate exists;
add a regression test covering reordered duplicate values such as [NM=1, NM=2]
versus [NM=2, NM=1].
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 5cc602f9-9027-484a-86c3-7c40f09ba3ae

📥 Commits

Reviewing files that changed from the base of the PR and between 63d7d30 and 13ae4ac.

📒 Files selected for processing (20)
  • docs/compare-cli.md
  • src/lib/commands/compare/bams.rs
  • src/lib/commands/compare/engines/content.rs
  • src/lib/commands/compare/engines/header.rs
  • src/lib/commands/compare/engines/keyjoin.rs
  • src/lib/commands/compare/engines/mod.rs
  • src/lib/commands/compare/engines/positional.rs
  • src/lib/commands/compare/engines/sort_verify.rs
  • src/lib/commands/compare/metrics.rs
  • src/lib/commands/compare/mod.rs
  • src/lib/commands/compare/raw_compare.rs
  • src/lib/commands/compare/record_key.rs
  • src/lib/commands/sort.rs
  • tests/integration/helpers/bam_generator.rs
  • tests/integration/helpers/mod.rs
  • tests/integration/main.rs
  • tests/integration/test_compare_bams.rs
  • tests/integration/test_compare_metrics_command.rs
  • tests/integration/test_compare_mutation.rs
  • tests/integration/test_e2e_regression.rs

Comment thread docs/compare-cli.md Outdated
Comment thread src/lib/commands/compare/bams.rs Outdated
Comment thread src/lib/commands/compare/bams.rs Outdated
Comment thread src/lib/commands/compare/engines/keyjoin.rs Outdated
Comment thread src/lib/commands/compare/engines/sort_verify.rs Outdated
Comment thread src/lib/commands/compare/engines/sort_verify.rs
Comment thread src/lib/commands/compare/metrics.rs
Comment thread src/lib/commands/compare/record_key.rs
Comment thread tests/integration/test_compare_mutation.rs Outdated
Comment thread tests/integration/test_compare_mutation.rs
@nh13
nh13 force-pushed the nh/compare-hardening branch from 13ae4ac to fe79633 Compare July 16, 2026 19:41
@nh13
nh13 temporarily deployed to github-actions July 16, 2026 19:41 — with GitHub Actions Inactive
@coderabbitai

coderabbitai Bot commented Jul 19, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 19, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 19, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 19, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 19, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 2

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
docs/compare-cli.md (1)

146-148: 🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win

Document tolerance-based equality, not only representation differences.

The implementation accepts any finite float difference within --precision, --rel-tol, and --abs-tol; formatting differences are only the motivating use case. Reword this to avoid promising exact comparison for other numeric differences.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docs/compare-cli.md` around lines 146 - 148, Revise the comparison-behavior
statement in docs/compare-cli.md to explain that finite float differences within
--precision, --rel-tol, or --abs-tol are tolerated, rather than limiting
tolerance to formatting differences. Preserve the statement that row sets and
non-key values are compared, but avoid promising exact equality for numeric
values covered by these tolerances.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@docs/compare-cli.md`:
- Around line 184-193: Update the Float Comparison documentation in
docs/compare-cli.md to scope tolerance formulas to finite float values, and
explicitly document that NaN compares equal to NaN and positive or negative
infinity compares equal only to the same-sign infinity. Keep the existing mixed
integer/float special-value behavior unchanged.

In `@tests/integration/test_e2e_regression.rs`:
- Around line 278-300: Move the BAM byte-identical assertion documentation block
from above read_bam_header to directly above assert_bams_record_byte_identical,
leaving the read_bam_header and hd_ss documentation attached to their respective
helpers.

---

Outside diff comments:
In `@docs/compare-cli.md`:
- Around line 146-148: Revise the comparison-behavior statement in
docs/compare-cli.md to explain that finite float differences within --precision,
--rel-tol, or --abs-tol are tolerated, rather than limiting tolerance to
formatting differences. Preserve the statement that row sets and non-key values
are compared, but avoid promising exact equality for numeric values covered by
these tolerances.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: a4d2cc52-8c5f-41bd-a492-c5a172b4b3db

📥 Commits

Reviewing files that changed from the base of the PR and between 84f74d9 and 772097e.

📒 Files selected for processing (10)
  • docs/compare-cli.md
  • src/lib/commands/compare/bams.rs
  • src/lib/commands/compare/engines/header.rs
  • src/lib/commands/compare/engines/molecule_join.rs
  • src/lib/commands/compare/engines/sort_verify.rs
  • src/lib/commands/compare/metrics.rs
  • src/lib/commands/compare/molecule.rs
  • tests/integration/test_compare_bams.rs
  • tests/integration/test_compare_mutation.rs
  • tests/integration/test_e2e_regression.rs

Comment thread docs/compare-cli.md
Comment thread tests/integration/test_e2e_regression.rs Outdated
nh13 added 10 commits July 19, 2026 09:25
values_equal compares Int/Int exactly but mixed Int/Float with the same
rel/abs tolerance as Float/Float. The module doc, CLI long_about, and
docs/compare-cli.md all previously said only "integers are compared
exactly", which reads as covering the mixed case too. Spell out the
three cases (int/int exact, mixed int/float tolerant, string exact) so
the docs match values_equal's actual behavior. No implementation change.
sort_order_from_header's SO:coordinate arm ignored the @hd SS tag
entirely, so a declared coordinate sub-order (e.g. a hypothetical
coordinate:foo) would falsely validate as plain coordinate even though
run_compare's equal-core-sort-key grouping only ever verifies
(tid, pos, reverse) -- weaker than whatever the sub-sort would promise.

Follow the existing SO:queryname arm's pattern: reuse ss_subsort to
validate the whole SS value (not just its suffix) against SO:coordinate.
No SS tag is still accepted (-> SortOrder::Coordinate); any sub-sort,
recognized or not, now bails, since this engine implements none.
SortOrder::header_ss_tag confirms fgumi itself never emits a coordinate
SS, so this cannot reject fgumi's own output.
mutation_catalog_covers_every_compared_dimension's REQUIRED_SUBSTRINGS
completeness list named cD but not cM/cE, even though the catalog has
"cM difference flagged" and "cE difference flagged" entries -- so
deleting those two rows would have silently passed the guard. Add both
substrings alongside the existing cD entries.

Verified empirically: temporarily deleting the cM/cE CatalogEntry
blocks makes this test fail with "no mutation catalog entry covers
required dimension \"cM difference flagged\"", confirming the guard
actually catches the deletion.
… diffs

compare_sq rendered BOTH entire @sq dictionaries into one diff string
whenever they diverged at all, regardless of --max-diffs (which can't
cap it since it's a single push_diff entry) -- a large reference
dictionary (whole-genome assemblies routinely carry thousands of
contigs/decoys/ALTs) allocated unbounded diagnostic text for even a
single differing entry.

Walk both dictionaries position-by-position instead: any divergence
anywhere (name/length/order/M5/etc., or a one-sided length difference)
is still counted and reported, but the rendered string shows only the
first MAX_SQ_DIFF_ENTRIES (5) differing positions plus an "and N more
differences" count. Existing differing_sq_* tests (which only assert
`starts_with("@sq")`) stay green; added a regression test building a
5000-entry dictionary with one differing entry and asserting the diff
string stays under 2KB, proving it's no longer O(dictionary size).
The same-seed determinism tests in test_e2e_regression.rs asserted
identity via CompareBams "content" mode, which normalizes tag order,
integer tag width, the tc tag, and @PG/@co (see
raw_compare::content_key_exact) -- so a same-seed writer nondeterminism
confined to exactly those dimensions would be masked and go undetected.

Add assert_bams_record_byte_identical, modeled on the strip_bin/
read_all_records helpers from the pre-redesign key-join canonicalization
test (removed by c156ba5 along with the key-join engine): it reads
every record's raw on-disk bytes via fgumi_sort::RawBamRecordReader
after skip_header(), zeroing only the non-semantic bin field (bytes
10-11). Header is excluded -- @pg carries per-run temp paths that
legitimately differ run-to-run even for deterministic record output.

Applied to all five same-seed determinism tests (sibling audit):
test_simulate_grouped_reads_deterministic, test_simplex_pipeline_
deterministic, test_simplex_filter_pipeline_deterministic,
test_full_pipeline_extract_to_filter, test_dedup_pipeline_deterministic.
Added alongside (not replacing) the existing content-mode assertion,
since content mode also exercises the full compare engine's header/
sort-order preconditions via an independent code path.

Verified empirically before wiring this up broadly: ran the new
assertion against all five tests repeatedly (4 full runs) and same-seed
output was record-byte-identical every time -- no real nondeterminism
found in tag order, integer tag width, or the tc tag for any of these
pipelines.
require_compatible_headers previously collapsed both branches of the
(Err, Err) sort_order_from_header case into a benign Ok(None), which
was correct for genuinely orderless headers (extract/fastq/zipper
output) but also silently swallowed headers that declare a recognized
orderable SO (coordinate or queryname) with an unrecognized SS
sub-sort. That let a header claiming a verifiable sort order skip
order verification entirely instead of being rejected.

Distinguish the two cases by reading the raw @hd SO tag directly: if
either header's SO names coordinate or queryname, propagate the
underlying error as a hard incompatibility; only fall back to the
SO/GO byte comparison when neither side claims an orderable SO.
…conditions

Close four soundness gaps in the streaming grouping comparison surfaced by
review:

- molecule_runs assumed same-MI-consecutive input but never checked it; a
  scattered MI base (records for one molecule split into two runs) now
  returns an Err instead of silently mis-grouping.
- require_compatible_headers' (Ok, Ok) arm only compared the normalized
  SortOrder, so two headers agreeing on sort order but declaring different
  @hd GO tags passed the gate; it now applies the same compare_hd check the
  (Err, Err) fallback already did.
- molecule_join_compare relied on its CLI caller having already run
  require_compatible_headers; it now enforces that precondition itself so
  the public API is sound for any caller.
- The fully-MI-less guard only fired when *both* inputs were MI-less, so a
  partial-MI-less pair (one side grouped, the other not) whose single
  spanning run happened to canonical-id-match a real molecule on the other
  side reported a false MATCH. The guard now rejects either non-empty
  MI-less side unconditionally.

Also updates the comment in bams.rs that claimed grouping mode needs no
order validation at all -- it is order-independent across molecules, but
molecule_runs now enforces same-MI-consecutive contiguity within each one.
…agnostics

index_by_key collapsed same-RecordKey members of one molecule into a
last-wins BTreeMap entry, so a duplicated or dropped record sharing a key
with another (e.g. two same-name/same-segment primaries at different
positions) could go unnoticed -- a genuine 3-vs-2 multiplicity difference
reported MATCH. It now maps each RecordKey to all of its member records and
compare_molecule compares per-key multiplicity, closing that false-MATCH
avenue. The duplex strand-partition sets are now multisets for the same
reason.

compare_molecule used to build its full Vec<String> of diffs before the
caller applied the max-diffs cap, so a hugely-divergent molecule allocated
unbounded diagnostic strings even under --max-diffs 0. It now takes the cap
and a diff sink directly, stops allocating past the cap, and returns an
uncapped bool so the matched-molecule verdict stays cap-independent.

The EOF residual drain iterated pending1/pending2 (AHashMap) via .drain(),
so which residual ids survived the max-diffs cap and their reported order
varied run to run. Residual canonical ids are now sorted before reporting.
… check

The five same-seed determinism tests in test_e2e_regression.rs asserted
both assert_bams_identical(..., "content", ...) and
assert_bams_record_byte_identical(...). The byte-exact assertion is
strictly stronger (content mode normalizes tag order, integer tag width,
and the tc tag), so the content-mode call added nothing; drop it and keep
determinism asserted byte-exact only, with content-mode compare reserved
for semantic-parity checks. assert_bams_identical becomes unused as a
result and is removed.

molecule_join_does_not_spill_to_disk counted directory entries beside the
BAM fixtures, which can't catch a spill routed through
tempfile/std::env::temp_dir() to the actual system temp directory, or one
written into a subdirectory. It now points TMPDIR/TEMP/TMP at an isolated,
empty directory for the compare (run via the real CLI subprocess, so the
env vars are scoped to the child process) and asserts that directory's
recursive contents stay empty.
… set)

Both fgumi group and fgbio GroupReadsByUmi assign the MI as a monotonically increasing counter in template-coordinate emission order, so a valid grouped file's base MI is strictly increasing across runs. Track only the largest base seen (O(1)) instead of a HashSet of all closed bases (O(molecules)): a new run whose base is not greater than one already seen is a reappearance or a non-monotonic id -- i.e. not grouped -- and is rejected. This also catches a base reappearing across an intervening MI-less run, which the set-membership form still handled but the adjacent-only form would miss. Reorder test fixtures are renumbered so each file's MI is monotonic in emission order (the molecule emitted first gets the lower MI), matching real tool output.
@nh13
nh13 force-pushed the nh/compare-hardening branch from 772097e to c00345c Compare July 19, 2026 13:26
@nh13
nh13 temporarily deployed to github-actions July 19, 2026 13:27 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 19, 2026

Copy link
Copy Markdown
Member Author

Addressed all three findings from the latest review (2 inline + 1 outside-diff), amended into their originating commits and force-pushed.

docs/compare-cli.md 146-148 (outside diff) — tolerance-based equality, not only representation differences. Reworded the compare metrics overview: numeric slack is now stated as finite floats being rounded to --precision and then accepted within --rel-tol/--abs-tol, with the defaults merely sized to absorb fgbio DecimalFormat("0.######") vs full-precision formatting — no longer promising that only reformatting is tolerated. The faithful-comparison claim is now scoped to the row set, integer/integer pairs, and strings, with mixed integer/float pointed at the detailed rules.

docs/compare-cli.md 184-193 — float/float special-value semantics. The tolerance formulas are now explicitly scoped to finite floats, plus a new list documenting the non-finite rules that values_equal actually implements: NaN == NaN (deliberately unlike IEEE 754), an infinity equal only to the same-sign infinity, and a non-finite value never equal to a finite one at any tolerance. These were already covered by test_nan_equality, the same-sign infinity cases, and test_nan_not_equal_to_number — the docs had just never stated them.

tests/integration/test_e2e_regression.rs 278-300 — doc comment on the wrong function. Moved the byte-identical assertion block down onto assert_bams_record_byte_identical and gave read_bam_header its own one-line doc; hd_ss is unchanged.

A pre-push self-review pass on the reworded prose caught two follow-ons in my own edit, both fixed before pushing: the overview still promised faithful comparison for "any non-key value" (which would have re-introduced the same over-promise, since non-key floats are exactly what the tolerance absorbs), and it conflated --precision with the tolerance window — --precision rounds at parse time in parse_value and is not a term in the rel_tol/abs_tol inequality.

Full suite green: 5480 passed, 26 skipped; lint and fmt clean.

@nh13

nh13 commented Jul 19, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 19, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

This branch was previously deployed

1 inactive deployment
github-actions — c00345cd Deployed Jul 19, 2026 by nh13 via coverage #2733
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant