Repository navigation
fix(sort): use monotonic counter for chunk file naming during consolidation - #178
Conversation
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #178 +/- ##
==========================================
+ Coverage 84.29% 85.65% +1.36%
==========================================
Files 128 128
Lines 51978 51990 +12
==========================================
+ Hits 43814 44533 +719
+ Misses 8164 7457 -707 ☔ View full report in Codecov by Sentry. 🚀 New features to boost your workflow:
|
8d932d1 to
e60e7d6
Compare
|
Warning Rate limit exceeded
⌛ How to resolve this issue?After the wait time has elapsed, a review can be triggered using the We recommend that you space out your commits to avoid hitting the rate limit. 🚦 How do rate limits work?CodeRabbit enforces hourly rate limits for each developer per organization. Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout. Please see our FAQ for further information. ℹ️ Review info⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: 📒 Files selected for processing (1)
📝 WalkthroughWalkthroughChunk filename generation in the raw sorter was changed to use a separate monotonically increasing 🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Inline comments:
In `@src/lib/sort/raw.rs`:
- Around line 2236-2243: The assertions that compare stats.output_records (which
is derived from input counts) can miss regressions that drop records; update the
tests that currently assert stats.output_records and stats.chunks_written
(references: stats, stats.output_records, stats.chunks_written) to instead open
and count records in the produced BAM file ("output.bam") and assert that the
on-disk record count equals the expected value (expected = (num_pairs * 2) as
u64). Replace the three assertion sites noted (around the current blocks using
stats.output_records) so they validate by reading/parsing output.bam and
counting records rather than relying on stats fields. Ensure the test fails if
the file is missing or the counted records != expected.
- Around line 2207-2309: The coordinate-consolidation test needs to exercise the
index-writing path; update
test_sort_coordinate_with_consolidation_preserves_all_records to accept a
write_index parameter (rstest cases false and true) and pass it into the sorter
builder (call .write_index(write_index) on the RawExternalSorter before .sort).
Keep the other consolidation tests unchanged; ensure the SortOrder::Coordinate
case is run with both write_index = false and write_index = true so the
coordinate index-writing path is covered.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: c17bce99-0d28-4c92-8fd8-92b2a6734cf0
📒 Files selected for processing (1)
src/lib/sort/raw.rs
There was a problem hiding this comment.
♻️ Duplicate comments (2)
src/lib/sort/raw.rs (2)
2207-2211:⚠️ Potential issue | 🟡 MinorMissing coverage for coordinate sort with index.
SortOrder::Coordinateexercisessort_coordinate_optimized, but thewrite_index=truepath (sort_coordinate_with_index) also received the fix. Add a case withwrite_index: truefor coordinate sort.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/sort/raw.rs` around lines 2207 - 2211, The coordinate sort path with index isn't covered: update the parameterization for test_sort_with_consolidation_preserves_all_records to include a case that exercises SortOrder::Coordinate with write_index = true so sort_coordinate_with_index gets tested; specifically add a test case (or modify the rstest cases) to pass a configuration/flag enabling write_index for the Coordinate case so the branch in sort_coordinate_optimized that delegates to sort_coordinate_with_index is executed and validated.
2239-2245:⚠️ Potential issue | 🟠 MajorAssertion won't catch dropped records.
stats.output_recordsis assigned fromstats.total_records(Lines 807, 956, 1085, 1225), not from actual writes. Count records from the output BAM to verify no data loss.🤖 Prompt for AI Agents
Verify each finding against the current code and only fix it if needed. In `@src/lib/sort/raw.rs` around lines 2239 - 2245, The assertion comparing stats.output_records to expected is unreliable because stats.output_records is derived from stats.total_records, not actual writes; update the test to open the produced BAM output (use the same output filename and your BAM reader used elsewhere) and count records from the file, then assert that that file-record count equals expected; keep the existing check that stats.chunks_written >= 4 but replace or augment assert_eq!(stats.output_records, expected, ...) with a read-and-count step that asserts file_record_count == expected and fail with a message indicating data loss if they differ (referencing stats, expected, stats.chunks_written, and stats.output_records to locate the code).
🤖 Prompt for all review comments with AI agents
Verify each finding against the current code and only fix it if needed.
Duplicate comments:
In `@src/lib/sort/raw.rs`:
- Around line 2207-2211: The coordinate sort path with index isn't covered:
update the parameterization for
test_sort_with_consolidation_preserves_all_records to include a case that
exercises SortOrder::Coordinate with write_index = true so
sort_coordinate_with_index gets tested; specifically add a test case (or modify
the rstest cases) to pass a configuration/flag enabling write_index for the
Coordinate case so the branch in sort_coordinate_optimized that delegates to
sort_coordinate_with_index is executed and validated.
- Around line 2239-2245: The assertion comparing stats.output_records to
expected is unreliable because stats.output_records is derived from
stats.total_records, not actual writes; update the test to open the produced BAM
output (use the same output filename and your BAM reader used elsewhere) and
count records from the file, then assert that that file-record count equals
expected; keep the existing check that stats.chunks_written >= 4 but replace or
augment assert_eq!(stats.output_records, expected, ...) with a read-and-count
step that asserts file_record_count == expected and fail with a message
indicating data loss if they differ (referencing stats, expected,
stats.chunks_written, and stats.output_records to locate the code).
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro
Run ID: 8be7b0e5-246a-415d-b5d3-b30a30cc8afe
📒 Files selected for processing (1)
src/lib/sort/raw.rs
…dation Chunk files were named using chunk_files.len(), which decreases after consolidation drains entries from the vector. This caused new chunks to collide with existing non-consolidated chunk files, silently overwriting them and losing records during the final merge. Replace chunk_files.len() with a monotonic chunk_counter in all four sort methods (coordinate, coordinate-with-index, queryname, and template-coordinate).
e60e7d6 to
61d6cde
Compare
Summary
chunk_files.len(), which decreases after consolidation drains entries from the vector viadrain(..merge_count). This caused new chunks to collide with existing non-consolidated chunk files, silently overwriting them and losing records during the final merge.chunk_files.len()with a monotonicchunk_counterin all four sort methods: coordinate, coordinate-with-index, queryname, and template-coordinate.max_temp_files(4) to force consolidation and verify all records survive.Test plan
test_sort_{coordinate,queryname,template_coordinate}_with_consolidation_preserves_all_recordscargo ci-test)cargo ci-fmtcleancargo ci-lintclean