Repository navigation
fix(fastq): narrow read-name sync suffix strip to fgbio parity - #512
Conversation
`strip_read_suffix` canonicalizes read names before checking that
paired/interleaved FASTQ streams are in sync (that the reads at the same
position across streams share a name). Its old rule was too lenient:
after stripping a trailing space-separated comment it also stripped a
trailing `separator + digit` for separator in {`/`, `.`, `_`, `:`} and
digit in {`1`, `2`}. That collapses a genuinely-mismatched pair such as
`read_1` (stream 0) and `read_2` (stream 1) to the same base name
`read`, so two different reads are silently accepted as "in sync".
Narrow the rule to match fgbio's `FastqSource` read-name
canonicalization (`com/fulcrumgenomics/fastq/FastqSource.scala`), which
strips only a trailing `/` followed by a single ASCII digit (0-9). The
space-comment strip is unchanged; `.`/`_`/`:` separators and multi-digit
runs (`read/12`) are now preserved. This is the same rule PR #489
established for the extract path, applied here to the shared
FASTQ-sync helper.
All three production callers are paired-FASTQ name-sync validation and
benefit uniformly with no call-site changes: `FastqGrouper`
(`grouper.rs`), `zip_fastq`, and `parse_zip_fastq`.
Behavior tradeoff: narrowing correctly rejects mismatched `read_1` /
`read_2` streams, but it ALSO now rejects non-standard paired FASTQs
that legitimately use `.1`/`.2`/`_1`/`_2`/`:1`/`:2` read-number
suffixes. This matches fgbio (which rejects them too), so it is correct
fgbio parity, but it is a real behavior change reviewers should note.
Future cleanup: `extract.rs` still carries its own broad
`strip_read_suffix_extract`; consolidating both onto the narrow shared
helper is a follow-up left out of this focused fix.
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (2)
WalkthroughRisk of collapsing distinct paired reads is reduced by narrowing suffix stripping to trailing ChangesFASTQ suffix canonicalization
Estimated code review effort: 2 (Simple) | ~10 minutes Possibly related PRs
Suggested labels: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #512 +/- ##
==========================================
- Coverage 91.12% 91.05% -0.07%
==========================================
Files 78 78
Lines 51540 51541 +1
==========================================
- Hits 46964 46933 -31
- Misses 4576 4608 +32 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
Problem
fastq_parse::strip_read_suffixcanonicalizes read names before validating that paired/interleaved FASTQ streams are in sync — i.e. that the reads at the same position across streams share a name. Its rule was too lenient: after stripping a trailing space-separated comment, it also stripped a trailingseparator + digitwhere the separator was one of/,.,_,:and the digit was1or2. That collapses a genuinely mismatched pair such asread_1(stream 0) andread_2(stream 1) to the same base nameread, so two different reads are silently accepted as "in sync" instead of being rejected.Fix
Narrow the rule to match fgbio's
FastqSourceread-name canonicalization incom/fulcrumgenomics/fastq/FastqSource.scala, which strips only a trailing/followed by a single ASCII digit (0-9):The trailing space-comment strip (a separate, correct concern) is unchanged. After it,
strip_read_suffixnow strips a suffix only when it is/+ a single ASCII digit. It no longer strips./_/:separators, and it no longer strips multi-digit runs (read/12is left intact). This is the same rule PR #489 established for the extract path (strip_read_number_suffix), now applied to the shared FASTQ-sync helper.New one-line spec: strip a trailing space comment, then strip a trailing
/+ single ASCII digit; leave everything else (./_/:separators, multi-digit runs, non-digit suffixes) intact.Affected callers
All three production callers are paired-FASTQ name-sync validation and benefit uniformly with no call-site changes:
FastqGrouper::drain_complete_templates(src/lib/grouper.rs)zip_fastqsource stepparse_zip_fastqsource step(On this branch's base only
grouper.rsactually calls the shared helper; thezip_fastq/parse_zip_fastqsource steps are not present onmainyet, but they use the same helper and get the fix for free when they land.)Behavior tradeoff (please note)
Narrowing correctly rejects mismatched
read_1/read_2streams that the old rule silently accepted. But it also now rejects non-standard paired FASTQs that legitimately use.1/.2/_1/_2/:1/:2read-number suffixes to distinguish R1/R2. This matches fgbio, which also rejects such names as out of sync, so it is correct fgbio parity — but it is a real, user-visible behavior change reviewers should be aware of. FASTQs using the standard/1//2(or no suffix) are unaffected.Tests
test_strip_read_suffix(fastq_parse.rs) as anrstesttable:/1,/2,/3,/0strip toread;.1/.2/_1/_2/:1/:2are preserved;read/12(multi-digit) andread/x(non-digit) are preserved; a trailing space comment is still stripped first (with and without a following/+digit).test_strip_read_suffix_rejects_mismatched_underscore_pair(fastq_parse.rs) assertingread_1andread_2no longer collide, with a comment documenting the tradeoff.test_fastq_grouper_rejects_mismatched_underscore_pair(grouper.rs) drivingFastqGrouperend-to-end to prove the sync-validation now rejects a mismatchedread_1/read_2pair as "out of sync"./1//2suffixes, which remain valid.Follow-up (out of scope here)
src/lib/commands/extract.rsstill carries its own broadstrip_read_suffix_extract. Consolidating both onto the narrow shared helper is a natural follow-up, intentionally left out to keep this fix focused onstrip_read_suffix.Summary by CodeRabbit
_1and_2from being incorrectly treated as identical.