Skip to content

fix(extract): match fgbio strict read-name UMI extraction (EXT-01/03/04) - #489

Merged
nh13 merged 1 commit into
mainfrom
nh/fix-extract-readname-umi-strict
Jul 10, 2026
Merged

nh13 merged 1 commit into
mainfrom
nh/fix-extract-readname-umi-strict

Conversation

@nh13

@nh13 nh13 commented Jul 8, 2026 •

Copy link
Copy Markdown
Member

Summary

fgumi extract --extract-umis-from-read-names is fgumi's port of fgbio's FastqToBam -n, which calls Umis.extractUmisFromReadName(name, strict=true). fgbio's strict path upper-cases the extracted UMI, reverse-complements r-prefixed segments, and throws on any character outside ACGTN-. fgumi previously wrote the last read-name field to RX verbatim — no upper-casing, no r-revcomp, and it silently accepted illegal UMI characters.

This closes three findings from the fgbio behavioral-parity tracker:

  • EXT-03 — upper-case the read-name UMI (Umis.scala:111-117). Case-sensitive UMIs otherwise mis-group vs fgbio.
  • EXT-04 — reverse-complement r-prefixed segments (Umis.scala:102-115, reverseComplementPrefixedUmis defaults true, as FastqToBam uses it). For a +-delimited dual UMI only the prefixed segment is reverse-complemented. Reuses the shared fgumi_dna::dna::reverse_complement helper rather than a new copy.
  • EXT-01 — reject (return an error) a UMI containing any character outside ACGTN-, matching fgbio strict mode (Umis.scala:121-123, UmisTest.scala:110-114). The surfaced error names the offending read.

Out of scope by decision: EXT-02 — fgumi's lenient >=8-field / last-field extraction (which also supports demultiplexers that append a sample index, producing 9+ fields) is intentionally kept (tracker Decisions Q2). The field-count logic is unchanged; only the transform + validation of the extracted field now matches fgbio.

The new is_valid_umi_char mirrors fgbio's Umis.isValidUmiCharacter (ACGTN-) and is deliberately distinct from fgumi_umi::validate_umi, which implements the more permissive GroupReadsByUmi counting rule (only rejects upper-case N).

fgbio ↔ fgumi parity evidence

Adversarial fixture (8-field read names, so fgbio strict and fgumi both select the same last field), run against fgbio 4.1.0:

read-name UMI fgbio FastqToBam -n fgumi before fgumi after
…:acgtn ACGTN acgtn ACGTN
…:rGGTTAA TTAACC rGGTTAA TTAACC
…:acgt+rGGTTAA ACGT-TTAACC acgt-rGGTTAA ACGT-TTAACC
…:ACGTXY (illegal) rejected (rc=1) accepted (rc=0) rejected (rc=1)

Before: RESULT: DIFFER. After: RX tags match fgbio exactly and the illegal UMI is rejected like fgbio:

Error: extracting UMI from read name 'inst:1:fc:1:1:4:1:ACGTXY'
Caused by:
    Invalid UMI 'ACGTXY' extracted from read name (illegal character 'X')

Testing

  • TDD: new unit tests for upper-casing, r-prefix revcomp (single + dual-UMI), illegal-char rejection (ACGTXY, ACGT-CCKC, CCKC-ACGT per UmisTest.scala), a valid-UMI boundary, and lowercase +-delimited normalization.
  • cargo ci-fmt, cargo ci-lint, cargo ci-test (2221 passed) all clean.
  • End-to-end fgbio-vs-fgumi reproduction captured before (DIFFER) and after (MATCH).

Refs: EXT-01, EXT-03, EXT-04 (fgbio behavioral parity tracker; PR 2 of the burn-down).

Summary by CodeRabbit

  • New Features

    • DNA complement utilities now treat RNA bases U/u as complements to A/a, including in reverse-complement behavior.
    • Read-name UMI extraction now applies stricter normalization, including uppercase handling, r-segment reverse-complementing, and dual-UMI separator translation.
  • Bug Fixes

    • Extraction now fails with clear errors when extracted UMIs contain characters outside the allowed ACGTN- set.
  • Tests & Documentation

    • Expanded unit tests and updated command documentation to reflect the new U/u complement behavior and stricter UMI semantics.

@nh13
nh13 temporarily deployed to github-actions July 8, 2026 20:44 — with GitHub Actions Inactive
@coderabbitai

coderabbitai Bot commented Jul 8, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: 2468fe21-bf01-4314-b44c-335fd7e89d2d

📥 Commits

Reviewing files that changed from the base of the PR and between f165260 and 9879c95.

📒 Files selected for processing (2)
  • crates/fgumi-dna/src/dna.rs
  • src/lib/commands/extract.rs

Walkthrough

Uracil now complements to adenine, while read-name UMI extraction uses fallible strict normalization with reverse-complement handling, delimiter conversion, upper-casing, and ACGTN- validation.

Changes

DNA complement U/u support

Layer / File(s) Summary
Complement functions and documentation
crates/fgumi-dna/src/dna.rs
Complement and reverse-complement behavior and documentation now include U/u mappings to A/a.
U/u complement tests
crates/fgumi-dna/src/dna.rs
Tests verify U/u handling in base and reverse-complement operations.

Read-name UMI extraction strictness

Layer / File(s) Summary
Normalization contract and extraction API
src/lib/commands/extract.rs
The helper now accepts &[u8], returns Result, and documents strict normalization behavior.
UMI normalization and validation
src/lib/commands/extract.rs
UMIs apply selective reverse-complementing, +-to-- conversion, upper-casing, and strict ACGTN- validation.
Caller propagation and regression coverage
src/lib/commands/extract.rs
Both record-building paths propagate errors, and tests cover parsing invariants, normalization, RNA U handling, and invalid characters.

Estimated code review effort: 4 (Complex) | ~45 minutes

Sequence Diagram(s)

sequenceDiagram
  participant RecordBuilder
  participant extract_read_name_and_umi
  participant reverse_complement

  RecordBuilder->>extract_read_name_and_umi: read name bytes
  alt r-prefixed segment
    extract_read_name_and_umi->>reverse_complement: segment bytes
    reverse_complement-->>extract_read_name_and_umi: complemented bytes
  end
  extract_read_name_and_umi->>extract_read_name_and_umi: replace + with -, uppercase, validate ACGTN-
  alt invalid character
    extract_read_name_and_umi-->>RecordBuilder: error
  else valid
    extract_read_name_and_umi-->>RecordBuilder: normalized UMI
  end
Loading

Possibly related PRs

Suggested labels: bug, fgumi extract

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: stricter fgbio-compatible read-name UMI extraction in extract.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch nh/fix-extract-readname-umi-strict

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Jul 8, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.59036% with 2 lines in your changes missing coverage. Please review.
✅ Project coverage is 91.08%. Comparing base (f3116d5) to head (9879c95).
⚠️ Report is 3 commits behind head on main.

Files with missing lines Patch % Lines
src/lib/commands/extract.rs 97.59% 2 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main     #489      +/-   ##
==========================================
- Coverage   91.14%   91.08%   -0.06%     
==========================================
  Files          78       78              
  Lines       51606    51726     +120     
==========================================
+ Hits        47034    47116      +82     
- Misses       4572     4610      +38     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@nh13
nh13 force-pushed the nh/fix-extract-readname-umi-strict branch from 0b4a291 to a9e4a9a Compare July 8, 2026 21:16
@nh13
nh13 temporarily deployed to github-actions July 8, 2026 21:17 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 8, 2026

Copy link
Copy Markdown
Member Author

Follow-up (amended into the same commit): made the r-prefix reverse-complement path exactly faithful to fgbio by teaching fgumi_dna's complement to map RNA uracil U/u → A/a (fgbio Sequences.complement). Previously reused reverse_complement complemented only ACGT/N, so an r-prefixed UMI containing U was rejected where fgbio reverse-complements it.

Verified against fgbio 4.1.0:

  • rAAU → ATT, rU → A (revcomp path, U→A): now MATCH.
  • non-r ACGU: rejected by both (no revcomp, so U stays and fails ACGTN- validation).

cargo ci-fmt / ci-lint / ci-test (2222 passed) all clean.

@nh13

nh13 commented Jul 9, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 9, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 9, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 9, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/lib/commands/extract.rs`:
- Around line 2517-2531: The current
`test_extract_read_name_and_umi_rejects_illegal_chars` uses a manual loop over
multiple illegal inputs, which hides the failing case and does not follow the
Rust test convention. Convert this into a parameterized `#[rstest]` test with
separate cases for each illegal UMI input, keeping the assertion against
`Extract::extract_read_name_and_umi` in the same test logic so each input is
reported independently.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: dc7766dc-44e9-41f1-9f07-5f2a57ceacd1

📥 Commits

Reviewing files that changed from the base of the PR and between f638abc and a9e4a9a.

📒 Files selected for processing (2)
  • crates/fgumi-dna/src/dna.rs
  • src/lib/commands/extract.rs

Comment thread src/lib/commands/extract.rs Outdated
@nh13
nh13 force-pushed the nh/fix-extract-readname-umi-strict branch from a9e4a9a to 50f36c4 Compare July 9, 2026 17:00
@nh13
nh13 temporarily deployed to github-actions July 9, 2026 17:00 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 9, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 9, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13
nh13 force-pushed the nh/fix-extract-readname-umi-strict branch from 50f36c4 to f165260 Compare July 9, 2026 21:10
@nh13
nh13 temporarily deployed to github-actions July 9, 2026 21:10 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 9, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 9, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/lib/commands/extract.rs`:
- Around line 2539-2546: The test for Extract::extract_read_name_and_umi
currently loops over multiple headers, which makes it hard to see which input
failed; convert it to parameterized #[rstest] cases like
test_extract_read_name_and_umi_rejects_illegal_chars. Keep the same assertions
for extract_read_name_and_umi, but split the inputs (`@sample_1`, `@foo.2`, `@bar`:1)
into separate rstest cases so each regression reports the exact header that
broke.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: c3f688a8-3b50-46d8-bb2d-67b03a3f318d

📥 Commits

Reviewing files that changed from the base of the PR and between 50f36c4 and f165260.

📒 Files selected for processing (2)
  • crates/fgumi-dna/src/dna.rs
  • src/lib/commands/extract.rs

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Caution

Inline review comments failed to post. This is likely due to GitHub's internal server error or limits when posting large numbers of comments. If you are seeing this consistently it is likely a permissions issue. Please check "Moderation" -> "Code review limits" under your organization settings.

Actionable comments posted: 1

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@src/lib/commands/extract.rs`:
- Around line 2539-2546: The test for Extract::extract_read_name_and_umi
currently loops over multiple headers, which makes it hard to see which input
failed; convert it to parameterized #[rstest] cases like
test_extract_read_name_and_umi_rejects_illegal_chars. Keep the same assertions
for extract_read_name_and_umi, but split the inputs (`@sample_1`, `@foo.2`, `@bar`:1)
into separate rstest cases so each regression reports the exact header that
broke.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: c3f688a8-3b50-46d8-bb2d-67b03a3f318d

📥 Commits

Reviewing files that changed from the base of the PR and between 50f36c4 and f165260.

📒 Files selected for processing (2)
  • crates/fgumi-dna/src/dna.rs
  • src/lib/commands/extract.rs
🛑 Comments failed to post (1)
src/lib/commands/extract.rs (1)

2539-2546: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Loop-over-inputs test hides which input regressed; convert to #[rstest] cases. Same convention already applied to test_extract_read_name_and_umi_rejects_illegal_chars.

♻️ Convert to rstest
-    #[test]
-    fn test_extract_read_name_and_umi_does_not_strip_dot_or_underscore_suffix() {
-        // Only `/` is a read-number separator; `.`/`_`/`:` may be part of a real
-        // name and must be preserved (unlike the broader validation helper).
-        for header in [b"`@sample_1`".as_slice(), b"`@foo.2`".as_slice(), b"`@bar`:1".as_slice()] {
-            let (name, _umi) = Extract::extract_read_name_and_umi(header, false).unwrap();
-            assert_eq!(name, header[1..].to_vec(), "should not strip: {header:?}");
-        }
-    }
+    // Only `/` is a read-number separator; `.`/`_`/`:` may be part of a real
+    // name and must be preserved (unlike the broader validation helper).
+    #[rstest]
+    #[case::underscore(b"`@sample_1`".as_slice())]
+    #[case::dot(b"`@foo.2`".as_slice())]
+    #[case::colon(b"`@bar`:1".as_slice())]
+    fn test_extract_read_name_and_umi_does_not_strip_dot_or_underscore_suffix(#[case] header: &[u8]) {
+        let (name, _umi) = Extract::extract_read_name_and_umi(header, false).unwrap();
+        assert_eq!(name, header[1..].to_vec(), "should not strip: {header:?}");
+    }

As per coding guidelines: "Use rstest for parameterized tests in Rust test files".

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/lib/commands/extract.rs` around lines 2539 - 2546, The test for
Extract::extract_read_name_and_umi currently loops over multiple headers, which
makes it hard to see which input failed; convert it to parameterized #[rstest]
cases like test_extract_read_name_and_umi_rejects_illegal_chars. Keep the same
assertions for extract_read_name_and_umi, but split the inputs (`@sample_1`,
`@foo.2`, `@bar`:1) into separate rstest cases so each regression reports the exact
header that broke.

Source: Coding guidelines

`fgumi extract --extract-umis-from-read-names` mirrors fgbio's `FastqToBam -n`,
which calls `Umis.extractUmisFromReadName(name, strict=true)`. That path
upper-cases the extracted UMI, reverse-complements `r`-prefixed segments, and
throws on any character outside `ACGTN-`. fgumi previously wrote the last
read-name field to `RX` verbatim: no upper-casing, no `r`-revcomp, and it
silently accepted illegal UMI characters.

- EXT-03: upper-case the read-name UMI.
- EXT-04: reverse-complement `r`-prefixed segments (for a `+`-delimited dual
  UMI, only the prefixed segment is reverse-complemented); reuse the shared
  `fgumi_dna::dna::reverse_complement` helper.
- EXT-01: reject (error) a UMI containing any character outside `ACGTN-`,
  matching fgbio strict mode, and surface the offending read name.

Also teach `fgumi_dna`'s complement to map RNA uracil `U`/`u` -> `A`/`a`
(fgbio `Sequences.complement`), so an `r`-prefixed UMI containing `U` is
reverse-complemented identically to fgbio (`rAAU` -> `ATT`). A `U` in a non-`r`
UMI is not reverse-complemented and is rejected by both tools.

EXT-02 (lenient `>=8`-field / last-field extraction) is intentionally kept, so
the field-count logic is unchanged. `is_valid_umi_char` mirrors fgbio's
`Umis.isValidUmiCharacter` (`ACGTN-`) and is deliberately distinct from
`fgumi_umi::validate_umi`, which implements the more permissive
`GroupReadsByUmi` counting rule.

fgbio<->fgumi parity (ext-01-03-04-readname-umi.sh, fgbio 4.1.0):
  before: RX `acgtn` / `rGGTTAA` / `acgt-rGGTTAA` vs fgbio
          `ACGTN` / `TTAACC` / `ACGT-TTAACC` (DIFFER); illegal `ACGTXY`
          accepted (rc=0) vs fgbio rejected (rc=1)
  after:  RX tags MATCH fgbio exactly (incl. `rAAU` -> `ATT`); illegal
          `ACGTXY` rejected (rc=1)

Refs: EXT-01, EXT-03, EXT-04 (fgbio behavioral parity tracker).
@nh13
nh13 force-pushed the nh/fix-extract-readname-umi-strict branch from f165260 to 9879c95 Compare July 9, 2026 22:20
@nh13
nh13 temporarily deployed to github-actions July 9, 2026 22:20 — with GitHub Actions Inactive
@nh13

nh13 commented Jul 9, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 9, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13

nh13 commented Jul 10, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Jul 10, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13
nh13 merged commit 0371ce2 into main Jul 10, 2026
11 checks passed
@nh13
nh13 deleted the nh/fix-extract-readname-umi-strict branch July 10, 2026 00:45
@nh13 nh13 mentioned this pull request Jul 9, 2026
nh13 added a commit that referenced this pull request Jul 11, 2026
`strip_read_suffix` canonicalizes read names before checking that
paired/interleaved FASTQ streams are in sync (that the reads at the same
position across streams share a name). Its old rule was too lenient:
after stripping a trailing space-separated comment it also stripped a
trailing `separator + digit` for separator in {`/`, `.`, `_`, `:`} and
digit in {`1`, `2`}. That collapses a genuinely-mismatched pair such as
`read_1` (stream 0) and `read_2` (stream 1) to the same base name
`read`, so two different reads are silently accepted as "in sync".

Narrow the rule to match fgbio's `FastqSource` read-name
canonicalization (`com/fulcrumgenomics/fastq/FastqSource.scala`), which
strips only a trailing `/` followed by a single ASCII digit (0-9). The
space-comment strip is unchanged; `.`/`_`/`:` separators and multi-digit
runs (`read/12`) are now preserved. This is the same rule PR #489
established for the extract path, applied here to the shared
FASTQ-sync helper.

All three production callers are paired-FASTQ name-sync validation and
benefit uniformly with no call-site changes: `FastqGrouper`
(`grouper.rs`), `zip_fastq`, and `parse_zip_fastq`.

Behavior tradeoff: narrowing correctly rejects mismatched `read_1` /
`read_2` streams, but it ALSO now rejects non-standard paired FASTQs
that legitimately use `.1`/`.2`/`_1`/`_2`/`:1`/`:2` read-number
suffixes. This matches fgbio (which rejects them too), so it is correct
fgbio parity, but it is a real behavior change reviewers should note.

Future cleanup: `extract.rs` still carries its own broad
`strip_read_suffix_extract`; consolidating both onto the narrow shared
helper is a follow-up left out of this focused fix.

This branch was previously deployed

1 inactive deployment
github-actions — 9879c955 Deployed Jul 9, 2026 by nh13 via coverage #2109
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant