Skip to content

feat(retag)!: add pair op to build a paired UMI from own/mate tags - #985

Merged
nh13 merged 1 commit into
mainfrom
984/nh/feat-retag-pair-op
Sep 27, 2026
Merged

nh13 merged 1 commit into
mainfrom
984/nh/feat-retag-pair-op

Conversation

@nh13

@nh13 nh13 commented Sep 27, 2026 •

Copy link
Copy Markdown
Member

Closes #984.

Motivation

Published NanoSeq data (e.g. from the Sanger pipeline) is often shared only as aligned CRAMs. In those files the inline barcode is already trimmed off each read and carried in per-read tags: rb (this read's barcode) and mb (its mate's), instead of a paired RX. There was no fgumi tool to turn that into the RX that group --strategy paired needs, so users had to convert back to FASTQ and re-run extraction and alignment.

What this adds

A pair operation for fgumi retag:

fgumi retag -i in.bam -o out.bam rb,mb::pair::RX rb::delete mb::delete

OWN,MATE::pair::DST joins the two string tags into the R1-first paired UMI: OWN-MATE on every record except the second read of a pair (paired and not first-of-pair, the same rule group uses to pick a template's R1), which gets MATE-OWN. Both mates of a template therefore carry the identical value that fgumi extract would have written. This matters beyond group: duplex consensus builds its consensus RX from every source read and reverses by segment, so mates must agree.

Semantics:

  • Sources are kept; compose with SRC::delete to drop them.
  • A record missing either source skips the op (counted in src_missing); no partial join is ever written.
  • A source that is present but not a string (Z) tag is an error that stops the run.
  • Exactly two distinct sources; DST must differ from both. Only pair accepts a comma-separated source list.
  • The delimiter is -, matching what group splits on.

The NanoSeq guide gains a "Starting from Shared NanoSeq CRAMs" section: samtools view -b → retag → sort --order template-coordinate → group --strategy paired --edits 0 → duplex, with a note that the input must be the full (duplicate-marked, not deduplicated) CRAM and carry MC.

Design notes

  • A new verb rather than letting copy/move take a source list: a fixed-order concat would give R2 the swapped UMI, and overloading copy would give it two different semantics depending on source count.
  • Degenerate flag combinations (e.g. 0x80 without 0x1) follow group's R1 rule, so pair and group always agree on which read is R1.

Breaking change

RetagOp (public in fgumi_lib) gains a Pair variant, and its public src() accessor is replaced by a crate-private sources(). No in-tree callers outside retag are affected.

Known limitation

If a run fails on a non-string pair source, the output written up to that point is left on disk and is incomplete. This is the pipeline engine's general behavior on a mid-run step failure (not specific to retag), and is documented in retag --help.

Testing

  • Unit tests: parsing and printing of pair, a message-level check for every malformed pair / source-list input, R1/R2/fragment/secondary/supplementary and degenerate-flag ordering, missing sources, DST overwrite, non-string sources (including a non-string own with a missing mate), and the zero-match warning wording.
  • Integration tests: the two example records from accept Nanoseq preprocessed reads from CRAM for grouping #984 end up with RX:Z:GTT-CTA on both mates at 1 and 2 threads, with the metrics row checked; a non-string source fails the run at both thread counts.
  • cargo ci-fmt, ci-lint, ci-tag-literals, and ci-test (10,475 passed) all pass locally.

Risk: retag output changes by adding paired RX tags; ordering is pinned by unit and integration tests. Unsafe code: none added or changed; CLAUDE.md needs no allowlist update. Memory bounds, queue capacity, and thread/backpressure policy: none changed.

  • Add OWN,MATE::pair::DST to fgumi retag. It joins string tags in R1-first order, reverses their order on paired second reads, skips records with a missing source, and errors if a present source is not a string. Sources remain unless separately deleted.
  • Document NanoSeq CRAM processing with rb and mb tags. The guide specifies retaining the full duplicate-marked CRAM and requiring MC tags.
  • Add public RetagOp::Pair; replace the public src() accessor with crate-private sources().
  • The author reports that formatting, lint, tag-literal, and test checks pass.

Shared NanoSeq CRAMs carry each read's barcode in `rb` and its mate's in
`mb` rather than a paired `RX`, so they could not be grouped without
converting back to FASTQ and re-running extraction and alignment.

`OWN,MATE::pair::DST` joins the two string tags into the R1-first paired
UMI: `OWN-MATE` on every record except the second read of a pair (paired,
not first-of-pair, the same rule `group` uses to pick R1), which gets
`MATE-OWN`. Both mates therefore carry the value `fgumi extract` would
have written, which `group --strategy paired` and duplex consensus rely
on. Sources are kept (compose with `SRC::delete`); a record missing either
source skips the op; a non-string source is an error. Only `pair` accepts
a comma-separated source list.

Also adds a NanoSeq guide section for starting from shared CRAMs.

Closes #984

BREAKING CHANGE: `RetagOp` gains a `Pair` variant, and its public `src()`
accessor is replaced by a crate-private `sources()`.
@nh13
nh13 deployed to github-actions September 27, 2026 03:00 — with GitHub Actions Active
@coderabbitai

coderabbitai Bot commented Sep 27, 2026 •

Copy link
Copy Markdown

Review in Change Stack →

Navigate logical layers of code changes, visualize relationships, and explore their blast radius.

Note

Reviews paused

Use the following commands to manage reviews:

  • @coderabbitai resume to resume automatic reviews.
  • @coderabbitai review to trigger a single review.

Use the checkboxes below for quick actions:

  • ▶️ Resume reviews
  • 🔍 Trigger review

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Repository: fulcrumgenomics/fgumi/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Essentials

Run ID: 456131b2-7433-4eec-a4ef-2241ec4c2ed3

📥 Commits

Reviewing files that changed from the base of the PR and between 816d649 and 5e33ef2.

⛔ Files ignored due to path filters (1)
  • CHANGELOG.md is excluded by !**/CHANGELOG.md
📒 Files selected for processing (6)
  • README.md
  • docs/src/guide/nanoseq.md
  • docs/src/index.md
  • src/lib/commands/retag.rs
  • src/lib/pipeline/chains/commands/retag.rs
  • tests/integration/test_retag_command.rs

Included review availability: This review used your included allowance. 1 included review remains after this review. Your included PR review attempts over the past 7 days set your current allowance at 2 reviews per hour.


Walkthrough

The retag command now accepts paired source tags, joins their string values into a destination tag, and propagates invalid-value errors. The NanoSeq guide documents using this operation to build RX tags from shared CRAMs before grouping and duplex consensus calling.

Changes

Paired-tag retagging

Layer / File(s) Summary
Pair operation syntax and contract
src/lib/commands/retag.rs, README.md, docs/src/index.md
OWN,MATE::pair::DST accepts two distinct source tags and a distinct destination. CLI help and parser tests cover the syntax and validation rules.
Pair source values into destination tags
src/lib/commands/retag.rs
The operation joins string source values in read order, reverses the order for paired non-first reads, skips records with missing sources, and updates applied and overwrite counts.
Pipeline handling and NanoSeq workflow
src/lib/pipeline/chains/commands/retag.rs, tests/integration/test_retag_command.rs, docs/src/guide/nanoseq.md
The pipeline propagates pairing errors and reports all source tags in zero-match warnings. Integration tests cover successful pairing and non-string source errors. The NanoSeq guide documents pairing rb and mb into RX before grouping and consensus calling.

Priority: ⬇️ Low

Estimated code review effort: 3 (Moderate) | ~25 minutes

Change: Feature

Sequence Diagram(s)

sequenceDiagram
  participant RetagPipeline
  participant ApplyOp
  participant PairTags
  RetagPipeline->>ApplyOp: apply operation to record
  ApplyOp->>PairTags: pair source tag values
  PairTags->>PairTags: write joined destination tag or report missing source
  PairTags-->>ApplyOp: return result
  ApplyOp-->>RetagPipeline: return success or error
Loading

Merge Risk: ⚪ Minimal · up to 5e33e

The documented NanoSeq workflow retains duplicate-marked reads for family consensus. No actionable merge-blocking risk remains after normal checks.

🚥 Pre-merge checks | ✅ 3
✅ Passed checks (3 passed)
Check name Status Explanation
Title check ✅ Passed The title follows the required Conventional Commit format, uses the valid feat type and retag scope, starts with a lowercase imperative description, and accurately describes the main change.
Linked Issues check ✅ Passed Issue #984 requires reuse of aligned NanoSeq CRAM reads without FASTQ regeneration. RetagOp::Pair combines per-read rb and mate mb tags into R1-first RX values, reverses order for the second r…
Out of Scope Changes check ✅ Passed The reviewed changes stay connected to issue #984. The retag CLI and help updates expose the required operation. The sources() API change supports multiple sources for pair and crate-internal op…

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Sep 27, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 96.15%. Comparing base (816d649) to head (5e33ef2).

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #985      +/-   ##
==========================================
- Coverage   96.19%   96.15%   -0.04%     
==========================================
  Files         294      294              
  Lines      148023   148109      +86     
==========================================
+ Hits       142388   142412      +24     
- Misses       5635     5697      +62     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@nh13

nh13 commented Sep 27, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai pause

@nh13

nh13 commented Sep 27, 2026

Copy link
Copy Markdown
Member Author

@coderabbitai review

@coderabbitai

coderabbitai Bot commented Sep 27, 2026

Copy link
Copy Markdown
✅ Action performed

Reviews paused.

@coderabbitai

coderabbitai Bot commented Sep 27, 2026 •

Copy link
Copy Markdown
✅ Action performed

Review finished.

Note: CodeRabbit is an incremental review system and does not re-review already reviewed commits. This command is applicable only when automatic reviews are paused.

@nh13
nh13 added this pull request to the merge queue Sep 27, 2026
Merged via the queue into main with commit 1896e97 Sep 27, 2026
17 checks passed
@nh13
nh13 deleted the 984/nh/feat-retag-pair-op branch September 27, 2026 06:49

This branch was successfully deployed

1 active deployment
github-actions — 5e33ef29 Deployed Sep 27, 2026 by nh13 via coverage #4617
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

accept Nanoseq preprocessed reads from CRAM for grouping

1 participant