Repository navigation
feat(simulate): add the aligner replay subcommand - #935
Conversation
`fgumi simulate aligner` is a fake streaming aligner for benchmarking `fgumi runall`'s align stage. Invoked as the aligner subprocess via `--aligner::command`, it drains the interleaved FASTQ on stdin (discarding it) while concurrently streaming a pre-captured aligned BAM (`--replay-bam`) verbatim to stdout on independent threads, so it never deadlocks the align stage the way a single-threaded lockstep replay does. It does no real alignment; it replays alignments captured earlier (e.g. by bwa mem). The emitter locks stdout once and block-buffers it: the replay is BGZF, dense in newline bytes, and the default line-buffered stdout would split nearly every copy chunk at its last newline, roughly doubling the write syscall count in a tool whose whole purpose is to add minimal overhead. Wires the command into SimulateCommand, documents it in docs/simulate-cli.md (including the --replay-bam-before-positionals ordering constraint), and adds unit plus subprocess integration coverage of the concurrent drain/emit, broken-pipe, and error paths. The integration module is gated on the simulate feature so a non-simulate build compiles it out rather than failing.
|
Note Reviews pausedUse the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Essentials Run ID: 📒 Files selected for processing (5)
Included review availability: 0 reviews are currently available. Your included PR review attempts over the past 7 days set your current allowance at 1 review per hour. WalkthroughAdds ChangesSimulated aligner command
Priority: ⬇️ Low Estimated code review effort: 3 (Moderate) | ~30 minutes Change: Feature Merge Risk: ⚪ Minimal · up to This adds a simulated aligner that concurrently drains FASTQ input while replaying BAM output for benchmarking. Its intended streaming, error, and argument-handling behavior is covered, with no remaining merge-blocking risk identified. Sequence Diagram(s)sequenceDiagram
participant SimulateCommand
participant Aligner
participant StdinDrainer
participant Stdout
SimulateCommand->>Aligner: execute replay command
Aligner->>StdinDrainer: drain FASTQ stdin concurrently
Aligner->>Stdout: replay BAM bytes
Stdout-->>Aligner: flush or report pipe result
Aligner->>StdinDrainer: join after successful output
Possibly related PRs
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
Comment |
|
@coderabbitai pause |
✅ Action performedReviews paused. |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #935 +/- ##
==========================================
- Coverage 94.06% 94.06% -0.01%
==========================================
Files 302 303 +1
Lines 152561 152718 +157
==========================================
+ Hits 143503 143648 +145
- Misses 9058 9070 +12 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
|
@coderabbitai review |
✅ Action performedReview finished.
|
What
Adds
fgumi simulate aligner, a fake streaming aligner for benchmarkingfgumi runall's align stage. It is invoked as the aligner subprocess via--aligner::command: it drains the interleaved FASTQ on stdin (discarding it) while concurrently streaming a pre-captured aligned BAM (--replay-bam) verbatim to stdout. Reading and writing run on independent threads, so it never deadlocksAlignAndMergethe way a single-threaded lockstep replay (the shellreplay-aligner.shit replaces) does. It performs no real alignment — it replays alignments captured earlier (e.g. bybwa mem), letting a benchmark measure fgumi's own pipeline overhead without a real aligner's compute dominating every run.Why
The benchmark needs a drop-in for a real streaming aligner that adds as little overhead as possible. The concurrent drain/emit design is exactly what a real aligner does; the single-threaded shell replay it replaces deadlocked the align stage (the writer fills its in-flight gate feeding FASTQ and blocks, while the reader stalls waiting for aligner output the mid-drain shell can't emit).
Details
bwa mememits it), which is howAlignAndMergeregroups records into templates by queryname.{ref}/{threads}a command template substitutes) are accepted and ignored, so the command is a drop-in for real-aligner templates.--replay-bammust precede any positional token — documented on the field, indocs/simulate-cli.md, and pinned by a test.Testing
simulatefeature so a non-simulatebuild compiles it out rather than failing (matchingtest_simulate_sort.rs).cargo ci-fmt,cargo ci-lint(clippy pedantic), and the aligner unit + integration tests all pass.Scope
Ported from the
feat-runallbranch ontomain, deliberately limited to the aligner command: the existing delegatingregion_to_binis left untouched, the separatesimulate sortcommand is not included, and only the aligner-relevantdocs/simulate-cli.mdadditions are ported.Risk: output changes only for the new
fgumi simulate alignercommand, pinned by verbatim replay of the supplied BAM;unsafechanges: none; memory, queue, and thread/backpressure policy: changed through concurrent FASTQ draining and block-buffered BAM output.Adds
fgumi simulate alignerwith replay-file validation, trailing argument support, concurrent FASTQ draining, broken-pipe handling, and integration coverage. Updates the CLI documentation and command dispatch.