zipper's single-threaded fast path (no --threads, or --threads 1) got roughly 20% slower between v0.4.0 and v0.5.0 and has stayed there through v0.6.0, while the threaded path (--threads >= 2) got about 29% faster over the same window. The net effect on a threaded invocation is positive, which is why this has not blocked anything, but the non-threaded path is a real regression and it has now persisted across two releases.
Measurements
fgumi zipper, synthetic-pipeline-xlarge, wall-clock seconds, c7g.4xlarge, one job per instance:
| threads |
v0.4.0 (7adda90) |
v0.5.0 (3df1022) |
v0.6.0 (cf41215) |
v0.6.0 vs v0.4.0 |
| 0 (fast path) |
187.2 |
224.9 |
222.5 |
1.19x slower |
| 1 |
186.8 |
223.8 |
224.7 |
1.20x slower |
| 2 |
130.7 |
92.6 |
92.3 |
0.71x |
| 4 |
130.4 |
93.0 |
92.9 |
0.71x |
| 8 |
130.4 |
93.5 |
92.8 |
0.71x |
| 16 |
130.0 |
92.5 |
92.8 |
0.71x |
The split is clean at the threads-0/1 vs threads>=2 boundary, and the magnitude is stable to within ~1% across every thread count, so this is not measurement noise. It reproduced independently in the v0.5.0 benchmark run, in a second functionally identical v0.5.0 run (490a793, trees differ only in a CI workflow file), and now in the v0.6.0 run.
A second sample added in v0.5.0, synthetic-simplex-xlarge, shows the same shape at the same absolute level (threads-0: 218.3s in v0.5.0, 221.9s in v0.6.0; threads-2: 92.2s / 92.0s), so it is not specific to one input.
Worth noting the practical consequence: the fast path is now ~2.4x slower than simply passing --threads 2 on the same input (222.5s vs 92.3s). Before v0.5.0 that gap was 1.43x.
Scope
The regression window is v0.4.0 -> v0.5.0, which is 151 commits. Commits in that range touching zipper:
#644 looks like the most plausible starting point, since generalizing the input path to accept SAM and stdin is the kind of change that can introduce an indirection (trait object, extra buffering layer) whose cost is visible when a single thread does all the work and hidden once decoding is parallelised. That is a hypothesis from the commit titles, not something I have profiled.
Why it matters
--threads is opt-in, so anyone invoking zipper without it — which is the documented single-threaded optimized path — is paying the full 20%.
zipper's single-threaded fast path (no--threads, or--threads 1) got roughly 20% slower between v0.4.0 and v0.5.0 and has stayed there through v0.6.0, while the threaded path (--threads >= 2) got about 29% faster over the same window. The net effect on a threaded invocation is positive, which is why this has not blocked anything, but the non-threaded path is a real regression and it has now persisted across two releases.Measurements
fgumi zipper,synthetic-pipeline-xlarge, wall-clock seconds, c7g.4xlarge, one job per instance:7adda90)3df1022)cf41215)The split is clean at the threads-0/1 vs threads>=2 boundary, and the magnitude is stable to within ~1% across every thread count, so this is not measurement noise. It reproduced independently in the v0.5.0 benchmark run, in a second functionally identical v0.5.0 run (
490a793, trees differ only in a CI workflow file), and now in the v0.6.0 run.A second sample added in v0.5.0,
synthetic-simplex-xlarge, shows the same shape at the same absolute level (threads-0: 218.3s in v0.5.0, 221.9s in v0.6.0; threads-2: 92.2s / 92.0s), so it is not specific to one input.Worth noting the practical consequence: the fast path is now ~2.4x slower than simply passing
--threads 2on the same input (222.5s vs 92.3s). Before v0.5.0 that gap was 1.43x.Scope
The regression window is v0.4.0 -> v0.5.0, which is 151 commits. Commits in that range touching zipper:
eb8d71a5feat(io): accept uncompressed SAM and stdin anywhere BAM is accepted (feat(io): accept uncompressed SAM and stdin anywhere BAM is accepted #644)1db2c5b9fix(zipper): name the tag-skipping flag after the tag it actually writes (fix(zipper): name the tag-skipping flag after the tag it actually writes #617)7ce17634docs: clear fgbio-parity doc tail across commands (W11) (docs: clear fgbio-parity doc tail across commands (W11) #558)8cdf1b82fix(zipper): correct consensus reverse/revcomp tag sets to match fgbio ConsensusTags (fix(zipper): correct consensus reverse/revcomp tag sets to match fgbio ConsensusTags #488)#644looks like the most plausible starting point, since generalizing the input path to accept SAM and stdin is the kind of change that can introduce an indirection (trait object, extra buffering layer) whose cost is visible when a single thread does all the work and hidden once decoding is parallelised. That is a hypothesis from the commit titles, not something I have profiled.Why it matters
--threadsis opt-in, so anyone invokingzipperwithout it — which is the documented single-threaded optimized path — is paying the full 20%.