fix(pipeline): admit the gap-filler serial to Decode under memory-high - #787
Conversation
|
Note Reviews pausedUse the following commands to manage reviews:
Use the checkboxes below for quick actions:
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Pro Run ID: 📒 Files selected for processing (2)
Included review availability: 0 reviews are currently available. Based on recent review activity, included reviews refill at 1 per hour. WalkthroughDecode now checks physical Q3 capacity before popping Q2b. It then applies per-serial admission, requeues future serials, and directly decodes when requeueing fails. Integration tests add read-batch control and validate budget-scaled backpressure. ChangesDecode admission flow
Estimated code review effort: 3 (Moderate) | ~20 minutes Merge Risk: ⚪ Minimal · up to The change allows the Decode step to admit the required gap-filler while continuing to backpressure future work under high memory. No actionable merge-blocking risk remains beyond normal checks and review. Possibly related issues
Possibly related PRs
🚥 Pre-merge checks | ✅ 3✅ Passed checks (3 passed)
Comment |
|
@coderabbitai pause |
✅ Action performedReviews paused. |
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## main #787 +/- ##
=======================================
Coverage 94.36% 94.36%
=======================================
Files 186 186
Lines 113659 113713 +54
=======================================
+ Hits 107251 107304 +53
- Misses 6408 6409 +1 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
b18eb89 to
87c6537
Compare
f840cd7 to
e69b17a
Compare
|
@coderabbitai review |
✅ Action performedReview finished.
|
87c6537 to
d0836df
Compare
e69b17a to
20658c6
Compare
20658c6 to
9bde1d5
Compare
#746) The Decode step gated new work with a blanket `q3_reorder_state.is_memory_high()` check before popping Q2b. Unlike the Decompress step it was copied from — where FIFO Q1 guarantees the reorder buffer's `next_seq` was already produced and is sitting downstream or in a held slot — this is not deadlock-safe for Decode: when a slow writer backs the pipeline up and the Q3 reorder buffer reaches its high-water mark, the guard also stops Decode from decoding the very serial the reorder buffer is waiting on. That serial never reaches Q3, Group starves on it, and the whole pipeline wedges with the reader gated — the exact OOM-avoidance path this PR adds. Replace the blanket guard with the reorder buffer's own per-serial `can_proceed`, applied after the pop: it always admits the gap-filler (`serial == next_seq`), even over the memory limit, and backpressures only future serials once the buffer is half full. A future batch rejected by `can_proceed` is returned to Q2b so another worker can still reach the gap-filler, which avoids the "all workers hold a non-next_seq batch and nobody can produce next_seq" deadlock the code hit when `can_proceed` was previously wired into Decompress. On the healthy path (memory below the high-water mark) `can_proceed` is always true, so throughput is unchanged. The deadlock is scheduling-sensitive: CI's lighter load did not trigger it, so the three backpressure tests passed in CI while reliably hanging on a loaded host. A controlled A/B on that host (same code, only this guard toggled) confirmed the guard as the cause. With this change the four backpressure tests pass on that host across repeated runs.
9bde1d5 to
c95e22a
Compare
|
@coderabbitai review |
✅ Action performedReview finished.
|
The multi-threaded pipeline serialized decode under load: at 16 threads the decode stage ran at ~3.6x effective parallelism instead of ~15x, making consensus/group/filter/sort 1.7x (t8) to 3.3x avg / 8.5x worst (t16) slower than v0.6.0. #787 (fixing the #746 deadlock) gated future-serial admission at `heap_bytes < effective_limit / 2` with `effective_limit = min(--max-memory, 512 MiB)`. The Q3 decoded-record reorder buffer therefore stopped admitting anything but next_seq once it held ~256 MiB -- a ceiling that scales with neither --max-memory nor thread count; with N decode workers it is crossed almost immediately and decode collapses to serial. Bound Q3 admission by serial skew instead, as the reorder buffer's own docstring prescribes: admit next_seq unconditionally (the #746 core), and a future serial iff within [next_seq, next_seq + W) and under a raw memory backstop, with W = clamp(4 * num_threads, .., queue_capacity). Total in-flight memory stays bounded upstream by the Read admission gate, so decode no longer needs a per-stage byte throttle that doubled as a parallelism cap. Scoped to Q3 via a new `window` field (0 = legacy byte-threshold behavior). Output is record-identical to v0.6.0; the #746 backpressure suite passes and a tight --max-memory run stays bounded without wedging. Adds test_reorder_buffer_state_windowed_admission.
The multi-threaded pipeline serialized decode under load: at 16 threads the decode stage ran at ~3.6x effective parallelism instead of ~15x, making consensus/group/filter/sort 1.7x (t8) to 3.3x avg / 8.5x worst (t16) slower than v0.6.0. #787 (fixing the #746 deadlock) gated future-serial admission at `heap_bytes < effective_limit / 2` with `effective_limit = min(--max-memory, 512 MiB)`. The Q3 decoded-record reorder buffer therefore stopped admitting anything but next_seq once it held ~256 MiB -- a ceiling that scales with neither --max-memory nor thread count; with N decode workers it is crossed almost immediately and decode collapses to serial. Bound Q3 admission by serial skew instead, as the reorder buffer's own docstring prescribes: admit next_seq unconditionally (the #746 core), and a future serial iff within [next_seq, next_seq + W) and under a raw memory backstop, with W = clamp(4 * num_threads, .., queue_capacity). Total in-flight memory stays bounded upstream by the Read admission gate, so decode no longer needs a per-stage byte throttle that doubled as a parallelism cap. Scoped to Q3 via a new `window` field (0 = legacy byte-threshold behavior). Output is record-identical to v0.6.0; the #746 backpressure suite passes and a tight --max-memory run stays bounded without wedging. Adds test_reorder_buffer_state_windowed_admission.
The problem: #764's deadlock is not fixed, and green CI is masking it
#764's three
test_pipeline_memory_backpressuretests pass in CI but reliably hang on a loaded host. The hang is a real, scheduling-sensitive deadlock in the Decode step, not flaky tests:q3_reorder_state.is_memory_high()check before popping Q2b (bam.rs,try_step_decode). It was copied from the Decompress step, whose comment argues the guard is deadlock-safe because "Q1 is FIFO, sonext_seqfor the reorder buffer has already been produced."Evidence (controlled A/B on the loaded host, same commit, only this guard toggled):
is_memory_highguardCI's lighter scheduling doesn't trip the race, so #764 goes green — a false negative for this deadlock class.
The fix
Replace the blanket guard with the reorder buffer's own per-serial
can_proceed, applied after the pop:serial == next_seq), even over the memory limit, so the serial Group is waiting on can never be starved.can_proceedis returned to Q2b so another worker can still reach the gap-filler — avoiding the "all workers hold a non-next_seqbatch and nobody can producenext_seq" deadlock the code hit whencan_proceedwas previously wired into Decompress (see that step's comment).On the healthy path (memory below the high-water mark)
can_proceedis alwaystrue, so throughput is unchanged — the new backpressure engages only under the exact memory-high condition that used to deadlock.Validation
cargo ci-fmtandcargo ci-lintclean.Caveat worth raising separately
CI cannot currently catch this deadlock class (it needs constrained scheduling to reproduce). A follow-up that reproduces it under load in CI would keep it from regressing silently.
Risk: output for grouping, consensus, sort order, corrected UMIs, and metrics: none;
unsafe: none, and no CLAUDE.md allowlist update; memory bound, queue capacity, and thread/backpressure policy: changed for Decode under memory-high conditions.Fix: Decode applies per-serial admission after popping Q2b. It requeues future serials and admits the required Q3 gap-filler serial. This prevents stalls while preserving serial order.
Validation: