Skip to content

test(sort): fix flaky active_worker_limit_caps_then_reactivates - #646

Merged
nh13 merged 1 commit into
mainfrom
nh/fix-flaky-active-worker-limit-test
Jul 22, 2026
Merged

nh13 merged 1 commit into
mainfrom
nh/fix-flaky-active-worker-limit-test

Conversation

@nh13

@nh13 nh13 commented Jul 22, 2026 •

Copy link
Copy Markdown
Member

active_worker_limit_caps_then_reactivates (added in #614) is flaky and is failing CI on unrelated PRs — it turned up on #616, whose diff is a long_about string in codec.rs and cannot reach fgumi-sort:

assertion `left == right` failed: all jobs across both batches accounted for;
per-worker counts [231, 129, 57, 17, 134, 31]
  left: 599
 right: 600

Root cause

The test drains the compress result channel, joins the submit thread, and then reads per_thread_step_counts. Draining that channel does not mean the counters are up to date.

A worker sends the job's result — and drops the job, which closes that job's sender — inside handle_compress_job, i.e. inside execute_step. The counter is incremented afterwards, by record_step, once execute_step has returned:

let result = Self::execute_step(shared, worker, step);   // sends result, drops job's sender
if result == StepResult::Success {
    pstats.record_step(worker.worker_id, step, ...);     // counter incremented here
    did_work = true;
}

So result_rx.recv() returns Err — the loop the test uses to know the batch is done — as soon as the last job's sender is dropped, which is strictly before that job's counter increment. The main thread can then read the counters while a worker is still in the gap, and the sum comes up short by one per worker caught mid-gap. The counts in the CI failure sum to exactly 599.

Evidence

Not inferred from reading alone. Inserting a 200µs sleep at precisely that gap — between execute_step returning Success and record_step — makes the test fail on every run:

left: 298
right: 300

With this change in place and that sleep still inserted, the test passes. The sleep was then removed.

(The window is narrow enough that the test passed 180/180 unmodified runs locally, including 120 under 8-way CPU contention, which is why widening the window was necessary to demonstrate it rather than fish for a failure.)

Fix

Wait for the counters to settle instead of reading a value that is still in flight: after draining the results, poll until the summed counters reach the cumulative job total, with a 30s deadline that fails with the observed per-worker counts if they never do.

The counters are stats-only, and in production they are read after shutdown() has joined the workers — so this is a test synchronization defect, not a pool bug. Reordering record_step before the send to suit the test would change production behaviour (and the recorded duration would then exclude the send) for no benefit, so the fix stays in the test.

cargo nextest run -p fgumi-sort: 559 passed. cargo ci-fmt, cargo ci-lint clean.

Summary by CodeRabbit

  • Tests
    • Improved reliability of worker pool tests by allowing time for compression activity counters to reach their expected cumulative values.
    • Added timeout-based waiting to reduce flaky test results during asynchronous processing.

`active_worker_limit_caps_then_reactivates` reads
`per_thread_step_counts` as soon as it has drained the compress result
channel, but that channel closing does not mean the counters are up to
date.

A worker sends the job's result — and drops the job, closing its sender —
inside `handle_compress_job`, i.e. inside `execute_step`. The counter is
incremented by `record_step` only after `execute_step` returns. So the
main thread can observe every result, join the submit thread, and read
the counters while the last worker is still between the send and the
increment, leaving the sum one short.

That is what CI hit: `599` vs `600`, per-worker counts
[231, 129, 57, 17, 134, 31].

Confirmed by inserting a 200us sleep between `execute_step` returning
Success and `record_step`, which makes the test fail every run
(298 vs 300); with this change it passes with that sleep still in place.

The counters are stats-only and are read after `shutdown()` joins the
workers in production, so this is a test synchronization defect, not a
pool bug — the fix is to wait for the counters to settle rather than to
reorder the worker.

Full fgumi-sort suite: 559 passed.
@nh13
nh13 temporarily deployed to github-actions July 22, 2026 22:05 — with GitHub Actions Inactive
@coderabbitai

coderabbitai Bot commented Jul 22, 2026 •

Copy link
Copy Markdown

Review Change Stack

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Pro

Run ID: e7925f75-9179-4d57-b189-21220415089e

📥 Commits

Reviewing files that changed from the base of the PR and between bb6c8db and 11d2ed4.

📒 Files selected for processing (1)
  • crates/fgumi-sort/src/worker_pool.rs

Walkthrough

The worker pool test now waits for cumulative CompressSpill counters to reach expected totals, with a 30-second timeout and diagnostic failure output. The second batch passes a cumulative expectation of 600 jobs.

Changes

Worker counter synchronization

Layer / File(s) Summary
Cumulative counter wait
crates/fgumi-sort/src/worker_pool.rs
The test helper accepts cumulative job counts, polls per-worker CompressSpill counters until convergence or timeout, and passes 600 for the second batch.

Estimated code review effort: 2 (Simple) | ~10 minutes

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly describes the main change: fixing a flaky sort test named active_worker_limit_caps_then_reactivates.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch nh/fix-flaky-active-worker-limit-test

Comment @coderabbitai help to get the list of available commands.

@codecov

codecov Bot commented Jul 22, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 80.00000% with 4 lines in your changes missing coverage. Please review.
✅ Project coverage is 93.49%. Comparing base (bb6c8db) to head (11d2ed4).

Files with missing lines Patch % Lines
crates/fgumi-sort/src/worker_pool.rs 80.00% 4 Missing ⚠️

❌ Your patch check has failed because the patch coverage (80.00%) is below the target coverage (90.00%). You can increase the patch coverage or adjust the target coverage.

Additional details and impacted files
@@            Coverage Diff             @@
##             main     #646      +/-   ##
==========================================
- Coverage   93.53%   93.49%   -0.05%     
==========================================
  Files         175      175              
  Lines      105806   105820      +14     
==========================================
- Hits        98969    98939      -30     
- Misses       6837     6881      +44     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@nh13
nh13 merged commit 02ad7c1 into main Jul 22, 2026
13 of 14 checks passed
@nh13
nh13 deleted the nh/fix-flaky-active-worker-limit-test branch July 22, 2026 22:34
@nh13 nh13 mentioned this pull request Jul 22, 2026

This branch was previously deployed

1 inactive deployment
github-actions — 11d2ed4e Deployed Jul 22, 2026 by nh13 via coverage #2914
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant