Repository navigation
DistributedData multi-node: de-flake ReplicatorSpec (per-attempt probes for Get reads) - #8749
Merged
Aaronontheweb merged 2 commits intoOct 4, 2026
Conversation
…es for Get reads) A 50 ms ExpectMsg inside AwaitAssert on the shared TestActor desyncs its queue after one late reply. Use a fresh TestProbe per attempt with a 1 s bound, and await the TestConductor partition changes.
Every step is now async Task using RunOnAsync, ExpectMsgAsync, AwaitAssertAsync, WithinAsync and EnterBarrierAsync, so nothing blocks a thread-pool thread on sync-over-async. Get-read loops use a fresh probe per attempt, and the KeyB loops get a 10 s budget.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
ReplicatorSpecreads CRDTs with a 50 msExpectMsginsideAwaitAsserton the shared TestActor, so one reply later than 50 ms desyncs the queue (CI:GetSuccess(D1:GCounter(1)) did not match, or a strayGetSuccess(KeyC)that breaks the next step). Each attempt now uses a fresh probe, the same fix #8509 made forReplicatorChaosSpec.Changes
ReplicatorSpecTestsand everyCluster_CRDT_should_*step are nowasync Task, awaited in order. They useRunOnAsync,ExpectMsgAsync,ExpectNoMsgAsync,AwaitAssertAsync,WithinAsyncandEnterBarrierAsync(newJoinAsyncandEnterBarrierAfterTestStepAsynchelpers), so no step blocks a pool thread with sync-over-async. Assertions and values are unchanged.D1..D30, KeyA, replica counts) useAwaitAssertAsyncwith a freshTestProbeper attempt, a 1 sAttemptTimeoutand a 200 ms interval. The converge step sends all 30 Gets, then reads the 30 replies in order. Overall budgets (5 s, 10 s) are unchanged.TestConductor.Blackhole/PassThroughare awaited instead of.Wait(...), so a failed partition change fails the test.The
Cluster_CRDT_should_support_prefer_oldest_membersstep is converted too, butReplicatorSpecTestsdoes not call it.Testing
Builds clean with
-warnaserror. On head 1aa39e9, all four multi-node jobs (Linux, Linux Artery, Windows, Windows Artery) and the unit-test jobs pass in CI.Breaking changes
None (tests only).