Skip to content

DistributedData multi-node: de-flake ReplicatorSpec (per-attempt probes for Get reads) - #8749

Merged
Aaronontheweb merged 2 commits into
akkadotnet:devfrom
Aaronontheweb:fix/deflake-ddata-replicatorspec
Oct 4, 2026
Merged

Aaronontheweb merged 2 commits into
akkadotnet:devfrom
Aaronontheweb:fix/deflake-ddata-replicatorspec

Conversation

@Aaronontheweb

@Aaronontheweb Aaronontheweb commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

ReplicatorSpec reads CRDTs with a 50 ms ExpectMsg inside AwaitAssert on the shared TestActor, so one reply later than 50 ms desyncs the queue (CI: GetSuccess(D1:GCounter(1)) did not match, or a stray GetSuccess(KeyC) that breaks the next step). Each attempt now uses a fresh probe, the same fix #8509 made for ReplicatorChaosSpec.

Changes

  • ReplicatorSpecTests and every Cluster_CRDT_should_* step are now async Task, awaited in order. They use RunOnAsync, ExpectMsgAsync, ExpectNoMsgAsync, AwaitAssertAsync, WithinAsync and EnterBarrierAsync (new JoinAsync and EnterBarrierAfterTestStepAsync helpers), so no step blocks a pool thread with sync-over-async. Assertions and values are unchanged.
  • The Get-read loops (KeyC, D1..D30, KeyA, replica counts) use AwaitAssertAsync with a fresh TestProbe per attempt, a 1 s AttemptTimeout and a 200 ms interval. The converge step sends all 30 Gets, then reads the 30 replies in order. Overall budgets (5 s, 10 s) are unchanged.
  • TestConductor.Blackhole/PassThrough are awaited instead of .Wait(...), so a failed partition change fails the test.
  • The three KeyB loops in the 2-node step get a 10 s budget (was the 3 s default, equal to the inner expect), with a fresh probe per attempt.

The Cluster_CRDT_should_support_prefer_oldest_members step is converted too, but ReplicatorSpecTests does not call it.

Testing

Builds clean with -warnaserror. On head 1aa39e9, all four multi-node jobs (Linux, Linux Artery, Windows, Windows Artery) and the unit-test jobs pass in CI.

Breaking changes

None (tests only).

…es for Get reads)

A 50 ms ExpectMsg inside AwaitAssert on the shared TestActor desyncs its queue after one late reply. Use a fresh TestProbe per attempt with a 1 s bound, and await the TestConductor partition changes.
@Aaronontheweb Aaronontheweb added this to the 1.6.0 milestone Oct 4, 2026
@Aaronontheweb Aaronontheweb removed this from the 1.6.0 milestone Oct 4, 2026
Every step is now async Task using RunOnAsync, ExpectMsgAsync, AwaitAssertAsync, WithinAsync and EnterBarrierAsync, so nothing blocks a thread-pool thread on sync-over-async. Get-read loops use a fresh probe per attempt, and the KeyB loops get a 10 s budget.
@Aaronontheweb
Aaronontheweb merged commit d70c03e into akkadotnet:dev Oct 4, 2026
16 checks passed
@Aaronontheweb
Aaronontheweb deleted the fix/deflake-ddata-replicatorspec branch October 4, 2026 17:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant