Skip to content

De-flake ClusterSingletonManagerLeave2Spec: watch the proxy from a dedicated probe, as upstream does - #8569

Merged
Aaronontheweb merged 1 commit into
fix/zero-count-event-filter-inside-withinfrom
fix/singleton-leave2-spec-probe-watch
Sep 10, 2026
Merged

Aaronontheweb merged 1 commit into
fix/zero-count-event-filter-inside-withinfrom
fix/singleton-leave2-spec-probe-watch

Conversation

@Aaronontheweb

@Aaronontheweb Aaronontheweb commented Sep 10, 2026 •

Copy link
Copy Markdown
Member

Stacked on #8566.

What changes

ClusterSingletonManagerLeave2Spec watches its singleton proxy from a dedicated probe instead of the test actor, the way Akka JVM, Pekko, and the sibling ClusterSingletonManagerLeaveSpec in this repo already do. The proxy is created on first use through a lazy async factory that registers the watch once. The file moves to the async TestKit API: every RunOn, Within, AwaitAssert, ExpectMsg, EnterBarrier, and ExpectTerminated becomes its async form, a ForEach that would have hidden an async void becomes a for loop, and the one Thread.Sleep before a negative assertion becomes a delay with a comment saying it is a deliberate soak. No assertion or timeout changes.

Why

Build 131334, Linux Artery, on a PR that touched a different spec: node first expected the "MemberRemoved" string and received the proxy's Terminated instead, at 02:36:33.32.

One event drives both messages. When first finishes leaving, the cluster publishes MemberRemoved for it. The spec's member-status listener turns that into the "MemberRemoved" string for the test actor. The proxy stops itself on the same event, and because the test actor watched the proxy, its Terminated lands in the same queue. Two independent chains race into one mailbox, the event stream fans out in hash order, and nothing orders them. The spec asserted the string first, so the notice arriving first failed it. This is a general TestKit hazard: a watch from the test actor makes every later ExpectMsg an implicit ordering assertion over Terminated as well.

The leave hand-over itself is correct; the singleton's PostStop runs before removal, as the spec's comment claims. A separate product nit found on the way is #8568: the proxy's identification-failure timer starts with a zero initial delay, so every proxy start logs a false failure warning.

How it was checked

dotnet build src/contrib/cluster/Akka.Cluster.Tools.Tests.MultiNode -c Release -warnaserror clean. The spec run three times on each transport: 5 of 5 nodes passed every time, 10 to 11 s. The grep for synchronous TestKit calls, sleeps, blocking waits, and ForEach on the file returns nothing. Test-only change. No ledger entry.

…dicated probe, as upstream does; migrate the file to the async TestKit API

Cluster.Leave triggers one MemberRemoved(self) EventStream publication that
drives two independent, unordered subscribers: the RegisterOnMemberRemoved
callback that tells the test actor "MemberRemoved", and the proxy's own
self-stop on seeing MemberRemoved for its node, whose death-watch Terminated
lands in whatever queue is watching it. When the test actor did the watching,
both messages competed for the same FIFO queue, so ExpectMsg("MemberRemoved")
could dequeue the Terminated meant for the following ExpectTerminated instead
(observed on the Linux Artery lane, node "first"). Watching echoProxy from a
dedicated TestProbe, as upstream Akka JVM/Pekko and the sibling
ClusterSingletonManagerLeaveSpec already do, gives Terminated its own queue
and removes the race without touching any timeout, barrier, or assertion.

The lazy proxy accessor becomes Lazy<Task<IActorRef>> so the watch can be
registered asynchronously while keeping create-on-first-use, single-execution
semantics. The rest of the file moves to the async TestKit API (RunOnAsync,
WithinAsync, AwaitAssertAsync, ExpectMsgAsync, EnterBarrierAsync) to support
that; the ForEach-based negative-assertion loop becomes a for loop, and its
Thread.Sleep(1000) becomes Task.Delay(1000), unchanged as a deliberate 1s
soak between probes rather than a stand-in for a real event.
@Aaronontheweb
Aaronontheweb force-pushed the fix/singleton-leave2-spec-probe-watch branch from 82b5acb to 2166870 Compare September 10, 2026 14:54
@Aaronontheweb
Aaronontheweb merged commit b2524f6 into dev Sep 10, 2026
4 checks passed
@Aaronontheweb
Aaronontheweb deleted the fix/singleton-leave2-spec-probe-watch branch September 10, 2026 15:27
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
…dicated probe, as upstream does; migrate the file to the async TestKit API (#8569)

Cluster.Leave triggers one MemberRemoved(self) EventStream publication that
drives two independent, unordered subscribers: the RegisterOnMemberRemoved
callback that tells the test actor "MemberRemoved", and the proxy's own
self-stop on seeing MemberRemoved for its node, whose death-watch Terminated
lands in whatever queue is watching it. When the test actor did the watching,
both messages competed for the same FIFO queue, so ExpectMsg("MemberRemoved")
could dequeue the Terminated meant for the following ExpectTerminated instead
(observed on the Linux Artery lane, node "first"). Watching echoProxy from a
dedicated TestProbe, as upstream Akka JVM/Pekko and the sibling
ClusterSingletonManagerLeaveSpec already do, gives Terminated its own queue
and removes the race without touching any timeout, barrier, or assertion.

The lazy proxy accessor becomes Lazy<Task<IActorRef>> so the watch can be
registered asynchronously while keeping create-on-first-use, single-execution
semantics. The rest of the file moves to the async TestKit API (RunOnAsync,
WithinAsync, AwaitAssertAsync, ExpectMsgAsync, EnterBarrierAsync) to support
that; the ForEach-based negative-assertion loop becomes a for loop, and its
Thread.Sleep(1000) becomes Task.Delay(1000), unchanged as a deliberate 1s
soak between probes rather than a stand-in for a real event.

(cherry picked from commit b2524f6)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant