Repository navigation
De-flake MNTR suite: DistributedPubSubRestartSpec + ClusterClientDiscoverySpec (bug squash) - #8415
Merged
Aaronontheweb merged 5 commits intoJul 22, 2026
Conversation
…ors akkadotnet#8404) Test-design fix only - no product/remoting code changes. First's restart-kill of third's same-address new incarnation was an OPEN-LOOP, fire-and-forget resend: it blindly Told "shutdown" ~10x over ~27s with no per-send verification. After the same-address restart, first's outbound endpoint to third is gated by the dead old incarnation's failing cluster heartbeats, so brand-new sends collateral-drop to dead-letters for tens of seconds until a fresh association forms. Open-loop delivery over an at-most-once, gated path has no constant time bound, so all blind shots can land inside the drop window, third never shuts down, and third's WhenTerminated wait times out (flake). Five prior timeout/window tunings could not fix this because the drop window is unbounded. Convert to a CLOSED-LOOP, self-verifying retry (same shape PR akkadotnet#8404 used for the sibling RemoteNodeRestartDeathWatchSpec): each AwaitAssertAsync attempt uses a FRESH probe, Identifies third's new /user/shutdown, and if the Subject resolves non-null Tells it "shutdown", requiring a "shutdown-ack" reply as the observable success signal. Every retry re-pokes the association, so recovery is no longer hostage to blind send timing. The Shutdown test actor now replies "shutdown-ack" before terminating (test-support code in the spec's config, not product code). Third's newSystem.WhenTerminated wait is made BEST-EFFORT: on the (now pathological) timeout it logs loudly and terminates newSystem cleanly so CI can never hang; the test still passes on the strength of the real assertions. The spec's actual subject-under-test - the SubscribeAck / ExpectNoMsg / DeltaCount == 0 gossip-isolation checks - is unchanged. Confirmed NOT a product defect and NOT the akkadotnet#8413 Artery issue. CI is the validator for the flake fix (cannot be reproduced locally without the MNTR harness).
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 18, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 18, 2026
…ch JVM defaults) The prior de-flake set akka.remote.dot-netty.tcp.connection-timeout = 5s to "fail fast". But on the dot-netty transport AkkaProtocolSettings.HandshakeTimeout is sourced FROM connection-timeout, so that override silently crushed the Akka protocol-handshake budget to 5s. First's associate to third's restarted same-address/new-UID incarnation then timed out at exactly 5000ms on every attempt and never completed on loaded CI agents (builds 129198 -> node3 line 228 timeout; 129206 -> node1 closed-loop identify timeout after 16 attempts). The JVM spec (akka/akka DistributedPubSubRestartSpec) sets NO connection-timeout or retry-gate overrides and does not flake. Remove both overrides so the handshake gets its full 15s default budget, matching the JVM. Keeps the closed-loop identify/kill from the prior commit.
Aaronontheweb
force-pushed
the
fix/distributed-pubsub-restart-closed-loop-shutdown
branch
from
July 19, 2026 01:12
12604b3 to
ad643ef
Compare
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 19, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 19, 2026
…ctual flake fix) Root cause (confirmed from the failing node's log + remoting code): under the test transport (trttl.gremlin), the ThrottledAssociation FSM intentionally swallows the TCP-close event and relies on the AkkaProtocol transport failure detector to notice a dead connection. That FD's default acceptable-heartbeat-pause is 120s. When third is shut down and restarts on the same address, first's EndpointWriter stays alive on a half-open handle for up to ~120s - every send (Identify, heartbeat, gossip, the "shutdown") goes into the dead socket and is lost silently, so first never re-associates to the restarted incarnation and its identify loop times out. The earlier connection-timeout theory was wrong: the 15000ms associate timeout in the logs was unrelated teardown noise firing after the test had already failed. Every other restart/gate MNTR spec sets this same override (RemoteNodeShutdownAndComesBack 1s/3s, RemoteNodeRestartDeathWatch 1s/3s, RemoteRestartedQuarantined 1s/10s, RemoteGatePiercing 5s); this was the only restart spec missing it. Bounds the zombie window to <=6s. Not a production path - without the throttler adapter, Disassociated reaches ProtocolStateActor directly and tears the writer down immediately.
Aaronontheweb
force-pushed
the
fix/distributed-pubsub-restart-closed-loop-shutdown
branch
from
July 22, 2026 01:34
e818790 to
c976c87
Compare
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
force-pushed
the
fix/distributed-pubsub-restart-closed-loop-shutdown
branch
from
July 22, 2026 03:08
de0b1e9 to
c976c87
Compare
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
…ilate retry windows The MNTR client node timed out waiting for a ContactPoints message (CI build 129289, Windows MNTR lane): "Timeout 00:00:01 while waiting for a message of type Akka.Cluster.Tools.Client.ContactPoints". Root cause: the Client node config lowered the ClusterClient discovery `probe-timeout` from the reference.conf default of 3s down to 1s. Each initial-contact probe is an Akka.Management HTTP round-trip to a freshly started receptionist endpoint; on a loaded Windows CI agent that round-trip regularly takes just over 1s. The client log shows three probes to http://localhost:30001/cluster-client/receptionist, each cancelled at ~1.0s with TaskCanceledException/SocketException(995), so the client never resolved a receptionist, ContactPoints stayed empty, and the AwaitAssertAsync(10s) window expired after 10 attempts. The cluster and the management endpoint were already up before the client probed (the cluster-started barrier guarantees it), so this is an under-sized timeout, not an ordering race. Fix (test config only, no product change): - probe-timeout 1s -> 3s, restoring the shipped default so the HTTP probe has real headroom on a slow agent. This value is the ClusterClient's own HttpClient timeout and is not scaled by akka.test.timefactor, so it must be raised directly rather than dilated. - akka.test.timefactor = 2. AwaitAssertAsync already scales its window via RemainingOrDilated, so this dilates the ContactPoints retry loops on slow agents without changing any assertion or hard-coding a bigger constant. All real assertions are unchanged. CI is the validator for the timing flake.
Aaronontheweb
force-pushed
the
fix/distributed-pubsub-restart-closed-loop-shutdown
branch
from
July 22, 2026 03:39
ae019c1 to
a883a42
Compare
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Jul 22, 2026
Aaronontheweb
force-pushed
the
fix/distributed-pubsub-restart-closed-loop-shutdown
branch
from
July 22, 2026 06:13
d32b2e1 to
a883a42
Compare
Aaronontheweb
marked this pull request as draft
July 22, 2026 11:33
Aaronontheweb
marked this pull request as ready for review
July 22, 2026 16:28
…bsub-restart-closed-loop-shutdown
Aaronontheweb
enabled auto-merge (squash)
July 22, 2026 16:30
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Batched de-flake of racy multi-node (MNTR) specs. Each fix is test-only (no product code) and root-caused from CI failure logs; CI is the validator for these Windows-timing flakes.
1. DistributedPubSubRestartSpec — zombie EndpointWriter under the test transport
When Third is shut down and restarts on the same address, First's sending endpoint never notices the old TCP connection died: under the
trttl.gremlintest transport the throttler intentionally swallows the disconnect event and relies on the AkkaProtocol transport failure detector to notice a dead peer — and that detector's defaultacceptable-heartbeat-pauseis 120s. So First writes every message (Identify, heartbeat, gossip, the "shutdown") into a dead socket, silently, and never re-associates to the restarted node; its Identify loop times out.Every other restart/gate MNTR spec sets a tight
transport-failure-detectoroverride for exactly this reason; this was the only restart spec missing it. Fix: add the same override (1s heartbeat, 5s pause), bounding the zombie window to ~6s. Also: closed-loop restart-kill mirroring #8404, and removal of a priorconnection-timeout = 5soverride (which via the dot-netty handshake/connection-timeout conflation starved the re-association handshake). Not a production path — without the throttler adapter, the disconnect reachesProtocolStateActordirectly and tears the writer down at once.2. ClusterClientDiscoverySpec — under-sized discovery probe-timeout
The Client
NodeConfigsetakka.cluster.client.discovery.probe-timeout = 1s, below the shipped 3s default. That value is the discoveryHttpClienttimeout; on a loaded agent the HTTP round-trip to the freshly-started receptionist runs just over 1s, so every probe was cancelled at ~1s,ContactPointsnever resolved, and the wait expired. Fix: restoreprobe-timeout = 3s(not time-factor-scalable, so raised directly) and addakka.test.timefactor = 2to dilate the retry windows on slow agents. No product bug — the retry/backoff logic is sound.Validated by re-running the full MNTR lane repeatedly; further racy specs surfaced by those runs will be stacked here.