Repository navigation
De-flake RemoteNodeRestartDeathWatchSpec: out-wait the association gate - #8404
Merged
Merged
Conversation
The reconnect loop retried inside a Within(5s) window that opens at the same unreachability event that gates the restarted address for retry-gate-closed-for = 5s, leaving ~0ms of usable budget — every retry dead-lettered inside the gate. Widen the window to 30s with a fresh probe and bounded expect per attempt, drop the gate to 1s in the spec config, and convert the spec to async TestKit methods.
Aaronontheweb
added a commit
that referenced
this pull request
Jul 22, 2026
…overySpec (bug squash) (#8415) * De-flake DistributedPubSubRestartSpec: closed-loop restart-kill (mirrors #8404) Test-design fix only - no product/remoting code changes. First's restart-kill of third's same-address new incarnation was an OPEN-LOOP, fire-and-forget resend: it blindly Told "shutdown" ~10x over ~27s with no per-send verification. After the same-address restart, first's outbound endpoint to third is gated by the dead old incarnation's failing cluster heartbeats, so brand-new sends collateral-drop to dead-letters for tens of seconds until a fresh association forms. Open-loop delivery over an at-most-once, gated path has no constant time bound, so all blind shots can land inside the drop window, third never shuts down, and third's WhenTerminated wait times out (flake). Five prior timeout/window tunings could not fix this because the drop window is unbounded. Convert to a CLOSED-LOOP, self-verifying retry (same shape PR #8404 used for the sibling RemoteNodeRestartDeathWatchSpec): each AwaitAssertAsync attempt uses a FRESH probe, Identifies third's new /user/shutdown, and if the Subject resolves non-null Tells it "shutdown", requiring a "shutdown-ack" reply as the observable success signal. Every retry re-pokes the association, so recovery is no longer hostage to blind send timing. The Shutdown test actor now replies "shutdown-ack" before terminating (test-support code in the spec's config, not product code). Third's newSystem.WhenTerminated wait is made BEST-EFFORT: on the (now pathological) timeout it logs loudly and terminates newSystem cleanly so CI can never hang; the test still passes on the strength of the real assertions. The spec's actual subject-under-test - the SubscribeAck / ExpectNoMsg / DeltaCount == 0 gossip-isolation checks - is unchanged. Confirmed NOT a product defect and NOT the #8413 Artery issue. CI is the validator for the flake fix (cannot be reproduced locally without the MNTR harness). * DistributedPubSubRestartSpec: remove connection-timeout override (match JVM defaults) The prior de-flake set akka.remote.dot-netty.tcp.connection-timeout = 5s to "fail fast". But on the dot-netty transport AkkaProtocolSettings.HandshakeTimeout is sourced FROM connection-timeout, so that override silently crushed the Akka protocol-handshake budget to 5s. First's associate to third's restarted same-address/new-UID incarnation then timed out at exactly 5000ms on every attempt and never completed on loaded CI agents (builds 129198 -> node3 line 228 timeout; 129206 -> node1 closed-loop identify timeout after 16 attempts). The JVM spec (akka/akka DistributedPubSubRestartSpec) sets NO connection-timeout or retry-gate overrides and does not flake. Remove both overrides so the handshake gets its full 15s default budget, matching the JVM. Keeps the closed-loop identify/kill from the prior commit. * DistributedPubSubRestartSpec: bound transport-failure-detector (the actual flake fix) Root cause (confirmed from the failing node's log + remoting code): under the test transport (trttl.gremlin), the ThrottledAssociation FSM intentionally swallows the TCP-close event and relies on the AkkaProtocol transport failure detector to notice a dead connection. That FD's default acceptable-heartbeat-pause is 120s. When third is shut down and restarts on the same address, first's EndpointWriter stays alive on a half-open handle for up to ~120s - every send (Identify, heartbeat, gossip, the "shutdown") goes into the dead socket and is lost silently, so first never re-associates to the restarted incarnation and its identify loop times out. The earlier connection-timeout theory was wrong: the 15000ms associate timeout in the logs was unrelated teardown noise firing after the test had already failed. Every other restart/gate MNTR spec sets this same override (RemoteNodeShutdownAndComesBack 1s/3s, RemoteNodeRestartDeathWatch 1s/3s, RemoteRestartedQuarantined 1s/10s, RemoteGatePiercing 5s); this was the only restart spec missing it. Bounds the zombie window to <=6s. Not a production path - without the throttler adapter, Disassociated reaches ProtocolStateActor directly and tears the writer down immediately. * De-flake ClusterClientDiscoverySpec: raise discovery probe-timeout, dilate retry windows The MNTR client node timed out waiting for a ContactPoints message (CI build 129289, Windows MNTR lane): "Timeout 00:00:01 while waiting for a message of type Akka.Cluster.Tools.Client.ContactPoints". Root cause: the Client node config lowered the ClusterClient discovery `probe-timeout` from the reference.conf default of 3s down to 1s. Each initial-contact probe is an Akka.Management HTTP round-trip to a freshly started receptionist endpoint; on a loaded Windows CI agent that round-trip regularly takes just over 1s. The client log shows three probes to http://localhost:30001/cluster-client/receptionist, each cancelled at ~1.0s with TaskCanceledException/SocketException(995), so the client never resolved a receptionist, ContactPoints stayed empty, and the AwaitAssertAsync(10s) window expired after 10 attempts. The cluster and the management endpoint were already up before the client probed (the cluster-started barrier guarantees it), so this is an under-sized timeout, not an ordering race. Fix (test config only, no product change): - probe-timeout 1s -> 3s, restoring the shipped default so the HTTP probe has real headroom on a slow agent. This value is the ClusterClient's own HttpClient timeout and is not scaled by akka.test.timefactor, so it must be raised directly rather than dilated. - akka.test.timefactor = 2. AwaitAssertAsync already scales its window via RemainingOrDilated, so this dilates the ContactPoints retry loops on slow agents without changing any assertion or hard-coding a bigger constant. All real assertions are unchanged. CI is the validator for the timing flake.
Aaronontheweb
added a commit
that referenced
this pull request
Sep 12, 2026
…te (#8404) The reconnect loop retried inside a Within(5s) window that opens at the same unreachability event that gates the restarted address for retry-gate-closed-for = 5s, leaving ~0ms of usable budget — every retry dead-lettered inside the gate. Widen the window to 30s with a fresh probe and bounded expect per attempt, drop the gate to 1s in the spec config, and convert the spec to async TestKit methods. (cherry picked from commit 2679bef)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The spec's reconnect loop retries inside a Within(5s) window that opens at the same moment the endpoint manager gates the restarted node's address for retry-gate-closed-for = 5s — both are triggered by the same unreachability event, so on a slow box every retry dead-letters inside the gate and the window closes with ~0ms of usable budget (build 129150, Windows MNTR).
This widens the retry window to 30s with a fresh probe and bounded 3s expect per attempt, drops retry-gate-closed-for to 1s in the spec config to restore the sub-second-retry assumption the loop was originally sized for, and converts the spec to async TestKit methods.