Skip to content

De-flake RemoteNodeRestartDeathWatchSpec: out-wait the association gate - #8404

Merged
Aaronontheweb merged 1 commit into
devfrom
fix/remote-node-restart-deathwatch-deflake
Jul 18, 2026
Merged

Aaronontheweb merged 1 commit into
devfrom
fix/remote-node-restart-deathwatch-deflake

Conversation

@Aaronontheweb

Copy link
Copy Markdown
Member

The spec's reconnect loop retries inside a Within(5s) window that opens at the same moment the endpoint manager gates the restarted node's address for retry-gate-closed-for = 5s — both are triggered by the same unreachability event, so on a slow box every retry dead-letters inside the gate and the window closes with ~0ms of usable budget (build 129150, Windows MNTR).

This widens the retry window to 30s with a fresh probe and bounded 3s expect per attempt, drops retry-gate-closed-for to 1s in the spec config to restore the sub-second-retry assumption the loop was originally sized for, and converts the spec to async TestKit methods.

The reconnect loop retried inside a Within(5s) window that opens at the same unreachability event that gates the restarted address for retry-gate-closed-for = 5s, leaving ~0ms of usable budget — every retry dead-lettered inside the gate. Widen the window to 30s with a fresh probe and bounded expect per attempt, drop the gate to 1s in the spec config, and convert the spec to async TestKit methods.
@Aaronontheweb
Aaronontheweb merged commit 2679bef into dev Jul 18, 2026
11 checks passed
@Aaronontheweb
Aaronontheweb deleted the fix/remote-node-restart-deathwatch-deflake branch July 18, 2026 01:49
Aaronontheweb added a commit that referenced this pull request Jul 22, 2026
…overySpec (bug squash) (#8415)

* De-flake DistributedPubSubRestartSpec: closed-loop restart-kill (mirrors #8404)

Test-design fix only - no product/remoting code changes. First's restart-kill of
third's same-address new incarnation was an OPEN-LOOP, fire-and-forget resend: it
blindly Told "shutdown" ~10x over ~27s with no per-send verification. After the
same-address restart, first's outbound endpoint to third is gated by the dead old
incarnation's failing cluster heartbeats, so brand-new sends collateral-drop to
dead-letters for tens of seconds until a fresh association forms. Open-loop delivery
over an at-most-once, gated path has no constant time bound, so all blind shots can
land inside the drop window, third never shuts down, and third's WhenTerminated wait
times out (flake). Five prior timeout/window tunings could not fix this because the
drop window is unbounded.

Convert to a CLOSED-LOOP, self-verifying retry (same shape PR #8404 used for the
sibling RemoteNodeRestartDeathWatchSpec): each AwaitAssertAsync attempt uses a FRESH
probe, Identifies third's new /user/shutdown, and if the Subject resolves non-null
Tells it "shutdown", requiring a "shutdown-ack" reply as the observable success
signal. Every retry re-pokes the association, so recovery is no longer hostage to
blind send timing. The Shutdown test actor now replies "shutdown-ack" before
terminating (test-support code in the spec's config, not product code).

Third's newSystem.WhenTerminated wait is made BEST-EFFORT: on the (now pathological)
timeout it logs loudly and terminates newSystem cleanly so CI can never hang; the
test still passes on the strength of the real assertions. The spec's actual
subject-under-test - the SubscribeAck / ExpectNoMsg / DeltaCount == 0 gossip-isolation
checks - is unchanged.

Confirmed NOT a product defect and NOT the #8413 Artery issue. CI is the validator
for the flake fix (cannot be reproduced locally without the MNTR harness).

* DistributedPubSubRestartSpec: remove connection-timeout override (match JVM defaults)

The prior de-flake set akka.remote.dot-netty.tcp.connection-timeout = 5s to
"fail fast". But on the dot-netty transport AkkaProtocolSettings.HandshakeTimeout
is sourced FROM connection-timeout, so that override silently crushed the Akka
protocol-handshake budget to 5s. First's associate to third's restarted
same-address/new-UID incarnation then timed out at exactly 5000ms on every
attempt and never completed on loaded CI agents (builds 129198 -> node3 line
228 timeout; 129206 -> node1 closed-loop identify timeout after 16 attempts).

The JVM spec (akka/akka DistributedPubSubRestartSpec) sets NO connection-timeout
or retry-gate overrides and does not flake. Remove both overrides so the
handshake gets its full 15s default budget, matching the JVM. Keeps the
closed-loop identify/kill from the prior commit.

* DistributedPubSubRestartSpec: bound transport-failure-detector (the actual flake fix)

Root cause (confirmed from the failing node's log + remoting code): under the
test transport (trttl.gremlin), the ThrottledAssociation FSM intentionally
swallows the TCP-close event and relies on the AkkaProtocol transport failure
detector to notice a dead connection. That FD's default acceptable-heartbeat-pause
is 120s. When third is shut down and restarts on the same address, first's
EndpointWriter stays alive on a half-open handle for up to ~120s - every send
(Identify, heartbeat, gossip, the "shutdown") goes into the dead socket and is
lost silently, so first never re-associates to the restarted incarnation and its
identify loop times out. The earlier connection-timeout theory was wrong: the
15000ms associate timeout in the logs was unrelated teardown noise firing after
the test had already failed.

Every other restart/gate MNTR spec sets this same override (RemoteNodeShutdownAndComesBack
1s/3s, RemoteNodeRestartDeathWatch 1s/3s, RemoteRestartedQuarantined 1s/10s,
RemoteGatePiercing 5s); this was the only restart spec missing it. Bounds the zombie
window to <=6s. Not a production path - without the throttler adapter, Disassociated
reaches ProtocolStateActor directly and tears the writer down immediately.

* De-flake ClusterClientDiscoverySpec: raise discovery probe-timeout, dilate retry windows

The MNTR client node timed out waiting for a ContactPoints message (CI build
129289, Windows MNTR lane): "Timeout 00:00:01 while waiting for a message of
type Akka.Cluster.Tools.Client.ContactPoints".

Root cause: the Client node config lowered the ClusterClient discovery
`probe-timeout` from the reference.conf default of 3s down to 1s. Each
initial-contact probe is an Akka.Management HTTP round-trip to a freshly
started receptionist endpoint; on a loaded Windows CI agent that round-trip
regularly takes just over 1s. The client log shows three probes to
http://localhost:30001/cluster-client/receptionist, each cancelled at ~1.0s
with TaskCanceledException/SocketException(995), so the client never resolved
a receptionist, ContactPoints stayed empty, and the AwaitAssertAsync(10s)
window expired after 10 attempts. The cluster and the management endpoint were
already up before the client probed (the cluster-started barrier guarantees
it), so this is an under-sized timeout, not an ordering race.

Fix (test config only, no product change):
- probe-timeout 1s -> 3s, restoring the shipped default so the HTTP probe has
  real headroom on a slow agent. This value is the ClusterClient's own
  HttpClient timeout and is not scaled by akka.test.timefactor, so it must be
  raised directly rather than dilated.
- akka.test.timefactor = 2. AwaitAssertAsync already scales its window via
  RemainingOrDilated, so this dilates the ContactPoints retry loops on slow
  agents without changing any assertion or hard-coding a bigger constant.

All real assertions are unchanged. CI is the validator for the timing flake.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
…te (#8404)

The reconnect loop retried inside a Within(5s) window that opens at the same unreachability event that gates the restarted address for retry-gate-closed-for = 5s, leaving ~0ms of usable budget — every retry dead-lettered inside the gate. Widen the window to 30s with a fresh probe and bounded expect per attempt, drop the gate to 1s in the spec config, and convert the spec to async TestKit methods.

(cherry picked from commit 2679bef)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant