Skip to content

De-flake MNTR suite: DistributedPubSubRestartSpec + ClusterClientDiscoverySpec (bug squash) - #8415

Merged
Aaronontheweb merged 5 commits into
akkadotnet:devfrom
Aaronontheweb:fix/distributed-pubsub-restart-closed-loop-shutdown
Jul 22, 2026
Merged

Aaronontheweb merged 5 commits into
akkadotnet:devfrom
Aaronontheweb:fix/distributed-pubsub-restart-closed-loop-shutdown

Conversation

@Aaronontheweb

@Aaronontheweb Aaronontheweb commented Jul 18, 2026 •

Copy link
Copy Markdown
Member

Batched de-flake of racy multi-node (MNTR) specs. Each fix is test-only (no product code) and root-caused from CI failure logs; CI is the validator for these Windows-timing flakes.

1. DistributedPubSubRestartSpec — zombie EndpointWriter under the test transport

When Third is shut down and restarts on the same address, First's sending endpoint never notices the old TCP connection died: under the trttl.gremlin test transport the throttler intentionally swallows the disconnect event and relies on the AkkaProtocol transport failure detector to notice a dead peer — and that detector's default acceptable-heartbeat-pause is 120s. So First writes every message (Identify, heartbeat, gossip, the "shutdown") into a dead socket, silently, and never re-associates to the restarted node; its Identify loop times out.

Every other restart/gate MNTR spec sets a tight transport-failure-detector override for exactly this reason; this was the only restart spec missing it. Fix: add the same override (1s heartbeat, 5s pause), bounding the zombie window to ~6s. Also: closed-loop restart-kill mirroring #8404, and removal of a prior connection-timeout = 5s override (which via the dot-netty handshake/connection-timeout conflation starved the re-association handshake). Not a production path — without the throttler adapter, the disconnect reaches ProtocolStateActor directly and tears the writer down at once.

2. ClusterClientDiscoverySpec — under-sized discovery probe-timeout

The Client NodeConfig set akka.cluster.client.discovery.probe-timeout = 1s, below the shipped 3s default. That value is the discovery HttpClient timeout; on a loaded agent the HTTP round-trip to the freshly-started receptionist runs just over 1s, so every probe was cancelled at ~1s, ContactPoints never resolved, and the wait expired. Fix: restore probe-timeout = 3s (not time-factor-scalable, so raised directly) and add akka.test.timefactor = 2 to dilate the retry windows on slow agents. No product bug — the retry/backoff logic is sound.


Validated by re-running the full MNTR lane repeatedly; further racy specs surfaced by those runs will be stacked here.

…ors akkadotnet#8404)

Test-design fix only - no product/remoting code changes. First's restart-kill of
third's same-address new incarnation was an OPEN-LOOP, fire-and-forget resend: it
blindly Told "shutdown" ~10x over ~27s with no per-send verification. After the
same-address restart, first's outbound endpoint to third is gated by the dead old
incarnation's failing cluster heartbeats, so brand-new sends collateral-drop to
dead-letters for tens of seconds until a fresh association forms. Open-loop delivery
over an at-most-once, gated path has no constant time bound, so all blind shots can
land inside the drop window, third never shuts down, and third's WhenTerminated wait
times out (flake). Five prior timeout/window tunings could not fix this because the
drop window is unbounded.

Convert to a CLOSED-LOOP, self-verifying retry (same shape PR akkadotnet#8404 used for the
sibling RemoteNodeRestartDeathWatchSpec): each AwaitAssertAsync attempt uses a FRESH
probe, Identifies third's new /user/shutdown, and if the Subject resolves non-null
Tells it "shutdown", requiring a "shutdown-ack" reply as the observable success
signal. Every retry re-pokes the association, so recovery is no longer hostage to
blind send timing. The Shutdown test actor now replies "shutdown-ack" before
terminating (test-support code in the spec's config, not product code).

Third's newSystem.WhenTerminated wait is made BEST-EFFORT: on the (now pathological)
timeout it logs loudly and terminates newSystem cleanly so CI can never hang; the
test still passes on the strength of the real assertions. The spec's actual
subject-under-test - the SubscribeAck / ExpectNoMsg / DeltaCount == 0 gossip-isolation
checks - is unchanged.

Confirmed NOT a product defect and NOT the akkadotnet#8413 Artery issue. CI is the validator
for the flake fix (cannot be reproduced locally without the MNTR harness).
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 18, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 18, 2026
…ch JVM defaults)

The prior de-flake set akka.remote.dot-netty.tcp.connection-timeout = 5s to
"fail fast". But on the dot-netty transport AkkaProtocolSettings.HandshakeTimeout
is sourced FROM connection-timeout, so that override silently crushed the Akka
protocol-handshake budget to 5s. First's associate to third's restarted
same-address/new-UID incarnation then timed out at exactly 5000ms on every
attempt and never completed on loaded CI agents (builds 129198 -> node3 line
228 timeout; 129206 -> node1 closed-loop identify timeout after 16 attempts).

The JVM spec (akka/akka DistributedPubSubRestartSpec) sets NO connection-timeout
or retry-gate overrides and does not flake. Remove both overrides so the
handshake gets its full 15s default budget, matching the JVM. Keeps the
closed-loop identify/kill from the prior commit.
@Aaronontheweb
Aaronontheweb force-pushed the fix/distributed-pubsub-restart-closed-loop-shutdown branch from 12604b3 to ad643ef Compare July 19, 2026 01:12
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 19, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 19, 2026
…ctual flake fix)

Root cause (confirmed from the failing node's log + remoting code): under the
test transport (trttl.gremlin), the ThrottledAssociation FSM intentionally
swallows the TCP-close event and relies on the AkkaProtocol transport failure
detector to notice a dead connection. That FD's default acceptable-heartbeat-pause
is 120s. When third is shut down and restarts on the same address, first's
EndpointWriter stays alive on a half-open handle for up to ~120s - every send
(Identify, heartbeat, gossip, the "shutdown") goes into the dead socket and is
lost silently, so first never re-associates to the restarted incarnation and its
identify loop times out. The earlier connection-timeout theory was wrong: the
15000ms associate timeout in the logs was unrelated teardown noise firing after
the test had already failed.

Every other restart/gate MNTR spec sets this same override (RemoteNodeShutdownAndComesBack
1s/3s, RemoteNodeRestartDeathWatch 1s/3s, RemoteRestartedQuarantined 1s/10s,
RemoteGatePiercing 5s); this was the only restart spec missing it. Bounds the zombie
window to <=6s. Not a production path - without the throttler adapter, Disassociated
reaches ProtocolStateActor directly and tears the writer down immediately.
@Aaronontheweb
Aaronontheweb force-pushed the fix/distributed-pubsub-restart-closed-loop-shutdown branch from e818790 to c976c87 Compare July 22, 2026 01:34
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
@Aaronontheweb
Aaronontheweb force-pushed the fix/distributed-pubsub-restart-closed-loop-shutdown branch from de0b1e9 to c976c87 Compare July 22, 2026 03:08
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
…ilate retry windows

The MNTR client node timed out waiting for a ContactPoints message (CI build
129289, Windows MNTR lane): "Timeout 00:00:01 while waiting for a message of
type Akka.Cluster.Tools.Client.ContactPoints".

Root cause: the Client node config lowered the ClusterClient discovery
`probe-timeout` from the reference.conf default of 3s down to 1s. Each
initial-contact probe is an Akka.Management HTTP round-trip to a freshly
started receptionist endpoint; on a loaded Windows CI agent that round-trip
regularly takes just over 1s. The client log shows three probes to
http://localhost:30001/cluster-client/receptionist, each cancelled at ~1.0s
with TaskCanceledException/SocketException(995), so the client never resolved
a receptionist, ContactPoints stayed empty, and the AwaitAssertAsync(10s)
window expired after 10 attempts. The cluster and the management endpoint were
already up before the client probed (the cluster-started barrier guarantees
it), so this is an under-sized timeout, not an ordering race.

Fix (test config only, no product change):
- probe-timeout 1s -> 3s, restoring the shipped default so the HTTP probe has
  real headroom on a slow agent. This value is the ClusterClient's own
  HttpClient timeout and is not scaled by akka.test.timefactor, so it must be
  raised directly rather than dilated.
- akka.test.timefactor = 2. AwaitAssertAsync already scales its window via
  RemainingOrDilated, so this dilates the ContactPoints retry loops on slow
  agents without changing any assertion or hard-coding a bigger constant.

All real assertions are unchanged. CI is the validator for the timing flake.
@Aaronontheweb
Aaronontheweb force-pushed the fix/distributed-pubsub-restart-closed-loop-shutdown branch from ae019c1 to a883a42 Compare July 22, 2026 03:39
@Aaronontheweb Aaronontheweb changed the title De-flake DistributedPubSubRestartSpec: open-loop -> closed-loop restart-kill (mirrors #8404) De-flake MNTR suite: DistributedPubSubRestartSpec + ClusterClientDiscoverySpec (bug squash) Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
Aaronontheweb added a commit to Aaronontheweb/akka.net that referenced this pull request Jul 22, 2026
@Aaronontheweb
Aaronontheweb force-pushed the fix/distributed-pubsub-restart-closed-loop-shutdown branch from d32b2e1 to a883a42 Compare July 22, 2026 06:13
@Aaronontheweb
Aaronontheweb marked this pull request as draft July 22, 2026 11:33
@Aaronontheweb
Aaronontheweb marked this pull request as ready for review July 22, 2026 16:28
@Aaronontheweb
Aaronontheweb enabled auto-merge (squash) July 22, 2026 16:30
@Aaronontheweb
Aaronontheweb merged commit b603b04 into akkadotnet:dev Jul 22, 2026
11 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant