Repository navigation
Backport September 2026 de-flakes, product fixes, and TestKit changes from dev to v1.5 - #8588
Merged
Merged
Conversation
…wn (#8499) * Add TestKitBase.ShutdownAsync and use it in StressSpec churn teardown Closes #8497. TestKitBase.Shutdown is a blocking `Terminate().Wait(duration)`. On a loaded agent that pins a thread pool thread for the whole wait, and the work the shutdown is waiting on - coordinated shutdown phases, cluster heartbeats - competes for the threads that are left. In StressSpec's churn phase the wait ran its full 10s and starved the hosting node's heartbeat sender for 12+ seconds, so the survivors marked it unreachable and SBR downed it. The wait made its own timeout more likely. ShutdownAsync races Terminate() against a timer instead of blocking, using the AwaitWithTimeout helper Akka.TestKit already ships. Everything else matches the sync overloads: same default duration, same forced guardian stop, same verifySystemShutdown throw-vs-log, same message. Both paths now share the force-stop tail, so the two cannot drift. The sync Shutdown keeps its own Terminate().Wait rather than blocking on the async one - sync-over-async here would need two continuations off the pool while a thread is pinned, which is worse under the starvation this fixes. StressSpec's post-churn teardown now awaits ShutdownAsync. Same wait budget, same fallback; the wait just no longer holds a thread. * Await the in-loop churn teardown in StressSpec too The post-loop teardown already awaits ShutdownAsync. This is the other blocking Shutdown in ExerciseJoinRemoveAsync - it tears down the previous round's system at the top of each churn round, so it is the one that blocks while the cluster is under load. Leaving it on Terminate().Wait would keep pinning a thread pool thread for the whole wait, which is what starves the hosting node's heartbeat sender. Same wait budget, same forced-guardian fallback, same ordering: the teardown still finishes before the next ActorSystem is created. * Drop the redundant happy-path ShutdownAsync test The verify-throws test already proves ShutdownAsync waits: a fire-and-forget implementation cannot observe that the system failed to stop, so it cannot throw the TimeoutException the test requires. StressSpec exercises the happy path on every MNTR run. The two remaining tests stay separate because each call runs the forced-guardian-stop tail, which kills the stuck system - so the throw branch and the log branch each need their own system. (cherry picked from commit 81289e8)
…hin (#8406) * De-flake ClusterShardingSpec recovery phase: keep barriers out of Within One WithinAsync(50s) spanned five barriers plus an unbounded ExpectMsgAsync, so a slow remember-entities recovery drained the shared budget and the final barrier's timeout clamped to zero. Bound the recovery check in a retrying probe-per-attempt AwaitAssert and move the barriers outside the Within so each gets the full barrier-timeout. * Remove the phase's outer Within entirely The remaining Within(50s) still wrapped three barriers and bounded steps whose worst cases sum past it. Every wait now carries its own explicit bound; all five barriers run outside any Within. (cherry picked from commit f13a8fa)
…out flakiness (#8286) NodeChurnSpec intermittently failed with "Failed to stop [NodeChurnSpec] within [00:00:05]" on slower CI agents. The synchronous Shutdown(node, verifySystemShutdown:true) helper blocks a thread-pool thread inside Task.Wait() for its 5s budget while the coordinated-shutdown pipeline itself needs the thread pool to make progress — a sync-over-async self-starvation. Migrate the spec to the async, task-returning TestKit APIs (WithinAsync, AwaitMembersUpAsync, EnterBarrierAsync, AwaitAssertAsync, ExpectNoMsgAsync) and replace the blocking per-system Shutdown loop with a concurrent `await Task.WhenAll(systems.Select(s => s.Terminate())).WaitAsync(30s)`, which frees the thread and preserves verify-shutdown semantics (throws TimeoutException if a system fails to stop). Same idiom as QuickRestartSpec / DistributedPubSubRestartSpec. Verified locally across 4 consecutive runs (3/3 node roles pass each, ~45s). (cherry picked from commit 4b69aa5)
…ting initialization; the counter was seeded at -1 (#8570) ConnectionSourceStageLogic's _connectionFlowsAwaitingInitialization counter used AtomicCounterLong's default constructor, which seeds an id generator at -1 so the first Next() returns 0. JVM Akka's TcpStages.scala uses `new AtomicLong()`, which starts at 0. Because the counter never rests at 0 in this port, the UnbindCompleted() fast path (`if (_connectionFlowsAwaitingInitialization.Current == 0) CompleteStage();`) was unreachable, so every graceful ServerBinding.Unbind() fell through to the BindShutdownTimer and waited the full akka.stream.materializer.subscription-timeout.timeout (5s by default) even when no connection was ever accepted. Seed the counter at 0 to match upstream semantics and restore the fast path. Adds a regression test that binds, never connects, and asserts Unbind() completes well inside the subscription timeout. (cherry picked from commit 91556a0)
…ng-state-timeout, as the config documents (#8574) * Cluster.Sharding: arm the remember-entities write timeout with updating-state-timeout, as the config documents; the read setting had been used since #6479 Shard.SendToRememberStore armed the remember-entities WRITE timeout with TuningParameters.WaitingForStateTimeout (2s), while the exception it throws on timeout, reference.conf's own doc comment for updating-state-timeout ("Also used as timeout for writes of remember entities when that is enabled"), and DDataRememberEntitiesShardStore's write-majority retry sizing (3 retries at updating-state-timeout / 4) all agree the write should use UpdatingStateTimeout (5s). Commit dde36e9 (#6479) swapped both the read and write sites in one hunk but only described fixing the read path, leaving the write path armed with the wrong, shorter timeout and giving the store's retries zero time to run before the shard gave up and restarted, dropping any buffered messages for the entity. Restores UpdatingStateTimeout on the write arm; the read path (already using UpdatingStateTimeout, more leniently than Pekko) is untouched. Adds RememberEntitiesWriteTimeoutSpec, which fails on unpatched code (the shard restarts and the buffered message is dropped) and passes with the fix (the store's delayed-but-valid write ack arrives and no restart occurs). * RememberEntitiesFailureSpec: prove the write timeout fix with the existing fake store instead of a new spec Deletes RememberEntitiesWriteTimeoutSpec.cs's standalone 180-line spec and its bespoke delayed-ack store, replacing it with one fact on RememberEntitiesFailureSpec that reuses the shared FakeStore/FakeShardStoreActor harness. The fact delays a remember-entities write's ack by 3s (strictly between the 2s waiting-for-state-timeout and 5s updating-state-timeout reference defaults) and asserts the buffered message is still delivered and the shard does not restart. FakeShardStoreActor's existing Delay failure mode replies to a delayed write with the wrong payload type, so the write never actually completes - every existing Delay fact only relies on that to force a timeout. Adds a new DelayedSuccess failure mode, additive and reusing the same Delayed/timer plumbing, that acknowledges the write correctly so a write can genuinely finish late but within its timeout - which the existing Delay mode could not exercise. Verified this fact discriminates: reverting Shard.cs's one-line fix by hand made it fail (shard restarts at ~2s per the log, buffered message lost); restoring the fix made it pass. All 26 facts in RememberEntitiesFailureSpec pass in three consecutive runs. (cherry picked from commit 2d8fa21)
…ting/Downed (#8582) * Fix Replicator.IsKnownNode to trust members first seen as Leaving/Exiting InitialStateAsEvents replays each current member at its CURRENT status, so a Replicator that starts after a member has already moved to Leaving or Exiting receives MemberLeft/MemberExited and never MemberUp. IsKnownNode inferred membership from MemberUp alone, so every Write and gossip message from that member was silently dropped for as long as it stayed in the cluster. In cluster sharding this can stall a departing shard coordinator's ddata write for its entire updating-state-timeout during a rolling restart. _exitingNodes covers MemberExited. MemberLeft/MemberDowned already land in _leader via ReceiveOtherMemberEvent. Both sets are pruned in ReceiveMemberRemoved, so a removed node stops being trusted immediately. This reuses state the actor already maintains rather than adding a new membership set. Replica set composition is unchanged: AllNodes remains _nodes + _weaklyUpNodes, so a dying node's Write must still win a majority of current replicas. Only who may ask is widened, not who may vote. * Add MemberDowned coverage to ReplicatorKnownNodeSpec MemberDowned shares the generic ReceiveOtherMemberEvent path with MemberLeft, so it is the third status a late subscriber can be handed on InitialStateAsEvents replay and was the one status the spec did not cover. Note the enum value is MemberStatus.Down while the cluster event is MemberDowned. (cherry picked from commit 2c2a406)
…egion when the node is leaving (#8586) ClusterSingletonManager sends the coordinator its termination message on every hand-over, and a hand-over is not always a shutdown: while a cluster forms, two members of the same role can both reach Oldest, and the loser hands over while staying Up. ShardCoordinator.HandleTerminate read every termination message as a node shutdown, sent GracefulShutdown to the node's own live region, and deferred its stop until the region was gone. ClusterShardingGuardian then evicted the dead region from ClusterSharding's cache, so ShardRegion(typeName) threw on that node for the rest of the process. The coordinator still stops on every hand-over, which is what lets the singleton manager complete it. It only takes the local region down with it when that region is already handing off, which is the coordinated shutdown path, or when the node itself is leaving: the system is terminating, CoordinatedShutdown is running, or the self member is Leaving, Exiting, or Down. The branch that sends GracefulShutdown to a live local region came from JVM Akka PR 30338, a rolling-update latency optimization; before it a termination message was a bare stop, which is what a benign hand-over now gets again. ShardCoordinatorHandOverSpec sends the singleton's termination message the way the manager does and asserts that the coordinator stops and the region does not. It fails on the old code with the region's Terminated. Fixes #8583. (cherry picked from commit c0600b7)
…the transport reads it (re-land of #8546 with the RemoteConfigSpec fix) (#8561) * TestKit: put the DotNetty batching override under akka.remote, where the transport reads it (#8546) The key sat under akka.test and never took effect. A config test now guards it. (cherry picked from commit d88eec8) * RemoteConfigSpec: read batching defaults from the Akka.Remote reference config, and pin the TestKit override separately Since #8546 moved the DotNetty batching override to where the transport reads it, the running test system's config now carries the TestKit's batching-disabled override, so it no longer reflects the Akka.Remote defaults. The BatchWriter fact now reads straight from Akka.Remote.Configuration.RemoteConfigFactory to pin the library default, and a new sibling fact reads from the running system to pin the TestKit override, so a future misplacement of either key fails a test. (cherry picked from commit adacc74) * TestKit_Config_Tests: drop the stale copy of the misplaced batching key from the custom test config (cherry picked from commit 3f966aa)
…ia stdout sentinel (#8515) * Remove the MNTR conductor port race: bind port 0 and propagate The multi-node runner picked the TestConductor's port by binding a temporary socket, reading the port, closing the socket, and handing the number to node 1. The number is in the ephemeral range by construction, so any outbound connection on the machine can take it between the probe and the conductor's real bind. When that happens node 1 dies with "Address already in use" and every other node waits out a 30 second attach timeout against a conductor that never existed. The runner now starts node 1 alone with multinode.server-port=0. The node binds a free port, prints it as a sentinel line, and the runner starts the remaining nodes with the port the conductor actually holds. There is no window to lose. - ConductorPortSentinel: the stdout contract between a conductor node and the runner, with a strict parser. - Controller reports its bind result through a TaskCompletionSource instead of an ask. A failed bind now throws ConductorBindException naming the port, straight away, rather than surfacing as a query timeout. - MultiNodeSpec accepts server-port=0 on the conductor node and rejects it on client nodes, and emits the sentinel once the conductor is bound. - When the conductor node exits early or never publishes a port, the runner kills it and fails the nodes it never started with a message naming the cause. A spec that used to hang for minutes now fails in under a second. Explicitly configured ports are unchanged: the node binds that exact port and emits the sentinel anyway, so running node processes by hand still works. Applied to both the xUnit v3 and xUnit v2 adapters. * Model the conductor bind as a Controller behavior switch The Controller constructor blocked on the DotNetty bind, so a dispatcher thread sat in the actor's constructor waiting on I/O and a failed bind escaped as an exception out of construction. The constructor now only assigns fields. PreStart starts the bind and PipeTo's the outcome back into the mailbox as Bound or BindFailed, so the bind never blocks a thread and its result is serialized with every other message. - Binding (initial behavior): Bound assigns the connection, creates the BarrierCoordinator, completes the bind TCS - barrier first, report second, as before - and becomes Ready. BindFailed reports the failure and stops. It must stop rather than throw: a throw from a message handler runs the default supervision directive, Restart, which re-runs PreStart and re-binds the same dead port in a loop. - Anything else arriving during Binding is stashed. The socket starts accepting the moment the bind completes, which is before the actor processes Bound, so a client that connects in that window can get CreateServerFSM in ahead of it. - Ready is the previous receive, now entitled to a non-null connection and barrier by construction of the state machine. - The bind failure is wrapped into ConductorBindException at the pipe, and the single-exception AggregateException a piped failure arrives in is unwrapped, so the reported cause is the socket error itself. - PostStop guards a null connection and no longer blocks on the pool drain. ReleaseAll detaches the pools before it returns, so a later CreateConnection builds fresh ones either way. The caller contract is unchanged: StartControllerAsync still awaits the same TCS, still gets the endpoint only after the barrier exists, and still sees a failure carrying the named error. Node 1 emits its port sentinel at the same point, between the bind and the wait for players. (cherry picked from commit 2ece2c6) v1.5 backport note: the conflicting hunk in MultiNodeTestRunner.cs bundled this change's own _conductorPortSource field together with DefaultNodeExitTimeout, which was pure context here but is itself the payload of dev's earlier #8431 (a node-exit backstop that kills a hung node process after a timeout). v1.5's own #8492 backport of #8431 had dropped this file's hunk, so DefaultNodeExitTimeout did not exist on v1.5 before this pick. Taking dev's whole file for this merge - the only way to give the field a real consumer - therefore also lands the #8431 node-exit backstop on v1.5 for the first time, not just the #8515 port-sentinel change.
…c instead of blocking a pool worker in the constructor (#8558) * De-flake ClusterShardingLeaseSpec: join the cluster in InitializeAsync instead of blocking a pool worker in the constructor The constructor's blocking wait for Up against a flat 3s default missed by 462ms on CI, because the join's six dispatches queued behind the parked worker on a starved thread pool. Cluster formation now happens in IAsyncLifetime.InitializeAsync via JoinAsync/StartAsync, which waits on the real MemberUp signal without parking a thread. The file is also migrated to the async TestKit API per the repo's standing rule. * ClusterShardingLeaseSpec: drop the redundant TestActor touch; the ambient-context attribute already pins the cell (cherry picked from commit 408e008)
…Async no longer leaks the ActorSystem; de-flake StreamRefsSpec (#8545) * TestKit.Xunit: implement the async dispose chain so derived DisposeAsync no longer leaks the ActorSystem; de-flake StreamRefsSpec Fixes #8191: xUnit v3 calls IAsyncDisposable.DisposeAsync() in preference to IDisposable.Dispose() whenever a type implements both, so a derived spec's own no-op DisposeAsync (e.g. StreamRefsSpec) skipped TestKit's whole synchronous dispose chain and silently leaked a remoting-enabled ActorSystem per test. Akka.TestKit.Xunit.TestKit now implements IAsyncLifetime with a virtual InitializeAsync/DisposeAsync pair; DisposeAsync runs the sync chain and then shuts the system down with the non-blocking ShutdownAsync(). Ports the TestKit.cs shape from stalled PR #8217 (targeted v1.5) onto dev; its TestKitBase.cs hunk is dropped because ShutdownAsync already landed on dev via 81289e8, and EventFilterTestBase.cs is hand-merged onto dev's current AwaitAssertAsync retry body. Five other TestKit-derived specs that declared their own DisposeAsync/InitializeAsync are updated to override and chain to base: BugFixSpec, Bugfix8144Spec (Xunit v3), ParallelAmbientContextSpec (both base classes), and EventFilterTestBase. StreamRefsSpec.SinkRef_must_receive_elements_via_remoting is de-flaked on top of the leak fix: the test now gates on WatchTermination(Keep.Right) instead of asserting on a flat 3s wall-clock wait, uses async TestKit calls throughout, and gives every cross-boundary wait a dilated 15s budget. Remote system port is now 0 and loglevel is DEBUG for stage-level tracing. * StreamRefsSpec, BugFixSpec, Bugfix8144Spec: migrate to the async TestKit API Convert ExpectMsg, ExpectNoMsg, and .Wait()-on-stream-completion calls to ExpectMsgAsync, ExpectNoMsgAsync, and awaited AwaitWithTimeout across all facts in the three touched files. * StreamRefsSpec, BugFixSpec: finish the async migration; restore INFO logging * ClusterShardingLeaseSpec: override the TestKit's async lifecycle instead of declaring its own * TestKit.Xunit: one disposed flag; Shutdown moves out of Dispose(bool) so the async path needs no mode switch Dispose(bool) now only runs AfterAll() (and any derived teardown); it no longer terminates the ActorSystem and no longer guards re-entrancy. The guard collapses to a single _disposed flag checked once, in each public entry point: Dispose() runs the chain then calls the blocking Shutdown(), DisposeAsync() runs the chain then awaits ShutdownAsync(). This drops the now-redundant _disposing/_disposingAsync flags and the nested try/finally that used to suppress Shutdown() inside the async path. Verified: no Dispose(bool) override in the tree depends on the ActorSystem being torn down when base.Dispose(disposing) returns (the only overrides in the repo are on System.IO.Stream/GraphStageLogic, unrelated to this TestKit); StreamRefsSpec, BugFixSpec's DisposeAsync overrides still tear down their own systems before chaining to base.DisposeAsync(), and ClusterShardingLeaseSpec no longer declares one at all. Added a Bugfix8191Spec case covering Dispose() called after DisposeAsync(). (cherry picked from commit 450bebf) v1.5 backport note: re-added .WithFallback(ConfigurationFactory.Load()) to StreamRefsSpec's Config(), which the hand-merge of the port/hostname change had dropped. Akka.Streams.Tests still targets net48 on v1.5 (dev has no netfx lane), and its app.config sets stream.materializer.debug.fuzzing-mode = on for that lane. ActorSystem.Create only calls ConfigurationFactory.Load() when no config is supplied, so without this fallback an explicitly configured ActorSystem - like the ones this spec creates - never saw app.config's fuzzing setting, silently losing that coverage on the net48 lane.
…WithinAsync block (#8540) * De-flake ClusterSpec: stop passing an already-cancelled token to the WithinAsync block A_cluster_must_cancel_LeaveAsync_task_if_CancellationToken_fired_before_node_left cancels `cts` to prove LeaveAsync(cts.Token) is cancelled, then reused that same already-cancelled token as the `cancellationToken:` argument to the 10s WithinAsync block that drives Leaving -> Exiting -> Removed. WithinAsync races the block against a delay bound to that token; with the token born cancelled, any real suspension point inside the block before the race resolves throws OperationCanceledException out of WithinAsync (seen on CI as a 71ms failure). Fast/synchronous runs hid the bug; slower runs exposed it. Fix: drop the cancellationToken: argument so the block uses the default token. While here, switch the synchronous ExpectMsg<ClusterEvent.MemberRemoved>() call inside the block to the async ExpectMsgAsync, per the repo's async-only TestKit convention. * ClusterSpec: migrate the remaining synchronous TestKit calls to the async API Migrates all 7 TestKit Shutdown(ActorSystem) calls in ClusterSpec.cs to ShutdownAsync, awaited in their enclosing async Task facts. (cherry picked from commit 9705d36)
…e join is ordered first; one leader pass, then MemberRemoved (#8579) Both LeaveAsync facts hand-ticked LeaderActions() right after Join, with no event wait in between. Cluster.ClusterCore re-targets from the top-level daemon to the core daemon once its ref is published, so a command issued in that window can overtake a still-in-flight join and get dropped. The facts also fired two leader ticks for Leaving -> Removed, though the second one is always a no-op: Exiting -> Removed runs through the coordinated-shutdown round trip the first tick starts, not through another tick. Fix: subscribe before joining and wait for MemberUp as a real barrier before issuing Leave (no tick needed there - a bootstrap self-join already runs a leader pass inline). After Leave, fire exactly one LeaderActions() call and wait for MemberRemoved with an explicit timeout. Drops the WithinAsync wrapper and the ReceiveOneAsync-based retry loop in favor of direct, bounded event waits. (cherry picked from commit ec983f6)
…ted probe waits instead of one flat latch (#8541) The latch built with `new TestLatch(...)` does not dilate with akka.test.timefactor, and the test parked a pool worker in CountdownEvent.Wait while the scheduler's clock loop and the mailbox shared the same thread pool. Rewrote the test to await a TestProbe through three per-phase, dilated ExpectMsgAsync budgets instead of one flat 5s wall, stopping the actor in a finally. The probe sequence ("timeout", "tick", "timeout") also proves the transparent tick was actually delivered, rather than just inferring it from a latch count. Also converts the companion negative-assertion test (no receive-timeout ever set) from a blocking TestLatch wait to ExpectNoMsgAsync on a probe, so it stops parking a pool worker for the duration of the wait. (cherry picked from commit 86e479d)
… its own 3 s window instead of the Within's remaining time (#8566) A zero-count EventFilter.ExpectAsync(0, action) with no explicit timeout, run inside a WithinAsync(max) block, captures the Within's remaining time as its own quiet window, runs the action, then waits out that whole window to prove nothing was logged. The block therefore ends at start + T_action + max by construction, while WithinAsync abandons a still-running block at max + 200ms and fails the fact (since #8516). On a cold CI agent the action's async part (cluster self-join in BugFix3724Spec) doesn't leave enough margin. Pekko's filter always uses akka.test.filter-leeway (3s, dilated) for its quiet window, never the remaining Within time. Pass that window explicitly via the ExpectAsync(count, timeout, action) overload in both specs so the block runs in roughly T_action + 3s instead of racing the outer deadline. (cherry picked from commit 3432d6f)
…t for a flat second (#8587) CI build 131484 (Windows unit tests, 2-vCPU agent) failed Props_must_create_actor_by_producer with a TestLatch.Ready(1s) timeout. The test blocks on a CountdownEvent that is counted down when the actor cell instantiates the actor on a dispatcher thread, and a flat, undilated 1s wait is the whole budget for one dispatcher hop on a stalled agent. Make the fact async and replace the blocking Ready(1s) call with AwaitConditionAsync polling latchActor.IsOpen over a 3s dilated bound (the same budget ExpectMsg gets from akka.test.single-expect-default for a single dispatcher hop). (cherry picked from commit 147a810)
…e surfaces as TimeoutException (#8564) * De-flake InboxSpec: receive through ReceiveAsync so a deadline failure surfaces as TimeoutException, and restore the exact-type assertions Inbox.Receive(timeout) waits on the wall clock while the inbox actor schedules its own deadline for the same duration; when that deadline wins the race, Task.Wait rethrows the actor's Status.Failure(TimeoutException) wrapped in AggregateException instead of the bare TimeoutException the tests expect. Awaiting ReceiveAsync instead surfaces the exception directly with no race, so this migrates every InboxSpec receive call to the async API and drops the IsTimeoutException either-type workaround added by an earlier fix. * InboxSpec: queued-queries fact awaits delays instead of sleeping; exact exception type on the negative-deadline fact Task.Run + Task.Delay staggers the two ReceiveWhere calls without parking a pool thread; ReceiveWhere itself has no async overload and isn't in the race, so it stays as-is. Traced the negative-deadline fact through Inbox.Actor.cs's Kick handling and FutureActorRef<object>.TellInternal: a deadline already in the past always produces a bare TimeoutException("Deadline passed"), never anything else, so ThrowsAnyAsync<Exception> is tightened to ThrowsAsync<TimeoutException>. (cherry picked from commit 989e299)
…and wait out the restart backoff without blocking a pool thread (#8584) * De-flake EntityTerminationSpec: anchor on the shard's own entity set, wait out the restart backoff without blocking a pool thread The passivation fact slept a fixed 400 ms after the test actor saw the entity's Terminated and then read the shard state once. The shard learns of the death through a system message it re-queues as a user message, so the test actor seeing Terminated says nothing about the shard having processed it; on a two-core agent the sleep itself held one of the two pool workers, the shard's mailbox item sat in the other worker's local queue, and the pool only moved again when the sleep expired. The state read then saw the entity still active. Each fact now waits for the shard's active entity set to reach the expected value, which is the observable the assertions read. The passivation fact then waits out the 250 ms entity-restart-backoff with Task.Delay and reads the state once more, deliberately not polled, so a wrongful restart still fails it. The restart fact waits on the shard's own "Started entity" line for the restart instead of a sleep, since a state poll alone can pass on the dead ref before the shard has processed the Terminated. The non-remembering fact drops its sleep and the comment about a backoff that does not exist on that path. The file is migrated to the async TestKit API; the one-node join moves from the synchronous AtStartup into an awaited helper. * EntityTerminationSpec: query the region through a fresh probe per attempt; check for a wrongful restart with a zero-count filter on the shard's own restart line The polling helper's 1s per-attempt bound is shorter than the region's 3s query timeout, so a timed-out attempt left its reply in the test actor's queue, where the passivation fact's final single read could take a snapshot from before the wait as a current one, and the non-remembering fact's "pong-2" expect could dequeue it. Each request now goes through a fresh probe, and the region's Failed set is asserted empty. The passivation fact replaces Task.Delay plus a state read with a zero-count EventFilter on "Started entity" held open for 600 ms, which observes a wrongful restart directly, then a final state read through a probe with a 5 s bound. (cherry picked from commit bef497a)
…fresh probe per attempt under a sharding-derived budget (#8573) * De-flake ShardingBufferAdapterSpec: re-send the cold messages with a fresh probe per attempt under a sharding-derived budget; per-system settings The first message through a fresh region has to survive coordinator allocation and shard start, all ddata majority writes/reads that degrade to "all nodes" on this test's 2-node cluster. If the shard's remember-entities write stalls past its deadline, the Shard restarts and its buffered messages die with it - no dead letter, nothing left to re-deliver. Sharding is at-most-once, so a bare ExpectMsgAsync on the flat akka.test.single-expect-default can time out waiting for a message that no longer exists; waiting longer never helps. FirstMessageThrough re-sends each of the three cold sends through AwaitAssertAsync with a fresh TestProbe per attempt (so a late reply to an abandoned attempt can't satisfy a later one, or leak into the warm phase below), budgeted at ShardStartTimeout + UpdatingStateTimeout (the cold path plus one full Shard restart cycle) with each attempt capped at RetryInterval so the loop actually iterates. The counter snapshots are taken after the loops, so BeGreaterOrEqualTo still holds under retries. The warm-phase sends stay single sends (entities are live; a shard restart here would mean no remember-entities write ever happened, so retrying would hide a real bug) bounded by UpdatingStateTimeout instead of the flat probe default, plus a LastSender check against the original entity ref so a restart between phases can't hide behind the assertions. Also: StartShard was building both regions from Sys's settings instead of the settings of the system whose region is being started (harmless today since sysB is created from Sys.Settings.Config, but wrong); and AfterAll shut Sys down twice (once directly, once via TestKit.Dispose's finally) since _sysA is Sys - removed the redundant call and left a TODO(#8545) for migrating both shutdowns to the async TestKit dispose chain once that lands. * ShardingBufferAdapterSpec: the identity check's comment names the right equality (cherry picked from commit 808de09)
…DistributedPubSubRestartSpec (#8498) * De-flake ClusterSingletonProxySpec and ClusterSingletonRestartSpec Both specs build multi-node clusters in one process and only fail on loaded CI agents: ClusterSingletonRestartSpec.Restarting_cluster_node_with_same_hostname_and_port_must_handover_to_next_oldest on Windows/net10 (build 130838, at 15s) and ClusterSingletonProxySpec.ClusterSingletonProxy_must_correctly_identify_the_singleton on Linux/net10 (build 130872, ~36s in-test). Neither reproduces in isolation - 8/8 and 6/6 green locally, and 30/30 and 24/24 green here under constrained thread pools and CPU pinning. So the work is to take the load sensitivity out of the structure, not to widen anything. No timeout was raised, no sleep and no backoff was added. ClusterSingletonProxySpec The identify test was synchronous throughout: TestProxy blocked on ExpectMsg for up to 25 seconds per node, and the finally block sat in Task.WhenAll(...).Wait(30s). On a two-core agent the pool starts at two worker threads and injects more slowly, so a blocked test thread is a large fraction of the pool that the five ActorSystems need in order to gossip, heartbeat and deliver the very reply being waited on. The test is now async: TestProxyAsync awaits ExpectMsgAsync and the teardown awaits WhenAll. Awaiting the teardown also matters on its own - the discarded Wait(30s) result meant five cluster systems could still be shutting down while the next test in the assembly ran. The bigger correctness gap was that nothing waited for the singleton to exist. The first TestProxy started before the cluster had formed, so cluster formation, singleton startup and proxy identification all had to fit inside a message timeout. Two gates replace that: every node must see all five members Up, and then each proxy must publish IdentifySingletonResult.Success. The subscription is made in the ActorSys constructor before the proxy actor is created, so the event cannot be missed, and the wait fishes for Success because the proxy also publishes Timeout results on a timer that starts with no initial delay. ClusterSingletonProxy_with_zero_buffering_should_work had the same gap with sharper teeth. It waited for membership and then sent one message into a proxy configured with BufferSize 0, which drops anything it cannot forward (ClusterSingletonProxy.Buffer). Membership does not imply identification, so that send could vanish and the test would wait out its full 25 seconds for a reply that was never coming. It now waits on the identification result. It also terminates its seed node, which was left running for the remainder of the assembly. Two waits in ClusterSingletonProxySingletonTimeoutTest2 read the identify result with AwaitAssertAsync around ExpectMsgAsync, both unbounded. AwaitAssertAsync resolved akka.test.single-expect-default (3s) and so did the inner ExpectMsgAsync on a different TestKit instance, so the budgets did not nest and the loop got about one attempt - against a queue holding the Timeout results the proxy emits every 500ms under that config. Both are now FishForMessageAsync with an explicit 30s bound. The AwaitConditionAsync in ClusterSingletonProxySingletonTimeoutTest gets an explicit bound for the same reason; with none it inherited the 3s default for a two-node join. ClusterSingletonRestartSpec The spec is now async end to end. Three points of substance: Shutdown(_sys1) waits Terminate().Wait(Dilated(5s)) and, when that expires, force-stops the user guardian and logs a warning - verifySystemShutdown defaults to false, so nothing fails. That stops /user but not /system, so remoting keeps sys1's listener bound; the very next statements create sys3 on that same host:port. dot-netty's tcp-reuse-addr is off-for-windows, which is the kind of asymmetry that produces a Windows-only failure. sys1 also owns the singleton at that moment, so a truncated shutdown cuts the hand-over to sys2 short. await _sys1.Terminate() lets CoordinatedShutdown finish: the hand-over completes and the port is released. Each proxy assertion created a TestProbe inside its retry loop. CreateTestProbe blocks the caller until the probe's PreStart has run on the test-actor dispatcher (TestKitBase.cs:739) and replaces the calling thread's SynchronizationContext (TestKitBase.cs:194) - so every attempt added a blocking wait on work that needs a pool thread, a context swap and a leaked system actor. AwaitProxyReplyAsync builds one probe per phase and retries only the send. Replies stay useful across attempts because the message is a plain echo. The 15s window that CI reported waits for sys2 to be fully Removed from sys3's view. Getting there runs sys2's cluster-exiting CoordinatedShutdown phase, which blocks on the singleton hand-over (ClusterSingletonManager.SetupCoordinatedShutdown) and gives up after 10s. JoinAsync only proves sys3's own view of the cluster, so sys2 could still see sys3 as Joining when it left - no hand-over target, phase runs to its timeout, and the removal no longer fits in 15s. A convergence gate now requires both sys2 and sys3 to see the same two-member Up cluster before the Leave. Smaller items: the Within wrappers are gone and each AwaitAssertAsync carries the same bound explicitly, so no wait resolves an ambient deadline; the join assertion reads one Cluster.State snapshot instead of two; the join retry runs at 500ms rather than 100ms, because a JoinTo arriving while the daemon is in TryingToJoin drops it back to Uninitialized and restarts the handshake (ClusterDaemon.cs:1273) - ten of those a second is load on the daemon the test is waiting for. The join must still be re-issued, since sys3 reuses sys1's address and is refused until the old incarnation is removed; a single Join would then wait out retry-unsuccessful-join-after (10s). sys1/sys2/sys3 logs now reach the test output via InitializeLogger, as in ClusterSingletonRestart2Spec. Verified: both specs green in constrained-pool loops (DOTNET_ThreadPool_ForceMaxWorkerThreads=3), full Akka.Cluster.Tools.Tests suite green, build clean with -warnaserror. * De-flake DistributedPubSubRestartSpec: resolve the association before asserting first's restart-kill loop asserted on ActorIdentity.Subject after a 2s Identify window. Use ActorSelection.ResolveOne instead, inside the same retry: the Identify round trip is the delivery confirmation, its temp actor is fresh per attempt, and a final failure raises ActorNotFoundException naming the path that never resolved rather than a bare timeout on a null Subject. Resolve and kill stay in one retry on purpose. Splitting them lets artery's ordinary outbound stream drop the kill in the gap between a successful resolve and the Tell, with no resend - a local artery soak reproduced exactly that. Also: - read the baseline DeltaCount on a probe with an explicit bound instead of on TestActor, which is subscribed to topic1 - name the same-port rebind dependency when the old system fails to stop inside third's WhenTerminated wait - fresh probe and an explicit 1s bound per Count attempt, so the 10s gossip window retries on the 500ms tick instead of spending itself on three 3s waits (cherry picked from commit 9b0fc9f)
…d-off (#8500) `ClusterShardingSpec` (base class for the Persistent/DData/WithEntityRecovery variants) failed in CI on both transports with the same shape: one node timed out waiting for an `Int32` counter reply and every other node then failed the next barrier. Mechanism. `HandOffStopper` replies `ShardStopped` the moment the last entity terminates - before the `Shard` actor stops, and before the `ShardRegion` observes `Terminated` and drops the shard from `_shards`/`_regionByShard`. The spec treats `ShardStopped` as "the shard is gone" and immediately sends a `Get` through the region, so the region forwards it into the dying shard: [.../RememberCounterEntitiesRegion/1] DeadLetter from [.../system/testActor1] to [.../RememberCounterEntitiesRegion/1]: Get { CounterId = 13 } Sharding is at-most-once, so the request has to be re-sent, not merely re-awaited. Two sites did not do that: * `recover_entities_upon_restart` wrapped `Get(13)` in `AwaitAssertAsync(5s)` but the inner `ExpectMsgAsync(0)` carried no bound, so it inherited `akka.test.single-expect-default` (5s) - the entire budget. The loop made exactly one attempt: AwaitAssert failed, timeout [00:00:05] is over after [1] attempts and [00:00:05.0055266] elapsed time * `permanently_stop_entities_which_passivate` sent `Get(25)` as a bare one-shot, so a single dropped message was an unconditional failure. Both now re-send with a fresh probe per attempt and a 1s per-attempt bound, so the 5s budget admits four sends instead of one. A fresh probe per attempt also keeps a late reply from a timed-out attempt out of the next attempt's queue; the same defect made the `Identify(4)` retry loop in the second phase a no-op, and it is bounded now too. The second phase also ran its three barriers inside `WithinAsync(15s)`. `EnterBarrierAsync` derives its timeout from `RemainingOr(barrier-timeout)`, so the rendezvous got the Within remainder instead of the configured 70s: EnterBarrier(Name: after-13, Role: [RoleName(sixth)], Timeout:00:00:13.8157551) timeout while waiting for barrier 'after-13' The Within is removed, matching the treatment already applied to `recover_entities_upon_restart`; its 15s is re-attached to the two waits it was actually protecting, so no wait ends up with a smaller budget than before. Verified by widening the hand-off window in a local throwaway build (delaying the Shard's stop after `ShardStopped`), which reproduces both CI signatures exactly - `Timeout 00:00:05 ... System.Int32` with `barrier failed:after-shard-restart`, and `Timeout 00:00:13.9 ... System.Int32` with `barrier failed:after-13` - and passes with these changes. Test-only change. (cherry picked from commit e6e014d)
… DistributedPubSubRestartSpec baseline (#8509) * Pin the legacy allocation strategy in ClusterShardingLeavingSpec The spec asserts that entities on surviving nodes keep the same actor incarnation across a graceful leave. Only the legacy threshold strategy has that property. The bounded default may move survivor shards during the leave window, which restarts their entities under new refs - that is the optimizer working as designed, not a regression, and the reference implementation behaves the same way. Two CI failures carried this exact signature: the DData variant in build 130784 and the Persistent variant in build 130956, both an entity ref identity mismatch followed by an after-4 barrier cascade. Pinning rebalance-absolute-limit = 0 selects the legacy strategy, so the identity assertion tests the algorithm that guarantees it. The spec becomes a regression test for that strategy's leave behavior instead of a coin flip against the bounded default. Closes #8481. * De-flake ReplicatorChaosSpec: rendezvous the replicators before the write burst `ReplicatorChaosSpec` failed three times across two CI builds, on the Artery Linux and Artery Windows lanes, with one fingerprint every time: [Node1:first] Timeout (00:00:03) while expecting 15 messages. Only got 10 [Node2:second] Timeout 00:00:00.0000044 while waiting for a message of type Akka.DistributedData.GetSuccess [Node3..5] barrier failed:update-during-split-verified Root cause. Nothing makes the five replicators ready at the same time. Each node checks `ReplicaCount(5)` against its own replicator and then walks straight into the update phase, so `first` can start writing while another replicator has not yet processed `first`'s MemberUp. A replicator drops a Write whose sender it does not know: [.../user/replicator] Ignoring message [Write] from [.../user/replicator/$a] unknown node [UniqueAddress: ...] and sends no ack. `WriteAll` cannot recover from that. Under WriteAll every node is primary, so `SendToSecondary` re-sends to nobody, and a GCounter update leaves `WriteAggregator._delta` null, so the delta re-send does not run either. One dropped Write costs the whole update. Build 130934 shows `first` issuing its first Update at 40.136 and `fifth` receiving its Welcome at 40.977; `fifth` ignored all five Writes 12ms after that, and all five KeyC updates timed out. Every node now proves `ReplicaCount(5)` and meets at a new `replicas-ready` barrier before anything is written. Second defect, and the reason the failure read as "Only got 10" instead of naming `UpdateTimeout`. `ReceiveN(15)` sat outside any `Within` and inherited `akka.test.single-expect-default` - 3s, the same 3s as the `WriteAll` deadline of the updates it collects. An aggregator answers no earlier than its own deadline, counted from when the replicator picks the Update up, which is at or after the `ReceiveN` call. The collect can therefore never see that answer; it is a dead heat by construction. In build 130934 the aggregators replied at 43.226, roughly 70ms too late, and the ten replies that did arrive were exactly the WriteLocal ones, which need no aggregator. Every expect that waits on a `WriteTo`/`WriteAll` reply now gets `WriteReplyTimeout` - the aggregator deadline plus one single-expect-default of slack, with the arithmetic stated in a comment on the field. Third defect. `AssertValue` ran `Within(10s)` around an `AwaitAssert` whose inner `ExpectMsg<GetSuccess>()` carried no bound, so each attempt could consume the whole remaining budget and the final attempt got almost none - the logged `00:00:00.0000044`. That reports a value which never converged as a bogus zero-budget timeout. `AssertDeleted` and the `ReplicaCount` loop had the same shape. All three now use a fresh probe and a 1s bound per attempt, the treatment `ClusterShardingSpec` got in #8500. Also swept. The `TestConductor.Blackhole`, `PassThrough` and `Exit` calls used `.Wait(timeout)` and discarded the returned bool, so a conductor round-trip slower than the 1s cap let the test walk on as though the partition were already installed. They are awaited now, bounded by the conductor's own 30s query-timeout, and a failure raises instead of vanishing. The spec is converted to the async TestKit API throughout, which removes the remaining `.Wait()` calls. Verified fail-first. A throwaway build that holds one replicator's membership events back by 4s reproduces the CI fingerprint on all five nodes exactly, with the same five ignored Writes; the same injection passes with these changes and produces zero ignored Writes. Bounding `ReceiveN` on its own converts the failure into the honest assertion `Expected 'UpdateSuccess' got 'UpdateSuccess','UpdateTimeout'`, and bounding the inner expect makes all five nodes report `Assert.Equal() Failure: Values differ` where four of five previously reported the zero-budget timeout. Test-only change. * De-flake DistributedPubSubRestartSpec: read the DeltaCount baseline after gossip goes quiet Build 130927 (artery, Windows) failed on second with "Expected value to be 3L, but found 4L" - one extra Delta arrived between the baseline read and the post-restart read. This is a different signature from the one PR #8498 fixed on this spec, and it is not caused by third's restart at all. Mechanism, confirmed by instrumenting the mediator locally and reading the counter it actually exports: DeltaCount counts Delta MESSAGES received, not registry changes. The mediator's Delta handler increments _deltaCount and only then checks _nodes.Contains(bucket.Owner), so a Delta whose payload is thrown away is still counted. While second is still waiting for third's MemberUp, first pushes third's bucket on every 500ms gossip tick and second discards each push - all counted. Convergence therefore arrives as a burst, and redundant Deltas trail it, because a peer keeps re-sending until second's next outbound gossip tells it that second caught up. CountAsync unblocks on the first push second actually merges, so it returns while the burst is still draining. The old code read the baseline immediately after that. Over 20 local runs the last Delta landed just 104-1010ms before that read, against a 500ms gossip tick - one delayed tick puts a trailing Delta on the far side of the read and the invariance assertion fails through no fault of the restart. So the assertion was right and the baseline was wrong: "no new Deltas during third's restart" only holds if the baseline is taken once gossip has settled. ReadStableDeltaCountAsync now requires two samples 2s apart to agree before accepting a baseline. 2s is 4 ticks of this spec's 500ms gossip-interval - enough for our outbound tick plus the peer's reaction, with slack. Fresh probe per sample, matching CountAsync, so a late reply cannot be misread as the next sample. Verified by fault injection rather than by luck. Delivering exactly one extra Delta to the mediator 1s after the baseline query - precisely the event the CI log implies - reproduces the reported failure verbatim on the old code ("Expected value to be 3L, but found 4L" on second) and passes on the new code, where the gate re-samples and absorbs it. The post-fix margin between the last Delta and the baseline read is 3s instead of ~150ms. Soak: artery 8/8 under CPU load, classic 4/4. No timeout was raised, no sleep and no backoff added; the only change is when the baseline is sampled. (cherry picked from commit e340146)
… contact, and the survivor waits on it (#8552) * De-flake DistributedPubSubRestartSpec: the restarted node makes first contact, and the survivor waits on it first's association to third's old incarnation carries no signal that a same-address restart even started; the survivor's association only learns the new incarnation exists from a frame the restarted node itself sends. The old test started its 45s clock at Shutdown()'s return, before any of third's restart cost had begun, and none of the ten prior fixes changed who makes first contact after the restart - they all kept first polling a stale association instead. Now third's Shutdown actor pings a forwarder on first the moment it exists (PreStart), so first waits on that ping (dilated 60s) before running a short 20s closed-loop kill over ActorSelection.Tell (only an ActorSelectionMessage pierces a quarantined association). The ping's own inbound handshake is what heals first's association as a side effect. barrier-timeout goes to 180s to fit the new worst case with headroom, and a bound-address assertion on third settles, on the next Linux failure, whether the fresh listener came up on the pinned port. * DistributedPubSubRestartSpec: repeat the ready ping until acknowledged, use the conductor's async shutdown, keep the port assertion inside the cleanup Addresses review findings on PR #8552: - Shutdown.PreStart no longer sends the ready ping once; it resends every 500ms on an undilated self-scheduled timer until first's ReadyPingForwarder acks it back to the sender, cancelling the timer on ack (and defensively in PostStop). first still waits once for the first ping it sees and ignores later resends. This removes the dependency on any Artery handshake-stage fix - a lost ping is simply retried instead of failing the 60s wait. - Replace TestConductor.Shutdown(...) with the ShutdownAsync(RoleName, ...) overload, keeping the same 30s bound. - Move the restarted-address port assertion inside the try/finally that owns newSystem, so a failed assertion still terminates it. - State the measured worst case behind the 60s ready-ping wait and re-derive newSystem.WhenTerminated's bound (120s -> 155s) against the current 110s worst-case pipeline, restoring its original 45s of slack. * DistributedPubSubRestartSpec: let the shutdown ack leave before the restarted system terminates Build 131332 (this PR's own CI run, Linux Artery) showed third's Shutdown actor replying "shutdown-ack" and calling Context.System.Terminate() inline in the same handler; Artery aborted third's outbound streams before the ack could flush, so it never reached first. First's closed-loop kill then retried "shutdown" every 500ms against an already-exited process and its 20s bound timed out. Local runs only passed because the ack happened to win that race on a fast machine. Slide the terminate behind a 2s timer instead of calling it inline - same shape as the Subject actor in RemoteNodeRestartDeathWatchSpec (PR #8557) - so a live system, and the warmed-up lane under it, always outlasts whichever "shutdown" attempt first's loop last sends. Cancel the timer in PostStop. * DistributedPubSubRestartSpec: terminate timer outlives the retry cycle; restore the 120 s termination wait; last sync call migrated (cherry picked from commit 2cc9d1e)
…ers the loop (#8542) 'second' and 'third' do no work of their own, so they reach the "after-1" barrier about a second into the run while 'first' is still sending its 500 letters. BarrierCoordinator arms a barrier's clock on the first arrival and only ever shortens it, so the 30s default barrier-timeout - not the 5s per-letter wait - is what bounds the whole loop. On a saturated CI agent the loop needs longer than that, the barrier fails, the relay nodes stop, and 'first' then fails its next per-letter wait because the route it was using has gone away. EnterBarrierAsync asks the coordinator for RemainingOr(barrier-timeout), so a Within around the final barrier sets that budget. 300s over 500 letters permits a 600ms mean round trip, about 3x the slowest mean measured on such an agent and far below the 5s per-letter deadline, which remains the assertion that reports a real drop. Scoped to the barrier rather than set as akka.testconductor.barrier-timeout, because that key also drives the teardown poll in MultiNodeSpecAfterAll, where AwaitCondition polls at max/10 and the first poll always misses. Raising the key to 300s adds 30s of wall clock to every run of this spec; measured here at 30s, 60s, 100s and 300s, the tax tracks barrier-timeout/10 exactly. The loop, its route, its per-letter assertion and the letter count are unchanged, so the spec proves exactly what it proved before. The rest of the diff is the file's migration to the async TestKit API. (cherry picked from commit e62aa19)
…the spec's own budget instead of a probe's flat 5 s default (#8544) * De-flake ClusterShardingRegistrationCoordinatedShutdownSpec: wait on the spec's own budget instead of a probe's flat 5 s default On the Artery lane the coordinator singleton hand-off from `second` to `first` takes about 5.6 s: 5 s of that is a replicator on `first` that never learned `second` was a cluster member (its subscription lands after `second` is already `Leaving`), so the leaving coordinator burns its whole `updating-state-timeout` on a ddata write that can never reach quorum. The shutdown task's `probe.ExpectMsg(1)` only waited a flat 5 s, because a `TestProbe` is its own `TestKitBase` with its own deadline and never saw this spec's `Within(30s)`. The JVM spec sends from its own test actor and waits on the enclosing `within`, so it never had this gap. Send with `_region.Value.Tell(1, TestActor)` and wait with `ExpectMsg(1)` on the spec itself, so the wait inherits the spec's 30 s budget. Pin `before-cluster-shutdown`'s phase timeout to 30s so an incomplete task can't be cut short by the default 5s phase timeout either. * ClusterShardingRegistrationCoordinatedShutdownSpec: migrate to the async TestKit API Moves Within, Join, RunOn, AwaitAssert, AwaitCondition, EnterBarrier, and ExpectMsg calls to their awaited Async variants; the coordinated-shutdown task on `third` becomes an async body, and the csTaskDone probe wait gets an explicit 20s dilated bound (measured 5.6s handoff, ~3.5x margin). * ClusterShardingRegistrationCoordinatedShutdownSpec: bound the registration wait explicitly and size the outer window above its contents * ClusterShardingRegistrationCoordinatedShutdownSpec: pass the explicit bounds undilated, since ExpectMsgAsync dilates them itself (cherry picked from commit 9c75211)
…turn lane before sending End (#8563) * De-flake ClusterDeathWatchSpec: resolve first's end actor over the return lane before sending End; migrate the file to the async TestKit API On the Windows Artery lane, node fourth's fresh EndSystem sent End to first's /user/end and waited 15s for EndAck. The ack has to ride a brand-new outbound Artery lane from first to EndSystem, and first was observed terminating its ActorSystem ~340ms after replying, before that lane finished materializing, handshaking and connecting - so the ack lost the race and the wait timed out. Fix: before sending End, resolve first's /user/end from EndSystem with ActorSelection.ResolveOne. The ActorIdentity reply can only travel back over the same outbound lane the EndAck will use, so a successful resolve proves that lane is already up. This is safe because first is parked in ExpectMsgAsync<End>() and cannot start tearing down until the End we have not sent yet arrives. Bound the resolve at a dilated 8s and the subsequent EndAck wait at an explicit, undilated 5s - 8 + 5 = 13s, strictly narrower than the 15s single-expect-default the probe used to inherit unbounded. Also switch the EndSystem teardown to the async ShutdownAsync instead of the blocking Shutdown. While in the file, migrated every remaining synchronous TestKit call to its async equivalent: the RemoteWatcher property became an async GetRemoteWatcherAsync helper, TestLatch.Ready() became an AwaitConditionAsync poll on TestLatch.IsOpen with the same 5s budget, and the two remaining RunOn calls became RunOnAsync. No timeout was widened and no assertion was loosened. * ClusterDeathWatchSpec: state the shutdown bound honestly against the enclosing window; drop a dead null check (cherry picked from commit a00df61)
…dicated probe, as upstream does; migrate the file to the async TestKit API (#8569) Cluster.Leave triggers one MemberRemoved(self) EventStream publication that drives two independent, unordered subscribers: the RegisterOnMemberRemoved callback that tells the test actor "MemberRemoved", and the proxy's own self-stop on seeing MemberRemoved for its node, whose death-watch Terminated lands in whatever queue is watching it. When the test actor did the watching, both messages competed for the same FIFO queue, so ExpectMsg("MemberRemoved") could dequeue the Terminated meant for the following ExpectTerminated instead (observed on the Linux Artery lane, node "first"). Watching echoProxy from a dedicated TestProbe, as upstream Akka JVM/Pekko and the sibling ClusterSingletonManagerLeaveSpec already do, gives Terminated its own queue and removes the race without touching any timeout, barrier, or assertion. The lazy proxy accessor becomes Lazy<Task<IActorRef>> so the watch can be registered asynchronously while keeping create-on-first-use, single-execution semantics. The rest of the file moves to the async TestKit API (RunOnAsync, WithinAsync, AwaitAssertAsync, ExpectMsgAsync, EnterBarrierAsync) to support that; the ForEach-based negative-assertion loop becomes a for loop, and its Thread.Sleep(1000) becomes Task.Delay(1000), unchanged as a deliberate 1s soak between probes rather than a stand-in for a real event. (cherry picked from commit b2524f6)
…esend-derived bound, identify asserts the actor exists (#8581) * De-flake RemoteNodeDeathWatchSpec: barrier after the subject stops so the watcher's wait starts on the real event; resend-derived bound; identify asserts the actor exists Phase 1 had first waiting a flat 3s (single-expect-default) for WrappedTerminated while second independently slept 3s before calling Sys.Stop, with no barrier between the two sleeps and the stop. The window had to absorb both timers' drift plus the remote notification hop, and a lost DeathWatchNotification is only resent every 2s. second now does a local WatchAsync + Stop + ExpectTerminatedAsync before entering a new "subject-stopped-1" barrier; TellWatchersWeDied notifies remote watchers before local ones, so the local Terminated proves the remote notification already left. first's wait now opens after that barrier and uses an explicit 6s bound derived from akka.remote.resend-interval (two resend cycles plus a hop), not the failure detector (the peer stays alive and heartbeating here, so that path never fires). Applied the same derived bound to the equivalent wait in phase 4, which already had the correct barrier shape. Restored the identify helper's dropped assertion (ActorIdentity.Subject must not be null, naming the actor path) so a missing actor fails at the identify instead of surfacing as an unrelated Ack timeout downstream. Also migrated the file's remaining synchronous TestKit calls to their async counterparts: the remote-watcher lookup, the identify helper, RunOn, and ReceiveN, plus an explicit bound on AssertCleanup's previously-unbounded nested expect. * RemoteNodeDeathWatchSpec: name the real remote watcher in the comment, note Artery's resend interval, drop the AssertCleanup bound The remote watcher recorded on the subject is first's /system/remote-watcher, which re-issues the ProbeActor's watch over the wire; the comment named the local watcher. The 6s bound reads the classic resend-interval; Artery resends at 1s, so the same budget covers it. The 1s bound on AssertCleanup's nested expect is removed: the predicate overload asserts rather than filters, so the retry loop was already content-driven, and a timed-out attempt could offset later requests from their replies. (cherry picked from commit 1ad1abe)
…t pause over the measured stall, and make the payload listener match real log messages (#8556) * De-flake NodeChurnSpec: turn off legacy auto-down, widen the heartbeat pause over the measured stall, and make the payload listener match real log messages MultiNodeSpec.BaseConfig empties cluster.downing-provider-class, so this spec's auto-down-unreachable-after setting always selected the legacy AutoDowning provider (no quorum logic) rather than the keep-majority SBR its deprecation warning promises. On build 131192 a transient-system teardown produced a 5.6s heartbeat gap on one fixed node; the other two marked it unreachable, auto-downed it, and it auto-downed them right back before its heartbeats resumed - a permanent split that made round 2 wait 25s for 9 members in a cluster that was 8 + 1. Turn auto-down off (this spec already removes every transient member with an explicit Leave or Down and does not need a downing provider) and widen acceptable-heartbeat-pause to 10s as a second guard over the same measured stall, as a stopgap until the underlying product-side blocking wait is fixed (issue #8549). Also add prune-gossip-tombstones-after = 1s, matching Pekko, so gossip tombstones actually prune inside a 5-round run. Separately, LogListener's `info.Message is string` guard could never match: a parameterized Info log is wrapped in LogMessage<LogValues<T1,T2>>, never a string, so the assertions this spec exists to make (gossip payload size growth) could never fail. Match on info.Message?.ToString() instead, like RemoteMetricsSpec already does. Also add Pekko's missing end-of-round barrier so a fast node cannot start creating new transient ActorSystems while a slow peer still holds two dying ones, and replace the stale comment above Task.WhenAll with the corrected explanation of why sequential termination would not help: the scheduler's blocking stop is registered as the first termination callback and therefore runs after the awaited task completes, whether Terminate() calls are sequential or concurrent. * NodeChurnSpec: match the .NET type name in the payload listener; fix the tombstone and auto-down comments; keep the teardown inside the barrier (cherry picked from commit 986e473)
…te (#8404) The reconnect loop retried inside a Within(5s) window that opens at the same unreachability event that gates the restarted address for retry-gate-closed-for = 5s, leaving ~0ms of usable budget — every retry dead-lettered inside the gate. Widen the window to 30s with a fresh probe and bounded expect per attempt, drop the gate to 1s in the spec config, and convert the spec to async TestKit methods. (cherry picked from commit 2679bef)
Aaronontheweb
force-pushed
the
backport/dev-2026-09-to-v1.5
branch
from
September 12, 2026 02:23
46c8068 to
afae265
Compare
Member
Author
|
/azp run |
|
Azure Pipelines successfully started running 1 pipeline(s). |
…ssSpec CI deadlock) (#8359) * Akka.Cluster: make Cluster extension startup non-blocking The Cluster constructor blocked on GetClusterCoreRef().Result (an Ask to /system/cluster bounded by akka.actor.creation-timeout), and the internal ClusterCore getter re-issued that blocking ask whenever it observed an unresolved core ref. Actors spawned during construction (cluster event subscribers, the SBR downing provider) hit the getter from PreStart and parked dispatcher threads; on small pools the parked waiters starved the very daemon that had to reply, the redundant asks timed out, and the timeout path called Shutdown() on an already-healthy node. This is the root cause of the StressSpec nondeterminism on 2-core Windows CI agents and reproduces as a hard deadlock with a 1-thread default dispatcher. - Replace the constructor ask with fire-and-forget InternalClusterAction.Init(cluster); the constructor is now straight-line code whose completion depends on no other thread - ClusterDaemon and ClusterCoreSupervisor become explicit Uninitialized/Initialized state machines (Become + unbounded stash): pre-Init traffic is buffered, post-Init unmatched traffic forwards to the core daemon; ordering is guaranteed by mailbox FIFO, not blocking - ClusterCoreSupervisor hands the core ref back via Cluster.SetClusterCoreRef(); the ClusterCore getter is now unconditionally non-blocking (_clusterCore ?? _clusterDaemons) - Delete GetClusterCoreRef() and its Shutdown()-on-timeout path; akka.actor.creation-timeout is no longer consulted by Akka.Cluster; genuine core-startup failures still shut the node down via the existing ClusterCoreSupervisor supervision strategy - Add ClusterStartupFuzzSpec: randomized startup fuzz under pinned 1-2-thread dispatcher pools; killed 63-91% of iterations pre-fix and gates this class of regression at zero kills - Record the behavioral change in BREAKING_CHANGES_V1.6.md - dotnet format applied to the touched files (includes pre-existing whitespace cleanup; diff is unchanged with whitespace ignored) Verification: fuzz spec 1,980/1,980 iterations reached Up (0 kills, 0 soft-timeouts) across baseline / 2-core / 1-core / CPU-saturated / combined conditions; StressSpec passes 3/3 under taskset -c 0,1 where it previously failed in ~90s with the CI signature; Akka.Cluster.Tests 379/379; Cluster.Tools 98/98; Cluster.Metrics 44/44; Sharding 191/192 (the 1 failure is pre-existing on dev, unrelated); no public API changes (Akka.API.Tests 18/18). * Address slopwatch SW003 findings in ClusterStartupFuzzSpec Suppress six SW003 (empty catch block) findings in the fuzz spec's teardown/best-effort paths via slopwatch-ignore comments with justifications. These catches are intentional: cancellation on the hard test deadline, best-effort probing of a cluster mid-race, and bounded/best-effort teardown of burner tasks and the actor system. No behavior change. (cherry picked from commit 26e32af)
…n's lifetime so their order survives startup; LeaveAsync re-sends its Leave (#8580) Cluster.ClusterCore used to switch from /system/cluster to a direct reference to the resolved core daemon the instant that ref was published during startup. Akka's FIFO guarantee holds only per (sender, receiver) pair, so a Join and a Leave issued back-to-back by the same caller with no await between them could ride different mailboxes and arrive at the core daemon out of order. A Leave that overtook its own JoinTo was dead-lettered by ClusterCoreDaemon's Uninitialized behavior, and because LeaveSelf only ever sent its Leave once, that loss was permanent. ClusterCore now always targets /system/cluster, keeping the (sender, receiver) pair fixed for the life of the extension; LeaveSelf also now re-sends its Leave command on every call (Leaving(address) is idempotent), so a lost Leave no longer wedges LeaveAsync forever. Adds a regression fact in ClusterSpec covering Join immediately followed by Leave with no await between them, and amends the existing (cherry picked from commit 59d5da6)
Aaronontheweb
force-pushed
the
backport/dev-2026-09-to-v1.5
branch
from
September 12, 2026 13:57
afae265 to
d60e87d
Compare
Member
Author
|
/azp run |
|
Azure Pipelines successfully started running 1 pipeline(s). |
… member-up signal (#8398) Root cause: MemberUp cannot arrive before the 1s periodic-tasks-initial-delay leader tick, but the spec waited on a raw non-dilated TestLatch.Ready(1s) -- a 100-400ms margin Windows CI load routinely eats. Also a non-dilated 1s ResolveOne raced fire-and-forget cluster-extension init, and a fixed DotNetty port (15508) could collide across processes on shared agents. - ResolveOne wrapped in dilated AwaitAssertAsync (still fails if the downing provider is dead -- the #5962 regression signal) - TestLatch -> TaskCompletionSource(RunContinuationsAsynchronously) awaited with a dilated timeout OUTSIDE the EventFilter block (no continuation inlining onto the member-listener's channel-executor thread) - EventFilter.DeadLetter quiet window via the explicit dilated 2s overload (> SBR's 1s self-tick, so a dead resolver is still detected; test ~3s faster) - port = 0 + explicit cluster.Join(SelfAddress) replaces fixed port/seed-nodes Verified 50/50 consecutive Release runs (net10.0). (cherry picked from commit 5f225ed)
Cherry-picking #8545 from dev recreated this file (its modify/delete conflict was resolved toward dev's content so the pick stayed faithful to what it was on dev). v1.5 has no v1.6 release cycle and no equivalent ledger, so the file does not belong here. Removing it once, on top of the cherry-picks, rather than dropping the hunk from the pick itself, keeps that cherry-pick's diff identical to its dev counterpart.
DistributedPubSubRestartSpec.cs (backported from #8552) dual-pinned the fresh ActorSystem's port for both transports, including akka.remote.artery.canonical.port, and its surrounding comments explained the fix using Artery-specific classes (OutboundHandshakeStage, InboundHandshakeStage.HandleReq, ArteryRemoting.Send) and a CI run captured on the Linux Artery lane. v1.5 has no Artery transport at all (src/core/Akka.Remote/Artery does not exist on this branch), so the second config line is dead HOCON and the class names in the comments point at code that isn't there. Dropped the artery.canonical.port line (only dot-netty.tcp.port is needed on a classic-only branch) and reworded the affected comments to describe the same reasoning in transport-neutral terms, without naming Artery classes. No behavior change: the spec exercises the classic DotNetty transport exactly as it did before this commit. Checked every other file touched by this backport for "artery"/"Artery" mentions; the remaining ones (ClusterDeathWatchSpec.cs, NodeChurnSpec.cs, RemoteDeliverySpec.cs, RemoteNodeDeathWatchSpec.cs) only reference Artery in passing, as comparison context for a classic-transport fix, and set no Artery configuration - left unchanged.
#8499 added TestKitBase.ShutdownAsync (public virtual) and its protected ActorSystem overload. The API-approval snapshots kept during that cherry-pick were deliberately v1.5's own pre-existing files (not dev's - the two branches' TestKit surfaces differ by roughly 300 lines elsewhere), so CoreAPISpec.ApproveTestKit was failing after the pick: v1.5's public API now has the two new methods but the approved snapshot did not. Ran `dotnet test src/core/Akka.API.Tests --filter ApproveTestKit` and accepted the resulting .received.txt as the new DotNet.verified.txt. The only delta from the previous baseline is the two new ShutdownAsync overloads and the compiler's incidental renumbering of nearby async-iterator state-machine class names (adding members shifts Roslyn's sequential ordinals for unrelated nested types declared after them; this is expected and harmless). The Net.verified.txt (net48/netstandard2.0) counterpart could not be regenerated by running the net48 test host in this environment (the vstest agent fails to negotiate over the .NET Framework runtime here). Derived it instead: TestKitBase.cs has no #if/target-framework-conditional code, and diffing the previous DotNet/Net verified pair showed they differed only in the single assembly-level TargetFrameworkAttribute line - every other line, including all compiler-generated class ordinals, was already identical. Applied the same content with only that one line swapped back to ".NETStandard,Version=v2.0", and confirmed the result still differs from the new DotNet file by exactly that one line, matching the pre-existing pattern. Verified: `dotnet test src/core/Akka.API.Tests --framework net10.0` - 18/18 passed, including both ApproveTestKit and ApproveTestKitXunit2.
…ard2.0 has no process-tree kill overload) The #8515 cherry-pick's MultiNodeTestRunner.cs conflict was resolved by taking dev's full file (needed to bring in #8431's node-exit-timeout backstop alongside #8515's port-sentinel change, so the new DefaultNodeExitTimeout field would actually be used - see that commit's message). Building Akka.Cluster.Tests.MultiNode surfaced a genuine v1.5 incompatibility that dry-run text diffing could not catch: dev compiles Akka.MultiNode.TestAdapter against net10.0, where `Process.Kill(bool entireProcessTree)` exists, but this project's target on v1.5 is netstandard2.0 (a single TFM), whose Process surface only has the parameterless `Kill()`. Switched to `process.Kill()`, matching the convention already used for the same reason in Akka.MultiNode.RemoteHost/RemoteHost.cs. The only behavior difference is that a hung node's child processes (if it spawned any) are no longer force-killed alongside it - the backstop still kills the node process itself and still unblocks the runner.
…wn (port of #8543) Hand-port of dev's #8543 substance, not a cherry-pick - #8543 is the last commit of a five-deep stack (#8372, #8427, #8429, #8499) and its auto-merged parts would not compile as-is on v1.5 (BuildConfig's arithmetic and the ClusterResultAggregatorAsync call sites are written against dev's 10-node/1-per-phase config, which v1.5 does not have). What changed, and why each piece is still correct on v1.5: * acceptable-heartbeat-pause raised from 3s to 20s, with the arithmetic comment explaining the phi-accrual crossing-threshold math (1s + 20s + 3*0.1s = 21.3s of detection). v1.5's failure-detector defaults (heartbeat-interval=1s, min-std-deviation=100ms) and split-brain-resolver stable-after (10s) already match dev's, so the same 3s-was-too-tight problem applies here: a churn round abruptly tearing down an ActorSystem can starve the surviving node's own heartbeat sender for longer than a 3s pause tolerates. * akka.test.single-expect-default = 10s and akka.test.timefactor = 3, ported from #8372 (the first commit of the same dev stack), which the earlier port of this commit had missed. Every Within/WithinAsync bound in this file is dilated by timefactor - including RemoveOneAsync's removal budget (TimeSpan.FromSeconds(25) + ConvergenceWithin(3s, NbrUsedRoles - 1), which #8543 never widened) - so raising acceptable-heartbeat-pause to 20s without also porting timefactor left that budget structurally unable to cover the ~34.3s abrupt-removal path this spec's own ChurnMemberRemovalWithin() arithmetic predicts at the node counts the 7-node CI phase sequence reaches. Without timefactor, RemoveOneAsync's ceiling at NbrUsedRoles=3 is a flat 31s; with it, 93s. Confirmed against three consecutive 7-node runs: the abrupt-removal phase measured 32.9-33.4s in each, and the AwaitAssert measured a 30.98s give-up in each, both matching the undilated 31s ceiling to within noise. All three runs pass with timefactor ported. * ChurnMemberRemovalWithin(), ported byte-for-byte from dev (it only calls Cluster.Settings.FailureDetectorConfig/HeartbeatInterval/GossipInterval, all present on v1.5). Computes how long an abruptly-terminated churn member takes to actually leave the ring: detection + stable-after + a leader-gossip margin = 34.3s at this spec's config. * ExerciseJoinRemoveAsync's loopDuration now includes ChurnMemberRemovalWithin() so each round's Within budget covers the full removal path, not just the new join. The abrupt-shutdown behavior itself (ShutdownAsync, no cluster Leave) was already in place from #8499 - this keeps exactly what the phase proves (abrupt loss), it only fixes the budget around it. * Async TestKit migration of the call sites #8543 touches: the two RunOn sends inside the churn Loop become RunOnAsync, and ClusterResultAggregator (a one-shot, non-retried Identify/ExpectMsg lookup) gains a ClusterResultAggregatorAsync sibling (fresh-probe-per-attempt, retried over 30s) ported from dev, since a lone lost reply under this phase's deliberate churn must not be fatal. Repointed CreateResultAggregatorAsync, AwaitClusterResultAsync, and the async ReportResult<T> overload - the three call chains ExerciseJoinRemoveAsync depends on - at the new async lookup. Every remaining call site in the file passes an async lambda with an explicit return statement, so all of them already bound to the async ReportResult<T>(Func<Task<T>>) overload; the sync ClusterResultAggregator() and the sync ReportResult<T>(Func<T>) overload it served had no callers left after that repointing and are removed here. (Not ported: dev's extra "result-aggregator-identified" barrier in CreateResultAggregatorAsync, a related but separate hardening against a PartitionSeveral race - open item below.) * RemoveOneAsync's watchee lookup, hand-ported from dev (#8372/#8429): a fresh CreateTestProbe() per attempt instead of the shared IdentifyProbe (the shared probe kept a timed-out attempt's late ActorIdentity queued, so the next attempt consumed that stale reply, and a reply resolved before the watchee existed carries a null Subject that the retry could never recover from); an explicit identity.Subject.Should().NotBeNull() guard before WatchAsync; WatchAsync bounded with WaitAsync(Dilated(3s)) instead of inheriting the outer Within's RemainingOrDefault, since a hung/slow watch Ask would otherwise burn the whole retry budget in a single attempt; and an explicit AwaitAssertAsync(10s, 1.25s) bound instead of an unbounded retry loop. This is the exact method that was failing in the runs above, so it is ported alongside the budget fix rather than left as a separate follow-up. * StressSpecConfig node-count env override already existed on v1.5 (MNTR_STRESSSPEC_NODECOUNT, default 13). What it lacked was BuildConfig's shrink arithmetic: below the reference count, phase sizes must shrink or Settings' constructor throws. v1.5's reference config is heavier than dev's - every joining phase defaults to 2 nodes here, not 1 - so reaching the same practical floor of 7 requires shrinking both the joining side (halve every 2 back to 1) and the leaving/shutdown side (drop the "-large" one-by-one phases, halve the simultaneous counts), the same technique dev's config uses on one more group of phases. Derived and documented in StressSpecConfigSpec.cs (ported alongside, adapted from dev's 10-node version to v1.5's 13-node/2-per-phase defaults): 7 is confirmed the practical floor (6 throws because the joining phases alone need 7 regardless of how far leaving/shutdown shrinks). Also fixed while building this: MultiNodeTestRunner.cs's Process.Kill(bool) call from the earlier #8515 hand-merge doesn't compile against netstandard2.0 - see the preceding commit. Verified: dotnet build src/core/Akka.Cluster.Tests.MultiNode -warnaserror clean; StressSpecConfigSpec 9/9; StressSpec run three times at MNTR_STRESSSPEC_NODECOUNT=7, all three passed (see the PR body for full timings). Open items for the maintainer: - dev's extra CreateResultAggregatorAsync barrier (guards a PartitionSeveral aggregator- identification race) was not ported; the retrying ClusterResultAggregatorAsync lookup narrows that race but does not close it the way the extra barrier does.
…ressSpec node-count doc Three comment-only corrections, none change behavior: * ClusterSpec.cs (from the #8579 pick): the comment on the MemberUp wait described dev's Cluster.ClusterCore re-targeting from a supervisor to the core daemon when the ref is published - a mechanism that no longer matches this branch's own code now that #8359 and #8580 are in. ClusterCore always targets /system/cluster for the life of the extension, so two commands from the same sender are already delivered in order without this wait; the MemberUp wait is still the stronger barrier, since it proves the join was processed, not merely enqueued, before Leave is sent. Reworded to say so. * DistributedPubSubRestartSpec.cs (from the earlier Artery-removal tailoring commit): the reworded comment claimed a plain Tell to a quarantined peer is dropped at the transport layer. On v1.5's classic remoting it is not: EndpointManager's Quarantined case creates a brand-new writing endpoint for any Send that reaches it. It is the separate Gated policy that dead-letters a Send while its release deadline is unexpired. Reworded to describe what EndpointManager.cs actually does, and to explain ActorSelection.Tell's real advantage here (re-resolving by path on every send reaches whichever incarnation is live, rather than a stale cached ref). * StressSpec.cs BuildConfig doc comment: said 13 was "also the smallest count that fits every phase" and then derived >= 11 two sentences later - contradicting itself. Reworded to say what the arithmetic says: 11 is the true minimum at full phase size; 13 is v1.5's original default, kept as the boundary below which this method starts shrinking phases. Verified: dotnet build on Akka.Cluster.Tests, Akka.Cluster.Tools.Tests.MultiNode, and Akka.Cluster.Tests.MultiNode -warnaserror, all clean.
…pec at 7 nodes Mirrors dev's #8553. build-system/azure-pipeline.mntr-template.yaml gained no way to pass extra environment variables into the test-execution step, so build-system/pr-validation.yaml's single MNTR lane (net_mntr_windows) could not set MNTR_STRESSSPEC_NODECOUNT even though StressSpec.cs already reads that variable. Added an `env` parameter (defaulting to `{}`, matching dev's template) to the MNTR template and wired it into the `${{ parameters.command }}` step, then set `MNTR_STRESSSPEC_NODECOUNT: "7"` on the v1.5 MNTR lane so StressSpec runs at 7 nodes instead of its 13-node default on the 2-vCPU hosted agents. 7 is validated as a safe floor by the preceding StressSpec commit's StressSpecConfigSpec tests (StressSpecConfig.BuildConfig shrinks every phase to fit at exactly 7 nodes; 6 still throws).
#8580 is now fully in this branch (both the ClusterCore routing fix and the LeaveSelf re-send fix cherry-picked earlier), so this commit is no longer a carve-out - it only adds the one ClusterSpec fact dev's own #8580 pick didn't need, because dev's ported fact (A_cluster_must_process_a_Leave_issued_immediately_after_Join_from_the_same_thread) exercises the ClusterCore-routing race, not the LeaveSelf re-send: it calls Leave() exactly once, and that first call behaves identically before and after the re-send fix. Adds A_cluster_must_resend_Leave_on_a_later_LeaveAsync_call_after_an_earlier_send_was_lost, which does discriminate: LeaveAsync() before any Join reaches ClusterCoreDaemon's Uninitialized behavior (via the pre-Init stash in ClusterDaemon, which buffers traffic until its own Init completes, then forwards it down), which has no case for ClusterUserAction.Leave, so the command is unhandled and dropped; then Join(self), LeaderActions(), await MemberUp via a subscription, then LeaveAsync() again and assert the returned task completes. Verified by temporarily reverting LeaveSelf() to send Leave only once (its pre-fix shape): this fact times out after 10s on that code and passes in under a second on the fix. Review follow-up (Opus review, findings 8 and 9): the fact now asserts its own premise instead of assuming it - a probe subscribes to UnhandledMessage on the EventStream before the first LeaveAsync() and asserts one arrives whose Message is a ClusterUserAction.Leave addressed to .../system/cluster/core/daemon, so a future change that gives Uninitialized a Leave case fails this fact here instead of leaving it silently non-discriminating. The final WaitAsync bound is now Dilated(10s) instead of a raw TimeSpan, consistent with the sibling facts' RemainingOrDefault/ dilated ExpectMsgAsync usage (equivalent to the old value today, since this spec sets no akka.test.timefactor). Verified: `dotnet build src/core/Akka.Cluster.Tests -warnaserror` clean; `dotnet test src/core/Akka.Cluster.Tests --filter ClusterSpec` run 3 times - all facts passed every time, including both Leave facts. Re-confirmed the fact still discriminates after this change: with LeaveSelf() temporarily reverted to its pre-fix one-shot send, the fact fails with a TimeoutException (the dead-lettered/unhandled first Leave is visible in the log); restoring Cluster.cs returns it to passing.
Adds a `1.5.72 TBD` section covering every user-visible change in this PR, including the non-blocking Cluster startup fix and a migration note for #8574's write-timeout key change. Review follow-up (Opus review, findings 2, 3, 5, 10, 11): - The #8359 bullet now states outright that the cluster's internal actor tree (/system/cluster/core, /system/cluster/core/daemon, /system/cluster/heartbeatReceiver) is created after Cluster.Get() returns rather than before it - the one sharp edge of this change, and previously the only thing this note left unsaid - and says code resolving those paths right after Cluster.Get() should await JoinAsync/RegisterOnMemberUp or a cluster event instead. - Says outright that akka.actor.creation-timeout is no longer consulted by Akka.Cluster at all (stronger and more accurate than "no longer blocks on"), while noting Akka.Cluster.Sharding still uses it to bound its own start-up asks, so nobody reads this as license to revert that setting. - Replaced "a small dedicated dispatcher pool" with wording that also covers a starved default pool, since the deadlock reproduces on either. - The #8580 bullet no longer says "during startup" - the routing through /system/cluster is permanent for the extension's lifetime (two extra local mailbox hops on the low-traffic command path), and now says a `Leave`/`LeaveAsync()` call issued after the node has already shut down produces a dead-lettered ClusterUserAction.Leave, relevant to dead-letter alerting. - Adds a Testing bullet for #8398 (De-flake Bugfix5962Spec), the dev de-flake this branch had been missing for the spec #8359 makes newly racy.
The #8359 cherry-pick brought in ClusterStartupFuzzSpec.cs, which uses Task.IsCompletedSuccessfully at two call sites, and the #8398 cherry-pick (5f225ed, applied earlier in this branch as the Opus-review follow-up for finding 1) uses the non-generic `new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously)` / `TrySetResult()` in Bugfix5962Spec.cs. On dev, Akka.Cluster.Tests no longer targets net48, so both compiled fine there. On v1.5, Akka.Cluster.Tests still targets net48 alongside net10.0: `Task.IsCompletedSuccessfully` does not exist on .NET Framework 4.8 (CI build 131554 failed with CS1061 on every net48 lane), and the non-generic `TaskCompletionSource` does not exist before .NET 5 (CS0305). - ClusterStartupFuzzSpec.cs: switched both call sites to `joinTask.Status == TaskStatus.RanToCompletion`, the exact equivalent available on all target frameworks (the `using System.Threading.Tasks;` needed for `TaskStatus` was already present). - Bugfix5962Spec.cs: switched to `TaskCompletionSource<Done>` / `TrySetResult(Done.Instance)`, the same pattern already used by ClusterDaemon's own `_clusterPromise`/`_selfExiting` fields - available on every target framework this project builds for. No behavior change: `Done` carries no data, so the task's completion is still the only thing the spec observes. Verified: - `dotnet build src/core/Akka.Cluster.Tests -c Release -f net48 -warnaserror` and `-f net10.0 -warnaserror`: both 0 warnings/errors. - `dotnet test -f net10.0 --filter FullyQualifiedName~ClusterStartupFuzzSpec` passes both tests: aggressive run completed=72, kills(delayed=0, immediate=0), pass=72 (reachedUp=72); moderate run completed=60, kills(delayed=0, immediate=0), pass=60 (reachedUp=60). - `dotnet test -f net10.0 --filter FullyQualifiedName~Bugfix5962Spec`: passes.
Aaronontheweb
force-pushed
the
backport/dev-2026-09-to-v1.5
branch
from
September 12, 2026 16:18
9152ed7 to
8ab2713
Compare
Aaronontheweb
added a commit
that referenced
this pull request
Sep 12, 2026
…s the new oldest, so a lost HandOverInProgress is named instead of timed out In PR #8588 build 131555 (Windows unit lane), this spec failed with only "Timeout 00:00:01 while waiting for a message of type System.String" from the final AwaitProxyReplyAsync(_sys3, proxy3, "hello3", 5s). The real cause was upstream: sys2 wrote HandOverInProgress, HandOverDone and ExitingConfirmed to the socket before closing it during CoordinatedShutdown, but on Windows a queued inbound frame from sys3 turned the close into a TCP RST, and the RST made sys3's stack discard the bytes it had not yet read (see #8589). sys3 never saw the hand-over confirmation, never cancelled its retry timer, and fell back to the removal-margin / hand-over-retries path - which does recover the singleton, but only after 20s or more, well outside the 5s budget the test is actually checking. Wrap both hand-overs (sys1 -> sys2 and sys2 -> sys3) in an EventFilter on "Hand-over in progress at" (ClusterSingletonManager.cs, logged by the new oldest when it receives HandOverToMe's confirmation) so a lost confirmation fails the fact at the point the hand-over actually broke, with the filter's own message, instead of surfacing as an unrelated-looking timeout later. No timeout, window, or config value is changed. The 5s AwaitProxyReplyAsync budget is deliberately left alone: the loss path needs roughly 21s or more to resolve (6s failure-detector + 20s removal margin, or ~27s via hand-over-retry exhaustion and a manager crash-restart), and widening the budget to cover it would only make the test pass on a path where the hand-over failed and the manager had to crash to recover - defeating the point of a spec named after the hand-over. Verification: - dotnet build src/contrib/cluster/Akka.Cluster.Tools.Tests -c Release -f net10.0 -warnaserror -p:SuppressTfmSupportBuildWarnings=true: clean (project targets only net10.0 on dev) - Spec run 10x: dotnet test ... --filter FullyQualifiedName~ClusterSingletonRestartSpec: 10/10 passed, ~6-7s each - Negative check: temporarily changed the sys3 filter's start: text to "Hand-over in progress at NOWHERE" (a string that never logs); the fact failed with "Timeout (00:00:03) while waiting for messages. Only received 0/1 messages that matched filter [Info when Message starts with 'Hand-over in progress at NOWHERE']" - confirming the filter discriminates. Restored the real text; dotnet build and a follow-up 10x run both green again, and git diff shows only the intended change. - dotnet format --verify-no-changes on the touched file flags one pre-existing, unrelated whitespace issue ("if(_sys3 != null)" at the original line 166, in AfterAll) that predates this change; left untouched to keep the diff focused.
…0 unit-test host hangs The Linux ".NET Unit Tests" lane hangs inside src/core/Akka.Tests now and then, with no test named and no artifact to read: Incrementalist (.incrementalist/testsOnly.json, timeoutMinutes: 20) cancels the dotnet test command at its 20-minute per-project cap before vstest can write a trx. Seen on v1.5 build 130811 (2026-08-26, before this PR), and again on this PR's own builds 131545 and 131555 - each time going silent about 75s into the project with nothing pointing at the stuck test. Append vstest's blame-hang flags (--blame-hang --blame-hang-timeout 15m --blame-hang-dump-type mini) to the two net10.0 unit-test commands in build-system/pr-validation.yaml (the Windows and Linux ".NET Unit Tests" lanes, both using .incrementalist/testsOnly.json). When a test goes quiet for 15 minutes, vstest writes a Sequence_<guid>.xml naming the tests that started and did not finish, plus a minidump of the test host and its children, into TestResults - which the lane already publishes as a build artifact - then kills the host so the lane fails with the test named instead of just timing out silently. 15 minutes sits below Incrementalist's 20-minute cap, so vstest's own hang timer fires and writes its artifacts before Incrementalist kills the process tree; no single test in these projects legitimately runs anywhere near 15 minutes (the whole Akka.Tests project takes about 6-7 minutes on CI). Left out: - The net48 lane (.incrementalist/testsOnlyNetFx.json): vstest's hang-dump collector needs procdump on the agent for a full-framework test host, which these agents don't have. - The multi-node lane (.incrementalist/mutliNodeOnly.json): its node processes are child processes the MNTR runner already manages and tears down itself. Verified locally on this SDK (dotnet 10.0.103, Linux): built src/core/Akka.Tests/Akka.Tests.csproj for net10.0, then ran Akka.Tests.Actor.TimerSpec.Must_replace_timer (a ~3s test) with a deliberately tiny --blame-hang-timeout to confirm the hang path fires: - 1s timeout: fired during test-host startup/discovery, before the test itself began running - produced two hang-dump .dmp files (test host + child process) but no Sequence file, since no test was yet in progress when the collector ran. - 2s timeout: fired mid-test - produced Sequence_<guid>.xml with <Test Name="Akka.Tests.Actor.TimerSpec.Must_replace_timer" Completed="False" />, plus the same two .dmp files, confirming the mechanism names the actual stuck test once one is running. - 15m timeout: passed normally in ~3s, no Sequence file or dumps produced - confirms the flags are inert on a healthy run.
Member
Author
|
/azp run |
|
Azure Pipelines successfully started running 1 pipeline(s). |
Aaronontheweb
added a commit
that referenced
this pull request
Sep 21, 2026
…s the new oldest, so a lost HandOverInProgress is named instead of timed out (#8590) In PR #8588 build 131555 (Windows unit lane), this spec failed with only "Timeout 00:00:01 while waiting for a message of type System.String" from the final AwaitProxyReplyAsync(_sys3, proxy3, "hello3", 5s). The real cause was upstream: sys2 wrote HandOverInProgress, HandOverDone and ExitingConfirmed to the socket before closing it during CoordinatedShutdown, but on Windows a queued inbound frame from sys3 turned the close into a TCP RST, and the RST made sys3's stack discard the bytes it had not yet read (see #8589). sys3 never saw the hand-over confirmation, never cancelled its retry timer, and fell back to the removal-margin / hand-over-retries path - which does recover the singleton, but only after 20s or more, well outside the 5s budget the test is actually checking. Wrap both hand-overs (sys1 -> sys2 and sys2 -> sys3) in an EventFilter on "Hand-over in progress at" (ClusterSingletonManager.cs, logged by the new oldest when it receives HandOverToMe's confirmation) so a lost confirmation fails the fact at the point the hand-over actually broke, with the filter's own message, instead of surfacing as an unrelated-looking timeout later. No timeout, window, or config value is changed. The 5s AwaitProxyReplyAsync budget is deliberately left alone: the loss path needs roughly 21s or more to resolve (6s failure-detector + 20s removal margin, or ~27s via hand-over-retry exhaustion and a manager crash-restart), and widening the budget to cover it would only make the test pass on a path where the hand-over failed and the manager had to crash to recover - defeating the point of a spec named after the hand-over. Verification: - dotnet build src/contrib/cluster/Akka.Cluster.Tools.Tests -c Release -f net10.0 -warnaserror -p:SuppressTfmSupportBuildWarnings=true: clean (project targets only net10.0 on dev) - Spec run 10x: dotnet test ... --filter FullyQualifiedName~ClusterSingletonRestartSpec: 10/10 passed, ~6-7s each - Negative check: temporarily changed the sys3 filter's start: text to "Hand-over in progress at NOWHERE" (a string that never logs); the fact failed with "Timeout (00:00:03) while waiting for messages. Only received 0/1 messages that matched filter [Info when Message starts with 'Hand-over in progress at NOWHERE']" - confirming the filter discriminates. Restored the real text; dotnet build and a follow-up 10x run both green again, and git diff shows only the intended change. - dotnet format --verify-no-changes on the touched file flags one pre-existing, unrelated whitespace issue ("if(_sys3 != null)" at the original line 166, in AfterAll) that predates this change; left untouched to keep the diff focused.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Merge with a merge commit, not a squash, so each cherry-pick keeps its
-xreference to the dev commit.Summary
Backports the de-flakes, product fixes, TestKit changes, and a Cluster-startup fix merged into
devsince 2026-09-01 (plus four pre-window prerequisites) tov1.5. 33 commits are faithfulcherry-picks (
git cherry-pick -x) in dependency order; 11 further commits adapt or hand-port thepieces that need v1.5-specific work, each kept separate so it can be reviewed or dropped on its
own.
Cherry-picked from dev
TestKitBase.ShutdownAsync- non-blocking system shutdown, used by later picksWithinTcpUnbind now completes at once instead of always waiting out the subscription timeoutupdating-state-timeout, matching docsReplicator.IsKnownNodenow trusts members first seen Leaving/Exiting/Downedakka.remote, where the transport reads itInitializeAsyncinstead of blocking a pool threadDisposeAsyncno longer leaks the ActorSystemWithinAsyncReceiveAsyncso a deadline surfaces asTimeoutException/system/clusterfor the extension's lifetime;LeaveAsyncre-sends itsLeave(fixes an ordering regression #8359 introduces)ResolveOneraced the now-asynchronous cluster-extension init); found by the downstream audit and the Opus reviewNote on #8515: the conflicting hunk in
MultiNodeTestRunner.csbundled #8515's own changewith
DefaultNodeExitTimeout, which turned out to be the payload of dev's earlier #8431 (a20-minute backstop that kills a hung node process) - v1.5's own prior #8431 backport had dropped
this file's hunk, so the field did not exist on v1.5 before this pick. Taking dev's whole file for
the merge, the only way to give the field a real consumer, lands the #8431 node-exit backstop on
v1.5 for the first time alongside #8515's port-sentinel change. Details in that commit's message.
v1.5 tailoring commits
These are not cherry-picks - each adapts a picked change, or ports one dev change by hand,
to v1.5's own code shape. Listed in commit order:
BREAKING_CHANGES_V1.6.md- the TestKit.Xunit: implement the async dispose chain so a derived DisposeAsync no longer leaks the ActorSystem; de-flake StreamRefsSpec #8545 cherry-pick touches this file on dev; v1.5 has no v1.6 release ledger, so it is removed again after the picks land.DistributedPubSubRestartSpec.csreferenced Artery, which does not exist on v1.5; removed/reworded, no behavior change.ShutdownAsync- the API-approval snapshot needed re-accepting for the two new methods Add TestKitBase.ShutdownAsync and adopt it in StressSpec churn teardown #8499 adds; the net48 counterpart was derived by reasoning rather than executed (see Verification).Process.Kill(bool)call surfaced by the MNTR: eliminate the conductor port race - bind port 0 and propagate via stdout sentinel #8515 hand-merge - netstandard2.0 lacks the process-tree-kill overload dev's file uses; switched to the parameterless overload, matching an existing convention elsewhere in the codebase.akka.test.timefactor/single-expect-default(from De-flake StressSpec: retry the aggregator Identify, fix blocking WatchAsync, add timefactor #8372, the first commit of the same stack - without these two lines everyWithinbound in the file stays undilated while the heartbeat-pause widening triples the detection time it has to cover, see "StressSpec root cause and fix" below), theChurnMemberRemovalWithinbudget term, the async call-site migration, the hardenedRemoveOneAsyncwatchee lookup (fresh probe per attempt, null-Subject guard, boundedWatchAsync), and adaptedStressSpecConfig's node-count floor arithmetic to v1.5's heavier default config (13 nodes / 2-per-phase vs dev's 10/1) - still reaches the same practical floor of 7 nodes. Removed the syncClusterResultAggregator()/ReportResult<T>(Func<T>)overload pair, which had no callers left once the async chains were repointed. A newStressSpecConfigSpec.csdocuments and asserts the node-count derivation.#8579pick'sClusterSpec.cscomment on theMemberUpwait described dev'sClusterCorere-targeting from a supervisor to the core daemon mid-startup, a mechanism this branch no longer has once Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) #8359/Cluster: route user commands through /system/cluster for the extension's lifetime so their order survives startup; LeaveAsync re-sends its Leave #8580 land (ClusterCorealways targets/system/cluster) - reworded to say the wait is still the stronger barrier because it proves the join was processed, not merely enqueued; a comment introduced by tailoring commit 2 wrongly claimed a Tell to a quarantined peer is dropped at the transport layer (on v1.5's classic remoting it is not - seeEndpointManager.cs); and theStressSpec.csBuildConfigdoc comment contradicted its own arithmetic. No behavior change.envparameter dev's CI template has, and setsMNTR_STRESSSPEC_NODECOUNT=7on v1.5's one MNTR lane.ClusterSpecfact forLeaveSelf's re-send - now that Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) #8359 and Cluster: route user commands through /system/cluster for the extension's lifetime so their order survives startup; LeaveAsync re-sends its Leave #8580 are both cherry-picked in full, dev's own ported fact (A_cluster_must_process_a_Leave_issued_immediately_after_Join_from_the_same_thread) already covers theClusterCore-routing half; it callsLeave()exactly once and behaves identically before and after the re-send fix, so it doesn't discriminate the re-send itself. This commit adds one more fact,A_cluster_must_resend_Leave_on_a_later_LeaveAsync_call_after_an_earlier_send_was_lost, that does (see "Testing" below).1.5.72 TBDsection covering every user-visible change in this PR, including the non-blocking Cluster startup fix and a migration note for Cluster.Sharding: arm the remember-entities write timeout with updating-state-timeout, as the config documents #8574's write-timeout key change.ClusterStartupFuzzSpecusedTask.IsCompletedSuccessfully, and the De-flake Bugfix5962Spec.SBR_Should_work_with_channel_executor #8398 cherry-pick'sBugfix5962Specused the non-genericTaskCompletionSource/TrySetResult()- neither exists on .NET Framework 4.8 (v1.5'sAkka.Cluster.Testsstill targets net48, unlike dev); switched toStatus == TaskStatus.RanToCompletionandTaskCompletionSource<Done>/TrySetResult(Done.Instance)respectively.--blame-hang --blame-hang-timeout 15m --blame-hang-dump-type minito the Windows and Linux.NET Unit Testscommands inbuild-system/pr-validation.yaml(thetestsOnly.json/net10.0 lanes only, not the net48 lane or the multi-node lane). See "CI evidence for a pre-existing Linux hang" below for why.Not backported
WatchAsyncfixes are superseded by tailoring commit 5's own port of the same mechanism; its two config lines (akka.test.timefactor = 3,akka.test.single-expect-default = 10s) are ported directly into that same commitCircuitBreakerSpecde-flake plus a retryingReportResultaggregator lookup for StressSpec) - the StressSpec half is superseded by tailoring commit 5; theCircuitBreakerSpechalf is out of scope for this PRWithin/AwaitAssertAsyncbounds in StressSpec andClusterShardingQueriesSpec) - the StressSpec half is superseded by tailoring commit 5; theClusterShardingQueriesSpechalf is out of scope for this PRsrc/core/Akka.Remote/Arterydoes not exist on v1.5Akka.Serialization.V2generator stack and docs; that project does not exist on v1.5Verification
Environment note: this verification environment's SDK pulls in a transitive package that refuses to build the
net6.0library target under strict warnings; reproduced identically on a pristine, unmodifiedv1.5checkout, so it is a local tooling artifact, not a backport defect(
-p:SuppressTfmSupportBuildWarnings=trueused to work around it for local verification only).Directory.Build.propsalready setsTreatWarningsAsErrors=trueunconditionally, so every testrun below already enforces warnings-as-errors without a separate
-warnaserrorbuild step.dotnet build -c Release- succeeded, 0 errors.dotnet test src/core/Akka.API.Tests- 18/18 passed, no approval diff.-warnaserrorbuilds: Akka.Cluster.Tests, Akka.Cluster.Tests.MultiNode,Akka.Cluster.Tools.Tests.MultiNode, Akka.API.Tests - all clean, 0/0.
Akka.Cluster.Tests, full project: 368/368 passed.ClusterSpec, run 3 times: 22/22 passed every time, including both Leave facts.ClusterStartupFuzzSpec, run once: both profiles passed, 0 kills across 132 simulatedstartup iterations (72 aggressive + 60 moderate).
Akka.Cluster.Testsbuilt
-warnaserroron bothnet48andnet10.0(0 warnings/errors each), the two targetframeworks that project ships on v1.5;
ClusterStartupFuzzSpecandBugfix5962Specre-run onnet10.0afterward -ClusterStartupFuzzSpec0 kills across 132 iterations,Bugfix5962Spec20/20 consecutive runs.
Tests (Windows)" job succeeded on this PR's own
refs/pull/8588/mergeref, running the wholeAkka.Cluster.Testsproject - includingClusterStartupFuzzSpecandBugfix5962Spec- onWindows/net48:
Passed! - Failed: 0, Passed: 367, Skipped: 1, Total: 368, Duration: 6 m 37 s - Akka.Cluster.Tests.exe (net48)(the one skip is a pre-existing, unrelatedDowningProviderSpecfact). The fuzz spec's concurrency design is proven on the exact lane thisPR adds it to, not just on net10.0.
Akka.Testsintermittently, unrelated to this PR's own changes - seen on v1.5 build 130811(2026-08-26, before this PR existed) and again on this PR's own builds 131545 and 131555. Each
time, Incrementalist (
.incrementalist/testsOnly.json,timeoutMinutes: 20) cancels thedotnet testprocess at its 20-minute per-project cap before vstest can write a trx or name atest, so the hung test itself was never identified. Tailoring commit 11 adds vstest's
--blame-hangflags (15-minute per-test hang timeout, comfortably under Incrementalist's20-minute cap and well above this project's ~6-7 minute normal run time) to both net10.0
unit-test lanes, so the next occurrence names the stuck test in a
Sequence_*.xmland leaves aprocess minidump in the
TestResultsartifact the lane already publishes.must-fix and several nits, all applied here - cherry-picked De-flake Bugfix5962Spec.SBR_Should_work_with_channel_executor #8398 (the de-flake this branch was
missing for the one spec Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) #8359 makes newly racy, with the mandatory net48 adaptation folded into
tailoring commit 10 above), reworded the RELEASE_NOTES Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) #8359/Cluster: route user commands through /system/cluster for the extension's lifetime so their order survives startup; LeaveAsync re-sends its Leave #8580 entries and added a Testing
bullet for De-flake Bugfix5962Spec.SBR_Should_work_with_channel_executor #8398, and had the
ClusterSpecfact from tailoring commit 8 assert its own premise(an
UnhandledMessageprobe confirming the pre-JoinLeaveis actually dropped) and use adilated final wait instead of a raw one.
(
--filter FullyQualifiedName~<Spec>), not as a whole-project run, acrossAkka.Cluster.Sharding.Tests, Akka.Cluster.Tools.Tests, Akka.DistributedData.Tests,
Akka.DependencyInjection.Tests, Akka.Cluster.Tests, Akka.Remote.Tests, Akka.Remote.TestKit.Tests,
Akka.Streams.Tests, Akka.Tests, and Akka.Cluster.Tests.MultiNode (
StressSpecConfigSpec) - allpassed; a handful of pre-existing skips unrelated to this backport. The full
Akka.TestKit.Xunit.Tests and Akka.TestKit.Tests projects were run in full, not filtered.
(v1.5 has no Artery lane): RemoteDeliverySpec, RemoteNodeDeathWatchSpec,
RemoteNodeRestartDeathWatchSpec, ClusterDeathWatchSpec, NodeChurnSpec,
ClusterShardingLeavingSpec, ClusterShardingRegistrationCoordinatedShutdownSpec,
ClusterShardingSpec, DistributedPubSubRestartSpec, ClusterSingletonManagerLeave2Spec,
ReplicatorChaosSpec - all passed.
MNTR_STRESSSPEC_NODECOUNT=7, matching the new CI setting): faileddeterministically before the churn-phase fix below; passes after it (see the next section
for the root cause, timings, and this PR's Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) #8359/Cluster: route user commands through /system/cluster for the extension's lifetime so their order survives startup; LeaveAsync re-sends its Leave #8580 addendum).
StressSpec root cause and fix (see tailoring commit 5)
Tailoring commit 5 ports #8543's
acceptable-heartbeat-pausewidening (3s -> 20s) intoStressSpec, together with #8372's two companion config lines,
akka.test.timefactor = 3andakka.test.single-expect-default = 10s(#8372 is the first commit of the same 5-deep stack#8543 sits on). Both lines matter: every
Within/WithinAsyncbound in this spec is dilated byakka.test.timefactor- includingRemoveOneAsync's removal budget,TimeSpan.FromSeconds(25) + ConvergenceWithin(TimeSpan.FromSeconds(3), NbrUsedRoles - 1), which#8543never widened directly. Without the factor, that budget is a flat 31s atNbrUsedRoles = 3- the node count the 7-node phase sequence reaches duringMustShutdownNodesOneByOneFromSmallClusterAsync.ChurnMemberRemovalWithin()'s own arithmeticputs the actual abrupt-removal detection time at 34.3s at this spec's config, which structurally
cannot fit inside a 31s ceiling. With the factor restored, the same budget is 93s. Also ported
dev's hardened
RemoveOneAsyncwatchee lookup (fresh probe per attempt, null-Subject guard,bounded
WatchAsync), which is the method that was failing.Re-ran
StressSpecthree times atMNTR_STRESSSPEC_NODECOUNT=7after the fix, in full each time:shutdown 1 in 4 nodes clustershutdown one from 3 nodes cluster(the previously-failingRemoveOneAsync)shutdown one from 3 nodes clusteris the exact phase that used to fail atMustShutdownNodesOneByOneFromSmallClusterAsync->RemoveOneAsync. AtNbrUsedRoles = 3,RemoveOneAsync's budget is now(25 + 3 * 1 * 2) * 3 = 93s(dilated by the restoredtimefactor), comfortably covering the measured 32.6-33.5s cost - previously that same phase hada flat, undilated 31s ceiling against this same ~33s cost, which is why it failed 3/3 times
before this fix.
partition 2 in 7 nodes cluster(33.8-34.4s across runs 2-3) independentlyconfirms
ChurnMemberRemovalWithin()'s predicted 34.3s detection time.#8359 / #8580: a second StressSpec failure mode, found in this PR's own CI
After this PR was first opened, CI build 131547 (re-run on this PR's previous head, after the
churn-phase fix above had already landed) failed
StressSpecon the Windows MNTR lane with adifferent signature than the one above:
JoinSeveralAsynctimed out waiting for node-6 to join("Expected: 6, Actual: 5"). Node-6's own log showed "Started up successfully" at 04:06:01.605
(the
Clusterconstructor's ask answered in ~140 ms that time), then two "Failed to startupCluster ... AskTimeoutException: Timeout after 20.00 seconds" at 04:06:21.57 - one from the
cluster event-bus listener's
PreStart, one from the SBR downing provider'sPreStart- followedby
Shutdown().StressSpecpins the default dispatcher's fork-join pool toparallelism-min=2/factor=1(unchanged by this PR, and identical todev's own setting), so onthe 2-core CI agent both of those dedicated threads parked on the
ClusterCoregetter's re-issuedblocking
GetClusterCoreRef().Resultasks (the constructor's own ask had already been answered), andthe daemon that had to answer them could never get scheduled:
a hard deadlock, resolved only by the 20s ask timeout. Raising
akka.actor.creation-timeoutwouldonly delay that timeout, not fix the deadlock. This is exactly the failure mode dev's #8359 fixes
by removing the blocking ask; #8580 rides along in the same commit set because #8359 itself
introduces a command-ordering regression that #8580 fixes: #8359's non-blocking
ClusterCoregetter switched from
/system/clusterto a direct core-daemon reference the instant that ref waspublished mid-startup, so a Join and a Leave issued back-to-back by the same caller could arrive
at the core daemon out of order.
Local verification before adding these two commits to this PR (full logs in
8359-v15-verification.md, not included in the PR itself):ClusterStartupFuzzSpec(randomized concurrent-startup fuzzing under pinned 1-2-threaddispatcher pools, Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) #8359's own regression test): 0 kills across 396 simulated iterations
with the fix; 33 immediate kills across 42 completed iterations without it, reproducing the
exact "Failed to startup Cluster" /
AskTimeoutExceptionsignature from build 131547.StressSpecatMNTR_STRESSSPEC_NODECOUNT=7, pinned to two cores (taskset -c 0,1, to matchthe CI agent's core count): 3/3 passed with the fix.
This session's own verification, after rebuilding this PR's cherry-pick/tailoring stack to
include #8359 and #8580 (see "Verification" above for the full list):
Akka.Cluster.Testsfullproject 368/368;
ClusterSpec22/22 across three separate runs, including both Leave facts;ClusterStartupFuzzSpec0 kills across 132 iterations; API approvals 18/18 with no diff (thechange is entirely internal); and
StressSpecatMNTR_STRESSSPEC_NODECOUNT=7, pinned to twocores, 16/16 passed in 3m 14s.
Testing
ClusterSpecrun 8 times total across this PR's history (22/22 passing every time), includingtwo
Leave-ordering facts: dev's own#8580fact,A_cluster_must_process_a_Leave_issued_immediately_after_Join_from_the_same_thread, portedunchanged (it calls
Leave()exactly once and behaves identically before and after theLeaveSelfre-send fix, so it's a Join-then-Leave smoke test, not proof of the re-send); andtailoring commit 8's own fact,
A_cluster_must_resend_Leave_on_a_later_LeaveAsync_call_after_an_earlier_send_was_lost, whichdoes discriminate the re-send. That fact issues
LeaveAsync()before any Join (stashed byClusterDaemonuntil its ownInitcompletes, then unhandled and dropped byClusterCoreDaemon'sUninitializedbehavior, which has no case forClusterUserAction.Leave),joins, waits for
MemberUp, then callsLeaveAsync()again and asserts it completes. Confirmedit actually discriminates: reverting
LeaveSelf()to the pre-fix one-shot version and re-runningjust this fact fails it with a
TimeoutExceptionafter 10s (the dead-lettered firstLeaveislogged:
Message [Leave] from [.../testActor1] ... was unhandled. [1] dead letters encountered);restoring the file returns it to passing in under a second.
StreamRefsSpecrun twice (16/16 passing each time, one pre-existing skip unrelated to thisbackport).
ClusterStartupFuzzSpec(Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) #8359's own regression test) run once after rebuilding this PR'sstack: both profiles passed, 0 kills across 132 simulated startup iterations.
Open items for the maintainer
Net.verified.txt(net48/netstandard2.0) API-approval file forApproveTestKitwasupdated by derivation, not by running the net48 test host (this environment's vstest runner cannot
negotiate over .NET Framework). Please confirm on a Windows/.NET Framework-capable runner.
StressSpec's churn-phase port (tailoring commit 5) intentionally does not port dev's extraCreateResultAggregatorAsyncbarrier, a separate hardening against aPartitionSeveralaggregator-identification race; the retrying lookup narrows that race but doesn't fully close
it the way the barrier does.
akka.cluster.prune-gossip-tombstones-after, a HOCON key v1.5'sCluster.confdoes not have. It is inert on v1.5 (unknown keys are ignored) but that part of thespec's intent does not transfer; left as-is since it is a faithful cherry-pick and does no harm.
scope here since Bump FsCheck.Xunit from 3.3.3 to 3.4.0 #8517 doesn't apply as-is.