Skip to content

Backport September 2026 de-flakes, product fixes, and TestKit changes from dev to v1.5 - #8588

Merged
Aaronontheweb merged 44 commits into
v1.5from
backport/dev-2026-09-to-v1.5
Sep 12, 2026
Merged

Aaronontheweb merged 44 commits into
v1.5from
backport/dev-2026-09-to-v1.5

Conversation

@Aaronontheweb

@Aaronontheweb Aaronontheweb commented Sep 12, 2026 •

Copy link
Copy Markdown
Member

Merge with a merge commit, not a squash, so each cherry-pick keeps its -x reference to the dev commit.

Summary

Backports the de-flakes, product fixes, TestKit changes, and a Cluster-startup fix merged into
dev since 2026-09-01 (plus four pre-window prerequisites) to v1.5. 33 commits are faithful
cherry-picks (git cherry-pick -x) in dependency order; 11 further commits adapt or hand-port the
pieces that need v1.5-specific work, each kept separate so it can be reviewed or dropped on its
own.

Cherry-picked from dev

# Fixes Category
#8499 TestKitBase.ShutdownAsync - non-blocking system shutdown, used by later picks TestKit
#8406 ClusterShardingSpec recovery phase: keep barriers out of Within De-flake
#8286 NodeChurnSpec: migrate to async TestKit APIs De-flake
#8570 Streams: Tcp Unbind now completes at once instead of always waiting out the subscription timeout Product fix
#8574 Sharding: remember-entities write timeout now uses updating-state-timeout, matching docs Product fix
#8582 Replicator.IsKnownNode now trusts members first seen Leaving/Exiting/Downed Product fix
#8586 ShardCoordinator hand-over only stops the local ShardRegion when the node is actually leaving Product fix
#8561 TestKit: DotNetty batching override moved under akka.remote, where the transport reads it TestKit
#8515 MNTR: conductor port race eliminated (bind port 0, propagate via stdout sentinel) Test infra
#8558 ClusterShardingLeaseSpec: join in InitializeAsync instead of blocking a pool thread De-flake
#8545 TestKit.Xunit: async dispose chain so a derived DisposeAsync no longer leaks the ActorSystem TestKit
#8540 ClusterSpec: stop passing an already-cancelled token to WithinAsync De-flake
#8579 ClusterSpec: wait for Up before leaving, one leader pass instead of two De-flake
#8541 ReceiveTimeoutSpec: dilated probe waits instead of a flat latch De-flake
#8566 BugFix3724Spec/HubSpec: zero-count event filter gets its own window De-flake
#8587 PropsSpec: await the producer latch instead of blocking on it De-flake
#8564 InboxSpec: receive through ReceiveAsync so a deadline surfaces as TimeoutException De-flake
#8584 EntityTerminationSpec: anchor on the shard's own entity set De-flake
#8573 ShardingBufferAdapterSpec: fresh probe per re-send attempt De-flake
#8498 ClusterSingletonProxySpec/ClusterSingletonRestartSpec de-flaked De-flake
#8500 ClusterShardingSpec family: re-send Gets that race shard hand-off De-flake
#8509 ReplicatorChaosSpec/ClusterShardingLeavingSpec de-flaked De-flake
#8552 DistributedPubSubRestartSpec: restarted node makes first contact, survivor waits on it De-flake
#8542 RemoteDeliverySpec: final barrier gets a budget covering the loop De-flake
#8544 ClusterShardingRegistrationCoordinatedShutdownSpec: wait on the spec's own budget De-flake
#8563 ClusterDeathWatchSpec: resolve the end actor over the return lane before sending End De-flake
#8569 ClusterSingletonManagerLeave2Spec: watch the proxy from a dedicated probe De-flake
#8581 RemoteNodeDeathWatchSpec: resend-derived bound, Identify asserts the actor exists De-flake
#8556 NodeChurnSpec: turn off legacy auto-down, widen the heartbeat pause De-flake
#8404 RemoteNodeRestartDeathWatchSpec: out-wait the association gate (classic-lane fix, in place of the Artery-only #8557) De-flake
#8359 Akka.Cluster: make Cluster extension startup non-blocking (fixes a StressSpec CI deadlock) Product fix
#8580 Cluster: route user commands through /system/cluster for the extension's lifetime; LeaveAsync re-sends its Leave (fixes an ordering regression #8359 introduces) Product fix
#8398 Bugfix5962Spec: dilated async waits, dynamic port, async member-up signal - dev's own follow-up to #8359 (a non-dilated 1 second ResolveOne raced the now-asynchronous cluster-extension init); found by the downstream audit and the Opus review De-flake

Note on #8515: the conflicting hunk in MultiNodeTestRunner.cs bundled #8515's own change
with DefaultNodeExitTimeout, which turned out to be the payload of dev's earlier #8431 (a
20-minute backstop that kills a hung node process) - v1.5's own prior #8431 backport had dropped
this file's hunk, so the field did not exist on v1.5 before this pick. Taking dev's whole file for
the merge, the only way to give the field a real consumer, lands the #8431 node-exit backstop on
v1.5 for the first time alongside #8515's port-sentinel change. Details in that commit's message.

v1.5 tailoring commits

These are not cherry-picks - each adapts a picked change, or ports one dev change by hand,
to v1.5's own code shape. Listed in commit order:

  1. Drop BREAKING_CHANGES_V1.6.md - the TestKit.Xunit: implement the async dispose chain so a derived DisposeAsync no longer leaks the ActorSystem; de-flake StreamRefsSpec #8545 cherry-pick touches this file on dev; v1.5 has no v1.6 release ledger, so it is removed again after the picks land.
  2. Remove Artery-only configuration from backported specs - one HOCON line and a few comments in the backported DistributedPubSubRestartSpec.cs referenced Artery, which does not exist on v1.5; removed/reworded, no behavior change.
  3. Regenerate TestKit API approvals for ShutdownAsync - the API-approval snapshot needed re-accepting for the two new methods Add TestKitBase.ShutdownAsync and adopt it in StressSpec churn teardown #8499 adds; the net48 counterpart was derived by reasoning rather than executed (see Verification).
  4. Fix a Process.Kill(bool) call surfaced by the MNTR: eliminate the conductor port race - bind port 0 and propagate via stdout sentinel #8515 hand-merge - netstandard2.0 lacks the process-tree-kill overload dev's file uses; switched to the parameterless overload, matching an existing convention elsewhere in the codebase.
  5. StressSpec churn phase - abrupt removal budget and async teardown (v1.5-native port of De-flake StressSpec: leave the cluster before terminating a churn system; run it with 7 nodes on the 2-vCPU hosted agents #8543) - not a cherry-pick: De-flake StressSpec: leave the cluster before terminating a churn system; run it with 7 nodes on the 2-vCPU hosted agents #8543 sits on a 5-deep dev-only stack. Hand-ported the raised heartbeat-pause, akka.test.timefactor/single-expect-default (from De-flake StressSpec: retry the aggregator Identify, fix blocking WatchAsync, add timefactor #8372, the first commit of the same stack - without these two lines every Within bound in the file stays undilated while the heartbeat-pause widening triples the detection time it has to cover, see "StressSpec root cause and fix" below), the ChurnMemberRemovalWithin budget term, the async call-site migration, the hardened RemoveOneAsync watchee lookup (fresh probe per attempt, null-Subject guard, bounded WatchAsync), and adapted StressSpecConfig's node-count floor arithmetic to v1.5's heavier default config (13 nodes / 2-per-phase vs dev's 10/1) - still reaches the same practical floor of 7 nodes. Removed the sync ClusterResultAggregator()/ReportResult<T>(Func<T>) overload pair, which had no callers left once the async chains were repointed. A new StressSpecConfigSpec.cs documents and asserts the node-count derivation.
  6. Reword backported test comments for classic remoting and the StressSpec node-count doc - three comment-only corrections: the #8579 pick's ClusterSpec.cs comment on the MemberUp wait described dev's ClusterCore re-targeting from a supervisor to the core daemon mid-startup, a mechanism this branch no longer has once Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) #8359/Cluster: route user commands through /system/cluster for the extension's lifetime so their order survives startup; LeaveAsync re-sends its Leave #8580 land (ClusterCore always targets /system/cluster) - reworded to say the wait is still the stronger barrier because it proves the join was processed, not merely enqueued; a comment introduced by tailoring commit 2 wrongly claimed a Tell to a quarantined peer is dropped at the transport layer (on v1.5's classic remoting it is not - see EndpointManager.cs); and the StressSpec.cs BuildConfig doc comment contradicted its own arithmetic. No behavior change.
  7. Allow the multi-node pipeline template to pass env, run StressSpec at 7 nodes - adds the env parameter dev's CI template has, and sets MNTR_STRESSSPEC_NODECOUNT=7 on v1.5's one MNTR lane.
  8. Add a discriminating ClusterSpec fact for LeaveSelf's re-send - now that Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) #8359 and Cluster: route user commands through /system/cluster for the extension's lifetime so their order survives startup; LeaveAsync re-sends its Leave #8580 are both cherry-picked in full, dev's own ported fact (A_cluster_must_process_a_Leave_issued_immediately_after_Join_from_the_same_thread) already covers the ClusterCore-routing half; it calls Leave() exactly once and behaves identically before and after the re-send fix, so it doesn't discriminate the re-send itself. This commit adds one more fact, A_cluster_must_resend_Leave_on_a_later_LeaveAsync_call_after_an_earlier_send_was_lost, that does (see "Testing" below).
  9. Release notes for the September backport - adds a 1.5.72 TBD section covering every user-visible change in this PR, including the non-blocking Cluster startup fix and a migration note for Cluster.Sharding: arm the remember-entities write timeout with updating-state-timeout, as the config documents #8574's write-timeout key change.
  10. Fix net48 build breaks from the Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) #8359/De-flake Bugfix5962Spec.SBR_Should_work_with_channel_executor #8398 cherry-picks - the Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) #8359 cherry-pick's ClusterStartupFuzzSpec used Task.IsCompletedSuccessfully, and the De-flake Bugfix5962Spec.SBR_Should_work_with_channel_executor #8398 cherry-pick's Bugfix5962Spec used the non-generic TaskCompletionSource/TrySetResult() - neither exists on .NET Framework 4.8 (v1.5's Akka.Cluster.Tests still targets net48, unlike dev); switched to Status == TaskStatus.RanToCompletion and TaskCompletionSource<Done>/TrySetResult(Done.Instance) respectively.
  11. Collect a hang dump and the in-progress test list when a net10.0 unit-test host hangs - appends --blame-hang --blame-hang-timeout 15m --blame-hang-dump-type mini to the Windows and Linux .NET Unit Tests commands in build-system/pr-validation.yaml (the testsOnly.json/net10.0 lanes only, not the net48 lane or the multi-node lane). See "CI evidence for a pre-existing Linux hang" below for why.

Not backported

# Reason
#8372 First commit of the #8543 StressSpec stack. Its retrying-aggregator and blocking-WatchAsync fixes are superseded by tailoring commit 5's own port of the same mechanism; its two config lines (akka.test.timefactor = 3, akka.test.single-expect-default = 10s) are ported directly into that same commit
#8427 Second commit of the #8543 stack (a CircuitBreakerSpec de-flake plus a retrying ReportResult aggregator lookup for StressSpec) - the StressSpec half is superseded by tailoring commit 5; the CircuitBreakerSpec half is out of scope for this PR
#8429 Third commit of the #8543 stack (fixes double-dilated Within/AwaitAssertAsync bounds in StressSpec and ClusterShardingQueriesSpec) - the StressSpec half is superseded by tailoring commit 5; the ClusterShardingQueriesSpec half is out of scope for this PR
#8504, #8554 Artery-only; src/core/Akka.Remote/Artery does not exist on v1.5
#8557 Artery-lane-only fix stacked on #8404; the classic-lane fix (#8404) is cherry-picked instead
#8575 Artery-lane-only, stacked on an Artery-only rewrite of the spec
#8512-8514, #8518-8520, #8525-8528, #8530, #8532-8534, #8536-8538 Akka.Serialization.V2 generator stack and docs; that project does not exist on v1.5
#8517 Bumps FsCheck 3.3.3->3.4.0; v1.5 is on a different FsCheck major (2.16.6) - not the same edit
#8546, #8560 A change and its same-day revert on dev; net-zero, superseded by #8561
#8553 CI env-var change folded into #8543's squash; covered by tailoring commit 7
#8510, #8516, #8524, #8535 Already on v1.5 via earlier backport PRs

Verification

Environment note: this verification environment's SDK pulls in a transitive package that refuses to build the
net6.0 library target under strict warnings; reproduced identically on a pristine, unmodified
v1.5 checkout, so it is a local tooling artifact, not a backport defect
(-p:SuppressTfmSupportBuildWarnings=true used to work around it for local verification only).
Directory.Build.props already sets TreatWarningsAsErrors=true unconditionally, so every test
run below already enforces warnings-as-errors without a separate -warnaserror build step.

StressSpec root cause and fix (see tailoring commit 5)

Tailoring commit 5 ports #8543's acceptable-heartbeat-pause widening (3s -> 20s) into
StressSpec, together with #8372's two companion config lines, akka.test.timefactor = 3 and
akka.test.single-expect-default = 10s (#8372 is the first commit of the same 5-deep stack
#8543 sits on). Both lines matter: every Within/WithinAsync bound in this spec is dilated by
akka.test.timefactor - including RemoveOneAsync's removal budget,
TimeSpan.FromSeconds(25) + ConvergenceWithin(TimeSpan.FromSeconds(3), NbrUsedRoles - 1), which
#8543 never widened directly. Without the factor, that budget is a flat 31s at
NbrUsedRoles = 3 - the node count the 7-node phase sequence reaches during
MustShutdownNodesOneByOneFromSmallClusterAsync. ChurnMemberRemovalWithin()'s own arithmetic
puts the actual abrupt-removal detection time at 34.3s at this spec's config, which structurally
cannot fit inside a 31s ceiling. With the factor restored, the same budget is 93s. Also ported
dev's hardened RemoveOneAsync watchee lookup (fresh probe per attempt, null-Subject guard,
bounded WatchAsync), which is the method that was failing.

Re-ran StressSpec three times at MNTR_STRESSSPEC_NODECOUNT=7 after the fix, in full each time:

Run Result Total time shutdown 1 in 4 nodes cluster shutdown one from 3 nodes cluster (the previously-failing RemoveOneAsync)
1 16/16 passed 3m 10s (minimal logging) (minimal logging)
2 16/16 passed 3m 10s 34712 ms 33453 ms
3 16/16 passed 3m 14s 33468 ms 32621 ms

shutdown one from 3 nodes cluster is the exact phase that used to fail at
MustShutdownNodesOneByOneFromSmallClusterAsync -> RemoveOneAsync. At NbrUsedRoles = 3,
RemoveOneAsync's budget is now (25 + 3 * 1 * 2) * 3 = 93s (dilated by the restored
timefactor), comfortably covering the measured 32.6-33.5s cost - previously that same phase had
a flat, undilated 31s ceiling against this same ~33s cost, which is why it failed 3/3 times
before this fix. partition 2 in 7 nodes cluster (33.8-34.4s across runs 2-3) independently
confirms ChurnMemberRemovalWithin()'s predicted 34.3s detection time.

#8359 / #8580: a second StressSpec failure mode, found in this PR's own CI

After this PR was first opened, CI build 131547 (re-run on this PR's previous head, after the
churn-phase fix above had already landed) failed StressSpec on the Windows MNTR lane with a
different signature than the one above: JoinSeveralAsync timed out waiting for node-6 to join
("Expected: 6, Actual: 5"). Node-6's own log showed "Started up successfully" at 04:06:01.605
(the Cluster constructor's ask answered in ~140 ms that time), then two "Failed to startup
Cluster ... AskTimeoutException: Timeout after 20.00 seconds" at 04:06:21.57 - one from the
cluster event-bus listener's PreStart, one from the SBR downing provider's PreStart - followed
by Shutdown(). StressSpec pins the default dispatcher's fork-join pool to
parallelism-min=2/factor=1 (unchanged by this PR, and identical to dev's own setting), so on
the 2-core CI agent both of those dedicated threads parked on the ClusterCore getter's re-issued
blocking GetClusterCoreRef().Result asks (the constructor's own ask had already been answered), and
the daemon that had to answer them could never get scheduled:
a hard deadlock, resolved only by the 20s ask timeout. Raising akka.actor.creation-timeout would
only delay that timeout, not fix the deadlock. This is exactly the failure mode dev's #8359 fixes
by removing the blocking ask; #8580 rides along in the same commit set because #8359 itself
introduces a command-ordering regression that #8580 fixes: #8359's non-blocking ClusterCore
getter switched from /system/cluster to a direct core-daemon reference the instant that ref was
published mid-startup, so a Join and a Leave issued back-to-back by the same caller could arrive
at the core daemon out of order.

Local verification before adding these two commits to this PR (full logs in
8359-v15-verification.md, not included in the PR itself):

  • ClusterStartupFuzzSpec (randomized concurrent-startup fuzzing under pinned 1-2-thread
    dispatcher pools, Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) #8359's own regression test): 0 kills across 396 simulated iterations
    with the fix; 33 immediate kills across 42 completed iterations without it, reproducing the
    exact "Failed to startup Cluster" / AskTimeoutException signature from build 131547.
  • StressSpec at MNTR_STRESSSPEC_NODECOUNT=7, pinned to two cores (taskset -c 0,1, to match
    the CI agent's core count): 3/3 passed with the fix.

This session's own verification, after rebuilding this PR's cherry-pick/tailoring stack to
include #8359 and #8580 (see "Verification" above for the full list): Akka.Cluster.Tests full
project 368/368; ClusterSpec 22/22 across three separate runs, including both Leave facts;
ClusterStartupFuzzSpec 0 kills across 132 iterations; API approvals 18/18 with no diff (the
change is entirely internal); and StressSpec at MNTR_STRESSSPEC_NODECOUNT=7, pinned to two
cores, 16/16 passed in 3m 14s.

Testing

  • ClusterSpec run 8 times total across this PR's history (22/22 passing every time), including
    two Leave-ordering facts: dev's own #8580 fact,
    A_cluster_must_process_a_Leave_issued_immediately_after_Join_from_the_same_thread, ported
    unchanged (it calls Leave() exactly once and behaves identically before and after the
    LeaveSelf re-send fix, so it's a Join-then-Leave smoke test, not proof of the re-send); and
    tailoring commit 8's own fact,
    A_cluster_must_resend_Leave_on_a_later_LeaveAsync_call_after_an_earlier_send_was_lost, which
    does discriminate the re-send. That fact issues LeaveAsync() before any Join (stashed by
    ClusterDaemon until its own Init completes, then unhandled and dropped by
    ClusterCoreDaemon's Uninitialized behavior, which has no case for ClusterUserAction.Leave),
    joins, waits for MemberUp, then calls LeaveAsync() again and asserts it completes. Confirmed
    it actually discriminates: reverting LeaveSelf() to the pre-fix one-shot version and re-running
    just this fact fails it with a TimeoutException after 10s (the dead-lettered first Leave is
    logged: Message [Leave] from [.../testActor1] ... was unhandled. [1] dead letters encountered);
    restoring the file returns it to passing in under a second.
  • StreamRefsSpec run twice (16/16 passing each time, one pre-existing skip unrelated to this
    backport).
  • ClusterStartupFuzzSpec (Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) #8359's own regression test) run once after rebuilding this PR's
    stack: both profiles passed, 0 kills across 132 simulated startup iterations.

Open items for the maintainer

  • The Net.verified.txt (net48/netstandard2.0) API-approval file for ApproveTestKit was
    updated by derivation, not by running the net48 test host (this environment's vstest runner cannot
    negotiate over .NET Framework). Please confirm on a Windows/.NET Framework-capable runner.
  • StressSpec's churn-phase port (tailoring commit 5) intentionally does not port dev's extra
    CreateResultAggregatorAsync barrier, a separate hardening against a PartitionSeveral
    aggregator-identification race; the retrying lookup narrows that race but doesn't fully close
    it the way the barrier does.
  • De-flake NodeChurnSpec: turn off legacy auto-down, widen the heartbeat pause over the measured stall, and make the payload listener match real log messages #8556's cherry-pick sets akka.cluster.prune-gossip-tombstones-after, a HOCON key v1.5's
    Cluster.conf does not have. It is inert on v1.5 (unknown keys are ignored) but that part of the
    spec's intent does not transfer; left as-is since it is a faithful cherry-pick and does no harm.
  • Whether a v1.5 FsCheck bump is wanted at all (v1.5: 2.16.6/3.3.2, dev: 3.4.0/3.3.3) - out of
    scope here since Bump FsCheck.Xunit from 3.3.3 to 3.4.0 #8517 doesn't apply as-is.

Aaronontheweb and others added 30 commits September 11, 2026 23:39
…wn (#8499)

* Add TestKitBase.ShutdownAsync and use it in StressSpec churn teardown

Closes #8497.

TestKitBase.Shutdown is a blocking `Terminate().Wait(duration)`. On a loaded
agent that pins a thread pool thread for the whole wait, and the work the
shutdown is waiting on - coordinated shutdown phases, cluster heartbeats -
competes for the threads that are left. In StressSpec's churn phase the wait
ran its full 10s and starved the hosting node's heartbeat sender for 12+
seconds, so the survivors marked it unreachable and SBR downed it. The wait
made its own timeout more likely.

ShutdownAsync races Terminate() against a timer instead of blocking, using the
AwaitWithTimeout helper Akka.TestKit already ships. Everything else matches the
sync overloads: same default duration, same forced guardian stop, same
verifySystemShutdown throw-vs-log, same message. Both paths now share the
force-stop tail, so the two cannot drift.

The sync Shutdown keeps its own Terminate().Wait rather than blocking on the
async one - sync-over-async here would need two continuations off the pool
while a thread is pinned, which is worse under the starvation this fixes.

StressSpec's post-churn teardown now awaits ShutdownAsync. Same wait budget,
same fallback; the wait just no longer holds a thread.

* Await the in-loop churn teardown in StressSpec too

The post-loop teardown already awaits ShutdownAsync. This is the other blocking
Shutdown in ExerciseJoinRemoveAsync - it tears down the previous round's system
at the top of each churn round, so it is the one that blocks while the cluster
is under load. Leaving it on Terminate().Wait would keep pinning a thread pool
thread for the whole wait, which is what starves the hosting node's heartbeat
sender.

Same wait budget, same forced-guardian fallback, same ordering: the teardown
still finishes before the next ActorSystem is created.

* Drop the redundant happy-path ShutdownAsync test

The verify-throws test already proves ShutdownAsync waits: a fire-and-forget
implementation cannot observe that the system failed to stop, so it cannot
throw the TimeoutException the test requires. StressSpec exercises the happy
path on every MNTR run. The two remaining tests stay separate because each
call runs the forced-guardian-stop tail, which kills the stuck system - so
the throw branch and the log branch each need their own system.

(cherry picked from commit 81289e8)
…hin (#8406)

* De-flake ClusterShardingSpec recovery phase: keep barriers out of Within

One WithinAsync(50s) spanned five barriers plus an unbounded ExpectMsgAsync, so a slow remember-entities recovery drained the shared budget and the final barrier's timeout clamped to zero. Bound the recovery check in a retrying probe-per-attempt AwaitAssert and move the barriers outside the Within so each gets the full barrier-timeout.

* Remove the phase's outer Within entirely

The remaining Within(50s) still wrapped three barriers and bounded steps whose worst cases sum past it. Every wait now carries its own explicit bound; all five barriers run outside any Within.

(cherry picked from commit f13a8fa)
…out flakiness (#8286)

NodeChurnSpec intermittently failed with "Failed to stop [NodeChurnSpec] within
[00:00:05]" on slower CI agents. The synchronous Shutdown(node,
verifySystemShutdown:true) helper blocks a thread-pool thread inside Task.Wait()
for its 5s budget while the coordinated-shutdown pipeline itself needs the thread
pool to make progress — a sync-over-async self-starvation.

Migrate the spec to the async, task-returning TestKit APIs (WithinAsync,
AwaitMembersUpAsync, EnterBarrierAsync, AwaitAssertAsync, ExpectNoMsgAsync) and
replace the blocking per-system Shutdown loop with a concurrent
`await Task.WhenAll(systems.Select(s => s.Terminate())).WaitAsync(30s)`, which frees
the thread and preserves verify-shutdown semantics (throws TimeoutException if a
system fails to stop). Same idiom as QuickRestartSpec / DistributedPubSubRestartSpec.

Verified locally across 4 consecutive runs (3/3 node roles pass each, ~45s).

(cherry picked from commit 4b69aa5)
…ting initialization; the counter was seeded at -1 (#8570)

ConnectionSourceStageLogic's _connectionFlowsAwaitingInitialization counter used
AtomicCounterLong's default constructor, which seeds an id generator at -1 so
the first Next() returns 0. JVM Akka's TcpStages.scala uses `new AtomicLong()`,
which starts at 0. Because the counter never rests at 0 in this port, the
UnbindCompleted() fast path (`if (_connectionFlowsAwaitingInitialization.Current == 0) CompleteStage();`)
was unreachable, so every graceful ServerBinding.Unbind() fell through to the
BindShutdownTimer and waited the full akka.stream.materializer.subscription-timeout.timeout
(5s by default) even when no connection was ever accepted.

Seed the counter at 0 to match upstream semantics and restore the fast path.
Adds a regression test that binds, never connects, and asserts Unbind()
completes well inside the subscription timeout.

(cherry picked from commit 91556a0)
…ng-state-timeout, as the config documents (#8574)

* Cluster.Sharding: arm the remember-entities write timeout with updating-state-timeout, as the config documents; the read setting had been used since #6479

Shard.SendToRememberStore armed the remember-entities WRITE timeout with
TuningParameters.WaitingForStateTimeout (2s), while the exception it throws on
timeout, reference.conf's own doc comment for updating-state-timeout ("Also
used as timeout for writes of remember entities when that is enabled"), and
DDataRememberEntitiesShardStore's write-majority retry sizing (3 retries at
updating-state-timeout / 4) all agree the write should use
UpdatingStateTimeout (5s). Commit dde36e9 (#6479) swapped both the read and
write sites in one hunk but only described fixing the read path, leaving the
write path armed with the wrong, shorter timeout and giving the store's
retries zero time to run before the shard gave up and restarted, dropping any
buffered messages for the entity.

Restores UpdatingStateTimeout on the write arm; the read path (already using
UpdatingStateTimeout, more leniently than Pekko) is untouched. Adds
RememberEntitiesWriteTimeoutSpec, which fails on unpatched code (the shard
restarts and the buffered message is dropped) and passes with the fix (the
store's delayed-but-valid write ack arrives and no restart occurs).

* RememberEntitiesFailureSpec: prove the write timeout fix with the existing fake store instead of a new spec

Deletes RememberEntitiesWriteTimeoutSpec.cs's standalone 180-line spec and its bespoke
delayed-ack store, replacing it with one fact on RememberEntitiesFailureSpec that reuses
the shared FakeStore/FakeShardStoreActor harness. The fact delays a remember-entities
write's ack by 3s (strictly between the 2s waiting-for-state-timeout and 5s
updating-state-timeout reference defaults) and asserts the buffered message is still
delivered and the shard does not restart.

FakeShardStoreActor's existing Delay failure mode replies to a delayed write with the
wrong payload type, so the write never actually completes - every existing Delay fact
only relies on that to force a timeout. Adds a new DelayedSuccess failure mode, additive
and reusing the same Delayed/timer plumbing, that acknowledges the write correctly so a
write can genuinely finish late but within its timeout - which the existing Delay mode
could not exercise.

Verified this fact discriminates: reverting Shard.cs's one-line fix by hand made it fail
(shard restarts at ~2s per the log, buffered message lost); restoring the fix made it
pass. All 26 facts in RememberEntitiesFailureSpec pass in three consecutive runs.

(cherry picked from commit 2d8fa21)
…ting/Downed (#8582)

* Fix Replicator.IsKnownNode to trust members first seen as Leaving/Exiting

InitialStateAsEvents replays each current member at its CURRENT status, so a
Replicator that starts after a member has already moved to Leaving or Exiting
receives MemberLeft/MemberExited and never MemberUp. IsKnownNode inferred
membership from MemberUp alone, so every Write and gossip message from that
member was silently dropped for as long as it stayed in the cluster. In cluster
sharding this can stall a departing shard coordinator's ddata write for its
entire updating-state-timeout during a rolling restart.

_exitingNodes covers MemberExited. MemberLeft/MemberDowned already land in
_leader via ReceiveOtherMemberEvent. Both sets are pruned in
ReceiveMemberRemoved, so a removed node stops being trusted immediately. This
reuses state the actor already maintains rather than adding a new membership
set.

Replica set composition is unchanged: AllNodes remains _nodes + _weaklyUpNodes,
so a dying node's Write must still win a majority of current replicas. Only who
may ask is widened, not who may vote.

* Add MemberDowned coverage to ReplicatorKnownNodeSpec

MemberDowned shares the generic ReceiveOtherMemberEvent path with MemberLeft,
so it is the third status a late subscriber can be handed on
InitialStateAsEvents replay and was the one status the spec did not cover.

Note the enum value is MemberStatus.Down while the cluster event is
MemberDowned.

(cherry picked from commit 2c2a406)
…egion when the node is leaving (#8586)

ClusterSingletonManager sends the coordinator its termination message on
every hand-over, and a hand-over is not always a shutdown: while a cluster
forms, two members of the same role can both reach Oldest, and the loser
hands over while staying Up. ShardCoordinator.HandleTerminate read every
termination message as a node shutdown, sent GracefulShutdown to the
node's own live region, and deferred its stop until the region was gone.
ClusterShardingGuardian then evicted the dead region from ClusterSharding's
cache, so ShardRegion(typeName) threw on that node for the rest of the
process.

The coordinator still stops on every hand-over, which is what lets the
singleton manager complete it. It only takes the local region down with
it when that region is already handing off, which is the coordinated
shutdown path, or when the node itself is leaving: the system is
terminating, CoordinatedShutdown is running, or the self member is
Leaving, Exiting, or Down. The branch that sends GracefulShutdown to a
live local region came from JVM Akka PR 30338, a rolling-update latency
optimization; before it a termination message was a bare stop, which is
what a benign hand-over now gets again.

ShardCoordinatorHandOverSpec sends the singleton's termination message the
way the manager does and asserts that the coordinator stops and the region
does not. It fails on the old code with the region's Terminated.

Fixes #8583.

(cherry picked from commit c0600b7)
…the transport reads it (re-land of #8546 with the RemoteConfigSpec fix) (#8561)

* TestKit: put the DotNetty batching override under akka.remote, where the transport reads it (#8546)

The key sat under akka.test and never took effect. A config test now guards it.

(cherry picked from commit d88eec8)

* RemoteConfigSpec: read batching defaults from the Akka.Remote reference config, and pin the TestKit override separately

Since #8546 moved the DotNetty batching override to where the transport reads it, the running test system's config now carries the TestKit's batching-disabled override, so it no longer reflects the Akka.Remote defaults. The BatchWriter fact now reads straight from Akka.Remote.Configuration.RemoteConfigFactory to pin the library default, and a new sibling fact reads from the running system to pin the TestKit override, so a future misplacement of either key fails a test.

(cherry picked from commit adacc74)

* TestKit_Config_Tests: drop the stale copy of the misplaced batching key from the custom test config

(cherry picked from commit 3f966aa)
…ia stdout sentinel (#8515)

* Remove the MNTR conductor port race: bind port 0 and propagate

The multi-node runner picked the TestConductor's port by binding a temporary
socket, reading the port, closing the socket, and handing the number to node 1.
The number is in the ephemeral range by construction, so any outbound connection
on the machine can take it between the probe and the conductor's real bind.
When that happens node 1 dies with "Address already in use" and every other node
waits out a 30 second attach timeout against a conductor that never existed.

The runner now starts node 1 alone with multinode.server-port=0. The node binds a
free port, prints it as a sentinel line, and the runner starts the remaining nodes
with the port the conductor actually holds. There is no window to lose.

- ConductorPortSentinel: the stdout contract between a conductor node and the
  runner, with a strict parser.
- Controller reports its bind result through a TaskCompletionSource instead of an
  ask. A failed bind now throws ConductorBindException naming the port, straight
  away, rather than surfacing as a query timeout.
- MultiNodeSpec accepts server-port=0 on the conductor node and rejects it on
  client nodes, and emits the sentinel once the conductor is bound.
- When the conductor node exits early or never publishes a port, the runner kills
  it and fails the nodes it never started with a message naming the cause. A spec
  that used to hang for minutes now fails in under a second.

Explicitly configured ports are unchanged: the node binds that exact port and
emits the sentinel anyway, so running node processes by hand still works.

Applied to both the xUnit v3 and xUnit v2 adapters.

* Model the conductor bind as a Controller behavior switch

The Controller constructor blocked on the DotNetty bind, so a dispatcher thread
sat in the actor's constructor waiting on I/O and a failed bind escaped as an
exception out of construction.

The constructor now only assigns fields. PreStart starts the bind and PipeTo's
the outcome back into the mailbox as Bound or BindFailed, so the bind never
blocks a thread and its result is serialized with every other message.

- Binding (initial behavior): Bound assigns the connection, creates the
  BarrierCoordinator, completes the bind TCS - barrier first, report second, as
  before - and becomes Ready. BindFailed reports the failure and stops. It must
  stop rather than throw: a throw from a message handler runs the default
  supervision directive, Restart, which re-runs PreStart and re-binds the same
  dead port in a loop.
- Anything else arriving during Binding is stashed. The socket starts accepting
  the moment the bind completes, which is before the actor processes Bound, so a
  client that connects in that window can get CreateServerFSM in ahead of it.
- Ready is the previous receive, now entitled to a non-null connection and
  barrier by construction of the state machine.
- The bind failure is wrapped into ConductorBindException at the pipe, and the
  single-exception AggregateException a piped failure arrives in is unwrapped, so
  the reported cause is the socket error itself.
- PostStop guards a null connection and no longer blocks on the pool drain.
  ReleaseAll detaches the pools before it returns, so a later CreateConnection
  builds fresh ones either way.

The caller contract is unchanged: StartControllerAsync still awaits the same TCS,
still gets the endpoint only after the barrier exists, and still sees a failure
carrying the named error. Node 1 emits its port sentinel at the same point,
between the bind and the wait for players.

(cherry picked from commit 2ece2c6)

v1.5 backport note: the conflicting hunk in MultiNodeTestRunner.cs bundled this change's own
_conductorPortSource field together with DefaultNodeExitTimeout, which was pure context here but
is itself the payload of dev's earlier #8431 (a node-exit backstop that kills a hung node process
after a timeout). v1.5's own #8492 backport of #8431 had dropped this file's hunk, so
DefaultNodeExitTimeout did not exist on v1.5 before this pick. Taking dev's whole file for this
merge - the only way to give the field a real consumer - therefore also lands the #8431 node-exit
backstop on v1.5 for the first time, not just the #8515 port-sentinel change.
…c instead of blocking a pool worker in the constructor (#8558)

* De-flake ClusterShardingLeaseSpec: join the cluster in InitializeAsync instead of blocking a pool worker in the constructor

The constructor's blocking wait for Up against a flat 3s default missed
by 462ms on CI, because the join's six dispatches queued behind the
parked worker on a starved thread pool. Cluster formation now happens
in IAsyncLifetime.InitializeAsync via JoinAsync/StartAsync, which waits
on the real MemberUp signal without parking a thread. The file is also
migrated to the async TestKit API per the repo's standing rule.

* ClusterShardingLeaseSpec: drop the redundant TestActor touch; the ambient-context attribute already pins the cell

(cherry picked from commit 408e008)
…Async no longer leaks the ActorSystem; de-flake StreamRefsSpec (#8545)

* TestKit.Xunit: implement the async dispose chain so derived DisposeAsync no longer leaks the ActorSystem; de-flake StreamRefsSpec

Fixes #8191: xUnit v3 calls IAsyncDisposable.DisposeAsync() in preference to
IDisposable.Dispose() whenever a type implements both, so a derived spec's own
no-op DisposeAsync (e.g. StreamRefsSpec) skipped TestKit's whole synchronous
dispose chain and silently leaked a remoting-enabled ActorSystem per test.
Akka.TestKit.Xunit.TestKit now implements IAsyncLifetime with a virtual
InitializeAsync/DisposeAsync pair; DisposeAsync runs the sync chain and then
shuts the system down with the non-blocking ShutdownAsync(). Ports the
TestKit.cs shape from stalled PR #8217 (targeted v1.5) onto dev; its
TestKitBase.cs hunk is dropped because ShutdownAsync already landed on dev via
81289e8, and EventFilterTestBase.cs is hand-merged onto dev's current
AwaitAssertAsync retry body. Five other TestKit-derived specs that declared
their own DisposeAsync/InitializeAsync are updated to override and chain to
base: BugFixSpec, Bugfix8144Spec (Xunit v3), ParallelAmbientContextSpec (both
base classes), and EventFilterTestBase.

StreamRefsSpec.SinkRef_must_receive_elements_via_remoting is de-flaked on top
of the leak fix: the test now gates on WatchTermination(Keep.Right) instead of
asserting on a flat 3s wall-clock wait, uses async TestKit calls throughout,
and gives every cross-boundary wait a dilated 15s budget. Remote system port
is now 0 and loglevel is DEBUG for stage-level tracing.

* StreamRefsSpec, BugFixSpec, Bugfix8144Spec: migrate to the async TestKit API

Convert ExpectMsg, ExpectNoMsg, and .Wait()-on-stream-completion calls to ExpectMsgAsync, ExpectNoMsgAsync, and awaited AwaitWithTimeout across all facts in the three touched files.

* StreamRefsSpec, BugFixSpec: finish the async migration; restore INFO logging

* ClusterShardingLeaseSpec: override the TestKit's async lifecycle instead of declaring its own

* TestKit.Xunit: one disposed flag; Shutdown moves out of Dispose(bool) so the async path needs no mode switch

Dispose(bool) now only runs AfterAll() (and any derived teardown); it no
longer terminates the ActorSystem and no longer guards re-entrancy. The
guard collapses to a single _disposed flag checked once, in each public
entry point: Dispose() runs the chain then calls the blocking Shutdown(),
DisposeAsync() runs the chain then awaits ShutdownAsync(). This drops the
now-redundant _disposing/_disposingAsync flags and the nested try/finally
that used to suppress Shutdown() inside the async path.

Verified: no Dispose(bool) override in the tree depends on the ActorSystem
being torn down when base.Dispose(disposing) returns (the only overrides in
the repo are on System.IO.Stream/GraphStageLogic, unrelated to this
TestKit); StreamRefsSpec, BugFixSpec's DisposeAsync overrides still tear
down their own systems before chaining to base.DisposeAsync(), and
ClusterShardingLeaseSpec no longer declares one at all. Added a
Bugfix8191Spec case covering Dispose() called after DisposeAsync().

(cherry picked from commit 450bebf)

v1.5 backport note: re-added .WithFallback(ConfigurationFactory.Load()) to StreamRefsSpec's
Config(), which the hand-merge of the port/hostname change had dropped. Akka.Streams.Tests still
targets net48 on v1.5 (dev has no netfx lane), and its app.config sets
stream.materializer.debug.fuzzing-mode = on for that lane. ActorSystem.Create only calls
ConfigurationFactory.Load() when no config is supplied, so without this fallback an explicitly
configured ActorSystem - like the ones this spec creates - never saw app.config's fuzzing setting,
silently losing that coverage on the net48 lane.
…WithinAsync block (#8540)

* De-flake ClusterSpec: stop passing an already-cancelled token to the WithinAsync block

A_cluster_must_cancel_LeaveAsync_task_if_CancellationToken_fired_before_node_left
cancels `cts` to prove LeaveAsync(cts.Token) is cancelled, then reused that same
already-cancelled token as the `cancellationToken:` argument to the 10s
WithinAsync block that drives Leaving -> Exiting -> Removed. WithinAsync races
the block against a delay bound to that token; with the token born cancelled,
any real suspension point inside the block before the race resolves throws
OperationCanceledException out of WithinAsync (seen on CI as a 71ms failure).
Fast/synchronous runs hid the bug; slower runs exposed it.

Fix: drop the cancellationToken: argument so the block uses the default token.
While here, switch the synchronous ExpectMsg<ClusterEvent.MemberRemoved>() call
inside the block to the async ExpectMsgAsync, per the repo's async-only TestKit
convention.

* ClusterSpec: migrate the remaining synchronous TestKit calls to the async API

Migrates all 7 TestKit Shutdown(ActorSystem) calls in ClusterSpec.cs to ShutdownAsync, awaited in their enclosing async Task facts.

(cherry picked from commit 9705d36)
…e join is ordered first; one leader pass, then MemberRemoved (#8579)

Both LeaveAsync facts hand-ticked LeaderActions() right after Join, with no
event wait in between. Cluster.ClusterCore re-targets from the top-level
daemon to the core daemon once its ref is published, so a command issued in
that window can overtake a still-in-flight join and get dropped. The facts
also fired two leader ticks for Leaving -> Removed, though the second one is
always a no-op: Exiting -> Removed runs through the coordinated-shutdown
round trip the first tick starts, not through another tick.

Fix: subscribe before joining and wait for MemberUp as a real barrier before
issuing Leave (no tick needed there - a bootstrap self-join already runs a
leader pass inline). After Leave, fire exactly one LeaderActions() call and
wait for MemberRemoved with an explicit timeout. Drops the WithinAsync
wrapper and the ReceiveOneAsync-based retry loop in favor of direct,
bounded event waits.

(cherry picked from commit ec983f6)
…ted probe waits instead of one flat latch (#8541)

The latch built with `new TestLatch(...)` does not dilate with akka.test.timefactor,
and the test parked a pool worker in CountdownEvent.Wait while the scheduler's clock
loop and the mailbox shared the same thread pool. Rewrote the test to await a
TestProbe through three per-phase, dilated ExpectMsgAsync budgets instead of one flat
5s wall, stopping the actor in a finally. The probe sequence ("timeout", "tick",
"timeout") also proves the transparent tick was actually delivered, rather than just
inferring it from a latch count.

Also converts the companion negative-assertion test (no receive-timeout ever set)
from a blocking TestLatch wait to ExpectNoMsgAsync on a probe, so it stops parking a
pool worker for the duration of the wait.

(cherry picked from commit 86e479d)
… its own 3 s window instead of the Within's remaining time (#8566)

A zero-count EventFilter.ExpectAsync(0, action) with no explicit timeout,
run inside a WithinAsync(max) block, captures the Within's remaining time
as its own quiet window, runs the action, then waits out that whole window
to prove nothing was logged. The block therefore ends at
start + T_action + max by construction, while WithinAsync abandons a
still-running block at max + 200ms and fails the fact (since #8516). On a
cold CI agent the action's async part (cluster self-join in BugFix3724Spec)
doesn't leave enough margin.

Pekko's filter always uses akka.test.filter-leeway (3s, dilated) for its
quiet window, never the remaining Within time. Pass that window explicitly
via the ExpectAsync(count, timeout, action) overload in both specs so the
block runs in roughly T_action + 3s instead of racing the outer deadline.

(cherry picked from commit 3432d6f)
…t for a flat second (#8587)

CI build 131484 (Windows unit tests, 2-vCPU agent) failed Props_must_create_actor_by_producer
with a TestLatch.Ready(1s) timeout. The test blocks on a CountdownEvent that is counted down
when the actor cell instantiates the actor on a dispatcher thread, and a flat, undilated 1s
wait is the whole budget for one dispatcher hop on a stalled agent.

Make the fact async and replace the blocking Ready(1s) call with AwaitConditionAsync polling
latchActor.IsOpen over a 3s dilated bound (the same budget ExpectMsg gets from
akka.test.single-expect-default for a single dispatcher hop).

(cherry picked from commit 147a810)
…e surfaces as TimeoutException (#8564)

* De-flake InboxSpec: receive through ReceiveAsync so a deadline failure surfaces as TimeoutException, and restore the exact-type assertions

Inbox.Receive(timeout) waits on the wall clock while the inbox actor schedules its own deadline for the same duration; when that deadline wins the race, Task.Wait rethrows the actor's Status.Failure(TimeoutException) wrapped in AggregateException instead of the bare TimeoutException the tests expect. Awaiting ReceiveAsync instead surfaces the exception directly with no race, so this migrates every InboxSpec receive call to the async API and drops the IsTimeoutException either-type workaround added by an earlier fix.

* InboxSpec: queued-queries fact awaits delays instead of sleeping; exact exception type on the negative-deadline fact

Task.Run + Task.Delay staggers the two ReceiveWhere calls without parking
a pool thread; ReceiveWhere itself has no async overload and isn't in the
race, so it stays as-is.

Traced the negative-deadline fact through Inbox.Actor.cs's Kick handling
and FutureActorRef<object>.TellInternal: a deadline already in the past
always produces a bare TimeoutException("Deadline passed"), never
anything else, so ThrowsAnyAsync<Exception> is tightened to
ThrowsAsync<TimeoutException>.

(cherry picked from commit 989e299)
…and wait out the restart backoff without blocking a pool thread (#8584)

* De-flake EntityTerminationSpec: anchor on the shard's own entity set, wait out the restart backoff without blocking a pool thread

The passivation fact slept a fixed 400 ms after the test actor saw the
entity's Terminated and then read the shard state once. The shard learns
of the death through a system message it re-queues as a user message, so
the test actor seeing Terminated says nothing about the shard having
processed it; on a two-core agent the sleep itself held one of the two
pool workers, the shard's mailbox item sat in the other worker's local
queue, and the pool only moved again when the sleep expired. The state
read then saw the entity still active.

Each fact now waits for the shard's active entity set to reach the
expected value, which is the observable the assertions read. The
passivation fact then waits out the 250 ms entity-restart-backoff with
Task.Delay and reads the state once more, deliberately not polled, so a
wrongful restart still fails it. The restart fact waits on the shard's
own "Started entity" line for the restart instead of a sleep, since a
state poll alone can pass on the dead ref before the shard has processed
the Terminated. The non-remembering fact drops its sleep and the comment
about a backoff that does not exist on that path.

The file is migrated to the async TestKit API; the one-node join moves
from the synchronous AtStartup into an awaited helper.

* EntityTerminationSpec: query the region through a fresh probe per attempt; check for a wrongful restart with a zero-count filter on the shard's own restart line

The polling helper's 1s per-attempt bound is shorter than the region's 3s
query timeout, so a timed-out attempt left its reply in the test actor's
queue, where the passivation fact's final single read could take a
snapshot from before the wait as a current one, and the non-remembering
fact's "pong-2" expect could dequeue it. Each request now goes through a
fresh probe, and the region's Failed set is asserted empty.

The passivation fact replaces Task.Delay plus a state read with a
zero-count EventFilter on "Started entity" held open for 600 ms, which
observes a wrongful restart directly, then a final state read through a
probe with a 5 s bound.

(cherry picked from commit bef497a)
…fresh probe per attempt under a sharding-derived budget (#8573)

* De-flake ShardingBufferAdapterSpec: re-send the cold messages with a fresh probe per attempt under a sharding-derived budget; per-system settings

The first message through a fresh region has to survive coordinator
allocation and shard start, all ddata majority writes/reads that degrade
to "all nodes" on this test's 2-node cluster. If the shard's
remember-entities write stalls past its deadline, the Shard restarts and
its buffered messages die with it - no dead letter, nothing left to
re-deliver. Sharding is at-most-once, so a bare ExpectMsgAsync on the
flat akka.test.single-expect-default can time out waiting for a message
that no longer exists; waiting longer never helps.

FirstMessageThrough re-sends each of the three cold sends through
AwaitAssertAsync with a fresh TestProbe per attempt (so a late reply to
an abandoned attempt can't satisfy a later one, or leak into the warm
phase below), budgeted at ShardStartTimeout + UpdatingStateTimeout (the
cold path plus one full Shard restart cycle) with each attempt capped at
RetryInterval so the loop actually iterates. The counter snapshots are
taken after the loops, so BeGreaterOrEqualTo still holds under retries.

The warm-phase sends stay single sends (entities are live; a shard
restart here would mean no remember-entities write ever happened, so
retrying would hide a real bug) bounded by UpdatingStateTimeout instead
of the flat probe default, plus a LastSender check against the original
entity ref so a restart between phases can't hide behind the
assertions.

Also: StartShard was building both regions from Sys's settings instead
of the settings of the system whose region is being started (harmless
today since sysB is created from Sys.Settings.Config, but wrong); and
AfterAll shut Sys down twice (once directly, once via TestKit.Dispose's
finally) since _sysA is Sys - removed the redundant call and left a
TODO(#8545) for migrating both shutdowns to the async TestKit dispose
chain once that lands.

* ShardingBufferAdapterSpec: the identity check's comment names the right equality

(cherry picked from commit 808de09)
…DistributedPubSubRestartSpec (#8498)

* De-flake ClusterSingletonProxySpec and ClusterSingletonRestartSpec

Both specs build multi-node clusters in one process and only fail on loaded CI agents:
ClusterSingletonRestartSpec.Restarting_cluster_node_with_same_hostname_and_port_must_handover_to_next_oldest
on Windows/net10 (build 130838, at 15s) and
ClusterSingletonProxySpec.ClusterSingletonProxy_must_correctly_identify_the_singleton
on Linux/net10 (build 130872, ~36s in-test). Neither reproduces in isolation - 8/8 and 6/6 green
locally, and 30/30 and 24/24 green here under constrained thread pools and CPU pinning. So the work
is to take the load sensitivity out of the structure, not to widen anything. No timeout was raised,
no sleep and no backoff was added.

ClusterSingletonProxySpec

The identify test was synchronous throughout: TestProxy blocked on ExpectMsg for up to 25 seconds
per node, and the finally block sat in Task.WhenAll(...).Wait(30s). On a two-core agent the pool
starts at two worker threads and injects more slowly, so a blocked test thread is a large fraction
of the pool that the five ActorSystems need in order to gossip, heartbeat and deliver the very reply
being waited on. The test is now async: TestProxyAsync awaits ExpectMsgAsync and the teardown awaits
WhenAll. Awaiting the teardown also matters on its own - the discarded Wait(30s) result meant five
cluster systems could still be shutting down while the next test in the assembly ran.

The bigger correctness gap was that nothing waited for the singleton to exist. The first TestProxy
started before the cluster had formed, so cluster formation, singleton startup and proxy
identification all had to fit inside a message timeout. Two gates replace that: every node must see
all five members Up, and then each proxy must publish IdentifySingletonResult.Success. The subscription
is made in the ActorSys constructor before the proxy actor is created, so the event cannot be missed,
and the wait fishes for Success because the proxy also publishes Timeout results on a timer that
starts with no initial delay.

ClusterSingletonProxy_with_zero_buffering_should_work had the same gap with sharper teeth. It waited
for membership and then sent one message into a proxy configured with BufferSize 0, which drops
anything it cannot forward (ClusterSingletonProxy.Buffer). Membership does not imply identification,
so that send could vanish and the test would wait out its full 25 seconds for a reply that was never
coming. It now waits on the identification result. It also terminates its seed node, which was left
running for the remainder of the assembly.

Two waits in ClusterSingletonProxySingletonTimeoutTest2 read the identify result with
AwaitAssertAsync around ExpectMsgAsync, both unbounded. AwaitAssertAsync resolved
akka.test.single-expect-default (3s) and so did the inner ExpectMsgAsync on a different TestKit
instance, so the budgets did not nest and the loop got about one attempt - against a queue holding
the Timeout results the proxy emits every 500ms under that config. Both are now
FishForMessageAsync with an explicit 30s bound. The AwaitConditionAsync in
ClusterSingletonProxySingletonTimeoutTest gets an explicit bound for the same reason; with none it
inherited the 3s default for a two-node join.

ClusterSingletonRestartSpec

The spec is now async end to end. Three points of substance:

Shutdown(_sys1) waits Terminate().Wait(Dilated(5s)) and, when that expires, force-stops the user
guardian and logs a warning - verifySystemShutdown defaults to false, so nothing fails. That stops
/user but not /system, so remoting keeps sys1's listener bound; the very next statements create sys3
on that same host:port. dot-netty's tcp-reuse-addr is off-for-windows, which is the kind of
asymmetry that produces a Windows-only failure. sys1 also owns the singleton at that moment, so a
truncated shutdown cuts the hand-over to sys2 short. await _sys1.Terminate() lets CoordinatedShutdown
finish: the hand-over completes and the port is released.

Each proxy assertion created a TestProbe inside its retry loop. CreateTestProbe blocks the caller
until the probe's PreStart has run on the test-actor dispatcher (TestKitBase.cs:739) and replaces the
calling thread's SynchronizationContext (TestKitBase.cs:194) - so every attempt added a blocking wait
on work that needs a pool thread, a context swap and a leaked system actor. AwaitProxyReplyAsync
builds one probe per phase and retries only the send. Replies stay useful across attempts because the
message is a plain echo.

The 15s window that CI reported waits for sys2 to be fully Removed from sys3's view. Getting there
runs sys2's cluster-exiting CoordinatedShutdown phase, which blocks on the singleton hand-over
(ClusterSingletonManager.SetupCoordinatedShutdown) and gives up after 10s. JoinAsync only proves
sys3's own view of the cluster, so sys2 could still see sys3 as Joining when it left - no hand-over
target, phase runs to its timeout, and the removal no longer fits in 15s. A convergence gate now
requires both sys2 and sys3 to see the same two-member Up cluster before the Leave.

Smaller items: the Within wrappers are gone and each AwaitAssertAsync carries the same bound
explicitly, so no wait resolves an ambient deadline; the join assertion reads one Cluster.State
snapshot instead of two; the join retry runs at 500ms rather than 100ms, because a JoinTo arriving
while the daemon is in TryingToJoin drops it back to Uninitialized and restarts the handshake
(ClusterDaemon.cs:1273) - ten of those a second is load on the daemon the test is waiting for. The
join must still be re-issued, since sys3 reuses sys1's address and is refused until the old
incarnation is removed; a single Join would then wait out retry-unsuccessful-join-after (10s).
sys1/sys2/sys3 logs now reach the test output via InitializeLogger, as in
ClusterSingletonRestart2Spec.

Verified: both specs green in constrained-pool loops (DOTNET_ThreadPool_ForceMaxWorkerThreads=3),
full Akka.Cluster.Tools.Tests suite green, build clean with -warnaserror.

* De-flake DistributedPubSubRestartSpec: resolve the association before asserting

first's restart-kill loop asserted on ActorIdentity.Subject after a 2s Identify
window. Use ActorSelection.ResolveOne instead, inside the same retry: the Identify
round trip is the delivery confirmation, its temp actor is fresh per attempt, and
a final failure raises ActorNotFoundException naming the path that never resolved
rather than a bare timeout on a null Subject.

Resolve and kill stay in one retry on purpose. Splitting them lets artery's
ordinary outbound stream drop the kill in the gap between a successful resolve
and the Tell, with no resend - a local artery soak reproduced exactly that.

Also:
- read the baseline DeltaCount on a probe with an explicit bound instead of on
  TestActor, which is subscribed to topic1
- name the same-port rebind dependency when the old system fails to stop inside
  third's WhenTerminated wait
- fresh probe and an explicit 1s bound per Count attempt, so the 10s gossip window
  retries on the 500ms tick instead of spending itself on three 3s waits

(cherry picked from commit 9b0fc9f)
…d-off (#8500)

`ClusterShardingSpec` (base class for the Persistent/DData/WithEntityRecovery
variants) failed in CI on both transports with the same shape: one node timed
out waiting for an `Int32` counter reply and every other node then failed the
next barrier.

Mechanism. `HandOffStopper` replies `ShardStopped` the moment the last entity
terminates - before the `Shard` actor stops, and before the `ShardRegion`
observes `Terminated` and drops the shard from `_shards`/`_regionByShard`. The
spec treats `ShardStopped` as "the shard is gone" and immediately sends a `Get`
through the region, so the region forwards it into the dying shard:

  [.../RememberCounterEntitiesRegion/1] DeadLetter from [.../system/testActor1]
    to [.../RememberCounterEntitiesRegion/1]: Get { CounterId = 13 }

Sharding is at-most-once, so the request has to be re-sent, not merely
re-awaited. Two sites did not do that:

* `recover_entities_upon_restart` wrapped `Get(13)` in `AwaitAssertAsync(5s)`
  but the inner `ExpectMsgAsync(0)` carried no bound, so it inherited
  `akka.test.single-expect-default` (5s) - the entire budget. The loop made
  exactly one attempt:

    AwaitAssert failed, timeout [00:00:05] is over after [1] attempts
      and [00:00:05.0055266] elapsed time

* `permanently_stop_entities_which_passivate` sent `Get(25)` as a bare
  one-shot, so a single dropped message was an unconditional failure.

Both now re-send with a fresh probe per attempt and a 1s per-attempt bound, so
the 5s budget admits four sends instead of one. A fresh probe per attempt also
keeps a late reply from a timed-out attempt out of the next attempt's queue;
the same defect made the `Identify(4)` retry loop in the second phase a no-op,
and it is bounded now too.

The second phase also ran its three barriers inside `WithinAsync(15s)`.
`EnterBarrierAsync` derives its timeout from `RemainingOr(barrier-timeout)`, so
the rendezvous got the Within remainder instead of the configured 70s:

  EnterBarrier(Name: after-13, Role: [RoleName(sixth)], Timeout:00:00:13.8157551)
  timeout while waiting for barrier 'after-13'

The Within is removed, matching the treatment already applied to
`recover_entities_upon_restart`; its 15s is re-attached to the two waits it was
actually protecting, so no wait ends up with a smaller budget than before.

Verified by widening the hand-off window in a local throwaway build (delaying
the Shard's stop after `ShardStopped`), which reproduces both CI signatures
exactly - `Timeout 00:00:05 ... System.Int32` with `barrier failed:after-shard-restart`,
and `Timeout 00:00:13.9 ... System.Int32` with `barrier failed:after-13` - and
passes with these changes.

Test-only change.

(cherry picked from commit e6e014d)
… DistributedPubSubRestartSpec baseline (#8509)

* Pin the legacy allocation strategy in ClusterShardingLeavingSpec

The spec asserts that entities on surviving nodes keep the same actor
incarnation across a graceful leave. Only the legacy threshold strategy
has that property. The bounded default may move survivor shards during
the leave window, which restarts their entities under new refs - that is
the optimizer working as designed, not a regression, and the reference
implementation behaves the same way. Two CI failures carried this exact
signature: the DData variant in build 130784 and the Persistent variant
in build 130956, both an entity ref identity mismatch followed by an
after-4 barrier cascade.

Pinning rebalance-absolute-limit = 0 selects the legacy strategy, so the
identity assertion tests the algorithm that guarantees it. The spec
becomes a regression test for that strategy's leave behavior instead of
a coin flip against the bounded default.

Closes #8481.

* De-flake ReplicatorChaosSpec: rendezvous the replicators before the write burst

`ReplicatorChaosSpec` failed three times across two CI builds, on the Artery
Linux and Artery Windows lanes, with one fingerprint every time:

  [Node1:first]  Timeout (00:00:03) while expecting 15 messages. Only got 10
  [Node2:second] Timeout 00:00:00.0000044 while waiting for a message of type
                 Akka.DistributedData.GetSuccess
  [Node3..5]     barrier failed:update-during-split-verified

Root cause. Nothing makes the five replicators ready at the same time. Each node
checks `ReplicaCount(5)` against its own replicator and then walks straight into
the update phase, so `first` can start writing while another replicator has not
yet processed `first`'s MemberUp. A replicator drops a Write whose sender it
does not know:

  [.../user/replicator] Ignoring message [Write] from
    [.../user/replicator/$a] unknown node [UniqueAddress: ...]

and sends no ack. `WriteAll` cannot recover from that. Under WriteAll every node
is primary, so `SendToSecondary` re-sends to nobody, and a GCounter update
leaves `WriteAggregator._delta` null, so the delta re-send does not run either.
One dropped Write costs the whole update. Build 130934 shows `first` issuing its
first Update at 40.136 and `fifth` receiving its Welcome at 40.977; `fifth`
ignored all five Writes 12ms after that, and all five KeyC updates timed out.
Every node now proves `ReplicaCount(5)` and meets at a new `replicas-ready`
barrier before anything is written.

Second defect, and the reason the failure read as "Only got 10" instead of
naming `UpdateTimeout`. `ReceiveN(15)` sat outside any `Within` and inherited
`akka.test.single-expect-default` - 3s, the same 3s as the `WriteAll` deadline
of the updates it collects. An aggregator answers no earlier than its own
deadline, counted from when the replicator picks the Update up, which is at or
after the `ReceiveN` call. The collect can therefore never see that answer; it
is a dead heat by construction. In build 130934 the aggregators replied at
43.226, roughly 70ms too late, and the ten replies that did arrive were exactly
the WriteLocal ones, which need no aggregator. Every expect that waits on a
`WriteTo`/`WriteAll` reply now gets `WriteReplyTimeout` - the aggregator
deadline plus one single-expect-default of slack, with the arithmetic stated in
a comment on the field.

Third defect. `AssertValue` ran `Within(10s)` around an `AwaitAssert` whose
inner `ExpectMsg<GetSuccess>()` carried no bound, so each attempt could consume
the whole remaining budget and the final attempt got almost none - the logged
`00:00:00.0000044`. That reports a value which never converged as a bogus
zero-budget timeout. `AssertDeleted` and the `ReplicaCount` loop had the same
shape. All three now use a fresh probe and a 1s bound per attempt, the treatment
`ClusterShardingSpec` got in #8500.

Also swept. The `TestConductor.Blackhole`, `PassThrough` and `Exit` calls used
`.Wait(timeout)` and discarded the returned bool, so a conductor round-trip
slower than the 1s cap let the test walk on as though the partition were already
installed. They are awaited now, bounded by the conductor's own 30s
query-timeout, and a failure raises instead of vanishing. The spec is converted
to the async TestKit API throughout, which removes the remaining `.Wait()`
calls.

Verified fail-first. A throwaway build that holds one replicator's membership
events back by 4s reproduces the CI fingerprint on all five nodes exactly, with
the same five ignored Writes; the same injection passes with these changes and
produces zero ignored Writes. Bounding `ReceiveN` on its own converts the
failure into the honest assertion `Expected 'UpdateSuccess' got
'UpdateSuccess','UpdateTimeout'`, and bounding the inner expect makes all five
nodes report `Assert.Equal() Failure: Values differ` where four of five
previously reported the zero-budget timeout.

Test-only change.

* De-flake DistributedPubSubRestartSpec: read the DeltaCount baseline after gossip goes quiet

Build 130927 (artery, Windows) failed on second with "Expected value to be 3L, but
found 4L" - one extra Delta arrived between the baseline read and the post-restart
read. This is a different signature from the one PR #8498 fixed on this spec, and it
is not caused by third's restart at all.

Mechanism, confirmed by instrumenting the mediator locally and reading the counter it
actually exports:

DeltaCount counts Delta MESSAGES received, not registry changes. The mediator's Delta
handler increments _deltaCount and only then checks _nodes.Contains(bucket.Owner), so
a Delta whose payload is thrown away is still counted. While second is still waiting
for third's MemberUp, first pushes third's bucket on every 500ms gossip tick and
second discards each push - all counted. Convergence therefore arrives as a burst,
and redundant Deltas trail it, because a peer keeps re-sending until second's next
outbound gossip tells it that second caught up.

CountAsync unblocks on the first push second actually merges, so it returns while the
burst is still draining. The old code read the baseline immediately after that. Over
20 local runs the last Delta landed just 104-1010ms before that read, against a 500ms
gossip tick - one delayed tick puts a trailing Delta on the far side of the read and
the invariance assertion fails through no fault of the restart.

So the assertion was right and the baseline was wrong: "no new Deltas during third's
restart" only holds if the baseline is taken once gossip has settled. ReadStableDeltaCountAsync
now requires two samples 2s apart to agree before accepting a baseline. 2s is 4 ticks
of this spec's 500ms gossip-interval - enough for our outbound tick plus the peer's
reaction, with slack. Fresh probe per sample, matching CountAsync, so a late reply
cannot be misread as the next sample.

Verified by fault injection rather than by luck. Delivering exactly one extra Delta to
the mediator 1s after the baseline query - precisely the event the CI log implies -
reproduces the reported failure verbatim on the old code ("Expected value to be 3L,
but found 4L" on second) and passes on the new code, where the gate re-samples and
absorbs it. The post-fix margin between the last Delta and the baseline read is 3s
instead of ~150ms.

Soak: artery 8/8 under CPU load, classic 4/4. No timeout was raised, no sleep and no
backoff added; the only change is when the baseline is sampled.

(cherry picked from commit e340146)
… contact, and the survivor waits on it (#8552)

* De-flake DistributedPubSubRestartSpec: the restarted node makes first contact, and the survivor waits on it

first's association to third's old incarnation carries no signal that a same-address restart
even started; the survivor's association only learns the new incarnation exists from a frame
the restarted node itself sends. The old test started its 45s clock at Shutdown()'s return,
before any of third's restart cost had begun, and none of the ten prior fixes changed who makes
first contact after the restart - they all kept first polling a stale association instead.

Now third's Shutdown actor pings a forwarder on first the moment it exists (PreStart), so first
waits on that ping (dilated 60s) before running a short 20s closed-loop kill over
ActorSelection.Tell (only an ActorSelectionMessage pierces a quarantined association). The
ping's own inbound handshake is what heals first's association as a side effect. barrier-timeout
goes to 180s to fit the new worst case with headroom, and a bound-address assertion on third
settles, on the next Linux failure, whether the fresh listener came up on the pinned port.

* DistributedPubSubRestartSpec: repeat the ready ping until acknowledged, use the conductor's async shutdown, keep the port assertion inside the cleanup

Addresses review findings on PR #8552:
- Shutdown.PreStart no longer sends the ready ping once; it resends every
  500ms on an undilated self-scheduled timer until first's ReadyPingForwarder
  acks it back to the sender, cancelling the timer on ack (and defensively in
  PostStop). first still waits once for the first ping it sees and ignores
  later resends. This removes the dependency on any Artery handshake-stage
  fix - a lost ping is simply retried instead of failing the 60s wait.
- Replace TestConductor.Shutdown(...) with the ShutdownAsync(RoleName, ...)
  overload, keeping the same 30s bound.
- Move the restarted-address port assertion inside the try/finally that owns
  newSystem, so a failed assertion still terminates it.
- State the measured worst case behind the 60s ready-ping wait and re-derive
  newSystem.WhenTerminated's bound (120s -> 155s) against the current
  110s worst-case pipeline, restoring its original 45s of slack.

* DistributedPubSubRestartSpec: let the shutdown ack leave before the restarted system terminates

Build 131332 (this PR's own CI run, Linux Artery) showed third's Shutdown
actor replying "shutdown-ack" and calling Context.System.Terminate() inline
in the same handler; Artery aborted third's outbound streams before the ack
could flush, so it never reached first. First's closed-loop kill then
retried "shutdown" every 500ms against an already-exited process and its
20s bound timed out. Local runs only passed because the ack happened to win
that race on a fast machine.

Slide the terminate behind a 2s timer instead of calling it inline - same
shape as the Subject actor in RemoteNodeRestartDeathWatchSpec (PR #8557) -
so a live system, and the warmed-up lane under it, always outlasts whichever
"shutdown" attempt first's loop last sends. Cancel the timer in PostStop.

* DistributedPubSubRestartSpec: terminate timer outlives the retry cycle; restore the 120 s termination wait; last sync call migrated

(cherry picked from commit 2cc9d1e)
…ers the loop (#8542)

'second' and 'third' do no work of their own, so they reach the "after-1" barrier
about a second into the run while 'first' is still sending its 500 letters.
BarrierCoordinator arms a barrier's clock on the first arrival and only ever
shortens it, so the 30s default barrier-timeout - not the 5s per-letter wait - is
what bounds the whole loop. On a saturated CI agent the loop needs longer than
that, the barrier fails, the relay nodes stop, and 'first' then fails its next
per-letter wait because the route it was using has gone away.

EnterBarrierAsync asks the coordinator for RemainingOr(barrier-timeout), so a
Within around the final barrier sets that budget. 300s over 500 letters permits a
600ms mean round trip, about 3x the slowest mean measured on such an agent and far
below the 5s per-letter deadline, which remains the assertion that reports a real
drop.

Scoped to the barrier rather than set as akka.testconductor.barrier-timeout,
because that key also drives the teardown poll in MultiNodeSpecAfterAll, where
AwaitCondition polls at max/10 and the first poll always misses. Raising the key
to 300s adds 30s of wall clock to every run of this spec; measured here at 30s,
60s, 100s and 300s, the tax tracks barrier-timeout/10 exactly.

The loop, its route, its per-letter assertion and the letter count are unchanged,
so the spec proves exactly what it proved before. The rest of the diff is the
file's migration to the async TestKit API.

(cherry picked from commit e62aa19)
…the spec's own budget instead of a probe's flat 5 s default (#8544)

* De-flake ClusterShardingRegistrationCoordinatedShutdownSpec: wait on the spec's own budget instead of a probe's flat 5 s default

On the Artery lane the coordinator singleton hand-off from `second` to
`first` takes about 5.6 s: 5 s of that is a replicator on `first` that
never learned `second` was a cluster member (its subscription lands
after `second` is already `Leaving`), so the leaving coordinator burns
its whole `updating-state-timeout` on a ddata write that can never
reach quorum. The shutdown task's `probe.ExpectMsg(1)` only waited a
flat 5 s, because a `TestProbe` is its own `TestKitBase` with its own
deadline and never saw this spec's `Within(30s)`. The JVM spec sends
from its own test actor and waits on the enclosing `within`, so it
never had this gap.

Send with `_region.Value.Tell(1, TestActor)` and wait with
`ExpectMsg(1)` on the spec itself, so the wait inherits the spec's 30 s
budget. Pin `before-cluster-shutdown`'s phase timeout to 30s so an
incomplete task can't be cut short by the default 5s phase timeout
either.

* ClusterShardingRegistrationCoordinatedShutdownSpec: migrate to the async TestKit API

Moves Within, Join, RunOn, AwaitAssert, AwaitCondition, EnterBarrier, and ExpectMsg calls to their awaited Async variants; the coordinated-shutdown task on `third` becomes an async body, and the csTaskDone probe wait gets an explicit 20s dilated bound (measured 5.6s handoff, ~3.5x margin).

* ClusterShardingRegistrationCoordinatedShutdownSpec: bound the registration wait explicitly and size the outer window above its contents

* ClusterShardingRegistrationCoordinatedShutdownSpec: pass the explicit bounds undilated, since ExpectMsgAsync dilates them itself

(cherry picked from commit 9c75211)
…turn lane before sending End (#8563)

* De-flake ClusterDeathWatchSpec: resolve first's end actor over the return lane before sending End; migrate the file to the async TestKit API

On the Windows Artery lane, node fourth's fresh EndSystem sent End to
first's /user/end and waited 15s for EndAck. The ack has to ride a
brand-new outbound Artery lane from first to EndSystem, and first was
observed terminating its ActorSystem ~340ms after replying, before
that lane finished materializing, handshaking and connecting - so the
ack lost the race and the wait timed out.

Fix: before sending End, resolve first's /user/end from EndSystem with
ActorSelection.ResolveOne. The ActorIdentity reply can only travel back
over the same outbound lane the EndAck will use, so a successful
resolve proves that lane is already up. This is safe because first is
parked in ExpectMsgAsync<End>() and cannot start tearing down until the
End we have not sent yet arrives. Bound the resolve at a dilated 8s and
the subsequent EndAck wait at an explicit, undilated 5s - 8 + 5 = 13s,
strictly narrower than the 15s single-expect-default the probe used to
inherit unbounded. Also switch the EndSystem teardown to the async
ShutdownAsync instead of the blocking Shutdown.

While in the file, migrated every remaining synchronous TestKit call to
its async equivalent: the RemoteWatcher property became an async
GetRemoteWatcherAsync helper, TestLatch.Ready() became an
AwaitConditionAsync poll on TestLatch.IsOpen with the same 5s budget,
and the two remaining RunOn calls became RunOnAsync. No timeout was
widened and no assertion was loosened.

* ClusterDeathWatchSpec: state the shutdown bound honestly against the enclosing window; drop a dead null check

(cherry picked from commit a00df61)
…dicated probe, as upstream does; migrate the file to the async TestKit API (#8569)

Cluster.Leave triggers one MemberRemoved(self) EventStream publication that
drives two independent, unordered subscribers: the RegisterOnMemberRemoved
callback that tells the test actor "MemberRemoved", and the proxy's own
self-stop on seeing MemberRemoved for its node, whose death-watch Terminated
lands in whatever queue is watching it. When the test actor did the watching,
both messages competed for the same FIFO queue, so ExpectMsg("MemberRemoved")
could dequeue the Terminated meant for the following ExpectTerminated instead
(observed on the Linux Artery lane, node "first"). Watching echoProxy from a
dedicated TestProbe, as upstream Akka JVM/Pekko and the sibling
ClusterSingletonManagerLeaveSpec already do, gives Terminated its own queue
and removes the race without touching any timeout, barrier, or assertion.

The lazy proxy accessor becomes Lazy<Task<IActorRef>> so the watch can be
registered asynchronously while keeping create-on-first-use, single-execution
semantics. The rest of the file moves to the async TestKit API (RunOnAsync,
WithinAsync, AwaitAssertAsync, ExpectMsgAsync, EnterBarrierAsync) to support
that; the ForEach-based negative-assertion loop becomes a for loop, and its
Thread.Sleep(1000) becomes Task.Delay(1000), unchanged as a deliberate 1s
soak between probes rather than a stand-in for a real event.

(cherry picked from commit b2524f6)
…esend-derived bound, identify asserts the actor exists (#8581)

* De-flake RemoteNodeDeathWatchSpec: barrier after the subject stops so the watcher's wait starts on the real event; resend-derived bound; identify asserts the actor exists

Phase 1 had first waiting a flat 3s (single-expect-default) for WrappedTerminated
while second independently slept 3s before calling Sys.Stop, with no barrier
between the two sleeps and the stop. The window had to absorb both timers'
drift plus the remote notification hop, and a lost DeathWatchNotification is
only resent every 2s. second now does a local WatchAsync + Stop +
ExpectTerminatedAsync before entering a new "subject-stopped-1" barrier;
TellWatchersWeDied notifies remote watchers before local ones, so the local
Terminated proves the remote notification already left. first's wait now
opens after that barrier and uses an explicit 6s bound derived from
akka.remote.resend-interval (two resend cycles plus a hop), not the failure
detector (the peer stays alive and heartbeating here, so that path never
fires). Applied the same derived bound to the equivalent wait in phase 4,
which already had the correct barrier shape.

Restored the identify helper's dropped assertion (ActorIdentity.Subject must
not be null, naming the actor path) so a missing actor fails at the identify
instead of surfacing as an unrelated Ack timeout downstream.

Also migrated the file's remaining synchronous TestKit calls to their async
counterparts: the remote-watcher lookup, the identify helper, RunOn, and
ReceiveN, plus an explicit bound on AssertCleanup's previously-unbounded
nested expect.

* RemoteNodeDeathWatchSpec: name the real remote watcher in the comment, note Artery's resend interval, drop the AssertCleanup bound

The remote watcher recorded on the subject is first's /system/remote-watcher,
which re-issues the ProbeActor's watch over the wire; the comment named the
local watcher. The 6s bound reads the classic resend-interval; Artery resends
at 1s, so the same budget covers it. The 1s bound on AssertCleanup's nested
expect is removed: the predicate overload asserts rather than filters, so the
retry loop was already content-driven, and a timed-out attempt could offset
later requests from their replies.

(cherry picked from commit 1ad1abe)
…t pause over the measured stall, and make the payload listener match real log messages (#8556)

* De-flake NodeChurnSpec: turn off legacy auto-down, widen the heartbeat pause over the measured stall, and make the payload listener match real log messages

MultiNodeSpec.BaseConfig empties cluster.downing-provider-class, so this spec's
auto-down-unreachable-after setting always selected the legacy AutoDowning provider
(no quorum logic) rather than the keep-majority SBR its deprecation warning promises.
On build 131192 a transient-system teardown produced a 5.6s heartbeat gap on one
fixed node; the other two marked it unreachable, auto-downed it, and it auto-downed
them right back before its heartbeats resumed - a permanent split that made round 2
wait 25s for 9 members in a cluster that was 8 + 1. Turn auto-down off (this spec
already removes every transient member with an explicit Leave or Down and does not
need a downing provider) and widen acceptable-heartbeat-pause to 10s as a second
guard over the same measured stall, as a stopgap until the underlying product-side
blocking wait is fixed (issue #8549). Also add prune-gossip-tombstones-after = 1s,
matching Pekko, so gossip tombstones actually prune inside a 5-round run.

Separately, LogListener's `info.Message is string` guard could never match: a
parameterized Info log is wrapped in LogMessage<LogValues<T1,T2>>, never a string,
so the assertions this spec exists to make (gossip payload size growth) could never
fail. Match on info.Message?.ToString() instead, like RemoteMetricsSpec already does.

Also add Pekko's missing end-of-round barrier so a fast node cannot start creating
new transient ActorSystems while a slow peer still holds two dying ones, and replace
the stale comment above Task.WhenAll with the corrected explanation of why sequential
termination would not help: the scheduler's blocking stop is registered as the first
termination callback and therefore runs after the awaited task completes, whether
Terminate() calls are sequential or concurrent.

* NodeChurnSpec: match the .NET type name in the payload listener; fix the tombstone and auto-down comments; keep the teardown inside the barrier

(cherry picked from commit 986e473)
…te (#8404)

The reconnect loop retried inside a Within(5s) window that opens at the same unreachability event that gates the restarted address for retry-gate-closed-for = 5s, leaving ~0ms of usable budget — every retry dead-lettered inside the gate. Widen the window to 30s with a fresh probe and bounded expect per attempt, drop the gate to 1s in the spec config, and convert the spec to async TestKit methods.

(cherry picked from commit 2679bef)
@Aaronontheweb
Aaronontheweb force-pushed the backport/dev-2026-09-to-v1.5 branch from 46c8068 to afae265 Compare September 12, 2026 02:23
@Aaronontheweb

Copy link
Copy Markdown
Member Author

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

Aaronontheweb and others added 2 commits September 12, 2026 13:37
…ssSpec CI deadlock) (#8359)

* Akka.Cluster: make Cluster extension startup non-blocking

The Cluster constructor blocked on GetClusterCoreRef().Result (an Ask to
/system/cluster bounded by akka.actor.creation-timeout), and the internal
ClusterCore getter re-issued that blocking ask whenever it observed an
unresolved core ref. Actors spawned during construction (cluster event
subscribers, the SBR downing provider) hit the getter from PreStart and
parked dispatcher threads; on small pools the parked waiters starved the
very daemon that had to reply, the redundant asks timed out, and the
timeout path called Shutdown() on an already-healthy node. This is the
root cause of the StressSpec nondeterminism on 2-core Windows CI agents
and reproduces as a hard deadlock with a 1-thread default dispatcher.

- Replace the constructor ask with fire-and-forget
  InternalClusterAction.Init(cluster); the constructor is now
  straight-line code whose completion depends on no other thread
- ClusterDaemon and ClusterCoreSupervisor become explicit
  Uninitialized/Initialized state machines (Become + unbounded stash):
  pre-Init traffic is buffered, post-Init unmatched traffic forwards to
  the core daemon; ordering is guaranteed by mailbox FIFO, not blocking
- ClusterCoreSupervisor hands the core ref back via
  Cluster.SetClusterCoreRef(); the ClusterCore getter is now
  unconditionally non-blocking (_clusterCore ?? _clusterDaemons)
- Delete GetClusterCoreRef() and its Shutdown()-on-timeout path;
  akka.actor.creation-timeout is no longer consulted by Akka.Cluster;
  genuine core-startup failures still shut the node down via the
  existing ClusterCoreSupervisor supervision strategy
- Add ClusterStartupFuzzSpec: randomized startup fuzz under pinned
  1-2-thread dispatcher pools; killed 63-91% of iterations pre-fix and
  gates this class of regression at zero kills
- Record the behavioral change in BREAKING_CHANGES_V1.6.md
- dotnet format applied to the touched files (includes pre-existing
  whitespace cleanup; diff is unchanged with whitespace ignored)

Verification: fuzz spec 1,980/1,980 iterations reached Up (0 kills, 0
soft-timeouts) across baseline / 2-core / 1-core / CPU-saturated /
combined conditions; StressSpec passes 3/3 under taskset -c 0,1 where it
previously failed in ~90s with the CI signature; Akka.Cluster.Tests
379/379; Cluster.Tools 98/98; Cluster.Metrics 44/44; Sharding 191/192
(the 1 failure is pre-existing on dev, unrelated); no public API changes
(Akka.API.Tests 18/18).

* Address slopwatch SW003 findings in ClusterStartupFuzzSpec

Suppress six SW003 (empty catch block) findings in the fuzz spec's
teardown/best-effort paths via slopwatch-ignore comments with
justifications. These catches are intentional: cancellation on the
hard test deadline, best-effort probing of a cluster mid-race, and
bounded/best-effort teardown of burner tasks and the actor system.
No behavior change.

(cherry picked from commit 26e32af)
…n's lifetime so their order survives startup; LeaveAsync re-sends its Leave (#8580)

Cluster.ClusterCore used to switch from /system/cluster to a direct
reference to the resolved core daemon the instant that ref was
published during startup. Akka's FIFO guarantee holds only per
(sender, receiver) pair, so a Join and a Leave issued back-to-back by
the same caller with no await between them could ride different
mailboxes and arrive at the core daemon out of order. A Leave that
overtook its own JoinTo was dead-lettered by ClusterCoreDaemon's
Uninitialized behavior, and because LeaveSelf only ever sent its Leave
once, that loss was permanent. ClusterCore now always targets
/system/cluster, keeping the (sender, receiver) pair fixed for the
life of the extension; LeaveSelf also now re-sends its Leave command
on every call (Leaving(address) is idempotent), so a lost Leave no
longer wedges LeaveAsync forever.

Adds a regression fact in ClusterSpec covering Join immediately
followed by Leave with no await between them, and amends the existing

(cherry picked from commit 59d5da6)
@Aaronontheweb
Aaronontheweb force-pushed the backport/dev-2026-09-to-v1.5 branch from afae265 to d60e87d Compare September 12, 2026 13:57
@Aaronontheweb

Copy link
Copy Markdown
Member Author

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

… member-up signal (#8398)

Root cause: MemberUp cannot arrive before the 1s periodic-tasks-initial-delay
leader tick, but the spec waited on a raw non-dilated TestLatch.Ready(1s) --
a 100-400ms margin Windows CI load routinely eats. Also a non-dilated 1s
ResolveOne raced fire-and-forget cluster-extension init, and a fixed DotNetty
port (15508) could collide across processes on shared agents.

- ResolveOne wrapped in dilated AwaitAssertAsync (still fails if the downing
  provider is dead -- the #5962 regression signal)
- TestLatch -> TaskCompletionSource(RunContinuationsAsynchronously) awaited
  with a dilated timeout OUTSIDE the EventFilter block (no continuation
  inlining onto the member-listener's channel-executor thread)
- EventFilter.DeadLetter quiet window via the explicit dilated 2s overload
  (> SBR's 1s self-tick, so a dead resolver is still detected; test ~3s
  faster)
- port = 0 + explicit cluster.Join(SelfAddress) replaces fixed port/seed-nodes

Verified 50/50 consecutive Release runs (net10.0).

(cherry picked from commit 5f225ed)
Cherry-picking #8545 from dev recreated this file (its modify/delete conflict was resolved
toward dev's content so the pick stayed faithful to what it was on dev). v1.5 has no v1.6
release cycle and no equivalent ledger, so the file does not belong here. Removing it once, on
top of the cherry-picks, rather than dropping the hunk from the pick itself, keeps that
cherry-pick's diff identical to its dev counterpart.
DistributedPubSubRestartSpec.cs (backported from #8552) dual-pinned the
fresh ActorSystem's port for both transports, including
akka.remote.artery.canonical.port, and its surrounding comments explained
the fix using Artery-specific classes (OutboundHandshakeStage,
InboundHandshakeStage.HandleReq, ArteryRemoting.Send) and a CI run captured
on the Linux Artery lane. v1.5 has no Artery transport at all
(src/core/Akka.Remote/Artery does not exist on this branch), so the second
config line is dead HOCON and the class names in the comments point at code
that isn't there.

Dropped the artery.canonical.port line (only dot-netty.tcp.port is needed
on a classic-only branch) and reworded the affected comments to describe
the same reasoning in transport-neutral terms, without naming Artery
classes. No behavior change: the spec exercises the classic DotNetty
transport exactly as it did before this commit.

Checked every other file touched by this backport for "artery"/"Artery"
mentions; the remaining ones (ClusterDeathWatchSpec.cs, NodeChurnSpec.cs,
RemoteDeliverySpec.cs, RemoteNodeDeathWatchSpec.cs) only reference Artery
in passing, as comparison context for a classic-transport fix, and set no
Artery configuration - left unchanged.
#8499 added TestKitBase.ShutdownAsync (public virtual) and its protected
ActorSystem overload. The API-approval snapshots kept during that cherry-pick
were deliberately v1.5's own pre-existing files (not dev's - the two
branches' TestKit surfaces differ by roughly 300 lines elsewhere), so
CoreAPISpec.ApproveTestKit was failing after the pick: v1.5's public API
now has the two new methods but the approved snapshot did not.

Ran `dotnet test src/core/Akka.API.Tests --filter ApproveTestKit` and
accepted the resulting .received.txt as the new DotNet.verified.txt. The
only delta from the previous baseline is the two new ShutdownAsync
overloads and the compiler's incidental renumbering of nearby
async-iterator state-machine class names (adding members shifts Roslyn's
sequential ordinals for unrelated nested types declared after them; this is
expected and harmless).

The Net.verified.txt (net48/netstandard2.0) counterpart could not be
regenerated by running the net48 test host in this environment (the vstest
agent fails to negotiate over the .NET Framework runtime here). Derived it
instead: TestKitBase.cs has no #if/target-framework-conditional code, and
diffing the previous DotNet/Net verified pair showed they differed only in
the single assembly-level TargetFrameworkAttribute line - every other line,
including all compiler-generated class ordinals, was already identical.
Applied the same content with only that one line swapped back to
".NETStandard,Version=v2.0", and confirmed the result still differs from
the new DotNet file by exactly that one line, matching the pre-existing
pattern.

Verified: `dotnet test src/core/Akka.API.Tests --framework net10.0` -
18/18 passed, including both ApproveTestKit and ApproveTestKitXunit2.
…ard2.0 has no process-tree kill overload)

The #8515 cherry-pick's MultiNodeTestRunner.cs conflict was resolved by
taking dev's full file (needed to bring in #8431's node-exit-timeout
backstop alongside #8515's port-sentinel change, so the new
DefaultNodeExitTimeout field would actually be used - see that commit's
message). Building Akka.Cluster.Tests.MultiNode surfaced a genuine v1.5
incompatibility that dry-run text diffing could not catch: dev compiles
Akka.MultiNode.TestAdapter against net10.0, where `Process.Kill(bool
entireProcessTree)` exists, but this project's target on v1.5 is
netstandard2.0 (a single TFM), whose Process surface only has the
parameterless `Kill()`.

Switched to `process.Kill()`, matching the convention already used for the
same reason in Akka.MultiNode.RemoteHost/RemoteHost.cs. The only behavior
difference is that a hung node's child processes (if it spawned any) are no
longer force-killed alongside it - the backstop still kills the node
process itself and still unblocks the runner.
…wn (port of #8543)

Hand-port of dev's #8543 substance, not a cherry-pick - #8543 is the last commit of a five-deep
stack (#8372, #8427, #8429, #8499) and its auto-merged parts would not compile as-is on v1.5
(BuildConfig's arithmetic and the ClusterResultAggregatorAsync call sites are written against
dev's 10-node/1-per-phase config, which v1.5 does not have).

What changed, and why each piece is still correct on v1.5:

* acceptable-heartbeat-pause raised from 3s to 20s, with the arithmetic comment explaining the
  phi-accrual crossing-threshold math (1s + 20s + 3*0.1s = 21.3s of detection). v1.5's
  failure-detector defaults (heartbeat-interval=1s, min-std-deviation=100ms) and
  split-brain-resolver stable-after (10s) already match dev's, so the same 3s-was-too-tight
  problem applies here: a churn round abruptly tearing down an ActorSystem can starve the
  surviving node's own heartbeat sender for longer than a 3s pause tolerates.

* akka.test.single-expect-default = 10s and akka.test.timefactor = 3, ported from #8372 (the
  first commit of the same dev stack), which the earlier port of this commit had missed. Every
  Within/WithinAsync bound in this file is dilated by timefactor - including RemoveOneAsync's
  removal budget (TimeSpan.FromSeconds(25) + ConvergenceWithin(3s, NbrUsedRoles - 1), which
  #8543 never widened) - so raising acceptable-heartbeat-pause to 20s without also porting
  timefactor left that budget structurally unable to cover the ~34.3s abrupt-removal path this
  spec's own ChurnMemberRemovalWithin() arithmetic predicts at the node counts the 7-node CI
  phase sequence reaches. Without timefactor, RemoveOneAsync's ceiling at NbrUsedRoles=3 is a flat
  31s; with it, 93s. Confirmed against three consecutive 7-node runs: the abrupt-removal phase
  measured 32.9-33.4s in each, and the AwaitAssert measured a 30.98s give-up in each, both
  matching the undilated 31s ceiling to within noise. All three runs pass with timefactor ported.

* ChurnMemberRemovalWithin(), ported byte-for-byte from dev (it only calls
  Cluster.Settings.FailureDetectorConfig/HeartbeatInterval/GossipInterval, all present on v1.5).
  Computes how long an abruptly-terminated churn member takes to actually leave the ring:
  detection + stable-after + a leader-gossip margin = 34.3s at this spec's config.

* ExerciseJoinRemoveAsync's loopDuration now includes ChurnMemberRemovalWithin() so each round's
  Within budget covers the full removal path, not just the new join. The abrupt-shutdown behavior
  itself (ShutdownAsync, no cluster Leave) was already in place from #8499 - this keeps exactly
  what the phase proves (abrupt loss), it only fixes the budget around it.

* Async TestKit migration of the call sites #8543 touches: the two RunOn sends inside the churn
  Loop become RunOnAsync, and ClusterResultAggregator (a one-shot, non-retried Identify/ExpectMsg
  lookup) gains a ClusterResultAggregatorAsync sibling (fresh-probe-per-attempt, retried over 30s)
  ported from dev, since a lone lost reply under this phase's deliberate churn must not be fatal.
  Repointed CreateResultAggregatorAsync, AwaitClusterResultAsync, and the async ReportResult<T>
  overload - the three call chains ExerciseJoinRemoveAsync depends on - at the new async lookup.
  Every remaining call site in the file passes an async lambda with an explicit return statement,
  so all of them already bound to the async ReportResult<T>(Func<Task<T>>) overload; the sync
  ClusterResultAggregator() and the sync ReportResult<T>(Func<T>) overload it served had no
  callers left after that repointing and are removed here. (Not ported: dev's extra
  "result-aggregator-identified" barrier in CreateResultAggregatorAsync, a related but separate
  hardening against a PartitionSeveral race - open item below.)

* RemoveOneAsync's watchee lookup, hand-ported from dev (#8372/#8429): a fresh CreateTestProbe()
  per attempt instead of the shared IdentifyProbe (the shared probe kept a timed-out attempt's
  late ActorIdentity queued, so the next attempt consumed that stale reply, and a reply resolved
  before the watchee existed carries a null Subject that the retry could never recover from); an
  explicit identity.Subject.Should().NotBeNull() guard before WatchAsync; WatchAsync bounded with
  WaitAsync(Dilated(3s)) instead of inheriting the outer Within's RemainingOrDefault, since a
  hung/slow watch Ask would otherwise burn the whole retry budget in a single attempt; and an
  explicit AwaitAssertAsync(10s, 1.25s) bound instead of an unbounded retry loop. This is the exact
  method that was failing in the runs above, so it is ported alongside the budget fix rather than
  left as a separate follow-up.

* StressSpecConfig node-count env override already existed on v1.5 (MNTR_STRESSSPEC_NODECOUNT,
  default 13). What it lacked was BuildConfig's shrink arithmetic: below the reference count,
  phase sizes must shrink or Settings' constructor throws. v1.5's reference config is heavier
  than dev's - every joining phase defaults to 2 nodes here, not 1 - so reaching the same
  practical floor of 7 requires shrinking both the joining side (halve every 2 back to 1) and the
  leaving/shutdown side (drop the "-large" one-by-one phases, halve the simultaneous counts), the
  same technique dev's config uses on one more group of phases. Derived and documented in
  StressSpecConfigSpec.cs (ported alongside, adapted from dev's 10-node version to v1.5's
  13-node/2-per-phase defaults): 7 is confirmed the practical floor (6 throws because the joining
  phases alone need 7 regardless of how far leaving/shutdown shrinks).

Also fixed while building this: MultiNodeTestRunner.cs's Process.Kill(bool) call from the earlier
#8515 hand-merge doesn't compile against netstandard2.0 - see the preceding commit.

Verified: dotnet build src/core/Akka.Cluster.Tests.MultiNode -warnaserror clean;
StressSpecConfigSpec 9/9; StressSpec run three times at MNTR_STRESSSPEC_NODECOUNT=7, all three
passed (see the PR body for full timings).

Open items for the maintainer:
- dev's extra CreateResultAggregatorAsync barrier (guards a PartitionSeveral aggregator-
  identification race) was not ported; the retrying ClusterResultAggregatorAsync lookup narrows
  that race but does not close it the way the extra barrier does.
…ressSpec node-count doc

Three comment-only corrections, none change behavior:

* ClusterSpec.cs (from the #8579 pick): the comment on the MemberUp wait described dev's
  Cluster.ClusterCore re-targeting from a supervisor to the core daemon when the ref is
  published - a mechanism that no longer matches this branch's own code now that #8359 and
  #8580 are in. ClusterCore always targets /system/cluster for the life of the extension, so
  two commands from the same sender are already delivered in order without this wait; the
  MemberUp wait is still the stronger barrier, since it proves the join was processed, not
  merely enqueued, before Leave is sent. Reworded to say so.

* DistributedPubSubRestartSpec.cs (from the earlier Artery-removal tailoring commit): the
  reworded comment claimed a plain Tell to a quarantined peer is dropped at the transport layer.
  On v1.5's classic remoting it is not: EndpointManager's Quarantined case creates a brand-new
  writing endpoint for any Send that reaches it. It is the separate Gated policy that dead-letters
  a Send while its release deadline is unexpired. Reworded to describe what EndpointManager.cs
  actually does, and to explain ActorSelection.Tell's real advantage here (re-resolving by path on
  every send reaches whichever incarnation is live, rather than a stale cached ref).

* StressSpec.cs BuildConfig doc comment: said 13 was "also the smallest count that fits every
  phase" and then derived >= 11 two sentences later - contradicting itself. Reworded to say what
  the arithmetic says: 11 is the true minimum at full phase size; 13 is v1.5's original default,
  kept as the boundary below which this method starts shrinking phases.

Verified: dotnet build on Akka.Cluster.Tests, Akka.Cluster.Tools.Tests.MultiNode, and
Akka.Cluster.Tests.MultiNode -warnaserror, all clean.
…pec at 7 nodes

Mirrors dev's #8553. build-system/azure-pipeline.mntr-template.yaml gained
no way to pass extra environment variables into the test-execution step,
so build-system/pr-validation.yaml's single MNTR lane (net_mntr_windows)
could not set MNTR_STRESSSPEC_NODECOUNT even though StressSpec.cs already
reads that variable.

Added an `env` parameter (defaulting to `{}`, matching dev's template) to
the MNTR template and wired it into the `${{ parameters.command }}` step,
then set `MNTR_STRESSSPEC_NODECOUNT: "7"` on the v1.5 MNTR lane so
StressSpec runs at 7 nodes instead of its 13-node default on the 2-vCPU
hosted agents. 7 is validated as a safe floor by the preceding StressSpec
commit's StressSpecConfigSpec tests (StressSpecConfig.BuildConfig shrinks
every phase to fit at exactly 7 nodes; 6 still throws).
#8580 is now fully in this branch (both the ClusterCore routing fix and the LeaveSelf re-send
fix cherry-picked earlier), so this commit is no longer a carve-out - it only adds the one
ClusterSpec fact dev's own #8580 pick didn't need, because dev's ported fact
(A_cluster_must_process_a_Leave_issued_immediately_after_Join_from_the_same_thread) exercises
the ClusterCore-routing race, not the LeaveSelf re-send: it calls Leave() exactly once, and that
first call behaves identically before and after the re-send fix.

Adds A_cluster_must_resend_Leave_on_a_later_LeaveAsync_call_after_an_earlier_send_was_lost,
which does discriminate: LeaveAsync() before any Join reaches ClusterCoreDaemon's Uninitialized
behavior (via the pre-Init stash in ClusterDaemon, which buffers traffic until its own Init
completes, then forwards it down), which has no case for ClusterUserAction.Leave, so the command
is unhandled and dropped; then Join(self), LeaderActions(), await MemberUp via a subscription,
then LeaveAsync() again and assert the returned task completes. Verified by temporarily
reverting LeaveSelf() to send Leave only once (its pre-fix shape): this fact times out after 10s
on that code and passes in under a second on the fix.

Review follow-up (Opus review, findings 8 and 9): the fact now asserts its own premise instead of
assuming it - a probe subscribes to UnhandledMessage on the EventStream before the first
LeaveAsync() and asserts one arrives whose Message is a ClusterUserAction.Leave addressed to
.../system/cluster/core/daemon, so a future change that gives Uninitialized a Leave case fails
this fact here instead of leaving it silently non-discriminating. The final WaitAsync bound is
now Dilated(10s) instead of a raw TimeSpan, consistent with the sibling facts' RemainingOrDefault/
dilated ExpectMsgAsync usage (equivalent to the old value today, since this spec sets no
akka.test.timefactor).

Verified: `dotnet build src/core/Akka.Cluster.Tests -warnaserror` clean; `dotnet test
src/core/Akka.Cluster.Tests --filter ClusterSpec` run 3 times - all facts passed every time,
including both Leave facts. Re-confirmed the fact still discriminates after this change: with
LeaveSelf() temporarily reverted to its pre-fix one-shot send, the fact fails with a
TimeoutException (the dead-lettered/unhandled first Leave is visible in the log); restoring
Cluster.cs returns it to passing.
Adds a `1.5.72 TBD` section covering every user-visible change in this PR, including the
non-blocking Cluster startup fix and a migration note for #8574's write-timeout key change.

Review follow-up (Opus review, findings 2, 3, 5, 10, 11):
- The #8359 bullet now states outright that the cluster's internal actor tree
  (/system/cluster/core, /system/cluster/core/daemon, /system/cluster/heartbeatReceiver) is
  created after Cluster.Get() returns rather than before it - the one sharp edge of this change,
  and previously the only thing this note left unsaid - and says code resolving those paths
  right after Cluster.Get() should await JoinAsync/RegisterOnMemberUp or a cluster event instead.
- Says outright that akka.actor.creation-timeout is no longer consulted by Akka.Cluster at all
  (stronger and more accurate than "no longer blocks on"), while noting Akka.Cluster.Sharding
  still uses it to bound its own start-up asks, so nobody reads this as license to revert that
  setting.
- Replaced "a small dedicated dispatcher pool" with wording that also covers a starved default
  pool, since the deadlock reproduces on either.
- The #8580 bullet no longer says "during startup" - the routing through /system/cluster is
  permanent for the extension's lifetime (two extra local mailbox hops on the low-traffic command
  path), and now says a `Leave`/`LeaveAsync()` call issued after the node has already shut down
  produces a dead-lettered ClusterUserAction.Leave, relevant to dead-letter alerting.
- Adds a Testing bullet for #8398 (De-flake Bugfix5962Spec), the dev de-flake this branch had been
  missing for the spec #8359 makes newly racy.
The #8359 cherry-pick brought in ClusterStartupFuzzSpec.cs, which uses
Task.IsCompletedSuccessfully at two call sites, and the #8398 cherry-pick (5f225ed, applied
earlier in this branch as the Opus-review follow-up for finding 1) uses the non-generic
`new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously)` /
`TrySetResult()` in Bugfix5962Spec.cs. On dev, Akka.Cluster.Tests no longer targets net48, so
both compiled fine there. On v1.5, Akka.Cluster.Tests still targets net48 alongside net10.0:
`Task.IsCompletedSuccessfully` does not exist on .NET Framework 4.8 (CI build 131554 failed with
CS1061 on every net48 lane), and the non-generic `TaskCompletionSource` does not exist before
.NET 5 (CS0305).

- ClusterStartupFuzzSpec.cs: switched both call sites to
  `joinTask.Status == TaskStatus.RanToCompletion`, the exact equivalent available on all target
  frameworks (the `using System.Threading.Tasks;` needed for `TaskStatus` was already present).
- Bugfix5962Spec.cs: switched to `TaskCompletionSource<Done>` / `TrySetResult(Done.Instance)`,
  the same pattern already used by ClusterDaemon's own `_clusterPromise`/`_selfExiting` fields -
  available on every target framework this project builds for. No behavior change: `Done`
  carries no data, so the task's completion is still the only thing the spec observes.

Verified:
- `dotnet build src/core/Akka.Cluster.Tests -c Release -f net48 -warnaserror` and
  `-f net10.0 -warnaserror`: both 0 warnings/errors.
- `dotnet test -f net10.0 --filter FullyQualifiedName~ClusterStartupFuzzSpec` passes both tests:
  aggressive run completed=72, kills(delayed=0, immediate=0), pass=72 (reachedUp=72); moderate
  run completed=60, kills(delayed=0, immediate=0), pass=60 (reachedUp=60).
- `dotnet test -f net10.0 --filter FullyQualifiedName~Bugfix5962Spec`: passes.
@Aaronontheweb
Aaronontheweb force-pushed the backport/dev-2026-09-to-v1.5 branch from 9152ed7 to 8ab2713 Compare September 12, 2026 16:18
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
…s the new oldest, so a lost HandOverInProgress is named instead of timed out

In PR #8588 build 131555 (Windows unit lane), this spec failed with only
"Timeout 00:00:01 while waiting for a message of type System.String" from
the final AwaitProxyReplyAsync(_sys3, proxy3, "hello3", 5s). The real cause
was upstream: sys2 wrote HandOverInProgress, HandOverDone and
ExitingConfirmed to the socket before closing it during CoordinatedShutdown,
but on Windows a queued inbound frame from sys3 turned the close into a TCP
RST, and the RST made sys3's stack discard the bytes it had not yet read
(see #8589). sys3 never saw the hand-over confirmation, never cancelled its
retry timer, and fell back to the removal-margin / hand-over-retries path -
which does recover the singleton, but only after 20s or more, well outside
the 5s budget the test is actually checking.

Wrap both hand-overs (sys1 -> sys2 and sys2 -> sys3) in an EventFilter on
"Hand-over in progress at" (ClusterSingletonManager.cs, logged by the new
oldest when it receives HandOverToMe's confirmation) so a lost confirmation
fails the fact at the point the hand-over actually broke, with the filter's
own message, instead of surfacing as an unrelated-looking timeout later.

No timeout, window, or config value is changed. The 5s AwaitProxyReplyAsync
budget is deliberately left alone: the loss path needs roughly 21s or more
to resolve (6s failure-detector + 20s removal margin, or ~27s via
hand-over-retry exhaustion and a manager crash-restart), and widening the
budget to cover it would only make the test pass on a path where the
hand-over failed and the manager had to crash to recover - defeating the
point of a spec named after the hand-over.

Verification:
- dotnet build src/contrib/cluster/Akka.Cluster.Tools.Tests -c Release
  -f net10.0 -warnaserror -p:SuppressTfmSupportBuildWarnings=true: clean
  (project targets only net10.0 on dev)
- Spec run 10x: dotnet test ... --filter
  FullyQualifiedName~ClusterSingletonRestartSpec: 10/10 passed, ~6-7s each
- Negative check: temporarily changed the sys3 filter's start: text to
  "Hand-over in progress at NOWHERE" (a string that never logs); the fact
  failed with "Timeout (00:00:03) while waiting for messages. Only received
  0/1 messages that matched filter [Info when Message starts with
  'Hand-over in progress at NOWHERE']" - confirming the filter discriminates.
  Restored the real text; dotnet build and a follow-up 10x run both green
  again, and git diff shows only the intended change.
- dotnet format --verify-no-changes on the touched file flags one
  pre-existing, unrelated whitespace issue ("if(_sys3 != null)" at the
  original line 166, in AfterAll) that predates this change; left untouched
  to keep the diff focused.
…0 unit-test host hangs

The Linux ".NET Unit Tests" lane hangs inside src/core/Akka.Tests now and then, with
no test named and no artifact to read: Incrementalist (.incrementalist/testsOnly.json,
timeoutMinutes: 20) cancels the dotnet test command at its 20-minute per-project cap
before vstest can write a trx. Seen on v1.5 build 130811 (2026-08-26, before this PR),
and again on this PR's own builds 131545 and 131555 - each time going silent about 75s
into the project with nothing pointing at the stuck test.

Append vstest's blame-hang flags (--blame-hang --blame-hang-timeout 15m
--blame-hang-dump-type mini) to the two net10.0 unit-test commands in
build-system/pr-validation.yaml (the Windows and Linux ".NET Unit Tests" lanes, both
using .incrementalist/testsOnly.json). When a test goes quiet for 15 minutes, vstest
writes a Sequence_<guid>.xml naming the tests that started and did not finish, plus a
minidump of the test host and its children, into TestResults - which the lane already
publishes as a build artifact - then kills the host so the lane fails with the test
named instead of just timing out silently.

15 minutes sits below Incrementalist's 20-minute cap, so vstest's own hang timer fires
and writes its artifacts before Incrementalist kills the process tree; no single test
in these projects legitimately runs anywhere near 15 minutes (the whole Akka.Tests
project takes about 6-7 minutes on CI).

Left out:
- The net48 lane (.incrementalist/testsOnlyNetFx.json): vstest's hang-dump collector
  needs procdump on the agent for a full-framework test host, which these agents don't
  have.
- The multi-node lane (.incrementalist/mutliNodeOnly.json): its node processes are
  child processes the MNTR runner already manages and tears down itself.

Verified locally on this SDK (dotnet 10.0.103, Linux): built
src/core/Akka.Tests/Akka.Tests.csproj for net10.0, then ran
Akka.Tests.Actor.TimerSpec.Must_replace_timer (a ~3s test) with a deliberately tiny
--blame-hang-timeout to confirm the hang path fires:
- 1s timeout: fired during test-host startup/discovery, before the test itself began
  running - produced two hang-dump .dmp files (test host + child process) but no
  Sequence file, since no test was yet in progress when the collector ran.
- 2s timeout: fired mid-test - produced Sequence_<guid>.xml with
  <Test Name="Akka.Tests.Actor.TimerSpec.Must_replace_timer" Completed="False" />,
  plus the same two .dmp files, confirming the mechanism names the actual stuck test
  once one is running.
- 15m timeout: passed normally in ~3s, no Sequence file or dumps produced - confirms
  the flags are inert on a healthy run.
@Aaronontheweb

Copy link
Copy Markdown
Member Author

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@Aaronontheweb
Aaronontheweb merged commit 9d659bd into v1.5 Sep 12, 2026
12 checks passed
@Aaronontheweb
Aaronontheweb deleted the backport/dev-2026-09-to-v1.5 branch September 12, 2026 19:16
Aaronontheweb added a commit that referenced this pull request Sep 21, 2026
…s the new oldest, so a lost HandOverInProgress is named instead of timed out (#8590)

In PR #8588 build 131555 (Windows unit lane), this spec failed with only
"Timeout 00:00:01 while waiting for a message of type System.String" from
the final AwaitProxyReplyAsync(_sys3, proxy3, "hello3", 5s). The real cause
was upstream: sys2 wrote HandOverInProgress, HandOverDone and
ExitingConfirmed to the socket before closing it during CoordinatedShutdown,
but on Windows a queued inbound frame from sys3 turned the close into a TCP
RST, and the RST made sys3's stack discard the bytes it had not yet read
(see #8589). sys3 never saw the hand-over confirmation, never cancelled its
retry timer, and fell back to the removal-margin / hand-over-retries path -
which does recover the singleton, but only after 20s or more, well outside
the 5s budget the test is actually checking.

Wrap both hand-overs (sys1 -> sys2 and sys2 -> sys3) in an EventFilter on
"Hand-over in progress at" (ClusterSingletonManager.cs, logged by the new
oldest when it receives HandOverToMe's confirmation) so a lost confirmation
fails the fact at the point the hand-over actually broke, with the filter's
own message, instead of surfacing as an unrelated-looking timeout later.

No timeout, window, or config value is changed. The 5s AwaitProxyReplyAsync
budget is deliberately left alone: the loss path needs roughly 21s or more
to resolve (6s failure-detector + 20s removal margin, or ~27s via
hand-over-retry exhaustion and a manager crash-restart), and widening the
budget to cover it would only make the test pass on a path where the
hand-over failed and the manager had to crash to recover - defeating the
point of a spec named after the hand-over.

Verification:
- dotnet build src/contrib/cluster/Akka.Cluster.Tools.Tests -c Release
  -f net10.0 -warnaserror -p:SuppressTfmSupportBuildWarnings=true: clean
  (project targets only net10.0 on dev)
- Spec run 10x: dotnet test ... --filter
  FullyQualifiedName~ClusterSingletonRestartSpec: 10/10 passed, ~6-7s each
- Negative check: temporarily changed the sys3 filter's start: text to
  "Hand-over in progress at NOWHERE" (a string that never logs); the fact
  failed with "Timeout (00:00:03) while waiting for messages. Only received
  0/1 messages that matched filter [Info when Message starts with
  'Hand-over in progress at NOWHERE']" - confirming the filter discriminates.
  Restored the real text; dotnet build and a follow-up 10x run both green
  again, and git diff shows only the intended change.
- dotnet format --verify-no-changes on the touched file flags one
  pre-existing, unrelated whitespace issue ("if(_sys3 != null)" at the
  original line 166, in AfterAll) that predates this change; left untouched
  to keep the diff focused.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant