Skip to content

Akka.Cluster: make Cluster extension startup non-blocking (fixes StressSpec CI deadlock) - #8359

Merged
Aaronontheweb merged 3 commits into
akkadotnet:devfrom
Aaronontheweb:fix/cluster-nonblocking-startup
Jul 10, 2026
Merged

Aaronontheweb merged 3 commits into
akkadotnet:devfrom
Aaronontheweb:fix/cluster-nonblocking-startup

Conversation

@Aaronontheweb

Copy link
Copy Markdown
Member

Root cause: a decade-old startup deadlock disguised as CI flakiness

StressSpec failed nondeterministically on the 2-processor Windows CI agents (e.g. builds 128865 and 128869 on #8352's PR validation — different victim node and join phase each time, identical signature). Investigation traced it to Cluster extension startup, not the spec:

  1. The Cluster constructor blocks on GetClusterCoreRef().Result — an Ask to /system/cluster bounded by akka.actor.creation-timeout (20s, not test-dilated).
  2. The internal ClusterCore property getter re-issued that blocking ask whenever it observed _clusterCore == null. Every call site reads as fire-and-forget (ClusterCore.Tell(...)), but the block hides inside the property access — resolving the target of the Tell.
  3. Actors spawned during construction (the read-view's event-bus listener, the SBR downing provider) call Cluster.Subscribe from PreStart, hit the getter mid-window, and park dispatcher threads on redundant 20s asks.
  4. On small pools, the parked waiters starve the very ClusterDaemon/ClusterCoreSupervisor that must reply. The asks time out at exactly 20.000s, and the timeout path calls Shutdown() — terminating a node whose cluster had already started successfully ("Started up successfully" appears in the CI logs ~70–190ms before the fatal timeouts). JoinAsync then throws ClusterJoinFailedException: Cluster has already been terminated and the other 12 StressSpec nodes cascade.

This is a hard deadlock, not load flakiness: a purpose-built fuzz spec kills current dev 3-for-3 on an idle machine with a 1-thread default dispatcher, and 63–91% of iterations under randomized 1–2-thread pools. The 2-core CI agents merely sample it probabilistically. Both ClusterDaemon and ClusterCoreSupervisor have carried warning comments about this exact constructor deadlock for years; partial fixes (reordering construction, single-flighting the ask through a shared task) were prototyped and empirically killed by the same fuzz harness — any design where startup completion depends on another thread being scheduled retains the failure class.

The change

The constructor stops asking. Startup coordination moves from blocking to mailbox causality:

  • InternalClusterAction.Init(cluster) (fire-and-forget) replaces the GetClusterCoreRef ask; the constructor is now straight-line code whose completion depends on no other thread. The extension-registration Lazy window collapses to microseconds, which also dissolves the Error when using ChannelTaskScheduler for internal-dispatcher in akka cluster node #5498-style re-entrancy hazards (Cluster.Get during startup can no longer park on a constructor that's waiting for a message).
  • ClusterDaemon and ClusterCoreSupervisor become explicit Uninitialized → Initialized state machines (Become + unbounded stash): pre-Init traffic is buffered, post-Init unmatched traffic forwards down to the core daemon. Init is provably the first message in the daemon's mailbox, so forwarded messages can never race the core tree's creation; UnstashAll preserves arrival order. A PreRestart re-sends Init so a restarted daemon re-initializes.
  • ClusterCoreSupervisor hands the core ref back via Cluster.SetClusterCoreRef(); the ClusterCore getter is now unconditionally non-blocking (_clusterCore ?? _clusterDaemons) — the forwarding hop exists only during the ms-scale window.
  • GetClusterCoreRef() and its Shutdown()-on-timeout path are deleted. akka.actor.creation-timeout is no longer consulted by Akka.Cluster. Genuine core-startup failures are unchanged: ClusterCoreSupervisor's existing supervision strategy (stop + _cluster?.Shutdown()) already owned that policy — the timer path was redundant for real failures and harmful for spurious ones.
  • ClusterStartupFuzzSpec (new) is the regression gate: randomized startup hammering under pinned 1–2-thread dispatcher pools, seed-reproducible (AKKA_FUZZ_SEED). It killed 63–91% of iterations pre-fix and killed both rejected prototype fixes — only a genuinely non-blocking startup passes it.

Behavioral note (recorded in BREAKING_CHANGES_V1.6.md): Cluster.Get() now returns immediately with core initialization completing asynchronously (previously it blocked up to 20s). All public APIs are message-based or event-driven and unaffected; code treating Cluster.Get() as a "core is started" synchronization point should await JoinAsync/RegisterOnMemberUp instead. No public API surface changes (Akka.API.Tests green, everything added is internal). Wire format untouched. Deliberate divergence from JVM Akka, which still blocks in its constructor; protocols, actor paths, and child structure remain aligned.

Verification

Gate Pre-fix baseline (measured) Post-fix
ClusterStartupFuzzSpec, hostile matrix (baseline / taskset -c 0,1 / 1-core / 6-spinner CPU load / combined) 63–91% of iterations killed, 5/5 runs; 3/3 kills at pool=1 on an idle box 0 kills, 0 soft-timeouts in 1,980/1,980 iterations (15 runs, fixed + random seeds)
StressSpec under taskset -c 0,1 (CI-signature repro) Fails in ~90s, 13/13 nodes, doubled GetClusterCoreRef timeout signature 3/3 passes, 13/13 nodes
Bugfix5962Spec (channel-executor constructor-tail canary; detected a rejected prototype 5/5) green 5/5 unconstrained + 5/5 under 2-core constraint
StartupWithOneThreadSpec green green, incl. under 1-core taskset
Akka.Cluster.Tests — 379/379
Akka.Cluster.Tools.Tests / Akka.Cluster.Metrics.Tests — 98/98 / 44/44
Akka.Cluster.Sharding.Tests — 191/192 — the 1 failure (RememberEntitiesStarterSpec constant-strategy throttling) is pre-existing on dev (fails 3/3 on unfixed HEAD), unrelated; will be filed separately
Akka.API.Tests — 18/18, no approval diff

dotnet format was applied to the touched files; the diff is unchanged with whitespace ignored (use hide-whitespace when reviewing).

Related: #4146, #5498 (this removes the blocking-Lazy window implicated there), the StressSpec failures on #8352's PR validation runs.

Possible follow-ups (deliberately not in this PR): pinning the cluster daemon tree to the internal dispatcher (JVM parity, defense-in-depth — no longer load-bearing once nothing blocks); an Akka.Analyzers rule flagging .Result/.Wait() inside property getters, which would have caught this mechanically.

The Cluster constructor blocked on GetClusterCoreRef().Result (an Ask to
/system/cluster bounded by akka.actor.creation-timeout), and the internal
ClusterCore getter re-issued that blocking ask whenever it observed an
unresolved core ref. Actors spawned during construction (cluster event
subscribers, the SBR downing provider) hit the getter from PreStart and
parked dispatcher threads; on small pools the parked waiters starved the
very daemon that had to reply, the redundant asks timed out, and the
timeout path called Shutdown() on an already-healthy node. This is the
root cause of the StressSpec nondeterminism on 2-core Windows CI agents
and reproduces as a hard deadlock with a 1-thread default dispatcher.

- Replace the constructor ask with fire-and-forget
  InternalClusterAction.Init(cluster); the constructor is now
  straight-line code whose completion depends on no other thread
- ClusterDaemon and ClusterCoreSupervisor become explicit
  Uninitialized/Initialized state machines (Become + unbounded stash):
  pre-Init traffic is buffered, post-Init unmatched traffic forwards to
  the core daemon; ordering is guaranteed by mailbox FIFO, not blocking
- ClusterCoreSupervisor hands the core ref back via
  Cluster.SetClusterCoreRef(); the ClusterCore getter is now
  unconditionally non-blocking (_clusterCore ?? _clusterDaemons)
- Delete GetClusterCoreRef() and its Shutdown()-on-timeout path;
  akka.actor.creation-timeout is no longer consulted by Akka.Cluster;
  genuine core-startup failures still shut the node down via the
  existing ClusterCoreSupervisor supervision strategy
- Add ClusterStartupFuzzSpec: randomized startup fuzz under pinned
  1-2-thread dispatcher pools; killed 63-91% of iterations pre-fix and
  gates this class of regression at zero kills
- Record the behavioral change in BREAKING_CHANGES_V1.6.md
- dotnet format applied to the touched files (includes pre-existing
  whitespace cleanup; diff is unchanged with whitespace ignored)

Verification: fuzz spec 1,980/1,980 iterations reached Up (0 kills, 0
soft-timeouts) across baseline / 2-core / 1-core / CPU-saturated /
combined conditions; StressSpec passes 3/3 under taskset -c 0,1 where it
previously failed in ~90s with the CI signature; Akka.Cluster.Tests
379/379; Cluster.Tools 98/98; Cluster.Metrics 44/44; Sharding 191/192
(the 1 failure is pre-existing on dev, unrelated); no public API changes
(Akka.API.Tests 18/18).

@Aaronontheweb Aaronontheweb left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM - need to fix slopwatch though

/// down. <c>ClusterStartupFuzzSpec</c> is the regression gate for that class of failure.
/// </para>
/// </summary>
internal IActorRef ClusterCore => _clusterCore ?? _clusterDaemons;

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Forwards messages down to the core ref in the event that the _clusterDaemons gets cached. Otherwise, we return the volatile _clusterCore type after initialization is complete.

// Everything else (Subscribe, JoinTo, cluster state queries, ...) is forwarded to the
// core supervisor, which forwards it on to the core daemon. Registration order matters
// in a ReceiveActor: this catch-all MUST be registered last.
ReceiveAny(msg => _coreSupervisor.Forward(msg));

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is what allows any races with Cluster to resolve naturally - forward any unhandled messages to the _coreSupervisor instead.

Suppress six SW003 (empty catch block) findings in the fuzz spec's
teardown/best-effort paths via slopwatch-ignore comments with
justifications. These catches are intentional: cancellation on the
hard test deadline, best-effort probing of a cluster mid-race, and
bounded/best-effort teardown of burner tasks and the actor system.
No behavior change.
@Aaronontheweb
Aaronontheweb enabled auto-merge (squash) July 10, 2026 20:42
@Aaronontheweb Aaronontheweb added the akka.net v1.6 Akka.NET v1.6-related issues label Jul 10, 2026
@Aaronontheweb
Aaronontheweb merged commit 26e32af into akkadotnet:dev Jul 10, 2026
11 checks passed
Aaronontheweb added a commit that referenced this pull request Sep 10, 2026
…n's lifetime so their order survives startup; LeaveAsync re-sends its Leave (#8580)

Cluster.ClusterCore used to switch from /system/cluster to a direct
reference to the resolved core daemon the instant that ref was
published during startup. Akka's FIFO guarantee holds only per
(sender, receiver) pair, so a Join and a Leave issued back-to-back by
the same caller with no await between them could ride different
mailboxes and arrive at the core daemon out of order. A Leave that
overtook its own JoinTo was dead-lettered by ClusterCoreDaemon's
Uninitialized behavior, and because LeaveSelf only ever sent its Leave
once, that loss was permanent. ClusterCore now always targets
/system/cluster, keeping the (sender, receiver) pair fixed for the
life of the extension; LeaveSelf also now re-sends its Leave command
on every call (Leaving(address) is idempotent), so a lost Leave no
longer wedges LeaveAsync forever.

Adds a regression fact in ClusterSpec covering Join immediately
followed by Leave with no await between them, and amends the existing
#8359 row in BREAKING_CHANGES_V1.6.md to describe the routing change.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
…carve-out of #8580)

Carves out only the LeaveSelf() hunk from dev's #8580 - not its ClusterCore routing change,
which fixes a regression from #8359 (non-blocking cluster startup) that is 1.6-only and does not
exist on v1.5 (v1.5's Cluster.cs still resolves ClusterCore once, synchronously, in the
constructor, and never switches receivers).

v1.5's LeaveSelf() had the same one-shot defect #8580 fixes regardless: it sent
ClusterUserAction.Leave exactly once, then memoized the resulting Task in a field forever. If
that single send were ever lost - dropped in transit, or issued before the core daemon had
finished starting - every later LeaveAsync()/Leave() call for the life of the process returned
the same never-completing memoized task, with no way to recover short of restarting the process.

Changed LeaveSelf() so only the MemberRemoved subscription is guarded by the CAS (it must run
exactly once), while the Leave send itself happens on every call, outside that guard - so a
later Leave() naturally retries. Safe because ClusterCoreDaemon.Leaving(address) is idempotent
once a member is no longer Joining/WeaklyUp/Up, so a redundant Leave delivered after the member
already left is a no-op.

Ported the matching ClusterSpec fact from dev, adapted: dev's version documents a specific
ClusterCore-routing race (Join immediately followed by Leave, no await between them, reordered
across two different receiver refs) that cannot happen on v1.5 since ClusterCore is a single
fixed ref here. The adapted comment says so. This fact does NOT exercise the re-send this commit
adds - it calls Leave() exactly once, and that first call behaves identically before and after
the fix - so it stays only as a Join-then-Leave smoke test kept aligned with dev.

Added a second fact, A_cluster_must_resend_Leave_on_a_later_LeaveAsync_call_after_an_earlier_send_was_lost,
that does discriminate: LeaveAsync() before any Join (dead-lettered by ClusterCoreDaemon's
Uninitialized behavior, which has no case for ClusterUserAction.Leave), then Join(self),
LeaderActions(), await MemberUp via a subscription, then LeaveAsync() again and assert the
returned task completes. Verified by temporarily reverting to the pre-fix one-shot LeaveSelf()
locally: this fact times out after 10s on the old code (the lost first Leave wedges the memoized
task forever) and passes in under a second on the fix.

Verified: `dotnet build src/core/Akka.Cluster.Tests -warnaserror` clean on both net48 and
net10.0; `dotnet test src/core/Akka.Cluster.Tests --filter ClusterSpec` run 5 times - 22/22
passed every time, including both facts.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
…ssSpec CI deadlock) (#8359)

* Akka.Cluster: make Cluster extension startup non-blocking

The Cluster constructor blocked on GetClusterCoreRef().Result (an Ask to
/system/cluster bounded by akka.actor.creation-timeout), and the internal
ClusterCore getter re-issued that blocking ask whenever it observed an
unresolved core ref. Actors spawned during construction (cluster event
subscribers, the SBR downing provider) hit the getter from PreStart and
parked dispatcher threads; on small pools the parked waiters starved the
very daemon that had to reply, the redundant asks timed out, and the
timeout path called Shutdown() on an already-healthy node. This is the
root cause of the StressSpec nondeterminism on 2-core Windows CI agents
and reproduces as a hard deadlock with a 1-thread default dispatcher.

- Replace the constructor ask with fire-and-forget
  InternalClusterAction.Init(cluster); the constructor is now
  straight-line code whose completion depends on no other thread
- ClusterDaemon and ClusterCoreSupervisor become explicit
  Uninitialized/Initialized state machines (Become + unbounded stash):
  pre-Init traffic is buffered, post-Init unmatched traffic forwards to
  the core daemon; ordering is guaranteed by mailbox FIFO, not blocking
- ClusterCoreSupervisor hands the core ref back via
  Cluster.SetClusterCoreRef(); the ClusterCore getter is now
  unconditionally non-blocking (_clusterCore ?? _clusterDaemons)
- Delete GetClusterCoreRef() and its Shutdown()-on-timeout path;
  akka.actor.creation-timeout is no longer consulted by Akka.Cluster;
  genuine core-startup failures still shut the node down via the
  existing ClusterCoreSupervisor supervision strategy
- Add ClusterStartupFuzzSpec: randomized startup fuzz under pinned
  1-2-thread dispatcher pools; killed 63-91% of iterations pre-fix and
  gates this class of regression at zero kills
- Record the behavioral change in BREAKING_CHANGES_V1.6.md
- dotnet format applied to the touched files (includes pre-existing
  whitespace cleanup; diff is unchanged with whitespace ignored)

Verification: fuzz spec 1,980/1,980 iterations reached Up (0 kills, 0
soft-timeouts) across baseline / 2-core / 1-core / CPU-saturated /
combined conditions; StressSpec passes 3/3 under taskset -c 0,1 where it
previously failed in ~90s with the CI signature; Akka.Cluster.Tests
379/379; Cluster.Tools 98/98; Cluster.Metrics 44/44; Sharding 191/192
(the 1 failure is pre-existing on dev, unrelated); no public API changes
(Akka.API.Tests 18/18).

* Address slopwatch SW003 findings in ClusterStartupFuzzSpec

Suppress six SW003 (empty catch block) findings in the fuzz spec's
teardown/best-effort paths via slopwatch-ignore comments with
justifications. These catches are intentional: cancellation on the
hard test deadline, best-effort probing of a cluster mid-race, and
bounded/best-effort teardown of burner tasks and the actor system.
No behavior change.

(cherry picked from commit 26e32af)
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
…ressSpec node-count doc

Three comment-only corrections, none change behavior:

* ClusterSpec.cs (from the #8579 pick): the comment on the MemberUp wait described dev's
  Cluster.ClusterCore re-targeting from a supervisor to the core daemon when the ref is
  published - a mechanism that no longer matches this branch's own code now that #8359 and
  #8580 are in. ClusterCore always targets /system/cluster for the life of the extension, so
  two commands from the same sender are already delivered in order without this wait; the
  MemberUp wait is still the stronger barrier, since it proves the join was processed, not
  merely enqueued, before Leave is sent. Reworded to say so.

* DistributedPubSubRestartSpec.cs (from the earlier Artery-removal tailoring commit): the
  reworded comment claimed a plain Tell to a quarantined peer is dropped at the transport layer.
  On v1.5's classic remoting it is not: EndpointManager's Quarantined case creates a brand-new
  writing endpoint for any Send that reaches it. It is the separate Gated policy that dead-letters
  a Send while its release deadline is unexpired. Reworded to describe what EndpointManager.cs
  actually does, and to explain ActorSelection.Tell's real advantage here (re-resolving by path on
  every send reaches whichever incarnation is live, rather than a stale cached ref).

* StressSpec.cs BuildConfig doc comment: said 13 was "also the smallest count that fits every
  phase" and then derived >= 11 two sentences later - contradicting itself. Reworded to say what
  the arithmetic says: 11 is the true minimum at full phase size; 13 is v1.5's original default,
  kept as the boundary below which this method starts shrinking phases.

Verified: dotnet build on Akka.Cluster.Tests, Akka.Cluster.Tools.Tests.MultiNode, and
Akka.Cluster.Tests.MultiNode -warnaserror, all clean.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
Adds a new #### 1.5.72 TBD #### section covering every user-visible change in this backport:
non-blocking Cluster extension startup and its startup-ordering follow-up (#8359, #8580), the
TestKit.Xunit async dispose chain and its CS0114 source break (#8545), the DotNetty batching
override fix (#8561), the new TestKitBase.ShutdownAsync API (#8499), the MultiNodeTestRunner
conductor port race fix and its node-exit backstop (#8515), the Streams Tcp Unbind timing fix
(#8570), the Sharding remember-entities write-timeout fix with a migration note (#8574), the
Replicator.IsKnownNode fix (#8582), the ShardCoordinator hand-over fix (#8586), and a one-line
summary of the de-flaked specs.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
…sfully)

The #8359 cherry-pick brought in ClusterStartupFuzzSpec.cs, which used
Task.IsCompletedSuccessfully at two call sites. On dev, Akka.Cluster.Tests
no longer targets net48, so this compiled fine there. On v1.5,
Akka.Cluster.Tests still targets net48 alongside net10.0, and
Task.IsCompletedSuccessfully does not exist on .NET Framework 4.8 - CI
build 131554 failed with CS1061 on every net48 lane.

Switched both call sites to `joinTask.Status == TaskStatus.RanToCompletion`,
the exact equivalent available on all target frameworks (the `using
System.Threading.Tasks;` needed for TaskStatus was already present).

Verified:
- `dotnet build -f net48 -warnaserror` and `-f net10.0 -warnaserror` both
  succeed with 0 warnings/errors for Akka.Cluster.Tests.
- `dotnet test -f net10.0 --filter FullyQualifiedName~ClusterStartupFuzzSpec`
  passes both tests: aggressive run completed=72, kills(delayed=0,
  immediate=0), pass=72 (reachedUp=72); moderate run completed=60,
  kills(delayed=0, immediate=0), pass=60 (reachedUp=60).
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
…ressSpec node-count doc

Three comment-only corrections, none change behavior:

* ClusterSpec.cs (from the #8579 pick): the comment on the MemberUp wait described dev's
  Cluster.ClusterCore re-targeting from a supervisor to the core daemon when the ref is
  published - a mechanism that no longer matches this branch's own code now that #8359 and
  #8580 are in. ClusterCore always targets /system/cluster for the life of the extension, so
  two commands from the same sender are already delivered in order without this wait; the
  MemberUp wait is still the stronger barrier, since it proves the join was processed, not
  merely enqueued, before Leave is sent. Reworded to say so.

* DistributedPubSubRestartSpec.cs (from the earlier Artery-removal tailoring commit): the
  reworded comment claimed a plain Tell to a quarantined peer is dropped at the transport layer.
  On v1.5's classic remoting it is not: EndpointManager's Quarantined case creates a brand-new
  writing endpoint for any Send that reaches it. It is the separate Gated policy that dead-letters
  a Send while its release deadline is unexpired. Reworded to describe what EndpointManager.cs
  actually does, and to explain ActorSelection.Tell's real advantage here (re-resolving by path on
  every send reaches whichever incarnation is live, rather than a stale cached ref).

* StressSpec.cs BuildConfig doc comment: said 13 was "also the smallest count that fits every
  phase" and then derived >= 11 two sentences later - contradicting itself. Reworded to say what
  the arithmetic says: 11 is the true minimum at full phase size; 13 is v1.5's original default,
  kept as the boundary below which this method starts shrinking phases.

Verified: dotnet build on Akka.Cluster.Tests, Akka.Cluster.Tools.Tests.MultiNode, and
Akka.Cluster.Tests.MultiNode -warnaserror, all clean.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
Adds a `1.5.72 TBD` section covering every user-visible change in this PR, including the
non-blocking Cluster startup fix and a migration note for #8574's write-timeout key change.

Review follow-up (Opus review, findings 2, 3, 5, 10, 11):
- The #8359 bullet now states outright that the cluster's internal actor tree
  (/system/cluster/core, /system/cluster/core/daemon, /system/cluster/heartbeatReceiver) is
  created after Cluster.Get() returns rather than before it - the one sharp edge of this change,
  and previously the only thing this note left unsaid - and says code resolving those paths
  right after Cluster.Get() should await JoinAsync/RegisterOnMemberUp or a cluster event instead.
- Says outright that akka.actor.creation-timeout is no longer consulted by Akka.Cluster at all
  (stronger and more accurate than "no longer blocks on"), while noting Akka.Cluster.Sharding
  still uses it to bound its own start-up asks, so nobody reads this as license to revert that
  setting.
- Replaced "a small dedicated dispatcher pool" with wording that also covers a starved default
  pool, since the deadlock reproduces on either.
- The #8580 bullet no longer says "during startup" - the routing through /system/cluster is
  permanent for the extension's lifetime (two extra local mailbox hops on the low-traffic command
  path), and now says a `Leave`/`LeaveAsync()` call issued after the node has already shut down
  produces a dead-lettered ClusterUserAction.Leave, relevant to dead-letter alerting.
- Adds a Testing bullet for #8398 (De-flake Bugfix5962Spec), the dev de-flake this branch had been
  missing for the spec #8359 makes newly racy.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
The #8359 cherry-pick brought in ClusterStartupFuzzSpec.cs, which uses
Task.IsCompletedSuccessfully at two call sites, and the #8398 cherry-pick (5f225ed, applied
earlier in this branch as the Opus-review follow-up for finding 1) uses the non-generic
`new TaskCompletionSource(TaskCreationOptions.RunContinuationsAsynchronously)` /
`TrySetResult()` in Bugfix5962Spec.cs. On dev, Akka.Cluster.Tests no longer targets net48, so
both compiled fine there. On v1.5, Akka.Cluster.Tests still targets net48 alongside net10.0:
`Task.IsCompletedSuccessfully` does not exist on .NET Framework 4.8 (CI build 131554 failed with
CS1061 on every net48 lane), and the non-generic `TaskCompletionSource` does not exist before
.NET 5 (CS0305).

- ClusterStartupFuzzSpec.cs: switched both call sites to
  `joinTask.Status == TaskStatus.RanToCompletion`, the exact equivalent available on all target
  frameworks (the `using System.Threading.Tasks;` needed for `TaskStatus` was already present).
- Bugfix5962Spec.cs: switched to `TaskCompletionSource<Done>` / `TrySetResult(Done.Instance)`,
  the same pattern already used by ClusterDaemon's own `_clusterPromise`/`_selfExiting` fields -
  available on every target framework this project builds for. No behavior change: `Done`
  carries no data, so the task's completion is still the only thing the spec observes.

Verified:
- `dotnet build src/core/Akka.Cluster.Tests -c Release -f net48 -warnaserror` and
  `-f net10.0 -warnaserror`: both 0 warnings/errors.
- `dotnet test -f net10.0 --filter FullyQualifiedName~ClusterStartupFuzzSpec` passes both tests:
  aggressive run completed=72, kills(delayed=0, immediate=0), pass=72 (reachedUp=72); moderate
  run completed=60, kills(delayed=0, immediate=0), pass=60 (reachedUp=60).
- `dotnet test -f net10.0 --filter FullyQualifiedName~Bugfix5962Spec`: passes.
This was referenced Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant