Skip to content

Add TestKitBase.ShutdownAsync and adopt it in StressSpec churn teardown - #8499

Merged
Aaronontheweb merged 4 commits into
akkadotnet:devfrom
Aaronontheweb:feature/testkit-shutdown-async
Aug 28, 2026
Merged

Aaronontheweb merged 4 commits into
akkadotnet:devfrom
Aaronontheweb:feature/testkit-shutdown-async

Conversation

@Aaronontheweb

Copy link
Copy Markdown
Member

Closes #8497.

Problem

TestKitBase.Shutdown is a blocking Terminate().Wait(duration). On a loaded agent that pins a thread-pool thread for the whole wait, while the work the shutdown is waiting on — coordinated-shutdown phases, cluster heartbeats — competes for the threads that remain. In StressSpec's join-remove churn phase the wait ran its full 10s and starved the hosting node's own heartbeat sender for 12+ seconds, so the survivors marked it unreachable and SBR downed it mid-test. The wait made its own timeout more likely.

Change

New API, purely additive — two overloads on TestKitBase, mirroring the sync pair exactly (same access modifiers, parameter names, defaults, forced-guardian-stop fallback, and throw-vs-log rule for verifySystemShutdown):

public virtual Task ShutdownAsync(TimeSpan? duration = null, bool verifySystemShutdown = false);
protected virtual Task ShutdownAsync(ActorSystem system, TimeSpan? duration = null, bool verifySystemShutdown = false);

The timeout race reuses AwaitWithTimeout from Akka.TestKit.Extensions — no new timer plumbing. The only observable difference from the sync path is that a faulted Terminate() surfaces its exception unwrapped instead of inside an AggregateException.

Both paths share one private force-stop tail, so the sync and async behavior cannot drift. The sync Shutdown keeps its own Terminate().Wait rather than blocking on the async one: sync-over-async would pin a thread and need two pool continuations scheduled while it is held — strictly worse under the starvation this fixes.

StressSpec adopts it at both churn teardown sites (in-loop and post-loop). Same wait budget, same fallback, same ordering relative to the settle block and the churn barriers — the wait just no longer holds a thread. StressSpec now contains no calls to the blocking Shutdown.

Tests

Three new deterministic tests in ShutdownAsyncTests. The timeout cases are not time-bounded guesses: a coordinated-shutdown task returns a TaskCompletionSource the test controls, so Terminate() genuinely cannot complete until released.

  1. Awaiting ShutdownAsync leaves WhenTerminated completed — catches a fire-and-forget regression.
  2. verifySystemShutdown: true + a stuck system throws TimeoutException naming the system.
  3. verifySystemShutdown: false + a stuck system logs the same message as a Warning and does not throw.

Verification

  • ShutdownAsync tests: 3/3, repeated 5×, all green.
  • Full Akka.TestKit.Tests: 327 passed, 1 pre-existing skip, 0 failed.
  • Akka.API.Tests: 18/18 after the additive approval update (two lines added per approval file, nothing removed or changed).
  • Akka.TestKit, Akka.TestKit.Tests, Akka.Cluster.Tests.MultiNode at -warnaserror: 0 warnings.

A full local StressSpec run is impractical — it needs ten node processes and reproduces only under the 2-vCPU starvation the issue describes — so the churn-site change rides on the clean build plus the API-behavior tests, and CI's MNTR lanes carry the live proof.

Closes akkadotnet#8497.

TestKitBase.Shutdown is a blocking `Terminate().Wait(duration)`. On a loaded
agent that pins a thread pool thread for the whole wait, and the work the
shutdown is waiting on - coordinated shutdown phases, cluster heartbeats -
competes for the threads that are left. In StressSpec's churn phase the wait
ran its full 10s and starved the hosting node's heartbeat sender for 12+
seconds, so the survivors marked it unreachable and SBR downed it. The wait
made its own timeout more likely.

ShutdownAsync races Terminate() against a timer instead of blocking, using the
AwaitWithTimeout helper Akka.TestKit already ships. Everything else matches the
sync overloads: same default duration, same forced guardian stop, same
verifySystemShutdown throw-vs-log, same message. Both paths now share the
force-stop tail, so the two cannot drift.

The sync Shutdown keeps its own Terminate().Wait rather than blocking on the
async one - sync-over-async here would need two continuations off the pool
while a thread is pinned, which is worse under the starvation this fixes.

StressSpec's post-churn teardown now awaits ShutdownAsync. Same wait budget,
same fallback; the wait just no longer holds a thread.
The post-loop teardown already awaits ShutdownAsync. This is the other blocking
Shutdown in ExerciseJoinRemoveAsync - it tears down the previous round's system
at the top of each churn round, so it is the one that blocks while the cluster
is under load. Leaving it on Terminate().Wait would keep pinning a thread pool
thread for the whole wait, which is what starves the hosting node's heartbeat
sender.

Same wait budget, same forced-guardian fallback, same ordering: the teardown
still finishes before the next ActorSystem is created.
@Aaronontheweb Aaronontheweb added tests akka-testkit Akka.NET Testkit issues labels Aug 27, 2026
@Aaronontheweb

Copy link
Copy Markdown
Member Author

Triage for the Linux Artery MNTR lane failure (build 130892): pre-existing flake, unrelated to this diff.

The lane failed on exactly one spec, Akka.Cluster.Sharding.Tests.DDataClusterShardingSpec. Node third hit the real failure — Timeout 00:00:05 while waiting for a message of type System.Int32 in the shard-restart phase — and the other four nodes then failed the after-shard-restart barrier behind it.

Evidence this is ambient:

  • The same spec failed the same lane with the same shape on dev build 130792, before this branch existed.
  • This PR touches TestKitBase (additive ShutdownAsync overloads) and StressSpec. DDataClusterShardingSpec uses neither; the sync Shutdown path it does use is byte-identical in behavior (only the shared force-stop tail was factored out).
  • StressSpec — the one MNTR spec this PR modifies — is absent from the failure artifact, i.e. it passed on this lane.

The classic Linux MNTR lane published a failure-logs artifact too, but its TRX files contain no failed results and the lane reported green — a first-attempt failure that passed on retry.

Logging DDataClusterShardingSpec (artery Linux, 2 recent occurrences: 130792, 130892) as a candidate for the next de-flake wave.

The verify-throws test already proves ShutdownAsync waits: a fire-and-forget
implementation cannot observe that the system failed to stop, so it cannot
throw the TimeoutException the test requires. StressSpec exercises the happy
path on every MNTR run. The two remaining tests stay separate because each
call runs the forced-guardian-stop tail, which kills the stuck system - so
the throw branch and the log branch each need their own system.
@Aaronontheweb

Copy link
Copy Markdown
Member Author

Triage for the two Linux MNTR failures on build 130907: same pre-existing flake family, both lanes, unrelated to this diff.

Both Linux lanes failed on one spec each — PersistentClusterShardingSpec, the persistence-mode variant of the ClusterSharding_specs mega-spec:

  • Artery lane: node third, Timeout 00:00:05 waiting for Int32, barrier after-shard-restart — byte-for-byte the signature DDataClusterShardingSpec produced on build 130892 (and 130792 on dev), just in the sibling variant.
  • Classic lane: node fourth, Timeout 00:00:13.8 waiting for Int32, barrier after-13 — a different phase of the same spec, same counter-round-trip shape.

Why this cannot be the PR:

Running tally for the ClusterShardingSpec family: DData variant on artery Linux (130792, 130892), Persistent variant on both Linux lanes (130907). This is the strongest candidate for the next de-flake wave; the recurring shape is an Int32 counter read racing shard recovery after a region restart.

Will retrigger CI once this build settles.

@Aaronontheweb
Aaronontheweb merged commit 81289e8 into akkadotnet:dev Aug 28, 2026
13 of 15 checks passed
@Aaronontheweb
Aaronontheweb deleted the feature/testkit-shutdown-async branch August 28, 2026 01:20
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
#8499 added TestKitBase.ShutdownAsync (public virtual) and its protected
ActorSystem overload. The API-approval snapshots kept during that cherry-pick
were deliberately v1.5's own pre-existing files (not dev's - the two
branches' TestKit surfaces differ by roughly 300 lines elsewhere), so
CoreAPISpec.ApproveTestKit was failing after the pick: v1.5's public API
now has the two new methods but the approved snapshot did not.

Ran `dotnet test src/core/Akka.API.Tests --filter ApproveTestKit` and
accepted the resulting .received.txt as the new DotNet.verified.txt. The
only delta from the previous baseline is the two new ShutdownAsync
overloads and the compiler's incidental renumbering of nearby
async-iterator state-machine class names (adding members shifts Roslyn's
sequential ordinals for unrelated nested types declared after them; this is
expected and harmless).

The Net.verified.txt (net48/netstandard2.0) counterpart could not be
regenerated by running the net48 test host in this environment (the vstest
agent fails to negotiate over the .NET Framework runtime here). Derived it
instead: TestKitBase.cs has no #if/target-framework-conditional code, and
diffing the previous DotNet/Net verified pair showed they differed only in
the single assembly-level TargetFrameworkAttribute line - every other line,
including all compiler-generated class ordinals, was already identical.
Applied the same content with only that one line swapped back to
".NETStandard,Version=v2.0", and confirmed the result still differs from
the new DotNet file by exactly that one line, matching the pre-existing
pattern.

Verified: `dotnet test src/core/Akka.API.Tests --framework net10.0` -
18/18 passed, including both ApproveTestKit and ApproveTestKitXunit2.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
…wn (port of #8543)

Hand-port of dev's #8543 substance, not a cherry-pick - #8543 is the last commit of a five-deep
stack (#8372, #8427, #8429, #8499) and its auto-merged parts would not compile as-is on v1.5
(BuildConfig's arithmetic and the ClusterResultAggregatorAsync call sites are written against
dev's 10-node/1-per-phase config, which v1.5 does not have).

What changed, and why each piece is still correct on v1.5:

* acceptable-heartbeat-pause raised from 3s to 20s, with the arithmetic comment explaining the
  phi-accrual crossing-threshold math (1s + 20s + 3*0.1s = 21.3s of detection). v1.5's
  failure-detector defaults (heartbeat-interval=1s, min-std-deviation=100ms) and
  split-brain-resolver stable-after (10s) already match dev's, so the same 3s-was-too-tight
  problem applies here: a churn round abruptly tearing down an ActorSystem can starve the
  surviving node's own heartbeat sender for longer than a 3s pause tolerates.

* akka.test.single-expect-default = 10s and akka.test.timefactor = 3, ported from #8372 (the
  first commit of the same dev stack), which the earlier port of this commit had missed. Every
  Within/WithinAsync bound in this file is dilated by timefactor - including RemoveOneAsync's
  removal budget (TimeSpan.FromSeconds(25) + ConvergenceWithin(3s, NbrUsedRoles - 1), which
  #8543 never widened) - so raising acceptable-heartbeat-pause to 20s without also porting
  timefactor left that budget structurally unable to cover the ~34.3s abrupt-removal path this
  spec's own ChurnMemberRemovalWithin() arithmetic predicts at the node counts the 7-node CI
  phase sequence reaches. Without timefactor, RemoveOneAsync's ceiling at NbrUsedRoles=3 is a flat
  31s; with it, 93s. Confirmed against three consecutive 7-node runs: the abrupt-removal phase
  measured 32.9-33.4s in each, and the AwaitAssert measured a 30.98s give-up in each, both
  matching the undilated 31s ceiling to within noise. All three runs pass with timefactor ported.

* ChurnMemberRemovalWithin(), ported byte-for-byte from dev (it only calls
  Cluster.Settings.FailureDetectorConfig/HeartbeatInterval/GossipInterval, all present on v1.5).
  Computes how long an abruptly-terminated churn member takes to actually leave the ring:
  detection + stable-after + a leader-gossip margin = 34.3s at this spec's config.

* ExerciseJoinRemoveAsync's loopDuration now includes ChurnMemberRemovalWithin() so each round's
  Within budget covers the full removal path, not just the new join. The abrupt-shutdown behavior
  itself (ShutdownAsync, no cluster Leave) was already in place from #8499 - this keeps exactly
  what the phase proves (abrupt loss), it only fixes the budget around it.

* Async TestKit migration of the call sites #8543 touches: the two RunOn sends inside the churn
  Loop become RunOnAsync, and ClusterResultAggregator (a one-shot, non-retried Identify/ExpectMsg
  lookup) gains a ClusterResultAggregatorAsync sibling (fresh-probe-per-attempt, retried over 30s)
  ported from dev, since a lone lost reply under this phase's deliberate churn must not be fatal.
  Repointed CreateResultAggregatorAsync, AwaitClusterResultAsync, and the async ReportResult<T>
  overload - the three call chains ExerciseJoinRemoveAsync depends on - at the new async lookup.
  Every remaining call site in the file passes an async lambda with an explicit return statement,
  so all of them already bound to the async ReportResult<T>(Func<Task<T>>) overload; the sync
  ClusterResultAggregator() and the sync ReportResult<T>(Func<T>) overload it served had no
  callers left after that repointing and are removed here. (Not ported: dev's extra
  "result-aggregator-identified" barrier in CreateResultAggregatorAsync, a related but separate
  hardening against a PartitionSeveral race - open item below.)

* RemoveOneAsync's watchee lookup, hand-ported from dev (#8372/#8429): a fresh CreateTestProbe()
  per attempt instead of the shared IdentifyProbe (the shared probe kept a timed-out attempt's
  late ActorIdentity queued, so the next attempt consumed that stale reply, and a reply resolved
  before the watchee existed carries a null Subject that the retry could never recover from); an
  explicit identity.Subject.Should().NotBeNull() guard before WatchAsync; WatchAsync bounded with
  WaitAsync(Dilated(3s)) instead of inheriting the outer Within's RemainingOrDefault, since a
  hung/slow watch Ask would otherwise burn the whole retry budget in a single attempt; and an
  explicit AwaitAssertAsync(10s, 1.25s) bound instead of an unbounded retry loop. This is the exact
  method that was failing in the runs above, so it is ported alongside the budget fix rather than
  left as a separate follow-up.

* StressSpecConfig node-count env override already existed on v1.5 (MNTR_STRESSSPEC_NODECOUNT,
  default 13). What it lacked was BuildConfig's shrink arithmetic: below the reference count,
  phase sizes must shrink or Settings' constructor throws. v1.5's reference config is heavier
  than dev's - every joining phase defaults to 2 nodes here, not 1 - so reaching the same
  practical floor of 7 requires shrinking both the joining side (halve every 2 back to 1) and the
  leaving/shutdown side (drop the "-large" one-by-one phases, halve the simultaneous counts), the
  same technique dev's config uses on one more group of phases. Derived and documented in
  StressSpecConfigSpec.cs (ported alongside, adapted from dev's 10-node version to v1.5's
  13-node/2-per-phase defaults): 7 is confirmed the practical floor (6 throws because the joining
  phases alone need 7 regardless of how far leaving/shutdown shrinks).

Also fixed while building this: MultiNodeTestRunner.cs's Process.Kill(bool) call from the earlier
#8515 hand-merge doesn't compile against netstandard2.0 - see the preceding commit.

Verified: dotnet build src/core/Akka.Cluster.Tests.MultiNode -warnaserror clean;
StressSpecConfigSpec 9/9; StressSpec run three times at MNTR_STRESSSPEC_NODECOUNT=7, all three
passed (see the PR body for full timings).

Open items for the maintainer:
- dev's extra CreateResultAggregatorAsync barrier (guards a PartitionSeveral aggregator-
  identification race) was not ported; the retrying ClusterResultAggregatorAsync lookup narrows
  that race but does not close it the way the extra barrier does.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
Adds a new #### 1.5.72 TBD #### section covering every user-visible change in this backport:
the TestKit.Xunit async dispose chain and its CS0114 source break (#8545), the DotNetty batching
override fix (#8561), the new TestKitBase.ShutdownAsync API (#8499), the MultiNodeTestRunner
conductor port race fix and its node-exit backstop (#8515), the Streams Tcp Unbind timing fix
(#8570), the Sharding remember-entities write-timeout fix with a migration note (#8574), the
Replicator.IsKnownNode fix (#8582), the ShardCoordinator hand-over fix (#8586), the LeaveSelf
re-send fix, and a one-line summary of the de-flaked specs.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
#8499 added TestKitBase.ShutdownAsync (public virtual) and its protected
ActorSystem overload. The API-approval snapshots kept during that cherry-pick
were deliberately v1.5's own pre-existing files (not dev's - the two
branches' TestKit surfaces differ by roughly 300 lines elsewhere), so
CoreAPISpec.ApproveTestKit was failing after the pick: v1.5's public API
now has the two new methods but the approved snapshot did not.

Ran `dotnet test src/core/Akka.API.Tests --filter ApproveTestKit` and
accepted the resulting .received.txt as the new DotNet.verified.txt. The
only delta from the previous baseline is the two new ShutdownAsync
overloads and the compiler's incidental renumbering of nearby
async-iterator state-machine class names (adding members shifts Roslyn's
sequential ordinals for unrelated nested types declared after them; this is
expected and harmless).

The Net.verified.txt (net48/netstandard2.0) counterpart could not be
regenerated by running the net48 test host in this environment (the vstest
agent fails to negotiate over the .NET Framework runtime here). Derived it
instead: TestKitBase.cs has no #if/target-framework-conditional code, and
diffing the previous DotNet/Net verified pair showed they differed only in
the single assembly-level TargetFrameworkAttribute line - every other line,
including all compiler-generated class ordinals, was already identical.
Applied the same content with only that one line swapped back to
".NETStandard,Version=v2.0", and confirmed the result still differs from
the new DotNet file by exactly that one line, matching the pre-existing
pattern.

Verified: `dotnet test src/core/Akka.API.Tests --framework net10.0` -
18/18 passed, including both ApproveTestKit and ApproveTestKitXunit2.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
…wn (port of #8543)

Hand-port of dev's #8543 substance, not a cherry-pick - #8543 is the last commit of a five-deep
stack (#8372, #8427, #8429, #8499) and its auto-merged parts would not compile as-is on v1.5
(BuildConfig's arithmetic and the ClusterResultAggregatorAsync call sites are written against
dev's 10-node/1-per-phase config, which v1.5 does not have).

What changed, and why each piece is still correct on v1.5:

* acceptable-heartbeat-pause raised from 3s to 20s, with the arithmetic comment explaining the
  phi-accrual crossing-threshold math (1s + 20s + 3*0.1s = 21.3s of detection). v1.5's
  failure-detector defaults (heartbeat-interval=1s, min-std-deviation=100ms) and
  split-brain-resolver stable-after (10s) already match dev's, so the same 3s-was-too-tight
  problem applies here: a churn round abruptly tearing down an ActorSystem can starve the
  surviving node's own heartbeat sender for longer than a 3s pause tolerates.

* akka.test.single-expect-default = 10s and akka.test.timefactor = 3, ported from #8372 (the
  first commit of the same dev stack), which the earlier port of this commit had missed. Every
  Within/WithinAsync bound in this file is dilated by timefactor - including RemoveOneAsync's
  removal budget (TimeSpan.FromSeconds(25) + ConvergenceWithin(3s, NbrUsedRoles - 1), which
  #8543 never widened) - so raising acceptable-heartbeat-pause to 20s without also porting
  timefactor left that budget structurally unable to cover the ~34.3s abrupt-removal path this
  spec's own ChurnMemberRemovalWithin() arithmetic predicts at the node counts the 7-node CI
  phase sequence reaches. Without timefactor, RemoveOneAsync's ceiling at NbrUsedRoles=3 is a flat
  31s; with it, 93s. Confirmed against three consecutive 7-node runs: the abrupt-removal phase
  measured 32.9-33.4s in each, and the AwaitAssert measured a 30.98s give-up in each, both
  matching the undilated 31s ceiling to within noise. All three runs pass with timefactor ported.

* ChurnMemberRemovalWithin(), ported byte-for-byte from dev (it only calls
  Cluster.Settings.FailureDetectorConfig/HeartbeatInterval/GossipInterval, all present on v1.5).
  Computes how long an abruptly-terminated churn member takes to actually leave the ring:
  detection + stable-after + a leader-gossip margin = 34.3s at this spec's config.

* ExerciseJoinRemoveAsync's loopDuration now includes ChurnMemberRemovalWithin() so each round's
  Within budget covers the full removal path, not just the new join. The abrupt-shutdown behavior
  itself (ShutdownAsync, no cluster Leave) was already in place from #8499 - this keeps exactly
  what the phase proves (abrupt loss), it only fixes the budget around it.

* Async TestKit migration of the call sites #8543 touches: the two RunOn sends inside the churn
  Loop become RunOnAsync, and ClusterResultAggregator (a one-shot, non-retried Identify/ExpectMsg
  lookup) gains a ClusterResultAggregatorAsync sibling (fresh-probe-per-attempt, retried over 30s)
  ported from dev, since a lone lost reply under this phase's deliberate churn must not be fatal.
  Repointed CreateResultAggregatorAsync, AwaitClusterResultAsync, and the async ReportResult<T>
  overload - the three call chains ExerciseJoinRemoveAsync depends on - at the new async lookup.
  Every remaining call site in the file passes an async lambda with an explicit return statement,
  so all of them already bound to the async ReportResult<T>(Func<Task<T>>) overload; the sync
  ClusterResultAggregator() and the sync ReportResult<T>(Func<T>) overload it served had no
  callers left after that repointing and are removed here. (Not ported: dev's extra
  "result-aggregator-identified" barrier in CreateResultAggregatorAsync, a related but separate
  hardening against a PartitionSeveral race - open item below.)

* RemoveOneAsync's watchee lookup, hand-ported from dev (#8372/#8429): a fresh CreateTestProbe()
  per attempt instead of the shared IdentifyProbe (the shared probe kept a timed-out attempt's
  late ActorIdentity queued, so the next attempt consumed that stale reply, and a reply resolved
  before the watchee existed carries a null Subject that the retry could never recover from); an
  explicit identity.Subject.Should().NotBeNull() guard before WatchAsync; WatchAsync bounded with
  WaitAsync(Dilated(3s)) instead of inheriting the outer Within's RemainingOrDefault, since a
  hung/slow watch Ask would otherwise burn the whole retry budget in a single attempt; and an
  explicit AwaitAssertAsync(10s, 1.25s) bound instead of an unbounded retry loop. This is the exact
  method that was failing in the runs above, so it is ported alongside the budget fix rather than
  left as a separate follow-up.

* StressSpecConfig node-count env override already existed on v1.5 (MNTR_STRESSSPEC_NODECOUNT,
  default 13). What it lacked was BuildConfig's shrink arithmetic: below the reference count,
  phase sizes must shrink or Settings' constructor throws. v1.5's reference config is heavier
  than dev's - every joining phase defaults to 2 nodes here, not 1 - so reaching the same
  practical floor of 7 requires shrinking both the joining side (halve every 2 back to 1) and the
  leaving/shutdown side (drop the "-large" one-by-one phases, halve the simultaneous counts), the
  same technique dev's config uses on one more group of phases. Derived and documented in
  StressSpecConfigSpec.cs (ported alongside, adapted from dev's 10-node version to v1.5's
  13-node/2-per-phase defaults): 7 is confirmed the practical floor (6 throws because the joining
  phases alone need 7 regardless of how far leaving/shutdown shrinks).

Also fixed while building this: MultiNodeTestRunner.cs's Process.Kill(bool) call from the earlier
#8515 hand-merge doesn't compile against netstandard2.0 - see the preceding commit.

Verified: dotnet build src/core/Akka.Cluster.Tests.MultiNode -warnaserror clean;
StressSpecConfigSpec 9/9; StressSpec run three times at MNTR_STRESSSPEC_NODECOUNT=7, all three
passed (see the PR body for full timings).

Open items for the maintainer:
- dev's extra CreateResultAggregatorAsync barrier (guards a PartitionSeveral aggregator-
  identification race) was not ported; the retrying ClusterResultAggregatorAsync lookup narrows
  that race but does not close it the way the extra barrier does.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
Adds a new #### 1.5.72 TBD #### section covering every user-visible change in this backport:
non-blocking Cluster extension startup and its startup-ordering follow-up (#8359, #8580), the
TestKit.Xunit async dispose chain and its CS0114 source break (#8545), the DotNetty batching
override fix (#8561), the new TestKitBase.ShutdownAsync API (#8499), the MultiNodeTestRunner
conductor port race fix and its node-exit backstop (#8515), the Streams Tcp Unbind timing fix
(#8570), the Sharding remember-entities write-timeout fix with a migration note (#8574), the
Replicator.IsKnownNode fix (#8582), the ShardCoordinator hand-over fix (#8586), and a one-line
summary of the de-flaked specs.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
#8499 added TestKitBase.ShutdownAsync (public virtual) and its protected
ActorSystem overload. The API-approval snapshots kept during that cherry-pick
were deliberately v1.5's own pre-existing files (not dev's - the two
branches' TestKit surfaces differ by roughly 300 lines elsewhere), so
CoreAPISpec.ApproveTestKit was failing after the pick: v1.5's public API
now has the two new methods but the approved snapshot did not.

Ran `dotnet test src/core/Akka.API.Tests --filter ApproveTestKit` and
accepted the resulting .received.txt as the new DotNet.verified.txt. The
only delta from the previous baseline is the two new ShutdownAsync
overloads and the compiler's incidental renumbering of nearby
async-iterator state-machine class names (adding members shifts Roslyn's
sequential ordinals for unrelated nested types declared after them; this is
expected and harmless).

The Net.verified.txt (net48/netstandard2.0) counterpart could not be
regenerated by running the net48 test host in this environment (the vstest
agent fails to negotiate over the .NET Framework runtime here). Derived it
instead: TestKitBase.cs has no #if/target-framework-conditional code, and
diffing the previous DotNet/Net verified pair showed they differed only in
the single assembly-level TargetFrameworkAttribute line - every other line,
including all compiler-generated class ordinals, was already identical.
Applied the same content with only that one line swapped back to
".NETStandard,Version=v2.0", and confirmed the result still differs from
the new DotNet file by exactly that one line, matching the pre-existing
pattern.

Verified: `dotnet test src/core/Akka.API.Tests --framework net10.0` -
18/18 passed, including both ApproveTestKit and ApproveTestKitXunit2.
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
…wn (port of #8543)

Hand-port of dev's #8543 substance, not a cherry-pick - #8543 is the last commit of a five-deep
stack (#8372, #8427, #8429, #8499) and its auto-merged parts would not compile as-is on v1.5
(BuildConfig's arithmetic and the ClusterResultAggregatorAsync call sites are written against
dev's 10-node/1-per-phase config, which v1.5 does not have).

What changed, and why each piece is still correct on v1.5:

* acceptable-heartbeat-pause raised from 3s to 20s, with the arithmetic comment explaining the
  phi-accrual crossing-threshold math (1s + 20s + 3*0.1s = 21.3s of detection). v1.5's
  failure-detector defaults (heartbeat-interval=1s, min-std-deviation=100ms) and
  split-brain-resolver stable-after (10s) already match dev's, so the same 3s-was-too-tight
  problem applies here: a churn round abruptly tearing down an ActorSystem can starve the
  surviving node's own heartbeat sender for longer than a 3s pause tolerates.

* akka.test.single-expect-default = 10s and akka.test.timefactor = 3, ported from #8372 (the
  first commit of the same dev stack), which the earlier port of this commit had missed. Every
  Within/WithinAsync bound in this file is dilated by timefactor - including RemoveOneAsync's
  removal budget (TimeSpan.FromSeconds(25) + ConvergenceWithin(3s, NbrUsedRoles - 1), which
  #8543 never widened) - so raising acceptable-heartbeat-pause to 20s without also porting
  timefactor left that budget structurally unable to cover the ~34.3s abrupt-removal path this
  spec's own ChurnMemberRemovalWithin() arithmetic predicts at the node counts the 7-node CI
  phase sequence reaches. Without timefactor, RemoveOneAsync's ceiling at NbrUsedRoles=3 is a flat
  31s; with it, 93s. Confirmed against three consecutive 7-node runs: the abrupt-removal phase
  measured 32.9-33.4s in each, and the AwaitAssert measured a 30.98s give-up in each, both
  matching the undilated 31s ceiling to within noise. All three runs pass with timefactor ported.

* ChurnMemberRemovalWithin(), ported byte-for-byte from dev (it only calls
  Cluster.Settings.FailureDetectorConfig/HeartbeatInterval/GossipInterval, all present on v1.5).
  Computes how long an abruptly-terminated churn member takes to actually leave the ring:
  detection + stable-after + a leader-gossip margin = 34.3s at this spec's config.

* ExerciseJoinRemoveAsync's loopDuration now includes ChurnMemberRemovalWithin() so each round's
  Within budget covers the full removal path, not just the new join. The abrupt-shutdown behavior
  itself (ShutdownAsync, no cluster Leave) was already in place from #8499 - this keeps exactly
  what the phase proves (abrupt loss), it only fixes the budget around it.

* Async TestKit migration of the call sites #8543 touches: the two RunOn sends inside the churn
  Loop become RunOnAsync, and ClusterResultAggregator (a one-shot, non-retried Identify/ExpectMsg
  lookup) gains a ClusterResultAggregatorAsync sibling (fresh-probe-per-attempt, retried over 30s)
  ported from dev, since a lone lost reply under this phase's deliberate churn must not be fatal.
  Repointed CreateResultAggregatorAsync, AwaitClusterResultAsync, and the async ReportResult<T>
  overload - the three call chains ExerciseJoinRemoveAsync depends on - at the new async lookup.
  Every remaining call site in the file passes an async lambda with an explicit return statement,
  so all of them already bound to the async ReportResult<T>(Func<Task<T>>) overload; the sync
  ClusterResultAggregator() and the sync ReportResult<T>(Func<T>) overload it served had no
  callers left after that repointing and are removed here. (Not ported: dev's extra
  "result-aggregator-identified" barrier in CreateResultAggregatorAsync, a related but separate
  hardening against a PartitionSeveral race - open item below.)

* RemoveOneAsync's watchee lookup, hand-ported from dev (#8372/#8429): a fresh CreateTestProbe()
  per attempt instead of the shared IdentifyProbe (the shared probe kept a timed-out attempt's
  late ActorIdentity queued, so the next attempt consumed that stale reply, and a reply resolved
  before the watchee existed carries a null Subject that the retry could never recover from); an
  explicit identity.Subject.Should().NotBeNull() guard before WatchAsync; WatchAsync bounded with
  WaitAsync(Dilated(3s)) instead of inheriting the outer Within's RemainingOrDefault, since a
  hung/slow watch Ask would otherwise burn the whole retry budget in a single attempt; and an
  explicit AwaitAssertAsync(10s, 1.25s) bound instead of an unbounded retry loop. This is the exact
  method that was failing in the runs above, so it is ported alongside the budget fix rather than
  left as a separate follow-up.

* StressSpecConfig node-count env override already existed on v1.5 (MNTR_STRESSSPEC_NODECOUNT,
  default 13). What it lacked was BuildConfig's shrink arithmetic: below the reference count,
  phase sizes must shrink or Settings' constructor throws. v1.5's reference config is heavier
  than dev's - every joining phase defaults to 2 nodes here, not 1 - so reaching the same
  practical floor of 7 requires shrinking both the joining side (halve every 2 back to 1) and the
  leaving/shutdown side (drop the "-large" one-by-one phases, halve the simultaneous counts), the
  same technique dev's config uses on one more group of phases. Derived and documented in
  StressSpecConfigSpec.cs (ported alongside, adapted from dev's 10-node version to v1.5's
  13-node/2-per-phase defaults): 7 is confirmed the practical floor (6 throws because the joining
  phases alone need 7 regardless of how far leaving/shutdown shrinks).

Also fixed while building this: MultiNodeTestRunner.cs's Process.Kill(bool) call from the earlier
#8515 hand-merge doesn't compile against netstandard2.0 - see the preceding commit.

Verified: dotnet build src/core/Akka.Cluster.Tests.MultiNode -warnaserror clean;
StressSpecConfigSpec 9/9; StressSpec run three times at MNTR_STRESSSPEC_NODECOUNT=7, all three
passed (see the PR body for full timings).

Open items for the maintainer:
- dev's extra CreateResultAggregatorAsync barrier (guards a PartitionSeveral aggregator-
  identification race) was not ported; the retrying ClusterResultAggregatorAsync lookup narrows
  that race but does not close it the way the extra barrier does.
This was referenced Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

akka.net v1.6 Akka.NET v1.6-related issues akka-testkit Akka.NET Testkit issues api-change multi node spec tests

Projects

None yet

Development

Successfully merging this pull request may close these issues.

StressSpec churn: blocking Shutdown starves the churn host's heartbeats on 2-vCPU agents until SBR downs it (TestKit lacks ShutdownAsync)

1 participant