Repository navigation
Add TestKitBase.ShutdownAsync and adopt it in StressSpec churn teardown - #8499
Conversation
Closes akkadotnet#8497. TestKitBase.Shutdown is a blocking `Terminate().Wait(duration)`. On a loaded agent that pins a thread pool thread for the whole wait, and the work the shutdown is waiting on - coordinated shutdown phases, cluster heartbeats - competes for the threads that are left. In StressSpec's churn phase the wait ran its full 10s and starved the hosting node's heartbeat sender for 12+ seconds, so the survivors marked it unreachable and SBR downed it. The wait made its own timeout more likely. ShutdownAsync races Terminate() against a timer instead of blocking, using the AwaitWithTimeout helper Akka.TestKit already ships. Everything else matches the sync overloads: same default duration, same forced guardian stop, same verifySystemShutdown throw-vs-log, same message. Both paths now share the force-stop tail, so the two cannot drift. The sync Shutdown keeps its own Terminate().Wait rather than blocking on the async one - sync-over-async here would need two continuations off the pool while a thread is pinned, which is worse under the starvation this fixes. StressSpec's post-churn teardown now awaits ShutdownAsync. Same wait budget, same fallback; the wait just no longer holds a thread.
The post-loop teardown already awaits ShutdownAsync. This is the other blocking Shutdown in ExerciseJoinRemoveAsync - it tears down the previous round's system at the top of each churn round, so it is the one that blocks while the cluster is under load. Leaving it on Terminate().Wait would keep pinning a thread pool thread for the whole wait, which is what starves the hosting node's heartbeat sender. Same wait budget, same forced-guardian fallback, same ordering: the teardown still finishes before the next ActorSystem is created.
|
Triage for the Linux Artery MNTR lane failure (build 130892): pre-existing flake, unrelated to this diff. The lane failed on exactly one spec, Evidence this is ambient:
The classic Linux MNTR lane published a failure-logs artifact too, but its TRX files contain no failed results and the lane reported green — a first-attempt failure that passed on retry. Logging |
The verify-throws test already proves ShutdownAsync waits: a fire-and-forget implementation cannot observe that the system failed to stop, so it cannot throw the TimeoutException the test requires. StressSpec exercises the happy path on every MNTR run. The two remaining tests stay separate because each call runs the forced-guardian-stop tail, which kills the stuck system - so the throw branch and the log branch each need their own system.
|
Triage for the two Linux MNTR failures on build 130907: same pre-existing flake family, both lanes, unrelated to this diff. Both Linux lanes failed on one spec each —
Why this cannot be the PR:
Running tally for the Will retrigger CI once this build settles. |
#8499 added TestKitBase.ShutdownAsync (public virtual) and its protected ActorSystem overload. The API-approval snapshots kept during that cherry-pick were deliberately v1.5's own pre-existing files (not dev's - the two branches' TestKit surfaces differ by roughly 300 lines elsewhere), so CoreAPISpec.ApproveTestKit was failing after the pick: v1.5's public API now has the two new methods but the approved snapshot did not. Ran `dotnet test src/core/Akka.API.Tests --filter ApproveTestKit` and accepted the resulting .received.txt as the new DotNet.verified.txt. The only delta from the previous baseline is the two new ShutdownAsync overloads and the compiler's incidental renumbering of nearby async-iterator state-machine class names (adding members shifts Roslyn's sequential ordinals for unrelated nested types declared after them; this is expected and harmless). The Net.verified.txt (net48/netstandard2.0) counterpart could not be regenerated by running the net48 test host in this environment (the vstest agent fails to negotiate over the .NET Framework runtime here). Derived it instead: TestKitBase.cs has no #if/target-framework-conditional code, and diffing the previous DotNet/Net verified pair showed they differed only in the single assembly-level TargetFrameworkAttribute line - every other line, including all compiler-generated class ordinals, was already identical. Applied the same content with only that one line swapped back to ".NETStandard,Version=v2.0", and confirmed the result still differs from the new DotNet file by exactly that one line, matching the pre-existing pattern. Verified: `dotnet test src/core/Akka.API.Tests --framework net10.0` - 18/18 passed, including both ApproveTestKit and ApproveTestKitXunit2.
…wn (port of #8543) Hand-port of dev's #8543 substance, not a cherry-pick - #8543 is the last commit of a five-deep stack (#8372, #8427, #8429, #8499) and its auto-merged parts would not compile as-is on v1.5 (BuildConfig's arithmetic and the ClusterResultAggregatorAsync call sites are written against dev's 10-node/1-per-phase config, which v1.5 does not have). What changed, and why each piece is still correct on v1.5: * acceptable-heartbeat-pause raised from 3s to 20s, with the arithmetic comment explaining the phi-accrual crossing-threshold math (1s + 20s + 3*0.1s = 21.3s of detection). v1.5's failure-detector defaults (heartbeat-interval=1s, min-std-deviation=100ms) and split-brain-resolver stable-after (10s) already match dev's, so the same 3s-was-too-tight problem applies here: a churn round abruptly tearing down an ActorSystem can starve the surviving node's own heartbeat sender for longer than a 3s pause tolerates. * akka.test.single-expect-default = 10s and akka.test.timefactor = 3, ported from #8372 (the first commit of the same dev stack), which the earlier port of this commit had missed. Every Within/WithinAsync bound in this file is dilated by timefactor - including RemoveOneAsync's removal budget (TimeSpan.FromSeconds(25) + ConvergenceWithin(3s, NbrUsedRoles - 1), which #8543 never widened) - so raising acceptable-heartbeat-pause to 20s without also porting timefactor left that budget structurally unable to cover the ~34.3s abrupt-removal path this spec's own ChurnMemberRemovalWithin() arithmetic predicts at the node counts the 7-node CI phase sequence reaches. Without timefactor, RemoveOneAsync's ceiling at NbrUsedRoles=3 is a flat 31s; with it, 93s. Confirmed against three consecutive 7-node runs: the abrupt-removal phase measured 32.9-33.4s in each, and the AwaitAssert measured a 30.98s give-up in each, both matching the undilated 31s ceiling to within noise. All three runs pass with timefactor ported. * ChurnMemberRemovalWithin(), ported byte-for-byte from dev (it only calls Cluster.Settings.FailureDetectorConfig/HeartbeatInterval/GossipInterval, all present on v1.5). Computes how long an abruptly-terminated churn member takes to actually leave the ring: detection + stable-after + a leader-gossip margin = 34.3s at this spec's config. * ExerciseJoinRemoveAsync's loopDuration now includes ChurnMemberRemovalWithin() so each round's Within budget covers the full removal path, not just the new join. The abrupt-shutdown behavior itself (ShutdownAsync, no cluster Leave) was already in place from #8499 - this keeps exactly what the phase proves (abrupt loss), it only fixes the budget around it. * Async TestKit migration of the call sites #8543 touches: the two RunOn sends inside the churn Loop become RunOnAsync, and ClusterResultAggregator (a one-shot, non-retried Identify/ExpectMsg lookup) gains a ClusterResultAggregatorAsync sibling (fresh-probe-per-attempt, retried over 30s) ported from dev, since a lone lost reply under this phase's deliberate churn must not be fatal. Repointed CreateResultAggregatorAsync, AwaitClusterResultAsync, and the async ReportResult<T> overload - the three call chains ExerciseJoinRemoveAsync depends on - at the new async lookup. Every remaining call site in the file passes an async lambda with an explicit return statement, so all of them already bound to the async ReportResult<T>(Func<Task<T>>) overload; the sync ClusterResultAggregator() and the sync ReportResult<T>(Func<T>) overload it served had no callers left after that repointing and are removed here. (Not ported: dev's extra "result-aggregator-identified" barrier in CreateResultAggregatorAsync, a related but separate hardening against a PartitionSeveral race - open item below.) * RemoveOneAsync's watchee lookup, hand-ported from dev (#8372/#8429): a fresh CreateTestProbe() per attempt instead of the shared IdentifyProbe (the shared probe kept a timed-out attempt's late ActorIdentity queued, so the next attempt consumed that stale reply, and a reply resolved before the watchee existed carries a null Subject that the retry could never recover from); an explicit identity.Subject.Should().NotBeNull() guard before WatchAsync; WatchAsync bounded with WaitAsync(Dilated(3s)) instead of inheriting the outer Within's RemainingOrDefault, since a hung/slow watch Ask would otherwise burn the whole retry budget in a single attempt; and an explicit AwaitAssertAsync(10s, 1.25s) bound instead of an unbounded retry loop. This is the exact method that was failing in the runs above, so it is ported alongside the budget fix rather than left as a separate follow-up. * StressSpecConfig node-count env override already existed on v1.5 (MNTR_STRESSSPEC_NODECOUNT, default 13). What it lacked was BuildConfig's shrink arithmetic: below the reference count, phase sizes must shrink or Settings' constructor throws. v1.5's reference config is heavier than dev's - every joining phase defaults to 2 nodes here, not 1 - so reaching the same practical floor of 7 requires shrinking both the joining side (halve every 2 back to 1) and the leaving/shutdown side (drop the "-large" one-by-one phases, halve the simultaneous counts), the same technique dev's config uses on one more group of phases. Derived and documented in StressSpecConfigSpec.cs (ported alongside, adapted from dev's 10-node version to v1.5's 13-node/2-per-phase defaults): 7 is confirmed the practical floor (6 throws because the joining phases alone need 7 regardless of how far leaving/shutdown shrinks). Also fixed while building this: MultiNodeTestRunner.cs's Process.Kill(bool) call from the earlier #8515 hand-merge doesn't compile against netstandard2.0 - see the preceding commit. Verified: dotnet build src/core/Akka.Cluster.Tests.MultiNode -warnaserror clean; StressSpecConfigSpec 9/9; StressSpec run three times at MNTR_STRESSSPEC_NODECOUNT=7, all three passed (see the PR body for full timings). Open items for the maintainer: - dev's extra CreateResultAggregatorAsync barrier (guards a PartitionSeveral aggregator- identification race) was not ported; the retrying ClusterResultAggregatorAsync lookup narrows that race but does not close it the way the extra barrier does.
Adds a new #### 1.5.72 TBD #### section covering every user-visible change in this backport: the TestKit.Xunit async dispose chain and its CS0114 source break (#8545), the DotNetty batching override fix (#8561), the new TestKitBase.ShutdownAsync API (#8499), the MultiNodeTestRunner conductor port race fix and its node-exit backstop (#8515), the Streams Tcp Unbind timing fix (#8570), the Sharding remember-entities write-timeout fix with a migration note (#8574), the Replicator.IsKnownNode fix (#8582), the ShardCoordinator hand-over fix (#8586), the LeaveSelf re-send fix, and a one-line summary of the de-flaked specs.
#8499 added TestKitBase.ShutdownAsync (public virtual) and its protected ActorSystem overload. The API-approval snapshots kept during that cherry-pick were deliberately v1.5's own pre-existing files (not dev's - the two branches' TestKit surfaces differ by roughly 300 lines elsewhere), so CoreAPISpec.ApproveTestKit was failing after the pick: v1.5's public API now has the two new methods but the approved snapshot did not. Ran `dotnet test src/core/Akka.API.Tests --filter ApproveTestKit` and accepted the resulting .received.txt as the new DotNet.verified.txt. The only delta from the previous baseline is the two new ShutdownAsync overloads and the compiler's incidental renumbering of nearby async-iterator state-machine class names (adding members shifts Roslyn's sequential ordinals for unrelated nested types declared after them; this is expected and harmless). The Net.verified.txt (net48/netstandard2.0) counterpart could not be regenerated by running the net48 test host in this environment (the vstest agent fails to negotiate over the .NET Framework runtime here). Derived it instead: TestKitBase.cs has no #if/target-framework-conditional code, and diffing the previous DotNet/Net verified pair showed they differed only in the single assembly-level TargetFrameworkAttribute line - every other line, including all compiler-generated class ordinals, was already identical. Applied the same content with only that one line swapped back to ".NETStandard,Version=v2.0", and confirmed the result still differs from the new DotNet file by exactly that one line, matching the pre-existing pattern. Verified: `dotnet test src/core/Akka.API.Tests --framework net10.0` - 18/18 passed, including both ApproveTestKit and ApproveTestKitXunit2.
…wn (port of #8543) Hand-port of dev's #8543 substance, not a cherry-pick - #8543 is the last commit of a five-deep stack (#8372, #8427, #8429, #8499) and its auto-merged parts would not compile as-is on v1.5 (BuildConfig's arithmetic and the ClusterResultAggregatorAsync call sites are written against dev's 10-node/1-per-phase config, which v1.5 does not have). What changed, and why each piece is still correct on v1.5: * acceptable-heartbeat-pause raised from 3s to 20s, with the arithmetic comment explaining the phi-accrual crossing-threshold math (1s + 20s + 3*0.1s = 21.3s of detection). v1.5's failure-detector defaults (heartbeat-interval=1s, min-std-deviation=100ms) and split-brain-resolver stable-after (10s) already match dev's, so the same 3s-was-too-tight problem applies here: a churn round abruptly tearing down an ActorSystem can starve the surviving node's own heartbeat sender for longer than a 3s pause tolerates. * akka.test.single-expect-default = 10s and akka.test.timefactor = 3, ported from #8372 (the first commit of the same dev stack), which the earlier port of this commit had missed. Every Within/WithinAsync bound in this file is dilated by timefactor - including RemoveOneAsync's removal budget (TimeSpan.FromSeconds(25) + ConvergenceWithin(3s, NbrUsedRoles - 1), which #8543 never widened) - so raising acceptable-heartbeat-pause to 20s without also porting timefactor left that budget structurally unable to cover the ~34.3s abrupt-removal path this spec's own ChurnMemberRemovalWithin() arithmetic predicts at the node counts the 7-node CI phase sequence reaches. Without timefactor, RemoveOneAsync's ceiling at NbrUsedRoles=3 is a flat 31s; with it, 93s. Confirmed against three consecutive 7-node runs: the abrupt-removal phase measured 32.9-33.4s in each, and the AwaitAssert measured a 30.98s give-up in each, both matching the undilated 31s ceiling to within noise. All three runs pass with timefactor ported. * ChurnMemberRemovalWithin(), ported byte-for-byte from dev (it only calls Cluster.Settings.FailureDetectorConfig/HeartbeatInterval/GossipInterval, all present on v1.5). Computes how long an abruptly-terminated churn member takes to actually leave the ring: detection + stable-after + a leader-gossip margin = 34.3s at this spec's config. * ExerciseJoinRemoveAsync's loopDuration now includes ChurnMemberRemovalWithin() so each round's Within budget covers the full removal path, not just the new join. The abrupt-shutdown behavior itself (ShutdownAsync, no cluster Leave) was already in place from #8499 - this keeps exactly what the phase proves (abrupt loss), it only fixes the budget around it. * Async TestKit migration of the call sites #8543 touches: the two RunOn sends inside the churn Loop become RunOnAsync, and ClusterResultAggregator (a one-shot, non-retried Identify/ExpectMsg lookup) gains a ClusterResultAggregatorAsync sibling (fresh-probe-per-attempt, retried over 30s) ported from dev, since a lone lost reply under this phase's deliberate churn must not be fatal. Repointed CreateResultAggregatorAsync, AwaitClusterResultAsync, and the async ReportResult<T> overload - the three call chains ExerciseJoinRemoveAsync depends on - at the new async lookup. Every remaining call site in the file passes an async lambda with an explicit return statement, so all of them already bound to the async ReportResult<T>(Func<Task<T>>) overload; the sync ClusterResultAggregator() and the sync ReportResult<T>(Func<T>) overload it served had no callers left after that repointing and are removed here. (Not ported: dev's extra "result-aggregator-identified" barrier in CreateResultAggregatorAsync, a related but separate hardening against a PartitionSeveral race - open item below.) * RemoveOneAsync's watchee lookup, hand-ported from dev (#8372/#8429): a fresh CreateTestProbe() per attempt instead of the shared IdentifyProbe (the shared probe kept a timed-out attempt's late ActorIdentity queued, so the next attempt consumed that stale reply, and a reply resolved before the watchee existed carries a null Subject that the retry could never recover from); an explicit identity.Subject.Should().NotBeNull() guard before WatchAsync; WatchAsync bounded with WaitAsync(Dilated(3s)) instead of inheriting the outer Within's RemainingOrDefault, since a hung/slow watch Ask would otherwise burn the whole retry budget in a single attempt; and an explicit AwaitAssertAsync(10s, 1.25s) bound instead of an unbounded retry loop. This is the exact method that was failing in the runs above, so it is ported alongside the budget fix rather than left as a separate follow-up. * StressSpecConfig node-count env override already existed on v1.5 (MNTR_STRESSSPEC_NODECOUNT, default 13). What it lacked was BuildConfig's shrink arithmetic: below the reference count, phase sizes must shrink or Settings' constructor throws. v1.5's reference config is heavier than dev's - every joining phase defaults to 2 nodes here, not 1 - so reaching the same practical floor of 7 requires shrinking both the joining side (halve every 2 back to 1) and the leaving/shutdown side (drop the "-large" one-by-one phases, halve the simultaneous counts), the same technique dev's config uses on one more group of phases. Derived and documented in StressSpecConfigSpec.cs (ported alongside, adapted from dev's 10-node version to v1.5's 13-node/2-per-phase defaults): 7 is confirmed the practical floor (6 throws because the joining phases alone need 7 regardless of how far leaving/shutdown shrinks). Also fixed while building this: MultiNodeTestRunner.cs's Process.Kill(bool) call from the earlier #8515 hand-merge doesn't compile against netstandard2.0 - see the preceding commit. Verified: dotnet build src/core/Akka.Cluster.Tests.MultiNode -warnaserror clean; StressSpecConfigSpec 9/9; StressSpec run three times at MNTR_STRESSSPEC_NODECOUNT=7, all three passed (see the PR body for full timings). Open items for the maintainer: - dev's extra CreateResultAggregatorAsync barrier (guards a PartitionSeveral aggregator- identification race) was not ported; the retrying ClusterResultAggregatorAsync lookup narrows that race but does not close it the way the extra barrier does.
Adds a new #### 1.5.72 TBD #### section covering every user-visible change in this backport: non-blocking Cluster extension startup and its startup-ordering follow-up (#8359, #8580), the TestKit.Xunit async dispose chain and its CS0114 source break (#8545), the DotNetty batching override fix (#8561), the new TestKitBase.ShutdownAsync API (#8499), the MultiNodeTestRunner conductor port race fix and its node-exit backstop (#8515), the Streams Tcp Unbind timing fix (#8570), the Sharding remember-entities write-timeout fix with a migration note (#8574), the Replicator.IsKnownNode fix (#8582), the ShardCoordinator hand-over fix (#8586), and a one-line summary of the de-flaked specs.
#8499 added TestKitBase.ShutdownAsync (public virtual) and its protected ActorSystem overload. The API-approval snapshots kept during that cherry-pick were deliberately v1.5's own pre-existing files (not dev's - the two branches' TestKit surfaces differ by roughly 300 lines elsewhere), so CoreAPISpec.ApproveTestKit was failing after the pick: v1.5's public API now has the two new methods but the approved snapshot did not. Ran `dotnet test src/core/Akka.API.Tests --filter ApproveTestKit` and accepted the resulting .received.txt as the new DotNet.verified.txt. The only delta from the previous baseline is the two new ShutdownAsync overloads and the compiler's incidental renumbering of nearby async-iterator state-machine class names (adding members shifts Roslyn's sequential ordinals for unrelated nested types declared after them; this is expected and harmless). The Net.verified.txt (net48/netstandard2.0) counterpart could not be regenerated by running the net48 test host in this environment (the vstest agent fails to negotiate over the .NET Framework runtime here). Derived it instead: TestKitBase.cs has no #if/target-framework-conditional code, and diffing the previous DotNet/Net verified pair showed they differed only in the single assembly-level TargetFrameworkAttribute line - every other line, including all compiler-generated class ordinals, was already identical. Applied the same content with only that one line swapped back to ".NETStandard,Version=v2.0", and confirmed the result still differs from the new DotNet file by exactly that one line, matching the pre-existing pattern. Verified: `dotnet test src/core/Akka.API.Tests --framework net10.0` - 18/18 passed, including both ApproveTestKit and ApproveTestKitXunit2.
…wn (port of #8543) Hand-port of dev's #8543 substance, not a cherry-pick - #8543 is the last commit of a five-deep stack (#8372, #8427, #8429, #8499) and its auto-merged parts would not compile as-is on v1.5 (BuildConfig's arithmetic and the ClusterResultAggregatorAsync call sites are written against dev's 10-node/1-per-phase config, which v1.5 does not have). What changed, and why each piece is still correct on v1.5: * acceptable-heartbeat-pause raised from 3s to 20s, with the arithmetic comment explaining the phi-accrual crossing-threshold math (1s + 20s + 3*0.1s = 21.3s of detection). v1.5's failure-detector defaults (heartbeat-interval=1s, min-std-deviation=100ms) and split-brain-resolver stable-after (10s) already match dev's, so the same 3s-was-too-tight problem applies here: a churn round abruptly tearing down an ActorSystem can starve the surviving node's own heartbeat sender for longer than a 3s pause tolerates. * akka.test.single-expect-default = 10s and akka.test.timefactor = 3, ported from #8372 (the first commit of the same dev stack), which the earlier port of this commit had missed. Every Within/WithinAsync bound in this file is dilated by timefactor - including RemoveOneAsync's removal budget (TimeSpan.FromSeconds(25) + ConvergenceWithin(3s, NbrUsedRoles - 1), which #8543 never widened) - so raising acceptable-heartbeat-pause to 20s without also porting timefactor left that budget structurally unable to cover the ~34.3s abrupt-removal path this spec's own ChurnMemberRemovalWithin() arithmetic predicts at the node counts the 7-node CI phase sequence reaches. Without timefactor, RemoveOneAsync's ceiling at NbrUsedRoles=3 is a flat 31s; with it, 93s. Confirmed against three consecutive 7-node runs: the abrupt-removal phase measured 32.9-33.4s in each, and the AwaitAssert measured a 30.98s give-up in each, both matching the undilated 31s ceiling to within noise. All three runs pass with timefactor ported. * ChurnMemberRemovalWithin(), ported byte-for-byte from dev (it only calls Cluster.Settings.FailureDetectorConfig/HeartbeatInterval/GossipInterval, all present on v1.5). Computes how long an abruptly-terminated churn member takes to actually leave the ring: detection + stable-after + a leader-gossip margin = 34.3s at this spec's config. * ExerciseJoinRemoveAsync's loopDuration now includes ChurnMemberRemovalWithin() so each round's Within budget covers the full removal path, not just the new join. The abrupt-shutdown behavior itself (ShutdownAsync, no cluster Leave) was already in place from #8499 - this keeps exactly what the phase proves (abrupt loss), it only fixes the budget around it. * Async TestKit migration of the call sites #8543 touches: the two RunOn sends inside the churn Loop become RunOnAsync, and ClusterResultAggregator (a one-shot, non-retried Identify/ExpectMsg lookup) gains a ClusterResultAggregatorAsync sibling (fresh-probe-per-attempt, retried over 30s) ported from dev, since a lone lost reply under this phase's deliberate churn must not be fatal. Repointed CreateResultAggregatorAsync, AwaitClusterResultAsync, and the async ReportResult<T> overload - the three call chains ExerciseJoinRemoveAsync depends on - at the new async lookup. Every remaining call site in the file passes an async lambda with an explicit return statement, so all of them already bound to the async ReportResult<T>(Func<Task<T>>) overload; the sync ClusterResultAggregator() and the sync ReportResult<T>(Func<T>) overload it served had no callers left after that repointing and are removed here. (Not ported: dev's extra "result-aggregator-identified" barrier in CreateResultAggregatorAsync, a related but separate hardening against a PartitionSeveral race - open item below.) * RemoveOneAsync's watchee lookup, hand-ported from dev (#8372/#8429): a fresh CreateTestProbe() per attempt instead of the shared IdentifyProbe (the shared probe kept a timed-out attempt's late ActorIdentity queued, so the next attempt consumed that stale reply, and a reply resolved before the watchee existed carries a null Subject that the retry could never recover from); an explicit identity.Subject.Should().NotBeNull() guard before WatchAsync; WatchAsync bounded with WaitAsync(Dilated(3s)) instead of inheriting the outer Within's RemainingOrDefault, since a hung/slow watch Ask would otherwise burn the whole retry budget in a single attempt; and an explicit AwaitAssertAsync(10s, 1.25s) bound instead of an unbounded retry loop. This is the exact method that was failing in the runs above, so it is ported alongside the budget fix rather than left as a separate follow-up. * StressSpecConfig node-count env override already existed on v1.5 (MNTR_STRESSSPEC_NODECOUNT, default 13). What it lacked was BuildConfig's shrink arithmetic: below the reference count, phase sizes must shrink or Settings' constructor throws. v1.5's reference config is heavier than dev's - every joining phase defaults to 2 nodes here, not 1 - so reaching the same practical floor of 7 requires shrinking both the joining side (halve every 2 back to 1) and the leaving/shutdown side (drop the "-large" one-by-one phases, halve the simultaneous counts), the same technique dev's config uses on one more group of phases. Derived and documented in StressSpecConfigSpec.cs (ported alongside, adapted from dev's 10-node version to v1.5's 13-node/2-per-phase defaults): 7 is confirmed the practical floor (6 throws because the joining phases alone need 7 regardless of how far leaving/shutdown shrinks). Also fixed while building this: MultiNodeTestRunner.cs's Process.Kill(bool) call from the earlier #8515 hand-merge doesn't compile against netstandard2.0 - see the preceding commit. Verified: dotnet build src/core/Akka.Cluster.Tests.MultiNode -warnaserror clean; StressSpecConfigSpec 9/9; StressSpec run three times at MNTR_STRESSSPEC_NODECOUNT=7, all three passed (see the PR body for full timings). Open items for the maintainer: - dev's extra CreateResultAggregatorAsync barrier (guards a PartitionSeveral aggregator- identification race) was not ported; the retrying ClusterResultAggregatorAsync lookup narrows that race but does not close it the way the extra barrier does.
Closes #8497.
Problem
TestKitBase.Shutdownis a blockingTerminate().Wait(duration). On a loaded agent that pins a thread-pool thread for the whole wait, while the work the shutdown is waiting on — coordinated-shutdown phases, cluster heartbeats — competes for the threads that remain. In StressSpec's join-remove churn phase the wait ran its full 10s and starved the hosting node's own heartbeat sender for 12+ seconds, so the survivors marked it unreachable and SBR downed it mid-test. The wait made its own timeout more likely.Change
New API, purely additive — two overloads on
TestKitBase, mirroring the sync pair exactly (same access modifiers, parameter names, defaults, forced-guardian-stop fallback, and throw-vs-log rule forverifySystemShutdown):The timeout race reuses
AwaitWithTimeoutfromAkka.TestKit.Extensions— no new timer plumbing. The only observable difference from the sync path is that a faultedTerminate()surfaces its exception unwrapped instead of inside anAggregateException.Both paths share one private force-stop tail, so the sync and async behavior cannot drift. The sync
Shutdownkeeps its ownTerminate().Waitrather than blocking on the async one: sync-over-async would pin a thread and need two pool continuations scheduled while it is held — strictly worse under the starvation this fixes.StressSpec adopts it at both churn teardown sites (in-loop and post-loop). Same wait budget, same fallback, same ordering relative to the settle block and the churn barriers — the wait just no longer holds a thread.
StressSpecnow contains no calls to the blockingShutdown.Tests
Three new deterministic tests in
ShutdownAsyncTests. The timeout cases are not time-bounded guesses: a coordinated-shutdown task returns aTaskCompletionSourcethe test controls, soTerminate()genuinely cannot complete until released.ShutdownAsyncleavesWhenTerminatedcompleted — catches a fire-and-forget regression.verifySystemShutdown: true+ a stuck system throwsTimeoutExceptionnaming the system.verifySystemShutdown: false+ a stuck system logs the same message as aWarningand does not throw.Verification
Akka.TestKit.Tests: 327 passed, 1 pre-existing skip, 0 failed.Akka.API.Tests: 18/18 after the additive approval update (two lines added per approval file, nothing removed or changed).Akka.TestKit,Akka.TestKit.Tests,Akka.Cluster.Tests.MultiNodeat-warnaserror: 0 warnings.A full local StressSpec run is impractical — it needs ten node processes and reproduces only under the 2-vCPU starvation the issue describes — so the churn-site change rides on the clean build plus the API-behavior tests, and CI's MNTR lanes carry the live proof.