Repository navigation
TestKit.Xunit: implement the async dispose chain so a derived DisposeAsync no longer leaks the ActorSystem; de-flake StreamRefsSpec - #8545
Merged
Conversation
Aaronontheweb
added a commit
that referenced
this pull request
Sep 9, 2026
… fact; name it distinctly Each fact now owns its peer ActorSystem and shuts it down with ShutdownAsync in a finally block, so teardown no longer pins a thread pool thread in the synchronous AfterAll and no DisposeAsync is declared on the class, which would collide with the async dispose chain the TestKit gains in #8545. The peer system gets a distinct name; it self-joins its own cluster and Sys never joins it, so nothing depends on the two names matching.
Aaronontheweb
force-pushed
the
fix/testkit-xunit-async-dispose-chain
branch
from
September 9, 2026 18:18
4a56b3c to
b18522d
Compare
Aaronontheweb
added a commit
that referenced
this pull request
Sep 9, 2026
… fact; name it distinctly Each fact now owns its peer ActorSystem and shuts it down with ShutdownAsync in a finally block, so teardown no longer pins a thread pool thread in the synchronous AfterAll and no DisposeAsync is declared on the class, which would collide with the async dispose chain the TestKit gains in #8545. The peer system gets a distinct name; it self-joins its own cluster and Sys never joins it, so nothing depends on the two names matching.
Aaronontheweb
force-pushed
the
fix/testkit-xunit-async-dispose-chain
branch
from
September 10, 2026 04:59
b18522d to
4698131
Compare
Aaronontheweb
changed the base branch from
dev
to
fix/sharding-lease-spec-async-join
September 10, 2026 04:59
Aaronontheweb
added this pull request to stack #8578
September 10, 2026 12:59
Aaronontheweb
added a commit
that referenced
this pull request
Sep 10, 2026
…fresh probe per attempt under a sharding-derived budget; per-system settings The first message through a fresh region has to survive coordinator allocation and shard start, all ddata majority writes/reads that degrade to "all nodes" on this test's 2-node cluster. If the shard's remember-entities write stalls past its deadline, the Shard restarts and its buffered messages die with it - no dead letter, nothing left to re-deliver. Sharding is at-most-once, so a bare ExpectMsgAsync on the flat akka.test.single-expect-default can time out waiting for a message that no longer exists; waiting longer never helps. FirstMessageThrough re-sends each of the three cold sends through AwaitAssertAsync with a fresh TestProbe per attempt (so a late reply to an abandoned attempt can't satisfy a later one, or leak into the warm phase below), budgeted at ShardStartTimeout + UpdatingStateTimeout (the cold path plus one full Shard restart cycle) with each attempt capped at RetryInterval so the loop actually iterates. The counter snapshots are taken after the loops, so BeGreaterOrEqualTo still holds under retries. The warm-phase sends stay single sends (entities are live; a shard restart here would mean no remember-entities write ever happened, so retrying would hide a real bug) bounded by UpdatingStateTimeout instead of the flat probe default, plus a LastSender check against the original entity ref so a restart between phases can't hide behind the assertions. Also: StartShard was building both regions from Sys's settings instead of the settings of the system whose region is being started (harmless today since sysB is created from Sys.Settings.Config, but wrong); and AfterAll shut Sys down twice (once directly, once via TestKit.Dispose's finally) since _sysA is Sys - removed the redundant call and left a TODO(#8545) for migrating both shutdowns to the async TestKit dispose chain once that lands.
Aaronontheweb
added a commit
that referenced
this pull request
Sep 10, 2026
…fresh probe per attempt under a sharding-derived budget; per-system settings The first message through a fresh region has to survive coordinator allocation and shard start, all ddata majority writes/reads that degrade to "all nodes" on this test's 2-node cluster. If the shard's remember-entities write stalls past its deadline, the Shard restarts and its buffered messages die with it - no dead letter, nothing left to re-deliver. Sharding is at-most-once, so a bare ExpectMsgAsync on the flat akka.test.single-expect-default can time out waiting for a message that no longer exists; waiting longer never helps. FirstMessageThrough re-sends each of the three cold sends through AwaitAssertAsync with a fresh TestProbe per attempt (so a late reply to an abandoned attempt can't satisfy a later one, or leak into the warm phase below), budgeted at ShardStartTimeout + UpdatingStateTimeout (the cold path plus one full Shard restart cycle) with each attempt capped at RetryInterval so the loop actually iterates. The counter snapshots are taken after the loops, so BeGreaterOrEqualTo still holds under retries. The warm-phase sends stay single sends (entities are live; a shard restart here would mean no remember-entities write ever happened, so retrying would hide a real bug) bounded by UpdatingStateTimeout instead of the flat probe default, plus a LastSender check against the original entity ref so a restart between phases can't hide behind the assertions. Also: StartShard was building both regions from Sys's settings instead of the settings of the system whose region is being started (harmless today since sysB is created from Sys.Settings.Config, but wrong); and AfterAll shut Sys down twice (once directly, once via TestKit.Dispose's finally) since _sysA is Sys - removed the redundant call and left a TODO(#8545) for migrating both shutdowns to the async TestKit dispose chain once that lands.
Aaronontheweb
added a commit
that referenced
this pull request
Sep 10, 2026
…fresh probe per attempt under a sharding-derived budget (#8573) * De-flake ShardingBufferAdapterSpec: re-send the cold messages with a fresh probe per attempt under a sharding-derived budget; per-system settings The first message through a fresh region has to survive coordinator allocation and shard start, all ddata majority writes/reads that degrade to "all nodes" on this test's 2-node cluster. If the shard's remember-entities write stalls past its deadline, the Shard restarts and its buffered messages die with it - no dead letter, nothing left to re-deliver. Sharding is at-most-once, so a bare ExpectMsgAsync on the flat akka.test.single-expect-default can time out waiting for a message that no longer exists; waiting longer never helps. FirstMessageThrough re-sends each of the three cold sends through AwaitAssertAsync with a fresh TestProbe per attempt (so a late reply to an abandoned attempt can't satisfy a later one, or leak into the warm phase below), budgeted at ShardStartTimeout + UpdatingStateTimeout (the cold path plus one full Shard restart cycle) with each attempt capped at RetryInterval so the loop actually iterates. The counter snapshots are taken after the loops, so BeGreaterOrEqualTo still holds under retries. The warm-phase sends stay single sends (entities are live; a shard restart here would mean no remember-entities write ever happened, so retrying would hide a real bug) bounded by UpdatingStateTimeout instead of the flat probe default, plus a LastSender check against the original entity ref so a restart between phases can't hide behind the assertions. Also: StartShard was building both regions from Sys's settings instead of the settings of the system whose region is being started (harmless today since sysB is created from Sys.Settings.Config, but wrong); and AfterAll shut Sys down twice (once directly, once via TestKit.Dispose's finally) since _sysA is Sys - removed the redundant call and left a TODO(#8545) for migrating both shutdowns to the async TestKit dispose chain once that lands. * ShardingBufferAdapterSpec: the identity check's comment names the right equality
Aaronontheweb
force-pushed
the
fix/testkit-xunit-async-dispose-chain
branch
from
September 10, 2026 15:27
4698131 to
d682a57
Compare
Aaronontheweb
force-pushed
the
fix/testkit-xunit-async-dispose-chain
branch
from
September 10, 2026 15:46
d682a57 to
26d8a57
Compare
Aaronontheweb
added a commit
that referenced
this pull request
Sep 10, 2026
… fact; name it distinctly Each fact now owns its peer ActorSystem and shuts it down with ShutdownAsync in a finally block, so teardown no longer pins a thread pool thread in the synchronous AfterAll and no DisposeAsync is declared on the class, which would collide with the async dispose chain the TestKit gains in #8545. The peer system gets a distinct name; it self-joins its own cluster and Sys never joins it, so nothing depends on the two names matching.
Aaronontheweb
force-pushed
the
fix/testkit-xunit-async-dispose-chain
branch
from
September 11, 2026 14:25
36bd569 to
84a5a6d
Compare
Member
Author
|
/azp run |
|
Azure Pipelines successfully started running 1 pipeline(s). |
…ync no longer leaks the ActorSystem; de-flake StreamRefsSpec Fixes #8191: xUnit v3 calls IAsyncDisposable.DisposeAsync() in preference to IDisposable.Dispose() whenever a type implements both, so a derived spec's own no-op DisposeAsync (e.g. StreamRefsSpec) skipped TestKit's whole synchronous dispose chain and silently leaked a remoting-enabled ActorSystem per test. Akka.TestKit.Xunit.TestKit now implements IAsyncLifetime with a virtual InitializeAsync/DisposeAsync pair; DisposeAsync runs the sync chain and then shuts the system down with the non-blocking ShutdownAsync(). Ports the TestKit.cs shape from stalled PR #8217 (targeted v1.5) onto dev; its TestKitBase.cs hunk is dropped because ShutdownAsync already landed on dev via 81289e8, and EventFilterTestBase.cs is hand-merged onto dev's current AwaitAssertAsync retry body. Five other TestKit-derived specs that declared their own DisposeAsync/InitializeAsync are updated to override and chain to base: BugFixSpec, Bugfix8144Spec (Xunit v3), ParallelAmbientContextSpec (both base classes), and EventFilterTestBase. StreamRefsSpec.SinkRef_must_receive_elements_via_remoting is de-flaked on top of the leak fix: the test now gates on WatchTermination(Keep.Right) instead of asserting on a flat 3s wall-clock wait, uses async TestKit calls throughout, and gives every cross-boundary wait a dilated 15s budget. Remote system port is now 0 and loglevel is DEBUG for stage-level tracing.
…Kit API Convert ExpectMsg, ExpectNoMsg, and .Wait()-on-stream-completion calls to ExpectMsgAsync, ExpectNoMsgAsync, and awaited AwaitWithTimeout across all facts in the three touched files.
…ead of declaring its own
… so the async path needs no mode switch Dispose(bool) now only runs AfterAll() (and any derived teardown); it no longer terminates the ActorSystem and no longer guards re-entrancy. The guard collapses to a single _disposed flag checked once, in each public entry point: Dispose() runs the chain then calls the blocking Shutdown(), DisposeAsync() runs the chain then awaits ShutdownAsync(). This drops the now-redundant _disposing/_disposingAsync flags and the nested try/finally that used to suppress Shutdown() inside the async path. Verified: no Dispose(bool) override in the tree depends on the ActorSystem being torn down when base.Dispose(disposing) returns (the only overrides in the repo are on System.IO.Stream/GraphStageLogic, unrelated to this TestKit); StreamRefsSpec, BugFixSpec's DisposeAsync overrides still tear down their own systems before chaining to base.DisposeAsync(), and ClusterShardingLeaseSpec no longer declares one at all. Added a Bugfix8191Spec case covering Dispose() called after DisposeAsync().
Aaronontheweb
force-pushed
the
fix/testkit-xunit-async-dispose-chain
branch
from
September 11, 2026 17:34
84a5a6d to
c2b01e7
Compare
This was referenced Oct 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Stacked on #8558.
What changes
Akka.TestKit.Xunit.TestKitimplements the async lifecycle. It gainsvirtual InitializeAsync()andvirtual DisposeAsync(). Both public entry points run the same chain,Dispose(bool)withAfterAlland any override inside it, and differ only in how they terminate the ActorSystem:Dispose()blocks onShutdown(),DisposeAsync()awaitsShutdownAsync(). A derived test class that overridesDisposeAsyncand chains to the base no longer skips teardown. This is stalled PR fix: implement IAsyncLifetime on Akka.TestKit.Xunit (v3) to stop ActorSystem leaks (#8191) #8217 ported to dev, minus itsTestKitBasehunk, which landed in Add TestKitBase.ShutdownAsync and adopt it in StressSpec churn teardown #8499. Its regression test,Bugfix8191Spec, comes with it. Extend-only.DisposeAsyncorInitializeAsyncnow useoverrideand chain to the base.BREAKING_CHANGES_V1.6.mdrecords the compile-warning risk for user code in the same shape.StreamRefsSpecoverridesDisposeAsyncto shut the remote system down asynchronously and then chain to the base, binds its remote system to port 0 instead of a probed port, and its remoting fact is now async, gated onWatchTerminationso the test waits for the stream to finish rather than for a clock, with a dilated 15 s budget on all four remote waits.Why
Build 131174 failed
SinkRef_must_receive_elements_via_remotingafter a flat 3 s wait, on a PR that changed one package version. The stream-ref handshake is correct and matches Pekko; no element can leave before the partner is watched. The test was fragile for two reasons the analysis established:IAsyncDisposable.DisposeAsync()in preference toDispose()when a class has both.StreamRefsSpechad a no-opDisposeAsync, so its teardown never ran and every fact leaked two remoting-enabled ActorSystems, each keeping DotNetty threads and a 100 Hz scheduler on the shared thread pool. Twenty-eight leaked systems by the end of the class..GetAwaiter().GetResult()while the handshake needed about ten thread-pool dispatches, on a 2-vCPU agent whose pool floor is two. That is the shape every other thread-pool flake this week shared.No thread pool setting is changed. Two product items from the analysis are not in this PR and are worth their own issues:
SinkRefImpllacks Pekko's grace period on early partner termination, andActorMaterializerImpl.ActorOfblocks on a task result. A separate small PR fixes a TestKit config key that never took effect.How it was checked
Akka.TestKit.XunitandAkka.Streams.Testsbuild with warnings as errors, as do the four other touched test projects.Akka.API.Tests: 18 passed, no approval change, since the v3 TestKit has no approval file andTestKitBaseis untouched.StreamRefsSpecrun five times: 16 passed, 1 skipped, every time.Bugfix8191Spec, all ofAkka.TestKit.Xunit.Tests(33), andAkka.TestKit.Tests(220) pass.Second commit: the touched files move to the async TestKit API
Per the maintainer's rule that a touched test file migrates in the same PR: the fourteen remaining facts in
StreamRefsSpecareasync Taskand awaitExpectMsgAsync,ExpectNoMsgAsync, and the stream-completion tasks (three.Wait(8s)calls become awaits with the same timeout). One fact each in the DIBugFixSpecandBugfix8144Specmigrate the same way, since this PR already touched them for theoverride. No assertion or skip marker changes. A grep for the synchronous forms overStreamRefsSpecreturns nothing.StreamRefsSpecrun three times after the migration: 16 passed, 1 skipped, each time.Third commit: what the adversarial review found
EnsureSubscription,ExpectNext,ExpectNextN,ExpectError. Three more turned up on a full scan,ExpectCancellationtwice andExpectRequestonce. All ten now await the async form.BugFixSpecstill shut its DI-managed system down with a blocking call inAfterAll; it now overridesDisposeAsync, shuts that system down asynchronously, and chains to the base, the same shapeStreamRefsSpecuses.StreamRefsSpecrun three times: 16 passed, 1 skipped, each time.BugFixSpecpasses. Both projects build with warnings as errors. The grep for blocking probe and TestKit calls on both files returns only comments.Fourth commit: the lease spec joins the chain
ClusterShardingLeaseSpecon #8558 declares its ownInitializeAsyncandDisposeAsyncthroughIAsyncLifetime, and this PR makes both virtual on the TestKit, so whichever landed second would have failed the build with two hidden-member errors. This PR is now stacked on #8558, and its fourth commit converts that spec the way the other derived classes were converted:InitializeAsyncbecomes an override that chains to the base, the dispose bridge goes away because the base now runs the synchronous dispose and the async shutdown itself, and the interface declaration is dropped. The sharding test project builds with warnings as errors, which is the build that would have broken, and the lease spec passes 15 of 15 twice.Fifth commit: one flag, one chain
The maintainer asked why the TestKit carried three flags. Two predated this PR, and one of those was already redundant since it was never cleared; the third was a mode switch this PR added so
DisposeAsynccould run the virtualDispose(bool)while suppressing the blockingShutdown()inside it. The chain is now the smallest correct shape:Dispose(bool)runsAfterAll()and nothing else, so it is the one place overriders hook; a single_disposedflag is checked and set at the top of each public entry point;Dispose()runs the chain thenShutdown()in afinally,DisposeAsync()runs it thenawait ShutdownAsync()in afinally; a second call by either route, includingDispose()afterDisposeAsync(), is a no-op.What changes for an overrider:
base.Dispose(disposing)no longer terminates the system; the public entry point does, after the whole chain. EveryDispose(bool)override in this repo is on a stream or stage type, not on this TestKit, and the derivedDisposeAsyncoverrides inStreamRefsSpecand the DIBugFixSpecshut their own second system down and then chain to base, which is unaffected. A new fact inBugfix8191Specproves the chain runs once whenDispose()followsDisposeAsync().Checked:
Akka.TestKit.Xunitbuilds with warnings as errors;Akka.TestKit.Xunit.Tests34 passed;Akka.TestKit.Tests331 passed, 1 skipped;StreamRefsSpec16 passed, 1 skipped;Akka.API.Tests18 passed, no approval diff.