Repository navigation
De-flake ParallelTestActorDeadlockSpec: wait for the actor, dilate the budgets - #8595
Merged
Aaronontheweb merged 2 commits intoSep 30, 2026
Conversation
…e budgets The spec starts 16 TestKits with Task.Run, and that is deliberate: it is what Akka.Hosting.TestKit does (the TestKit is built inside a StartActors callback on a host-startup continuation), so the bodies must stay on ThreadPool threads or the spec stops guarding the condition it exists for. Why it flaked: building a TestKit blocks its caller more than once (LoggingBus.StartDefaultLoggers waits up to akka.logger-startup-timeout, TestKitBase waits for the TestActor's PreStart), and every one of those waits is for work that runs on the same pool. On a two-core agent the pool injects roughly one thread per 500 ms, so sixteen blocked pool threads starve it. Measured at two cores over 160 TestKits: construction up to 5.3 s (the loggers hit their 5 s timeout), and a PingerActor up to 3.1 s to exist after ActorOf returns, against a flat 5 s ExpectMsgAsync budget. akkadotnet#7974, akkadotnet#8016 and akkadotnet#8023 raised that budget or lowered the count; the wait they were sizing was never made explicit. - ActorOf returns a RepointableActorRef before the actor exists. The spec now resolves the PingerActor first, then expects its ping, so the wait is on the thing that is actually slow. - The per-step budgets are 10 s and Dilated, so akka.test.timefactor applies. Construction plus one wait still fits under the 20 s [Fact] timeout, which remains the deadlock detector. - Teardown is await using, so Terminate().Wait no longer blocks a pool thread during the shutdown cascade. - The spec passed $"test-{id}" to the single-string TestKit overload, which parses its argument as HOCON: every system was named "test". It now uses the (Config, name, output) overload, so the log output is attributable. - One Stopwatch line per TestKit is kept for the next investigator. Validation on a build pinned to two cores: 30 runs, 0 failures; worst wait for the actor 3.1 s of the 10 s budget, ping within 26 ms once it exists. Akka.TestKit.Tests passes in full. Follow-ups not made here: TestKit construction blocking on pool work is the real cause and belongs in TestKitBase (an async construction path, or the test loggers on a dispatcher that is not the shared pool); the ManualResetEventSlim created with `using var` in CreateInitialTestActor is kept by InternalTestActor and would be Set() after disposal on a TestActor restart.
Aaronontheweb
force-pushed
the
fix/parallel-testkit-deadlock-spec
branch
from
September 21, 2026 21:22
b2d7821 to
d9eeed5
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
De-flakes
ParallelTestActorDeadlockSpec, which timed out on the Windows unit-test lane for #8594 (Timeout 00:00:05 while waiting for a message of type System.String, one of sixteen TestKits). This is the fourth PR against this spec; #7974, #8016 and #8023 each raised a budget or lowered the concurrency and the failure kept returning.The spec keeps its 16
Task.Runbodies on ThreadPool threads. That is deliberate: it is what Akka.Hosting.TestKit does (the TestKit is built inside aStartActorscallback on a host-startup continuation), and it is the condition the spec exists to guard. An earlier revision of this PR moved the bodies to dedicated threads, which made the numbers pretty and would have hidden a real regression class. Reverted.Root cause of the flake
Building a TestKit blocks its caller more than once:
LoggingBus.StartDefaultLoggerswaits up toakka.logger-startup-timeout, andTestKitBasewaits for the TestActor'sPreStart. Every one of those waits is for work that runs on the same CLR ThreadPool. On a two-core CI agent the pool injects roughly one thread per 500 ms, so sixteen blocked pool threads starve it.Measured on a build pinned to two cores, 10 runs of 16 TestKits, with this change applied:
PingerActorexists afterActorOfreturnsThe construction ceiling of about 5 s is the loggers hitting
akka.logger-startup-timeouton the starved pool. The old spec put a flat 5 sExpectMsgAsyncbudget on the sum of the second and third rows without ever waiting for the actor explicitly; the three earlier fixes re-sized that budget without changing what it waited on. Raising the pool floor (ThreadPool.SetMinThreads(64,64)) collapses every row to under 30 ms, which rules out everything except pool starvation.The CI log from the failed run agrees: the Akka timestamps show 2.5 to 5 s holes in pool progress, and nine
CoordinatedShutdownlines land within 4 ms of each other when the pool finally frees up.Change
One file,
src/core/Akka.TestKit.Tests/TestActorRefTests/ParallelTestActorDeadlockSpec.cs.ActorOfreturns aRepointableActorRefbefore the actor exists. The spec now resolves thePingerActorfirst, then expects its ping, so the wait is on the thing that is actually slow.Dilated, soakka.test.timefactorapplies. Construction plus one wait still fits under the 20 s[Fact]timeout, which remains the deadlock detector: a genuine circular wait during startup, the TestKit deadlock during parallel test execution with TestActor initialization #7770 class, hangs regardless of budget and trips that timeout.await usingfor teardown, soTerminate().Waitno longer blocks a pool thread during the shutdown cascade.$"test-{id}"to the single-stringTestKitoverload, which parses its argument as HOCON. Every system was namedtest, which is why three PRs' worth of log output could not tell the TestKits apart. It now uses the(Config, name, output)overload.Validation
Akka.TestKit.Testswhole projectFollow-ups not in this PR
TestKitBase: an async construction path, or the test loggers on a dispatcher that is not the shared pool. Until then this spec, and any user running many Akka.Hosting.TestKit tests in parallel on a small box, will see construction stall for the logger timeout under load.TestKitBase.CreateInitialTestActorcreates itsManualResetEventSlimwithusing var, butInternalTestActorkeeps the reference; a TestActor restart would callSet()on a disposed handle.TestKit(string config, ...)overload silently accepting a system name as HOCON is an easy trap.Akka.Streams.Tests.TCK, the other flake on the same lane, uses a hard-coded 800 msTimeouts.DefaultTimeoutMillisthat nothing scales for CI, even though the sibling timeout in the same constructors already reads an environment variable. Separate change.