Skip to content

De-flake ParallelTestActorDeadlockSpec: wait for the actor, dilate the budgets - #8595

Merged
Aaronontheweb merged 2 commits into
akkadotnet:devfrom
Aaronontheweb:fix/parallel-testkit-deadlock-spec
Sep 30, 2026
Merged

Aaronontheweb merged 2 commits into
akkadotnet:devfrom
Aaronontheweb:fix/parallel-testkit-deadlock-spec

Conversation

@Aaronontheweb

@Aaronontheweb Aaronontheweb commented Sep 21, 2026 •

Copy link
Copy Markdown
Member

Summary

De-flakes ParallelTestActorDeadlockSpec, which timed out on the Windows unit-test lane for #8594 (Timeout 00:00:05 while waiting for a message of type System.String, one of sixteen TestKits). This is the fourth PR against this spec; #7974, #8016 and #8023 each raised a budget or lowered the concurrency and the failure kept returning.

The spec keeps its 16 Task.Run bodies on ThreadPool threads. That is deliberate: it is what Akka.Hosting.TestKit does (the TestKit is built inside a StartActors callback on a host-startup continuation), and it is the condition the spec exists to guard. An earlier revision of this PR moved the bodies to dedicated threads, which made the numbers pretty and would have hidden a real regression class. Reverted.

Root cause of the flake

Building a TestKit blocks its caller more than once: LoggingBus.StartDefaultLoggers waits up to akka.logger-startup-timeout, and TestKitBase waits for the TestActor's PreStart. Every one of those waits is for work that runs on the same CLR ThreadPool. On a two-core CI agent the pool injects roughly one thread per 500 ms, so sixteen blocked pool threads starve it.

Measured on a build pinned to two cores, 10 runs of 16 TestKits, with this change applied:

Stage p90 max
TestKit construction 4975 ms 5262 ms
PingerActor exists after ActorOf returns 21 ms 3116 ms
Ping arrives once the actor exists 8 ms 26 ms

The construction ceiling of about 5 s is the loggers hitting akka.logger-startup-timeout on the starved pool. The old spec put a flat 5 s ExpectMsgAsync budget on the sum of the second and third rows without ever waiting for the actor explicitly; the three earlier fixes re-sized that budget without changing what it waited on. Raising the pool floor (ThreadPool.SetMinThreads(64,64)) collapses every row to under 30 ms, which rules out everything except pool starvation.

The CI log from the failed run agrees: the Akka timestamps show 2.5 to 5 s holes in pool progress, and nine CoordinatedShutdown lines land within 4 ms of each other when the pool finally frees up.

Change

One file, src/core/Akka.TestKit.Tests/TestActorRefTests/ParallelTestActorDeadlockSpec.cs.

  • ActorOf returns a RepointableActorRef before the actor exists. The spec now resolves the PingerActor first, then expects its ping, so the wait is on the thing that is actually slow.
  • Per-step budgets are 10 s and go through Dilated, so akka.test.timefactor applies. Construction plus one wait still fits under the 20 s [Fact] timeout, which remains the deadlock detector: a genuine circular wait during startup, the TestKit deadlock during parallel test execution with TestActor initialization #7770 class, hangs regardless of budget and trips that timeout.
  • await using for teardown, so Terminate().Wait no longer blocks a pool thread during the shutdown cascade.
  • Bug fix: the spec passed $"test-{id}" to the single-string TestKit overload, which parses its argument as HOCON. Every system was named test, which is why three PRs' worth of log output could not tell the TestKits apart. It now uses the (Config, name, output) overload.
  • One timing line per TestKit for the next investigator.

Validation

Run Result
30 runs, pinned to two cores 0 failures, worst actor wait 3.1 s of 10 s
Akka.TestKit.Tests whole project 332 passed, 0 failed

Follow-ups not in this PR

  • The real cause is TestKit construction blocking on pool work. That belongs in TestKitBase: an async construction path, or the test loggers on a dispatcher that is not the shared pool. Until then this spec, and any user running many Akka.Hosting.TestKit tests in parallel on a small box, will see construction stall for the logger timeout under load.
  • TestKitBase.CreateInitialTestActor creates its ManualResetEventSlim with using var, but InternalTestActor keeps the reference; a TestActor restart would call Set() on a disposed handle.
  • The TestKit(string config, ...) overload silently accepting a system name as HOCON is an easy trap.
  • Spec 317 in Akka.Streams.Tests.TCK, the other flake on the same lane, uses a hard-coded 800 ms Timeouts.DefaultTimeoutMillis that nothing scales for CI, even though the sibling timeout in the same constructors already reads an environment variable. Separate change.

…e budgets

The spec starts 16 TestKits with Task.Run, and that is deliberate: it is
what Akka.Hosting.TestKit does (the TestKit is built inside a StartActors
callback on a host-startup continuation), so the bodies must stay on
ThreadPool threads or the spec stops guarding the condition it exists for.

Why it flaked: building a TestKit blocks its caller more than once
(LoggingBus.StartDefaultLoggers waits up to akka.logger-startup-timeout,
TestKitBase waits for the TestActor's PreStart), and every one of those
waits is for work that runs on the same pool. On a two-core agent the pool
injects roughly one thread per 500 ms, so sixteen blocked pool threads
starve it. Measured at two cores over 160 TestKits: construction up to
5.3 s (the loggers hit their 5 s timeout), and a PingerActor up to 3.1 s to
exist after ActorOf returns, against a flat 5 s ExpectMsgAsync budget.
akkadotnet#7974, akkadotnet#8016 and akkadotnet#8023 raised that budget or lowered the count; the wait
they were sizing was never made explicit.

- ActorOf returns a RepointableActorRef before the actor exists. The spec
  now resolves the PingerActor first, then expects its ping, so the wait is
  on the thing that is actually slow.
- The per-step budgets are 10 s and Dilated, so akka.test.timefactor
  applies. Construction plus one wait still fits under the 20 s [Fact]
  timeout, which remains the deadlock detector.
- Teardown is await using, so Terminate().Wait no longer blocks a pool
  thread during the shutdown cascade.
- The spec passed $"test-{id}" to the single-string TestKit overload, which
  parses its argument as HOCON: every system was named "test". It now uses
  the (Config, name, output) overload, so the log output is attributable.
- One Stopwatch line per TestKit is kept for the next investigator.

Validation on a build pinned to two cores: 30 runs, 0 failures; worst wait
for the actor 3.1 s of the 10 s budget, ping within 26 ms once it exists.
Akka.TestKit.Tests passes in full.

Follow-ups not made here: TestKit construction blocking on pool work is the
real cause and belongs in TestKitBase (an async construction path, or the
test loggers on a dispatcher that is not the shared pool); the
ManualResetEventSlim created with `using var` in CreateInitialTestActor is
kept by InternalTestActor and would be Set() after disposal on a TestActor
restart.
@Aaronontheweb
Aaronontheweb force-pushed the fix/parallel-testkit-deadlock-spec branch from b2d7821 to d9eeed5 Compare September 21, 2026 21:22
@Aaronontheweb Aaronontheweb changed the title De-flake ParallelTestActorDeadlockSpec: keep the blocking off the ThreadPool De-flake ParallelTestActorDeadlockSpec: wait for the actor, dilate the budgets Sep 21, 2026

@Aaronontheweb Aaronontheweb left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@Aaronontheweb
Aaronontheweb merged commit 7516241 into akkadotnet:dev Sep 30, 2026
15 checks passed
@Aaronontheweb
Aaronontheweb deleted the fix/parallel-testkit-deadlock-spec branch September 30, 2026 18:47
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant