Skip to content

De-flake EntityTerminationSpec: anchor on the shard's own entity set and wait out the restart backoff without blocking a pool thread - #8584

Merged
Aaronontheweb merged 2 commits into
devfrom
fix/entity-termination-spec-anchor-on-shard
Sep 10, 2026
Merged

Aaronontheweb merged 2 commits into
devfrom
fix/entity-termination-spec-anchor-on-shard

Conversation

@Aaronontheweb

@Aaronontheweb Aaronontheweb commented Sep 10, 2026 •

Copy link
Copy Markdown
Member

What changes

EntityTerminationSpec stops asserting shard state at a fixed wall-clock offset from an event the shard has not necessarily processed yet.

  • All three facts wait for the shard's active entity set to reach the expected value before reading anything else. That set is the observable the assertions read, so it is the right anchor. An entity's Terminated reaches the shard as a system message that the cell re-queues as a user message behind whatever is already in the shard's mailbox, so the test actor seeing Terminated says nothing about the shard having processed it.
  • Passivation fact. After the entity has left the active set, which is also the moment a shard that mistook the passivation for a crash would arm the 250 ms entity-restart-backoff, the fact holds a zero-count EventFilter on the shard's own "Started entity" line open for 600 ms, so a wrongful restart is observed directly rather than inferred from a later state read. Then it reads the state once more and asserts it is still empty. That read is deliberately a single read, not a poll: polling for "empty" would accept the first empty reading and stop proving that nothing restarted in the meantime.
  • Restart fact. The restart is the shard's own event: when the backoff expires it starts the entity again and logs the same "Started entity" line it logged the first time. The fact waits on that line instead of sleeping, because until the shard processes the Terminated its active set still holds the dead ref, so a state poll on its own could pass on the old incarnation.
  • Non-remembering fact. Its 400 ms sleep and the comment about a restart backoff are gone. With remember-entities off there is no restart path, so there is nothing to wait out.
  • Every state query goes through a fresh test probe, and the region's Failed set is asserted empty. The polling helper bounds each attempt at 1 s while the region's own query timeout is 3 s, so a timed-out attempt's late reply must not land in the test actor's queue, where a later expect would take it for a current snapshot.
  • The file is migrated to the async TestKit API. The one-node join moves from the synchronous AtStartup into an awaited helper each fact calls first.

Why

Build 131434, Windows unit tests, on #8554 (an Artery change that is not on this code path): the passivation fact failed with Expected collection to be empty, but found {"1"}. The log shows the entity passivating at 42.814, the test actor's ExpectTerminated returning at once, then nothing from the shard until 43.222, when it answered the state query with the entity still active, and 43.223, when it processed the entity's Terminated. The shard's mailbox did not run for 408 ms.

Nothing in the shard or the mailbox lost that run. The test body runs on a thread-pool worker, the pool's worker floor on the two-core hosted agent is two, and Thread.Sleep is the one form of blocking the pool cannot see. The shard's mailbox item had been queued to the other worker's local queue; with one worker asleep in the test and the other not returning to its dispatch loop, nothing could pick it up until the sleep expired, which is why the stall is 408 ms and not the gate thread's 500 ms starvation check. Both ends of the stall line up with the sleep to within a millisecond. The test actor's prompt Terminated proves nothing about pool health, because the TestKit's test actor runs on the calling-thread dispatcher and receives it inline on the dying entity's thread.

The fixed sleep was also anchored to the wrong event even when the test passed. A wrongful restart timer would be armed when the shard processes the Terminated, not when the test actor sees it, so the 400 ms window often did not cover the backoff at all. Pekko's copy of this spec has the same shape.

Not related to #8574, which changed the remember-entities write timeout: the log shows the 5 s value in effect, and the writes here completed in about a millisecond.

Later commit

The second commit answers review. The first version read the state on the test actor with a 1 s bound per polling attempt, shorter than the region's 3 s query timeout, so a timed-out attempt could leave a stale reply that the passivation fact's final read or the non-remembering fact's pong-2 expect would dequeue. Queries now go through a fresh probe per attempt. The passivation fact's negative check moved from a Task.Delay plus a state read to the zero-count filter described above. Rebased onto dev.

How it was checked

dotnet build src/contrib/cluster/Akka.Cluster.Sharding.Tests -c Release -warnaserror clean. The three facts run ten times through dotnet test after each commit: 60 of 60 passed, about 5 s per run. The grep for synchronous TestKit calls, blocking waits, and sleeps on the file returns nothing. Test-only change. No ledger entry.

… wait out the restart backoff without blocking a pool thread

The passivation fact slept a fixed 400 ms after the test actor saw the
entity's Terminated and then read the shard state once. The shard learns
of the death through a system message it re-queues as a user message, so
the test actor seeing Terminated says nothing about the shard having
processed it; on a two-core agent the sleep itself held one of the two
pool workers, the shard's mailbox item sat in the other worker's local
queue, and the pool only moved again when the sleep expired. The state
read then saw the entity still active.

Each fact now waits for the shard's active entity set to reach the
expected value, which is the observable the assertions read. The
passivation fact then waits out the 250 ms entity-restart-backoff with
Task.Delay and reads the state once more, deliberately not polled, so a
wrongful restart still fails it. The restart fact waits on the shard's
own "Started entity" line for the restart instead of a sleep, since a
state poll alone can pass on the dead ref before the shard has processed
the Terminated. The non-remembering fact drops its sleep and the comment
about a backoff that does not exist on that path.

The file is migrated to the async TestKit API; the one-node join moves
from the synchronous AtStartup into an awaited helper.
…empt; check for a wrongful restart with a zero-count filter on the shard's own restart line

The polling helper's 1s per-attempt bound is shorter than the region's 3s
query timeout, so a timed-out attempt left its reply in the test actor's
queue, where the passivation fact's final single read could take a
snapshot from before the wait as a current one, and the non-remembering
fact's "pong-2" expect could dequeue it. Each request now goes through a
fresh probe, and the region's Failed set is asserted empty.

The passivation fact replaces Task.Delay plus a state read with a
zero-count EventFilter on "Started entity" held open for 600 ms, which
observes a wrongful restart directly, then a final state read through a
probe with a 5 s bound.
@Aaronontheweb
Aaronontheweb force-pushed the fix/entity-termination-spec-anchor-on-shard branch from 2a3be35 to 80697cd Compare September 10, 2026 21:35
@Aaronontheweb
Aaronontheweb merged commit bef497a into dev Sep 10, 2026
15 checks passed
@Aaronontheweb
Aaronontheweb deleted the fix/entity-termination-spec-anchor-on-shard branch September 10, 2026 21:59
Aaronontheweb added a commit that referenced this pull request Sep 12, 2026
…and wait out the restart backoff without blocking a pool thread (#8584)

* De-flake EntityTerminationSpec: anchor on the shard's own entity set, wait out the restart backoff without blocking a pool thread

The passivation fact slept a fixed 400 ms after the test actor saw the
entity's Terminated and then read the shard state once. The shard learns
of the death through a system message it re-queues as a user message, so
the test actor seeing Terminated says nothing about the shard having
processed it; on a two-core agent the sleep itself held one of the two
pool workers, the shard's mailbox item sat in the other worker's local
queue, and the pool only moved again when the sleep expired. The state
read then saw the entity still active.

Each fact now waits for the shard's active entity set to reach the
expected value, which is the observable the assertions read. The
passivation fact then waits out the 250 ms entity-restart-backoff with
Task.Delay and reads the state once more, deliberately not polled, so a
wrongful restart still fails it. The restart fact waits on the shard's
own "Started entity" line for the restart instead of a sleep, since a
state poll alone can pass on the dead ref before the shard has processed
the Terminated. The non-remembering fact drops its sleep and the comment
about a backoff that does not exist on that path.

The file is migrated to the async TestKit API; the one-node join moves
from the synchronous AtStartup into an awaited helper.

* EntityTerminationSpec: query the region through a fresh probe per attempt; check for a wrongful restart with a zero-count filter on the shard's own restart line

The polling helper's 1s per-attempt bound is shorter than the region's 3s
query timeout, so a timed-out attempt left its reply in the test actor's
queue, where the passivation fact's final single read could take a
snapshot from before the wait as a current one, and the non-remembering
fact's "pong-2" expect could dequeue it. Each request now goes through a
fresh probe, and the region's Failed set is asserted empty.

The passivation fact replaces Task.Delay plus a state read with a
zero-count EventFilter on "Started entity" held open for 600 ms, which
observes a wrongful restart directly, then a final state read through a
probe with a 5 s bound.

(cherry picked from commit bef497a)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant