Repository navigation
De-flake EntityTerminationSpec: anchor on the shard's own entity set and wait out the restart backoff without blocking a pool thread - #8584
Merged
Conversation
… wait out the restart backoff without blocking a pool thread The passivation fact slept a fixed 400 ms after the test actor saw the entity's Terminated and then read the shard state once. The shard learns of the death through a system message it re-queues as a user message, so the test actor seeing Terminated says nothing about the shard having processed it; on a two-core agent the sleep itself held one of the two pool workers, the shard's mailbox item sat in the other worker's local queue, and the pool only moved again when the sleep expired. The state read then saw the entity still active. Each fact now waits for the shard's active entity set to reach the expected value, which is the observable the assertions read. The passivation fact then waits out the 250 ms entity-restart-backoff with Task.Delay and reads the state once more, deliberately not polled, so a wrongful restart still fails it. The restart fact waits on the shard's own "Started entity" line for the restart instead of a sleep, since a state poll alone can pass on the dead ref before the shard has processed the Terminated. The non-remembering fact drops its sleep and the comment about a backoff that does not exist on that path. The file is migrated to the async TestKit API; the one-node join moves from the synchronous AtStartup into an awaited helper.
…empt; check for a wrongful restart with a zero-count filter on the shard's own restart line The polling helper's 1s per-attempt bound is shorter than the region's 3s query timeout, so a timed-out attempt left its reply in the test actor's queue, where the passivation fact's final single read could take a snapshot from before the wait as a current one, and the non-remembering fact's "pong-2" expect could dequeue it. Each request now goes through a fresh probe, and the region's Failed set is asserted empty. The passivation fact replaces Task.Delay plus a state read with a zero-count EventFilter on "Started entity" held open for 600 ms, which observes a wrongful restart directly, then a final state read through a probe with a 5 s bound.
Aaronontheweb
force-pushed
the
fix/entity-termination-spec-anchor-on-shard
branch
from
September 10, 2026 21:35
2a3be35 to
80697cd
Compare
Aaronontheweb
deleted the
fix/entity-termination-spec-anchor-on-shard
branch
September 10, 2026 21:59
This was referenced Sep 11, 2026
Merged
Aaronontheweb
added a commit
that referenced
this pull request
Sep 12, 2026
…and wait out the restart backoff without blocking a pool thread (#8584) * De-flake EntityTerminationSpec: anchor on the shard's own entity set, wait out the restart backoff without blocking a pool thread The passivation fact slept a fixed 400 ms after the test actor saw the entity's Terminated and then read the shard state once. The shard learns of the death through a system message it re-queues as a user message, so the test actor seeing Terminated says nothing about the shard having processed it; on a two-core agent the sleep itself held one of the two pool workers, the shard's mailbox item sat in the other worker's local queue, and the pool only moved again when the sleep expired. The state read then saw the entity still active. Each fact now waits for the shard's active entity set to reach the expected value, which is the observable the assertions read. The passivation fact then waits out the 250 ms entity-restart-backoff with Task.Delay and reads the state once more, deliberately not polled, so a wrongful restart still fails it. The restart fact waits on the shard's own "Started entity" line for the restart instead of a sleep, since a state poll alone can pass on the dead ref before the shard has processed the Terminated. The non-remembering fact drops its sleep and the comment about a backoff that does not exist on that path. The file is migrated to the async TestKit API; the one-node join moves from the synchronous AtStartup into an awaited helper. * EntityTerminationSpec: query the region through a fresh probe per attempt; check for a wrongful restart with a zero-count filter on the shard's own restart line The polling helper's 1s per-attempt bound is shorter than the region's 3s query timeout, so a timed-out attempt left its reply in the test actor's queue, where the passivation fact's final single read could take a snapshot from before the wait as a current one, and the non-remembering fact's "pong-2" expect could dequeue it. Each request now goes through a fresh probe, and the region's Failed set is asserted empty. The passivation fact replaces Task.Delay plus a state read with a zero-count EventFilter on "Started entity" held open for 600 ms, which observes a wrongful restart directly, then a final state read through a probe with a 5 s bound. (cherry picked from commit bef497a)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changes
EntityTerminationSpecstops asserting shard state at a fixed wall-clock offset from an event the shard has not necessarily processed yet.Terminatedreaches the shard as a system message that the cell re-queues as a user message behind whatever is already in the shard's mailbox, so the test actor seeingTerminatedsays nothing about the shard having processed it.entity-restart-backoff, the fact holds a zero-countEventFilteron the shard's own "Started entity" line open for 600 ms, so a wrongful restart is observed directly rather than inferred from a later state read. Then it reads the state once more and asserts it is still empty. That read is deliberately a single read, not a poll: polling for "empty" would accept the first empty reading and stop proving that nothing restarted in the meantime.Terminatedits active set still holds the dead ref, so a state poll on its own could pass on the old incarnation.Failedset is asserted empty. The polling helper bounds each attempt at 1 s while the region's own query timeout is 3 s, so a timed-out attempt's late reply must not land in the test actor's queue, where a later expect would take it for a current snapshot.AtStartupinto an awaited helper each fact calls first.Why
Build 131434, Windows unit tests, on #8554 (an Artery change that is not on this code path): the passivation fact failed with
Expected collection to be empty, but found {"1"}. The log shows the entity passivating at 42.814, the test actor'sExpectTerminatedreturning at once, then nothing from the shard until 43.222, when it answered the state query with the entity still active, and 43.223, when it processed the entity'sTerminated. The shard's mailbox did not run for 408 ms.Nothing in the shard or the mailbox lost that run. The test body runs on a thread-pool worker, the pool's worker floor on the two-core hosted agent is two, and
Thread.Sleepis the one form of blocking the pool cannot see. The shard's mailbox item had been queued to the other worker's local queue; with one worker asleep in the test and the other not returning to its dispatch loop, nothing could pick it up until the sleep expired, which is why the stall is 408 ms and not the gate thread's 500 ms starvation check. Both ends of the stall line up with the sleep to within a millisecond. The test actor's promptTerminatedproves nothing about pool health, because the TestKit's test actor runs on the calling-thread dispatcher and receives it inline on the dying entity's thread.The fixed sleep was also anchored to the wrong event even when the test passed. A wrongful restart timer would be armed when the shard processes the
Terminated, not when the test actor sees it, so the 400 ms window often did not cover the backoff at all. Pekko's copy of this spec has the same shape.Not related to #8574, which changed the remember-entities write timeout: the log shows the 5 s value in effect, and the writes here completed in about a millisecond.
Later commit
The second commit answers review. The first version read the state on the test actor with a 1 s bound per polling attempt, shorter than the region's 3 s query timeout, so a timed-out attempt could leave a stale reply that the passivation fact's final read or the non-remembering fact's
pong-2expect would dequeue. Queries now go through a fresh probe per attempt. The passivation fact's negative check moved from aTask.Delayplus a state read to the zero-count filter described above. Rebased onto dev.How it was checked
dotnet build src/contrib/cluster/Akka.Cluster.Sharding.Tests -c Release -warnaserrorclean. The three facts run ten times throughdotnet testafter each commit: 60 of 60 passed, about 5 s per run. The grep for synchronous TestKit calls, blocking waits, and sleeps on the file returns nothing. Test-only change. No ledger entry.