Repository navigation
fix: make ClusterShardingQueriesSpec stats/region-state queries converge-then-assert (async) - #8310
Merged
Aaronontheweb merged 2 commits intoJul 3, 2026
Conversation
…rge-then-assert (async) The stats (GetClusterShardingStats) and region-state (GetShardRegionState) assertions in ClusterShardingQueriesSpec were one-shot reads with no retry, racing sharding's internal 3s shard-region-query-timeout on the Second/Third regions. Under CI load a Shard actor can momentarily miss that timeout and be reported in Failed instead of Stats, transiently undercounting the Stats sum (4 -> 3) and failing the assertion. The product is correct by design (partial results are expected); the defect was purely the one-shot assertion. Wrap the stats query + both sum assertions, and each GetShardRegionState block, in AwaitAssertAsync (converge-then-assert) with a fresh TestProbe per attempt so a late reply can't pollute the next expect. Mirrors ClusterShardingGetStatsSpec. Also converts the spec to async TestKit (RunOnAsync/AwaitAssertAsync/ ExpectMsgAsync/EnterBarrierAsync). Busy's 0ms shard-region-query-timeout (the feature under test) is intentionally kept. Test-only; multi-node spec so compile-verified only (CI validates behavior).
Aaronontheweb
deleted the
fix/clustershardingqueriesspec-converge-then-assert
branch
July 3, 2026 14:18
Aaronontheweb
added a commit
that referenced
this pull request
Sep 9, 2026
…e allocating shards (#8524) `ClusterShardingQueriesSpec` failed on all four nodes in the Artery Windows MNTR lane of build 131139 (PR #8519, which only moved source-generator code). `third` reported the primary failure after exhausting a full 30s converge-then-assert budget: AwaitAssert failed, timeout [00:00:30] is over after [30] attempts Expected regions.Values.Select(i => i.Stats.Count).Sum() to be 4, but found 5. `second` and `busy` then timed out 10s waiting for `ClusterShardingStats` (the coordinator's fan-out waits on the region `third` had just torn down), and `controller` failed the `received failed stats from timed out shards vs empty` barrier. Mechanism. The spec asserts a 2/2/2 shard layout (`timeouts = NumberOfShards / regions.Count`, Stats == 4, Failed == 2, and later `Shards.HaveCount(2)` / `Failed.HaveCount(2)` per region) but nothing guaranteed it. The `sharding started` barrier only proves that each node called `StartSharding`; it says nothing about which regions the coordinator has on its books. In the failing run the coordinator on `third` was still reading its initial state from DData when `second` and `busy` sent their first `Register` (20.318 and 20.339; the state load completed at 20.354). `DDataShardCoordinator.WaitingForInitialState` drops a `Register` rather than stashing it: DatatypeA: ShardRegion tried to register but ShardCoordinator not initialized yet: [[akka://...@localhost:54620/system/sharding/DatatypeA]] and a region only re-sends on its `RegisterRetry` timer (250ms, doubling towards `retry-interval`). `third`'s own region won the retry race (registered at 20.557), the controller's 20 pings reached the coordinator at 20.643, and `LeastShardAllocationStrategy` could only pick from the regions it knew about: 20.669 Shard [0] allocated at [.../DatatypeA#1012486927] (third) 20.739 Shard [5] allocated at [.../DatatypeA#1012486927] (third) 20.797 Shard [1] allocated at [.../DatatypeA#1012486927] (third) 20.833 Shard [3] allocated at [.../DatatypeA#1012486927] (third) 20.833 ShardRegion registered: [...@localhost:54620/...] (second) 20.840 ShardRegion registered: [...@localhost:54619/...] (busy) 20.957 Shard [4] allocated at [...@localhost:54620/...] (second) 20.991 Shard [2] allocated at [...@localhost:54619/...] (busy) 4/1/1: five shards answer the stats query and only `busy`'s single shard fails on its 0ms `shard-region-query-timeout`, hence "found 5". The spec disables rebalancing (`rebalance-interval = 120s`), so that layout is permanent and the `AwaitAssertAsync` loops from #8310 cannot converge - they re-read a state that will never change. The race is between two remote round trips - the coordinator's majority read of its state and the regions' identification of the singleton - so the transport only shifts the odds. Artery's first-contact path is shorter than DotNetty's association handshake, which is why the remote Registers land that much earlier relative to the state load in the Artery lane. Fix. Before the controller sends the pings that trigger allocation, it polls `GetCurrentRegions` through its proxy until the coordinator reports all three regions - fresh probe and 1s bound per attempt inside a 30s dilated budget, the same gate `ClusterShardingRolePartitioningSpec` already uses. Once the coordinator knows all three regions the layout is 2/2/2 by construction: each `ShardHomeAllocated` update stashes the next `GetShardHome`, so allocations are sequential against fresh state and the least-shards ordering fills the regions round-robin. The product is doing what it is designed to do (allocate among the regions that have registered); the spec asserted a layout it had not waited for. Verified fail-first. The race does not open on its own on a Linux workstation (10/10 green before the change, ~6s per run). A throwaway build that holds the coordinator's initial-state result back by 500ms - so every first Register is dropped, as in CI - and gives `busy` and `second` a 3s registration retry reproduces the signature on the unfixed spec 5/5 times ("Expected ... Stats.Count).Sum() to be 4, but found 6": all six shards on `third`), and passes 5/5 with this gate on the same injected build. With the injection removed: 20/20 green under Artery, and the full Akka.Cluster.Sharding.Tests.MultiNode assembly passes under Artery (137/137, 9m11s). Test-only change.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Makes the flaky multi-node spec
ClusterShardingQueriesSpec.Querying_cluster_sharding_specsstable by converting its query assertions to converge-then-assert and to async TestKit methods. Test-only — no production code underAkka.Cluster.Sharding/is touched.The flake
Observed CI failure:
The
GetClusterShardingStats(andGetShardRegionState) reads were one-shot:region.Tell(query) → probe.ExpectMsg<...>() → .Sum().Should().Be(4)with no retry. That read is a point-in-time snapshot that races sharding's internalshard-region-query-timeout(3s on theSecond/Thirdregions). Under CI load aSecond/ThirdShardactor can momentarily miss that internal timeout and be reported inFailedinstead ofStats, transiently undercounting theStatssum from 4 to 3. It clears on the next query (warm mailbox).The product is correct by design — sharding stats/state queries return partial results when a shard doesn't answer within the query timeout. The defect was purely the one-shot assertion racing convergence.
The fix
GetShardRegionStateassertion block, inAwaitAssertAsync(converge-then-assert) that re-issues the query until the deterministic steady state is reached (Statssum == 4,Failedsum ==NumberOfShards / regions.Count).TestProbeper attempt so a lateClusterShardingStats/CurrentShardRegionStatereply from a prior iteration can't pollute the nextExpectMsg.ClusterShardingGetStatsSpecfor the identicalGetClusterShardingStats → Sum()==4assertion.[MultiNodeFact]is nowasync Task, usingRunOnAsync,AwaitClusterUpAsync,WithinAsync,AwaitAssertAsync,ExpectMsgAsync,ReceiveWhileAsync, andEnterBarrierAsync— no.Result/.Wait(), no synchronousExpectMsg/AwaitAssert/Within.Intentionally unchanged
Busy'sshard-region-query-timeout = 0ms— this is the feature under test (its 2 shards deterministically report asFailed, keepingStatssum at 4 andFailedsum atNumberOfShards / regions.Count). Kept as-is.NumberOfShards, theid % NumberOfShardsmapping, and the rebalance /min-nr-of-membersconfig are all unchanged.Validation
This is a multi-node (MNTR) spec, which cannot be run locally without the multi-node test runner, so it was compile-verified only:
The actual multi-node behavior is validated by CI's multi-node runner.
Backport
Candidate to backport to v1.5 (the same spec exists there).