Skip to content

fix: make ClusterShardingQueriesSpec stats/region-state queries converge-then-assert (async) - #8310

Merged
Aaronontheweb merged 2 commits into
akkadotnet:devfrom
Aaronontheweb:fix/clustershardingqueriesspec-converge-then-assert
Jul 3, 2026
Merged

Aaronontheweb merged 2 commits into
akkadotnet:devfrom
Aaronontheweb:fix/clustershardingqueriesspec-converge-then-assert

Conversation

@Aaronontheweb

Copy link
Copy Markdown
Member

Summary

Makes the flaky multi-node spec ClusterShardingQueriesSpec.Querying_cluster_sharding_specs stable by converting its query assertions to converge-then-assert and to async TestKit methods. Test-only — no production code under Akka.Cluster.Sharding/ is touched.

The flake

Observed CI failure:

Expected regions.Values.Select(i => i.Stats.Count).Sum() to be 4, but found 3

The GetClusterShardingStats (and GetShardRegionState) reads were one-shot: region.Tell(query) → probe.ExpectMsg<...>() → .Sum().Should().Be(4) with no retry. That read is a point-in-time snapshot that races sharding's internal shard-region-query-timeout (3s on the Second/Third regions). Under CI load a Second/Third Shard actor can momentarily miss that internal timeout and be reported in Failed instead of Stats, transiently undercounting the Stats sum from 4 to 3. It clears on the next query (warm mailbox).

The product is correct by design — sharding stats/state queries return partial results when a shard doesn't answer within the query timeout. The defect was purely the one-shot assertion racing convergence.

The fix

  • Wrap the stats query + both sum assertions, and each GetShardRegionState assertion block, in AwaitAssertAsync (converge-then-assert) that re-issues the query until the deterministic steady state is reached (Stats sum == 4, Failed sum == NumberOfShards / regions.Count).
  • Use a fresh TestProbe per attempt so a late ClusterShardingStats / CurrentShardRegionState reply from a prior iteration can't pollute the next ExpectMsg.
  • This mirrors the existing converge-then-assert pattern already used by the sibling ClusterShardingGetStatsSpec for the identical GetClusterShardingStats → Sum()==4 assertion.
  • Converted the spec to async TestKit (per project convention): the [MultiNodeFact] is now async Task, using RunOnAsync, AwaitClusterUpAsync, WithinAsync, AwaitAssertAsync, ExpectMsgAsync, ReceiveWhileAsync, and EnterBarrierAsync — no .Result/.Wait(), no synchronous ExpectMsg/AwaitAssert/Within.

Intentionally unchanged

  • Busy's shard-region-query-timeout = 0ms — this is the feature under test (its 2 shards deterministically report as Failed, keeping Stats sum at 4 and Failed sum at NumberOfShards / regions.Count). Kept as-is.
  • Entity/shard setup, NumberOfShards, the id % NumberOfShards mapping, and the rebalance / min-nr-of-members config are all unchanged.

Validation

This is a multi-node (MNTR) spec, which cannot be run locally without the multi-node test runner, so it was compile-verified only:

dotnet build src/contrib/cluster/Akka.Cluster.Sharding.Tests.MultiNode/Akka.Cluster.Sharding.Tests.MultiNode.csproj \
  -c Release -warnaserror --framework net10.0
# Build succeeded. 0 Warning(s) 0 Error(s)

The actual multi-node behavior is validated by CI's multi-node runner.

Backport

Candidate to backport to v1.5 (the same spec exists there).

…rge-then-assert (async)

The stats (GetClusterShardingStats) and region-state (GetShardRegionState)
assertions in ClusterShardingQueriesSpec were one-shot reads with no retry,
racing sharding's internal 3s shard-region-query-timeout on the Second/Third
regions. Under CI load a Shard actor can momentarily miss that timeout and be
reported in Failed instead of Stats, transiently undercounting the Stats sum
(4 -> 3) and failing the assertion. The product is correct by design (partial
results are expected); the defect was purely the one-shot assertion.

Wrap the stats query + both sum assertions, and each GetShardRegionState block,
in AwaitAssertAsync (converge-then-assert) with a fresh TestProbe per attempt so
a late reply can't pollute the next expect. Mirrors ClusterShardingGetStatsSpec.
Also converts the spec to async TestKit (RunOnAsync/AwaitAssertAsync/
ExpectMsgAsync/EnterBarrierAsync). Busy's 0ms shard-region-query-timeout (the
feature under test) is intentionally kept.

Test-only; multi-node spec so compile-verified only (CI validates behavior).
@Aaronontheweb
Aaronontheweb merged commit 4175b76 into akkadotnet:dev Jul 3, 2026
11 checks passed
@Aaronontheweb
Aaronontheweb deleted the fix/clustershardingqueriesspec-converge-then-assert branch July 3, 2026 14:18
Aaronontheweb added a commit that referenced this pull request Sep 9, 2026
…e allocating shards (#8524)

`ClusterShardingQueriesSpec` failed on all four nodes in the Artery Windows MNTR
lane of build 131139 (PR #8519, which only moved source-generator code). `third`
reported the primary failure after exhausting a full 30s converge-then-assert
budget:

  AwaitAssert failed, timeout [00:00:30] is over after [30] attempts
  Expected regions.Values.Select(i => i.Stats.Count).Sum() to be 4, but found 5.

`second` and `busy` then timed out 10s waiting for `ClusterShardingStats` (the
coordinator's fan-out waits on the region `third` had just torn down), and
`controller` failed the `received failed stats from timed out shards vs empty`
barrier.

Mechanism. The spec asserts a 2/2/2 shard layout (`timeouts = NumberOfShards /
regions.Count`, Stats == 4, Failed == 2, and later `Shards.HaveCount(2)` /
`Failed.HaveCount(2)` per region) but nothing guaranteed it. The `sharding
started` barrier only proves that each node called `StartSharding`; it says
nothing about which regions the coordinator has on its books. In the failing
run the coordinator on `third` was still reading its initial state from DData
when `second` and `busy` sent their first `Register` (20.318 and 20.339; the
state load completed at 20.354). `DDataShardCoordinator.WaitingForInitialState`
drops a `Register` rather than stashing it:

  DatatypeA: ShardRegion tried to register but ShardCoordinator not
    initialized yet: [[akka://...@localhost:54620/system/sharding/DatatypeA]]

and a region only re-sends on its `RegisterRetry` timer (250ms, doubling towards
`retry-interval`). `third`'s own region won the retry race (registered at
20.557), the controller's 20 pings reached the coordinator at 20.643, and
`LeastShardAllocationStrategy` could only pick from the regions it knew about:

  20.669  Shard [0] allocated at [.../DatatypeA#1012486927]        (third)
  20.739  Shard [5] allocated at [.../DatatypeA#1012486927]        (third)
  20.797  Shard [1] allocated at [.../DatatypeA#1012486927]        (third)
  20.833  Shard [3] allocated at [.../DatatypeA#1012486927]        (third)
  20.833  ShardRegion registered: [...@localhost:54620/...]        (second)
  20.840  ShardRegion registered: [...@localhost:54619/...]        (busy)
  20.957  Shard [4] allocated at [...@localhost:54620/...]         (second)
  20.991  Shard [2] allocated at [...@localhost:54619/...]         (busy)

4/1/1: five shards answer the stats query and only `busy`'s single shard fails
on its 0ms `shard-region-query-timeout`, hence "found 5". The spec disables
rebalancing (`rebalance-interval = 120s`), so that layout is permanent and the
`AwaitAssertAsync` loops from #8310 cannot converge - they re-read a state that
will never change.

The race is between two remote round trips - the coordinator's majority read of
its state and the regions' identification of the singleton - so the transport
only shifts the odds. Artery's first-contact path is shorter than DotNetty's
association handshake, which is why the remote Registers land that much earlier
relative to the state load in the Artery lane.

Fix. Before the controller sends the pings that trigger allocation, it polls
`GetCurrentRegions` through its proxy until the coordinator reports all three
regions - fresh probe and 1s bound per attempt inside a 30s dilated budget, the
same gate `ClusterShardingRolePartitioningSpec` already uses. Once the
coordinator knows all three regions the layout is 2/2/2 by construction: each
`ShardHomeAllocated` update stashes the next `GetShardHome`, so allocations are
sequential against fresh state and the least-shards ordering fills the regions
round-robin. The product is doing what it is designed to do (allocate among the
regions that have registered); the spec asserted a layout it had not waited
for.

Verified fail-first. The race does not open on its own on a Linux workstation
(10/10 green before the change, ~6s per run). A throwaway build that holds the
coordinator's initial-state result back by 500ms - so every first Register is
dropped, as in CI - and gives `busy` and `second` a 3s registration retry
reproduces the signature on the unfixed spec 5/5 times ("Expected ...
Stats.Count).Sum() to be 4, but found 6": all six shards on `third`), and
passes 5/5 with this gate on the same injected build. With the
injection removed: 20/20 green under Artery, and the full
Akka.Cluster.Sharding.Tests.MultiNode assembly passes under Artery
(137/137, 9m11s).

Test-only change.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant