Repository navigation
De-flake region-registration asserts in ClusterShardingRolePartitioningSpec - #8363
Merged
Aaronontheweb merged 1 commit intoJul 11, 2026
Conversation
…ngSpec The two region-count asserts on node Fourth used the shared TestActor mailbox with the default 5s expect per attempt, so only about six attempts fit in the 30s window and a late CurrentRegions reply from a timed-out attempt could pollute the next read. akkadotnet#8333 already widened the window to 30s and the flake still escaped on slow Windows MNTR agents. This switches both asserts to the ClusterShardingQueriesSpec pattern -- fresh probe per attempt, bounded 3s inner expect, explicit 1s retry interval -- which yields about thirty clean, independent samples in the same budget without widening anything.
Aaronontheweb
enabled auto-merge (squash)
July 11, 2026 19:24
Aaronontheweb
added a commit
that referenced
this pull request
Jul 12, 2026
…hAsync, add timefactor (#8372) * De-flake StressSpec: retry the aggregator Identify, fix blocking WatchAsync, add timefactor The per-phase aggregator lookup was a one-shot remote Identify with a 3s expect - the tightest of many undilated windows in this spec, and whichever window loses first names the run's failure. That lookup now retries with a fresh probe per attempt and a bounded inner expect, the same shape as the in-file watchee lookup and #8363. The watchee WatchAsync was an Ask bounded by the remaining outer Within budget, so one slow attempt consumed the entire retry window the enclosing AwaitAssert was supposed to spread across attempts, diverging from Pekko's non-blocking watch; it now carries an explicit small dilated bound. The watchee lookup also gained a null-Subject guard, since under load the selection can resolve to a null Subject and passing that into WatchAsync crashed the TestActor instead of letting the retry loop retry. Also sets single-expect-default 10s and timefactor 3, which reach the Dilated/single-expect paths but not raw Within bounds. * StressSpec: down-tune to Pekko parity (10 nodes, single-count churn phases) The .NET port drifted heavier than upstream: 13 nodes and double-count churn phases where Pekko runs 10 and 1. Re-converging on Pekko's reference block cuts both the load and the number of live timing windows on CI agents. nr-of-nodes-leaving also had to drop to Pekko's 1: at 10 nodes the Settings validation caps leaving+shutdown at 7 and the old value of 2 pushed the sum to 8, failing construction on every node.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes the Windows MNTR flake seen on #8357's validation run:
ClusterShardingMinMembersPerRoleNotConfiguredSpecfailing with "Expected value to be 3, but found 2" on node Fourth, which then cascades into node First's 3-secondInt32expect timing out (node Fourth dropping kills the R2 coordinator host).Root cause analysis: the region-count asserts polled through the shared TestActor mailbox with the default 5-second expect per attempt, so the 30-second window (already widened once by #8333, which is in this failure's base — widening is played out) only fits about six attempts, and a late
CurrentRegionsreply from a timed-out attempt can pollute the next attempt's read. The failing run just needed more clean samples of an eventually-consistent value, not more time.The fix applies the existing
ClusterShardingQueriesSpecpattern to both asserts: fresh probe per attempt, bounded 3-second inner expect, explicit 1-second retry interval. Same 30-second budget, roughly thirty clean independent samples instead of six polluted ones. Node First's block is untouched — it already uses a bounded inner expect, and fixing node Fourth removes the cascade that made it fail.Not caused by #8357: MNTR runs classic remoting (the env-var opt-in for Artery is unset in CI), and that PR's only classic-visible change is a field classic code never reads. Verified: spec passes 3x unconstrained and 1x under
taskset -c 0,1locally (~9-12s per run against the 30s budget).