Skip to content

De-flake region-registration asserts in ClusterShardingRolePartitioningSpec - #8363

Merged
Aaronontheweb merged 1 commit into
akkadotnet:devfrom
Aaronontheweb:fix/sharding-role-partitioning-mntr-flake
Jul 11, 2026
Merged

Aaronontheweb merged 1 commit into
akkadotnet:devfrom
Aaronontheweb:fix/sharding-role-partitioning-mntr-flake

Conversation

@Aaronontheweb

Copy link
Copy Markdown
Member

Fixes the Windows MNTR flake seen on #8357's validation run: ClusterShardingMinMembersPerRoleNotConfiguredSpec failing with "Expected value to be 3, but found 2" on node Fourth, which then cascades into node First's 3-second Int32 expect timing out (node Fourth dropping kills the R2 coordinator host).

Root cause analysis: the region-count asserts polled through the shared TestActor mailbox with the default 5-second expect per attempt, so the 30-second window (already widened once by #8333, which is in this failure's base — widening is played out) only fits about six attempts, and a late CurrentRegions reply from a timed-out attempt can pollute the next attempt's read. The failing run just needed more clean samples of an eventually-consistent value, not more time.

The fix applies the existing ClusterShardingQueriesSpec pattern to both asserts: fresh probe per attempt, bounded 3-second inner expect, explicit 1-second retry interval. Same 30-second budget, roughly thirty clean independent samples instead of six polluted ones. Node First's block is untouched — it already uses a bounded inner expect, and fixing node Fourth removes the cascade that made it fail.

Not caused by #8357: MNTR runs classic remoting (the env-var opt-in for Artery is unset in CI), and that PR's only classic-visible change is a field classic code never reads. Verified: spec passes 3x unconstrained and 1x under taskset -c 0,1 locally (~9-12s per run against the 30s budget).

…ngSpec

The two region-count asserts on node Fourth used the shared TestActor mailbox with the default 5s expect per attempt, so only about six attempts fit in the 30s window and a late CurrentRegions reply from a timed-out attempt could pollute the next read. akkadotnet#8333 already widened the window to 30s and the flake still escaped on slow Windows MNTR agents. This switches both asserts to the ClusterShardingQueriesSpec pattern -- fresh probe per attempt, bounded 3s inner expect, explicit 1s retry interval -- which yields about thirty clean, independent samples in the same budget without widening anything.

@Aaronontheweb Aaronontheweb left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

@Aaronontheweb
Aaronontheweb enabled auto-merge (squash) July 11, 2026 19:24
@Aaronontheweb
Aaronontheweb merged commit 30e749c into akkadotnet:dev Jul 11, 2026
11 checks passed
Aaronontheweb added a commit that referenced this pull request Jul 12, 2026
…hAsync, add timefactor (#8372)

* De-flake StressSpec: retry the aggregator Identify, fix blocking WatchAsync, add timefactor

The per-phase aggregator lookup was a one-shot remote Identify with a 3s expect - the tightest of many undilated windows in this spec, and whichever window loses first names the run's failure. That lookup now retries with a fresh probe per attempt and a bounded inner expect, the same shape as the in-file watchee lookup and #8363. The watchee WatchAsync was an Ask bounded by the remaining outer Within budget, so one slow attempt consumed the entire retry window the enclosing AwaitAssert was supposed to spread across attempts, diverging from Pekko's non-blocking watch; it now carries an explicit small dilated bound. The watchee lookup also gained a null-Subject guard, since under load the selection can resolve to a null Subject and passing that into WatchAsync crashed the TestActor instead of letting the retry loop retry. Also sets single-expect-default 10s and timefactor 3, which reach the Dilated/single-expect paths but not raw Within bounds.

* StressSpec: down-tune to Pekko parity (10 nodes, single-count churn phases)

The .NET port drifted heavier than upstream: 13 nodes and double-count churn phases where Pekko runs 10 and 1. Re-converging on Pekko's reference block cuts both the load and the number of live timing windows on CI agents. nr-of-nodes-leaving also had to drop to Pekko's 1: at 10 nodes the Settings validation caps leaving+shutdown at 7 and the old value of 2 pushed the sum to 8, failing construction on every node.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant