Skip to content

Serialization.V2: split the generator into phase files (move only, no behavior change) - #8519

Merged
Aaronontheweb merged 2 commits into
akkadotnet:devfrom
Aaronontheweb:feature/serialization-v2-generator-split
Sep 8, 2026
Merged

Aaronontheweb merged 2 commits into
akkadotnet:devfrom
Aaronontheweb:feature/serialization-v2-generator-split

Conversation

@Aaronontheweb

Copy link
Copy Markdown
Member

Summary

Part of #8384. A move-only split of the Akka.Serialization.V2 source generator, one file of about 4,900 lines, into six partial-class files by phase. No behavior change: no diagnostic id, title, severity, or message text changed, no pipeline stage changed, and the generated output is byte-identical.

This is the first step of the generator architecture pass (Decision H in the OpenSpec design record's working notes): make each phase findable so the next feature PRs each touch one file, and make the structural PRs that follow reviewable as diffs inside one file.

Layout

File Lines Holds
AkkaSerializerGenerator.cs 105 [Generator] entry point, Initialize, pipeline wiring, tracking names, and a comment that maps the phases to files
AkkaSerializerGenerator.Diagnostics.cs 368 every DiagnosticDescriptor, in id order
AkkaSerializerGenerator.Extraction.cs 1,095 symbol-to-model extraction, field-type mapping, KnownTypes, constructor matching
AkkaSerializerGenerator.Validation.cs 688 validation over collected models, the reachability walk, the Report* helpers
AkkaSerializerGenerator.Emission.cs 1,929 source emission, naming and folding, union helper planning
AkkaSerializerGenerator.Models.cs 842 the value-equatable model records

The only reordering is the diagnostics catalog, now in strict id order (AKKASG024 and AKKASG025 were declared in the other order).

Proof that it is a move

A script strips blank lines, using lines, #nullable, namespace and class declaration lines, braces, and file headers from both the original file and the concatenation of the six new files, sorts both, and diffs. Result: 4,217 lines on each side, identical.

Testing

  • dotnet build src/core/Akka.Serialization.V2.Generators -c Release -warnaserror: clean.
  • dotnet test src/core/Akka.Serialization.V2.Tests -c Release: 268 passed, 0 failed. No changes under GoldenOutput/ or WireSnapshots/.
  • GeneratorIncrementalCachingSpec: passes unchanged.
  • dotnet build src/core/Akka.Remote -c Release: clean.

What comes next

A structural pass, each PR behavior-neutral and gated by the golden and wire snapshot tests: a shared test harness with internal models, validation moved out of the emission callback, a cached per-serializer resolve stage, typed keys with one schema table, and a compilation-facts stage. Those give Decisions 16 through 21 a home with intact incremental caching.

@Aaronontheweb Aaronontheweb added serialization akka.net v1.6 Akka.NET v1.6-related issues labels Sep 8, 2026
@Aaronontheweb
Aaronontheweb merged commit 6cd8884 into akkadotnet:dev Sep 8, 2026
7 of 14 checks passed
Aaronontheweb added a commit that referenced this pull request Sep 9, 2026
…e allocating shards (#8524)

`ClusterShardingQueriesSpec` failed on all four nodes in the Artery Windows MNTR
lane of build 131139 (PR #8519, which only moved source-generator code). `third`
reported the primary failure after exhausting a full 30s converge-then-assert
budget:

  AwaitAssert failed, timeout [00:00:30] is over after [30] attempts
  Expected regions.Values.Select(i => i.Stats.Count).Sum() to be 4, but found 5.

`second` and `busy` then timed out 10s waiting for `ClusterShardingStats` (the
coordinator's fan-out waits on the region `third` had just torn down), and
`controller` failed the `received failed stats from timed out shards vs empty`
barrier.

Mechanism. The spec asserts a 2/2/2 shard layout (`timeouts = NumberOfShards /
regions.Count`, Stats == 4, Failed == 2, and later `Shards.HaveCount(2)` /
`Failed.HaveCount(2)` per region) but nothing guaranteed it. The `sharding
started` barrier only proves that each node called `StartSharding`; it says
nothing about which regions the coordinator has on its books. In the failing
run the coordinator on `third` was still reading its initial state from DData
when `second` and `busy` sent their first `Register` (20.318 and 20.339; the
state load completed at 20.354). `DDataShardCoordinator.WaitingForInitialState`
drops a `Register` rather than stashing it:

  DatatypeA: ShardRegion tried to register but ShardCoordinator not
    initialized yet: [[akka://...@localhost:54620/system/sharding/DatatypeA]]

and a region only re-sends on its `RegisterRetry` timer (250ms, doubling towards
`retry-interval`). `third`'s own region won the retry race (registered at
20.557), the controller's 20 pings reached the coordinator at 20.643, and
`LeastShardAllocationStrategy` could only pick from the regions it knew about:

  20.669  Shard [0] allocated at [.../DatatypeA#1012486927]        (third)
  20.739  Shard [5] allocated at [.../DatatypeA#1012486927]        (third)
  20.797  Shard [1] allocated at [.../DatatypeA#1012486927]        (third)
  20.833  Shard [3] allocated at [.../DatatypeA#1012486927]        (third)
  20.833  ShardRegion registered: [...@localhost:54620/...]        (second)
  20.840  ShardRegion registered: [...@localhost:54619/...]        (busy)
  20.957  Shard [4] allocated at [...@localhost:54620/...]         (second)
  20.991  Shard [2] allocated at [...@localhost:54619/...]         (busy)

4/1/1: five shards answer the stats query and only `busy`'s single shard fails
on its 0ms `shard-region-query-timeout`, hence "found 5". The spec disables
rebalancing (`rebalance-interval = 120s`), so that layout is permanent and the
`AwaitAssertAsync` loops from #8310 cannot converge - they re-read a state that
will never change.

The race is between two remote round trips - the coordinator's majority read of
its state and the regions' identification of the singleton - so the transport
only shifts the odds. Artery's first-contact path is shorter than DotNetty's
association handshake, which is why the remote Registers land that much earlier
relative to the state load in the Artery lane.

Fix. Before the controller sends the pings that trigger allocation, it polls
`GetCurrentRegions` through its proxy until the coordinator reports all three
regions - fresh probe and 1s bound per attempt inside a 30s dilated budget, the
same gate `ClusterShardingRolePartitioningSpec` already uses. Once the
coordinator knows all three regions the layout is 2/2/2 by construction: each
`ShardHomeAllocated` update stashes the next `GetShardHome`, so allocations are
sequential against fresh state and the least-shards ordering fills the regions
round-robin. The product is doing what it is designed to do (allocate among the
regions that have registered); the spec asserted a layout it had not waited
for.

Verified fail-first. The race does not open on its own on a Linux workstation
(10/10 green before the change, ~6s per run). A throwaway build that holds the
coordinator's initial-state result back by 500ms - so every first Register is
dropped, as in CI - and gives `busy` and `second` a 3s registration retry
reproduces the signature on the unfixed spec 5/5 times ("Expected ...
Stats.Count).Sum() to be 4, but found 6": all six shards on `third`), and
passes 5/5 with this gate on the same injected build. With the
injection removed: 20/20 green under Artery, and the full
Akka.Cluster.Sharding.Tests.MultiNode assembly passes under Artery
(137/137, 9m11s).

Test-only change.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

akka.net v1.6 Akka.NET v1.6-related issues serialization

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant