Skip to content

Akka.Cluster: gossip removal tombstones - stop resurrecting removed members - #8484

Merged
Aaronontheweb merged 13 commits into
akkadotnet:devfrom
Aaronontheweb:fix/gossip-removal-tombstones
Aug 27, 2026
Merged

Aaronontheweb merged 13 commits into
akkadotnet:devfrom
Aaronontheweb:fix/gossip-removal-tombstones

Conversation

@Aaronontheweb

Copy link
Copy Markdown
Member

Closes #8467.

A member the leader has removed can be merged back into the gossip by a lagging peer that still holds it as Leaving. The merge infers removal from member status, and status cannot distinguish "was removed" from "has not heard yet". The resurrected member is dead, can never mark gossip seen, and Leaving does not skip convergence - the cluster stops converging permanently and only a full restart clears it. Observed live on CI with a 185ms window between Exiting and removal; same failure family as #1458 and #2584.

Fix: gossip carries removal tombstones. When the leader removes a member it records the member's unique address and a timestamp. The merge then answers the one-sided question directly: absent-with-tombstone stays removed; absent-without-tombstone is a lagging peer and the member is kept. Down and Exiting remain evidence of removal even without a tombstone, so existing behavior is preserved where it was already correct.

Wire: one new repeated field on Gossip at the long-vacant field 7, plus a two-field Tombstone message (addressIndex into the existing address table, epoch-millis timestamp). proto3 skips unknown fields, so rolling upgrades are safe in both directions; the protection simply applies between upgraded nodes. Tombstones expire on the leader after akka.cluster.prune-gossip-tombstones-after (default 24h - the bound that matters is partition duration, not wire cost).

Four correctness details worth reviewer attention, each with a test that fails if it regresses (verified by sabotage runs):

  1. The merged vector clock is pruned for tombstoned nodes - the clock entry resurrects through the union the same way the member does, and filtering members alone leaves the removed node in the clock forever.
  2. All four ReceiveGossip comparison branches union tombstones, and the clock-bump invariant they lean on is now documented where it lives (ClusterCoreDaemon.UpdateLatestGossip).
  3. LeaderActionsOnConvergence prunes expired tombstones on quiet converged ticks via a reference-equality no-op guard, so idle clusters do not bump the clock forever.
  4. The serializer's address table is built from members; tombstoned addresses are appended explicitly, with a round-trip test for the case where a tombstoned address differs from a live member only by uid.

Also first-class in the suite: the inverse test asserting a one-sided Leaving member without a tombstone is kept - that both proves the resurrection test bites and blocks the tempting-but-unsound variant of adding Leaving to the status drop-set, which deletes live joiners.

Tests: Akka.Cluster.Tests 403/403 (including with the new members/tombstones disjointness assertion armed); API approvals additive only (one PickHighestPriority overload, one ClusterSettings property); Akka.Cluster builds clean under -warnaserror.

Deliberately not included, per the frozen minimal scope: a config kill switch, a second tombstone write site, a multinode repro spec (unit-level coverage constructs the resurrection directly from gossip values, which is stronger), and any change to the other uses of RemoveUnreachableWithMemberStatus - removal gating and convergence-skip semantics are untouched.

…merging

A member status on its own cannot tell "this member was removed" apart from "this
node has not heard about it yet". Merge had only the status to go on, so a member
the leader had removed was put back whenever a lagging peer still held it as
Leaving. The resurrected member is dead, so it can never mark itself seen, and
convergence stops for good.

Gossip now carries a tombstone per removed node - keyed by UniqueAddress, stamped
with the epoch milliseconds of the removal. That is positive evidence a removal
happened, and it travels with the gossip so any peer can apply it.

- Gossip.RemoveAll strips a node from members, seen, reachability and the vector
  clock and records its tombstone in one step, so no gossip that breaks the
  invariants is ever built.
- Gossip.Merge unions tombstones first, then prunes the merged vector clock by
  them. Union alone is not enough: a clock entry for a removed node is resurrected
  by the merge exactly the way the member is.
- Gossip.MergeTombstones unions tombstones without touching anything else, for the
  gossip-reception branches that pick a whole gossip as the winner.
- Gossip.PruneTombstones drops expired entries and returns the same instance when
  it drops nothing, so a caller can skip the update by reference check.
- Member.PickHighestPriority gains a tombstone-aware overload. The one-sided drop
  condition is widened with OR, not replaced: every member dropped before is still
  dropped, and no input keeps a member the old code dropped.
- AssertInvariants now requires members and tombstones to be disjoint. Merge picks
  a two-sided member by status alone, which is only correct while that holds.

Timestamps order nothing; they only decide when a tombstone expires. Union keeps
the later timestamp on collision, so two nodes merging the same pair of gossips in
opposite order reach the same state.
Adds `repeated Tombstone tombstones = 7` to message Gossip, plus the Tombstone
message itself. Field 7 was the one vacant slot in Gossip - allAppVersions already
sits at 8 - so nothing needs renumbering. Codegen runs at build time; there is no
checked-in generated file to update.

The serializer builds its address table from members only, and a tombstoned node
is by definition not a member. Tombstone addresses are appended to that table
after the member loop, so a tombstone can index into it like everything else.
Getting this wrong is silent: it produces tombstones pointing at other nodes'
addresses rather than an error.

Gossip written before this field decodes to an empty tombstone set, which reduces
the merge to its previous status-only rule.
The leader records a tombstone for every node it removes, on both removal paths -
unreachable Down/Exiting members and confirmed-Exiting members. Both already ran
through the same block in LeaderActionsOnConvergence; that block now calls
Gossip.RemoveAll instead of deriving members, seen, reachability and the vector
clock separately.

ReceiveGossip unions tombstones on all four comparison branches. Only the
concurrent branch merges; the other three pick one whole gossip as the winner and
would otherwise drop the loser's tombstones, letting removals decay out of the
cluster over time. That failure mode presents as the original bug returning
intermittently, which no unit test would catch.

Expired tombstones are dropped on the leader at the end of the same method.
Pruning has to run on converged ticks where nothing else changed, so it sits
outside the change guard: the method now computes the updated gossip - or the
local one when nothing changed - prunes it, and publishes only when the result
differs from the local gossip by reference. PruneTombstones returns the same
instance when it drops nothing, which is what makes that check work. Without it
every converged leader tick would bump the vector clock and reset the seen table,
and the cluster would never sit still.

Retention is akka.cluster.prune-gossip-tombstones-after, default 24h. The bound
that matters is how long a gossip carrying the stale member can survive before it
merges back in, and under a partition that is unbounded - so the value is set
against partition length, not bandwidth. A partition outlasting the window
re-opens the hole silently.

UpdateLatestGossip now records the invariant it carries: every change this node
makes to the gossip goes through it and bumps the vector clock, which is what lets
peers tell whose removal set is newer.
GossipSpec
- The removal-resurrection case built straight from Gossip values: the leader has
  removed a member and holds its tombstone, a lagging peer still holds it as
  Leaving. The merge must not put it back, in either direction.
- The same setup without the tombstone, asserting the member IS kept. That case
  both proves the test above bites and pins the correct behaviour: with no evidence
  of a removal, a one-sided Leaving member may be a node the other side has not
  heard about yet, and dropping it would strand a live process.
- Union is commutative, and keeps the later timestamp on collision.
- The merged vector clock is pruned for every tombstoned node. Each side carries a
  clock entry for both removed nodes and a tombstone the other lacks; with only the
  member filter in place, this is the single test that fails.
- One-sided Down and Exiting members are still dropped with no tombstones present,
  so the OR did not turn into a replacement.
- A new incarnation at the same host and port survives its predecessor's tombstone.
- RemoveAll strips members, seen, reachability as observer and as subject, and the
  vector clock entry, while leaving other nodes' clock entries alone.
- Pruning boundary, and the same-instance identity that the leader's no-op check
  depends on.
- The three non-merge gossip comparison branches each keep both sides' tombstones,
  and the union refuses to adopt a tombstone for a node it still holds as a member.

ClusterMessageSerializerSpec
- Round trip with tombstones, over GossipEnvelope and over Welcome.
- A tombstoned address that shares a host and port with a member survives with its
  own UID, which is what catches a serializer resolving tombstones through the
  member address table. That bug is silent - it yields the wrong address, not an
  error.
- A gossip proto with field 7 cleared decodes to an empty tombstone set with
  members and version intact.

ClusterConfigSpec asserts the 24h default. The API approval files gain the two new
public members: the PickHighestPriority overload and ClusterSettings
.PruneGossipTombstonesAfter.

Also fixes Gossip.Prune, which rebuilt the gossip through the constructor and so
dropped tombstones on the pre-merge clock pruning path.
Wire addition to cluster gossip, the new retention setting, and the behavior
change: a removed member can no longer be resurrected by stale gossip.
@Aaronontheweb Aaronontheweb added akka-cluster akka.net v1.6 Akka.NET v1.6-related issues labels Aug 26, 2026
@Aaronontheweb Aaronontheweb added this to the 1.6.0 milestone Aug 26, 2026
The example specs pin the cases a human thought of. These sample the same code
over a few hundred to a few thousand random gossips per property and check laws:
merge is commutative, idempotent and associative; a tombstone always beats a
stale member; a removal never comes back over a random sequence of exchanges.

CsCheck 4.8.0, referenced from Akka.Cluster.Tests only. Fixed iteration counts,
so the class runs in about five seconds. A failure prints the seed to replay.

Generators draw from a bounded universe of six nodes, two of which share a host
and port and differ only by UID - the case that catches a serializer resolving a
tombstone through the member address table. Member statuses honour the allowed
transition table, tombstone timestamps come from a four value pool so collisions
are frequent, and every timestamp is passed in rather than read off the clock.

P1-P3   merge is commutative, idempotent and associative over members,
        tombstones with their timestamps, reachability and the vector clock.
P4-P6   a tombstoned node loses; a one-sided live member with no tombstone is
        kept; a one-sided Down or Exiting member is still dropped.
P7      disjointness, split by reception branch: the concurrent path lets the
        tombstone win, the winner-picked paths let the winner's member win.
P8      no tombstoned node keeps a vector clock entry through a merge.
P9      merged tombstones are the union, collisions take the later timestamp.
P10     PruneTombstones drops what expired and hands back the same instance
        otherwise, which the leader's no-op check depends on.
P11     RemoveAll writes a tombstone for every removal and strips it elsewhere.
P12     proto round trip preserves members, tombstones, reachability and clock.
P13-P14 random histories over three to five nodes exchanging gossip through the
        same branch selection ReceiveGossip uses, checked against a plain set of
        removed UIDs. Every node converges on that set, and no node ever loses a
        tombstone outside a prune.

Three properties carry a documented caveat where the law they check is narrower
than it first looks. Merge with itself also drops the clock entries of tombstoned
nodes, so P2 compares against that form. The one-sided drop for Down and Exiting
is not associative on its own, predating tombstones, so P3 draws non-terminal
statuses and lets tombstones carry the removals. Reachability merge breaks an
equal-version tie by argument order, so the two sides observe from disjoint
observer sets.

Each property counts the iterations that hit the case it is about and fails if
that count is too low, so none of them can pass by never generating anything
interesting.
@Aaronontheweb
Aaronontheweb marked this pull request as draft August 26, 2026 19:49
…ed branches

ReceiveGossip unioned the receiver's tombstones into the winning gossip on all
four comparison branches, while PruneTombstones is called from one place - the
leader's LeaderActionsOnConvergence. A prune was therefore undone by the next
exchange with any peer, because peers never prune and always re-unioned. So
prune-gossip-tombstones-after reclaimed nothing, tombstones rode every gossip
message forever, and once one was past the window the leader looped on every
tick: prune, clock bump, seen reset, re-converge, peers hand it back, prune
again. That is exactly the never-settles failure the reference-equality guard in
PruneTombstones exists to prevent.

Tombstones are now unioned only where neither gossip descends from the other -
the equal-clock branch, and the concurrent branch inside Gossip.Merge. The two
branches that pick a strictly newer gossip keep the winner's tombstones and
nothing else. The winner is not behind on removals: a removal bumps the removing
node's own clock entry, so a gossip that dominates the loser's clock descends
from every removal the loser knows about. The tombstones it does not carry are
the ones its own leader pruned, and honouring those prunes is the point.

The branch selection moved out of ReceiveGossip into
ClusterCoreDaemon.SelectWinningGossip, unchanged apart from the two lines above.
The property specs now call it instead of re-implementing it - they had their own
copy of the branches, so deleting the production change left them all green.

Three new properties:

P15  a tombstone pruned on a converged tick does not come back. The pruning node
     stamps the prune the way UpdateLatestGossip does, which puts it strictly
     ahead of every peer, so every exchange afterwards takes a winner-picked
     branch - the branch that used to hand the tombstones straight back.
P16  a simulated week of virtual time: sparse jumps forward, a constant per-node
     clock offset up to five minutes, and a prune tick run by a different drawn
     node each time, measured against the shipped prune-gossip-tombstones-after
     default rather than a cutoff invented for the test. Expired tombstones are
     reclaimed and stay reclaimed, and no tombstone is dropped sooner than the
     window minus the widest clock disagreement.
P17  a removal holds when the node that recorded it is itself removed later.
     That is where Merge prunes the clock entry that recorded the removal, so it
     is the one place a gossip could come out strictly newer than a tombstone
     carrier without having descended from the removal.

P15 and P16 both fail if the union is restored on the Before and After branches.
P17 documents a resurrection the model can reach and a cluster cannot, because
removing a member is a leader action and the leader needs convergence first;
the same history resurrects the member with the union restored, so it is not
something the union was buying.

Also in the suite: the seen table joins the equivalence helper; reachability is
built from one shared op history per observer with each side replaying a prefix
of it, so sides share observers at different versions and Reachability.Merge's
version arbitration is actually exercised; histories gained a status op so peers
hold a victim at mixed statuses while its tombstone propagates; every op asserts
the gossip carries no clock entry for a non-member, mirroring the "too many
vector clock entries" check; P4, P7c and P11 assert no reachability record
survives whose subject was removed; the serializer property gained a sample
where a tombstone key is also a member, which is the arm of the address-table
writer nothing else reached; coverage guards count iterations rather than
occurrences; and the properties whose inputs the Gossip constructor rejects under
AKKA_CLUSTER_ASSERT=on skip there instead of failing.
@Aaronontheweb
Aaronontheweb marked this pull request as ready for review August 26, 2026 20:54

@Aaronontheweb Aaronontheweb left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

# back into the gossip. If the cluster is partitioned for longer than this, a
# removed node can be re-added by stale gossip, so keep this well above the
# longest partition the cluster is expected to heal from.
prune-gossip-tombstones-after = 24h

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

repeated Member members = 4;
GossipOverview overview = 5;
VectorClock version = 6;
repeated Tombstone tombstones = 7;

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM - extend-only design

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

// A tombstoned node is never a member, so its address is missing from the table the
// member loop just built. Append it, so the tombstone can index into the same table.
var tombstoneProtos = new List<Proto.Msg.Tombstone>(gossip.Tombstones.Count);
foreach (var t in gossip.Tombstones)

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

/// gossip. A partition that outlasts this window re-opens that hole, so the value should sit well
/// above the longest partition the cluster is expected to heal from. Defaults to 24 hours.
/// </summary>
public TimeSpan PruneGossipTombstonesAfter { get; }

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This needs to be asserted in our cluster settings spec, no? The default values here?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Already covered on the branch - ClusterConfigSpec.cs:40 asserts the 24h default (added with the prune fix commit, after your review snapshot). The ctor overload and the TBDs are being addressed now.

Comment thread src/core/Akka.Cluster/Gossip.cs Outdated
public Gossip(ImmutableSortedSet<Member> members, GossipOverview overview, VectorClock version)
: this(members, overview, version, EmptyTombstones) { }

/// <summary>

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

No new TBDs - we either need to use inheritdoc to reference previous constructors or we need to say what the variables do. In fact, we should clean up the TBDs in this file while we're at it.

/// <param name="version">TBD</param>
/// <param name="tombstones">The nodes this cluster has removed. See <see cref="Tombstones"/>.</param>
/// <exception cref="ArgumentException">TBD</exception>
public Gossip(ImmutableSortedSet<Member> members, GossipOverview overview, VectorClock version,

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we really need a new CTOR overload? This is an internal type, so there's no API risk here.

Comment thread src/core/Akka.Cluster/ClusterDaemon.cs Outdated

switch (comparison)
{
// Three of the four branches below pick one whole gossip as the winner, which on its own

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM

…t the file

Addresses two review comments on akkadotnet#8484.

Drop the four-argument Gossip constructor added by the tombstone work.
Gossip is internal, so there is no API risk in changing the existing
three-argument constructor instead: tombstones becomes an optional
parameter that defaults to no tombstones. Every existing call site
compiles unchanged, including the ones that pass tombstones positionally.

Replace all 62 TBD placeholders in Gossip.cs with docs that say what each
member actually does. The delegating constructors carry a one-line summary
noting their defaults and inherit the rest from the primary. Notable
details now written down: GetMember hands back a Removed placeholder for a
node that is not in the ring, AddMember is keyed by address so it will not
replace a member that differs only in status, YoungestMember treats a
member that is not up yet as up-number zero, and Prune clears a node's
vector clock entry without touching members or tombstones.
@Aaronontheweb
Aaronontheweb enabled auto-merge (squash) August 26, 2026 21:57
@Aaronontheweb

Copy link
Copy Markdown
Member Author

Bulk of the commits and lines here are some new CsCheck specs i added to enforce all sorts of Gossip properties. Found some issues with the current PR that we subsequently resolved as a result.

@Aaronontheweb

Copy link
Copy Markdown
Member Author

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines will not run the associated pipelines, because the pull request was updated after the run command was issued. Review the pull request again and issue a new run command.

# Conflicts:
#	BREAKING_CHANGES_V1.6.md
@Aaronontheweb

Copy link
Copy Markdown
Member Author

/azp run

@azure-pipelines

Copy link
Copy Markdown
Azure Pipelines successfully started running 1 pipeline(s).

@Aaronontheweb
Aaronontheweb merged commit b86a6a1 into akkadotnet:dev Aug 27, 2026
14 of 15 checks passed
@Aaronontheweb
Aaronontheweb deleted the fix/gossip-removal-tombstones branch August 27, 2026 21:02
Aaronontheweb added a commit that referenced this pull request Oct 2, 2026
* Akka.Cluster: gossip removal tombstones - stop resurrecting removed members (#8484)

* Akka.Cluster: add removal tombstones to Gossip and consult them when merging

A member status on its own cannot tell "this member was removed" apart from "this
node has not heard about it yet". Merge had only the status to go on, so a member
the leader had removed was put back whenever a lagging peer still held it as
Leaving. The resurrected member is dead, so it can never mark itself seen, and
convergence stops for good.

Gossip now carries a tombstone per removed node - keyed by UniqueAddress, stamped
with the epoch milliseconds of the removal. That is positive evidence a removal
happened, and it travels with the gossip so any peer can apply it.

- Gossip.RemoveAll strips a node from members, seen, reachability and the vector
  clock and records its tombstone in one step, so no gossip that breaks the
  invariants is ever built.
- Gossip.Merge unions tombstones first, then prunes the merged vector clock by
  them. Union alone is not enough: a clock entry for a removed node is resurrected
  by the merge exactly the way the member is.
- Gossip.MergeTombstones unions tombstones without touching anything else, for the
  gossip-reception branches that pick a whole gossip as the winner.
- Gossip.PruneTombstones drops expired entries and returns the same instance when
  it drops nothing, so a caller can skip the update by reference check.
- Member.PickHighestPriority gains a tombstone-aware overload. The one-sided drop
  condition is widened with OR, not replaced: every member dropped before is still
  dropped, and no input keeps a member the old code dropped.
- AssertInvariants now requires members and tombstones to be disjoint. Merge picks
  a two-sided member by status alone, which is only correct while that holds.

Timestamps order nothing; they only decide when a tombstone expires. Union keeps
the later timestamp on collision, so two nodes merging the same pair of gossips in
opposite order reach the same state.

* Akka.Cluster: put gossip tombstones on the wire

Adds `repeated Tombstone tombstones = 7` to message Gossip, plus the Tombstone
message itself. Field 7 was the one vacant slot in Gossip - allAppVersions already
sits at 8 - so nothing needs renumbering. Codegen runs at build time; there is no
checked-in generated file to update.

The serializer builds its address table from members only, and a tombstoned node
is by definition not a member. Tombstone addresses are appended to that table
after the member loop, so a tombstone can index into it like everything else.
Getting this wrong is silent: it produces tombstones pointing at other nodes'
addresses rather than an error.

Gossip written before this field decodes to an empty tombstone set, which reduces
the merge to its previous status-only rule.

* Akka.Cluster: write tombstones on removal and expire them on the leader

The leader records a tombstone for every node it removes, on both removal paths -
unreachable Down/Exiting members and confirmed-Exiting members. Both already ran
through the same block in LeaderActionsOnConvergence; that block now calls
Gossip.RemoveAll instead of deriving members, seen, reachability and the vector
clock separately.

ReceiveGossip unions tombstones on all four comparison branches. Only the
concurrent branch merges; the other three pick one whole gossip as the winner and
would otherwise drop the loser's tombstones, letting removals decay out of the
cluster over time. That failure mode presents as the original bug returning
intermittently, which no unit test would catch.

Expired tombstones are dropped on the leader at the end of the same method.
Pruning has to run on converged ticks where nothing else changed, so it sits
outside the change guard: the method now computes the updated gossip - or the
local one when nothing changed - prunes it, and publishes only when the result
differs from the local gossip by reference. PruneTombstones returns the same
instance when it drops nothing, which is what makes that check work. Without it
every converged leader tick would bump the vector clock and reset the seen table,
and the cluster would never sit still.

Retention is akka.cluster.prune-gossip-tombstones-after, default 24h. The bound
that matters is how long a gossip carrying the stale member can survive before it
merges back in, and under a partition that is unbounded - so the value is set
against partition length, not bandwidth. A partition outlasting the window
re-opens the hole silently.

UpdateLatestGossip now records the invariant it carries: every change this node
makes to the gossip goes through it and bumps the vector clock, which is what lets
peers tell whose removal set is newer.

* Akka.Cluster: tests for gossip removal tombstones

GossipSpec
- The removal-resurrection case built straight from Gossip values: the leader has
  removed a member and holds its tombstone, a lagging peer still holds it as
  Leaving. The merge must not put it back, in either direction.
- The same setup without the tombstone, asserting the member IS kept. That case
  both proves the test above bites and pins the correct behaviour: with no evidence
  of a removal, a one-sided Leaving member may be a node the other side has not
  heard about yet, and dropping it would strand a live process.
- Union is commutative, and keeps the later timestamp on collision.
- The merged vector clock is pruned for every tombstoned node. Each side carries a
  clock entry for both removed nodes and a tombstone the other lacks; with only the
  member filter in place, this is the single test that fails.
- One-sided Down and Exiting members are still dropped with no tombstones present,
  so the OR did not turn into a replacement.
- A new incarnation at the same host and port survives its predecessor's tombstone.
- RemoveAll strips members, seen, reachability as observer and as subject, and the
  vector clock entry, while leaving other nodes' clock entries alone.
- Pruning boundary, and the same-instance identity that the leader's no-op check
  depends on.
- The three non-merge gossip comparison branches each keep both sides' tombstones,
  and the union refuses to adopt a tombstone for a node it still holds as a member.

ClusterMessageSerializerSpec
- Round trip with tombstones, over GossipEnvelope and over Welcome.
- A tombstoned address that shares a host and port with a member survives with its
  own UID, which is what catches a serializer resolving tombstones through the
  member address table. That bug is silent - it yields the wrong address, not an
  error.
- A gossip proto with field 7 cleared decodes to an empty tombstone set with
  members and version intact.

ClusterConfigSpec asserts the 24h default. The API approval files gain the two new
public members: the PickHighestPriority overload and ClusterSettings
.PruneGossipTombstonesAfter.

Also fixes Gossip.Prune, which rebuilt the gossip through the constructor and so
dropped tombstones on the pre-merge clock pruning path.

* Record gossip removal tombstones in the v1.6 breaking changes ledger

Wire addition to cluster gossip, the new retention setting, and the behavior
change: a removed member can no longer be resurrected by stale gossip.

* Akka.Cluster: property-based specs for gossip removal tombstones

The example specs pin the cases a human thought of. These sample the same code
over a few hundred to a few thousand random gossips per property and check laws:
merge is commutative, idempotent and associative; a tombstone always beats a
stale member; a removal never comes back over a random sequence of exchanges.

CsCheck 4.8.0, referenced from Akka.Cluster.Tests only. Fixed iteration counts,
so the class runs in about five seconds. A failure prints the seed to replay.

Generators draw from a bounded universe of six nodes, two of which share a host
and port and differ only by UID - the case that catches a serializer resolving a
tombstone through the member address table. Member statuses honour the allowed
transition table, tombstone timestamps come from a four value pool so collisions
are frequent, and every timestamp is passed in rather than read off the clock.

P1-P3   merge is commutative, idempotent and associative over members,
        tombstones with their timestamps, reachability and the vector clock.
P4-P6   a tombstoned node loses; a one-sided live member with no tombstone is
        kept; a one-sided Down or Exiting member is still dropped.
P7      disjointness, split by reception branch: the concurrent path lets the
        tombstone win, the winner-picked paths let the winner's member win.
P8      no tombstoned node keeps a vector clock entry through a merge.
P9      merged tombstones are the union, collisions take the later timestamp.
P10     PruneTombstones drops what expired and hands back the same instance
        otherwise, which the leader's no-op check depends on.
P11     RemoveAll writes a tombstone for every removal and strips it elsewhere.
P12     proto round trip preserves members, tombstones, reachability and clock.
P13-P14 random histories over three to five nodes exchanging gossip through the
        same branch selection ReceiveGossip uses, checked against a plain set of
        removed UIDs. Every node converges on that set, and no node ever loses a
        tombstone outside a prune.

Three properties carry a documented caveat where the law they check is narrower
than it first looks. Merge with itself also drops the clock entries of tombstoned
nodes, so P2 compares against that form. The one-sided drop for Down and Exiting
is not associative on its own, predating tombstones, so P3 draws non-terminal
statuses and lets tombstones carry the removals. Reachability merge breaks an
equal-version tie by argument order, so the two sides observe from disjoint
observer sets.

Each property counts the iterations that hit the case it is about and fails if
that count is too low, so none of them can pass by never generating anything
interesting.

* Akka.Cluster: keep only the winner's tombstones on the causally-ordered branches

ReceiveGossip unioned the receiver's tombstones into the winning gossip on all
four comparison branches, while PruneTombstones is called from one place - the
leader's LeaderActionsOnConvergence. A prune was therefore undone by the next
exchange with any peer, because peers never prune and always re-unioned. So
prune-gossip-tombstones-after reclaimed nothing, tombstones rode every gossip
message forever, and once one was past the window the leader looped on every
tick: prune, clock bump, seen reset, re-converge, peers hand it back, prune
again. That is exactly the never-settles failure the reference-equality guard in
PruneTombstones exists to prevent.

Tombstones are now unioned only where neither gossip descends from the other -
the equal-clock branch, and the concurrent branch inside Gossip.Merge. The two
branches that pick a strictly newer gossip keep the winner's tombstones and
nothing else. The winner is not behind on removals: a removal bumps the removing
node's own clock entry, so a gossip that dominates the loser's clock descends
from every removal the loser knows about. The tombstones it does not carry are
the ones its own leader pruned, and honouring those prunes is the point.

The branch selection moved out of ReceiveGossip into
ClusterCoreDaemon.SelectWinningGossip, unchanged apart from the two lines above.
The property specs now call it instead of re-implementing it - they had their own
copy of the branches, so deleting the production change left them all green.

Three new properties:

P15  a tombstone pruned on a converged tick does not come back. The pruning node
     stamps the prune the way UpdateLatestGossip does, which puts it strictly
     ahead of every peer, so every exchange afterwards takes a winner-picked
     branch - the branch that used to hand the tombstones straight back.
P16  a simulated week of virtual time: sparse jumps forward, a constant per-node
     clock offset up to five minutes, and a prune tick run by a different drawn
     node each time, measured against the shipped prune-gossip-tombstones-after
     default rather than a cutoff invented for the test. Expired tombstones are
     reclaimed and stay reclaimed, and no tombstone is dropped sooner than the
     window minus the widest clock disagreement.
P17  a removal holds when the node that recorded it is itself removed later.
     That is where Merge prunes the clock entry that recorded the removal, so it
     is the one place a gossip could come out strictly newer than a tombstone
     carrier without having descended from the removal.

P15 and P16 both fail if the union is restored on the Before and After branches.
P17 documents a resurrection the model can reach and a cluster cannot, because
removing a member is a leader action and the leader needs convergence first;
the same history resurrects the member with the union restored, so it is not
something the union was buying.

Also in the suite: the seen table joins the equivalence helper; reachability is
built from one shared op history per observer with each side replaying a prefix
of it, so sides share observers at different versions and Reachability.Merge's
version arbitration is actually exercised; histories gained a status op so peers
hold a victim at mixed statuses while its tombstone propagates; every op asserts
the gossip carries no clock entry for a non-member, mirroring the "too many
vector clock entries" check; P4, P7c and P11 assert no reachability record
survives whose subject was removed; the serializer property gained a sample
where a tombstone key is also a member, which is the arm of the address-table
writer nothing else reached; coverage guards count iterations rather than
occurrences; and the properties whose inputs the Gossip constructor rejects under
AKKA_CLUSTER_ASSERT=on skip there instead of failing.

* ci: retrigger after superseded-run cancellations

* Akka.Cluster: fold tombstones into the Gossip constructor and document the file

Addresses two review comments on #8484.

Drop the four-argument Gossip constructor added by the tombstone work.
Gossip is internal, so there is no API risk in changing the existing
three-argument constructor instead: tombstones becomes an optional
parameter that defaults to no tombstones. Every existing call site
compiles unchanged, including the ones that pass tombstones positionally.

Replace all 62 TBD placeholders in Gossip.cs with docs that say what each
member actually does. The delegating constructors carry a one-line summary
noting their defaults and inherit the rest from the primary. Notable
details now written down: GetMember hands back a Removed placeholder for a
node that is not in the ring, AddMember is keyed by address so it will not
replace a member that differs only in status, YoungestMember treats a
member that is not up yet as up-number zero, and Prune clears a node's
vector clock entry without touching members or tombstones.

* Adapt gossip tombstone tests for v1.5 targets

* Rerun CI after flaky net48 test
This was referenced Oct 4, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

akka.net v1.6 Akka.NET v1.6-related issues akka-cluster

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Akka.Cluster: removed member resurrected by gossip merge, permanently blocking convergence

1 participant