Fix ClusterClientDiscovery: preserve contact-point subscriptions across rediscovery - #8426
Merged
Aaronontheweb merged 2 commits intoJul 24, 2026
Conversation
…ss rediscovery ClusterClientDiscovery supervises the real ClusterClient as a child and recreates that child from scratch on every rediscovery (when the current child exhausts its reconnect-timeout and stops). The contact-point subscriber list lives on the child, so a SubscribeContactPoints subscriber silently stopped receiving ContactPoints/ContactPointAdded/ContactPointRemoved once the client rediscovered — the fresh child had no record of it. Track contact-point subscribers at the supervisor level and re-subscribe them onto each newly-created child (on the subscriber's behalf, so the child registers them and replies with the current snapshot). DeathWatch each subscriber so a terminated subscriber is dropped from tracking (auto-unsubscribe). Subscribe/Unsubscribe are still forwarded to the current child, so live behavior is unchanged; only the across-rediscovery gap is closed.
…ents Replace the three GetContactPoints polling blocks (each with an orphaned, hand-tightened 1s inner ExpectMsg that spuriously timed out on loaded CI agents) with a single event-driven helper: subscribe via SubscribeContactPoints and wait on the client's own ContactPoints/ContactPointAdded/ContactPointRemoved stream until the contact points settle to exactly the expected node. No per-attempt reply timeout to race. This also covers the ClusterClientDiscovery subscription-survival fix: the second and third phases exercise rediscovery after a graceful down and a hard shutdown, so the subscription must survive the child being recreated for the helper to observe the new node. Verified locally, 4/4 across repeated runs.
Aaronontheweb
deleted the
fix/clusterclient-discovery-subscription-survival
branch
July 24, 2026 19:15
Aaronontheweb
added a commit
that referenced
this pull request
Aug 11, 2026
…ss rediscovery (#8426) * Fix ClusterClientDiscovery: preserve contact-point subscriptions across rediscovery ClusterClientDiscovery supervises the real ClusterClient as a child and recreates that child from scratch on every rediscovery (when the current child exhausts its reconnect-timeout and stops). The contact-point subscriber list lives on the child, so a SubscribeContactPoints subscriber silently stopped receiving ContactPoints/ContactPointAdded/ContactPointRemoved once the client rediscovered — the fresh child had no record of it. Track contact-point subscribers at the supervisor level and re-subscribe them onto each newly-created child (on the subscriber's behalf, so the child registers them and replies with the current snapshot). DeathWatch each subscriber so a terminated subscriber is dropped from tracking (auto-unsubscribe). Subscribe/Unsubscribe are still forwarded to the current child, so live behavior is unchanged; only the across-rediscovery gap is closed. * ClusterClientDiscoverySpec: verify contact points via subscription events Replace the three GetContactPoints polling blocks (each with an orphaned, hand-tightened 1s inner ExpectMsg that spuriously timed out on loaded CI agents) with a single event-driven helper: subscribe via SubscribeContactPoints and wait on the client's own ContactPoints/ContactPointAdded/ContactPointRemoved stream until the contact points settle to exactly the expected node. No per-attempt reply timeout to race. This also covers the ClusterClientDiscovery subscription-survival fix: the second and third phases exercise rediscovery after a graceful down and a hard shutdown, so the subscription must survive the child being recreated for the helper to observe the new node. Verified locally, 4/4 across repeated runs.
This was referenced Aug 27, 2026
Closed
Closed
This was referenced Aug 30, 2026
This was referenced Sep 9, 2026
This was referenced Sep 20, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #8425.
Problem
With initial-contacts discovery enabled,
ClusterClientDiscoverysupervises the realClusterClientas a child and recreates that child from scratch on every rediscovery (when the current child exhausts itsreconnect-timeoutand stops). The contact-point subscriber list lives on the child, and the supervisor kept no record of it — so aSubscribeContactPointssubscriber silently stopped receivingContactPoints/ContactPointAdded/ContactPointRemovedevents once the client rediscovered.GetContactPointswas unaffected (the supervisor forwards each query to the current child); only the durable subscription was dropped.Fix
ClusterClientDiscoverynow tracks contact-point subscribers at the supervisor level and re-establishes them on each newly-created child:SubscribeContactPoints/UnsubscribeContactPointsare recorded in a supervisor-level subscriber set (and DeathWatched, so a terminated subscriber is dropped — auto-unsubscribe), then still forwarded to the current child so live behavior is unchanged.SubscribeContactPointson each tracked subscriber's behalf, so the child registers them and replies with the current snapshot.No public API change; internal supervisor behavior only.
Testing
ClusterClientDiscoverySpecis reworked to verify contact points through the client's own subscription event stream (SubscribeContactPoints→ContactPoints/ContactPointAdded/ContactPointRemoved) instead of pollingGetContactPoints. Its second and third phases exercise rediscovery after a graceful down and a hard shutdown, so the subscription must survive the child being recreated for the test to observe the new node — i.e. the spec both de-flakes the old orphaned inner-timeout poll and directly covers this fix. Verified locally, 4/4 across repeated runs.