Repository navigation
De-flake RemotingTerminatorSpecs: don't depend on the best-effort graceful Terminated - #8622
Merged
Aaronontheweb merged 1 commit intoSep 24, 2026
Conversation
…rom a graceful shutdown RemotingTerminator_should_shutdown_properly_with_remotely_deployed_actor waited 10 s for the remote-deployed actor's Terminated after the deployed system terminated. That Terminated normally arrives in milliseconds over the graceful disassociation, but the delivery is best-effort: ShutdownAndFlush writes what's pending and closes without waiting for system-message acks. When it's lost (as on a 2-vCPU Windows agent), only the watch failure detector reports the death - after its 10 s acceptable pause plus the heartbeat interval, i.e. past the 10 s window (~14 s observed). - akka.remote.watch-failure-detector.acceptable-heartbeat-pause = 2s for this spec, so the guaranteed path lands well inside the 10 s windows. - WatchRemoteDeployedAsync: WatchAsync, then an Identify round trip with the local RemoteWatcher, so the remote Watch is on the wire before the Identify used to confirm the association (TestKit.Watch only waits for the TestActor, not the RemoteWatcher). Both tests with this shape get the change.
Aaronontheweb
commented
Sep 24, 2026
| private async Task WatchRemoteDeployedAsync(IActorRef remoteDeployed) | ||
| { | ||
| await WatchAsync(remoteDeployed); | ||
| await Sys.ActorSelection("/system/remote-watcher").Ask<ActorIdentity>(new Identify(null), RemainingOrDefault); |
Member
Author
There was a problem hiding this comment.
Good idea - the watch must have gone through in order for this to pass
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
RemotingTerminatorSpecs.RemotingTerminator_should_shutdown_properly_with_remotely_deployed_actorfailed on a 2-vCPU Windows agent (in a build of AOT-stack PR #8613, which doesn't touch remoting): it timed out after 10 s waiting for the remote-deployed actor'sTerminatedafterSystem2.Terminate().Cause
There are two ways the deployer can learn that the remote-deployed actor died.
DeathWatchNotificationsent over the association while System2 shuts down gracefully. The ordering is correct: the child notifies remote watchers first, and the notification reaches theEndpointManagerbeforeShutdownAndFlush. But delivery is best-effort:ShutdownAndFlushwrites whatever is pending and stops the writer.Endpoint.cs:776-786: "don't know if they were properly delivered").TcpAssociationHandle.Writedoesn't await the DotNetty write before the transport closes its channels.AddressTerminated, afteracceptable-heartbeat-pause(10 s by default) plus the heartbeat interval.In the failing run the fast path's notice never arrived. The failure detector fired about 14 s after shutdown began, past the test's 10 s window. The test had been widened once before, from 3 s to 10 s (#8328), which only made the window the same size as the pause.
I didn't find a product bug in the terminator's ordering. Classic remoting has no acked flush on shutdown, and I believe classic JVM Akka behaves the same way (not verified).
Change
akka.remote.watch-failure-detector.acceptable-heartbeat-pause = 2sin this spec's config. When the fast notice is lost, the guaranteed path now lands well inside the 10 s windows.WatchRemoteDeployedAsyncrunsWatchAsync, then anIdentifyround trip with the local/system/remote-watcher.TestKit.Watchonly waits for the TestActor. This round trip makes sure theRemoteWatcherhas sent its remoteWatchbefore the test'sIdentifyconfirms the association, which closes a smaller ordering race...._without_exception_logging_while_graceful_shutdownhas the same exposure.Verification (local)
RemotingTerminatorSpecspasses 4/4 across 3 runs, 3 s each.