Repository navigation
Classic remoting: half-close on disassociate so a Windows peer doesn't lose our last frames to an RST - #8635
Merged
Aaronontheweb merged 7 commits intoSep 25, 2026
Conversation
… the peer (akkadotnet#8589) Drives TcpTransport against a raw socket peer whose bytes we never read. On dev the disassociate closes over that unread data, so the peer's SO_ERROR reports a reset (10058) and Write still returns true after Disassociate.
…that can RST (akkadotnet#8589) DotNetty's close is Socket.Shutdown(Both) + Dispose. With unread inbound data that sends an RST, and a Windows peer then drops frames it has not read yet - the lost ClusterSingleton HandOverDone in akkadotnet#8589. TcpAssociationHandle.Disassociate now refuses further writes, waits for the last write to reach the kernel, shuts down only the send side, and keeps reading (and discarding) until the peer's FIN, when DotNetty closes the channel over an empty buffer. A force-close on the channel's event loop bounds this by akka.remote.flush-wait-on-shutdown, and DotNettyTransport.Shutdown waits up to that long for pending graceful closes before its force-close. The send-side shutdown goes through a second Socket over the same handle: TcpSocketChannel.ShutdownOutputAsync calls Socket.Shutdown, which clears Socket.Connected, and DotNetty's DoReadBytes then returns EOF on the next read and closes over the unread data anyway.
…se fallbacks (akkadotnet#8589) Remoting starts the transport Shutdown once the endpoint writers have stopped, but the protocol actor handles its DisassociateUnderlying asynchronously, so a graceful close may not have started yet when Shutdown looks. Tracking only the channels already draining missed those and force-closed them - the akkadotnet#8589 reset. Shutdown now waits up to flush-wait on the close of every non-server channel before its force-close. Also: log a warning when the half-close falls back to a full close, and once per transport if DotNetty's socket field is missing; fall back to CloseAsync if the event loop rejects the scheduled work.
…wo-system terminate, and the reflected DotNetty field The shutdown-before-disassociate case fails with SO_ERROR 10058 against the previous commit, which only waited on channels already draining.
…akkadotnet#8589) Waiting on every association channel cost a full flush-wait whenever a channel was never disassociated - the throttler adapter never disassociates its TCP handle, so every throttled spec and multi-node node paid 2 s more. The race it guarded against can't happen through the protocol transport: ActorTransportAdapter.Shutdown stops the protocol manager first, and each ProtocolStateActor's OnTermination disassociates its handle, which now registers the channel synchronously before any async hop. Drops the spec case that called Shutdown before Disassociate directly on TcpTransport, an ordering the protocol transport can't produce.
Aaronontheweb
marked this pull request as ready for review
September 24, 2026 20:34
Aaronontheweb
enabled auto-merge (squash)
September 24, 2026 21:25
This was referenced Sep 25, 2026
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Oct 2, 2026
…t lose our last frames to an RST (akkadotnet#8635) * Add DotNettyGracefulCloseSpec: a graceful disassociate must not reset the peer (akkadotnet#8589) Drives TcpTransport against a raw socket peer whose bytes we never read. On dev the disassociate closes over that unread data, so the peer's SO_ERROR reports a reset (10058) and Write still returns true after Disassociate. * Classic remoting: half-close on disassociate instead of a full close that can RST (akkadotnet#8589) DotNetty's close is Socket.Shutdown(Both) + Dispose. With unread inbound data that sends an RST, and a Windows peer then drops frames it has not read yet - the lost ClusterSingleton HandOverDone in akkadotnet#8589. TcpAssociationHandle.Disassociate now refuses further writes, waits for the last write to reach the kernel, shuts down only the send side, and keeps reading (and discarding) until the peer's FIN, when DotNetty closes the channel over an empty buffer. A force-close on the channel's event loop bounds this by akka.remote.flush-wait-on-shutdown, and DotNettyTransport.Shutdown waits up to that long for pending graceful closes before its force-close. The send-side shutdown goes through a second Socket over the same handle: TcpSocketChannel.ShutdownOutputAsync calls Socket.Shutdown, which clears Socket.Connected, and DotNetty's DoReadBytes then returns EOF on the next read and closes over the unread data anyway. * Wait on every association channel in transport Shutdown; log half-close fallbacks (akkadotnet#8589) Remoting starts the transport Shutdown once the endpoint writers have stopped, but the protocol actor handles its DisassociateUnderlying asynchronously, so a graceful close may not have started yet when Shutdown looks. Tracking only the channels already draining missed those and force-closed them - the akkadotnet#8589 reset. Shutdown now waits up to flush-wait on the close of every non-server channel before its force-close. Also: log a warning when the half-close falls back to a full close, and once per transport if DotNetty's socket field is missing; fall back to CloseAsync if the event loop rejects the scheduled work. * DotNettyGracefulCloseSpec: cover shutdown-before-disassociate, TLS, two-system terminate, and the reflected DotNetty field The shutdown-before-disassociate case fails with SO_ERROR 10058 against the previous commit, which only waited on channels already draining. * Transport Shutdown waits only on channels that are closing gracefully (akkadotnet#8589) Waiting on every association channel cost a full flush-wait whenever a channel was never disassociated - the throttler adapter never disassociates its TCP handle, so every throttled spec and multi-node node paid 2 s more. The race it guarded against can't happen through the protocol transport: ActorTransportAdapter.Shutdown stops the protocol manager first, and each ProtocolStateActor's OnTermination disassociates its handle, which now registers the channel synchronously before any async hop. Drops the spec case that called Shutdown before Disassociate directly on TcpTransport, an ordering the protocol transport can't produce. (cherry picked from commit b18540f)
Aaronontheweb
added a commit
to Aaronontheweb/akka.net
that referenced
this pull request
Oct 3, 2026
…t lose our last frames to an RST (akkadotnet#8635) * Add DotNettyGracefulCloseSpec: a graceful disassociate must not reset the peer (akkadotnet#8589) Drives TcpTransport against a raw socket peer whose bytes we never read. On dev the disassociate closes over that unread data, so the peer's SO_ERROR reports a reset (10058) and Write still returns true after Disassociate. * Classic remoting: half-close on disassociate instead of a full close that can RST (akkadotnet#8589) DotNetty's close is Socket.Shutdown(Both) + Dispose. With unread inbound data that sends an RST, and a Windows peer then drops frames it has not read yet - the lost ClusterSingleton HandOverDone in akkadotnet#8589. TcpAssociationHandle.Disassociate now refuses further writes, waits for the last write to reach the kernel, shuts down only the send side, and keeps reading (and discarding) until the peer's FIN, when DotNetty closes the channel over an empty buffer. A force-close on the channel's event loop bounds this by akka.remote.flush-wait-on-shutdown, and DotNettyTransport.Shutdown waits up to that long for pending graceful closes before its force-close. The send-side shutdown goes through a second Socket over the same handle: TcpSocketChannel.ShutdownOutputAsync calls Socket.Shutdown, which clears Socket.Connected, and DotNetty's DoReadBytes then returns EOF on the next read and closes over the unread data anyway. * Wait on every association channel in transport Shutdown; log half-close fallbacks (akkadotnet#8589) Remoting starts the transport Shutdown once the endpoint writers have stopped, but the protocol actor handles its DisassociateUnderlying asynchronously, so a graceful close may not have started yet when Shutdown looks. Tracking only the channels already draining missed those and force-closed them - the akkadotnet#8589 reset. Shutdown now waits up to flush-wait on the close of every non-server channel before its force-close. Also: log a warning when the half-close falls back to a full close, and once per transport if DotNetty's socket field is missing; fall back to CloseAsync if the event loop rejects the scheduled work. * DotNettyGracefulCloseSpec: cover shutdown-before-disassociate, TLS, two-system terminate, and the reflected DotNetty field The shutdown-before-disassociate case fails with SO_ERROR 10058 against the previous commit, which only waited on channels already draining. * Transport Shutdown waits only on channels that are closing gracefully (akkadotnet#8589) Waiting on every association channel cost a full flush-wait whenever a channel was never disassociated - the throttler adapter never disassociates its TCP handle, so every throttled spec and multi-node node paid 2 s more. The race it guarded against can't happen through the protocol transport: ActorTransportAdapter.Shutdown stops the protocol manager first, and each ProtocolStateActor's OnTermination disassociates its handle, which now registers the channel synchronously before any async hop. Drops the spec case that called Shutdown before Disassociate directly on TcpTransport, an ordering the protocol transport can't produce. (cherry picked from commit b18540f)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #8589
Problem
DotNetty 0.7.6 closes a channel with
Socket.Shutdown(Both)+Dispose(). If any inbound bytes are still unread at that point, the kernel sends an RST instead of a FIN. A Windows peer that gets an RST throws away everything in its receive buffer that it has not read yet. In #8589 that was sys2'sHandOverInProgress,HandOverDone, andExitingConfirmed. The actor layer wrote every frame into the kernel before the close.Two paths end in that close:
TcpAssociationHandle.Disassociate(per association) andDotNettyTransport.Shutdown(all channels).Fix
TcpAssociationHandle.Disassociatenow closes gracefully:Writereturnsfalseonce disassociated, because a write after the send side is shut down fails inDoWriteand hard-closes the socket.InboundPayloadto a stopped listener), turnAutoReadon, and shut down only the send side. The peer reads our frames, then a FIN.akka.remote.flush-wait-on-shutdown(default 2 s).DotNettyTransport.Shutdownwaits up toflush-wait-on-shutdownfor the channels that are closing gracefully, then runs its existing force-close loop. A channel registers as draining synchronously, as the first step ofDisassociate. Through the protocol transport that happens before the transport shutdown:ActorTransportAdapter.Shutdownstops the protocol manager first, and eachProtocolStateActordisassociates its handle inOnTermination. Channels that are never disassociated (the throttler adapter's, for example) don't make Shutdown wait.No wire format or public API change.
The half-close uses an alias socket
TcpSocketChannel.ShutdownOutputAsync()doesn't work here. It callsSocket.Shutdown(Send), which clearsSocket.Connected, and DotNetty'sDoReadBytestreats!Socket.Connectedas EOF. The channel then closes on the next read, over whatever is unread, and sends the RST anyway. (Checked: withShutdownOutputAsyncthe spec's first test still fails withSO_ERROR = 10058.)Instead, the send side is shut down through a second
Socketbuilt over the same handle (ownsHandle: false, so disposing it doesn't close anything). The originalSocketstaysConnected, so DotNetty keeps reading until the real FIN.AbstractSocketChannel.Socketfield. A test asserts the field exists, and the transport logs a warning if it's missing.Blocking = falsebeforeShutdown.Shutdownre-applies the alias's blocking mode to the shared handle, and on Windows aSocketbuilt from a handle assumes blocking.CloseAsync()and logs a warning.Where this differs from the plan in #8589
Disassociatetakes the graceful path, not justShutdown-reason ones. The handle does not know the reason. The cost for quarantine and failure paths is at mostflush-wait-on-shutdownbefore the socket closes.ChannelInactiveandExceptionCaughtstill notify the listener while draining. OnlyChannelReaddiscards.Disassociatedis alreadyIDeadLetterSuppression, andDotNettyTransportShutdownSpecrelies on the local listener gettingDisassociatedafter a local disassociate.Tests
New
DotNettyGracefulCloseSpec(Akka.Remote.Tests/Transport) drivesTcpTransportagainst a rawSocketpeer. Unless a test says otherwise, noReadHandlerSourceis set, so the peer's bytes sit unread on our side:SO_ERRORstays 0. After the peer closes, our channel closes within 1 s. Ondevthis fails withSO_ERROR = 10058(the reset that loses frames on Windows; Linux still delivers the bytes, so the error code is the signal).TlsHandler): the peer reads our frame, then EOF, with no reset.Shutdown()returns within 1.5 s (flush-wait 300 ms) while a drain is stuck.Terminate()finishes in under 3 s (about 130 ms locally), so a graceful close that stalls on flush-wait shows up.WritereturnsfalseafterDisassociate(fails ondev).AbstractSocketChannel.Socketfield exists, so a DotNetty upgrade can't turn the fix off unnoticed.Windows CI is the real check
This was developed and tested on Linux. The loss in #8589 happens on Windows, and two parts of this change are Windows-specific in ways I can't check locally:
Socketbuilt from a handle assumes blocking mode, whichBlocking = falseworks around. Disposing a non-owning alias doesn't close the handle. I read both behaviors from the .NET source; neither has run on Windows yet.Shutdown(Both)after the half-close. DotNetty's final close still callsSocket.Shutdown(Both)on the original socket. On Linux, .NET ignoresENOTCONNthere, with a comment that this "matches Winsock behavior", so I expect Winsock to return success too. If it did throw, DotNetty'sDoClose0still completes the channel's close future, and only the handle'sDisposewould move to the finalizer.The Windows unit-test lane on this PR is the real check.