Repository navigation
fix(testing): cancel draining after fatal silo errors - #11294
ReubenBond merged 2 commits into
Conversation
There was a problem hiding this comment.
Copilot review overview
🟡 Changes recommended
Cancellation aggregates may be logged at Error, and the regression test does not verify the required logging behavior.
Get a fresh assessment by requesting another Copilot review.
Review effort: Lite
Findings: 1
Open (1)
What changed in this PR
Updates test-host fatal-silo handling to cancel graceful draining promptly while preserving cleanup.
Changes:
- Uses a pre-canceled token for fatal-error shutdown.
- Adds regression coverage for cancellation and duplicate stop scheduling.
| File | Review summary |
|---|---|
test/TestInfrastructure/Orleans.TestingHost.Tests/TestClusterFatalErrorHandlerTests.cs |
Tests cancellation and single-stop behavior; logging assertions are still needed (nit, 1 vote). |
src/Orleans.TestingHost/TestClusterFatalErrorHandler.cs |
Needs aggregate cancellation handling to preserve Debug logging while retaining Error logging for real failures (moderate, 2 votes). |
💡 Add a code-review agent skill or configure MCP servers for context-aware, tailored reviews. Learn more in the docs.
Code coverage
Report-only conclusion: regressed. The current-main baseline is commit Coverage combines every CI test matrix job, including providers, CodeGen, .NET 8/10, Linux, Windows, and macOS, using canonical physical source and branch identities. The comparison remains report-only while normal line and branch variance is calibrated. Coverage details |

Fatal silo errors represent process loss. Graceful draining in the test host can wait on unreachable peers and delay cleanup.
Stop the affected host with a pre-canceled drain budget while preserving in-process cleanup. Log single and aggregated cancellation at Debug and actual stop failures at Error. The regression probe verifies canceled drain tokens, one host stop for repeated fatal notifications, and log classification for single, multiple, and nested stop failures.
Extracted from commit
48302ff632bbc7f99eadc0ce366b08c2ca4d9b88in #10236 as a standalone prerequisite for that PR. Addresses the fatal-drain policy implicated in #11286; end-to-end reproduction of that streaming shutdown remains a separate verification gap.